Powered by ZeroGPU (H200) | 4-bit AWQ quantized | 25.5 GB
Note: First inference takes ~2 min to load the model. Subsequent ones are faster.