Denali-AI Garment VLM Benchmark

Submit your Vision-Language Model for evaluation on our 3,500-sample hard garment classification benchmark.

Your model will be automatically downloaded, served via vLLM on our GPU cluster (2x RTX PRO 6000 Blackwell, 196 GB VRAM), and evaluated across 6 metrics on 9 garment attribute fields.

Supported models: Any VLM on HuggingFace compatible with vLLM — public, private, or gated.

Benchmark Results

Ranked by Weighted Field Score (SBERT+NLI combined, field weights: type=2.5x, defect=2.0x, brand/size=1.5x, others=1.0x)

Benchmark

3.5k hard = curated hard eval set; 18k peakbench = full peakbench eval