Five questions stand between a model page and a model that runs, and every one of them has an answer already sitting in the repo's metadata. These tools read it — in your browser, without downloading weights and without running anything.
Put a model in the box and each step below opens with it already filled in.
Leave it empty to browse, or try google/gemma-4-12B-it · aleada/Nemotron-3.5-…-W4A16
A model repo publishes its architecture, its attention topology, its tensor index and its chat template before you download a byte of weights. Almost every serving failure — it will not fit, the context is too long, the answers come back empty, the pack is not what its name says — is decidable from that alone. None of these tools loads a model, so none of them can tell you how good it is. They tell you whether it will work, and what it will cost.