Multi-Modal Foundation · ADSKAILab
Qwen3-VL fine-tuned to generate parametric CAD models directly from images — bridges vision-language reasoning and engineering geometry synthesis.
1–30s
Runs on On-demand GPU
Runs on our servers with on-demand compute. A first run needs time to load the model; active capacity can be reused and scales down when idle.Run this model
About this model
Qwen3-VL fine-tuned to generate parametric CAD models directly from images — bridges vision-language reasoning and engineering geometry synthesis. Served through the generic Modal runner, which pulls ADSKAILab/Zero-To-CAD-Qwen3-VL-2B from the Hugging Face Hub and runs it as a image-text-to-text pipeline. Inputs and outputs were derived from the repository's declared pipeline tag rather than written by hand — check the model card before relying on a result.
Standardized I/O contract
Every model in the Hub speaks the same contract, which is what lets the Router and the agent call any of them without special-casing.
Inputs
Image
Question
Outputs
The model's completion.