Corollary

Research

  • Papers

Library

  • Catalog

Account

  • Jobs
  • Billing
SettingsResearch demo · not for clinical or commercial use
Corollary
  1. Catalog
  2. Multi-Modal Foundation
  3. Zero-To-CAD Qwen3-VL 2B

Multi-Modal Foundation · ADSKAILab

Zero-To-CAD Qwen3-VL 2B

Qwen3-VL fine-tuned to generate parametric CAD models directly from images — bridges vision-language reasoning and engineering geometry synthesis.

Verify the licence before commercial useThe See model card licence carries terms worth reading before you depend on it.See model card
Availability
On demand
Typical latency
12s
Credits per run
~70
Compute tier
A · Fast

1–30s

Runs on On-demand GPU

Runs on our servers with on-demand compute. A first run needs time to load the model; active capacity can be reused and scales down when idle.

Run this model

Zero-To-CAD Qwen3-VL 2B

PNG, JPEG or TIFF. Sent to the model as-is.

Optional. Without one the model describes the image unprompted.

Needs image

About this model

Qwen3-VL fine-tuned to generate parametric CAD models directly from images — bridges vision-language reasoning and engineering geometry synthesis. Served through the generic Modal runner, which pulls ADSKAILab/Zero-To-CAD-Qwen3-VL-2B from the Hugging Face Hub and runs it as a image-text-to-text pipeline. Inputs and outputs were derived from the repository's declared pipeline tag rather than written by hand — check the model card before relying on a result.

Standardized I/O contract

Every model in the Hub speaks the same contract, which is what lets the Router and the agent call any of them without special-casing.

Inputs

  • imagefile · required

    Image

  • prompttextarea

    Question

Outputs

  • texttext

    The model's completion.

Specification

Hardware
1× A100 16GB
GPU memory
24 GB
Version
hub
Licence
See model card
Backend
On demandOn-demand GPU
MCP server
mcp-sandbox-server

Tasks

image-text-to-text

Source

  • Model card
  • Repository
engineeringscientific-reasoningimage-text-to-text