123 lines
7.1 KiB
Markdown
123 lines
7.1 KiB
Markdown
---
|
|
name: model-support-checker
|
|
description: Use when checking HuggingFace or ModelScope model support for SGLang, vLLM, or vLLM-Ascend, finding the first supporting version, or diagnosing "model not supported" errors.
|
|
---
|
|
|
|
# Model Support Checker
|
|
|
|
Check whether a HuggingFace / ModelScope model is supported by **SGLang**, **vLLM**, or **vLLM-Ascend**, and since which version.
|
|
|
|
## Methodology
|
|
|
|
The authoritative signal: does the framework's **source code** contain an implementation for the model's architecture string? Docs are supplementary.
|
|
|
|
1. **Architecture** — `architectures` from `config.json` (HF → ModelScope fallback). Override with `--arch`.
|
|
2. **Docs** (supplementary) — grep the framework's supported-models page.
|
|
3. **Source** (authoritative) — search the models directory for the architecture string.
|
|
4. **Reconcile** — if docs say YES but source says NO, manually inspect the framework's repo to resolve the discrepancy (see Step 4b).
|
|
5. **Version** — earliest commit on the implementation file → nearest release tag.
|
|
|
|
The tool is stateful; mode persists in `.state/state.json`. Local clone (recommended) greps on disk. Token mode uses GitHub API.
|
|
|
|
## Skill Workflow
|
|
|
|
**Step 1 — Doctor.** `python3 main.py --doctor`. `[ERROR]` → fix first, stop. `[WARN]` → ignore, continue.
|
|
|
|
**Step 2 — Setup (optional).** Only if no state exists. Recommend local: `python3 main.py --setup local`.
|
|
|
|
**Step 3 — Resolve model ID.** Infer full `<org>/<model_name>` from user input, matching the hosting platform convention. Parse URLs or short names. Query the web if uncertain.
|
|
|
|
**Step 4 — Run.**
|
|
```bash
|
|
python3 main.py [--framework <fw>] [--source <src>] [--arch <arch>] <model_id>
|
|
```
|
|
- All frameworks: `python3 main.py <model_id>`
|
|
- One framework: `python3 main.py --framework vllm <model_id>`
|
|
- Manual arch (no config.json needed): `python3 main.py --arch LlamaForCausalLM`
|
|
- Arch + model: `python3 main.py --framework vllm --arch DeepseekV3ForCausalLM <model_id>`
|
|
|
|
**Step 4b — Reconcile discrepancies.** If docs=YES but source=NO for any framework, the directory layout has likely changed. Do not assume known paths; instead:
|
|
1. Read the framework's registry file to find where the architecture is registered.
|
|
2. Extract the module path from the registry entry.
|
|
3. Trace that module path to locate the actual implementation file on disk.
|
|
4. Verify the architecture string exists in that file.
|
|
5. If the registry does not contain the architecture, search the entire framework repo for the architecture string (e.g. `rg <ArchName> <repo_root>`) to locate where it is referenced.
|
|
6. Update the result to YES with the correct file path.
|
|
|
|
**Step 5 — Report.** Parse Summary. Report: YES/NO per framework, minimum version, file path, vLLM category/module/class.
|
|
|
|
## Fallback Strategies
|
|
|
|
### Fallback 1: config.json fetch fails
|
|
|
|
Escalate in order:
|
|
|
|
1. **Switch platform** — retry with `--source hf` (default is modelscope).
|
|
2. **webfetch** — agent fetches `config.json` directly:
|
|
- HF: `https://huggingface.co/<org>/<model>/resolve/main/config.json`
|
|
- MS: `https://modelscope.cn/models/<org>/<model>/resolve/master/config.json`
|
|
Extract `architectures`, then: `python3 main.py --arch <ArchName> <model_id>`
|
|
3. **Guide user** — if webfetch fails (private/gated), provide links for the user to read `architectures` from `config.json`:
|
|
- HF: `https://huggingface.co/<org>/<model>/blob/main/config.json`
|
|
- MS: `https://modelscope.cn/models/<org>/<model>/files`
|
|
Then run with `--arch`.
|
|
|
|
### Fallback 2: Script cannot run
|
|
|
|
Bypass the script; follow methodology manually:
|
|
|
|
1. Get architecture (webfetch or ask user).
|
|
2. Search framework GitHub repos for the architecture string:
|
|
- vLLM (`vllm-project/vllm`): `vllm/model_executor/models/` or `vllm/models/<family>/`
|
|
- SGLang (`sgl-project/sglang`): `python/sglang/srt/models/`
|
|
- vLLM-Ascend (`vllm-project/vllm-ascend`): `vllm_ascend/models/` (may include subdirs like `minimax_m3/`)
|
|
3. Check official supported-models docs.
|
|
4. Earliest commit → nearest release tag for version.
|
|
5. Report in script Summary format.
|
|
|
|
### Fallback 3: Source=NO but model is recent / may ship support out-of-tree
|
|
|
|
The script's source check is authoritative for the **merged `main` branch checkout**, but framework teams frequently ship support through channels that are NOT in the default `main` checkout the script greps:
|
|
|
|
- **Official Docker images** with the implementation baked in (e.g. `lmsysorg/sglang:glm-5.3-flash`, `vllm/vllm-openai-rocm:glm53-flash`).
|
|
- **A separate support branch** not yet merged to main (e.g. `xinyuan/glm-5.3-flash-support`).
|
|
- **Vendor recipes / cookbook pages** (e.g. `recipes.vllm.ai/<org>/<model>`, `cookbook.sglang.io/...`).
|
|
|
|
This is the reverse of Step 4b (docs=YES, source=NO). When source=NO, do NOT report a flat "unsupported" until the official deployment docs are checked:
|
|
|
|
1. Fetch the model card on **both** HF and ModelScope (one is often gated/401):
|
|
- HF: `https://huggingface.co/<org>/<model>` and `.../resolve/main/README.md`
|
|
- MS: `https://modelscope.cn/models/<org>/<model>` and `.../resolve/master/README.md`
|
|
2. Find a "Deploy / Serve locally" section. If it lists vLLM / SGLang / vLLM-Ascend, the model IS deployable — capture the exact mechanism:
|
|
- Docker image tag(s) and required version (e.g. "vLLM 0.29.0+", "use docker before integration is in the public repo").
|
|
- Extra requirements: FlashInfer version, GPU arch (Hopper+/Blackwell/gfx950), `--tool-call-parser` / `--reasoning-parser`, MTP/speculative config.
|
|
- Caveats: features only on a specific branch/PR, image alone insufficient, etc.
|
|
3. Follow the recipe links to get the concrete launch command and image name.
|
|
4. Report a **nuanced** result instead of a bare NO. Distinguish these states:
|
|
- `merged` — in the public repo/PyPI, version-known.
|
|
- `image-only` — runs via vendor Docker image, not yet in public main.
|
|
- `branch-only` — needs a specific git branch/PR checkout.
|
|
- `unsupported` — no official path at all.
|
|
|
|
The official model card + recipes are authoritative about HOW to run, even when the implementation is not in the default `main` checkout. This supersedes a flat NO from the script.
|
|
|
|
## Quick Reference
|
|
|
|
| Flag | Description |
|
|
|------|-------------|
|
|
| `--framework` | `sglang` / `vllm` / `vllm-ascend` / `all` |
|
|
| `--source` | `auto` / `hf` / `modelscope` |
|
|
| `--arch` | Manual arch name(s); skips config.json |
|
|
| `--doctor` | Check setup state |
|
|
| `--setup` | `local` or `token` |
|
|
|
|
`python3 main.py --help` for full flag list.
|
|
|
|
## Common Mistakes
|
|
|
|
- **Skipping doctor** — always `--doctor` first.
|
|
- **Giving up on config.json failure** — escalate: switch platform → webfetch → guide user.
|
|
- **Not using `--arch`** — when config.json is unreachable, pass `--arch` instead of aborting.
|
|
- **Trusting source=NO when docs=YES** — frameworks may reorganize directories; do Step 4b reconciliation.
|
|
- **Trusting source=NO as final for recent models** — vendors often ship support via official Docker images / separate branches / recipe pages that are NOT in the default `main` checkout the script greps. Check the official model card + recipes (Fallback 3) before concluding "unsupported".
|