20231088/vllm

History

Brian Dellabetta 44bbca78d7

Based on a request by @mgoin , with @kylesayrs we have added an example
doc for int4 w4a16 quantization, following the pre-existing int8 w8a8
quantization example and the example available in
[`llm-compressor`](https://github.com/vllm-project/llm-compressor/blob/main/examples/quantization_w4a16/llama3_example.py)

FIX #n/a (no issue created)

@kylesayrs and I have discussed a couple additional improvements for the
quantization docs. We will revisit at a later date, possibly including:
- A section for "choosing the correct quantization scheme/ compression
technique"
- Additional vision or audio calibration datasets

---------

Signed-off-by: Brian Dellabetta <bdellabe@redhat.com>
Co-authored-by: Michael Goin <michael@neuralmagic.com>

2025-01-31 15:38:48 -08:00

source

[Doc] int4 w4a16 example (#12585 )

2025-01-31 15:38:48 -08:00

make.bat

Add initial sphinx docs (#120 )

2023-05-22 17:02:44 -07:00

Makefile

[Doc] Group examples into categories (#11782 )

2025-01-08 09:20:12 +08:00

README.md

[CI/Build] Add markdown linter (#11857 )

2025-01-12 00:17:13 -08:00

requirements-docs.txt

[Doc] Convert docs to use colon fences (#12471 )

2025-01-29 11:38:29 +08:00

README.md

vLLM documents

Build the docs

# Install dependencies.
pip install -r requirements-docs.txt

# Build the docs.
make clean
make html

Open the docs with your browser

python -m http.server -d build/html/

Launch your browser and open localhost:8000.