20231088/vllm - vllm - Luminance Code Repo

20231088/vllm

Author	SHA1	Message	Date
Roy	9e8744a545	[BugFix] Fix get tokenizer when using ray (#3301 )	2024-03-10 19:17:16 -07:00
Nick Hill	2efce05dc3	[Fix] Avoid pickling entire LLMEngine for Ray workers (#3207 ) Co-authored-by: Antoni Baum <antoni.baum@protonmail.com>	2024-03-06 00:17:20 +00:00
Nick Hill	8999ec3c16	Store `eos_token_id` in `Sequence` for easy access (#3166 )	2024-03-05 15:35:43 -08:00
Antoni Baum	ff578cae54	Add health check, make async Engine more robust (#3015 ) Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>	2024-03-04 22:01:40 +00:00
Antoni Baum	22de45235c	Push logprob generation to LLMEngine (#3065 ) Co-authored-by: Avnish Narayan <avnish@anyscale.com>	2024-03-04 19:54:06 +00:00
Philipp Moritz	17c3103c56	Make it easy to profile workers with nsight (#3162 ) Co-authored-by: Roger Wang <136131678+ywang96@users.noreply.github.com>	2024-03-03 16:19:13 -08:00
Jason Cox	d65fac2738	Add vLLM version info to logs and openai API server (#3161 )	2024-03-02 21:00:29 -08:00
Sage Moore	ce4f5a29fb	Add Automatic Prefix Caching (#2762 ) Co-authored-by: ElizaWszola <eliza@neuralmagic.com> Co-authored-by: Michael Goin <michael@neuralmagic.com>	2024-03-02 00:50:01 -08:00
Sherry	54d3544784	Fix: Output text is always truncated in some models (#3016 )	2024-03-01 07:52:22 +00:00
Nick Hill	29a8d6a554	[Fix] Don't deep-copy LogitsProcessors when copying SamplingParams (#3099 )	2024-02-29 19:20:42 +00:00
Allen.Dou	9289e577ec	add cache_config's info to prometheus metrics. (#3100 )	2024-02-29 06:15:18 +00:00
Liangfu Chen	3b7178cfa4	[Neuron] Support inference with transformers-neuronx (#2569 )	2024-02-28 09:34:34 -08:00
Roy	c1c0d00b88	Don't use cupy when `enforce_eager=True` (#3037 )	2024-02-26 17:33:38 -08:00
Harry Mellor	ef978fe411	Port metrics from `aioprometheus` to `prometheus_client` (#2730 )	2024-02-25 11:54:00 -08:00
Ronen Schaffer	4caf7044e0	Include tokens from prompt phase in `counter_generation_tokens` (#2802 )	2024-02-22 14:00:12 -08:00
Antoni Baum	017d9f1515	Add metrics to RequestOutput (#2876 )	2024-02-20 21:55:57 -08:00
Ronen Schaffer	e433c115bc	Fix `vllm:prompt_tokens_total` metric calculation (#2869 )	2024-02-18 23:55:41 -08:00
Nick Hill	185b2c29e2	Defensively copy `sampling_params` (#2881 ) If the SamplingParams object passed to LLMEngine.add_request() is mutated after it returns, it could affect the async sampling process for that request. Suggested by @Yard1 https://github.com/vllm-project/vllm/pull/2514#discussion_r1490106059	2024-02-17 11:18:04 -08:00
Woosuk Kwon	a463c333dd	Use CuPy for CUDA graphs (#2811 )	2024-02-13 11:32:06 -08:00
SangBin Cho	65b89d16ee	[Ray] Integration compiled DAG off by default (#2471 )	2024-02-08 09:57:25 -08:00
Rex	5a6c81b051	Remove eos tokens from output by default (#2611 )	2024-02-04 14:32:42 -08:00
Kunshang Ji	96b6f475dd	Remove hardcoded `device="cuda"` to support more devices (#2503 ) Co-authored-by: Jiang Li <jiang1.li@intel.com> Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>	2024-02-01 15:46:39 -08:00
Robert Shaw	93b38bea5d	Refactor Prometheus and Add Request Level Metrics (#2316 )	2024-01-31 14:58:07 -08:00
zhaoyang-star	b72af8f1ed	Fix error when tp > 1 (#2644 ) Co-authored-by: zhaoyang-star <zhao.yang16@zte.com.cn>	2024-01-28 22:47:39 -08:00
zhaoyang-star	9090bf02e7	Support FP8-E5M2 KV Cache (#2279 ) Co-authored-by: zhaoyang <zhao.yang16@zte.com.cn> Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>	2024-01-28 16:43:54 -08:00
Hanzhi Zhou	380170038e	Implement custom all reduce kernels (#2192 )	2024-01-27 12:46:35 -08:00
Antoni Baum	9b945daaf1	[Experimental] Add multi-LoRA support (#1804 ) Co-authored-by: Chen Shen <scv119@gmail.com> Co-authored-by: Shreyas Krishnaswamy <shrekris@anyscale.com> Co-authored-by: Avnish Narayan <avnish@anyscale.com>	2024-01-23 15:26:37 -08:00
Philipp Moritz	ab7e6006d6	Fix https://github.com/vllm-project/vllm/issues/2540 (#2545 )	2024-01-22 19:02:38 +01:00
Cade Daniel	18bfcdd05c	[Speculative decoding 2/9] Multi-step worker for draft model (#2424 )	2024-01-21 16:31:47 -08:00
shiyi.c_98	d10f8e1d43	[Experimental] Prefix Caching Support (#1669 ) Co-authored-by: DouHappy <2278958187@qq.com> Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>	2024-01-17 16:32:10 -08:00
Jiaxiang	6549aef245	[DOC] Add additional comments for LLMEngine and AsyncLLMEngine (#1011 )	2024-01-11 19:26:49 -08:00
Nadav Shmayovits	05921a9a7a	Changed scheduler to use deques instead of lists (#2290 ) Co-authored-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>	2024-01-07 09:48:07 -08:00
Iskren Ivov Chernev	d0215a58e7	Ensure metrics are logged regardless of requests (#2347 )	2024-01-05 05:24:42 -08:00
Zhuohan Li	fd4ea8ef5c	Use NCCL instead of ray for control-plane communication to remove serialization overhead (#2221 )	2024-01-03 11:30:22 -08:00
Zhuohan Li	e0ff920001	[BUGFIX] Do not return ignored sentences twice in async llm engine (#2258 )	2023-12-26 13:41:09 +08:00
Woosuk Kwon	3a4fd5ca59	Disable Ray usage stats collection (#2206 )	2023-12-20 21:52:08 -08:00
Woosuk Kwon	8041b7305e	[BugFix] Raise error when max_model_len is larger than KV cache (#2163 )	2023-12-17 17:08:23 -08:00
Woosuk Kwon	c3372e87be	Remove dependency on CuPy (#2152 )	2023-12-17 01:49:07 -08:00
Woosuk Kwon	37ca558103	Optimize model execution with CUDA graph (#1926 ) Co-authored-by: Chen Shen <scv119@gmail.com> Co-authored-by: Antoni Baum <antoni.baum@protonmail.com>	2023-12-16 21:12:08 -08:00
Yunfeng Bai	c06170cc8e	Add a flag to include stop string in output text (#1976 )	2023-12-15 00:45:58 -08:00
Woosuk Kwon	464dd985e3	Fix num_gpus when TP > 1 (#1852 )	2023-12-03 12:24:30 -08:00
Simon Mo	5313c2cb8b	Add Production Metrics in Prometheus format (#1890 )	2023-12-02 16:37:44 -08:00
Woosuk Kwon	27feead2f8	Refactor Worker & InputMetadata (#1843 )	2023-11-29 22:16:37 -08:00
FlorianJoncour	0229c386c5	Better integration with Ray Serve (#1821 ) Co-authored-by: FlorianJoncour <florian@zetta-sys.com>	2023-11-29 13:25:43 -08:00
Zhuohan Li	708e6c18b0	[FIX] Fix class naming (#1803 )	2023-11-28 14:08:01 -08:00
boydfd	4bb6b67188	fix RAM OOM when load large models in tensor parallel mode. (#1395 ) Co-authored-by: ran_lin <rlin@thoughtworks.com>	2023-11-20 19:02:42 -08:00
Simon Mo	5ffc0d13a2	Migrate linter from `pylint` to `ruff` (#1665 )	2023-11-20 11:58:01 -08:00
Simon Mo	cb08cd0d75	[Minor] Fix duplication of ignored seq group in engine step (#1666 )	2023-11-16 13:11:41 -08:00
Dan Lord	7013a80170	Add support for `spaces_between_special_tokens`	2023-10-30 16:52:56 -07:00
Zhuohan Li	9d9072a069	Implement prompt logprobs & Batched topk for computing logprobs (#1328 ) Co-authored-by: Yunmo Chen <16273544+wanmok@users.noreply.github.com>	2023-10-16 10:56:50 -07:00

1 2

82 Commits