vLLM optimization pack
Last updated
Was this helpful?
The vLLM optimization pack provides support for optimizing the serving configuration of Large Language Models served through vLLM, balancing latency, throughput and GPU memory efficiency.
The following component types are supported.
A vLLM inference server
Here’s the command to install the vLLM optimization pack using the Akamas CLI:
akamas install optimization-pack vLLMLast updated
Was this helpful?
Was this helpful?