EBook
Discover how to build smarter, more efficient AI inference systems.
![]()
Learn about quantization, sparsity, and advanced techniques like vLLM with Red Hat AI.
As AI models move into production, inference efficiency becomes essential for controlling infrastructure costs, reducing latency, and supporting growing demand. Organizations need to optimize both their models and serving environments to achieve scalable performance.
This eBook introduces the fundamentals of AI inference and explains how organizations can improve performance through runtime optimization and model compression. It covers vLLM, quantization, sparsity, memory management, and Red Hat AI tools designed to increase throughput while reducing compute and infrastructure requirements.
Key Takeaways:- How inference optimization reduces latency, memory use, and infrastructure costs
- Why vLLM improves throughput through continuous batching and PagedAttention
- What quantization and sparsity provide for model compression and efficiency
- How Red Hat AI supports optimized, scalable inference across hybrid environments
Topics
Artificial Intelligence (AI)
GenAI
DevOps Security
Emerging Technologies
