{"product_id":"hands-on-llm-serving-and-optimization","title":"Hands-On LLM Serving and Optimization","description":"Large language models (LLMs) are rapidly becoming the backbone of AI-driven applications. Without proper optimization, however, LLMs can be expensive to run, slow to serve, and prone to performance bottlenecks. As the demand for real-time AI applications grows, along comes Hands-On Serving and Optimizing LLM Models, a comprehensive guide to the complexities of deploying and optimizing LLMs at scale. In this hands-on book, authors Chi Wang and Peiheng Hu take a real-world approach backed by practical examples and code, and assemble essential strategies for designing robust infrastructures that are equal to the demands of modern AI applications. Whether you're building high-performance AI systems or looking to enhance your knowledge of LLM optimization, this indispensable book will serve as a pillar of your success. Learn the key principles for designing a model-serving system tailored to popular business scenariosUnderstand the common challenges of hosting LLMs at scale while minimizing costsPick up practical techniques for optimizing LLM serving performanceBuild a model-serving system that meets specific business requirementsImprove LLM serving throughput and reduce latencyHost LLMs in a cost-effective manner, balancing performance and resource efficiency","brand":"O'Reilly Media","offers":[{"title":"Default Title","offer_id":58278379585871,"sku":"9798341621497","price":81.95,"currency_code":"EUR","in_stock":true}],"thumbnail_url":"\/\/cdn.shopify.com\/s\/files\/1\/0061\/0372\/8217\/files\/9798341621497_1-hands-on-llm-serving-and-optimization.jpg?v=1784700489","url":"https:\/\/www.suomalainen.com\/products\/hands-on-llm-serving-and-optimization","provider":"Suomalainen.com","version":"1.0","type":"link"}