Prompt Caching
Prompt caching is a technique | that stores the results | of previous LLM prompts,
Prompt caching là một kỹ thuật | lưu trữ các kết quả | của các LLM prompts trước đó,
allowing you to quickly retrieve | and reuse them instead of | re-running the prompt every time.
cho phép bạn truy xuất nhanh chóng | và tái sử dụng chúng thay vì | chạy lại prompt mỗi lần.
This can significantly improve efficiency | and reduce costs when dealing | with frequently used or computationally expensive prompts.
Điều này có thể cải thiện đáng kể hiệu quả | và giảm chi phí khi xử lý | các prompts được sử dụng thường xuyên hoặc tốn kém về mặt tính toán.
Resources
- What is Prompt Caching? (article)
- What is Prompt Caching? Optimize LLM Latency with AI Transformers (video)
References
- https://roadmap.sh/ai-engineer (Node: Prompt Caching)
← Function Calling · AI Engineer Roadmap · Streaming Responses →