<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>KV缓存 :: 标签 :: x7peeps</title><link>https://x7peeps.com/tags/KV%E7%BC%93%E5%AD%98/index.html</link><description/><generator>Hugo</generator><language>zh-CN</language><lastBuildDate>Wed, 26 Aug 2026 14:28:00 +0800</lastBuildDate><atom:link href="https://x7peeps.com/tags/KV%E7%BC%93%E5%AD%98/index.xml" rel="self" type="application/rss+xml"/><item><title>MetalAnchor 实战：520 行 Python 把智能体语义 KV 缓存落地 Apple Silicon</title><link>https://x7peeps.com/AI/06-AI%E5%B7%A5%E7%A8%8B%E5%8C%96/MetalAnchor%E5%AE%9E%E6%88%98520%E8%A1%8CPython%E6%8A%8A%E6%99%BA%E8%83%BD%E4%BD%93%E8%AF%AD%E4%B9%89KV%E7%BC%93%E5%AD%98%E8%90%BD%E5%9C%B0AppleSilicon/index.html</link><pubDate>Wed, 26 Aug 2026 14:28:00 +0800</pubDate><guid>https://x7peeps.com/AI/06-AI%E5%B7%A5%E7%A8%8B%E5%8C%96/MetalAnchor%E5%AE%9E%E6%88%98520%E8%A1%8CPython%E6%8A%8A%E6%99%BA%E8%83%BD%E4%BD%93%E8%AF%AD%E4%B9%89KV%E7%BC%93%E5%AD%98%E8%90%BD%E5%9C%B0AppleSilicon/index.html</guid><description>MetalAnchor 实战：520 行 Python 把智能体语义 KV 缓存落地 Apple Silicon 导读：如果你在 Mac 上跑本地大模型做智能体（Agent），大概率遇到过这种"慢"——新会话要等模型加载很久、同一个系统提示词每次都要重新计算一遍、长对话越聊越慢。这不是模型能力问题，而是"重复计算"问题：模型把已经算过的内容又算了一遍。本文记录一个真实解法：用 520 行纯 Python 写一个透明代理层（MetalAnchor），把 FreeToken 的智能体语义 KV 缓存思想落地到 Apple Silicon，实测把缓存复用率做到 90%，首 Token 延迟（TTFT）提速 5-7 倍。</description></item></channel></rss>