<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>AWQ :: 标签 :: x7peeps</title><link>https://x7peeps.com/tags/AWQ/index.html</link><description/><generator>Hugo</generator><language>zh-CN</language><lastBuildDate>Mon, 17 Aug 2026 08:35:59 +0000</lastBuildDate><atom:link href="https://x7peeps.com/tags/AWQ/index.xml" rel="self" type="application/rss+xml"/><item><title>模型量化与压缩技术：GPTQ、AWQ、GGUF与FP8的工程实践指南</title><link>https://x7peeps.com/AI/10-%E6%A8%A1%E5%9E%8B%E8%AE%AD%E7%BB%83/%E6%A8%A1%E5%9E%8B%E9%87%8F%E5%8C%96%E4%B8%8E%E5%8E%8B%E7%BC%A9%E6%8A%80%E6%9C%AFGPTQAWQGGUF%E4%B8%8EFP8%E7%9A%84%E5%B7%A5%E7%A8%8B%E5%AE%9E%E8%B7%B5%E6%8C%87%E5%8D%97/index.html</link><pubDate>Mon, 17 Aug 2026 08:35:59 +0000</pubDate><guid>https://x7peeps.com/AI/10-%E6%A8%A1%E5%9E%8B%E8%AE%AD%E7%BB%83/%E6%A8%A1%E5%9E%8B%E9%87%8F%E5%8C%96%E4%B8%8E%E5%8E%8B%E7%BC%A9%E6%8A%80%E6%9C%AFGPTQAWQGGUF%E4%B8%8EFP8%E7%9A%84%E5%B7%A5%E7%A8%8B%E5%AE%9E%E8%B7%B5%E6%8C%87%E5%8D%97/index.html</guid><description>模型量化与压缩技术：GPTQ、AWQ、GGUF与FP8的工程实践指南 大语言模型的参数规模在过去五年呈指数级增长——从 GPT-2 的 15 亿参数到 LLaMA-3.1-405B 的 4050 亿参数，再到 DeepSeek-V3 的 6710 亿参数。然而 GPU 的显存容量增长远不及模型规模的扩张速度。以 FP16 精度存储一个 70B 模型需要约 140GB 显存，而一块 H100 SXM5 的显存仅为 80GB。量化（Quantization） 正是打破这一瓶颈的关键技术。</description></item></channel></rss>