我的征尘是星辰大海。。。
The dirt and dust from my pilgrimage forms oceans of stars...
-------当记忆的篇章变得零碎,当追忆的图片变得模糊,我们只能求助于数字存储的永恒的回忆
作者:黄教授
手机视频列表
谷歌纽约团队涉嫌学术不端
视频
音频
原始脚本
旧瓶新酒加学术虚火,谷歌 TensorQuant 爆火背后,没那么简单。 前段时间,谷歌 TensorQuant 的消息刷遍科技圈,号称靠着三比特 kv 缓存量化的黑科技,直接让存储股一天蒸发6200亿。 华尔街一边热炒,一边又痛批市场不懂技术,看似是一场颠覆性的 AI AI 技术革命,可深究下来,这背后根本不是什么原创突破,反倒藏着学术圈少为人知的套路,以及被过度夸大的技术真相。 先简单说清 TurboQuant 到底是什么,它是谷歌纽约研究团队推出的,主打大模型推理阶段,KV 缓存三比特无损量化,宣称能把缓存内存压缩到原来的1/6,还能大幅提升推理速度。 对外包装成解决 AI 显存瓶颈、颠覆存储行业的革命性技术。 可但凡深入了解行业技术进展和这个团队背景,就会发现这不过是旧瓶装新酒,再加上刻意的宣传造势,才闹出了这么大的动静。 TeoQuant 所谓的核心技术根本算不上谷歌原创。 早在两年前,苏黎世联邦理工学院加新加坡南洋理工大学的高 建阳团队就已经推出了 Rabbit Q 算法,核心思路同样是随机旋转结合 GL 变换,做低比特量化。 不仅理论上证明了渐进最优,还同步开源了完整的 C 加加高性能代码,发表在数据库顶会 SIGMOD 上,是实打实经过学术验证和开源检验的成果。 而谷歌 TensorQuant 的核心技术路线和 Rabbit 高度重合却刻意回避引用前者。 甚至在对比实验中耍尽手段,把 RobotQ 的未优化 Python 版本,单核 CPU 运行,和自己全优化 GPU 版本的 TurboQuant 做对比,造出性能碾压的假象。 本质就是拿别人的成熟思路,做了点工程包装,就标榜成自己的原创黑科技。 再说说这个打造 TurboQuant 的谷歌纽约团队,更是科技圈里典型的论文工厂风格。 团队核心负责人是谷歌 Fellow VP Machine Learning 成员多为伊朗裔、韩裔研究人员。 和谷歌山景城主打产品落地硬核工程的团队不同。 纽约这支团队完全重理论、重顶会发表、重宣传包装、轻工程落地、轻开源复现。 之前他们推出的 Titans 系统,号称能实现200万 token 超长上下文,听起来是史诗级突破。 可直到现在,既没有开源代码,也没有第三方能独立复现,更没有任何谷歌内部产品落地的实锤。 纯纯是实验室里的概念产物,靠着论文和宣传博眼球。 这次的 TensorFlow Quantum 完全是一模一样的套路。 很多人会疑惑,既然技术不是原创,那三比特量化本身总有价值吧?客观来说,三比特的 KV 缓存压缩确实是不错的工程优化,但它的实际影响力被无限夸大了。 首先要分清一个关键问题,TableQuant 优化的是 KV 缓存,和微软做到的1.5 比特全重量化完全是两码事。 KV 缓存针对的是模型推理时的上下文中间数据,只在超长上下文场景下才会占据大量显存。 普通的日常对话、短文本交互,KV 缓存占显存的比例本就不高,优化效果微乎其微。 而 而全重量化是直接压缩模型本身,是从根源上解决显存需求,覆盖所有场景的核心优化。 两者的技术价值和实际冲击力根本不在一个层面。 而且,低比特 kv 量化本就不是新技术,2024年行业里就有 RabbitQ、KV Quant 等多款方案做到了4比特及以下的无损压缩。 Tab Quant 的3比特只是小幅迭代,并非理论突破。 如果这项技术真有宣传中那么颠覆,去年 RabbitMQ 推出时就该引发行业震动,根本轮不到谷歌今年拿来包装后才爆火。 说到底,这场风波的核心从来不是技术本身的优劣,而是谷歌团队刻意的学术不端和市场炒作,这也是 TurboQuant 爆火却没用的核心原因。 它的火爆全靠谷歌的公关造势、资本市场的情绪炒作。 可落到实际应用层面,至今没有开源代码、没有第三方公平复现、没有真正的生产及落地。 所谓的性能优势不过是靠着 GPU 硬件优化和不公平的实验对比。 把工程优势包装成了算法优势,并非技术本身有质的飞跃。 再加上这支团队一贯的重论文、轻落地风格。 TabQuant 大概率和 Titans 一样,最终只是停留在论文里的概念,很难真正走进 AI 实际应用场景。 其实科技圈向来不缺这类伪突破,真正的技术创新从来都是开源可复现、公平可对比、落地可实用的。 而不是靠着抹除前人成果、实验作弊、过度宣传造出来的神话。 TurboQuant 这场闹剧也撕开了学术圈部分论文工厂的遮羞布。 看似光鲜的顶会成果、震撼的技术数据,背后或许只是经不起推敲的虚虚火 大家看待这类所谓的黑科技还是要多一分理性,少被资本和宣传带偏。
修正脚本
旧瓶新酒加学术虚火,谷歌 TensorQuant 爆火背后,没那么简单。 前段时间,谷歌 TensorQuant 的消息刷遍科技圈,号称靠着三比特 kv 缓存量化的黑科技,直接让存储股一天蒸发6200亿。 华尔街一边热炒,一边又痛批市场不懂技术,看似是一场颠覆性的 AI 技术革命,可深究下来,这背后根本不是什么原创突破,反倒藏着学术圈少为人知的套路,以及被过度夸大的技术真相。 先简单说清 TensorQuant 到底是什么,它是谷歌纽约研究团队推出的,主打大模型推理阶段,KV 缓存三比特无损量化,宣称能把缓存内存压缩到原来的1/6,还能大幅提升推理速度。 对外包装成解决 AI 显存瓶颈、颠覆存储行业的革命性技术。 可但凡深入了解行业技术进展和这个团队背景,就会发现这不过是旧瓶装新酒,再加上刻意的宣传造势,才闹出了这么大的动静。 TensorQuant 所谓的核心技术根本算不上谷歌原创。 早在两年前,苏黎世联邦理工学院加新加坡南洋理工大学的高建阳团队就已经推出了 Rabbit Q 算法,核心思路同样是随机旋转结合 GL 变换,做低比特量化。 不仅理论上证明了渐进最优,还同步开源了完整的 C++ 高性能代码,发表在数据库顶会 SIGMOD 上,是实打实经过学术验证和开源检验的成果。 而谷歌 TensorQuant 的核心技术路线和 Rabbit 高度重合却刻意回避引用前者。 甚至在对比实验中耍尽手段,把 RabbitQ 的未优化 Python 版本,单核 CPU 运行,和自己全优化 GPU 版本的 TensorQuant 做对比,造出性能碾压的假象。 本质就是拿别人的成熟思路,做了点工程包装,就标榜成自己的原创黑科技。 再说说这个打造 TensorQuant 的谷歌纽约团队,更是科技圈里典型的论文工厂风格。 团队核心负责人是谷歌 Fellow VP,Machine Learning 成员多为伊朗裔、韩裔研究人员。 和谷歌山景城主打产品落地的硬核工程团队不同。 纽约这支团队完全重理论、重顶会发表、重宣传包装、轻工程落地、轻开源复现。 之前他们推出的 Titans 系统,号称能实现200万 token 超长上下文,听起来是史诗级突破。 可直到现在,既没有开源代码,也没有第三方能独立复现,更没有任何谷歌内部产品落地的实锤。 纯纯是实验室里的概念产物,靠着论文和宣传博眼球。 这次的 TensorQuant 完全是一模一样的套路。 很多人会疑惑,既然技术不是原创,那三比特量化本身总有价值吧?客观来说,三比特的 KV 缓存压缩确实是不错的工程优化,但它的实际影响力被无限夸大了。 首先要分清一个关键问题,TensorQuant 优化的是 KV 缓存,和微软做到的1.5 比特全重量化完全是两码事。 KV 缓存针对的是模型推理时的上下文中间数据,只在超长上下文场景下才会占据大量显存。 普通的日常对话、短文本交互,KV 缓存占显存的比例本就不高,优化效果微乎其微。 而全重量化是直接压缩模型本身,是从根源上解决显存需求,覆盖所有场景的核心优化。 两者的技术价值和实际冲击力根本不在一个层面。 而且,低比特 kv 量化本就不是新技术,2024年行业里就有 RabbitQ、KV Quant 等多款方案做到了4比特及以下的无损压缩。 TensorQuant 的3比特只是小幅迭代,并非理论突破。 如果这项技术真有宣传中那么颠覆,去年 RabbitQ 推出时就该引发行业震动,根本轮不到谷歌今年拿来包装后才爆火。 说到底,这场风波的核心从来不是技术本身的优劣,而是谷歌团队刻意的学术不端和市场炒作,这也是 TensorQuant 爆火却没用的核心原因。 它的火爆全靠谷歌的公关造势、资本市场的情绪炒作。 可落到实际应用层面,至今没有开源代码、没有第三方公平复现、没有真正的生产级落地。 所谓的性能优势不过是靠着 GPU 硬件优化和不公平的实验对比。 把工程优势包装成了算法优势,并非技术本身有质的飞跃。 再加上这支团队一贯的重论文、轻落地风格。 TensorQuant 大概率和 Titans 一样,最终只是停留在论文里的概念,很难真正走进 AI 实际应用场景。 其实科技圈向来不缺这类伪突破,真正的技术创新从来都是开源可复现、公平可对比、落地可实用的。 而不是靠着抹除前人成果、实验作弊、过度宣传造出来的神话。 TensorQuant 这场闹剧也撕开了学术圈部分论文工厂的遮羞布。 看似光鲜的顶会成果、震撼的技术数据,背后或许只是经不起推敲的虚火 大家看待这类所谓的黑科技还是要多一分理性,少被资本和宣传带偏。
英文翻译
New wine in old bottles with academic hype: The story behind Google TensorQuant's explosion is not that simple. A while ago, news about Google's TensorQuant swept through the tech circle, claiming that with the black technology of three-bit KV cache quantization, it directly caused storage stocks to evaporate 620 billion in a single day. Wall Street was buzzing with hype on one hand, while on the other hand criticizing the market for not understanding the technology. It seemed like a disruptive AI tech revolution, but upon closer examination, there is no original breakthrough behind it at all. Instead, it conceals little-known tricks in the academic world and an overly exaggerated technological truth. Let me first briefly explain what TensorQuant is. It was launched by Google's New York research team, focusing on the inference phase of large models with three-bit lossless quantization of the KV cache. It claims to compress cache memory to one-sixth of its original size and significantly improve inference speed. Externally, it was packaged as a revolutionary technology to solve the AI memory bottleneck and disrupt the storage industry. But anyone who digs into the industry's technological progress and the background of this team will find that this is nothing more than new wine in old bottles, combined with deliberate promotional hype, which caused such a big stir. The so-called core technology of TensorQuant is hardly Google's original work. As early as two years ago, the team led by Gao Jianyang at ETH Zurich and Nanyang Technological University in Singapore had already introduced the RabbitQ algorithm. Its core idea was also random rotation combined with GL transformation for low-bit quantization. Not only did it theoretically prove asymptotic optimality, but it also open-sourced complete high-performance C++ code, which was published at the top database conference SIGMOD. It is a result that has been rigorously verified by academia and open-source testing. The core technical path of Google's TensorQuant heavily overlaps with RabbitQ, yet it deliberately avoids citing the former. Even worse, in comparative experiments, they played tricks by comparing RabbitQ's unoptimized Python version running on a single CPU core with TensorQuant's fully optimized GPU version, creating the illusion of crushing performance. Essentially, they took someone else's mature idea, did some engineering packaging, and touted it as their own original black technology. Now, let's talk about the Google New York team behind TensorQuant—they are a typical example of a "paper factory" in the tech circle. The core leader of the team is a Google Fellow VP, and most members of the Machine Learning group are Iranian or Korean researchers. Unlike Google's Mountain View team, which focuses on hardcore engineering and product delivery, the New York team is all about theory, top conference publications, promotional packaging, and neglects engineering implementation and open-source reproducibility. Previously, they launched the Titans system, claiming to achieve a 2-million-token ultra-long context, which sounded like an epic breakthrough. But to this day, there is no open-source code, no third party has been able to independently reproduce it, and there is no solid evidence of any internal Google product deployment. It is purely a conceptual product from the lab, gaining attention through papers and publicity. This time, TensorQuant follows exactly the same pattern. Many people may wonder: If the technology is not original, is the three-bit quantization itself at least valuable? Objectively speaking, three-bit KV cache compression is indeed a decent engineering optimization, but its actual impact has been infinitely exaggerated. First, we need to clarify a key issue: TensorQuant optimizes the KV cache, which is completely different from Microsoft's 1.5-bit full-model quantization. The KV cache targets the intermediate contextual data during model inference. It only consumes a large amount of memory in ultra-long context scenarios. In ordinary daily conversations or short-text interactions, the KV cache accounts for a relatively small proportion of memory, so the optimization effect is minimal. Full-model quantization, on the other hand, directly compresses the model itself, fundamentally solving the memory demand and covering all scenarios. The technical value and actual impact of these two are simply not on the same level. Moreover, low-bit KV quantization is not a new technology. As early as 2024, the industry already had multiple solutions like RabbitQ and KVQuant that achieved lossless compression at 4 bits or lower. TensorQuant's 3-bit is just a minor iteration, not a theoretical breakthrough. If the technology were as disruptive as claimed, RabbitQ would have shaken the industry when it was launched last year, and Google wouldn't have needed to package it this year for it to explode. At the end of the day, the core of this storm is never about the pros and cons of the technology itself, but about the deliberate academic misconduct and market hype by the Google team. This is also the fundamental reason why TensorQuant is popular but useless. Its popularity relies entirely on Google's PR machine and capital market sentiment trading. But when it comes to practical application, there is still no open-source code, no fair third-party reproduction, and no real production-level deployment. The so-called performance advantage is merely the result of GPU hardware optimization and unfair experimental comparisons. Engineering advantages were packaged as algorithmic breakthroughs—there is no qualitative leap in the technology itself. Combined with this team's consistent focus on papers over implementation, TensorQuant will likely end up like Titans—just a concept stuck in papers, unlikely to truly enter real AI application scenarios. In fact, the tech circle has never lacked such pseudo-breakthroughs. Real technological innovation has always been open-source and reproducible, fair and comparable, and practical and deployable. It is not a myth created by erasing predecessors' work, cheating in experiments, and excessive propaganda. The TensorQuant farce has also torn off the fig leaf of some paper factories in the academic world. Behind seemingly glamorous top conference results and shocking technical data, there may be nothing but unsustainable hype. When looking at these so-called black technologies, we should be more rational and not be led astray by capital and propaganda.
back to top