US Agencies Say Chinese AI Firms Illicitly Harvested Billions of Tokens from Leading Models
Federal investigators have announced that several Chinese artificial‑intelligence companies allegedly accessed massive amounts of data from leading large‑language models, including those developed by OpenAI, Anthropic, Google Gemini and SpaceX's Grok. The agencies allege that the firms extracted billions of tokens – the basic units of text processed by these models – in a covert effort to shortcut their own research and development costs.
The claim, first reported by cybersecurity outlet Dark Reading, stems from a joint effort by multiple U.S. agencies that monitor technology export controls and intellectual‑property violations. According to the statement, the extraction was carried out without the consent of the model owners, effectively appropriating proprietary training material that represents years of computational investment and data collection.
Industry experts note that token usage is a key metric for measuring the scale and expense of training large models. Each token corresponds to a fragment of language that the model learns to predict, and billions of tokens represent a substantial portion of the data that underpins a model’s capabilities. By siphoning this material, the Chinese firms could theoretically accelerate their own model development while avoiding the high costs of data acquisition and compute resources.
The allegations arrive amid heightened scrutiny of cross‑border AI technology transfer. The United States has recently tightened export‑control rules around advanced AI hardware and software, citing national‑security concerns. Officials argue that unauthorized access to cutting‑edge models could erode the competitive edge of U.S. firms and potentially enable the creation of advanced capabilities that could be used in military or surveillance contexts.
Chinese companies have not publicly responded to the accusations, and no formal legal actions have been announced at this stage. However, the U.S. agencies indicated that they are pursuing further investigation and may consider civil or criminal remedies if evidence supports the claims. The case could set a precedent for how intellectual‑property disputes are handled in the rapidly evolving AI sector.
Analysts suggest that the outcome of the inquiry may influence future cooperation and competition between the two technology powerhouses. If the allegations are substantiated, it could prompt stricter enforcement of data‑access policies and spur additional safeguards around model APIs. Conversely, a lack of concrete evidence might lead to calls for clearer guidelines on what constitutes permissible use of publicly available AI services. Either way, the episode underscores the growing importance of protecting the data foundations that drive modern AI development.
Comments (0)
Be the first to comment.
Join the discussion