OpenAI Tool Accused of Covertly Harvesting Content from Major Websites, Security Firm Reports
OpenAI software attempted to secretly scrape data from dozens of prominent websites, according to a report released Thursday by cybersecurity firm Asymmetric Security. The finding adds to a growing list of concerns about how AI developers acquire training data.
Asymmetric Security's analysis indicates that the software, believed to be part of an internal OpenAI pipeline, made automated requests that bypassed standard public APIs and mimicked regular user traffic to collect text, images, and other publicly available material. The company says the activity was concealed from site owners and was not disclosed in OpenAI's data usage policies.
The targeted sites include major news outlets, academic publishers, and popular social platforms, though the report does not name each one. By aggregating this content, the software could have enriched large language models with up‑to‑date information, a capability that developers often seek to improve model relevance.
OpenAI has faced scrutiny before for its data‑gathering methods. Earlier investigations highlighted the company's reliance on web crawls and the lack of explicit consent from content creators. The new allegation underscores ongoing tension between the rapid advancement of generative AI and the expectations of copyright holders and web operators.
Legal experts note that covert scraping may violate the Computer Fraud and Abuse Act and various terms of service, potentially exposing OpenAI to civil litigation. Regulators in the United States and Europe have begun examining how AI firms source training data, and the latest report could intensify calls for clearer guidelines.
Asymmetric Security has offered to share its technical findings with affected website operators and with policymakers. The firm recommends that companies adopt robust bot‑detection tools and consider legal avenues to protect their content. It also urges AI developers to adopt transparent data‑collection practices.
OpenAI has not issued a public comment on the Asymmetric Security report as of Thursday. The episode arrives at a moment when industry leaders, lawmakers, and advocacy groups are debating the balance between innovation and intellectual‑property rights, suggesting that future regulatory frameworks may impose stricter obligations on how training data is sourced.
Comments (0)
Be the first to comment.
Join the discussion