Hackers Exploit AI Safety Filters to Slip Malware Past LLM‑Based Scanners
Security researchers have uncovered a new tactic in which threat actors disguise malicious code by leveraging the safety guardrails built into large‑language‑model (LLM) scanners, allowing the payload to pass undetected through AI‑driven triage tools.
LLM‑powered security platforms have become popular for quickly analysing suspicious scripts and binaries, using natural‑language understanding to flag potentially harmful patterns. These systems rely on safety filters that block content deemed unsafe or offensive, a feature originally intended to prevent the models from generating dangerous instructions.
According to the latest findings, attackers are deliberately crafting malware that triggers those safety mechanisms, causing the model to label the code as benign or to suppress warning messages. By embedding malicious payloads within seemingly innocuous functions and wrapping them in language that the guardrails interpret as harmless, the code can slip past automated analysis that would otherwise flag it in traditional sandboxes or endpoint protection solutions.
The technique was traced by ESET researchers to the Russia‑aligned group known as UAC‑0099. The team observed a series of campaigns in which the actors deployed the evasion method across a variety of payloads, from credential‑stealing scripts to more sophisticated ransomware droppers. While the group’s exact motives remain unclear, the pattern matches previous UAC‑0099 operations that target high‑value enterprises in Europe and North America.
Industry experts warn that the development marks a significant escalation in the cat‑and‑mouse game between defenders and attackers. As LLMs become integral to security workflows, adversaries are adapting their techniques to exploit the very safeguards designed to protect users. Traditional signature‑based detection is already strained by rapid code obfuscation; the added AI‑evasion layer could widen the detection gap.
Vendors are responding by hardening model training pipelines, introducing adversarial testing, and incorporating complementary static and dynamic analysis that does not rely solely on language understanding. Analysts advise organisations to treat AI‑based scanning as an augment rather than a replacement for established security controls, and to monitor for anomalous behaviour that may indicate a successful bypass. The emerging threat underscores the need for a layered defence strategy as the security community continues to grapple with the dual‑use nature of advanced AI technologies.
Comments (0)
Be the first to comment.
Join the discussion