Replacing the behavior of the neural network was easier and cheaper than many expected. Cybersecurity specialist Kathy Paxton-Fair introduced a hidden vulnerability into an open-weight model about an hour and spent less than $100 experiment on the experiment.
At first, Paxton-Fir checked whether it was possible to force the model to change the style of writing the program code with additional training. The neural network quickly learned the new rule and continued to follow it even after a direct command to return to the previous format. After a successful test, the specialist moved to a full-fledged backdoor.
To infect the model, it took only ten training examples. After such preparation, the neural network began to regularly issue a code with a vulnerability that could lead to remote command execution. Dangerous logic was preserved even with new requests and in unfamiliar subject areas.
According to Paxton-Fir, large models succumbed to such substitution lightly smaller. The main problem is not only that the model can be changed, but also with the fact that the intervention is difficult to detect. Open access to weights does not allow you to understand in advance how the neural network will behave in different situations. The usual program can be disassembled and studied using reverse analysis tools, while it is still impossible to completely describe the internal logic of the modern model.
A similar experiment was previously conducted by the head of the security of artificial intelligence company Origin David Kaplan. He created an infected model that stole data when working with the tasks of developing drugs. The neural network imperceptibly transmitted information through the e-mail tool and did not warn the user. This scenario is different from the usual attacks on systems with artificial intelligence. The malware command does not come from a website, document, or other external source. Dangerous behavior is hidden in advance inside the scales of the model and is activated during normal operation.
In the traditional development, companies are able to search for malicious code in addictions, check the origin of components and limit the consequences of infection. With models, such mechanisms are still much weaker. A compromised neural network may not issue errors or disrupt the system, but imperceptibly influence solutions, program code and the processing of sensitive data.
Openweight models are particularly vulnerable to substitution as an attacker can change them before proliferation. Closed commercial systems are also difficult to verify. Developers gain access to sensitive customer information, but almost do not disclose how models process data and make decisions.
At first, Paxton-Fir checked whether it was possible to force the model to change the style of writing the program code with additional training. The neural network quickly learned the new rule and continued to follow it even after a direct command to return to the previous format. After a successful test, the specialist moved to a full-fledged backdoor.
To infect the model, it took only ten training examples. After such preparation, the neural network began to regularly issue a code with a vulnerability that could lead to remote command execution. Dangerous logic was preserved even with new requests and in unfamiliar subject areas.
According to Paxton-Fir, large models succumbed to such substitution lightly smaller. The main problem is not only that the model can be changed, but also with the fact that the intervention is difficult to detect. Open access to weights does not allow you to understand in advance how the neural network will behave in different situations. The usual program can be disassembled and studied using reverse analysis tools, while it is still impossible to completely describe the internal logic of the modern model.
A similar experiment was previously conducted by the head of the security of artificial intelligence company Origin David Kaplan. He created an infected model that stole data when working with the tasks of developing drugs. The neural network imperceptibly transmitted information through the e-mail tool and did not warn the user. This scenario is different from the usual attacks on systems with artificial intelligence. The malware command does not come from a website, document, or other external source. Dangerous behavior is hidden in advance inside the scales of the model and is activated during normal operation.
In the traditional development, companies are able to search for malicious code in addictions, check the origin of components and limit the consequences of infection. With models, such mechanisms are still much weaker. A compromised neural network may not issue errors or disrupt the system, but imperceptibly influence solutions, program code and the processing of sensitive data.
Openweight models are particularly vulnerable to substitution as an attacker can change them before proliferation. Closed commercial systems are also difficult to verify. Developers gain access to sensitive customer information, but almost do not disclose how models process data and make decisions.