CEOs of companies usually praise products that they create themselves, but the head of Microsoft suddenly warned the business to stay away from one feature of modern neural networks - the one on which his own company is built.
Satya Nadella in a personal post on the social network X said that companies that buy access to language models pay twice. First, with subscription money, and then a much more valuable asset: their own business data, which have to feed neural networks to benefit from it.
Nadella called it a “paradox of inverse information.” The bottom line is that the model eventually learns more and more about the client company, and the company itself learns almost nothing about what is happening on the developer’s side. According to him, neural networks are trained not only on obvious data, but also on the “exhaust” – the requests of employees, the actions of agents and especially on the corrections that people make to the model’s answers. Such information leaks almost imperceptibly, accumulating with each editing.
The irony of the situation is that Microsoft itself has been promoting cloud-based AI services collecting customer business data and invested billions in OpenAI at the dawn of generative neural networks. At the same time, in 2024, the Copilot service faced a similar problem: some large organizations suspended its use because of too broad access rights to internal data in SharePoint and Microsoft 365, which is why the assistant could accidentally open access to sensitive information.
Nadella sees the decision within the companies of its own isolated infrastructure to work with AI, the so-called line of trust, through which no data should be leaked without the owner’s consent. He suggested that business build their own model training environments, independent evaluation systems, retain the rights to data on the use of AI and the results of its work, and also separate the software shell to work with the model itself so that it can be changed without losing the accumulated knowledge.
According to Nadella, the problem is not solved by just competent data management and is structural in nature: any company that uses AI through third-party services is under threat. At the same time, the company directly called Copilot and Azure AI Foundry a ready-made solution to the described problem, although both services remain cloudy and work on the same model that Nadella himself criticized.
Fearing data leakage through neural networks, companies should limit the amount of internal data transmitted to third-party AI services, delimit the rights of assistants to corporate systems and, if possible, store the history of requests and edits separately from the infrastructure of the model supplier.
Satya Nadella in a personal post on the social network X said that companies that buy access to language models pay twice. First, with subscription money, and then a much more valuable asset: their own business data, which have to feed neural networks to benefit from it.
Nadella called it a “paradox of inverse information.” The bottom line is that the model eventually learns more and more about the client company, and the company itself learns almost nothing about what is happening on the developer’s side. According to him, neural networks are trained not only on obvious data, but also on the “exhaust” – the requests of employees, the actions of agents and especially on the corrections that people make to the model’s answers. Such information leaks almost imperceptibly, accumulating with each editing.
The irony of the situation is that Microsoft itself has been promoting cloud-based AI services collecting customer business data and invested billions in OpenAI at the dawn of generative neural networks. At the same time, in 2024, the Copilot service faced a similar problem: some large organizations suspended its use because of too broad access rights to internal data in SharePoint and Microsoft 365, which is why the assistant could accidentally open access to sensitive information.
Nadella sees the decision within the companies of its own isolated infrastructure to work with AI, the so-called line of trust, through which no data should be leaked without the owner’s consent. He suggested that business build their own model training environments, independent evaluation systems, retain the rights to data on the use of AI and the results of its work, and also separate the software shell to work with the model itself so that it can be changed without losing the accumulated knowledge.
According to Nadella, the problem is not solved by just competent data management and is structural in nature: any company that uses AI through third-party services is under threat. At the same time, the company directly called Copilot and Azure AI Foundry a ready-made solution to the described problem, although both services remain cloudy and work on the same model that Nadella himself criticized.
Fearing data leakage through neural networks, companies should limit the amount of internal data transmitted to third-party AI services, delimit the rights of assistants to corporate systems and, if possible, store the history of requests and edits separately from the infrastructure of the model supplier.