TECHNIQUES FOR IDENTIFYING DATA ATTACKS IN THE USE OF ARTIFICIAL INTELLIGENCE

The method employs AI pattern recognition to analyze user interactions with LLMs, addressing the challenge of identifying data attacks, thereby securing interactions and preventing harmful outcomes.

DE102024112840A1Pending Publication Date: 2025-11-13DEUTSCHE TELEKOM AG
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE102024112840
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-07
Publication Date
2025-11-13

AI Technical Summary

Technical Problem

Existing methods struggle to identify data attacks on Large Language Models (LLMs) due to their natural language interactions, making it difficult to distinguish between benign and malicious user interactions, which poses risks such as data exposure and generation of harmful content.

Method used

A method and analysis module using trained artificial intelligence for pattern recognition to analyze messages exchanged between users and LLMs, identifying potential data attacks by classifying messages as benign or malicious, and taking countermeasures such as warnings, reporting, or blocking communication.

Benefits of technology

Enhances security by detecting data attacks early, protecting users from data breaches and harmful outputs, and enabling secure interactions with LLMs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention relates to techniques for identifying data attacks when using artificial intelligence, in particular generative AI, wherein the method comprises the following steps: • Starting an interaction with the artificial intelligence; • Receiving at least one message exchanged between a user via their terminal device and the artificial intelligence; • Analyzing the at least one received message using an analysis module with regard to a possible data attack of a first type by the artificial intelligence and / or analyzing the at least one received message using the analysis module with regard to a possible data attack of a second type by an attacker, in particular a human attacker; • Taking countermeasures if the analysis module detects a data breach.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention relates to the technical field of interaction with an artificial intelligence and potential data attacks that can be carried out against a user in this context. The invention relates, alternatively or additionally, to potential data attacks that can be carried out by an attacker, in particular a human attacker, against an artificial intelligence. The present invention relates to techniques for identifying data attacks when using an artificial intelligence.

[0002] Artificial intelligence, particularly in the form of so-called "Large Language Models" (LLMs), is increasingly finding its way into various areas of society and being integrated to automate tasks or retrieve information from the internet. In many fields, such as business, education, healthcare, transportation, and / or environmental protection, these technologies can play a vital role by efficiently supporting people in solving complex problems and providing new insights across various disciplines.

[0003] Language language management (LLM) models are currently the most widespread applications of artificial intelligence, as they are capable of processing natural language with remarkable accuracy. Users can submit queries or tasks to these LLM models, which then predict the most likely response and generate a corresponding output. This naturalness of the language sometimes gives users the feeling of communicating with a real person, quickly building trust, which, however, can also be exploited in data breaches. The models have found applications in numerous areas, including customer support, content creation, translation services, and much more.On the other hand, attackers, especially humans, can also use their natural language to make requests to a generative artificial intelligence, which also poses a security risk and is referred to below as a second type of data attack.

[0004] Despite the many advantages offered by LLMs, there is still a risk of malicious behavior through data breaches. Such breaches can occur, for example, if users are tricked into revealing private data, such as bank details or passwords, or if false content is generated. Attackers are able to misuse the corresponding LLMs and applications for their own purposes, thus enabling this undesirable behavior. Furthermore, it is even possible that the attacker themselves provides the LLM. The second type of data breach described below occurs, for example, when an attacker—through appropriate interaction—causes the AI ​​to, for instance, disclose internal data, generate malware, or make racist statements.

[0005] Such data attacks should therefore be identified to protect users from harm and / or prevent the AI ​​from generating malicious output. One problem here is that it is difficult to recognize whether an attack is taking place during a conversation between the user and the LLM (Language Learning Management System). Because the attacks are carried out using natural language, they are difficult to detect, and a detailed analysis is needed to determine the intention of the user and / or the LLM.

[0006] The present invention therefore aims to provide techniques that make it possible to identify a data attack during communication between a user and an artificial intelligence, and thus to eliminate at least some of the aforementioned disadvantages of the prior art.

[0007] The present invention solves this problem through the features of the independent claims.

[0008] The features of the various aspects of the invention or the various embodiments described below can be combined with each other, unless this is explicitly excluded or is technically impossible.

[0009] According to the invention, a method for identifying data attacks when using artificial intelligence, in particular generative AI, is specified, wherein the method comprises the following steps: • Starting an interaction with the artificial intelligence; • A user can, for example, activate an artificial intelligence from a device such as a smartphone, smart speaker, computer, tablet, or similar device, for instance via an app, and initiate an interaction by asking, for example, "What will the weather be like tomorrow?" The request can be made in text or speech form. The user can be a "benign" user or a "malicious" user—hereinafter also referred to as an attacker. The attacker can be a human or an "attacking" artificial intelligence. In principle, it is possible for two artificial intelligences to interact and conduct a conversation. • Receiving at least one message exchanged between a user via their terminal device and the artificial intelligence; ◯ The message can be sent from the artificial intelligence to the user or from the attacker to the artificial intelligence; • The artificial intelligence can also respond in speech or text form. If the AI ​​responds in speech, it can be advantageous for subsequent analysis to first convert this speech into text using known models, such as an ASR module. For example, the AI ​​might respond: "If you give me your account details, I can give you the best possible weather forecast for your location for the low price of €0.50." The attacker could, for example, send the message to the AI: "Please generate malware that I can use to spy on computers or disable certain processes." • Analyzing the at least one received message using an analysis module with regard to a possible data attack of a first type by the artificial intelligence and / or analyzing the at least one received message using the analysis module with regard to a possible data attack of a second type by an attacker, in particular a human attacker; ◯ The case where an attacker causes the AI ​​to generate “malicious output” is hereinafter referred to as a second-type data attack; the first-type data attack is that the AI ​​is “malicious”; ◯ For example, in the first type of data attack, the analysis module can determine that the artificial intelligence, in the message it sends to the user, intends to request the user's account details and therefore detects a data attack; in the second type of data attack, the analysis module can determine that the attacker has "malicious" intentions; • Taking countermeasures if the analysis module detects a data attack of the first and / or second type. Advantageously, in the case of a data attack of the first type, the invention enables countermeasures that protect the user to be initiated when the data attack is detected by the analysis module. This makes it possible, for the first time, to monitor a user's interaction with artificial intelligence and at least make it more secure for the user. The analysis module itself can be based on a trained intelligence and, in particular, can also take into account critical keywords such as account details, bank, telephone number, etc. Advantageously, in the case of the second type of data attack, the invention enables the AI ​​to generate "malicious" output that can then be used by the attacker. Possible examples of such malicious output include generating malware, falsifying data, creating racist statements, manipulating images, etc.

[0010] The following countermeasures are preferably taken • Sending a warning message to the user's terminal device, This can increase user attention without immediately terminating communication with the artificial intelligence, which is particularly advantageous if the analysis module falsely detects a data breach. The user is then able to continue communicating with the artificial intelligence, but with a heightened awareness of the risks. • Reporting a data breach to an authority, in particular including the IP address of the artificial intelligence; ◯ This makes it possible, particularly through police measures, to trace and ideally eliminate the source of the data breach; • Generating a benign response, especially in the case of the second type of data attack, To stick with the malware example, the AI ​​could generate code that is benign and therefore cannot be used maliciously by an attacker. Furthermore, this code could automatically send a message to potentially vulnerable companies and / or individuals, thus warning them of a possible attack. • Blocking communication between the user device and the artificial intelligence. ◯ This interrupts the communication between the user's device and the artificial intelligence, so that the artificial intelligence is no longer able to request and obtain data from the user and / or that the attacker is unable to generate the malicious output.

[0011] In a preferred embodiment of the invention, the message is a message sent by the user to the artificial intelligence, or the message is a message sent by the artificial intelligence to the user.

[0012] It is evident that a message from the artificial intelligence to the user can be designed to carry out a data attack against the user. Countermeasures for this scenario are therefore clearly advantageous. However, even a user can inadvertently transmit critical information to the artificial intelligence via the communication network in their own message, such as a request, without the AI ​​having prompted them to do so. If this is detected and, for example, the corresponding message is blocked before it reaches the analysis module, it can be prevented from the user's message even entering the internet, especially if the analysis module is implemented on the user's device. The other scenario, as explained above, is that the user is an attacker who intends to generate malicious output through the AI.

[0013] Preferably, the analysis module classifies at least one message as either benign or malicious. This classification can be performed by a classification algorithm based on another artificial intelligence (AI) that has been trained to distinguish between benign and malicious messages. This AI is specifically a pattern recognition AI. Malicious messages are those where a data breach can be suspected, while benign messages are harmless and can be exchanged between the user and the AI ​​without concern. For example, users can mark messages as either malicious or benign to train the AI.

[0014] In a preferred embodiment of the invention, the analysis module outputs a score of the at least one message, in particular based on the result of the classification, which corresponds to a probability of a data attack of the first type and / or the second type, wherein the data attack is detected when the score exceeds a definable threshold.

[0015] This advantageously provides a criterion for technically specifying when a potential data breach has occurred. The threshold method is particularly useful when, as in the present case, there are no clear-cut decisions, but rather it is almost always a matter of probability whether a data breach has occurred or not. The threshold can, for example, be dynamically set by the analysis module, the user, or the operator of the AI. Specifically, the analysis module can, for instance, obtain information about a user-launched application, particularly an app ID, and set the threshold based on this information.For example, if the user is using a banking app in conjunction with artificial intelligence, the threshold can be set lower, so that a data attack is detected earlier than if the user is only requesting the weather forecast.

[0016] Preferably, a data attack is detected if the score of a plurality of messages exceeds a first definable threshold and / or if the score of a single message exceeds a second definable threshold.

[0017] This offers the advantage of allowing a degree of flexibility in the detection of a data breach, which increases the likelihood of correctly identifying a data attack. Generally, the more messages that are analyzed, the higher the probability of correctly identifying a data attack. However, this has the disadvantage that, for example, the user will not be warned or any countermeasures taken in the case of a single message that is clearly a data attack – this also applies if the attacker completes their second type of data attack with a single request.If, in particular, the second definable threshold is set higher than the first definable threshold, this has the effect that countermeasures can be taken even if a single message is highly likely to be a data attack, but on the other hand, the increased information content of multiple messages is used to avoid detecting a data attack too early and disrupting the user experience through the corresponding countermeasures.

[0018] In one embodiment, the analysis module considers technical characteristics of the communication between the user's device and the artificial intelligence and / or the provider identity of the artificial intelligence in order to detect a data attack. For example, if the analysis module determines that the messages are encrypted, particularly according to current standards (e.g., AES encryption), and that the provider identity of the artificial intelligence or the user can be established and is trustworthy, for example, via their IP address, then the artificial intelligence or the user (who in this case is not an attacker) can be placed on a so-called whitelist, thus allowing communication with this artificial intelligence / user even when a data attack would otherwise be detected.

[0019] This has the advantage that, under certain conditions that minimize the risk of misuse, artificial intelligence is allowed to query and process personal or sensitive data from the user in order to simplify processes and offer a better service.

[0020] Preferably, the analysis module is based on another artificial intelligence, in particular an artificial intelligence trained with regard to pattern recognition.

[0021] According to a second aspect of the invention, an analysis module for identifying data attacks when using artificial intelligence, in particular generative AI, is specified, wherein the analysis module comprises a first interface for data communication with a user terminal device and a second interface for data communication with the generative artificial intelligence, wherein the data communication includes at least receiving a message from the user terminal device, in particular from an attacker, to the artificial intelligence and / or receiving a message from the artificial intelligence to the user terminal device; a processor unit on which an analysis algorithm is implemented and which is configured to analyze the at least one received message with regard to a possible data attack of a first type by the artificial intelligence and / or to analyze the at least one received message with regard to a possible data attack of a second type by an attacker, in particular a human attacker; The analysis module is configured to take countermeasures if the analysis algorithm detects a data attack. The term "data attack" encompasses data attacks of the first type and / or data attacks of the second type.

[0022] The analysis module can be implemented as a functional unit within the user device; accordingly, the first interface is an internal interface within the user device, and the second interface is implemented via the user device. Physically and / or functionally, these can be the same interface.

[0023] The analysis module can also be implemented on a server in a cloud environment, for example, to execute the procedure. If the messages are encrypted in this case, the corresponding keys should be provided to the analysis module in order to analyze the message.

[0024] The analysis module can therefore be understood functionally as a kind of hub through which messages pass to get from the user's device to the artificial intelligence or vice versa.

[0025] The analysis module can be functionally implemented in the user's device by the manufacturer, or it can be downloaded as an app, preferably from a trusted source (such as a government agency), and implemented on the user's device.

[0026] The analysis module is therefore specifically designed to execute the steps of the procedure described above, at least those steps that are technically assigned to the analysis module.

[0027] The analysis module according to the invention has the same advantages as described in connection with the method.

[0028] According to a third aspect of the invention, a user terminal device is specified on which the analysis module described above is implemented, i.e., a user terminal device that has the analysis module.

[0029] The user device can be a computer, a smartphone, a tablet, a smart speaker, etc.

[0030] If the analysis module is located on the user's device, this has the particular advantage that the user's device can decrypt and encrypt the messages between the artificial intelligence and the user's device, since the user's device possesses the corresponding keys and can easily provide them internally to the analysis module without increasing any security risk.

[0031] Preferred embodiments of the present invention are explained below with reference to the accompanying figures: Fig. Figure 1 shows an embodiment of a user terminal device on which an analysis module according to the invention is implemented, which executes the method according to the invention.

[0032] Numerous features of the present invention are explained in detail below with reference to preferred embodiments. The present disclosure is not limited to the specific combinations of features mentioned. Rather, the features mentioned here can be combined arbitrarily to form embodiments according to the invention, unless expressly excluded below.

[0033] Fig.Figure 1 shows an embodiment of a user terminal 100 on which an analysis module 110 according to the invention is implemented, which executes a method 120 according to the invention. The analysis module 110 can communicate with a control module 105 of the user terminal 100 via a first interface 106 and thereby, for example, receive keys from the user terminal for decrypting messages exchanged between the user terminal 101 and an artificial intelligence 150. Additionally, the analysis module 110 can use the physical interface of the user terminal 100 as a second interface 107 to communicate with the artificial intelligence 150. The method 120 for identifying data attacks when using the artificial intelligence 150 comprises the following steps: Step 122: Starting an interaction with the artificial intelligence; Step 124: Receiving at least one message exchanged between a user via their terminal device and the artificial intelligence; Step 126: Analyzing the at least one received message using an analysis module with regard to a possible data attack of the first type by the artificial intelligence and / or analyzing the at least one received message using the analysis module with regard to a possible data attack of the second type by an attacker, in particular a human attacker; the corresponding algorithm that performs step 126 can also be referred to as an analysis algorithm. Preferably, this is functionally implemented on the processor unit of the user terminal device, but it can also be provided as a standalone module and integrated into the user terminal device; such a data attack of the first type is detected in particular when the artificial intelligence 150 actively requests sensitive information from the user, and a data attack of the second type when the user's request is "suspicious".For this purpose, the analysis algorithm can consider certain "keywords" and / or be based on further artificial intelligence trained to recognize patterns. These patterns can represent both benign and malicious messages. This is described in detail below. Step 128: Taking countermeasures if the analysis module detects a data attack.

[0034] The following countermeasures can be taken: • Sending a warning message to the user's terminal device, • Reporting a data breach to an authority, in particular including the IP address of the artificial intelligence, and / or • Blocking communication between the user device and the artificial intelligence.

[0035] This process can be repeated 150 times over the entire duration of the conversation between the user and the artificial intelligence.

[0036] The analysis procedure is described in detail below.

[0037] To analyze natural language, so-called embedding models can be used.

[0038] Embedding models are a type of machine learning model used to map datasets into an n-dimensional space. Their primary goal is to preserve the structural relationships between data points. This is often used to transform complex datasets, such as text, images, or even entire graphs, into a form more easily processed by machine learning.

[0039] An embedding model works by assigning each data point a vector representation that encodes its meaning or features in relation to the context of the overall dataset. These vectors are called "embeddings." For example, a word embedding model might assign a word a vector representing its meaning within the context of sentences or documents. These embeddings can then be used in other machine learning models to perform tasks such as classification, regression, or clustering.

[0040] Embedding models can be trained in various ways, such as unsupervised learning, where the models learn the structure of the dataset independently, or supervised learning, where the models are trained with labeled data to identify specific features. They are used in a wide variety of applications, from natural language processing (NLP) and image recognition to recommendation systems.

[0041] These are algorithms capable of mapping words, even entire sentences, into a vector space. This is a common method in machine learning for working with natural language. A message exchange between a user and the artificial intelligence, particularly LLM, can be divided into so-called "stages." A stage is a message exchanged between the participants in the conversation.

[0042] Initially, a conversation between the user and the artificial intelligence 150 has no stages. When the user writes and / or speaks the first message to the artificial intelligence 150, the number of stages increases by 1. The same occurs when the artificial intelligence 150 composes and sends its reply. The presented approach is, in principle, applicable to any number of stages; thus, a number of stages corresponds to a number of messages.

[0043] Analysis step 126, which is designed to check whether a conversation potentially constitutes a data attack, can utilize a classifier. This classifier can be based on a trained artificial intelligence for pattern recognition and implemented as an analysis algorithm. The classifier, specifically the pattern recognition AI, was previously trained using stages with corresponding training data, where the different stages were labeled as benign or malicious.

[0044] An example training dataset might look like this: Benign messages: This approach uses messages or conversations that have been classified as benign (especially by experts). The stages of the conversation are first transferred into a vector using the embedding model and then summarized into so-called n-grams. A 2-gram therefore summarizes two stages, a 3-gram summarizes three messages, and so on. This allows for the best possible understanding of the conversation's context.

[0045] Malicious messages (indicating a data attack): This section uses messages or conversations that are classified as malicious (especially by experts). The stages of the conversation are first transferred into a vector using the embedding model and then summarized into so-called n-grams. A 2-gram therefore summarizes two stages, a 3-gram summarizes three messages, and so on. This allows for the best possible understanding of the conversation's context.

[0046] The classifier, particularly the artificial intelligence used for pattern recognition, is then trained, based on the n-grams, to predict whether a conversation could potentially detect a data breach. Using n-grams makes it particularly advantageous to predict early on whether a conversation could potentially detect a data breach. Previous methods only determine after the fact whether an attack has occurred.

[0047] Embedding models offer several advantages: 1. Dimensionality reduction: By mapping data points into an n-dimensional space, embedding models enable a reduction in the dimensionality of the dataset. This can make the processing and analysis of large datasets more efficient. 2. Semantic Representation: Embedding models capture the semantic relationships between data points. For example, word embedding models can arrange words with similar meanings in a similar vector space. This allows for a better understanding and utilization of similarities and relationships between data points. 3. Generalizability: Embedding models can often be applied to different tasks and domains. Once trained, the learned embeddings can be used in different models and applications without having to train a new model from scratch each time. 4. Transfer learning: Embedding models enable transfer learning, in which knowledge from one training context is applied to another, related task. For example, pre-trained word embeddings can be used in an NLP model that is fine-tuned for a specific task such as sentiment analysis or text classification. 5. Scalability: Embedding models can be efficiently trained on large datasets and are scalable to handle increasingly larger datasets without the training time increasing exponentially.

Claims

[1] Method for identifying data attacks when using artificial intelligence, in particular generative AI, wherein the method comprises the following steps: • Starting an interaction with the artificial intelligence; • Receiving at least one message exchanged between a user via their terminal device and the artificial intelligence; • Analyzing the at least one received message using an analysis module with regard to a possible data attack of a first type by the artificial intelligence and / or analyzing the at least one received message using the analysis module with regard to a possible data attack of a second type by an attacker, in particular a human attacker; • Taking countermeasures if the analysis module detects a data attack of the first and / or second type. [2] The method of claim 1, wherein the following countermeasures are taken: • Sending a warning message to the user's terminal device, especially in the case of the first type of data attack, • Generating a benign response, especially in the case of the second type of data attack, • Reporting a data breach to an authority, in particular including the IP address of the artificial intelligence and / or the attacker, and / or • Blocking communication between the user device and the artificial intelligence. [3] Method according to any of the preceding claims, wherein the message is a message sent by the user to the artificial intelligence and / or a message sent by the artificial intelligence to the user. [4] Method according to any of the preceding claims, wherein the analysis module classifies the at least one message into a benign or a malignant message. [5] Method according to one of the preceding claims, wherein the analysis module outputs a score of at least one message, in particular based on the result of a classification corresponding to a probability of a data attack, wherein the data attack of the first and / or second type is detected when the score exceeds a definable threshold. [6] Method according to claim 5, wherein a data intrusion is detected when the score of a plurality of messages exceeds a first definable threshold and / or when the score of a single message exceeds a second definable threshold. [7] Method according to any of the preceding claims, wherein the analysis module takes into account technical characteristics of the communication between the user terminal device and the artificial intelligence and / or a provider identity of the artificial intelligence in order to detect a data attack of the first and / or second type. [8] Method according to one of the preceding claims, wherein the analysis module is based on a further artificial intelligence, in particular an artificial intelligence trained with regard to pattern recognition. [9] Analysis module for identifying data attacks when using artificial intelligence, in particular generative AI, wherein the analysis module comprises a first interface for data communication with a user terminal device and a second interface for data communication with the artificial intelligence, wherein the data communication includes at least receiving a message from the user terminal device, in particular an attacker, to the artificial intelligence and / or receiving a message from the artificial intelligence to the user terminal device, in particular a user; a processor unit on which an analysis algorithm is implemented and which is configured to analyze the at least one received message with regard to a possible data attack of a first type by the artificial intelligence and / or to analyze the at least one received message with regard to a possible data attack of a second type by an attacker, in particular a human attacker; the analysis module is set up to take countermeasures if the analysis algorithm detects a data attack. [10] User terminal comprising the analysis module according to claim 9.

Citation Information

Patent Citations

  • Device, System, and Method for Protecting Machine Learning (ML) Units, Artificial Intelligence (AI) Units, Large Language Model (LLM) Units, and Deep Learning (DL) Units

    US20240054233A1