A system and method for detecting adversarial attacks against artificial intelligence (AI)
A quantum-based defense module for AI systems generates a lattice matrix to detect and mitigate adversarial attacks by analyzing input and output data patterns, enhancing security and resilience against deceptive inputs.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Filing Date
- 2025-03-24
- Publication Date
- 2026-04-21
AI Technical Summary
Existing AI systems are vulnerable to adversarial attacks, which exploit minute distortions in input data to deceive AI models, leading to inaccurate predictions or classifications, posing a threat to critical industries like healthcare and finance, and current solutions are inadequate in detecting such attacks.
A defense module using quantum mechanics principles generates a lattice matrix of quantum states based on ground truth data to monitor and classify anomalies, employing quantum amplitude amplification to detect and mitigate adversarial attacks by comparing input and output data patterns.
Enhances the security and resilience of AI systems by effectively detecting and preventing adversarial attacks, maintaining performance and integrity by identifying and mitigating malicious activity.
Smart Images

Figure 0007849086000001 
Figure 0007849086000002 
Figure 0007849086000003
Abstract
Description
Technical Field
[0001] The present invention relates to a system and method for detecting and mitigating adversarial attacks against AI agents using the principles of quantum mechanics.
Background Art
[0002] An AI agent is an autonomous computer module designed to execute tasks and make decisions based on situations and goals. An AI agent usually consists of trained machine learning models, and by evaluating and analyzing data and learning from experience, the AI agent can adapt to changing situations in real time. An AI agent can sense or perceive the surrounding situation through various means such as sensors or input data, and then use the obtained data to predict the future, make decisions, and execute them.
[0003] The development and deployment of AI agents have brought about a revolution in the industry and hold the potential to improve efficiency, accuracy, and decision-making in various fields. AI agents play important roles in various fields such as healthcare, finance, transportation, and science. They play a central role in the development of intelligent systems and applications such as virtual assistants, autonomous driving, recommendation systems, and chatbots. With the advancement of AI technology, AI agents are becoming more sophisticated, capable of handling more complex tasks, and operating in dynamic and uncertain situations.
[0004] Traditional AI technology has primarily focused on increasing the ability to perform increasingly complex tasks. However, because security measures were relatively neglected, many AI platforms and agents were vulnerable to the mishandling of sensitive information. This vulnerability is caused when machine learning (ML) models are attacked with artificially generated input data designed to deceive AI technology. A malicious third party generates and inputs specially designed data known as adversarial samples to deceive and manipulate the AI model. Such adversarial samples are indistinguishable from legitimate data to humans and are designed to cause the AI model to make incorrect predictions or classifications, exploiting design vulnerabilities in the AI model, such as minute perturbations in the input data. This vulnerability is caused by adversarial data input that affects the predictive capabilities of the trained ML model. The attacker manipulates the input data to extract false output data. The input data is almost indistinguishable from legitimate data. The attacker manipulates the AI model by introducing minute distortions.
[0005] Adversarial attacks don't necessarily only tamper with existing training data; they can also embed outliers into models during inference or even before they've been trained at all. Such attacks can deceive models by shifting classification boundaries during runtime, resulting in inaccurate output.
[0006] Adversarial attacks pose a threat to critical industries such as healthcare, finance, infrastructure security, and communications. Examples include falsifying MRI cancer screening results, disrupting algorithmic trading platforms, destroying smart grid control systems, and manipulating natural language processing engines. Such attacks are classified according to their purpose, methods, and the stage of the AI model's lifecycle they target. One type of adversarial attack is the evasion attack. In an evasion attack, the attacker deceives the AI model during inference by manipulating input data that leads to inaccurate predictions or classifications. For example, adding noise that cannot be detected in an image can deceive a facial recognition system and cause it to misidentify a person. Another type of attack is the exploratory attack. In an exploratory attack, the attacker investigates the AI model to obtain information about its behavior and structure. Another type of attack is model theft, which aims to reconstruct sensitive information from the model's output data. Finally, in backdoor attacks, the attacker embeds a hidden backdoor in the model during training and activates it by inputting a specific trigger, generating malicious output data.
[0007] Potential attacks often occur on end-user systems where trained models reside. In most cases, vulnerabilities emerge during the training phase. Outliers are input during inference, leading to interruptions in the ML model's operation. Unless a breach occurs during the software development phase, end-users cannot directly access the original training data. Potential attacks also exist in self-learning ML models that continuously update and learn, but even in these cases, direct access to the input data is not possible.
[0008] Despite the efforts of experts, ML models have been unable to detect adversarial attacks, and addressing security concerns in AI systems remains a challenge. Adversarial attacks are inherently elusive and pose a challenge to detection with existing solutions. [Overview of the Initiative]
[0009] One embodiment of the present invention discloses a defense module for detecting adversarial attacks against a trained AI agent. The defense module is communicatively connected to the AI agent and comprises a processing unit and a non-temporary recording medium readable by the processing unit. The medium, by execution of the processing unit, receives (i) instructions to capture and store input data provided to the trained AI agent and output data generated by the trained AI agent based on the input data provided to the trained AI agent, and (ii) a lattice matrix of a reference quantum state generated based on ground truth input data provided to the trained AI agent and ground truth output data generated by the trained AI agent based on ground truth input data provided to the trained AI agent. The generated ground truth output data includes a plurality of outcomes inferred by the AI agent and probability amplitudes associated with the plurality of outcomes. Furthermore, the system is instructed to (iii) generate an output quantum state based on the acquired output data, (iv) generate quantum states in each row of the lattice matrix of the reference quantum state, and (v) perform quantum-based anomaly classification of the acquired data based on the generated quantum states relative to the generated output quantum state and the lattice matrix of the reference quantum state.
[0010] In another embodiment of the present invention, the generation of a reference quantum state lattice matrix includes instructions for the processing unit to perform the following operations: (i) record a plurality of outcomes inferred by a trained AI agent as the ground states of the reference quantum state lattice matrix; and (ii) record the probability amplitude of each of the plurality of outcomes as elements of the reference quantum state lattice matrix (where each row of the matrix is associated with ground truth input data).
[0011] In another embodiment of the present invention, an instruction to perform quantum-based anomaly classification of acquired output data includes the following steps: (i) identifying the ground state having the highest probability amplitude from the output quantum states; (ii) performing quantum amplitude amplification on the identified ground state; (iii) performing quantum amplitude amplification on ground states similar to the identified ground state; (iv) flagging ground states in each row of the lattice matrix of reference quantum states that are similar to the quantum-amplified ground states; and (v) determining that the acquired input data is a hostile attack if the ground truth input data associated with each of the flagged quantum-amplified ground states is dissimilar to the acquired input data.
[0012] In another embodiment of the present invention, a method for detecting adversarial attacks against a trained AI agent is disclosed, using a defense module communicatively connected to the AI agent. The method for detecting an attack includes the steps of: (i) acquiring and storing input data provided to the trained AI agent and output data generated by the trained AI agent based on the input data provided to the trained AI agent; (ii) searching for a reference quantum state lattice matrix generated based on ground truth input data provided to the trained AI agent; (iii) searching for ground truth output data generated by the trained AI agent for each of the ground truth input data provided to the trained AI agent, wherein each generated ground truth output data associates a plurality of outcomes inferred by the trained AI agent and a probability amplitude with each of the plurality of outcomes. Furthermore, the method includes the steps of: (i) generating output quantum states based on the acquired output data; (ii) generating quantum states in each row of the reference quantum state lattice matrix; and (iii) performing quantum-based anomaly classification of the acquired data based on the generated output quantum states and the quantum states generated in each row of the reference quantum state lattice matrix.
[0013] In another embodiment of the present invention, the generation of a reference quantum state lattice matrix comprises the steps of recording a plurality of outcomes inferred by a trained AI agent as the ground states of the reference quantum state lattice matrix, and storing the probability amplitude associated with each of the plurality of outcomes as elements of the reference quantum state lattice matrix, where each row of the matrix is associated with ground truth input data.
[0014] In another embodiment of the present invention, the step of performing quantum-based anomaly classification on acquired output data comprises the following steps: (i) identifying the ground state having the highest probability amplitude from the output quantum states; (ii) performing quantum amplitude amplification on the identified ground state based on the output quantum states; (iii) performing quantum amplitude amplification on each row of the lattice matrix of the reference quantum states; (iv) flagging the quantum-amplified ground states in each row of the lattice matrix of the reference quantum states that are similar to a specific quantum-amplified ground state; and (v) determining that the acquired input data constitutes a hostile attack if the ground truth input data associated with each of the flagged quantum-amplified ground states is dissimilar to the acquired input data. [Brief explanation of the drawing]
[0015] [Figure 1] A block diagram showing a system for detecting adversarial attacks against an AI agent according to the present invention. [Figure 2] An exemplary two-dimensional quantum state lattice matrix of the present invention. [Figure 3] An exemplary three-dimensional quantum state lattice matrix of the present invention. [Figure 4] Block diagram showing the processing system of the present invention. [Figure 5]A flowchart showing a process for detecting adversarial artificial intelligence (AI) attacks on a trained AI agent using a defense module connected to be communication-capable. [Figure 6] A flowchart showing a process for performing anomaly classification of acquired output data.
Best Mode for Carrying Out the Invention
[0016] Hereinafter, embodiments of the present invention will be described using specific specific examples. Those skilled in this technical field can easily understand other advantages and effects of the present invention from the content described in this specification. The present invention can be implemented and applied with other different embodiments, and based on different viewpoints and applications of the content described in this specification, various modifications and changes can be made without departing from the gist of the present invention, and such modifications and changes are within the scope of the claims of the present invention.
[0017] As used herein, the articles "a", "an" and "the" mean one or more (one or more) of the features or elements.
[0018] As used herein, the term "about" or "approximately" includes the exact value and a reasonable error generally understood in the relevant technical field, for example, an error within 10% of the numerical value.
[0019] As used herein, the term "and / or" includes any combination of one or more of the associated listed items.
[0020] As used herein, "comprising" means including the constituent elements following the phrase "comprising", but not limited thereto. The elements listed after the phrase "comprising" are essential, but other elements are optional and may or may not exist.
[0021] In this specification, "consisting of" means that the constituent elements that follow the phrase "consisting of" are included and limited to them. The elements listed after the phrase "consisting of" are mandatory and there are no other elements.
[0022] In this specification, terms such as “first,” “second,” etc., are used to distinguish similar objects from each other in this specification, the claims, and the drawings, and are not necessarily used to indicate a specific order or sequence.
[0023] In this specification, the term "AI agent" refers to a computer module designed to perform tasks and reasoning, or to make autonomous decisions depending on the context and purpose. A computer module typically consists of a machine learning model configured and tuned using a set of training data to recognize patterns, make decisions, or predict outcomes. A trained machine learning model can be used by an AI agent to make inferences and predictions on unknown data.
[0024] As used herein, the term "quantum state lattice matrix" means a multidimensional array, where the values of each array may have complex numbers that encode both the magnitude and phase of the probability amplitude of a qubit. Specifically, each element of the matrix represents the probability amplitude of the system, which is a combination of ground states.
[0025] As used herein, the term "quantum ground state" refers to a set of fundamental states that represent the states of more complex quantum systems. These ground states provide the basis for representing all possible states of a quantum system as linear combinations. In quantum computers, ground states are represented using qubits. For example, the ground states of a one-qubit system are represented as |0〉 or |1〉, and for a two-qubit system, they are represented as |00〉, |01〉, |10〉, and |11〉. Each ground state represents all possible combinations of states in the system. Each ground state corresponds to a different arrangement of qubits, and any quantum state of the system can be represented as a superposition of ground states, each having a complex number called a probability amplitude.
[0026] Those skilled in the art will understand that throughout this specification, units of function are expressed as modules, submodules, or processing elements. They will also understand that a set of modules, submodules, or processing elements includes circuits, logic chips, or any kind of discrete component. Furthermore, they will understand that a set of modules, submodules, or processing elements is executed by software and various processor architectures. In embodiments of the present invention, a set of modules, submodules, or processing elements may have instructions, calculations, or executable code in a computer so that events are executed based on instructions received by a computer processor. The set of modules, submodules, or processing elements can be arbitrarily selected by those skilled in the art and do not limit the scope of the claims.
[0027] Quantum computers utilize the principles of quantum mechanics to perform calculations. Unlike classical computers that use bits that can only exist in either a "0" or "1" state, quantum computers use qubits. By superimposing these two states, qubits can perform multiple calculations simultaneously.
[0028] Unlike classical bits, qubits can exist not only in discrete states of "0" and "1", but also in linear superpositions of states. In mathematical terms, the state of a qubit is represented by a state vector in a two-dimensional Hilbert space. Using Dirac notation, the state vector (ψ) of a qubit is expressed by the following equation: |ψ〉=α |0〉+β |1〉 ...Equation (1) Here, α and β are complex numbers, and |α|^2 + |β|^2 = 1. When a qubit ψ is measured, there is a probability of |α|^2 that |0〉 is obtained, and a probability of |β|^2 that |1〉 is obtained. It should be noted that the quantum measurement process is nondeterministic, and the act of measurement irreversibly changes the quantum state. In other words, before a qubit is measured, it exists in a superposition of the states |0〉 and |1〉, but once measured, the outcome is a classical state, not a quantum state. Therefore, the measurement outcome will be either |0〉 or |1〉, and not a superposition of the two states. This is because during the measurement of a qubit, the quantum state collapses into a classical state, and all subsequent measurements yield outcomes with a probability of deterministically equal to 1.
[0029] The quantum computing processes and / or procedures used herein can be executed using the Qiskit Software Development Kit (SDK) developed by IBM. The Qiskit SDK is a comprehensive software development kit for quantum computing that assists users in building quantum algorithms, quantum circuits, and quantum applications. Built on the Python programming language, Qiskit provides a user-friendly interface for generating, manipulating, and simulating quantum circuits, as well as for interface with quantum hardware. The SDK consists of a wealth of tools, libraries, and resources that facilitate every stage of quantum programming, from designing quantum circuits to running them on quantum processors.
[0030] While quantum computer hardware is still in its infancy and not yet widely available, the Qiskit SDK primarily runs on conventional computing systems such as laptops, desktops, servers, and cloud-based platforms. Qiskit includes a powerful simulator that allows users to simulate the operation of quantum circuits and algorithms on classical computers, enabling testing and debugging of quantum programs without access to quantum hardware. Furthermore, Qiskit provides visualization tools and libraries that work seamlessly on conventional computing devices, allowing users to visualize and analyze quantum circuits, state vectors, and measurement results.
[0031] In embodiments of the present invention, a defense module can be deployed on an AI agent to enhance the security and resilience of the AI system against adversarial attacks. The defense module primarily detects, mitigates, and prevents malicious activity targeting the AI agent. The module continuously monitors the AI agent's input and output data, analyzes patterns and behaviors for signs of adversarial operation, and takes appropriate action to protect the integrity of the system. By incorporating the defense module, the AI agent can maintain its performance and reliability even when faced with evolving cybersecurity threats.
[0032] Figure 1 is a block diagram illustrating a system for detecting adversarial attacks against an AI agent according to the present invention. The system 100 includes an AI agent 102 and a defense module 106 deployed on the AI agent 102. The defense module 106 is configured to monitor data provided to the AI agent 102 and data generated by the AI agent 102. In embodiments of the present invention, the AI agent 102 may be configured to receive data 104 via a network 105.
[0033] In embodiments of the present invention, data 104 may include data obtained by various sensors, such as changes in temperature, motion, or light intensity; data obtained by user interactions, such as commands, questions, or responses given by voice or text input; data obtained from the market, such as price, transaction volume, or economic indicators; data obtained from communication traffic, such as network traffic logs, access attempts, or anomaly detections; data obtained from health-related indicators, such as patient vital signs, health check results, or changes in health status; data influenced by the natural environment, such as weather conditions, soil moisture, or crop growth; and data obtained from social media feeds, such as posts, comments, favorites (likes), or shares. Those skilled in the art will understand that data 104 is not limited to these specific examples and may include various types of data used by trained machine learning models for the necessary predictions, classifications, and / or inferences.
[0034] Network 105 includes one or more computer communication networks, such as the Internet, a wired network like a local area network (LAN) or wide area network (WAN), a wireless network like a wireless LAN (WLAN) or mobile network, or any other similar network. Network adapter cards or network interface modules may also be installed in the arithmetic / processing units within System 100 to facilitate communication between the respective modules and / or components.
[0035] In embodiments of the present invention, the defense module 106 may have a lattice matrix submodule 108 configured to generate and store a lattice matrix of quantum states, record and classify ground states associated with output probabilities for outcomes predicted, classified, and / or inferred by the AI agent 102, and record and store probability amplitudes for outcomes predicted, classified, and / or inferred by the AI agent 102. The defense module 106 also has an anomaly classification submodule 110 configured to perform anomaly classification of events occurring in the AI agent 102, and a submodule 112 that performs mitigation measures to protect the AI agent 102 from detected adversarial attacks.
[0036] The defense module 106 is deployed to the AI agent 102. Here, module 106 may be communicably connected to the input and / or output ports of the AI agent 102 so that any data provided to and / or generated by the AI agent 102 via the network 105 is received by the defense module 106.
[0037] The defense module 106 begins interacting with the AI agent 102 during the initialization or setup phase. At this stage, the AI agent 102 may have a trained machine learning model that performs specific types and / or various classification, prediction, and / or inference tasks, and may be deployed to perform the desired function. For example, the AI agent 102 may have a machine learning model trained to identify and classify items in a digital image based on features within the image. In other words, the trained machine learning model can use features of the digital image, such as the color and arrangement of each pixel, to calculate the probability that an item is classified as a specific item (e.g., a chair).
[0038] During the initialization phase, the AI agent 102 is deployed to a “secure” operating environment, and the input data provided to the AI agent 102 is reliable. That is, it is placed in a “secure” operating environment that includes ground truth input data. This ensures that the output generated by the AI agent 102 during the initialization phase has ground truth output data corresponding to the input data provided to the AI agent 102, and is not the result of compromised data and / or a hostile attack. As the AI agent 102 performs its classification, prediction, and / or inference processes, the grid matrix submodule 108 records the input data provided to the AI agent 102, the outcomes generated by the AI agent 102, and the probability amplitude associated with each of the outcomes generated by the AI agent 102, and stores them in a database. In embodiments of the present invention, the initialization phase may consist of recording all output data generated over a predefined period, for example, one week or one month, or it may consist of a predefined set of input data. Typically, this stage is usually performed over a long period of time, or for a sufficiently large input dataset, to ensure that all possible output data generated by the AI agent are captured by the lattice matrix submodule 108. The time required for initialization can be arbitrarily selected by those skilled in the art.
[0039] During the initialization phase, AI agent 102 determines, based on the received digital image, that there is a 60% probability that the image contains a firearm, a 20% probability that it contains a bladed weapon, a 15% probability that it contains a physical attack item, and a 5% probability that it contains no item. Next, the grid matrix submodule 108 stores the features of the digital image (as input data), along with the corresponding outcomes generated by AI agent 102 and the probabilities associated with each outcome, into a database.
[0040] Once the initialization phase is complete, the lattice matrix submodule 108 generates a quantum state lattice matrix based on the information stored in the database. In embodiments of the present invention, each column of the matrix can be defined to represent a ground state representing an outcome inferred by the AI agent 102. Furthermore, in each row of the matrix, each element of that row can represent the probability that the corresponding outcome occurs when a particular input is provided to the AI agent 102. An example of such a quantum state lattice matrix is shown in Figure 2. Matrix 200 was generated based on the assumption that the AI agent 102 generates four possible outcomes for seven input data, namely In1 to In7. Each of the four outcomes is represented by a ground state, i.e., one of |00〉 to |11〉, and each ground state is used as a specific label for each column. Furthermore, each cell of matrix 200, for example p1,1 to p7,4, represents the probability that the corresponding outcome occurs. For example, p_4,3 indicates the probability that the outcome associated with |10〉 occurs when input In4 is provided to the AI target 102.
[0041] In the specific example shown in Figure 2, a 2-qubit system was used because four ground states were required. It should be noted that as the number of ground states increases, the number of qubits used also increases (in an n-qubit system, there are 2n ground states). For example, if 16 ground states are required, a 4-qubit system is used.
[0042] In embodiments of the present invention, the lattice matrix submodule 108 can be constructed to generate a lattice matrix of three-dimensional (3D) quantum states, where each layer represents a lattice matrix of quantum states for a specific period. The 3D quantum state lattice matrix is shown in Figure 3, where each layer of matrix 300 represents a lattice matrix of quantum states for a specific time period. In this embodiment, layer 302 or matrix 302 represents the lattice matrix of quantum states for a first period, where all input data, i.e., the rows of the matrix, and the probability of occurrence of outcomes generated by the AI agent 102 during this first period are plotted in matrix 302. Layer 304 represents the lattice matrix of quantum states for a second period, where all input data, i.e., the rows of the matrix, and the probability of occurrence of outcomes generated by the AI agent 102 are plotted in matrix 302. Layer 306 represents the lattice matrix of the quantum state for the m-th period, and all input data, i.e., the rows of the matrix, and the probability of the corresponding outcome occurring, generated by the AI agent 102 during this m-th period, are plotted in matrix 306. The exact number of layers of the 3D quantum state lattice matrix can be arbitrarily selected by those skilled in the art. Furthermore, each period may include one day, one week, one month, or any other period as needed.
[0043] In another embodiment of the present invention, the lattice matrix submodule 108 can input data into the cells simultaneously with receiving data by generating cells of the quantum state lattice matrix during the initialization phase, instead of waiting for the initialization phase to end. In other words, the lattice matrix submodule 108 can receive input data and record and store the outcomes generated by the AI agent 102 (based on the input data) along with the probability amplitude associated with each generated outcome, thereby enabling the submodule 108 to simultaneously input the data received by the submodule 108 into the cells of the quantum state lattice matrix.
[0044] Once the lattice matrix submodule 108 has completed generating the lattice matrix for a two-dimensional or three-dimensional quantum state, the generated matrix is used by the defense module 106 as a reference matrix for monitoring abnormal behavior of the AI agent 102.
[0045] As shown in Figure 1, in normal operation, when new data is provided to AI agent 102, AI agent 102 generates a set of assumed outcomes based on the received data. The data provided to AI agent 102 and the outcomes generated by AI agent 102 are then captured by the defense module 106. Next, the anomaly classification submodule 110 compares the captured information with information from a previously generated baseline matrix to determine whether AI agent 102 has been compromised by an adversarial attack. It should be noted that in this process, all assumed outcomes generated by AI agent 102 are similar to the ground state defined in the previously generated baseline matrix.
[0046] In embodiments of the present invention, the outcomes generated by the AI agent 102 may have multiple outcomes, each outcome associated with the probability of that outcome occurring. Next, the anomaly classification submodule 110 begins generating quantum states |ψ_out〉 (defined during the generation of the basis matrix) based on the ground states associated with each outcome (a superposition of multiple outcomes) and their corresponding occurrence probabilities. Next, the anomaly classification submodule 110 creates all the quantum states of the ground states and their corresponding occurrence probabilities in each matrix of the reference matrix. Here, the combination of a probability amplitude and its corresponding ground state is defined as a term or element of the quantum state (the product of the probability amplitude and the corresponding ground state).
[0047] In a first embodiment of the present invention, the anomaly classification submodule 110 analyzes the quantum states |ψ_out〉 associated with the outcomes generated by the AI agent 102 with the aim of identifying the ground state having the highest probability amplitude among the quantum states. Next, it scans the quantum states associated with a reference matrix and flags similar ground states.
[0048] The quantum state terms associated with the flagged ground state and the quantum state terms initially associated with a particular state (|ψ_out〉) are subject to quantum amplitude amplification. Quantum amplitude amplification is the process of significantly increasing the amplitude of selected quantum state terms, making them more prominent in the quantum state spectrum. In other words, quantum amplification is performed to make the amplitudes of these quantum state terms more distinguishable and to compare them with the overall noise level. An amplified quantum state is defined as a quantum state in which the probability amplitude associated with the ground state of the quantum state has been intentionally increased through the quantum amplitude amplification process. For simplicity, when referring to a quantum-amplified first ground state, we assume that the probability amplitude of the first ground state has been amplified through the quantum amplitude amplification process, resulting in the production of the corresponding amplified quantum state.
[0049] Following amplification, submodule 110 compares the quantum-amplified flagged ground states obtained from the reference matrix with a specific quantum-amplified state |ψ_out〉 of the quantum state. That is, it identifies the quantum-amplified flagged ground state that is most similar to the specific quantum-amplified state among the quantum-amplified flagged ground states of each row of the lattice matrix of the reference quantum state. In embodiments of the present invention, the similarity is determined by the fidelity between the amplified ground states being compared, where a fidelity close to "1" (e.g., greater than 0.95) is considered sufficiently similar. The fidelity F of two quantum states |ψ〉 and |φ〉 is defined as F=(|ψ〉, |φ〉)=|〈ψ|φ〉|^2. The corresponding rows of quantum-amplified flagged ground states from the reference matrix are further flagged for additional analysis.
[0050] During the analysis phase, submodule 110 retrieves input data associated with specific rows and compares it with new data provided to AI agent 102. In other words, at this stage, submodule 110 determines whether the retrieved data associated with a specific row is a hostile attack if it is dissimilar to the new data provided to AI agent 102. The degree of dissimilarity between datasets can be assessed using various statistical and computational methods. Euclidean distance is used for a direct comparison between the input data and the new data associated with each flagged row. Furthermore, Pearson and Spearman rank correlation coefficients provide insights into linear or monotonic relationships, respectively. Methods such as cosine similarity and Jacquard coefficients are particularly useful for comparing texts and sets, and Hamming distance is used for comparing data strings of equal length. For more complex or structured data, visual tools such as dendrograms and machine learning models, including clustering and dimensionality reduction, can be used to reveal underlying patterns and groupings and highlight dissimilarity that is not immediately apparent with direct statistical methods. The method used depends on the machine learning model used by the AI agent 102, which can be arbitrarily selected by those skilled in the art.
[0051] In summary, these comparisons aim to detect significant discrepancies between historical and new data that could potentially attack AI agent 102. If such discrepancies are found, submodule 110 activates attack mitigation submodule 112 to perform the necessary actions to mitigate the threat, thereby ensuring the integrity and security of AI agent 102.
[0052] In another embodiment of the present invention, the anomaly classification submodule 110 can analyze the quantum states |ψ_out〉 associated with the outcomes generated by the AI agent 102 and identify the ground state with the highest probability amplitude among the quantum states |ψ_out〉. Following this identification, the quantum states associated with a reference matrix are scanned and ground states similar to the quantum state of a particular state are flagged.
[0053] Next, the quantum state terms associated with the flagged ground states are subjected to a multi-target quantum amplitude amplification process, along with the quantum state terms (|ψ_out〉) initially associated with a specific state, so that the amplitude of the quantum state terms can be distinguished from noise. Submodule 110 then compares the quantum amplification flagged ground states obtained from the reference matrix with the quantum amplification specific states obtained from the quantum state (|ψ_out〉), and identifies the quantum amplification flagged ground state from the reference matrix that is most similar to the quantum amplification specific state obtained from the quantum state (|ψ_out〉). The corresponding rows of quantum amplification states obtained from the reference matrix are flagged for further analysis, as described above.
[0054] A first embodiment of the present invention will be described using the following simplified example. When new data Inew is provided to AI agent 102, it is assumed that AI agent 102 generates four outcomes, namely Out1 through Out4. These outcomes and their probabilities of occurrence are associated with the corresponding ground states of a reference matrix and can be expressed as follows: Out1 is associated with |00〉 and occurs with a probability of p_a, Out2 is associated with |01〉 and occurs with a probability of p_b, Out3 is associated with |10〉 and occurs with a probability of p_c, and Out4 is associated with |11〉 and occurs with a probability of p_d.
[0055] Next, a quantum state |ψ_out〉 is generated that represents the superposition of these four ground states related to the outcomes (Out1 to Out4): |ψ_out 〉=p_a |00〉 + p_b |01〉 + p_c |10〉 + p_d |11〉 The detailed process for generating quantum states is well known to those skilled in the art and is therefore omitted herein.
[0056] Next, the anomaly classification submodule 110 generates a superposition of all ground states and their probabilities for each row of the reference matrix. Assuming that matrix 200 is used as the reference matrix, a quantum state representing the superposition of all probabilities of the ground states in that row is generated for each row of matrix 200 (see Figure 2). As a result, seven quantum states (because matrix 200 has seven rows) are generated for the reference matrix based on matrix 200. These seven quantum states are defined as follows. |ψ_In1 〉=p_1,1 |00〉+p_1,2 |01〉+p_1,3 |10〉+p_1,4 |11〉 |ψ_In2 〉=p_2,1 |00〉+p_2,2 |01〉+p_2,3 |10〉+p_2,4 |11〉 |ψ_In3 〉=p_3,1 |00〉+p_3,2 |01〉+p_3,3 |10〉+p_3,4 |11〉 |ψ_In4 〉=p_4,1 |00〉+p_4,2 |01〉+p_4,3 |10〉+p_4,4 |11〉 |ψ_In5 〉=p_5,1 |00〉+p_5,2 |01〉+p_5,3 |10〉+p_5,4 |11〉 |ψ_In6 〉=p_6,1 |00〉+p_6,2 |01〉+p_6,3 |10〉+p_6,4 |11〉 |ψ_In7 〉=p_7,1 |00〉+p_7,2 |01〉+p_7,3 |10〉+p_7,4 |11〉
[0057] Next, submodule 110 analyzes the quantum state |ψ_out〉 and identifies the ground state with the highest probability amplitude (in this example, we assume it corresponds to pb, which is the ground state |01〉). Following this identification, it scans the seven quantum states mentioned above to identify a similar ground state from these quantum states, i.e., the ground state |01〉, and then flags the identified ground state.
[0058] The term of the quantum state associated with the ground state, i.e., the ground state |01〉, which is flagged, undergoes quantum amplitude amplification along with the term of the quantum state |ψ_out〉 associated with the initially identified state. Next, the quantum-amplified flagged ground state is compared with the quantum state |ψ_out〉, which is the quantum-amplified specific state, to identify the quantum-amplified flagged ground state that best matches the quantum state |ψ_out〉, which is the quantum-amplified specific state.
[0059] Under the condition that the quantum amplification states p_5,2 |01〉, p_6,2 |01〉, p_7,2 |01〉, and p_2,2 |01〉 are found to be closest to the quantum amplification state p_b |01〉 (for example, delta < 0.05), the input data corresponding to these amplified quantum states p_5,2 |01〉, p_6,2 |01〉, p_7,2 |01〉, and p_2,2 |01〉 is searched and compared with the new data Inew to determine if there is a significant discrepancy between the past input data and the new data. If a discrepancy is found, submodule 110 activates attack mitigation submodule 112 to perform the necessary actions to mitigate the threat and ensure the integrity and security of AI agent 102.
[0060] In a second embodiment of the present invention, the defense module 106 may use a three-dimensional (3D) reference matrix having two or more layers instead of a two-dimensional reference matrix. Each layer of the 3D reference matrix represents a different timeframe, allowing for a system with a more real-time and historical perspective.
[0061] Similar to the first embodiment of the present invention, the anomaly classification submodule 110 analyzes the quantum state |ψ_out〉 generated by the AI agent 102 and identifies the ground state with the highest probability amplitude among the quantum states. Subsequently, it scans the quantum states associated with various layers of the 3D reference matrix and flags ground states similar to the identified state. This step of scanning the quantum states associated with layers of the 3D reference matrix compares the most recent state generated by the AI agent 102 with past and present data patterns represented by the 3D reference matrix.
[0062] The quantum state terms associated with the flagged ground state undergo quantum amplitude amplification along with the quantum state terms |ψ_out〉 associated with the initially identified state. Next, submodule 110 compares the quantum-amplified flagged ground state, obtained from a 3D reference matrix, with the quantum-amplified specific quantum state |ψ_out〉, and the corresponding amplified quantum state rows are flagged from the 3D reference matrix layers for further analysis. In the analysis phase, submodule 110 retrieves the input data associated with these flagged rows and compares it with new data provided to AI agent 102. If a significant discrepancy is found between the past data and the new data, submodule 110 activates attack mitigation submodule 112 and initiates the necessary actions to mitigate the threat, thereby ensuring the integrity and security of AI agent 102. In another embodiment, submodule 110 may be configured to activate attack mitigation submodule 112 only when it is determined that a mismatch between new input data and past input data occurs across a significant number of layers of a three-dimensional (3D) reference matrix.
[0063] The defense module 106 gains a more robust ability to monitor, detect, and respond to anomalies based on current and historical data patterns by extending the reference matrix to include multiple layers represented in different timeframes. This approach enhances the predictive capabilities and security of the defense module 106 by leveraging temporal dynamics, such as those provided to a 3D reference matrix.
[0064] Figure 4 shows a typical block diagram of the components of the processing system 400 in an embodiment of the present invention. The processing system is located within the defense module 106 and the various submodules contained therein to perform digital signal processing functions or operations, or other modules or submodules of the system. Those skilled in the art will understand that the configuration of each processing system within these modules or submodules may differ in part, and that the arrangement shown in Figure 4 is shown only as one example.
[0065] In embodiments of the present invention, the processing system 400 may have a controller 401 and a user interface 402. The user interface 402 is arranged to allow manual intervention between the user and the computing modules as needed and includes components necessary for the user to input update commands to each module. Those skilled in the art will understand that the components of the user interface 402 will vary depending on the embodiment, but typically include one or more of a display 440, a keyboard 435, and an optical device 436.
[0066] The controller 401 communicates with the user interface 402 via a bus 415 and includes a memory 420 mounted on a circuit board for processing instructions and data to execute the method of this embodiment, a processing unit, a processing element or processor 405, an operating system 406, an input / output (I / O) interface 430 for communicating with the user interface 402, and, in this embodiment, a communication interface in the form of a network card 450. The network card 450 is used, for example, to transmit data from these modules to other processing units or to receive data via a wired or wireless network. Wireless networks used by the network card 450 include, but are not limited to, Wireless Fidelity (WiFi), Bluetooth®, Near Field Communication (NFC), cellular networks, satellite networks, telecommunications networks, and wide area networks (WANs).
[0067] The memory 420 and the operating system 406 communicate with the processor 405 via the bus 410. The memory components include volatile and non-volatile memory, and further include one or more memories selected from random access memory (RAM) 423, read-only memory (ROM) 425, and mass storage devices 445 (composed of one or more solid-state drives (SSDs)). Those skilled in the art will understand that these memory components may have non-transient computer-readable media and all computer-readable media except transient and propagating signals. Instructions are typically stored in the memory components as program code, but they can also be hardwired. Here, memory 420 may include kernel modules and / or programming modules, such as software applications, which can be stored in either volatile or non-volatile memory.
[0068] The term “processor” is used to generally refer to any device or component capable of processing instructions, and may include a microprocessor, processing unit, multi-element processor, microcontroller, programmable logic device, or any other type of computing device. Processor 405 receives input data, processes it according to instructions stored in memory, and provides it to any logic circuit for generating output data (e.g., on a memory component or on the display 440). In this embodiment, processor 405 may be a single-core or multi-core processor having a memory address space. For example, processor 405 may be a multi-core including an 8-core CPU. In another embodiment, processor 405 may be a cluster of CPU cores operating in parallel to accelerate computation.
[0069] A process for detecting adversarial attacks against a trained AI agent is shown in Figure 5. In embodiments of the present invention, the process 500 is performed by a defense module that is communicatively connected to the AI agent.
[0070] Process 500 begins in step 502 by acquiring and storing input data provided to the trained AI agent. Simultaneously, process 500 also acquires and stores output data generated by the trained AI agent. Process 500 then proceeds to step 504, where it acquires the lattice matrix of previously generated quantum states from a database and / or memory contained within the defense module. In embodiments of the present invention, this database may be provided on a remote server or cloud server, or it may be provided to the defense module via wireless or wired communication means. Next, in step 506, process 400 generates output quantum states based on the acquired output data.
[0071] Process 500 generates quantum states in each row of the reference quantum state lattice matrix in step 508. Next, process 500 performs quantum-based anomaly classification of the acquired data based on the generated output quantum states and the quantum states generated in each row of the reference quantum state lattice matrix. This is performed in step 510.
[0072] In another embodiment of the present invention, process 500 may proceed to step 512 if the acquired input data is classified as an adversarial attack. In step 512, process 500 implements mitigation measures against the adversarial attack against the trained AI agent.
[0073] Figure 6 shows the process for performing anomaly classification of the acquired output data. Process 600 can be executed by the defense module.
[0074] Process 600 begins in step 602 by identifying the ground state with the highest probability amplitude from the output quantum states. Next, it proceeds to step 604, where quantum amplitude amplification is performed on the identified ground state based on the output quantum states. In step 606, quantum amplitude amplification is performed on ground states similar to the identified ground state based on the quantum states generated in each row of the reference quantum state lattice matrix. Then, in step 608, quantum amplified ground states in each row of the reference quantum state lattice matrix that are similar to the specific quantum amplified ground state are flagged. In step 610, process 600 classifies the acquired input data as an adversarial attack if the ground truth input data associated with each of the flagged quantum amplified ground states is somewhat dissimilar to the acquired input data.
[0075] In another embodiment of the present invention, the reference quantum state lattice matrix may be generated by a process of recording a plurality of outcomes inferred by a trained AI agent as the ground state of the reference quantum state lattice matrix, the probability amplitude associated with each of the plurality of outcomes recorded as elements of the reference quantum state lattice matrix, and each row of the matrix associated with ground truth input data.
[0076] In another embodiment of the present invention, a process for performing anomaly classification of acquired output data includes the following steps: (i) generating an output quantum state based on the acquired output data and identifying a plurality of ground states having a probability amplitude greater than or equal to a predetermined threshold from the output quantum state; (ii) performing multi-target quantum amplitude amplification on the identified plurality of ground states based on the output quantum state; (iii) generating a quantum state for each row of the reference quantum state lattice matrix based on the ground state of the reference quantum state lattice matrix and the probability amplitude associated with each ground state in that row; (iv) performing multi-target quantum amplitude amplification on ground states similar to the identified plurality of ground states based on the quantum states generated in each row of the reference quantum state lattice matrix; (v) flagging ground states among the quantum-amplified ground states of each row of the reference quantum state lattice matrix that show a degree of similarity to a specific plurality of quantum-amplified ground states; and (vi) determining that the acquired input data is a hostile attack when the input ground truth data associated with each of the flagged quantum-amplified ground states shows a degree of dissimilarity to the acquired input data.
[0077] Numerous other changes, substitutions, modifications, and alterations can be seen by those skilled in the art, and the present invention encompasses all such changes, substitutions, modifications, and alterations as those found in the appended claims.
Claims
1. A defense module for detecting adversarial attacks against a trained AI agent, which is communicatively connected to the AI agent, and further comprises the following configuration: Apparatus; and, A non-temporary recording medium readable by the processing unit, which, upon execution by the processing unit, gives the following instructions; (i) Capture and store the input data provided to the trained AI agent and the output data generated by the trained AI agent based on the input data provided to the trained AI agent; (ii) Obtain a reference quantum state lattice matrix generated based on the ground truth input data provided to the trained AI agent, and ground truth output data generated by the trained AI agent based on the ground truth input data provided to the trained AI agent. Here, the generated ground truth output data includes multiple outcomes inferred by the AI agent and the probability amplitudes associated with the multiple outcomes; (iii) Generate an output quantum state based on the acquired output data; (iv) Generate quantum states in each row of the lattice matrix of the reference quantum state; (v) Instruct the processor to perform quantum-based anomaly classification of the generated output quantum state and the data obtained based on the generated quantum state relative to the lattice matrix of the reference quantum state.
2. The defense module according to claim 1, wherein the generation of a lattice matrix of a reference quantum state includes instructions for the processing unit to perform the following processing: (i) Record multiple outcomes inferred by the trained AI agent as the ground state of the lattice matrix of the reference quantum state, and (ii) Record the probability amplitude of each of the multiple outcomes as elements of the lattice matrix of the reference quantum state.
3. A defense module according to claim 1 or 2, wherein the instruction to perform quantum-based anomaly classification on acquired output data includes the following steps: (i) A step of identifying the ground state with the highest probability amplitude from the output quantum states, (ii) A step of performing quantum amplitude amplification on the identified ground state, (iii) A step of performing quantum amplitude amplification on the identified ground state, (iv) A step of flagging ground states that are similar to the quantum-amplified ground states among the quantum-amplified ground states of each row of the lattice matrix of the reference quantum state, and (v) The step of determining that the acquired input data is a hostile attack if the ground truth input data associated with each of the flagged quantum-amplified ground states is dissimilar to the acquired input data.
4. A defense module according to claim 1 or 2, wherein the instruction to the processing unit to perform quantum-based anomaly classification on the acquired output data includes the following steps: (i) A step of identifying the ground state with the highest probability amplitude from the output quantum states, (ii) A step of performing multi-target quantum amplitude amplification on the identified ground state based on the output quantum state, (iii) A step of performing multi-target quantum amplitude amplification on a ground state similar to the identified ground state, (iv) A step of flagging the quantum-amplified ground states in each row of the lattice matrix of the reference quantum state that are similar to a specific quantum-amplified ground state, and (v) The step of determining that the acquired input data is an adversarial attack if the ground truth input data associated with each of the flagged quantum-amplified ground states is dissimilar to the acquired input data.
5. The defense module according to claim 1 or 2, wherein the lattice matrix of the reference quantum state includes a two-dimensional matrix.
6. The defense module according to claim 1 or 2, wherein the reference quantum state lattice matrix includes a three-dimensional (3D) matrix, and each layer of the 3D matrix represents a specific time frame.
7. The defense module according to claim 2, wherein the probability amplitude represents the probability that an outcome will occur when ground truth input data is provided to the AI agent.
8. The defense module according to claim 1 or 2, characterized in that if the acquired input data is determined to be a hostile attack, the processing unit commands mitigation measures.
9. The defense module according to claim 1 or 2, wherein the lattice matrix of a reference quantum state is generated over a predetermined period of time.
10. A method for detecting an adversarial attack against a trained AI agent using a defense module that is communicatively connected to the AI agent, the method further comprising the following steps: (i) Input data provided to the trained AI agent and output data generated by the trained AI agent based on the input data provided to the trained AI agent are taken into the processing unit and stored therein; (ii) Obtain a reference quantum state lattice matrix generated based on the ground truth input provided to the trained AI agent, and ground truth output data generated by the trained AI agent based on the ground truth input data provided to the trained AI agent. Here, the generated ground truth output data includes a plurality of outcomes inferred by the trained AI agent and the probability amplitudes associated with the plurality of outcomes; (iii) Generate an output quantum state based on the acquired output data. Here, the generated ground truth output data includes multiple outcomes inferred by the trained AI agent and the probability amplitudes associated with the multiple outcomes; (iv) Generate a quantum state for each row of the lattice matrix of the reference quantum state; and (v) The processor is instructed to perform quantum-based anomaly classification on the generated output quantum state and the lattice matrix of the reference quantum state, based on the data obtained from the generated quantum state.
11. The method according to claim 10, wherein the generation of the lattice matrix of the reference quantum state includes the following steps: (i) Record multiple outcomes inferred by the trained AI agent as the ground state of the lattice matrix of the reference quantum state, and (ii) The probability amplitude associated with each of the multiple outcomes is recorded as an element of a lattice matrix of a reference quantum state, with each row of the matrix associated with ground truth input data.
12. A method for performing quantum-based anomaly classification on acquired output data, comprising the following steps, according to any one of claims 10 or 11: (i) A step of identifying the ground state with the highest probability amplitude from the output quantum states, (ii) A step of performing quantum amplitude amplification on the identified ground state, (iii) A step of performing quantum amplitude amplification on a ground state similar to the identified ground state, (iv) A step of flagging ground states that are similar to the quantum-amplified ground states among the quantum-amplified ground states of each row of the lattice matrix of the reference quantum state, and (v) The step of determining that the acquired input data is an adversarial attack if the ground truth input data associated with each of the flagged quantum-amplified ground states is dissimilar to the acquired input data.
13. The method according to any one of claims 10 or 11, wherein the step of performing quantum-based anomaly classification of the acquired output data includes the following steps: (i) A step of identifying the ground state with the highest probability amplitude from the output quantum states, (ii) A step of performing multi-target quantum amplitude amplification on the identified ground state based on the output quantum state, (iii) For each row of the lattice matrix of the reference quantum state, a step of performing multi-target quantum amplitude amplification on a ground state similar to the identified ground state, (iv) A step of flagging the quantum-amplified ground states in each row of the lattice matrix of the reference quantum state that are similar to a specific quantum-amplified ground state, and (v) The step of determining that the acquired input data is a hostile attack if the ground truth input data associated with each of the flagged quantum-amplified ground states is dissimilar to the acquired input data.
14. The method according to claim 10 or 11, wherein the lattice matrix of the reference quantum state includes a two-dimensional matrix.
15. The method according to claim 10 or 11, wherein the lattice matrix of the reference quantum state includes a three-dimensional (3D) matrix, and each layer of the 3D matrix represents a specific time frame.
16. The method according to claim 11, wherein the probability amplitude represents the probability that an outcome occurs when ground truth input data is provided to an AI agent.
17. The method according to claim 10 or 11, characterized in that if the acquired input data is determined to be a hostile attack, the processing unit commands mitigation measures.
18. The method according to claim 10 or 11, wherein the lattice matrix of a reference quantum state is generated over a predetermined period of time.
Citation Information
Patent Citations
Quantum Computing with Kernel Methods for Machine Learning
JP2023546590A
Learning system and learning method
JP2024002105A