System and method for detecting hostile attack for artificial intelligence (AI)
A quantum-based defense module for AI systems addresses adversarial attacks by generating lattice matrices and performing anomaly classification, ensuring robustness and reliability against deceptive inputs.
Patent Information
- Application Number
- JP2025048782
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Priority Date
- 2024-06-05
- Filing Date
- 2025-03-24
- Publication Date
- 2025-12-17
- Estimated Expiration
- 2045-03-24
AI Technical Summary
Existing AI systems are vulnerable to adversarial attacks, which manipulate input data to deceive AI models, leading to incorrect predictions or classifications, posing threats to critical industries like healthcare and finance, and current solutions fail to detect these attacks effectively.
A defense module using quantum mechanics principles generates a lattice matrix of quantum states based on ground truth data, performs quantum-based anomaly classification, and employs quantum amplitude amplification to identify and mitigate adversarial attacks by comparing input and output data patterns.
The system enhances AI security by detecting and mitigating adversarial attacks, maintaining system integrity and reliability against evolving cybersecurity threats.
Smart Images

Figure 2025183921000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a system and method for detecting and mitigating adversarial attacks against AI agents using the principles of quantum mechanics. [Background technology]
[0002] An AI agent is an autonomous computer module designed to perform tasks and make decisions based on context and goals. AI agents typically consist of trained machine learning models that allow the AI agent to evaluate and analyze data and learn from experience, allowing it to adapt to changing circumstances in real time. AI agents can sense or perceive their surroundings through various means, such as sensors or input data, and then use the resulting data to predict the future, make decisions, and act.
[0003] The development and deployment of AI agents has the potential to revolutionize industries and improve efficiency, accuracy, and decision-making in a variety of fields. AI agents play an important role in a variety of fields, including medicine, finance, transportation, and science. They are central to the development of intelligent systems and applications, such as virtual assistants, autonomous driving, recommendation systems, and chatbots. As AI technology advances, AI agents are becoming increasingly sophisticated, able to tackle more complex tasks and operate under dynamic and uncertain conditions.
[0004] Traditional AI technology has focused on improving the ability to perform increasingly sophisticated tasks. However, due to a relative lack of emphasis on security, many AI platforms and AI agents are vulnerable to the mismanipulation of confidential information. This vulnerability is caused when machine learning (ML) models are attacked by artificially generated input data intended to deceive AI technology. Malicious third parties generate and input specially designed data, known as adversarial samples, to deceive and manipulate AI models. These adversarial samples are indistinguishable from legitimate data to humans and are designed to cause AI models to make incorrect predictions or classifications. They exploit design vulnerabilities in AI models, such as small perturbations in the input data. This vulnerability is caused by adversarial data inputs that affect the predictive capabilities of a trained ML model. Attackers manipulate input data to produce incorrect output data. The input data is virtually indistinguishable from legitimate data. Attackers manipulate AI models by adding small distortions.
[0005] Adversarial attacks do not necessarily only alter existing training data, but can also embed outliers during inference or even into models that have not been trained at all. Such attacks can also trick models into outputting inaccurate data by shifting classification boundaries during runtime.
[0006] Adversarial attacks also pose a threat to critical industries, including healthcare, finance, infrastructure security, and communications. Examples include tampering with MRI cancer test results, disrupting algorithmic trading platforms, subverting smart grid control systems, and manipulating natural language processing engines. These attacks are classified by their purpose, method, and the stage of the AI model's lifecycle they target. One type of adversarial attack is an evasion attack. In an evasion attack, an attacker manipulates input data to deceive an AI model during inference, leading to inaccurate predictions or classifications. For example, an attacker can add inaudible noise to an image to trick a facial recognition system into misidentifying a person. Another type of attack is an exploratory attack. In an exploratory attack, an attacker investigates an AI model to gain information about its behavior and structure. Another attack is model theft, which aims to reconstruct sensitive information from the model's output data. Another type of attack is a backdoor attack, in which an attacker embeds a hidden backdoor in a model during training and activates it with specific inputs, generating malicious output data.
[0007] Potential attacks often occur on end-user systems where the trained model resides. Vulnerabilities often emerge during the training phase, where outliers are input during inference, causing the ML model to break. End users do not have direct access to the original training data unless a breach occurs during the software's development phase. Potential attacks also exist in self-learning ML, which is continuously updated and trained, but even in this case, the input data is not directly accessible.
[0008] Despite these efforts, ML models are unable to detect adversarial attacks, and addressing security concerns in AI systems remains a challenge. Adversarial attacks are inherently elusive and cannot be detected by existing solutions. Summary of the Invention
[0009] In one embodiment of the present invention, a defense module for detecting adversarial attacks against a trained AI agent is disclosed, wherein the defense module is communicatively coupled to the AI agent and includes a processor and a non-transitory storage medium readable by the processor, the storage medium including instructions for, when executed by the processor, (i) capturing and storing input data provided to the trained AI agent and output data generated by the trained AI agent based on the input data provided to the trained AI agent, and (ii) obtaining a lattice matrix of a reference quantum state generated based on ground truth input data provided to the trained AI agent and ground truth output data generated by the trained AI agent based on the ground truth input data provided to the trained AI agent, wherein the generated ground truth output data includes multiple outcomes inferred by the AI agent and probability amplitudes associated with the multiple outcomes. Further, the method instructs the processing device to (iii) generate output quantum states based on the acquired output data, (iv) generate quantum states for each row of a lattice matrix of a reference quantum state, and (v) perform quantum-based anomaly classification of the acquired data based on the generated output quantum states and the generated quantum states relative to the lattice matrix of the reference quantum state.
[0010] In another embodiment of the invention, generating the lattice matrix of the reference quantum state includes instructions for a processing unit to: (i) record the multiple outcomes inferred by the trained AI agent as basis states of the lattice matrix of the reference quantum state, and (ii) record the probability amplitude of each of the multiple outcomes as an element of the lattice matrix of the reference quantum state, where each row of the matrix is associated with ground truth input data.
[0011] In another embodiment of the present invention, instructions for performing quantum-based anomaly classification of acquired output data include the following steps: (i) identifying a basis state from the output quantum states that has the highest probability amplitude; (ii) performing quantum amplitude amplification on the identified basis state; (iii) performing quantum amplitude amplification on basis states similar to the identified basis state; (iv) flagging basis states in each row of a lattice matrix of a reference quantum state that are similar to the quantum-amplified basis state; and (v) determining that the acquired input data is an adversarial attack if the ground truth input data associated with each of the flagged quantum-amplified basis states is dissimilar to the acquired input data.
[0012] In another aspect of the present invention, a method for detecting adversarial attacks against a trained AI agent using a defense module communicatively coupled to the AI agent is disclosed. The method for detecting attacks includes the following steps: (i) capturing and storing input data provided to the trained AI agent and output data generated by the trained AI agent based on the input data provided to the trained AI agent; (ii) looking up a lattice matrix of reference quantum states generated based on ground truth input data provided to the trained AI agent; and (iii) looking up ground truth output data generated by the trained AI agent for each of the ground truth input data provided to the trained AI agent, wherein each of the generated ground truth output data associates a plurality of outcomes inferred by the trained AI agent and a probability amplitude with each of the plurality of outcomes. The method further includes (i) generating output quantum states based on the acquired output data, (ii) generating a quantum state for each row of the lattice matrix of the reference quantum states, and (iii) performing quantum-based anomaly classification of the acquired data based on the generated output quantum states and the generated quantum states for each row of the lattice matrix of the reference quantum states.
[0013] In another embodiment of the present invention, generating the reference quantum state lattice matrix comprises recording the outcomes inferred by the trained AI agent as basis states of a reference quantum state lattice matrix, and storing the probability amplitudes associated with each of the outcomes as elements of the reference quantum state lattice matrix, where each row of the matrix is associated with ground truth input data.
[0014] In another embodiment of the present invention, performing quantum-based anomaly classification on the acquired output data includes the steps of: (i) identifying a basis state from the output quantum state that has the highest probability amplitude, (ii) performing quantum amplitude amplification on the identified basis state based on the output quantum state, (iii) performing quantum amplitude amplification on each row of a lattice matrix of a reference quantum state, (iv) flagging quantum amplified basis states of each row of the lattice matrix of the reference quantum state that are similar to the identified quantum amplified basis state, and (v) determining that the acquired input data constitutes an adversarial attack if ground truth input data associated with each of the flagged quantum amplified basis states is dissimilar to the acquired input data. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a block diagram illustrating a system for detecting adversarial attacks against an AI agent of the present invention. [Figure 2] 1 is a lattice matrix of an exemplary two-dimensional quantum state of the present invention. [Figure 3] 1 is a lattice matrix of an exemplary three-dimensional quantum state of the present invention. [Figure 4] FIG. 1 is a block diagram illustrating a processing system of the present invention. [Figure 5]1 is a flowchart illustrating a process for detecting adversarial artificial intelligence (AI) attacks against a trained AI agent using communicatively connected defense modules. [Figure 6] 10 is a flow chart illustrating a process for performing anomaly classification of acquired output data. DETAILED DESCRIPTION OF THE INVENTION
[0016] Hereinafter, embodiments of the present invention will be described using specific examples. Those skilled in the art will easily understand other advantages and effects of the present invention from the contents of this specification. The present invention can be implemented and applied in other different embodiments, and the contents described in this specification can be modified and changed in various ways based on different perspectives and applications without departing from the spirit of the present invention. Such modifications and changes are within the scope of the claims of the present invention.
[0017] As used herein, the articles "a," "an," and "the" refer to one or more of a feature or element.
[0018] As used herein, the term "about" or "approximately" includes exact values as well as a reasonable margin of error commonly understood in the relevant technical field, for example, within 10% of the numerical value.
[0019] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0020] As used herein, the term "comprises" means that the elements following the word "comprises" are included, but are not limited to them. The elements listed after the word "comprises" are required, but other elements are optional and may or may not be present.
[0021] As used herein, the term "consisting of" means including and limited to the constituent elements following the word "consisting of." The elements listed after the word "consisting of" are required, and no other elements are present.
[0022] In this specification, terms such as "first" and "second" are used to distinguish between similar objects in the specification, claims, and drawings, and are not necessarily used to indicate a particular order or sequence.
[0023] As used herein, the term "AI agent" refers to a computer module designed to perform tasks and inferences or make decisions autonomously depending on a situation and a goal. The computer module typically consists of a machine learning model that is constructed and tuned using a set of training data to recognize patterns, make decisions, or predict outcomes. The trained machine learning model can be used by the AI agent to make inferences or predictions on unseen data.
[0024] As used herein, the term "quantum state lattice matrix" refers to a multidimensional array, where each array value may have a complex number that encodes both the magnitude and phase of the probability amplitude of a qubit. Specifically, each element of the matrix represents the probability amplitude of the system, which is a combination of basis states.
[0025] As used herein, the term "quantum basis state" refers to a set of fundamental states used to represent the state of a more complex quantum system. These basis states provide the basis for representing all possible states of a quantum system as a linear combination. In quantum computers, basis states are represented using qubits. For example, the basis state of a one-qubit system is represented as |0> or |1>, while a two-qubit system is represented as |00>, |01>, |10>, and |11>. Each basis state represents all possible combinations of states in the system. Each basis state corresponds to a different configuration of the qubits, and any quantum state of the system can be represented as a superposition of basis states with a complex number called the probability amplitude.
[0026] Those skilled in the art will understand that functional units are expressed as modules, sub-modules, or processing elements throughout this specification. Those skilled in the art will also understand that a module, sub-module, or set of processing elements may include a circuit, a logic chip, or any type of discrete component. Those skilled in the art will also understand that a module, sub-module, or set of processing elements may be implemented by software and various processor architectures. In embodiments of the present invention, a module, sub-module, or set of processing elements may include computer instructions, calculations, or executable code such that a computer processor executes an event based on instructions received by the computer. The module, sub-module, or set of processing elements may be arbitrarily selected by those skilled in the art and does not limit the scope of the claims.
[0027] Quantum computers use the principles of quantum mechanics to perform calculations. Unlike classical computers, which use bits that can only be in one of two states, "0" or "1," quantum computers use qubits, which can superimpose these two states to perform multiple calculations simultaneously.
[0028] Unlike classical bits, qubits can exist not only in the discrete states of "0" and "1", but also in a linear superposition of states. In mathematical terms, the state of a qubit is represented by a state vector in a two-dimensional Hilbert space. Using Dirac notation, the qubit state vector (ψ) is given by: |ψ〉=α |0〉+β |1〉 ...Equation (1) where α and β are complex numbers, with |α|^2+|β|^2=1. When a qubit ψ is measured, |0〉 is obtained with probability |α|^2 and |1〉 is obtained with probability |β|^2. It should be noted that the process of quantum measurement is non-deterministic, and the act of measurement irreversibly changes the quantum state. In other words, before a qubit is measured, it exists in a superposition of states |0〉 and |1〉, but once it is measured, the outcome is a classical state, not a quantum state. Therefore, the measurement outcome will be either |0〉 or |1〉, never a superposition of the two. This is because during the measurement of a qubit, its quantum state collapses into a classical state, and all subsequent measurements will deterministically result in an outcome with probability equal to 1.
[0029] The quantum computing processes and / or steps used herein can be implemented using the Qiskit software development kit (SDK) developed by IBM. The Qiskit SDK is a comprehensive software development kit for quantum computing that assists users in building quantum algorithms, quantum circuits, and quantum applications. Built on the Python programming language, Qiskit provides a user-friendly interface for creating, manipulating, and simulating quantum circuits, as well as interfacing with quantum hardware. The SDK consists of a wealth of tools, libraries, and resources that facilitate every stage of quantum programming, from designing quantum circuits to running them on a quantum processor.
[0030] While quantum computer hardware is still in its infancy and not widely available, the Qiskit SDK primarily runs on conventional computing systems, including laptops, desktops, servers, and cloud-based platforms. Qiskit implements a powerful simulator that allows users to simulate the behavior of quantum circuits and algorithms on classical computers, enabling them to test and debug quantum programs without access to quantum hardware. Additionally, Qiskit provides visualization tools and libraries that run seamlessly on classical computing devices, allowing users to visualize and analyze quantum circuits, state vectors, and measurement results.
[0031] In an embodiment of the present invention, a defense module can be deployed in an AI agent to enhance the security and resilience of the AI system against adversarial attacks. The defense module primarily detects, mitigates, and prevents malicious activity targeting the AI agent. This module continuously monitors the AI agent's input and output data, analyzes patterns and behaviors for signs of adversarial manipulation, and takes appropriate action to protect the integrity of the system. Equipped with a defense module, the AI agent can maintain its performance and reliability in the face of evolving cybersecurity threats.
[0032] 1 is a block diagram illustrating a system for detecting adversarial attacks against an AI agent according to the present invention. The system 100 includes an AI agent 102 and a defense module 106 deployed on the AI agent 102. The defense module 106 is configured to monitor data provided to and generated by the AI agent 102. In an embodiment of the present invention, the AI agent 102 may be configured to receive data 104 via a network 105.
[0033] In embodiments of the present invention, data 104 may include data obtained by various sensors, such as changes in temperature, motion, or light intensity; data obtained through user interactions, such as commands, questions, or responses given by voice or text input; market data, such as prices, trading volumes, or economic indicators; communication traffic data, such as network traffic logs, access attempts, or anomaly detections; health-related indicators, such as patient vital signs, medical checkup results, or changes in health status; environmental data, such as weather conditions, soil moisture, or crop growth; and social media feed data, such as posts, comments, likes, or shares. Those skilled in the art will recognize that data 104 is not limited to these specific examples and may include various types of data used by a trained machine learning model to make the desired predictions, classifications, and / or inferences.
[0034] The network 105 may include one or more computer communication networks, such as the Internet, a wired network such as a local area network (LAN) or wide area network (WAN), a wireless network such as a wireless LAN (WLAN) or mobile network, or any other similar network. Network adapter cards or network interface modules may also be installed in computing / processing devices within the system 100 to facilitate communication between the respective modules and / or components.
[0035] In embodiments of the present invention, the defense module 106 may include a lattice matrix sub-module 108 configured to generate and store lattice matrices of quantum states, record and classify basis states associated with output probabilities for outcomes predicted, classified, and / or inferred by the AI agent 102, and record and store probability amplitudes for outcomes predicted, classified, and / or inferred by the AI agent 102. The defense module 106 also includes an anomaly classification sub-module 110 configured to perform anomaly classification of events occurring in the AI agent 102, and a sub-module 112 that performs mitigation measures to protect the AI agent 102 from detected adversarial attacks.
[0036] The defense module 106 is deployed on the AI agent 102, where the module 106 may be communicatively coupled to input and / or output ports of the AI agent 102 such that any data provided to and / or generated by the AI agent 102 via the network 105 is received by the defense module 106.
[0037] The defense module 106 begins interacting with the AI agent 102 during an initialization or setup phase. At this stage, the AI agent 102 may have trained machine learning models to perform specific types and / or various classification, prediction, and / or inference tasks and may be deployed to perform a desired function. For example, the AI agent 102 may have a machine learning model trained to identify and classify items in a digital image based on features within the image. In other words, the trained machine learning model may use features of the digital image, such as the color and arrangement of each pixel, to calculate the probability that an item will be classified as a particular item (e.g., a chair).
[0038] During the initialization phase, the AI agent 102 is deployed to a “safe” operating environment, and the input data provided to the AI agent 102 is trusted, i.e., includes ground truth input data. This ensures that the outputs generated by the AI agent 102 during the initialization phase have ground truth output data that correspond to the input data provided to the AI agent 102 and are not the result of compromised data and / or adversarial attacks. As the AI agent 102 performs its classification, prediction, and / or inference processes, the lattice matrix submodule 108 records and stores in a database the input data provided to the AI agent 102, the outcomes generated by the AI agent 102, and the probability amplitudes associated with each outcome generated by the AI agent 102. In an embodiment of the present invention, the initialization phase may record all output data generated over a predefined period, e.g., a week or a month, or may consist of a predefined set of input data. Typically, this stage is usually performed over a long period of time or on a sufficiently large input data set to ensure that all possible output data generated by the AI agent is obtained by the lattice matrix submodule 108. The period required for initialization can be chosen arbitrarily by one skilled in the art.
[0039] During the initialization phase, the AI agent 102 determines, based on the received digital image, that there is a 60% probability that the image contains a firearm, a 20% probability that it contains an edged weapon, a 15% probability that it contains a physical attack item, and a 5% probability that it does not contain any items. The lattice matrix sub-module 108 then stores the features of the digital image (as input data) in a database along with the corresponding outcomes generated by the AI agent 102, along with the probabilities associated with each outcome.
[0040] After the initialization phase is complete, the lattice matrix submodule 108 generates a quantum state lattice matrix based on the information stored in the database. In an embodiment of the present invention, each column of the matrix can be defined to represent a basis state, which represents an outcome inferred by the AI agent 102. Furthermore, for each row of the matrix, each element of the row can represent the probability of the corresponding outcome occurring when a specific input is provided to the AI agent 102. An example of such a quantum state lattice matrix is shown in FIG. 2. Matrix 200 was generated based on the assumption that the AI agent 102 generates four possible outcomes for seven input data, i.e., In1 through In7. Each of the four outcomes is represented by one of the basis states, i.e., |00〉 through |11〉, and each basis state is used as a specific label for each column. Furthermore, each cell of matrix 200, e.g., p1,1 through p7,4, represents the probability of the corresponding outcome occurring. For example, p_4,3 indicates the probability of the outcome associated with |10〉 occurring when the AI target 102 is provided with input In4.
[0041] In the example shown in Figure 2, four basis states were required, so a two-qubit system was used. It should be noted that as the number of basis states increases, the number of qubits used increases (in an n-qubit system, there are 2n basis states). For example, if 16 basis states were required, a four-qubit system would be used.
[0042] In an embodiment of the present invention, the lattice matrix submodule 108 can be configured to generate a three-dimensional (3D) quantum state lattice matrix, where each layer represents a lattice matrix of quantum states for a particular period of time. A three-dimensional quantum state lattice matrix is shown in FIG. 3 , where each layer of matrix 300 represents a lattice matrix of quantum states for a particular period of time. In this embodiment, layer 302 or matrix 302 represents a quantum state lattice matrix for a first period of time, whereby all input data, i.e., rows of the matrix, and the occurrence probabilities of outcomes generated by the AI agent 102 during this first period are plotted in matrix 302. Layer 304 represents a quantum state lattice matrix for a second period of time, whereby all input data, i.e., rows of the matrix, and the occurrence probabilities of outcomes generated by the AI agent 102 are plotted in matrix 302. Layer 306 represents the quantum state lattice matrix for the mth time period, where all input data, i.e., rows of the matrix, and the corresponding probability of occurrence of outcomes generated by AI agent 102 during this mth time period are plotted in matrix 306. The exact number of layers of the three-dimensional quantum state lattice matrix can be arbitrarily selected by one skilled in the art. Furthermore, each time period may include a day, a week, a month, or any other period as desired.
[0043] In another embodiment of the present invention, instead of waiting for the end of the initialization phase, the lattice matrix sub-module 108 generates the cells of the quantum state lattice matrix during the initialization phase, thereby allowing the data to be entered into the cells simultaneously as the data is received. In other words, the lattice matrix sub-module 108 receives input data and records and stores the outcomes generated by the AI agent 102 (based on the input data), along with the probability amplitude associated with each generated outcome, allowing the sub-module 108 to enter the cells of the quantum state lattice matrix simultaneously with the data received by the sub-module 108.
[0044] Once the lattice matrix sub-module 108 has completed generating the lattice matrix of the two-dimensional or three-dimensional quantum state, the generated matrix is used by the defense module 106 as a reference matrix for monitoring the abnormal behavior of the AI agent 102.
[0045] 1, in normal operation, when new data is provided to the AI agent 102, the AI agent 102 generates a set of expected outcomes based on the received data. The data provided to the AI agent 102 and the outcomes generated by the AI agent 102 are then captured by the defense module 106. The anomaly classification submodule 110 then compares the captured information with information from previously generated baseline matrices to determine whether the AI agent 102 has been compromised by an adversarial attack. It is noted that in this process, all expected outcomes generated by the AI agent 102 are similar to the ground states defined in the previously generated baseline matrices.
[0046] In an embodiment of the present invention, the outcome generated by the AI agent 102 may include multiple outcomes, each associated with a probability of that outcome occurring. Next, the anomaly classification submodule 110 begins generating quantum states |ψ_out〉 (defined during the generation of the basis matrices) based on the basis states associated with each outcome (a superposition of multiple outcomes) and their corresponding occurrence probabilities. Next, the anomaly classification submodule 110 creates quantum states for all basis states and their corresponding occurrence probabilities for each matrix of the reference matrix. Here, a combination of a probability amplitude and its corresponding basis state is defined as a quantum state term or element (the product of the probability amplitude and the corresponding basis state).
[0047] In a first embodiment of the present invention, the anomaly classification submodule 110 analyzes the quantum states |ψ_out〉 associated with outcomes generated by the AI agent 102 with the goal of identifying the basis state among the quantum states that has the highest probability amplitude. It then scans the quantum states associated with the reference matrix and flags similar basis states.
[0048] The quantum state terms associated with the flagged ground states and the quantum state terms initially associated with a particular state (|ψ_out 〉) are subjected to quantum amplitude amplification. Quantum amplitude amplification is a process that significantly increases the amplitude of selected quantum state terms, making them more prominent within the quantum state spectrum. In other words, quantum amplification is performed to make the amplitudes of these quantum state terms more distinguishable and comparable to the overall noise level. An amplified quantum state is defined as a quantum state in which the probability amplitude associated with the quantum state ground state has been intentionally increased through a quantum amplitude amplification process. For simplicity, when referring to a quantum amplified first ground state, we mean that the probability amplitude of the first ground state has been amplified through a quantum amplitude amplification process, resulting in the creation of the corresponding amplified quantum state.
[0049] Following amplification, submodule 110 compares the quantum amplified flagged basis states obtained from the reference matrix with the quantum amplified specific state |ψ_out〉 of the quantum state. That is, among the quantum amplified flagged basis states of each row of the lattice matrix of the reference quantum state, the quantum amplified flagged basis state that is most similar to the quantum amplified specific state is identified. In an embodiment of the present invention, the similarity is determined by the fidelity between the compared amplified basis states. Here, a fidelity close to "1" (e.g., greater than 0.95) is considered sufficiently similar. The fidelity F of two quantum states |ψ〉 and |φ〉 is defined as F = (|ψ〉, |φ〉) = |〈ψ|φ〉|^2. The corresponding row of quantum amplified flagged basis states from the reference matrix is further flagged for further analysis.
[0050] During the analysis phase, submodule 110 retrieves input data associated with a particular row and compares it with new data provided to AI agent 102. In other words, during this phase, submodule 110 determines whether the retrieved data associated with a particular row represents an adversarial attack if the retrieved input data is dissimilar to the new data provided to AI agent 102. The degree of dissimilarity between data sets can be assessed using various statistical and computational techniques. Euclidean distance is used to directly compare the input data and new data associated with each flagged row. Additionally, Pearson and Spearman rank correlation coefficients provide insight into linear or monotonic relationships, respectively. Techniques such as cosine similarity and Jaccard coefficient are particularly useful for comparing text or sets, while Hamming distance is used to compare data strings of equal length. For more complex or structured data, visual tools such as dendrograms and machine learning models including clustering and dimensionality reduction can be used to reveal underlying patterns and groupings and highlight dissimilarities that are not readily apparent through direct statistical techniques. The method used will depend on the machine learning model used by the AI agent 102 and can be chosen at the discretion of one skilled in the art.
[0051] In summary, these comparisons aim to detect significant discrepancies between past and new data that could potentially attack the AI agent 102. If such discrepancies are found, submodule 110 invokes attack mitigation submodule 112, which takes the necessary action to mitigate the threat, thereby ensuring the integrity and security of the AI agent 102.
[0052] In another embodiment of the present invention, the anomaly classification submodule 110 can analyze the quantum states |ψ_out〉 associated with outcomes produced by the AI agent 102 and identify the basis state with the highest probability amplitude among the quantum states |ψ_out〉. Following this identification, the quantum states associated with the reference matrix are scanned to flag basis states that are similar to the quantum state of the particular state.
[0053] The quantum state terms associated with the flagged basis states are then subjected to a multi-target quantum amplitude amplification process, along with the quantum state terms initially associated with the particular state (|ψ_out〉), to distinguish the amplitudes of the quantum state terms from noise. Submodule 110 then compares the quantum amplification flagged basis states obtained from the reference matrix with the quantum amplification specific states obtained from the quantum state |ψ_out〉, and identifies the quantum amplification flagged basis state from the reference matrix that is most similar to the quantum amplification specific state obtained from the quantum state |ψ_out〉. The corresponding row of the quantum amplification state obtained from the reference matrix is flagged for further analysis, as described above.
[0054] The first embodiment of the present invention will be illustrated using the following simplified example: When new data Inew is provided to the AI agent 102, the AI agent 102 is assumed to generate four outcomes, Out1 through Out4. These outcomes and their probabilities of occurrence are associated with corresponding ground states of a reference matrix and can be expressed as follows: Out1 is associated with |00〉 and occurs with probability p_a; Out2 is associated with |01〉 and occurs with probability p_b; Out3 is associated with |10〉 and occurs with probability p_c; and Out4 is associated with |11〉 and occurs with probability p_d.
[0055] Subsequently, a quantum state |ψ_out 〉 is generated that represents a superposition of these four basis states associated with the outcomes (Out1 to Out4): |ψ_out 〉=p_a |00〉 + p_b |01〉 + p_c |10〉 + p_d |11〉 Detailed steps for generating quantum states are well known to those skilled in the art and will not be described here.
[0056] Next, the anomaly classification submodule 110 generates a superposition of all basis states and their occurrence probabilities for each row of the reference matrix. Assuming that matrix 200 is used as the reference matrix, for each row of matrix 200 (see FIG. 2), a quantum state is generated that represents a superposition of all occurrence probabilities of the basis states in that row. As a result, for a reference matrix based on matrix 200, seven quantum states are generated (because matrix 200 has seven rows). The seven quantum states are defined as follows: |ψ_In1 〉=p_1,1 |00〉+p_1,2 |01〉+p_1,3 |10〉+p_1,4 |11〉 |ψ_In2 〉=p_2,1 |00〉+p_2,2 |01〉+p_2,3 |10〉+p_2,4 |11〉 |ψ_In3 〉=p_3,1 |00〉+p_3,2 |01〉+p_3,3 |10〉+p_3,4 |11〉 |ψ_In4 〉=p_4,1 |00〉+p_4,2 |01〉+p_4,3 |10〉+p_4,4 |11〉 |ψ_In5 〉=p_5,1 |00〉+p_5,2 |01〉+p_5,3 |10〉+p_5,4 |11〉 |ψ_In6 〉=p_6,1 |00〉+p_6,2 |01〉+p_6,3 |10〉+p_6,4 |11〉 |ψ_In7 〉=p_7,1 |00〉+p_7,2 |01〉+p_7,3 |10〉+p_7,4 |11〉
[0057] Submodule 110 then analyzes the quantum state |ψ_out〉 to identify the ground state with the highest probability amplitude (assumed in this example to be pb corresponding to the ground state |01〉), then scans the seven quantum states to identify a similar ground state from these quantum states, i.e., the ground state |01〉, and then flags the identified ground state.
[0058] The quantum state term associated with this flagged ground state, i.e., ground state |01〉, undergoes quantum amplitude amplification along with the quantum state term associated with the first identified state, |ψ_out〉. The quantum amplified flagged ground state is then compared to the quantum amplified identified state, quantum state |ψ_out〉, to identify the quantum amplified flagged ground state that best matches the quantum amplified identified state, quantum state |ψ_out〉.
[0059] Provided that quantum amplified states p_5,2|01〉, p_6,2|01〉, p_7,2|01〉, and p_2,2|01〉 are found to be closest (e.g., delta<0.05) to quantum amplified state p_b|01〉, the input data corresponding to these amplified quantum states p_5,2|01〉, p_6,2|01〉, p_7,2|01〉, and p_2,2|01〉 are retrieved and compared with the new data Inew to determine whether there is a significant discrepancy between the past input data and the new data. If a discrepancy is found, submodule 110 invokes attack mitigation submodule 112, which takes the necessary action to mitigate the threat and ensure the integrity and safety of AI agent 102.
[0060] In a second embodiment of the present invention, the protection module 106 may use a three-dimensional (3D) reference matrix with two or more layers instead of a two-dimensional reference matrix, where each layer of the 3D reference matrix represents a different time frame, allowing the system to have a more real-time and historical perspective.
[0061] Similar to the first embodiment of the present invention, the anomaly classification submodule 110 analyzes the quantum states |ψ_out〉 generated by the AI agent 102 to identify the basis state with the highest probability amplitude among the quantum states. It then scans the quantum states associated with various layers of the 3D reference matrix to flag basis states that are similar to the identified state. This process of scanning the quantum states associated with layers of the 3D reference matrix compares the most recent state generated by the AI agent 102 with the past and present data patterns represented by the 3D reference matrix.
[0062] The quantum state terms associated with the flagged basis states undergo quantum amplitude amplification, along with the quantum state |ψ_out〉 terms associated with the initially identified state. Submodule 110 then compares the quantum-amplified flagged basis states obtained from the three-dimensional (3D) reference matrix with the quantum-amplified specific quantum state |ψ_out〉, and the corresponding amplified quantum state rows are flagged from the three-dimensional (3D) reference matrix layer for further analysis. During the analysis phase, submodule 110 retrieves the input data associated with these flagged rows and compares them with new data provided to AI agent 102. If significant discrepancies are found between the past data and the new data, submodule 110 invokes attack mitigation submodule 112, which initiates the necessary actions to mitigate the threat, thereby ensuring the integrity and security of AI agent 102. In another embodiment, sub-module 110 can be configured to invoke attack mitigation sub-module 112 only if it is determined that a mismatch between new input data and past input data occurs across a significant number of layers of three-dimensional (3D) reference matrices.
[0063] By expanding the reference matrix to include multiple layers represented in different time frames, the protection module 106 gains a more robust ability to monitor, detect, and react to anomalies based on current and historical data patterns. This approach enhances the predictive capabilities and security of the protection module 106 by taking advantage of the temporal dynamics provided by the 3D reference matrix.
[0064] A representative block diagram of components of a processing system 400 according to an embodiment of the present invention is shown in Figure 4. Processing systems are located within the protection module 106 and various sub-modules contained therein to perform digital signal processing functions or operations, or other modules or sub-modules of the system. Those skilled in the art will appreciate that the configuration of each processing system within these modules or sub-modules may vary, and the arrangement shown in Figure 4 is provided as an example only.
[0065] In an embodiment of the present invention, the processing system 400 may include a controller 401 and a user interface 402. The user interface 402 is configured to allow manual interaction between a user and the computing modules as needed and includes components necessary for a user to input update instructions to each module. Those skilled in the art will appreciate that the components of the user interface 402 vary depending on the embodiment, but typically include one or more of a display 440, a keyboard 435, and an optical device 436.
[0066] The controller 401 is in data communication with the user interface 402 via a bus 415 and includes an on-board memory 420 for processing instructions and data to execute the methods of the present embodiment, a processing unit, processing element, or processor 405, an operating system 406, an input / output (I / O) interface 430 for communicating with the user interface 402, and a communications interface, in the form of a network card 450 in this embodiment. The network card 450 may be used to transmit data from these modules to other processing units, for example, via a wired or wireless network, or to receive data via a wired or wireless network. Wireless networks that may be utilized by the network card 450 include, but are not limited to, wireless fidelity (WiFi), Bluetooth, near field communication (NFC), cellular networks, satellite networks, telecommunications networks, wide area networks (WANs), etc.
[0067] The memory 420 and operating system 406 are in data communication with the processor 405 via the bus 410. The memory components include volatile and nonvolatile memory, and further include one or more memories selected from random access memory (RAM) 423, read-only memory (ROM) 425, and mass storage device 445 (comprised of one or more solid-state drives (SSDs)). Those skilled in the art will appreciate that these memory components include non-transitory computer-readable media and may include all computer-readable media except for transient, propagating signals. Typically, instructions are stored in the memory components as program code, but may also be hardwired. Here, the memory 420 may include kernel modules and / or programming modules, such as software applications, which may be stored in either volatile or nonvolatile memory.
[0068] The term "processor" is used generally to refer to any device or component capable of processing instructions and may include a microprocessor, processing unit, multiple element processor, microcontroller, programmable logic device, or any other type of computing device. Processor 405 provides any logic circuitry for receiving input data, processing it according to instructions stored in memory, and generating output data (e.g., on a memory component or display 440). In this embodiment, processor 405 may be a single-core or multi-core processor with a memory address space. For example, processor 405 may be multi-core, including an eight-core CPU. In another embodiment, processor 405 may be a cluster of CPU cores operating in parallel to accelerate computations.
[0069] A process for detecting adversarial attacks against a trained AI agent is shown in Figure 5. In an embodiment of the invention, process 500 is performed by a defense module communicatively connected to the AI agent.
[0070] Process 500 begins in step 502 by obtaining and storing input data provided to the trained AI agent. Simultaneously, process 500 also obtains and stores output data generated by the trained AI agent. Process 500 then proceeds to step 504, where a lattice matrix of previously generated quantum states is obtained from a database and / or memory included within the defense module. In embodiments of the present invention, this database may be provided on a remote server or cloud server, and may be provided to the defense module via wireless or wired communication means. Process 400 then generates an output quantum state based on the obtained output data in step 506.
[0071] Process 500 generates a quantum state for each row of a lattice matrix of reference quantum states in step 508. Process 500 then performs quantum-based anomaly classification of the acquired data based on the generated output quantum states and the generated quantum states for each row of the lattice matrix of reference quantum states, which is performed in step 510.
[0072] In another embodiment of the present invention, once the acquired input data is classified as an adversarial attack, process 500 may proceed to step 512. In step 512, process 500 performs mitigation measures against the adversarial attack on the trained AI agent.
[0073] A process for performing anomaly classification of the obtained output data is shown in Figure 6. Process 600 may be performed by a protection module.
[0074] Process 600 begins in step 602 by identifying a basis state with the highest probability amplitude from the output quantum state. The process then proceeds to step 604, where quantum amplitude amplification is performed on the identified basis state based on the output quantum state. In step 606, quantum amplitude amplification is performed on basis states similar to the identified basis state based on the quantum states generated in each row of the lattice matrix of the reference quantum state. Next, in step 608, quantum amplified basis states in each row of the lattice matrix of the reference quantum state that are similar to the specific quantum amplified basis state are flagged. In step 610, process 600 classifies the acquired input data as an adversarial attack if the ground truth input data associated with each of the flagged quantum amplified basis states is dissimilar to the acquired input data to some degree.
[0075] In another embodiment of the present invention, the lattice matrix of the reference quantum state may be generated by a process of recording multiple outcomes inferred by the trained AI agent as basis states of the lattice matrix of the reference quantum state, and recording the probability amplitude associated with each of the multiple outcomes as an element of the lattice matrix of the reference quantum state, with each row of the matrix associated with ground truth input data.
[0076] In another embodiment of the present invention, a process for performing anomaly classification on acquired output data includes the following steps: (i) generating an output quantum state based on the acquired output data and identifying, from the output quantum state, a plurality of basis states having a probability amplitude equal to or greater than a predetermined threshold, (ii) performing multi-target quantum amplitude amplification on the identified plurality of basis states based on the output quantum state, (iii) generating a quantum state for each row of a lattice matrix of a reference quantum state based on the basis states of the lattice matrix of the reference quantum state and the probability amplitude associated with each basis state of the row, (iv) performing multi-target quantum amplitude amplification on basis states similar to the identified plurality of basis states based on the quantum states generated for each row of the lattice matrix of the reference quantum state, (v) flagging quantum-amplified basis states for each row of the lattice matrix of the reference quantum state that exhibit a degree of similarity to the identified plurality of quantum-amplified basis states, and (vi) determining that the acquired input data represents an adversarial attack if ground truth input data associated with each of the flagged quantum-amplified basis states exhibits a degree of dissimilarity to the acquired input data.
[0077] Many other changes, substitutions, variations, and modifications will be ascertainable to those skilled in the art, and the present invention includes all such changes, substitutions, variations, and modifications as fall within the scope of the appended claims.
Claims
1. 1. A defense module communicatively coupled to a trained AI agent for detecting adversarial attacks against the AI agent, the defense module further comprising: a processing unit; and a non-transitory storage medium readable by a processing device, which, when executed by the processing device, provides instructions for: (i) capturing and storing input data provided to the trained AI agent and output data generated by the trained AI agent based on the input data provided to the trained AI agent; (ii) obtaining a lattice matrix of a reference quantum state generated based on ground truth input data provided to the trained AI agent, and ground truth output data generated by the trained AI agent based on the ground truth input data provided to the trained AI agent, where the generated ground truth output data includes a plurality of outcomes inferred by the AI agent and probability amplitudes associated with the plurality of outcomes; (iii) generating an output quantum state based on the obtained output data; (iv) generating quantum states for each row of the lattice matrix of the reference quantum state; (v) instructing the processing device to perform quantum-based anomaly classification of the obtained data based on the generated output quantum state and the generated quantum state relative to a lattice matrix of a reference quantum state;
2. The defense module of claim 1 , wherein generating a lattice matrix of the reference quantum state includes instructions for a processing unit to perform the following processes: (i) recording the outcomes inferred by the trained AI agent as basis states of a lattice matrix of a reference quantum state; and (ii) The probability amplitude of each of the multiple outcomes is recorded as an element of the lattice matrix of the reference quantum state.
3. 3. The protection module of claim 1, wherein the instructions for performing quantum-based anomaly classification on the obtained output data include the steps of: (i) identifying a basis state from the output quantum states that has the highest probability amplitude; (ii) performing quantum amplitude amplification on the identified basis state; (iii) performing quantum amplitude amplification on the identified basis state; (iv) flagging quantum amplified basis states of each row of the lattice matrix of the reference quantum state that are similar to the quantum amplified basis state; and (v) determining that the acquired input data is an adversarial attack if the ground truth input data associated with each of the flagged quantum amplified basis states is dissimilar to the acquired input data;
4. 3. The protection module of claim 1, wherein the instructions for causing the processing device to perform quantum-based anomaly classification on the acquired output data comprise the steps of: (i) identifying a basis state from the output quantum states that has the highest probability amplitude; (ii) performing multi-target quantum amplitude amplification for the identified basis states based on the output quantum states; (iii) performing multi-target quantum amplitude amplification on a ground state similar to the identified ground state; (iv) flagging quantum amplified basis states of each row of the lattice matrix of the reference quantum state that are similar to the particular quantum amplified basis state; and (v) determining that the acquired input data is an adversarial attack if the ground truth input data associated with each of the flagged quantum amplified basis states is dissimilar to the acquired input data;
5. 5. A protection module according to any one of claims 1 to 4, wherein the lattice matrix of the reference quantum state comprises a two-dimensional matrix.
6. 5. The protection module of claim 1, wherein the lattice matrix of the reference quantum state comprises a three-dimensional (3D) matrix, each layer of the 3D matrix representing a particular time frame.
7. 3. The defensive module of claim 2, wherein the probability amplitude represents the probability of an outcome occurring when ground truth input data is provided to the AI agent.
8. The protection module according to claim 1 , characterized in that the processing unit commands mitigation measures if the acquired input data is determined to be a hostile attack.
9. 8. A protection module according to any one of claims 1 to 7, wherein the lattice matrix of the reference quantum state is generated over a predetermined period of time.
10. 1. A method for detecting adversarial attacks against a trained AI agent using a defense module communicatively connected to the AI agent, the method further comprising the steps of: (i) capturing and storing in a processing device input data provided to the trained AI agent and output data generated by the trained AI agent based on the input data provided to the trained AI agent; (ii) obtaining a lattice matrix of a reference quantum state generated based on ground truth inputs provided to the trained AI agent and ground truth output data generated by the trained AI agent based on the ground truth input data provided to the trained AI agent, where the generated ground truth output data includes a plurality of outcomes inferred by the trained AI agent and probability amplitudes associated with the plurality of outcomes; (iii) generating an output quantum state based on the obtained output data, where the generated ground truth output data includes a plurality of outcomes inferred by the trained AI agent and probability amplitudes associated with the plurality of outcomes; (iv) generating a quantum state for each row of the lattice matrix of the reference quantum state; and (v) instructing the processing device to perform quantum-based anomaly classification on the obtained data based on the generated output quantum state and the generated quantum state relative to a lattice matrix of a reference quantum state;
11. The method of claim 10, wherein generating a lattice matrix of a reference quantum state comprises the steps of: (i) recording the outcomes inferred by the trained AI agent as basis states of a lattice matrix of a reference quantum state; and (ii) The probability amplitudes associated with each of the multiple outcomes are recorded as elements of a lattice matrix of the reference quantum state, where each row of the matrix is associated with ground truth input data.
12. 12. The method of claim 10 or 11, wherein the method for performing quantum-based anomaly classification on the obtained output data comprises the following steps: (i) identifying a basis state from the output quantum states that has the highest probability amplitude; (ii) performing quantum amplitude amplification on the identified basis state; (iii) performing quantum amplitude amplification on a ground state similar to the identified ground state; (iv) flagging quantum amplified basis states of each row of the lattice matrix of the reference quantum state that are similar to the quantum amplified basis state; and (v) determining that the acquired input data is an adversarial attack if the ground truth input data associated with each of the flagged quantum amplified basis states is dissimilar to the acquired input data;
13. 12. The method of claim 10, wherein performing quantum-based anomaly classification of the obtained output data comprises: (i) identifying a basis state from the output quantum states that has the highest probability amplitude; (ii) performing multi-target quantum amplitude amplification for the identified basis states based on the output quantum states; (iii) for each row of the lattice matrix of the reference quantum state, performing multi-target quantum amplitude amplification for a basis state similar to the identified basis state; (iv) flagging quantum amplified basis states of each row of the lattice matrix of the reference quantum state that are similar to the particular quantum amplified basis state; and (v) determining that the acquired input data is an adversarial attack if the ground truth input data associated with each of the flagged quantum amplified basis states is dissimilar to the acquired input data;
14. 14. The method of any one of claims 10 to 13, wherein the lattice matrix of the reference quantum state comprises a two-dimensional matrix.
15. 14. The method of any one of claims 10 to 13, wherein the lattice matrix of the reference quantum state comprises a three-dimensional (3D) matrix, each layer of the 3D matrix representing a particular time frame.
16. 12. The method of claim 11, wherein the probability amplitude represents the probability that the outcome will occur when the ground truth input data is provided to the AI agent.
17. 17. The method according to any one of claims 11 to 16, characterized in that the processing unit commands mitigation measures if the acquired input data is determined to be a hostile attack.
18. 18. The method of any one of claims 11 to 17, wherein the lattice matrix of the reference quantum state is generated over a predetermined period of time.
Citation Information
Patent Citations
Quantum Computing with Kernel Methods for Machine Learning
JP2023546590A
Learning system and learning method
JP2024002105A
Cited By
Inter-agent communication control system
JP7904660B1
Information processing system that automatically encrypts and selectively discloses internally processed data.
JP7912368B1