System and method for detecting adversarial artificial intelligence attacks
By introducing a defense module based on quantum mechanics principles into the AI agent, generating a quantum state lattice matrix and classifying anomalies, the problem of detecting adversarial attacks is solved, achieving effective defense and security protection for the AI agent.
Patent Information
- Application Number
- CN202510277317.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-06-05
- Filing Date
- 2025-03-10
- Publication Date
- 2025-12-05
AI Technical Summary
Existing technologies are unable to effectively detect and defend against adversarial AI attacks, making AI agents vulnerable to being deceived by abnormal inputs during the prediction phase, thus affecting the accuracy and security of their decisions.
A defense module based on quantum mechanics is adopted. By generating a baseline quantum state lattice matrix, the input and output data of the AI agent and its probability amplitude are recorded. Quantum amplitude amplification technology is used to detect abnormal behavior and trigger mitigation measures to protect the AI agent.
It improves the AI agent's ability to detect adversarial attacks, ensuring the reliability and security of its decisions. The defense module against adversarial attacks can effectively identify and mitigate potential threats, maintaining the integrity and reliability of the system.
Smart Images

Figure CN121072802A_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to Singapore Patent Application No. 10202401602R, filed on June 5, 2024, the entire contents of which are incorporated herein by reference for all purposes. Technical Field
[0003] This application relates to a system and method for using quantum mechanical principles to detect and mitigate adversarial AI attacks on artificial intelligence (AI) agents. Background Technology
[0004] Artificial intelligence (AI) agents are autonomous computer modules designed to perform tasks or make decisions based on their environment and objectives. These AI agents typically include trained machine learning models that enable them to interpret and analyze data, learn from experience, and adapt to changing conditions. Through various means, such as sensors or data input, AI agents can perceive or understand their surroundings and then use this information to reason / infer, make decisions, and take actions.
[0005] The development and deployment of AI agents have the potential to revolutionize industries and improve efficiency, accuracy, and decision-making capabilities across various sectors. Therefore, AI agents play a crucial role in numerous fields, including healthcare, finance, transportation, and scientific research. They are central to the development of intelligent systems and applications such as virtual assistants, autonomous vehicles, recommendation systems, and chatbots. As AI technology continues to advance, AI agents are becoming increasingly sophisticated, enabling them to handle more complex tasks and operate in dynamic and uncertain environments.
[0006] To date, the rapid development of AI technology has primarily focused on enhancing its ability to perform increasingly complex tasks. However, a relative lack of emphasis on security measures has left many AI platforms and agents vulnerable to mishandling of sensitive information. This weakness stems from the susceptibility of machine learning (ML) models to synthetically constructed inputs that can deceive AI technology—a phenomenon known as adversarial AI attacks—which in turn disable AI agents. In such attacks, a malicious third party constructs and inputs specially designed data (called adversarial examples) to deceive and manipulate the AI model. These adversarial examples are often indistinguishable from legitimate data to human observers, but are designed to cause the AI model to make incorrect predictions or classifications, thus exploiting vulnerabilities in the AI model's design, such as its sensitivity to small perturbations in the input data. This weakness arises from adversarially constructed inputs that influence the predictive capabilities of a trained ML model. Attackers design these inputs to produce erroneous outputs that are virtually indistinguishable from legitimate data. Specifically, ML models are sensitive to small calibration perturbations in the inputs they are attempting to classify or predict. By adding tiny but carefully chosen distortions, attackers exploit this sensitivity to manipulate model conclusions.
[0007] It's important to note that adversarial attacks don't necessarily involve simply tampering with or altering the original training data. Instead, such attacks can also involve surgically implanting outliers into the unmodified input data fed into a trained model for inference. These attacks often trick the model by shifting the classification boundary during runtime prediction, leading to inaccurate outputs.
[0008] Adversarial AI attacks also pose a serious threat to decision-critical industries such as healthcare, finance, infrastructure security, and communications. For example, attacks could tamper with MRI cancer detections, deliberately disrupt algorithmic trading platforms, compromise smart grid controls, or manipulate natural language recommendation engines. These attacks can be broadly categorized based on their objectives, techniques, and the stage of the AI model's lifecycle targeted. One category of adversarial AI attacks is evasion attacks. In these attacks, attackers aim to deceive the AI model during inference by constructing input data that leads to incorrect predictions or classifications. For example, adding imperceptible noise to an image can fool a facial recognition system, causing misidentification of a person. Another category includes exploratory attacks. In these attacks, attackers probe the AI model to understand its behavior or structure. Another type of attack is model inversion attacks, which aim to reconstruct sensitive information from the model's output. In model extraction attacks, attackers attempt to replicate the target model by observing its inputs and outputs. In backdoor attacks, attackers implant hidden backdoors into the model during training, and these backdoors can be activated by specific trigger inputs to produce malicious outputs.
[0009] From a security perspective, it has been observed that potential attacks can occur on end-user systems hosting pre-configured, trained versions of the model, rather than during the training phase of the ML model. This vulnerability typically arises during the prediction phase, where outliers are injected, causing the ML model to break down. In some cases, end-users may not have direct access to the original training data that shapes the ML model, unless a leak occurs during software development. However, another conceivable scenario involves self-learning ML models that are continuously updated and retrained. Even in this case, outliers in the predicted output become the entry point, rather than direct access to the input training data.
[0010] Therefore, despite efforts by those skilled in the art, addressing security issues in ML models and / or AI systems remains challenging because these ML models cannot detect such adversarial attacks. These challenges stem from the inherent stealth of these attacks, which can manifest through subtle modifications to the original data, thus evading detection by conventional methods. Consequently, such attacks may not have been previously categorized or classified by existing solutions and therefore remain undetectable. Summary of the Invention
[0011] In one aspect, this application provides a defense module for detecting adversarial AI attacks on a trained artificial intelligence (AI) agent, the trained AI agent being communicatively coupled to the defense module. The application provides the following: the defense module includes a processing unit and a non-transitory storage medium readable by the processing unit, the storage medium storing instructions that, when executed by the processing unit, cause the processing unit to: acquire and store input data provided to the trained AI agent, and output data generated by the trained AI agent based on the input data provided to the trained AI agent; and retrieve a baseline quantum state lattice matrix, the baseline quantum state lattice matrix being generated based on reference truth inputs provided to the trained AI agent, and reference truth outputs generated by the trained AI agent for each of the reference truth inputs provided to the trained AI agent, wherein each of the generated reference truth outputs includes multiple results inferred by the trained AI agent, and a probability amplitude associated with each of the multiple results. The instruction then directs the processing unit to: generate an output quantum state based on the acquired output data; generate a quantum state for each row of the baseline quantum state lattice matrix; and perform quantum-based anomaly classification of the acquired data based on the generated output quantum state and the quantum states generated for each row of the baseline quantum state lattice matrix.
[0012] In a further embodiment of this aspect, the generation of the baseline quantum state lattice matrix includes instructions for guiding the processing unit to perform the following operations: recording multiple results inferred by a trained AI agent as the base states of the baseline quantum state lattice matrix; and recording the probability amplitude associated with each of the multiple results as an element of the baseline quantum state lattice matrix, wherein each row of the matrix is associated with a reference truth input.
[0013] In a further embodiment of this aspect, the instructions for the processing unit to perform quantum-based anomaly classification of the captured output data include instructions for instructing the processing unit to perform the following operations: identify the base state with the highest probability amplitude from the output quantum state; perform quantum amplitude amplification on the identified base state based on the output quantum state; for each row in the baseline quantum state lattice matrix, perform quantum amplitude amplification on base states similar to the identified base state based on the quantum states generated for that row; label each of the base state lattice matrix in that row as a quantum-amplified base state exhibiting a certain degree of similarity to the identified base state after quantum amplification; and determine that the acquired input data includes an AI adversarial attack when a reference truth input associated with each of the labeled, quantum-amplified base states exhibits a certain degree of dissimilarity to the acquired input data.
[0014] On the other hand, this application provides a method for detecting adversarial AI attacks on a trained artificial intelligence (AI) agent using a defense module communicatively coupled to the AI agent. The method provided by this application includes the following steps: acquiring and storing input data provided to the trained AI agent, and output data generated by the trained AI agent based on the input data provided to the trained AI agent; and retrieving a baseline quantum state lattice matrix, which is generated based on a baseline truth input provided to the trained AI agent, and baseline truth outputs generated by the trained AI agent for each of the baseline truth inputs provided to the trained AI agent, wherein each of the generated baseline truth outputs includes multiple results inferred by the trained AI agent, and a probability amplitude associated with each of the multiple results. The method then includes the following steps: generating output quantum states based on the acquired output data, generating quantum states for each row of the baseline quantum state lattice matrix, and performing quantum-based anomaly classification of the acquired data based on the generated output quantum states and the quantum states generated for each row of the baseline quantum state lattice matrix.
[0015] According to a further embodiment of this aspect, the generation of a baseline quantum state lattice matrix includes the following steps: recording multiple results inferred by a trained AI agent as the base states of the baseline quantum state lattice matrix; and recording the probability amplitude associated with each of the multiple results as an element of the baseline quantum state lattice matrix, wherein each row of the matrix is associated with a reference truth input.
[0016] According to a further embodiment of this aspect, the step of performing quantum-based anomaly classification of the captured output data includes the following steps: identifying the base state with the highest probability amplitude from the output quantum state; performing quantum amplitude amplification on the identified base state based on the output quantum state; for each row in the baseline quantum state lattice matrix, performing quantum amplitude amplification on base states similar to the identified base state based on the quantum state generated for that row; marking each of the base states in that row of the baseline quantum state lattice matrix as a quantum-amplified base state exhibiting a certain degree of similarity to the identified base state after quantum amplification; and determining that the acquired input data includes an artificial intelligence adversarial attack when a reference truth input associated with each of the marked, quantum-amplified base states exhibits a certain degree of dissimilarity to the acquired input data. Attached Figure Description
[0017] The various embodiments of this application will be described below with reference to the following drawings:
[0018] Figure 1 This diagram illustrates a system block diagram for detecting adversarial AI attacks on an AI agent, according to an embodiment of this application.
[0019] Figure 2 An exemplary two-dimensional quantum state lattice matrix is shown according to an embodiment of this application;
[0020] Figure 3 An exemplary three-dimensional quantum state lattice matrix is shown according to an embodiment of this application;
[0021] Figure 4 The diagram shows a block diagram illustrating a processing system for executing embodiments of this application;
[0022] Figure 5 A flowchart illustrating a process for detecting adversarial AI attacks at a trained artificial intelligence (AI) agent using a defense module communicatively coupled to the AI agent, according to an embodiment of this application, is shown.
[0023] Figure 6 A flowchart illustrating the process of performing anomaly classification of captured output data according to an embodiment of this application is shown. Detailed Implementation
[0024] The following detailed description, with reference to the accompanying drawings, illustrates the details and embodiments of this application for illustrative purposes. Features described in the context of one embodiment may also be applied to the same or similar features in other embodiments, even if not explicitly described in other embodiments. Additions and / or combinations and / or alternatives described with respect to features in the context of one embodiment may be correspondingly applied to the same or similar features in other embodiments.
[0025] In the context of the various embodiments, the articles “a,” “an,” and “the” used for a feature or element include references to one or more of that type of feature or element.
[0026] In the context of the various embodiments, when the terms “about” or “approximately” are applied to numerical values, they cover the exact numerical value and a reasonable range of deviation that is generally understood in the relevant art, such as a deviation range within 10% of the specified value.
[0027] The term “and / or” as used herein includes any and all combinations of one or more of the listed related items.
[0028] As used herein, the term "including" means, but is not limited to, the content that follows. Therefore, the use of the term "including" indicates that the listed elements are required or mandatory, but other elements are optional and may or may not be present.
[0029] As used in this document, the term "consisting of" means including and limited to the contents listed herein. Therefore, the use of the term "consisting of" indicates that the listed elements are required or mandatory, and that no other elements are present.
[0030] The terms “first,” “second,” and similar terms used herein are used to distinguish similar objects in the specification, claims, and drawings, and are not necessarily used to describe a particular order or chronological sequence.
[0031] As used herein, the term "AI agent" and similar terms refer to a computer module designed to autonomously perform tasks, reason, or make decisions based on its environment and objectives. This computer module typically includes a trained machine learning model, which has been set up and tuned using a training dataset to recognize patterns, make decisions, or predict outcomes. The trained machine learning model can then be used by the AI agent to infer or predict from new, unseen data.
[0032] As used herein, the term "quantum state lattice matrix" and similar terms refer to a multidimensional array in the specification, where each value in the matrix may include a complex number that encodes both the amplitude and phase of the probability amplitude of each state of a qubit. Specifically, the elements in the matrix represent the probability amplitudes of various combinations of the system's fundamental states.
[0033] The term "fundamental quantum state" and similar expressions used in this document refer to a subset of the principal states, which, after decomposition of the fundamental state (i.e., expressing the state in more complex states), will generate all potential superpositions of the quantized system. The determination of the fundamental state typically involves various methods—all aiming to capture a subset of states that can subsequently undergo a decomposition process that reveals all feasible superpositions, including the fundamental state applicable to a particular quantum system. In other words, the fundamental state is a basic vector representing the possible states of a quantum system. For example, the fundamental state of a single-qubit system can be represented as follows: or Or, the fundamental state of a two-qubit system can be represented as , , and Each fundamental state represents all possible combinations of states in the system. It is important to note that each fundamental state corresponds to a different setting of the qubit, and any quantum state of the system can be represented as a superposition (linear combination) of these fundamental states, which has complex coefficients called probability amplitudes.
[0034] Furthermore, those skilled in the art will recognize that certain functional units in this specification are labeled as modules, submodules, or sets of processing elements. They will also recognize that modules, submodules, or sets of processing elements can be implemented as circuits, logic chips, or any type of discrete component. Moreover, they will recognize that modules, submodules, or sets of processing elements can be implemented in software and then executed by various processor architectures. In embodiments of this application, modules, submodules, or sets of processing elements may also include computer instructions, computations, or executable code that instruct a computer processor to perform a series of events based on received instructions. The choice of implementation for modules, submodules, or sets of processing elements is left as a design choice to those skilled in the art and does not in any way limit the scope of the claimed subject matter.
[0035] Quantum computing attempts to perform computations using the principles of quantum mechanics. Unlike classical computers, which use bits (which can only be in one of two states, labeled 0 or 1, at any given time), quantum computers use qubits. These qubits can be in any superposition of these two states, allowing quantum computers to process multiple computational paths simultaneously.
[0036] Unlike classical bits, qubits can exist not only in the two discrete states, but also in all possible linear superpositions of these states. In mathematical terms, the state of a qubit can be represented by a state vector in a two-dimensional Hilbert space. Using Dirac notation, a qubit... The state vector can be written as:
[0037] …Formula (1)
[0038] in and It is a complex number, and When quantum bits When being measured, there is The probability will return ,have The probability will return It is important to note that quantum measurement is nondeterministic, and the act of measurement irreversibly alters the quantum state. In other words, before a qubit is measured, it will be in a certain state. and In the quantum superposition state. Once measured, the result will no longer be in a quantum state, but in a classical state. Therefore, the measurement result will be or It is not a superposition of two states. This is because during the measurement of a qubit, the quantum state collapses into its observed classical state, and all subsequent measurements deterministically yield the same result with a probability equal to 1.
[0039] In the embodiments of this application, the quantum computing processes and / or steps described herein can be implemented using the Qiskit Software Development Kit (SDK) developed by IBM. The Qiskit SDK is a comprehensive software development kit for quantum computing, enabling users to experiment with quantum algorithms, circuits, and applications. Built on the Python programming language, Qiskit provides a user-friendly interface for creating, manipulating, and simulating quantum circuits, as well as for interfacing with real quantum hardware. The SDK includes a rich set of tools, libraries, and resources supporting every stage from designing quantum circuits to executing them on actual quantum processors.
[0040] Although quantum computing hardware is still in its early stages and not yet widely available, the Qiskit SDK primarily runs on traditional computing systems such as laptops, desktops, servers, and cloud-based platforms. It includes a powerful simulator that allows users to simulate the behavior of quantum circuits and algorithms on classical computers, enabling them to test and debug quantum programs without access to the quantum hardware. Furthermore, Qiskit provides visualization tools and libraries that run seamlessly on traditional computing devices, allowing users to visualize and analyze quantum circuits, state vectors, and measurement results.
[0041] In embodiments of this application, a defense module can be deployed on the AI agent to enhance the security and resilience of the AI system against adversarial AI attacks. The primary function of this module is to detect, mitigate, and prevent malicious activities targeting the AI agent. It operates by continuously monitoring the AI agent's inputs and outputs, analyzing patterns and behaviors to detect signs of adversarial manipulation, and applying appropriate countermeasures to protect the system's integrity. By integrating the defense module, the AI agent can better maintain its performance and reliability in the face of evolving cybersecurity threats, ensuring it remains a trustworthy and effective tool in its respective applications.
[0042] Figure 1A system for detecting adversarial AI attacks at an AI agent is illustrated according to an embodiment of this application. System 100 includes an AI agent 102 and a defense module 106 deployed at the AI agent 102, wherein the defense module 106 is configured to monitor data provided to the AI agent 102 and data generated by the AI agent 102. In embodiments of this application, the AI agent 102 may be configured to receive data 104 via a network 105.
[0043] In embodiments of this application, data 104 may include: sensor data acquired from various sensors, such as temperature readings, motion detection, or changes in light intensity; data acquired from user interactions, such as user commands, questions, or responses provided via voice or text input; data acquired from market sources, such as prices, transaction volumes, or economic indicators; data acquired from network traffic, such as network traffic logs, access attempts, or detected anomalies; data acquired from health indicators, such as a patient's vital signs, medical test results, or changes in health status; or data acquired from environmental indicators, such as weather conditions, soil moisture levels, or crop growth data; or data acquired from social media, such as posts, comments, likes, or shares. Those skilled in the art will recognize that data 104 is not limited to the examples provided above, but may include all data that can be used by trained machine learning models to perform predictions, classifications, and / or inferences.
[0044] As for network 105, the network may include one or more computer communication networks, such as the Internet, wired networks (e.g., local area networks (LANs) or wide area networks (WANs)), wireless networks (e.g., wireless LANs (WLANs) or mobile networks), or any other similar networks. Each computing / processing device in system 100 may also be provided with a network adapter card or a network interface module to facilitate communication between its respective modules and / or components.
[0045] In embodiments of this application, the defense module 106 may include a lattice matrix submodule 108, which is configured to generate and store a quantum state lattice matrix to record and classify the underlying states associated with the output probabilities of the results predicted, classified, and / or inferred by the AI agent 102, and to record and store the probability amplitudes of the results predicted, classified, and / or inferred by the AI agent 102. The defense module 106 also has an anomaly classification submodule 110 and an attack mitigation submodule 112. The anomaly classification submodule 110 is configured to perform anomaly classification on events occurring at the AI agent 102, while the attack mitigation submodule 112 is configured to implement mitigation strategies to protect the AI agent 102 from detected adversarial AI attacks.
[0046] During operation, the defense module 106 will be deployed at the AI agent 102, where the module 106 can be communicatively coupled to the input and / or output ports of the AI agent 102 so that any data provided to the AI agent 102 via the network 105 and / or generated by the AI agent 102 will be received by the defense module 106.
[0047] During the initialization or setup phase, the defense module 106 will begin its interaction with the AI agent 102. It is important to note that during this phase, the AI agent 102 includes a machine learning model trained to perform specific and / or various types of classification, prediction, and / or inference, and is in a state where it can be deployed for its intended function. For example, the AI agent 102 may include a machine learning model trained to identify and classify objects in a digital image based on features in that digital image. In other words, the trained machine learning model can utilize features in the digital image, such as the color in each pixel and the arrangement of that pixel relative to its neighboring pixels, to calculate the probability that an object contained in the digital image is classified as a specific object (e.g., a chair).
[0048] During the initialization phase, AI agent 102 is deployed in a "secure" operating environment where the inputs provided to AI agent 102 come from trusted sources, i.e., including benchmark truth inputs. This is to ensure that the outputs generated by AI agent 102 during this initialization phase include the corresponding benchmark truth outputs for each of the inputs provided to AI agent 102, rather than the results of compromised data and / or adversarial attacks. While AI agent 102 is performing its classification, prediction, and / or inference processes, lattice matrix submodule 108 is configured to record and store in a database the inputs provided to AI agent 102, the results generated by AI agent 102, and the probability amplitude associated with each of the results generated by AI agent 102. In embodiments of this application, this initialization phase may include a predefined time period (e.g., all output data generated within a week or month may be recorded) or may include a predetermined input dataset. Typically, this phase tends to last for a period of time or until there is a sufficiently large input dataset to ensure that all possible outputs generated by the AI agent can be captured by lattice matrix submodule 108. The exact duration of the initialization phase is left to those skilled in the art as a design choice.
[0049] As an example, continuing from the previous example, based on the digital image received by AI agent 102, AI agent 102 can determine during this initialization phase that there is a 0.6 probability that the image contains a firearm, a 0.2 probability that the image contains a bladed weapon, a 0.15 probability that the image contains a physical assault item, and a 0.05 probability that the image contains no objects. The lattice matrix submodule 108 then stores the features of the digital image (as input data) along with the corresponding results generated by AI agent 102 and the probability associated with each of the results in a database.
[0050] Once the initialization phase is complete, the lattice matrix submodule 108 can then proceed to generate a quantum state lattice matrix based on information stored in the database. In embodiments of this application, each column of the matrix can be defined to represent a basic state, where each basic state in the matrix can represent a result inferred by the AI agent 102. Furthermore, for each row of the matrix, each element in that row can represent the probability of the corresponding result occurring when a specific input is provided to the AI agent 102. An example of such a quantum state lattice matrix is... Figure 2As shown in the figure. Matrix 200 is generated based on the hypothesis that AI agent 102 generates four (4) possible outcomes for seven (7) input datasets (i.e., In1 to In7). Each of the four possible outcomes is determined by the base state (i.e., to One of them) is represented, and each basic state is used as a unique header for each column. Furthermore, each cell in matrix 200 (e.g., p) is represented by a unique header for each column. 1,1 to p 7,4 () represents the probability of the corresponding outcome occurring. For example, This indicates that when input In4 is provided to AI agent 102, it is related to... The probability that the associated results will occur.
[0051] exist Figure 2 In the example shown, a two-qubit system is used because four fundamental states are required. It should be noted that as the number of fundamental states increases, the number of qubits used will increase according to the following defined relationship: for an n-qubit system, there will be 2... n There are 16 fundamental states. For example, if 16 fundamental states are needed, a 4-qubit system would be used.
[0052] In embodiments of this application, the lattice matrix submodule 108 can be configured to generate a three-dimensional (3D) quantum state lattice matrix, where each layer represents a quantum state lattice matrix for a specific time period. Such a three-dimensional quantum state lattice matrix... Figure 3 The diagram illustrates that each layer in matrix 300 represents a quantum state lattice matrix for a specific time period. In this embodiment, layer 302, or matrix 302, represents a quantum state lattice matrix for a first time period, where all input data (i.e., the rows of the matrix) and the probabilities of the corresponding results generated by AI agent 102 during this first time period are plotted in matrix 302; layer 304 represents a quantum state lattice matrix for a second time period, where all input data (i.e., the rows of the matrix) and the probabilities of the corresponding results generated by AI agent 102 during this second time period are plotted in layer 304; and layer 306 represents a quantum state lattice matrix for a m-th time period, where all input data (i.e., the rows of the matrix) and the probabilities of the corresponding results generated by AI agent 102 during this m-th time period are plotted in matrix 306. The exact number of layers in the three-dimensional quantum state lattice matrix is left to those skilled in the art as a design choice. Furthermore, each time period may include a day, a week, a month, or any other predetermined period, as needed.
[0053] In a further embodiment of this application, the lattice matrix submodule 108 does not need to wait for the initialization phase to end, but instead generates and fills the cells of the quantum state lattice matrix during the initialization phase—synchronously with receiving data. In other words, when the lattice matrix submodule 108 receives input data and records and stores the results generated by the AI agent 102 (based on the input data) and the probability amplitude associated with each of the generated results, the submodule 108 can simultaneously use this information received by the submodule 108 to fill the cells of the quantum state lattice matrix in this step.
[0054] Once the lattice matrix submodule 108 completes the generation of the two-dimensional or three-dimensional quantum state lattice matrix, the generated matrix can then be used by the defense module 106 as a baseline matrix to monitor the performance of the AI agent 102 against anomalous behavior.
[0055] Reference Figure 1 In normal operation, when new data is provided to AI agent 102, AI agent 102 continues to generate a set of possible results based on the received data. Defense module 106 then captures the data provided to AI agent 102 and the results generated by AI agent 102. Anomaly classification submodule 110 then compares the captured information with the information in the previously generated baseline matrix to determine whether AI agent 102 has been compromised by an adversarial AI attack. It is important to note that all possible results generated by AI agent 102 in this step will be similar to the underlying states defined in the previously generated baseline matrix.
[0056] In embodiments of this application, each of the results generated by AI agent 102 may include multiple results, each of which is associated with a specific probability representing the likelihood of that result occurring. Anomaly classification submodule 110 then generates quantum states based on the underlying state associated with each of the results (as defined during the generation of the baseline matrix) and its corresponding probability of occurrence. (This is a superposition of multiple results). Once this is done, the anomaly classification submodule 110 then generates quantum states for all fundamental states and their corresponding probabilities of occurrence for each row in the baseline matrix. It is important to note that the combination of probability amplitudes and their corresponding fundamental states can be defined as terms or components of quantum states—where each term in a quantum state represents the product of the probability amplitude and its corresponding fundamental state.
[0057] In a first embodiment of this application, the anomaly classification submodule 110 analyzes the quantum states associated with the results generated by the AI agent 102. This process precisely identifies the fundamental state with the highest corresponding probability amplitude within the quantum state. Following this identification, the anomaly classification submodule 110 scans the quantum states associated with the baseline matrix to label fundamental states similar to the fundamental state of the identified state.
[0058] Subsequently, terms of quantum states associated with the labeled fundamental states, and quantum states associated with the initially identified states. The terms of these quantum states undergo quantum amplitude amplification—a process that significantly enhances the probability amplitude of the terms in these selected quantum states, making them more prominent in the quantum state spectrum. In other words, quantum amplitude amplification is performed on the terms of these quantum states to make the amplitudes of these terms more distinguishable compared to the overall noise level. It is important to note that an amplified quantum state is defined as a quantum state whose probability amplitude associated with the fundamental state of the quantum state is intentionally increased through the quantum amplitude amplification process. For simplicity, when referring to the first fundamental state that has undergone quantum amplification... st When referring to the basis state, it should be understood that the probability amplitude of the first basis state has been amplified through a quantum amplitude amplification process, thus forming the corresponding amplified quantum state.
[0059] After amplification, submodule 110 will obtain the quantum-amplified labeled fundamental and quantum states from the baseline matrix. The identified state is compared with the quantum amplified state to identify the quantum state. The quantum-amplified labeled fundamental state is the one that best matches the identified quantum state after quantum amplification; that is, the quantum-amplified labeled fundamental state that exhibits a certain degree of similarity to the identified quantum state in each row of the quantum state lattice matrix. According to embodiments of this application, this similarity can be determined by measuring the fidelity between the compared amplified fundamental states—where a fidelity close to "1" (e.g., greater than 0.95) can be considered to exhibit sufficient similarity. It is worth noting that for two quantum states… and Its fidelity F can be defined as Then, the corresponding rows of these quantum-amplified labeled fundamental states from the baseline matrix are further labeled for use in additional analysis.
[0060] During the analysis phase, submodule 110 retrieves the input data associated with these identified rows and compares it with the new data provided to AI agent 102. In other words, in this step, when the acquired input data associated with the identified rows exhibits a certain degree of dissimilarity with the new data provided to AI agent 102, submodule 110 determines whether the input data includes AI adversarial attacks. The dissimilarity between datasets can be evaluated using various statistical and computational methods. For a direct comparison of the input data associated with each of the labeled rows with the new data, Euclidean distances can be used to quantify the dissimilarity in spatial or multivariate data. Furthermore, relevance measures such as Pearson or Spearman's rank provide insights into linear or monotonic relationships, respectively. Techniques such as cosine similarity and the Jaccard index are particularly useful in text and set comparisons, while Hamming distance is suitable for comparing data strings of equal length. For more complex or structured data, visualization tools such as dendrograms or machine learning models that incorporate clustering and dimensionality reduction techniques can be used to reveal underlying patterns and groupings, highlighting differences that are not immediately apparent through direct statistical methods. The choice of method depends on the type of machine learning model employed by the AI agent 102, and is therefore left to those skilled in the art for design considerations.
[0061] In summary, this comparison aims to detect significant differences between historical and new data, which would indicate potential harm to AI agent 102. If such differences are found, submodule 110 will trigger attack mitigation submodule 112, initiating necessary procedures to mitigate the threat, thereby ensuring the integrity and security of AI agent 102.
[0062] In a further embodiment of this application, the anomaly classification submodule 110 can analyze the quantum states associated with the results generated by the AI agent 102. In order to identify this quantum state The system identifies several fundamental states with the highest probability amplitude. Following this identification, the anomaly classification submodule 110 scans the quantum states associated with the baseline matrix to label fundamental states similar to those of the identified states.
[0063] Subsequently, terms of quantum states associated with the labeled fundamental states, and quantum states associated with the initially identified states. The terms of the aforementioned quantum states undergo a multi-target quantum amplitude amplification process to make the amplitudes of these quantum state terms more distinguishable compared to the overall noise level. After this amplification, submodule 110 obtains the quantum-amplified labeled base states from the baseline matrix and the derived quantum states... The obtained quantum-amplified identified state is compared to identify the quantum state. The quantum-amplified identified states found in the matrix are the best-matched quantum-amplified labeled base states from the baseline matrix. The corresponding rows of these quantum-amplified states from the baseline matrix will then be labeled for further analysis as mentioned in the previous embodiments.
[0064] The description of the first embodiment can be best illustrated by the following simplified example. When new data I new When provided to AI agent 102, assume that AI agent 102 generates four possible outcomes, namely Out1 to Out4. Each of these outcomes and its probability of occurrence is associated with its corresponding base state in the baseline matrix, which can be represented as follows: Out1 and and probability Related, Out2 and and probability Related, Out3 and and probability Related, and Out4 with and probability Related.
[0065] Subsequently, the quantum state represents the superposition of these four fundamental states related to the outcome (i.e., Out1 to Out4). Generated:
[0066]
[0067] Since the detailed steps for generating quantum states are known to those skilled in the art, they are omitted in this description for the sake of brevity.
[0068] The anomaly classification submodule 110 then generates a superposition of all basic states and their corresponding occurrence probabilities for each row in the baseline matrix. Assuming matrix 200 is used as the baseline matrix, this means that for each row in matrix 200 (e.g., ... Figure 2As shown), this generates quantum states representing the superposition of probabilities of occurrence of all fundamental states in that row. Therefore, for a baseline matrix based on matrix 200, seven quantum states will be generated (since matrix 200 comprises 7 rows). The seven quantum states in this example can be defined as:
[0069]
[0070]
[0071]
[0072]
[0073]
[0074]
[0075]
[0076] Submodule 110 then analyzes the quantum state. To identify the base state with the highest probability amplitude, in this example, it is assumed that the highest probability amplitude is the one with the base state. Related p b Following this identification, submodule 110 scans the seven quantum states listed above to identify similar fundamental states from among these quantum states, i.e., the fundamental states. Then, the identified basic state is marked.
[0077] Subsequently, this marked base state, i.e., the base state The associated quantum state term, and the quantum state associated with the initially identified state. The terms undergo a quantum amplitude amplification process together. Then, the quantum-amplified labeled fundamental states and quantum states are compared... The identified state is compared with the quantum amplified state to identify the quantum state. The quantum-amplified identified state is the best match of the quantum-amplified labeled fundamental state.
[0078] Assuming a state of quantum amplification , , and Considered to be a state amplified by quantum mechanics The closest match (e.g., difference less than 0.05) is then retrieved, and then the quantum-amplified states are retrieved. , , and Relevant input data, and compare it with new data I new The data is compared to determine if there are any significant differences between the historical input data and the new data. If such differences are found, submodule 110 will trigger attack mitigation submodule 112 to initiate necessary procedures to mitigate the threat, thereby ensuring the integrity and security of AI agent 102.
[0079] In a second embodiment of this application, the defense module 106 may use a three-dimensional (3D) baseline matrix with two or more layers instead of a two-dimensional baseline matrix. Each layer in the 3D baseline matrix represents a different time frame, thereby allowing for a more dynamic and historical perspective of the system over time.
[0080] Similar to the first embodiment, the anomaly classification submodule 110 will subsequently analyze the quantum states associated with the results produced by the AI agent 102. This process precisely identifies the fundamental state with the highest probability amplitude within the quantum state. Following this identification, the anomaly classification submodule 110 scans the quantum states associated with each layer in the 3D baseline matrix to label fundamental states similar to the identified state. This step of scanning the quantum states associated with the layers in the 3D baseline matrix effectively compares the latest results generated by the AI agent 102 with the historical and current data patterns represented by the 3D baseline matrix.
[0081] Subsequently, the terms of quantum states associated with this labeled fundamental state, and the quantum states associated with the initially identified state. The terms undergo quantum amplitude amplification together. Then, submodule 110 will obtain the quantum-amplified labeled fundamental states and quantum states from the 3D baseline matrix. The identified state is compared with the quantum amplified state to identify the quantum state. The quantum-amplified identified state is matched best by the quantum-amplified labeled base state. The corresponding rows of these amplified quantum states from each layer of the 3D baseline matrix are then labeled for further analysis. During the analysis phase, submodule 110 retrieves the input data associated with these labeled rows and compares it with new data provided to AI agent 102. If a significant difference is found between the historical data and the new data, submodule 110 will then trigger attack mitigation submodule 112, initiating necessary procedures to mitigate the threat and ensure the integrity and security of AI agent 102. In a further embodiment, submodule 110 may be configured to trigger attack mitigation submodule 112 only if it is determined that the difference between the historical data and the new data occurs on a substantial number of layers of the 3D baseline matrix.
[0082] By extending the baseline matrix to include multiple layers representing different time frames, the defense module 106 gains a more robust capability to monitor, detect, and respond to anomalies based on both current and historical data patterns. This approach leverages the temporal dynamics provided by the 3D baseline matrix to enhance the predictive power and security robustness of the defense module 106.
[0083] According to the embodiments of this application, Figure 4 A block diagram illustrating the components of a processing system 400 is shown. This processing system 400 may be provided within the defense module 106 and its various sub-modules to perform digital signal processing functions or calculations according to embodiments of this application, or the processing system 400 may be located in any other module or sub-module of the system. Those skilled in the art will recognize that the exact configuration of each processing system within these modules or sub-modules may differ, and the exact configuration of the processing system 400 may vary. Figure 4 The arrangement shown is provided as an example only.
[0084] In embodiments of this application, the processing system 400 may include a controller 401 and a user interface 402. The user interface 402 is arranged to enable manual interaction between the user and the computing modules when needed, and for this purpose, the user interface 402 includes user input instructions to provide each of these modules with the input / output components required for updates. Those skilled in the art will recognize that the components of the user interface 402 may vary from embodiment to embodiment, but will generally include one or more of a display 440, a keyboard 435, and optical devices 436.
[0085] The controller 401 communicates with the user interface 402 via a bus 415 and includes: a memory 420; a processing unit, processing element, or processor 405 mounted on a circuit board for processing instructions and data for executing the methods of this embodiment; an operating system 406; an input / output (I / O) interface 430 for communicating with the user interface 402; and a communication interface, which in this embodiment is in the form of a network interface card 450. The network interface card 450 can be used, for example, to send data to or receive data from other processing devices via wired or wireless networks. Wireless networks that the network interface card 450 can use include, but are not limited to, Wireless-Fidelity (Wi-Fi), Bluetooth, Near Field Communication (NFC), cellular networks, satellite networks, telecommunications networks, and Wide Area Networks (WAN).
[0086] Memory 420 and operating system 406 communicate with processor 405 via bus 410. The memory components include both volatile and non-volatile memory, and more than one of each type, including Random Access Memory (RAM) 423, Read Only Memory (ROM) 425, and mass storage devices 445, the latter including one or more solid-state drives (SSDs). Those skilled in the art will recognize that the aforementioned memory components include non-transitory computer-readable media and should be considered to include all computer-readable media except for transient propagation signals. Typically, instructions are stored in the memory components in the form of program code, but may also be hardwired. Memory 420 may include a kernel and / or programming modules, such as software applications that may be stored in volatile or non-volatile memory.
[0087] The term "processor" as used herein is used generically to refer to any device or component capable of processing such instructions, and may include: microprocessors, processing units, multiple processing elements, microcontrollers, programmable logic devices, or any other type of computing device. That is, processor 405 may be provided by any suitable logic circuitry for receiving input, processing instructions stored in memory, and generating output (e.g., output to a memory component or display 440). In this embodiment, processor 405 may be a single-core or multi-core processor with memory-addressable space. In one example, processor 405 may be multi-core, such as including an 8-core central processing unit (CPU). In another example, it may be a cluster of CPU cores running in parallel to accelerate computation.
[0088] Figure 5 A process for detecting adversarial AI attacks at a trained AI agent is shown, wherein process 500 can be performed by a defense module communicatively coupled to the AI agent, according to embodiments of this application.
[0089] Process 500 begins at step 502, where input data provided to the trained AI agent is acquired and stored. Simultaneously, process 500 also acquires and stores output data generated by the trained AI agent, which is based on the input data provided to the trained AI agent. Once this is complete, process 500 proceeds to step 504, where process 504 retrieves the previously generated quantum state lattice matrix from a database and / or memory contained within the defense module. In embodiments of this application, this database may be provided on a remote server or cloud server and may be provided to the defense module via wireless or wired communication. Then, at step 506, process 500 generates an output quantum state based on the acquired output data.
[0090] At step 508, process 500 then generates quantum states for each row of the baseline quantum state lattice matrix. Then, process 500 performs quantum-based anomaly classification of the acquired data based on the generated output quantum states and the quantum states of each row in the baseline quantum state lattice matrix. This occurs at step 510.
[0091] In other embodiments of this application, when process 500 determines at step 510 that the acquired input data is classified as an AI adversarial attack, process 500 may proceed to step 512. At step 512, process 500 will implement attack mitigation measures to resolve and / or delay AI adversarial attacks on the trained AI agent.
[0092] Figure 6 The execution process for anomaly classification of the captured output data is shown, wherein process 600 can be executed by the defense module according to an embodiment of this application.
[0093] Process 600 begins at step 602, identifying the fundamental state with the highest probability amplitude from the output quantum states. Then, process 600 proceeds to step 604, where it subsequently performs quantum amplitude amplification on the identified fundamental state based on the output quantum state. At step 606, process 600 performs quantum amplitude amplification on fundamental states similar to the identified fundamental state based on the quantum states generated for each row of the baseline quantum state lattice matrix. Process 600 then marks, at step 608, each of the rows of the baseline quantum state lattice matrix as a quantum-amplified fundamental state exhibiting a certain degree of similarity to the quantum-amplified identified fundamental state. Then, at step 610, when the reference truth input associated with each of the marked, quantum-amplified fundamental states exhibits a certain degree of dissimilarity from the acquired input data, process 600 classifies the acquired input data as an AI adversarial attack.
[0094] According to another embodiment of this application, a baseline quantum state lattice matrix can be generated by the following process: recording multiple results inferred by a trained AI agent as the base states of the baseline quantum state lattice matrix; and recording the probability amplitude associated with each of the multiple results as an element of the baseline quantum state lattice matrix, wherein each row of the matrix is associated with a reference truth input.
[0095] In another embodiment of this application, the process of performing anomaly classification of the captured output data may include the following steps: generating an output quantum state based on the acquired output data, and identifying multiple fundamental states with probability amplitudes higher than a predetermined threshold from the output quantum state; performing multi-target quantum amplitude amplification on the identified multiple fundamental states based on the output quantum state; generating a quantum state for each row of the baseline quantum state lattice matrix based on the fundamental states of the baseline quantum state lattice matrix and the probability amplitude associated with each of these fundamental states in that row; performing multi-target quantum amplitude amplification on the fundamental states similar to the identified multiple fundamental states based on the quantum states generated for that row of the baseline quantum state lattice matrix; marking each of the rows of the baseline quantum state lattice matrix as a quantum-amplified fundamental state exhibiting a certain degree of similarity to the identified multiple fundamental states after quantum amplification; and determining that the acquired input data includes an AI adversarial attack when the reference truth input associated with each of the marked, quantum-amplified fundamental states exhibits a certain degree of dissimilarity to the acquired input data.
[0096] Those skilled in the art can identify many other changes, substitutions, variations and modifications, and this application is intended to cover all such changes, substitutions, variations and modifications that fall within the scope of the appended claims.
Claims
1. A defense module for detecting adversarial artificial intelligence attacks at a trained artificial intelligence agent communicatively coupled to the defense module, the defense module comprising: a processing unit; and a non-transitory storage medium readable by the processing unit, the storage medium storing instructions that, when executed by the processing unit, cause the processing unit to: obtain and store input data provided to the trained artificial intelligence agent, and output data generated by the trained artificial intelligence agent based on the input data provided to the trained artificial intelligence agent; retrieve a baseline quantum state lattice matrix generated based on benchmark ground truth inputs provided to the trained artificial intelligence agent, and benchmark ground truth outputs generated by the trained artificial intelligence agent for each of the benchmark ground truth inputs provided to the trained artificial intelligence agent, and wherein each of the generated benchmark ground truth outputs comprises a plurality of outcomes inferred by the trained artificial intelligence agent, and a probability amplitude associated with each of the plurality of outcomes; generate an output quantum state based on the obtained output data, wherein the obtained output data comprises a plurality of outcomes inferred by the trained artificial intelligence agent for the obtained input data, and a probability amplitude associated with each of the plurality of outcomes; generate a quantum state for each row in the baseline quantum state lattice matrix; and perform a quantum-based anomaly classification of the obtained data based on the generated output quantum state and the quantum state generated for each row in the baseline quantum state lattice matrix.
2. The defense module of claim 1, wherein the generation of the baseline quantum state lattice matrix comprises instructions for directing the processing unit to: record the plurality of outcomes inferred by the trained artificial intelligence agent as a basis state of the baseline quantum state lattice matrix; and record the probability amplitude associated with each of the plurality of outcomes as an element of the baseline quantum state lattice matrix, wherein each row of the matrix is associated with a benchmark ground truth input.
3. The defense module of claim 1 or 2, wherein the instructions that cause the processing unit to perform a quantum-based anomaly classification of the captured output data comprises instructions for directing the processing unit to: identify a basis state having a highest probability amplitude from the output quantum state; perform quantum amplitude amplification on the identified basis state based on the output quantum state; for each row in the baseline quantum state lattice matrix, perform quantum amplitude amplification on a basis state similar to the identified basis state based on the quantum state generated for the row; quantum-amplified basis states of each of the rows of the baseline quantum state lattice matrix that exhibit a certain degree of similarity to the quantum-amplified identified basis states; and determining that the acquired input data includes an artificial intelligence adversarial attack when the ground truth inputs associated with each of the tagged quantum-amplified basis states exhibit a certain degree of dissimilarity to the acquired input data.
4. The defense module of claim 1 or 2, wherein the instructions that cause the processing unit to perform quantum-based anomaly classification of the captured output data include instructions to direct the processing unit to: identify, from the output quantum state, a plurality of basis states having probability amplitudes above a predetermined threshold; perform multi-objective quantum amplitude amplification on the identified plurality of basis states based on the output quantum state; for each row of the baseline quantum state lattice matrix, perform multi-objective quantum amplitude amplification on basis states similar to the identified plurality of basis states based on the quantum state generated for the row; tag quantum-amplified basis states of each of the rows of the baseline quantum state lattice matrix that exhibit a certain degree of similarity to the quantum-amplified identified plurality of basis states; and determine that the acquired input data includes an artificial intelligence adversarial attack when the ground truth inputs associated with each of the tagged quantum-amplified basis states exhibit a certain degree of dissimilarity to the acquired input data.
5. The defense module of any one of claims 1 to 4, wherein the baseline quantum state lattice matrix comprises a two-dimensional matrix.
6. The defense module of any one of claims 1 to 4, wherein the baseline quantum state lattice matrix comprises a three-dimensional matrix, wherein each layer of the three-dimensional matrix represents a unique time frame.
7. The defense module of claim 2, wherein each probability amplitude represents a likelihood that a result occurs when a ground truth input is provided to the artificial intelligence agent.
8. The defense module of any one of claims 1 to 6, further comprising instructions to direct the processing unit to: perform attack mitigation measures when the acquired input data is classified as an artificial intelligence adversarial attack.
9. The defense module of any one of claims 1 to 7, wherein the baseline quantum state lattice matrix is generated over a predetermined period of time.
10. A method of detecting an adversarial artificial intelligence attack at a trained artificial intelligence agent using a defense module communicatively coupled to the artificial intelligence agent, the method comprising: acquiring and storing input data provided to the trained artificial intelligence agent and output data generated by the trained artificial intelligence agent based on the input data provided to the trained artificial intelligence agent; performing quantum-based anomaly classification of the output data; retrieving a baseline quantum state lattice matrix generated based on ground truth inputs provided to the trained artificial intelligence agent and ground truth outputs generated by the trained artificial intelligence agent for each of the ground truth inputs provided to the trained artificial intelligence agent, wherein each of the generated ground truth outputs comprises a plurality of outcomes inferred by the trained artificial intelligence agent and a probability amplitude associated with each of the plurality of outcomes; generating an output quantum state based on the retrieved output data, wherein the retrieved output data comprises a plurality of outcomes inferred by the trained artificial intelligence agent for the retrieved input data and a probability amplitude associated with each of the plurality of outcomes; generating a quantum state for each row in the baseline quantum state lattice matrix; and performing quantum-based anomaly classification of the retrieved data based on the generated output quantum state and the quantum state generated for each row in the baseline quantum state lattice matrix.
11. The method of claim 10, wherein the generating of the baseline quantum state lattice matrix comprises the steps of: recording the plurality of outcomes inferred by the trained artificial intelligence agent as basis states of the baseline quantum state lattice matrix; recording the probability amplitude associated with each of the plurality of outcomes as an element of the baseline quantum state lattice matrix, wherein each row of the matrix is associated with a ground truth input.
12. The method of claim 10 or 11, wherein the step of performing quantum-based anomaly classification of the captured output data comprises the steps of: identifying a basis state having a highest probability amplitude from the output quantum state; performing quantum amplitude amplification on the identified basis state based on the output quantum state; performing quantum amplitude amplification on basis states similar to the identified basis state based on the quantum state generated for each row in the baseline quantum state lattice matrix; labeling quantum-amplified basis states of each of the rows of the baseline quantum state lattice matrix that exhibit a certain degree of similarity to the quantum-amplified identified basis state; when the ground truth input associated with each of the labeled quantum-amplified basis states exhibits a certain degree of dissimilarity to the retrieved input data, determining that the retrieved input data comprises an artificial intelligence adversarial attack.
13. The method of claim 10 or 11, wherein the step of performing quantum-based anomaly classification of the captured output data comprises the steps of: identifying a plurality of basis states having a probability amplitude higher than a predetermined threshold from the output quantum state; performing multi-objective quantum amplitude amplification on the identified plurality of basis states based on the output quantum state; performing multi-objective quantum amplitude amplification on basis states similar to the identified plurality of basis states based on the quantum states generated for each row of the baseline quantum state lattice matrix; labeling quantum-amplified basis states of each of the rows of the baseline quantum state lattice matrix that exhibit a certain degree of similarity to the quantum-amplified identified plurality of basis states; and determining that the acquired input data includes an artificial intelligence adversarial attack when the ground truth input associated with each of the labeled quantum-amplified basis states exhibits a certain degree of dissimilarity to the acquired input data.
14. The method of any one of claims 10 to 13, wherein the baseline quantum state lattice matrix comprises a two-dimensional matrix.
15. The method of any one of claims 10 to 13, wherein the baseline quantum state lattice matrix comprises a three-dimensional matrix, wherein each layer of the three-dimensional matrix represents a unique time frame.
16. The method of claim 11, wherein each probability amplitude represents a likelihood that a result occurs when a ground truth input is provided to the artificial intelligence agent.
17. The method of any one of claims 10 to 16, further comprising the step of: performing an attack mitigation measure when the acquired input data is classified as an artificial intelligence adversarial attack.
18. The method of any one of claims 10 to 17, wherein the baseline quantum state lattice matrix is generated over a predetermined period of time.