Identification method and device based on body active defense

By establishing interactive training methods for adversarial environment models and embodied intelligent models, the problem of low defense effect in the existing technology when facing unseen attack technologies and adaptive attack technologies is solved, and high-rosty adversarial patch defense is achieved.

CN119963896APending Publication Date: 2025-05-09启元实验室
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510022412.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art has low defense effect and insufficient robustness when facing unseen attack technology and adaptive attack technology generated by adversarial patches.

Method used

A recognition method based on embodied active defense is proposed. By establishing an adversarial environment model and embodied intelligent model, the training data set is constructed interactively, and the preset objective function is used to train the embodied intelligent model to improve the defense ability of adversarial patches.

Benefits of technology

Image recognition with active defense function for adversarial patches is realized, which improves the robustness of the model and can effectively defend against adversarial patches generated by unseen attack technologies and adaptive attack technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963896A_ABST
    Figure CN119963896A_ABST
Patent Text Reader

Abstract

The invention provides a recognition method and device based on body active defense, and relates to the technical field of image recognition. An identification method based on active defense with a body comprises the following steps: S1, establishing a confrontation environment model and an intelligent body model with the body according to a target identification task; s2, constructing a training data set based on interaction of the body-equipped agent model and the confrontation environment model; s3, based on the training data set, training the body-equipped agent model by using a preset objective function; s4, the steps S2-S3 are repeated until a preset ending condition is met, and a target body-equipped agent model is obtained; and S5, based on the target body agent model, predicting a to-be-predicted target of the target identification task to obtain a prediction result. According to the technical scheme provided by the embodiment of the invention, the body agent model is circularly trained according to the training data set, and the trained body agent model is utilized to predict the target recognition task, so that the robustness is good, and an active defense function is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image recognition technology, and in particular to a recognition method and device based on embodied active defense. Background Art

[0002] At present, the rapid development of artificial intelligence technology is profoundly changing the way of life and production of human beings. As a representative technology of artificial intelligence, deep neural networks have achieved remarkable achievements and demonstrated excellent performance in image recognition, speech processing, natural language understanding and other fields. By simulating the neural network structure of the human brain, deep neural networks can process complex data patterns and extract useful information from them. However, this powerful pattern recognition ability is also accompanied by certain limitations. In particular, when facing adversarial patches, the stability and reliability of deep neural networks are severely tested. Therefore, improving the robustness of deep neural networks has become an important direction of current artificial intelligence research.

[0003] Adversarial patches are maliciously designed patterns that can be covertly placed in complex real-world environments in the form of stickers, etc., in order to mislead deep neural networks into making incorrect predictions. In response to the adversarial patch problem, a series of technologies have been developed in the field of artificial intelligence security to defend against adversarial patch attacks. The relevant technologies can be mainly divided into two categories: adversarial training technology and input preprocessing technology. Among them, adversarial training technology introduces adversarial samples in the model training stage to enhance the robustness of the model to adversarial samples; input preprocessing technology detects and removes adversarial patches in the input image by training a detection network.

[0004] However, these techniques for defending against adversarial patches all rely on prior assumptions about adversarial patches. When faced with unseen attack techniques and adversarial patches generated by adaptive attack techniques, the defenses can easily fail and have low robustness. Summary of the invention

[0005] Based on this, the present application provides a recognition method and device based on embodied active defense to achieve robust image recognition with active defense function.

[0006] According to one aspect of the present application, a recognition method based on embodied active defense is proposed, including: S1: establishing an adversarial environment model and an embodied intelligent body model according to a target recognition task; S2: constructing a training data set based on the interaction between the embodied intelligent body model and the adversarial environment model; S3: based on the training data set, training the embodied intelligent body model using a preset objective function; S4: repeating steps S2-S3 until a preset end condition is reached to obtain a target embodied intelligent body model; S5: predicting a target to be predicted of the target recognition task based on the target embodied intelligent body model to obtain a prediction result.

[0007] According to some embodiments, an adversarial environment model and an embodied intelligent agent model are established according to a target recognition task, including: using a partially observable Markov decision process to construct an environment model related to the target recognition task, and adding an adversarial patch on the basis of the environment model to obtain an adversarial environment model; constructing an embodied intelligent agent model based on a neural network and a target recognition task; wherein the embodied intelligent agent model includes a perception model and a strategy model, the perception model includes a backbone visual feature extractor, a predictive attention fusion model and a cognitive attention fusion model, the perception model is used to perceive the adversarial environment model and perform recognition prediction based on the perception results, the strategy model is used to perform action prediction based on the perception results, and control the embodied intelligent agent model to move based on the results of the action prediction.

[0008] According to some embodiments, a training data set is constructed based on the interaction between the embodied intelligent agent model and the adversarial environment model, including: S21: obtaining an observed image data sample of the current observation perspective of the adversarial environment model at the current time step; S22: inputting the observed image data sample into the perception model of the embodied intelligent agent model to obtain a predicted recognition label sample of the current time step and an internal cognition sample of the current time step; S23: inputting the internal cognition sample of the current time step into the strategy model to obtain an optimal action sample of the current time step; S24: controlling the embodied intelligent agent model to move to the next observation perspective according to the optimal action sample; S25: repeating steps S21-S24 until a preset time step is reached; S26: calculating the interaction reward for each observation perspective according to the predicted recognition label sample and the true classification label sample of the observed image data using a pre-constructed dense reward function; S27: obtaining the interaction trajectory corresponding to each observation perspective as training data according to the observed image data sample, internal cognition sample, optimal action sample and interaction reward of each observation perspective, thereby constructing a training data set.

[0009] According to some embodiments, the observed image data samples are input into the perception model of the embodied intelligent agent model to obtain the predicted recognition label samples of the current time step and the internal cognitive samples of the current time step, including: inputting the observed image data samples into the backbone visual feature extractor to obtain image feature samples; inputting the image feature samples into the predictive attention fusion model to obtain the predicted recognition label samples of the current time step based on the internal cognitive samples and the image feature samples of the previous time step; inputting the image feature samples into the cognitive attention fusion model to obtain the internal cognitive samples of the current time step based on the internal cognitive samples and the image feature samples of the previous time step.

[0010] According to some embodiments, the preset objective function includes:

[0011]

[0012] Among them, θ is the optimal parameter of the perception model f(·; θ), φ is the optimal parameter of the policy model π(·; φ); It is expectation abbreviation of; is the true classification label sample of scene x; τ is the The interaction trajectory obtained by the policy model π in ; p represents the adversarial patch; represents the set of all adversarial patches under scenario x; A is the application function of the adversarial patch; Is the perceptual loss function, used to measure the prediction of the recognition label sample The distance from the true classification label sample y; H represents the maximum limit of the time step; is the conditional entropy, which is used to quantify the internal cognition sample b H-1 and observed image data sample o′ H Described under the conditions The amount of information required; λ is a hyperparameter constant that controls the degree of influence of conditional entropy, a t is the optimal action at the current time step, s t is the observation angle at time step t, b t is the internal cognitive sample at the current time step t; o t is the observed image data sample without adversarial patch interference, o′ t =A(o t ,p;s t ) is the observed image data sample with adversarial patch interference, is the state transition probability of scene x, is the observation function.

[0013] According to some embodiments, based on a training data set, an embodied intelligent body model is trained using a preset objective function, including: sampling data of a preset size from the training data set to obtain multiple training batches of sampled data; for each training batch of sampled data, calculating a trimmed proximal policy gradient objective function and a perception objective function; performing gradient ascent according to the trimmed proximal policy gradient objective function to update the model parameters of the policy model, and performing gradient descent according to the perception objective function to update the model parameters of the perception model, until all sampled data are traversed to obtain a target embodied intelligent body model.

[0014] According to some embodiments, the proximal policy gradient objective function is clipped as:

[0015]

[0016] In the formula, To tailor the proximal policy gradient objective function, φ is the optimal parameter of the policy model π(·; φ); clip is the clipping function; at is the optimal action at the current time step; b t is the internal cognitive sample at the current time step t; ∈ is the clipping range; is the advantage function; is the sampling data of the τth training batch.

[0017] According to some embodiments, the perceptual objective function is:

[0018]

[0019] In the formula, is the perception objective function; is the perceptual loss function, is the true classification label sample of scene x, o′ t is an observed image data sample with adversarial patch interference; λ is a hyperparameter constant that controls the degree of influence of conditional entropy; is the conditional entropy, which is used to quantify the internal cognition sample b H-1 and observed image data sample o′ H Described under the conditions The amount of information required.

[0020] According to some embodiments, adding an adversarial patch based on the environment model to obtain the adversarial environment model includes: performing uniform noise sampling on a superset to obtain an adversarial patch, and adding it to the environment model to obtain the adversarial environment model; or for each observation perspective, maximizing the perceptual loss of the perceptual model to obtain the adversarial patch, and adding it to the environment model to obtain the adversarial environment model, wherein the perceptual loss is used to measure the difference between the predicted recognition label sample and the true classification label sample at the current observation perspective.

[0021] According to some embodiments, based on the target embodied intelligent body model, the target to be predicted of the target recognition task is predicted to obtain a prediction result, including: obtaining the observed image data of the target to be predicted at the current observation angle; inputting the observed image data into the target embodied intelligent body model to obtain a predicted identification label and the optimal action for the current time step; controlling the embodied intelligent body model to move to the next observation angle according to the optimal action to obtain new observed image data until a preset time step is reached, and outputting the final predicted identification label as the prediction result.

[0022] According to one aspect of the present application, an identification device based on embodied active defense includes: a model building module, which is used to establish an adversarial environment model and an embodied intelligent body model according to a target recognition task; a data set construction module, which is used to construct a training data set based on the interaction between the embodied intelligent body model and the adversarial environment model; a model training module, which is used to train the embodied intelligent body model based on the training data set using a preset objective function; a loop iteration module, which is used to repeat steps S2-S3 until a preset end condition is reached to obtain a target embodied intelligent body model; and an identification prediction module, which is used to predict a target to be predicted of a target recognition task based on the target embodied intelligent body model to obtain a prediction result.

[0023] According to one aspect of the present application, an electronic device is proposed, which includes: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described above.

[0024] According to one aspect of the present application, a computer-readable medium is provided, on which a computer program is stored. When the program is executed by a processor, the method described above is implemented.

[0025] Through the above-mentioned embodiments provided in the present application, an adversarial environment model and an embodied intelligent agent model for training are constructed according to the target recognition task. The embodied intelligent agent model interacts with the adversarial environment model to obtain a training data set. The embodied intelligent agent model is cyclically trained according to the training data set to obtain a robust embodied intelligent agent model with defense capabilities against adversarial patches. The trained embodied intelligent agent model is used to predict the target recognition task, thereby realizing robust image recognition with active defense capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] It should be understood that the foregoing general description and the following detailed description are exemplary only and are not restrictive of the present application.

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without exceeding the scope of protection required by the present application.

[0028] Figure 1 A flowchart of an identification method based on embodied active defense provided in an embodiment of the present application;

[0029] Figure 2 A flowchart for constructing a training data set based on the interaction between the embodied intelligent agent model and the adversarial environment model provided in an embodiment of the present application;

[0030] Figure 3 A schematic diagram of the interaction between the embodied intelligent agent model and the adversarial environment model provided in an embodiment of the present application;

[0031] Figure 4 A flowchart of inputting observed image data samples into an embodied intelligent agent model to obtain predicted recognition label samples and optimal action samples for the current time step provided in an embodiment of the present application;

[0032] Figure 5 A flowchart of training an embodied intelligent agent model based on a training data set provided in an embodiment of the present application using a preset objective function until all sampled data are traversed to obtain a target embodied intelligent agent model;

[0033] Figure 6 A flowchart of generating an adversarial patch provided in an embodiment of the present application;

[0034] Figure 7 A flow chart of predicting a target to be predicted in a target recognition task based on a target embodied intelligent agent model provided in an embodiment of the present application to obtain a prediction result;

[0035] Figure 8 A schematic diagram of the defense effect under the object classification embodiment provided in the embodiment of the present application;

[0036] Fig. 9 A schematic diagram of the defense effect under the face recognition embodiment provided in the embodiment of the present application;

[0037] Fig.10 Schematic diagram of the defense effect under the target detection embodiment provided in the embodiment of the present application;

[0038] Fig.11 A block diagram of an identification device based on embodied active defense provided in an embodiment of the present application;

[0039] Fig.12 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0040] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0041] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present application. However, those skilled in the art will appreciate that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, known methods, devices, realizations or operations are not shown or described in detail to avoid blurring the various aspects of the application.

[0042] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0043] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0044] It should be understood that although the terms first, second, third, etc. may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another component. Therefore, the first component discussed below can be referred to as the second component without departing from the teachings of the concepts of the present application. As used herein, the term "and / or" includes any one of the associated listed items and all combinations of one or more.

[0045] For specific implementation methods, please refer to the following embodiments.

[0046] Figure 1 Flow chart of the identification method based on embodied active defense provided in the embodiment of the present application. Figure 1 As shown, the method includes steps S110 to S150, corresponding to steps S1 to S5.

[0047] In step S110, an adversarial environment model and an embodied agent model are established according to the target recognition task.

[0048] The target recognition task is a task waiting to be recognized, including the target to be predicted.

[0049] The adversarial environment model is obtained by adding adversarial patches to the environment model related to the target recognition task, including the three-dimensional object of the target required by the task (i.e., the target to be predicted).

[0050] The embodied agent model is a prediction model based on embodied agents. It can be understood that an embodied agent is an agent that has both perception and action capabilities. Unlike passive agents that passively receive information, embodied agents can actively explore their environment and interact.

[0051] In step S120, a training data set is constructed based on the interaction between the embodied intelligent agent model and the adversarial environment model.

[0052] Before training the embodied agent model, it is necessary to first perform multiple training data collections to build a training dataset. Specifically, data collection is achieved by combining the interaction of the embodied agent model that combines perception and action with the adversarial environment model.

[0053] What needs to be explained is that the embodied intelligent model can build historical cognition through continuous observation, and through strategic actions, further collect new perspective environmental information and record interaction trajectories to obtain training data. Repeating this process many times can obtain a training data set.

[0054] In step S130, the embodied intelligent agent model is trained based on the training data set using a preset objective function.

[0055] Using the collected training data set, the parameters of the embodied intelligent agent model are updated to minimize the preset objective function until all the data in the training data set have been trained, which is the embodied intelligent agent model that has completed the current training.

[0056] In step S140, steps S120-S130 are repeated until a preset end condition is reached to obtain a target embodied intelligent body model.

[0057] According to an example embodiment, the preset end condition may be set as the convergence of a preset objective function curve, so that an embodied intelligence model combining perception and action after defense may be obtained.

[0058] It should be noted that the embodied agent model for the first data collection is untrained, and its actions are random and non-optimal. Based on the feedback of the reward, the agent learns which random actions are useful. However, it is not possible to complete the training of the agent in one go by relying solely on the data set collected for the first time. Steps S120 and S130 need to be repeated continuously, and the agent continues to learn. After learning, the agent can collect better data (the actions gradually tend to be optimal).

[0059] In step S150, based on the target embodied intelligent agent model, the target to be predicted of the target recognition task is predicted to obtain a prediction result.

[0060] The trained target embodied agent model is used to perform recognition and prediction on the target recognition task. Specifically, the target embodied agent model interacts with the actual environment in multiple steps and finally outputs the prediction result.

[0061] This application constructs an adversarial environment model and an embodied intelligent agent model for training according to the target recognition task. The embodied intelligent agent model interacts with the adversarial environment model to obtain a training data set. The embodied intelligent agent model is cyclically trained according to the training data set to obtain a robust embodied intelligent agent model with defense capabilities against adversarial patches. The trained embodied intelligent agent model is used to predict target recognition tasks to achieve robust image recognition with active defense capabilities.

[0062] According to some embodiments, in step S110, a confrontation environment model and an embodied intelligent agent model are established according to the target recognition task, which can be specifically implemented through steps S210 to S220.

[0063] In step S210, a partially observable Markov decision process is used to construct an environment model related to the target recognition task, and an adversarial patch is added to the environment model to obtain an adversarial environment model.

[0064] This application designs a framework for an embodied intelligent agent and a confrontation environment that combines perception and action. The framework specifically relates to a modeling of an embodied intelligent agent that combines perception and action, and a modeling of a confrontation environment.

[0065] For modeling adversarial models, it is necessary to first build an environmental model. For example, the corresponding environmental model can be selected according to the target recognition task.

[0066] In a specific embodiment, the target recognition task is an object classification task, and an OmniObject3D scanned object dataset can be used as a three-dimensional environment to obtain an environment model.

[0067] In another specific embodiment, the target recognition task is a face recognition task, and the EG3D three-dimensional face generative model can be used as a three-dimensional environment to obtain an environment model.

[0068] In another specific embodiment, the target recognition task is a target detection task, and the CARLA autonomous driving simulation platform can be selected as the three-dimensional environment to obtain the environment model.

[0069] It should be emphasized that the constructed environment models must contain the three-dimensional objects required for the target recognition task (i.e., the targets to be predicted).

[0070] In order to facilitate the embodied intelligent body model to obtain the corresponding observed image data, according to an example embodiment, the environment model may also include a set of API interfaces that can obtain image visual observations, and move and rotate the simulated camera.

[0071] It can be understood that the constructed environment model is a non-adversarial environment and does not have an adversarial patch. This application adds an adversarial patch p to interfere with the observation o based on the environment model. t , construct adversarial observation o′ t adversarial environment.

[0072] Specifically, by observing o′ containing the adversarial patch p t Instead of t , strengthen the model's defense capability against adversarial patches, where adversarial patches p need to be obtained through calculation. Set of adversarial patches in scenario x The definition is Among them, H P , W P and C represent the height, width, and number of channels of the adversarial patch, respectively. The desired part of this definition indicates that the adversarial patch is aggressive under multiple views, i.e., it has multi-view robustness.

[0073] In step S210, the generation of the adversarial patch can be achieved by solving the optimization problem In actual operation, the technology such as PGD, EoT, MeshAdv, Face3DAdv or AdvYOLO can be used for implementation, and this application does not limit this.

[0074] The addition of the adversarial patch can be achieved by applying the adversarial patch function. Specifically, the application function is used to apply the adversarial patch according to the camera state s. t Replace the observed image o by affine transformation or perspective transformation t Part of the goal is to construct an adversarial observation o′ against the patch p t .

[0075] For modeling of adversarial environments, this embodiment uses a partially observable Markov decision process for modeling.

[0076] According to an example embodiment,

[0077] in, Represents the adversarial model under scenario x. and Indicates the state of the agent and the actions that the agent can perform. The state transition probability of scene x is given by express, is the observation space of the scene. Due to the partially observable nature, the agent cannot directly obtain the state s, but through the observation function Get the observation sample o.

[0078] It needs to be explained that for all normal scenarios There is a data distribution in is a true classification label sample of scene x, describing the content of the scene. The task of the embodied agent is to perform multi-step actions a 1:t And collect multiple observations o 1:t , and finally complete the correct prediction Where t is the time step.

[0079] In a specific embodiment, the target recognition task is an object classification task, x is different three-dimensional object models, and y is the category name of each object, such as shoes and hats; in another specific embodiment, the target recognition task is a face recognition task, x is different three-dimensional face models, and y is the corresponding identity name of each face, such as Tom and Jack; in another specific embodiment, the target recognition task is a target detection task, x is different three-dimensional vehicle models, and y is the detection box position of each vehicle in the image, such as [horizontal center coordinate, vertical center coordinate, horizontal width, vertical width] format.

[0080] In the adversarial setting, the adversarial patch p from the unknown adversarial attack technique perturbs the observation o t , leading to the adversarial observation o′ t =A(o t ,p;s t ), where A is the application function of the adversarial patch, which is used to apply the adversarial patch s t Arranged according to the observation angle t Among them, it should be noted that adversarial attack is an attack that misleads the neural network by injecting maliciously designed noise, patterns, etc. into the input. In order to deal with adversarial attacks, it is necessary to reinforce the neural network and improve the model's defense capabilities against adversarial attacks.

[0081] In the adversarial environment, the task of the embodied agent is to perform multiple actions a t And collect multiple observations o′ t , and finally complete the correct prediction

[0082] In step S220, an embodied intelligent body model is constructed based on the neural network and the target recognition task; wherein the embodied intelligent body model includes a perception model and a strategy model, the perception model includes a backbone visual feature extractor, a predictive attention fusion model and a cognitive attention fusion model, the perception model is used to perceive the adversarial environment model and perform recognition prediction based on the perception results, the strategy model is used to perform action prediction based on the perception results, and control the embodied intelligent body model to move according to the results of the action prediction.

[0083] The modeling of embodied agents that combine perception and action mentioned in this application involves the combination of two different models with specific coupling relationships and functions. The first model is the perception model f(·; θ), which is parameterized by θ and represents the neural circuits involved in visual perception. The second model is the strategy model π(·; φ), which is parameterized by φ and represents the neural circuits involved in visual motor control. This abstraction enables the embodied agent model of this application to effectively perceive its environment and respond appropriately to potential threats.

[0084] Among them, the perception model Responsible for continuously observing multi-view input o 1:t And output the predicted content and internal cognition t , Strategy Model Responsible for outputting actions based on internal cognition a t , driving the embodied agent to collect image inputs from new perspectives o t+1 ,Through action perception and loop prediction, the embodied agent continuously searches for the optimal ,observation perspective, and combines multi-perspective observations, constantly alleviates the hallucination brought by the adversarial ,patch and the information loss caused by occlusion, and finally completes the correct ,prediction.

[0085] Based on the above embodiment, the observation o can be obtained through the API interface t (RGB image) is input to the perception model, and the action a output by the policy model t (displacement and rotation) can also be injected into the API interface to control the camera state t (camera 3D coordinates, as well as roll, pitch, and yaw angles) to enable the agent’s observation and action.

[0086] Furthermore, the perception model contains a predictive attention fusion model f y , a cognitive attention fusion model f b , and the backbone visual feature extractor v shared by the predictive attention fusion model and the cognitive attention fusion model.

[0087] According to an example embodiment, the perception model is overall modeled as follows:

[0088]

[0089] The policy model is an action prediction model that accepts the internal historical knowledge b output from the perception model at each time step. t , output the optimal action a for the current time step t , to guide the embodied agent to the optimal observation perspective.

[0090] According to an example embodiment, the policy model is overall modeled as follows:

[0091] a t =π(b t ;φ).

[0092] According to some embodiments, reference Figure 2 In step S120, a training data set is constructed based on the interaction between the embodied intelligent agent model and the adversarial environment model, which can be specifically implemented through steps S21 to S27.

[0093] In step S21, an observation image data sample of the adversarial environment model at the current observation angle at the current time step is obtained.

[0094] The acquisition of observed image data samples may adopt the API interface based on the above embodiment.

[0095] It should be emphasized that the acquired observed image data samples are image data of the current observation perspective of the adversarial environment model at the current time step.

[0096] In step S22, the observed image data samples are input into the perception model of the embodied agent model to obtain the predicted recognition label samples of the current time step and the internal cognition samples of the current time step.

[0097] The observed image data samples obtained at the current time step are input into the perception model of the embodied agent model, and finally the predicted recognition label samples of the image content are output. and new internal historical cognition t .

[0098] In step S23, the internal cognitive samples of the current time step are input into the strategy model to obtain the optimal action samples of the current time step.

[0099] The internal cognitive sample b of the current time step t Input to the policy model of the embodied agent model and output the optimal action sample a at the current time step t , to guide the embodied intelligent agent to the optimal viewing angle. The implementation of step S22 and step S23 is as follows Figure 4 shown.

[0100] In step S24, the embodied intelligent body model is controlled to move to the next observation perspective according to the optimal action sample.

[0101] In step S25, steps S21-S24 are repeated until a preset time step is reached.

[0102] In some embodiments, the interaction between the embodied agent model and the adversarial environment model is as follows: Figure 3 shown.

[0103] According to the example embodiment, the embodied agent model and the adversarial environment model interact for H steps, that is, H is the maximum limit of the time step, that is, the defense is completed after a maximum of H steps of interaction. H can be set to any number of steps greater than 1. In some embodiments, the settings of H=4 and H=16 are used. Generally speaking, increasing H can introduce additional observation inputs, which can improve the defense capability of the model.

[0104] In step S26, the interaction reward for each observation perspective is calculated using a pre-constructed dense reward function based on the predicted recognition label samples and the true classification label samples of the observed image data.

[0105] Through the interaction between the embodied agent model and the adversarial environment model, the corresponding data set can be collected Then, the reward r for each step is obtained by calculating the dense reward function t Among them, the calculation of the dense reward function involves a perceptual loss function Metrics predict identification label samples The distance from the true classification label sample y; and a decay constant γ, which controls the impact of the time step on the reward. Perceptual loss function This varies from task to task and is not limited in this application.

[0106] The real classification label sample y may be obtained by manual annotation or by annotation using other annotation tools (such as a large language model), and the present invention does not impose any limitation on this.

[0107] According to an example embodiment, the dense reward r t for:

[0108]

[0109] in, Is a perceptual loss function that measures the predicted recognition label samples The distance from the true classification label sample y; γ is a decay constant that controls the impact of the time step on the reward.

[0110] In a specific embodiment, the target recognition task is an object classification task embodiment, Can be set as a cross entropy function.

[0111] In another specific embodiment, the target recognition task is a face recognition task. It can be set as the cosine similarity of a pair of facial features.

[0112] In another specific embodiment, the target recognition task is a target detection task. It can be set as the sum of bounding box regression loss, target loss and classification loss for target detection.

[0113] Furthermore, the decay constant γ is a hyperparameter, and its optimal value can be obtained by performing grid search through multiple training runs. In a specific embodiment, the decay constant γ can be set to 0.95 based on experience.

[0114] In step S27, based on the observed image data samples, internal cognitive samples, optimal action samples and interaction rewards of each observation perspective, the interaction trajectory corresponding to each observation perspective is obtained as training data, thereby constructing a training data set.

[0115] Through the calculation of dense reward function, we can construct the interaction trajectory τ between the agent and the environment, which includes observation (i.e., observation image data samples), cognition (i.e., internal cognition samples), action (i.e., optimal action samples), and reward (i.e., interaction reward). The interaction trajectory τ of each time step is used as training data, and the interaction trajectories τ of all time steps together construct the training data set.

[0116] This application uses multi-perspective training data to construct a training data set. It is understandable that in addition to maliciously designed pattern interference, the adversarial patch itself also causes a certain degree of occlusion on the target. The single-perspective defense in the existing technology cannot obtain information from additional perspectives to compensate for the loss of occluded information, and its defense capability limit is naturally weaker than multi-perspective defense.

[0117] According to some embodiments, reference Figure 4 In step S22, the observed image data samples are input into the perception model of the embodied intelligent agent model to obtain the predicted recognition label samples of the current time step and the internal cognitive samples of the current time step, which can be specifically implemented through steps S410-S430.

[0118] In step S410, the observed image data samples are input into the backbone visual feature extractor to obtain image feature samples.

[0119] The perception model includes the prediction attention fusion model f y , cognitive attention fusion modelf b , and the backbone visual feature extractor v shared by the predictive attention fusion model and the cognitive attention fusion model.

[0120] After the observation image data sample is input, the perception model receives the observation image data sample from the current observation angle o t , extract the feature samples of the image through the backbone visual feature extractor t .

[0121] According to an example embodiment, the backbone visual feature extractor may be composed of a visual model pre-trained on the original task. The pre-trained visual model itself has good feature extraction capabilities for normal samples. Since the method of the present application relies on multi-view fusion of perception and action to defend against adversarial patches, rather than modifying the visual backbone itself, during the training of the perception model f(·; θ), the backbone visual feature extractor v can be frozen, and only the predictive attention fusion model f can be trained. y , and cognitive attention fusion model f b , reducing the number of parameters to achieve efficient training.

[0122] On this basis, the image feature sample e t It is the feature vector before the last few layers (such as the fully connected layer) of the visual model, which contains the high-dimensional features of the image extracted by the visual model.

[0123] In a specific embodiment, the target recognition task is an object classification task, and the visual model may be an image classification model such as ResNet, SWIN, ViT, etc. pre-trained on the ImageNet dataset; in a specific embodiment, the target recognition task is a face recognition task, and the visual model may be an iResNet model pre-trained on the LFW face dataset; in a specific embodiment, the target recognition task is a target detection task, and the visual model may be MaskRCNN or YOLO, etc. pre-trained on the COCO dataset.

[0124] In step S420, the image feature samples are input into the predictive attention fusion model, and the predicted recognition label samples of the current time step are obtained according to the internal cognitive samples and the image feature samples of the previous time step.

[0125] Image feature sampling via predictive attention fusion model t With internal cognitive sample b t-1 Fusion, based on the fusion features, output the predicted recognition label sample

[0126] It needs to be explained that the internal cognitive sample can be constructed as a sample of image features e t Eigenvectors of the same size.

[0127] According to an example embodiment, a predictive attention fusion model may be constructed by a Decision Transformer to fuse internal cognitive and temporal feature sequences.

[0128] In step S430, the image feature samples are input into the cognitive attention fusion model, and the internal cognitive samples of the current time step are obtained according to the internal cognitive samples of the previous time step and the image feature samples.

[0129] Image feature sampling via cognitive attention fusion model t With internal cognitive sample b t-1 According to the fusion characteristics, the internal cognitive sample b of the current time step is output t .

[0130] According to an example embodiment, a cognitive attention fusion model may be constructed by a Decision Transformer to fuse internal cognitive and temporal feature sequences.

[0131] On this basis, the policy model can be constructed by a shallow fully connected neural network as the head of DecisionTransformer.

[0132] Based on the above embodiment, during the training process, based on the dense reward function, reinforcement learning technology is used to optimize the parameters θ and φ, so as to determine the optimal parameters θ and φ of the perception model f(·; θ) and the strategy model π(·; φ).

[0133] In the reinforcement learning process, the present application explores the objective function by accumulating information, and uses the determined preset objective function to determine the optimal parameters θ and φ for achieving the task.

[0134] According to some embodiments, the preset objective function includes:

[0135]

[0136] Among them, θ is the optimal parameter of the perception model f(·; θ), φ is the optimal parameter of the policy model π(·; φ); It is expectation abbreviation of; is the true classification label sample of scene x; τ is the The interaction trajectory obtained by the policy model π has the following form τ = (o1′, b1, a1, r1, …, o′ t ,b t ,a t ,r t ); p represents the adversarial patch; represents the set of all adversarial patches under scenario x; A is the application function of the adversarial patch; Is the perceptual loss function, used to measure the prediction of the recognition label sample The distance from the true classification label sample y; H represents the maximum limit of the time step; is the conditional entropy, which is used to quantify the internal cognition sample b H-1 and observed image data sample o′ H Described under the conditions The amount of information required; λ is a hyperparameter constant used to control the degree of influence of conditional entropy, a t is the optimal action at the current time step, s t is the observation angle at time step t, b t is the internal cognitive sample at the current time step t; o t is the observed image data sample without adversarial patch interference, o t ′=A(o t ,p;s t ) is the observed image data sample with adversarial patch interference, is the state transition probability of scene x, is the observation function.

[0137] In the training process of step S130, this embodiment can minimize the number of samples of the recognition label predicted by the perception model by minimizing the preset objective function. The difference between the actual classification label sample y and the conditional entropy is used as a regular term to encourage the policy model to output the optimal action sequence a 1:t , to obtain the observation sequence o that minimizes the conditional entropy 1:t In addition, by introducing the cumulative long-term multi-step concept H as the preset end condition, the disadvantage of only considering the single-step greedy strategy that it is easy to fall into the local suboptimal solution is avoided, and it can better converge to the global optimal solution.

[0138] Furthermore, the cumulative information exploration objective function (i.e., the preset objective function) includes the embodied agent and the environment model Direct optimization of such an objective function involves a technique of backpropagation through time to solve the gradient of the objective function with respect to the optimization parameters θ and φ. This requires the environment model It has a differentiable property, that is, it can be used to observe o t and action a t At the same time, back propagation along time requires chained derivation along time step t = H until t = 1, which requires expensive computing resources and faces optimization problems such as gradient vanishing.

[0139] Based on the above embodiment, the preset objective function is further optimized to eliminate the dependence on the differentiable properties of the environment, and there is no need for back propagation along time, so the calculation is efficient.

[0140] According to some embodiments, reference Figure 5 In step S130, the embodied intelligent body model is trained based on the training data set using a preset objective function, which can be specifically implemented through steps S510 to S530.

[0141] In step S510, data of a preset size is sampled from the training data set to obtain sampled data of multiple training batches.

[0142] Using the collected training dataset The perception model f(·; θ) and the policy model π(·; φ) are trained in multiple batches, each batch is from The sample data K τ ,until All data has been trained.

[0143] According to an example embodiment, K τ The size can be set to 64, 128, 256, etc. τ The larger it is, the more stable the training effect will be, but the more computing resources will be occupied.

[0144] In step S520, for each training batch of sampled data, the pruned proximal policy gradient objective function and the perceptual objective function are calculated.

[0145] For each K τ , calculate the clipped proximal policy gradient objective function and the perception objective function.

[0146] In step S530, gradient ascent is performed according to the clipped proximal policy gradient objective function to update the model parameters of the policy model, and gradient descent is performed according to the perception objective function to update the model parameters of the perception model until all sampled data are traversed to obtain the target embodied intelligent body model.

[0147] By reverse derivation, gradient ascent and descent are performed to update the parameters φ and θ to maximize the clipping proximal policy gradient objective function and minimize the perception objective function.

[0148] According to an example embodiment, reverse derivation and parameter update can be implemented by a differentiable automatic derivation programming framework, such as Pytorch and TensorFlow. In addition, the above maximization and minimization are performed N times to make full use of the data. In a specific embodiment, N can be set to 2.

[0149] According to an example embodiment, data sampling and training are performed M times in total, and M can be initially set to a larger value, such as 10000. Empirically, the optimal value of M is the number of iterations corresponding to the convergence of the objective function clipping proximal policy gradient objective function and the perception objective function curve.

[0150] Repeat the training until The training is completed when all the data have been trained.

[0151] This embodiment uses reinforcement learning technology for non-differentiable environments to solve the problem of unstable back propagation calculation along time during multi-step model optimization. It does not require the differentiable nature of the environment and is consistent with the characteristics of most simulation environments and the real three-dimensional world.

[0152] According to some embodiments, the proximal policy gradient objective function is clipped as:

[0153]

[0154] In the formula, To tailor the proximal policy gradient objective function, φ is the optimal parameter of the policy model π(·; φ); clip is the clipping function; a t is the optimal action at the current time step; b t is the internal cognitive sample at the current time step t; ∈ is the clipping range; is the advantage function; is the sampling data of the τth training batch.

[0155] It should be noted that the advantage function Used to estimate the advantage value, which is in the form of the difference between the action value function and the state value function. Its specific selection is related to the target recognition task and this application does not impose any restrictions on this.

[0156] This embodiment uses the clipping function clip to limit the difference between the current learning strategy and the actual interaction strategy to not be too large, thereby achieving a stable update of the parameter φ. The clipping range ∈ is usually set to a small value, such as 0.2.

[0157] According to some embodiments, the perceptual objective function is:

[0158]

[0159] In the formula, is the perception objective function; is the perceptual loss function, is the true classification label sample of scene x, o t ′ is the observed image data sample with adversarial patch interference; λ is a hyperparameter constant that controls the degree of influence of conditional entropy; is the conditional entropy, which is used to quantify the internal cognition sample b H-1 and observed image data sample o′ H Described under the conditions The amount of information required.

[0160] It should be emphasized that The losses involved Its choice is related to the reward function Be consistent.

[0161] Based on the above embodiment, in order to explain step S130 in more detail, a specific training embodiment is given, and the implementation process is shown in Table 1.

[0162] Table 1. Training steps of the target embodied agent model

[0163]

[0164] The embodiments of the present application achieve minimization of target uncertainty in multiple steps, thereby training the embodied intelligent agent to learn strategic actions to collect multi-view image sequences that minimize target uncertainty and obtain optimal prediction labels.

[0165] Based on the above embodiments, it can be understood that to solve the optimization problem There are multiple feasible solutions to the optimization problem when generating adversarial patches. Therefore, for the same scenario, different optimization methods, that is, different adversarial attack techniques, will generate different adversarial patches. In addition, the above calculation process is complicated, and the generated adversarial patches are highly aggressive, which can be used as the worst case to evaluate the defense effect of the model. During the training process, the following approximate method can be used to generate adversarial patches to save computing resources, improve standard accuracy, and increase defense capabilities.

[0166] According to some embodiments, in step S210, an adversarial patch is added to the environment model to obtain an adversarial environment model, which can be specifically implemented through step S610 or step S620.

[0167] In step S610, uniform noise sampling is performed on the superset to obtain an adversarial patch, which is added to the environment model to obtain an adversarial environment model.

[0168] like Figure 6 As shown, the approximation method of the superset adversarial patch mentioned in this step uses the following method of uniform noise sampling in the superset to approximate the adversarial patch:

[0169]

[0170] Among them, U represents uniform distribution, that is, random noise is used to approximate the adversarial patch.

[0171] The method of approximating adversarial patches involved in this step does not make any assumptions about the adversary's adversarial attack technology and covers all possible adversarial patches.

[0172] The generated adversarial patch is added to the environment model through the application function of the adversarial patch to obtain the adversarial environment model.

[0173] It should be noted that the technology of introducing random noise in adversarial training already exists, but the previous technology only used a single-view training method and could not achieve satisfactory defense effects. The present embodiment combines the approximation method of the superset adversarial patch with the modeling of an embodied intelligent agent that combines perception and action, the cumulative information exploration objective function, and the optimization method based on clipping the proximal policy gradient. It not only reduces the computational complexity, but also eliminates the prior assumptions about adversarial attack technology, thereby achieving the generalization of defense capabilities that are unrelated to adversarial attack technology.

[0174] In step S620, for each observation perspective, the perceptual loss of the perceptual model is maximized to obtain an adversarial patch, which is added to the environment model to obtain an adversarial environment model, wherein the perceptual loss is used to measure the difference between the predicted recognition label sample and the true classification label sample at the current observation perspective.

[0175] like Figure 6 As shown, this step involves an offline adversarial patch approximation method for further improving the approximation effect of the adversarial patch. The adversarial patch generated by the method involved in step S620 is more aggressive, which can increase learning efficiency and enhance the defense capability of the trained model.

[0176] In the specific implementation process, by optimizing the perceptual loss, that is, To generate adversarial patches, for each viewing state s t , only optimize the observation of the adversarial patch for the current view t The adversarial nature of the proposed method eliminates the expectation for multiple angles, and the generation of the adversarial patch is only performed once offline.

[0177] Since existing adversarial training techniques require online regeneration of new adversarial patches at each step of model training, the adversarial patch approximation method provided in this embodiment has a higher adversarial patch sampling efficiency and is much more computationally efficient than existing adversarial training techniques.

[0178] The generated adversarial patch is added to the environment model through the application function of the adversarial patch to obtain the adversarial environment model.

[0179] Based on the above embodiment, in order to explain step S620 in more detail, a specific training embodiment is given, and the implementation process is shown in Table 2.

[0180] Table 2. Generation steps of adversarial patches

[0181]

[0182] It should be noted that the perceptual loss function The task varies, and this application does not limit this. In a specific embodiment, the target recognition task is an object classification task. It can be set as a cross entropy function; in another specific embodiment, the target recognition task is a face recognition task, It can be set to cosine similarity; in another specific embodiment, the target recognition task is a target detection embodiment, according to the attacker's intention, It can be set to classification loss (making the target recognition category wrong) or target loss (making the object undetectable).

[0183] In addition, the number of iterations M can be set based on experience, for example, set to 30.

[0184] The existing adversarial training technology introduces adversarial patches during the training process. Although it strengthens the model's defense against adversarial patches, it also interferes with the learning of normal samples, resulting in a decrease in standard accuracy. The existing input preprocessing technology introduces image preprocessing (such as erasing some areas), which often interferes with the recognition accuracy of normal samples. This embodiment does not modify the original good visual feature extraction backbone, nor does it modify the input image itself, but introduces additional perspective observations. The new perspective introduces additional information and strengthens the model's predictive ability. The purpose of this embodiment is to help the model learn the key features required to perceive actions so that it can defend against more powerful attacks when applied. When evaluating the defense model of this case, adversarial samples generated by the most powerful existing attack technology are still used to evaluate the defense of this case. That is, training with a weaker opponent to obtain defense capabilities against a stronger opponent.

[0185] According to some embodiments, reference Figure 7 In step S140, based on the target embodied intelligent agent model, the target to be predicted of the target recognition task is predicted to obtain a prediction result, which can be specifically implemented through steps S710 to S730.

[0186] In step S710, observation image data of the target to be predicted at the current observation angle is obtained.

[0187] The acquisition of observed image data may adopt the API interface based on the above embodiment.

[0188] It should be emphasized that the acquired observed image data is the image data of the current observation angle of the target to be predicted at the current time step.

[0189] In step S720, the observed image data is input into the target embodied agent model to obtain the predicted recognition label and the optimal action for the current time step.

[0190] The observed image data is input into the trained target embodied agent model, the perception model outputs the predicted identification label and new internal historical cognition of the observed image data, and the strategy model outputs the optimal action for the current time step.

[0191] In step S730, the embodied intelligent body model is controlled to move to the next observation perspective according to the optimal action to obtain new observation image data until a preset time step is reached, and the final predicted identification label is output as the prediction result.

[0192] The embodied agent model builds internal cognition through continuous observation and further collects new perspective environment information through strategic actions, thereby continuously optimizing the prediction of the target. Although the model may make incorrect predictions due to the influence of adversarial patches in the early stages of interaction, as the interaction proceeds, the model continuously reduces its cognitive uncertainty and eventually outputs the correct prediction.

[0193] Through motion perception and cycle prediction, this application successfully overcomes the hallucinations caused by adversarial patches and the loss of key information caused by occlusion, which not only reduces the attack success rate of adversarial patches, but also improves the standard accuracy for normal samples.

[0194] In order to further illustrate the identification method based on embodied active defense provided by the present application, several specific embodiments are given. Figure 8 , 9 10 respectively show the defensive effects of the recognition method based on embodied active defense provided by the present application in the embodiments of object classification, face recognition, and target detection. Figure 8 The flowchart of object classification defense under the attack of patch stickers generated by two adversarial attack techniques, MeshAdv and N attack, is shown. Although hats and mangoes are misled into other categories in the initial observation, as the interaction proceeds, the embodied adversarial defense method that combines perception and action successfully corrects its predictions and finally obtains the correct prediction results with increased confidence. Fig. 9 The flow chart of face recognition defense under the attack of adversarial glasses generated by Face3DAdv adversarial attack technology is shown. The goal of the adversarial attack is to make the detector think that the face does not match the face in the database, so as to achieve the effect of evading recognition. At the initial perspective, the cosine similarity between the target face features and the face features in the database is very low, but with the interaction of the embodied adversarial defense method combining perception and action, the cosine similarity between the target face features predicted by the model and the face features in the database gradually increases, and finally successfully recognizes its identity. Fig.10The flowchart of object detection defense under paint attack generated by SIB and AdvCaT adversarial attack technology is shown. Under initial observation, the attack of adversarial patches makes the vehicle unrecognizable by the detection camera, but as the defense progresses, the detector successfully finds the "invisible" vehicle and continuously improves the accuracy of its bounding box and the confidence of the prediction.

[0195] The following describes an apparatus embodiment of the present application, which can be used to execute the method embodiment of the present application. For details not disclosed in the apparatus embodiment of the present application, reference can be made to the method embodiment of the present application.

[0196] Fig.11 A block diagram of an identification device based on embodied active defense according to an exemplary embodiment is shown.

[0197] Fig.11 The device shown can execute the aforementioned identification method based on embodied active defense according to the embodiment of the present application.

[0198] like Fig.11 As shown, the identification device based on embodied active defense may include:

[0199] See also Fig.11 And referring to the previous description, the model building module 1110 is used to build an adversarial environment model and an embodied intelligent agent model according to the target recognition task.

[0200] The data set construction module 1120 is used to construct a training data set based on the interaction between the embodied intelligent agent model and the adversarial environment model.

[0201] The model training module 1130 is used to train the embodied intelligent agent model based on the training data set using a preset objective function.

[0202] The loop iteration module 1140 is used to repeat steps S2-S3 until a preset end condition is reached to obtain a target embodied intelligent agent model;

[0203] The recognition prediction module 1150 is used to predict the target to be predicted of the target recognition task based on the target embodied intelligent agent model to obtain a prediction result.

[0204] The device performs functions similar to the method provided above. For other functions, please refer to the previous description and will not be repeated here.

[0205] Fig.12 An electronic device according to an exemplary embodiment of the present application is shown. Fig.12 1200 according to this embodiment of the present application is described. Fig. 9 The electronic device 1200 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0206] like Fig.12 As shown, the electronic device 1200 is in the form of a general computing device. The components of the electronic device 1200 may include, but are not limited to: at least one processing unit 1210, at least one storage unit 1220, a bus 1230 connecting different system components (including the storage unit 1220 and the processing unit 1210), a display unit 1240, etc.

[0207] The storage unit stores program codes, which can be executed by the processing unit 1210, so that the processing unit 1210 executes the methods described in this specification according to various exemplary embodiments of the present application. For example, the processing unit 1210 can execute the method described above.

[0208] The storage unit 1220 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 12201 and / or a cache storage unit 12202 , and may further include a read-only storage unit (ROM) 12203 .

[0209] The storage unit 1220 may also include a program / utility 12204 having a set (at least one) of program modules 12205, such program modules 12205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0210] Bus 1230 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0211] The electronic device 1200 may also communicate with one or more external devices 300 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 1200, and / or communicate with any device that enables the electronic device 1200 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 1250. In addition, the electronic device 1200 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 1260. The network adapter 1260 may communicate with other modules of the electronic device 1200 via the bus 1230. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1200, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0212] Through the description of the above implementation methods, it is easy for those skilled in the art to understand that the example implementation methods described here can be implemented by software or by combining software with necessary hardware. The technical solution according to the implementation method of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above method according to the implementation method of the present application.

[0213] The software product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0214] Computer readable storage media may include data signals propagated in baseband or as part of a carrier wave, wherein readable program codes are carried. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or device. The program codes contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.

[0215] Program code for performing the operations of the present application may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0216] The computer-readable medium carries one or more programs. When the one or more programs are executed by a device, the computer-readable medium implements the aforementioned functions.

[0217] Those skilled in the art will appreciate that the above modules can be distributed in the device according to the description of the embodiment, or can be changed accordingly and only used in one or more devices different from the embodiment. The modules of the above embodiments can be combined into one module, or further divided into multiple sub-modules.

[0218] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the embodiment of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiment of the present application.

[0219] The exemplary embodiments of the present application are specifically shown and described above. It should be understood that the present application is not limited to the detailed structures, configurations or implementations described herein; on the contrary, the present application is intended to cover various modifications and equivalent configurations included in the spirit and scope of the appended claims.

Claims

1. A recognition method based on embodied active defense, characterized in that: include: S1: Based on the target recognition task, establish the adversarial environment model and embodied agent model; S2: constructing a training data set based on the interaction between the embodied agent model and the adversarial environment model; S3: Based on the training data set, training the embodied intelligent agent model using a preset objective function; S4: Repeat steps S2-S3 until the preset end condition is reached, and the target embodied intelligent agent model is obtained; S5: Based on the target embodied intelligent agent model, predict the target to be predicted of the target recognition task to obtain a prediction result.

2. The method according to claim 1, characterized in that The adversarial environment model and the embodied agent model are established according to the target recognition task, including: Using a partially observable Markov decision process, constructing an environment model related to the target recognition task, and adding an adversarial patch to the environment model to obtain the adversarial environment model; Based on neural networks and target recognition tasks, an embodied intelligent body model is constructed; wherein the embodied intelligent body model includes a perception model and a strategy model, the perception model includes a backbone visual feature extractor, a predictive attention fusion model and a cognitive attention fusion model, the perception model is used to perceive the adversarial environment model and perform recognition prediction based on the perception results, the strategy model is used to perform action prediction based on the perception results, and control the embodied intelligent body model to move according to the results of the action prediction.

3. The method according to claim 2, characterized in that The constructing of a training data set based on the interaction between the embodied intelligent agent model and the adversarial environment model comprises: S21: Obtaining an observation image data sample of the current observation angle of the adversarial environment model at the current time step; S22: Inputting the observed image data sample into the perception model of the embodied intelligent agent model to obtain the predicted recognition label sample of the current time step and the internal cognition sample of the current time step; S23: inputting the internal cognitive sample of the current time step into the strategy model to obtain the optimal action sample of the current time step; S24: controlling the embodied intelligent body model to move to the next observation angle according to the optimal action sample; S25: Repeat steps S21-S24 until a preset time step is reached; S26: Calculate the interaction reward of each observation perspective using a pre-constructed dense reward function according to the predicted recognition label sample and the real classification label sample of the observed image data; S27: According to the observed image data samples, the internal cognition samples, the optimal action samples and the interaction rewards of each observation perspective, the interaction trajectory corresponding to each observation perspective is obtained as training data, thereby constructing a training data set.

4. The method according to claim 3, characterized in that The step of inputting the observed image data sample into the perception model of the embodied agent model to obtain the predicted recognition label sample of the current time step and the internal cognition sample of the current time step includes: Inputting the observed image data sample into the backbone visual feature extractor to obtain an image feature sample; Inputting the image feature sample into the predictive attention fusion model, and obtaining the predicted recognition label sample of the current time step according to the internal cognitive sample of the previous time step and the image feature sample; The image feature samples are input into the cognitive attention fusion model, and the internal cognitive samples of the current time step are obtained according to the internal cognitive samples of the previous time step and the image feature samples.

5. The method according to claim 3, characterized in that: The preset objective function includes: Among them, θ is the optimal parameter of the perception model f(·; θ), φ is the optimal parameter of the policy model π(·; φ); It is expectation abbreviation of; is the true classification label sample of scene x; τ is the The interaction trajectory obtained by the policy model π in ; p represents the adversarial patch; represents the set of all adversarial patches under scenario x; A is the application function of the adversarial patch; Is the perceptual loss function, used to measure the prediction of the recognition label sample The distance from the true classification label sample y; H represents the maximum limit of the time step; is the conditional entropy, which is used to quantify the internal cognition sample b H-1 and observed image data sample o′ H Described under the conditions The amount of information required; λ is a hyperparameter constant that controls the degree of influence of conditional entropy, a t is the optimal action at the current time step, s t is the observation angle at time step t, b t is the internal cognitive sample at the current time step t; o t is the observed image data sample without adversarial patch interference, o′ t =A(o t ,p;s t ) is the observed image data sample with adversarial patch interference, is the state transition probability of scene x, is the observation function.

6. The method according to claim 3, characterized in that The step of training the embodied agent model based on the training data set and using a preset objective function includes: Sampling data of a preset size from the training data set to obtain a plurality of training batches of sampled data; For each of the sampled data of the training batch, calculating a trimmed proximal policy gradient objective function and a perception objective function; Perform gradient ascent according to the trimmed proximal policy gradient objective function to update the model parameters of the policy model, perform gradient descent according to the perception objective function to update the model parameters of the perception model, until all sampled data are traversed to obtain the target embodied intelligent body model.

7. The method according to claim 6, characterized in that The clipped proximal policy gradient objective function is: In the formula, To tailor the proximal policy gradient objective function, φ is the optimal parameter of the policy model π(·; φ); clip is the clipping function; a t is the optimal action at the current time step; b t is the internal cognitive sample at the current time step t; ∈ is the clipping range; is the advantage function; is the sampling data of the τth training batch.

8. The method according to claim 6, characterized in that The perception objective function is: In the formula, is the perception objective function; is the perceptual loss function, is the true classification label sample of scene x, o′ t is an observed image data sample with adversarial patch interference; λ is a hyperparameter constant that controls the degree of influence of conditional entropy; is the conditional entropy, which is used to quantify the internal cognition sample b H-1 and observed image data sample o′ H Described under the conditions The amount of information required.

9. The method according to claim 3, characterized in that: Adding an adversarial patch to the environment model to obtain the adversarial environment model includes: Perform uniform noise sampling on the superset to obtain the adversarial patch, and add it to the environment model to obtain the adversarial environment model; or For each of the observation perspectives, the perceptual loss of the perceptual model is maximized to obtain the adversarial patch, and added to the environment model to obtain the adversarial environment model, wherein the perceptual loss is used to measure the difference between the predicted recognition label sample and the true classification label sample at the current observation perspective.

10. The method according to claim 1, characterized in that The step of predicting the target to be predicted of the target recognition task based on the target embodied intelligent agent model to obtain a prediction result includes: Obtaining observation image data of the target to be predicted at the current observation angle; Inputting the observed image data into the target embodied agent model to obtain a predicted recognition label and an optimal action at the current time step; According to the optimal action, the embodied intelligent body model is controlled to move to the next observation perspective to obtain new observation image data until a preset time step is reached, and a final predicted identification label is output as a prediction result.

11. An identification device based on embodied active defense, characterized in that: include: The model building module is used to build the adversarial environment model and the embodied agent model according to the target recognition task; A data set construction module, for constructing a training data set based on the interaction between the embodied intelligent agent model and the adversarial environment model; A model training module, used to train the embodied agent model based on the training data set using a preset objective function; A loop iteration module, used to repeat steps S2-S3 until a preset end condition is reached to obtain a target embodied intelligent agent model; The recognition prediction module is used to predict the target to be predicted of the target recognition task based on the target embodied intelligent agent model to obtain a prediction result.

12. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that: A computer program or instruction is stored thereon, and when the computer program or instruction is executed by a processor, the method according to any one of claims 1 to 10 is implemented.

Citation Information

Cited By

  • Enhanced adversarial learning-based self-owned intelligent model-oriented answer rejection method

    CN121543085A

  • Control method of intelligent body with body, model training method, equipment and storage medium

    CN121625165A