Physical patch generation method, device and equipment based on diffusion model

By generating physical patches that are highly integrated with the environment based on diffusion model, the problems of visual unnaturalness and attributes in the prior art are solved, and higher concealment and stability are achieved, and the attack effect is enhanced.

CN120339060APending Publication Date: 2025-07-18BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510204057.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The physical patches generated by the prior art appear unnatural visually, easily recognized, and the physical properties of the patch are not fully considered, resulting in unstable effects under different environments and lighting conditions, affecting the concealment and effectiveness of the attack.

Method used

Using a diffusion model-based method, initial physical patches are generated through forward diffusion and reverse diffusion, and patch attributes are optimized in combination with differential dynamic models to ensure consistency between the patch and the environment background. The diffusion model is used to generate physical patches that are highly integrated with the environment.

Benefits of technology

The generated physical patches are more in line with the environment, improving the concealment and stability of the attack, reducing the risk of being detected and identified, and enhancing the effectiveness of the attack.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339060A_ABST
    Figure CN120339060A_ABST
Patent Text Reader

Abstract

The invention provides a physical patch generation method, device and equipment based on a diffusion model, and the method comprises the steps: obtaining an original background image, carrying out the forward diffusion of the original background image through a diffusion model, and obtaining a noise-added background image; performing reverse diffusion processing on the noise-added background image by using a diffusion model to obtain an initial physical patch; adding the initial physical patch into the original background image to obtain a target background image; and obtaining an initial state and a preset action of the intelligent agent in the target background image, inputting the initial state and the preset action into the differential dynamic model, and outputting the target physical patch. The original background of the environment is used as the initial state of the physical patch, so that the patch can be ensured to be more consistent with the environment background in the generation process, and the abrupt feeling possibly caused by random disturbance in a traditional method is avoided. The initial physical patch is optimized through the differential dynamic model, so that the obtained physical patch is more fit with the original environment and can be more hidden into the environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and particularly to a method, apparatus, and device for generating physical patches based on a diffusion model. Background Art

[0002] The security of deep reinforcement learning (DRL) systems is threatened by adversarial attacks, where attackers interfere with the decision-making process of agents through carefully designed strategies, leading to a significant decline in their performance in the environment. Similar to adversarial attacks that add tiny perturbations to deep neural networks (DNNs) to deceive classification models, adversarial attacks in deep reinforcement learning also add perturbations to the environmental state or the agent's behavior to mislead the agent into making wrong decisions, which are usually referred to as "adversarial strategies" or "adversarial behaviors".

[0003] Currently, adversarial physical patches are mainly generated based on the way of iteratively optimizing on randomly initialized initial physical patches. However, the generated physical patches may appear unnatural and too obtrusive visually, and are easily recognizable. Summary of the Invention

[0004] In view of this, the purpose of the present disclosure is to propose a method, apparatus, and device for generating physical patches based on a diffusion model to solve or partially solve the above problems.

[0005] Based on the above purpose, the first aspect of the present disclosure provides a method for generating physical patches based on a diffusion model, the method comprising:

[0006] Obtain an original background image, and input the original background image into a pre-trained diffusion model;

[0007] Perform forward diffusion on the original background image by using the diffusion model to obtain a noisy background image;

[0008] Perform reverse diffusion processing on the noisy background image by using the diffusion model to obtain an initial physical patch;

[0009] Add the initial physical patch to the original background image to obtain a target background image;

[0010] Obtain the initial state and a preset action of an agent in the target background image, input the initial state and the preset action into a pre-trained differential dynamic model, and output a target physical patch after being processed by the differential dynamic model.

[0011] Based on the same inventive concept, the second aspect of the present disclosure proposes a device for generating physical patches based on a diffusion model, comprising:

[0012] A data acquisition module, configured to acquire an original background image and input the original background image into a pre-trained diffusion model;

[0013] A forward diffusion module, configured to perform forward diffusion on the original background image by using the diffusion model to obtain a noisy background image;

[0014] A reverse diffusion module, configured to perform reverse diffusion processing on the noisy background image by using the diffusion model to obtain an initial physical patch;

[0015] A target background image determination module, configured to add the initial physical patch to the original background image to obtain a target background image;

[0016] A target physical patch determination module, configured to acquire an initial state and a preset action of an agent in the target background image, input the initial state and the preset action into a pre-trained differential dynamics model, and output a target physical patch after being processed by the differential dynamics model.

[0017] Based on the same inventive concept, a third aspect of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable by the processor. When the processor executes the computer program, the method for generating a physical patch based on a diffusion model as described above is implemented.

[0018] Based on the same inventive concept, a fourth aspect of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method for generating a physical patch based on a diffusion model as described above.

[0019] As can be seen from the above, the present disclosure proposes a physical patch generation method, apparatus, and device based on a diffusion model. An original background image is obtained, and the original background image is input into a pre-trained diffusion model. The diffusion model is used to perform forward diffusion on the original background image to obtain a noisy background image. The diffusion model is used to perform reverse diffusion processing on the noisy background image to obtain an initial physical patch. Using the original background of the environment as the initial state of the physical patch can ensure that the patch is more consistent with the environmental background during the generation process, avoiding the abruptness that may be caused by random perturbations in traditional methods. The initial physical patch is added to the original background image to obtain a target background image. The initial state and a preset action of the agent in the target background image are obtained, and the initial state and the preset action are input into a pre-trained differential dynamic model. After being processed by the differential dynamic model, a target physical patch is output. By optimizing the initial physical patch through the differential dynamic model, the obtained physical patch fits better with the original environment and can be more concealed in the environment, thereby improving the concealment of the attack and reducing the risk of being detected and recognized. Description of the Drawings

[0020] To more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following description are only embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 It is a flowchart of the physical patch generation method based on the diffusion model according to the embodiment of the present disclosure;

[0022] Figure 2 It is a structural block diagram of the physical patch generation apparatus based on the diffusion model according to the embodiment of the present disclosure;

[0023] Figure 3 It is a schematic structural diagram of the electronic device according to the embodiment of the present disclosure. Detailed Embodiments

[0024] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the following further details the present disclosure with reference to specific embodiments and the accompanying drawings.

[0025] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the ordinary meanings understood by those of ordinary skill in the field to which the present disclosure belongs. The "first", "second" and similar terms used in the embodiments of the present disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "linked" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0026] The following are the explanations of the terms related to the present disclosure:

[0027] VAE: Variational Auto-Encoders (VAE), as a form of deep generative model, is a generative network structure based on variational Bayes (VB) inference proposed by Kingma et al. in 2014.

[0028] MD-RNN: Multi-Dimensional Recurrent Neural Networks (MD-RNN) extends the standard RNN and is suitable for processing sequence data in two dimensions and higher dimensions, such as images and videos.

[0029] The security of deep reinforcement learning (DRL) systems is threatened by adversarial attacks, where attackers interfere with the decision-making process of agents through carefully designed strategies, resulting in a significant decline in their performance in the environment. Similar to adversarial attacks that add small perturbations to deep neural networks (DNNs) to deceive classification models, adversarial attacks in deep reinforcement learning also add perturbations to the environmental state or the agent's behavior, thus misleading the agent into making wrong decisions, which are usually called "adversarial strategies" or "adversarial behaviors". These perturbations can significantly reduce the learning effect or decision-making quality of agents in reinforcement learning tasks. Adversarial attacks in deep reinforcement learning can be classified into white-box attacks and black-box attacks according to the attacker's knowledge of the target model.

[0030] In white-box attacks, the attacker has full access to the target model and can obtain the model's structure, parameters, and other relevant information. In this case, the attacker can generate adversarial strategies by directly leveraging gradient information or other optimization techniques to precisely control the behavior of the agent, thereby achieving the purpose of deception. The attacker can use policy gradient methods, Q-value functions, or value iteration methods to perform precise perturbation design and create adversarial strategies that can effectively deceive the agent.

[0031] In black-box attacks, the attacker cannot directly obtain the internal information of the target model and can only obtain the input-output correspondence through interaction with the environment. In state-based black-box attacks, the attacker can only observe the state input of the agent and the action selection of the model output, and relies on multiple interactions with the environment to gradually approximate the behavior pattern of the model. To generate adversarial samples, the attacker usually estimates the gradient and updates the policy by exploring the action distribution of the model output and combining zero-order optimization algorithms (such as the finite difference method), thereby forcing the agent to execute incorrect actions.

[0032] For state-based black-box attack methods, although the attacker cannot directly calculate the gradient of the model, by querying the output of the model (such as each action selection), some information about the agent's decision-making process can be indirectly inferred. Through a large number of state-action pair queries, the attacker can gradually approximate the behavior of the target model and generate adversarial strategies through random search or optimization methods for perturbations. State-based black-box attacks can use heuristic methods to estimate the change in the policy and design adversarial strategies based on these estimates.

[0033] Physical Patch Attack (PPA) mainly studies the security issues of deep reinforcement learning in the field of autonomous driving and conducts targeted attacks through visual patterns in the physical environment. In the field of adversarial attacks on deep reinforcement learning systems, generating realistic physical patches to mislead the behavior of the agent is a key challenge. An ideal physical patch should blend naturally with the environmental background to improve the concealment and effectiveness of the attack.

[0034] Existing Targeted Physical Adversarial Attacks mainly generate adversarial physical patches based on iterative optimization on randomly initialized initial physical patches. Although existing methods aim to generate adversarial samples through visual patterns on physical objects, the generated physical patches may appear unnatural visually, which limits their concealment and effectiveness in practical applications. Especially in situations where the physical patch needs to blend in with the surrounding environment without attracting attention, the patches generated by existing methods may be too obtrusive and easily recognizable.

[0035] At the same time, when generating physical patches, the existing methods do not fully consider the physical properties of the patches, such as size and rotation angle, which are crucial for the actual effect and concealment of the patches. The lack of fine-tuning of these physical properties results in the inability of the patches to achieve the expected attack effect during actual deployment, and the effect is unstable under different environments and lighting conditions.

[0036] Based on the above description, this embodiment proposes a physical patch generation method based on a diffusion model, as Figure 1 shown, the method includes:

[0037] Step 101, obtain the original background image, and input the original background image into a pre-trained diffusion model.

[0038] Step 102, perform forward diffusion on the original background image using the diffusion model to obtain a noisy background image.

[0039] Step 103, perform reverse diffusion processing on the noisy background image using the diffusion model to obtain an initial physical patch.

[0040] Step 104, add the initial physical patch to the original background image to obtain a target background image.

[0041] Step 105, obtain the initial state and preset actions of the agent in the target background image, input the initial state and the preset actions into a pre-trained differential dynamic model, and after being processed by the differential dynamic model, output the target physical patch.

[0042] Specifically, when implementing, obtain the original background image, and input the original background image into a pre-trained diffusion model. The diffusion model is a deep generative model based on a probabilistic generative framework, and its core idea is to approximate the data distribution by gradually adding noise (forward diffusion process) and reverse denoising (reverse generation process).

[0043] Perform forward diffusion on the original background image using the diffusion model, that is, perform noise addition processing on the original background image to obtain a noisy background image.

[0044] Perform reverse diffusion processing on the noisy background image using the diffusion model to obtain an initial physical patch. In this embodiment, the physical patch is in the form of an image. The specific process of determining the initial physical patch includes:

[0045] Sample initial noise from a standard Gaussian distribution as the starting point of the reverse process. To improve the generation diversity, latent space interpolation technology is used to perform a linear combination of multiple noise samples.

[0046] Iteratively execute the reverse process to gradually generate a patched image Δo that blends with the environmental background. At each step t, input the current noisy image x t and the time-step embedding into the denoising network to predict the noise ∈ θ (x t , t). Calculate x t-1 according to the reverse update formula, and impose constraint conditions (such as truncating pixel values to [0, 1]) to ensure the physical feasibility of the generated image. After generating the image x t-1 at each step, it needs to be input into the reverse process of the next time step until t = 0 to obtain the initial physical patch.

[0047] Add the initial physical patch to the original background image to obtain the target background image. Obtain the initial state and preset actions of the agent in the target background image, and input the initial state and the preset actions into a pre-trained differential dynamics model, where the differential dynamics model is based on the VAE-MDRNN joint architecture. Process the initial state and preset actions through the differential dynamics model to optimize the initial physical patch and output the target physical patch.

[0048] Through the above solution, obtain the original background image, input the original background image into a pre-trained diffusion model, use the diffusion model to perform forward diffusion on the original background image to obtain a noisy background image, and use the diffusion model to perform reverse diffusion processing on the noisy background image to obtain the initial physical patch. Using the original background of the environment as the initial state of the physical patch can ensure that the patch is more consistent with the environmental background during the generation process, avoiding the abruptness that may be caused by random perturbations in traditional methods. Add the initial physical patch to the original background image to obtain the target background image, obtain the initial state and preset actions of the agent in the target background image, and input the initial state and the preset actions into a pre-trained differential dynamics model. After being processed by the differential dynamics model, output the target physical patch. Optimize the initial physical patch through the differential dynamics model to make the obtained physical patch fit the original environment better, be able to blend into the environment more covertly, thereby improving the concealment of the attack and reducing the risk of being detected and recognized.

[0049] In some embodiments, the training process of the diffusion model includes:

[0050] Step 10A, obtain a first training dataset and a preset noise, where the first training dataset contains training background images;

[0051] Step 10B, input the training background image and the preset noise into the initial diffusion model, and use the initial diffusion model to perform forward diffusion processing on the training background image to obtain a training noisy background image;

[0052] Step 10C: Perform reverse diffusion processing on the training noisy background image using the initial diffusion model to obtain training noise and training physical patches;

[0053] Step 10D: Determine the diffusion loss function according to the preset noise and the training noise;

[0054] Step 10E: Train the initial diffusion model based on the diffusion loss function until the diffusion loss function converges to a first preset convergence threshold, and the initial diffusion model is trained to completion to obtain a diffusion model.

[0055] In specific implementation, obtain a first training dataset and a preset noise, where the first training dataset contains training background images. To achieve the natural fusion of physical patches and the target environment, a background image dataset covering diverse environmental contents needs to be collected and constructed, and this background image dataset is the first training dataset.

[0056] The construction process of the first training dataset specifically includes:

[0057] Step a: Collect a background image dataset from the target environment. In the target environment, interact with the environment through a pre-trained policy model π, collect the states observed by the agent, and save them to the dataset In the CarRacing environment of autonomous driving, the dataset contains multi-position images of the track, vehicle, and surrounding environment to improve the generalization ability of the model.

[0058] Step b: To improve the generalization ability of the model, apply enhanced random cropping to the original images to simulate the offset of different camera perspectives, and also inject Gaussian noise into the original images to enhance the robustness of the model to sensor noise and other operations.

[0059] Input the training background image and the preset noise into the initial diffusion model, and perform forward diffusion processing on the training background image using the initial diffusion model to obtain a training noisy background image.

[0060] In this embodiment, the specific process of forward diffusion includes:

[0061] Obtain a preset number of time steps. Within the preset number of time steps, gradually add the preset noise to the training background image using the initial diffusion model to obtain a training noisy background image, where the training noisy background image is represented by the formula:

[0062]

[0063] where, x tis the training noisy background image obtained after the t-th step, α t is the noise scheduling coefficient corresponding to the preset noise, t is the time step, and t ∈ {1, 2,..., T}, x t-1 is the training noisy background image corresponding to the (t - 1)-th step, x t-1 relative to x t is the training background image, ∈ is the preset noise, is a standard multi-dimensional Gaussian distribution, which means that each dimension of ∈ is independent and follows the standard normal distribution N(0, 1).

[0064] That is, in the forward diffusion process of the diffusion model, it is realized by gradually adding Gaussian noise to the input data of the model. At the time step t ∈ {1, 2,..., T}, the data x t is generated from the previous state x t-1 by adding Gaussian noise:

[0065]

[0066] where α t is a predefined noise scheduling coefficient that controls the rate of noise addition. As the time step t increases, the data gradually loses its original structural information and finally approximates pure Gaussian noise, is a standard multi-dimensional Gaussian distribution, which means that each dimension of ∈ is independent and follows the standard normal distribution N(0, 1).

[0067] In this embodiment, the architecture of the denoising network in the diffusion model is an improved U-Net architecture, and the specific architecture design is as follows:

[0068] The encoder extracts multi-scale features through 4 layers of downsampling; the decoder gradually restores the resolution through transposed convolution and fuses with the skip connections of the encoder to retain high-frequency details. The time step embedding encodes the time step t of the diffusion process into a 128-dimensional vector, which is mapped through a fully connected layer and added to the features of each layer to dynamically adjust the denoising behavior under different noise intensities. At the same time, a channel attention module is introduced in the middle layer of the decoder to enhance the model's ability to focus on features in key regions.

[0069] Using the initial diffusion model to perform reverse diffusion processing on the training noisy background image to obtain training noise and training physical patches. In this embodiment, the specific process of reverse diffusion includes:

[0070] Within the preset time step, use the denoising network to perform reverse diffusion processing on the training noisy background image to obtain training noise, denoised background image and training physical patches, where the denoised background image is represented by the formula:

[0071]

[0072] Among them, x t-1 ′ is the denoised background image, σ t is the noise variance parameter, z is a random noise variable subject to the standard normal distribution, which is used to adjust the generated samples during the denoising process, ∈ θ is the denoising network, is the standard multi-dimensional Gaussian distribution, which means that each dimension of z is independent and subject to the standard normal distribution N(0, 1).

[0073] That is, in this embodiment, the reverse denoising process trains the denoising network ∈ θ to predict the noise ∈ at each step, and gradually recover the original data x0 from the noise x T

[0074] The reverse update at each step can be expressed as:

[0075]

[0076] Among them σ t is the noise variance parameter, is the standard multi-dimensional Gaussian distribution, which means that each dimension of z is independent and subject to the standard normal distribution N(0, 1). The essence of the reverse process is to gradually reconstruct the data distribution by iteratively correcting the noise prediction error. The step-by-step optimization characteristic of the diffusion model makes it more stable when generating complex textures (such as track signs, vehicle appearances), avoiding the problem of mode collapse.

[0077] Determine the diffusion loss function according to the preset noise and the training noise, that is, optimize the network parameters by minimizing the mean square error between the predicted noise and the true noise. The diffusion loss function is expressed by the formula:

[0078]

[0079] Among them, L diff is the diffusion loss function, x0 is the training background image, ∈ θ (x t , t) is the training noise.

[0080] Train the initial diffusion model based on the diffusion loss function until the diffusion loss function converges to the first preset convergence threshold, and the initial diffusion model training is completed to obtain the diffusion model.

[0081] In this embodiment, the Adam optimizer is used during the training process, and the learning rate is set to 1×10 -3 , and the training is stopped in advance when the loss does not decrease for 10 consecutive epochs to prevent overfitting. ​

[0082] Through the above solution, an improved U-Net architecture is adopted as the denoising network, and its design takes into account both global semantic perception and local detail reconstruction capabilities. The core of the diffusion model is the design and training of the denoising network ∈ θ The design and training of, using the U-Net architecture as the denoising network ∈ θ , the high-frequency details of the image can be retained through skip connections to avoid blurring of the patch edges. The time step embedding encodes the time step t of the diffusion process into a vector, which is concatenated with the noisy image x t and then input into the network to dynamically adjust the denoising intensity. Through a data-driven generation method, it is ensured that the initial physical patches generated by the diffusion model are balanced in terms of visual realism, concealment, and environmental adaptability, providing high-quality initial physical patch samples for subsequent multi-objective optimization.

[0083] In some embodiments, step 105 specifically includes:

[0084] Step 1051, input the initial state and the preset action into a pre-trained differential dynamic model, and after being processed by the differential dynamic model, obtain an initial predicted state. Determine the target loss gradient according to the initial predicted state and the preset target state to complete the iteration process.

[0085] Step 1052, determine whether the target loss gradient is less than a preset gradient threshold.

[0086] Step 1053, in response to the target loss gradient being greater than or equal to the preset gradient threshold, obtain the physical patch attributes in the background image corresponding to the predicted state, adjust the physical patch attributes according to the target loss gradient to obtain a new physical patch, add the new initial physical patch to the original background image to determine a new initial state, and repeat the iteration process.

[0087] Step 1054, in response to the target loss gradient being less than the preset gradient threshold, exit the iteration process, use the initial predicted state as the target predicted state, and use the physical patch in the background image corresponding to the target predicted state as the target physical patch.

[0088] Specifically in implementation, input the initial state and the preset action into a pre-trained differential dynamic model, and after being processed by the differential dynamic model, obtain an initial predicted state. That is, add the initial physical patch obtained by the diffusion model to the environment, input the initial state and the predetermined action of the agent in the environment into the differential dynamic model, and output the predicted state of the agent after being processed by the differential dynamic model. Determine the target loss gradient according to the initial predicted state and the preset target state to complete the iteration process.

[0089] In this embodiment, the process of determining the target loss gradient specifically includes:

[0090] An attack loss function is calculated based on the initial prediction state and a preset target state. The attack loss function is used to measure the ability of a patch to mislead the agent's decision-making, and is defined as the cumulative difference between the target state and the true state. The attack loss function is expressed by the formula:

[0091]

[0092] where L attack is the attack loss function, o t (φ) is the initial prediction state, and o target is the predicted target state.

[0093] Determine the physical patch in the background image corresponding to the initial prediction state. Based on the physical patch in the background image corresponding to the initial prediction state and the original background image, a concealment loss function is determined. The concealment loss function is composed of the weighted sum of the L2 norm loss and the perceptual loss, taking into account both the pixel-level perturbation amplitude and the semantic-level visual consistency. The concealment loss function is expressed by the formula:

[0094]

[0095] where L stealth is the concealment loss function, Δo is the physical patch in the background image corresponding to the initial prediction state, o ori is the original background image, φ l (I generated ) is the input of the l-th layer network corresponding to the generated physical patch, φ l (I target ) is the input of the l-th layer network corresponding to the original background image, and λ L2 is the L2 norm.

[0096] The L2 norm loss directly constrains the difference amplitude between the patch pixel values and the background, suppressing overly conspicuous local perturbations. The perceptual loss can utilize the deep neural network to capture the high-level features of the image, making the generated image more in line with human visual perception. The concealment loss reflects the visual consistency and structural similarity between the physical patch and the covered environmental background.

[0097] Obtain the physical patch attributes in the background image corresponding to the prediction state. Based on the physical patch attributes and preset attribute constraints, a physical feasibility loss function is determined. The physical feasibility loss function is expressed by the formula:

[0098]

[0099] where L feasibleis the physical feasibility loss function, A represents the area of the physical patch corresponding to the physical patch attribute, θ represents the rotation angle of the physical patch corresponding to the physical patch attribute, A0 represents the area of the entire picture observed by the agent, P=(x,y) represents the coordinate position of the physical patch, is the function to calculate the impact of the physical patch on the concealment loss based on its area. The larger the area of the physical patch, the greater the corresponding interference to the environment. is to calculate the contribution of the physical patch to the concealment loss based on its rotation angle. is the function to calculate the impact of the physical patch on the concealment loss based on its position. Both the rotation angle and position of the physical patch will affect its degree of integration with the environmental background, thereby affecting concealment. λ A is to adjust the influence degree of the physical patch area on the final concealment loss function, λ P is the weight corresponding to the physical patch position influence function, λ θ is the weight corresponding to the physical patch rotation angle influence function.

[0100] The physical feasibility loss function reflects the impact of the physical patch on the environment, restricting the patch area A not to exceed the ratio threshold τ of the total area A0 of the observed picture A , and calculates the patch position through the indicator function, punishing the physical patches beyond the specified range.

[0101] Determine the target loss gradient according to the attack loss function, the concealment loss function and the physical feasibility loss function. Among them, the target loss gradient is represented by the formula:

[0102]

[0103] Among them, is the target loss gradient, w1 is the weight value corresponding to the attack loss function, w2 is the weight value corresponding to the concealment loss function, w3 is the weight value corresponding to the physical feasibility loss function, and the physical patch attribute φ=(p x ,p y ,s,θ), s is the physical patch size, θ is the rotation angle of the physical patch, p x ,p y is the coordinate position of the physical patch.

[0104] The attack loss function measures the impact of the patch on the agent's decision-making. The concealment loss function ensures the visual consistency between the patch and the environmental background. The physical feasibility loss function restricts the physical attributes of the patch. By determining the target loss gradient according to the attack loss function, the concealment loss function and the physical feasibility loss function, the optimal solution of the comprehensive performance is achieved through dynamic weight allocation.

[0105] Determine whether the target loss gradient is less than a preset gradient threshold. If the target loss gradient is greater than or equal to the preset gradient threshold, it indicates that iterative processing needs to continue. At this time, obtain the physical patch attributes in the background image corresponding to the predicted state. The physical patch attributes include physical patch content, physical patch size, physical patch position, and physical patch angle.

[0106] Adjust the physical patch attributes according to the target loss gradient to obtain a new physical patch. The specific adjustment process is as follows:

[0107] Perform attribute update according to the target loss gradient. The attribute update process is represented by the formula where, φ (k+1) is the physical patch attribute corresponding to the new physical patch, φ (k) is the physical patch attribute, and L total is the target loss gradient.

[0108] Add the new initial physical patch to the original background image to obtain a new target background image. Obtain the new initial state and preset action of the agent in the new target background image, input the new initial state and the preset action into the pre-trained differential dynamic model, and after being processed by the differential dynamic model, output a new initial predicted state. Determine the new target loss gradient according to the new initial predicted state and the preset target state.

[0109] If the target loss gradient is less than the preset gradient threshold, it indicates that the iteration is completed. Exit the iterative process, use the initial predicted state as the target predicted state, and use the physical patch in the background image corresponding to the target predicted state as the target physical patch. The target physical patch and its corresponding physical patch attributes are the optimal solutions.

[0110] Through the above solution, by constructing a differentiable calculation link, combining dynamic model prediction, loss calculation, and gradient update, the efficient optimization of patch attributes is achieved.

[0111] In some embodiments, the differential dynamic model includes a variational autoencoder, a variational decoder, and a mixture density recurrent neural network. In step 1051, inputting the initial state and the preset action into the pre-trained differential dynamic model and obtaining an initial predicted state through the processing of the differential dynamic model specifically includes:

[0112] Step 10511, input the initial state into the variational autoencoder, and perform compression processing using the variational autoencoder to obtain a first latent variable;

[0113] Step 10512: Input the first latent variable and the preset action into the mixture density recurrent neural network, process them using the mixture density recurrent neural network, and output the second latent variable corresponding to the next state of the initial state.

[0114] Step 10513: Input the second latent variable into the variational autoencoder, and perform reconstruction processing using the variational autoencoder to obtain the initial predicted state.

[0115] Specifically, the differential dynamic model is based on the VAE-MDRNN joint architecture. Its differentiable property allows the calculation of the gradient of the loss with respect to the patch attributes through automatic differentiation techniques. The state transition process of the environment is predicted through a pre-trained environment model. The differentiable dynamic model can introduce gradient calculation. Given the current state o t and action a t , the model predicts the next state o t+1 , and backpropagates the gradient through the chain rule.

[0116] Specifically, input the initial state o t into the variational autoencoder, and perform compression processing using the variational autoencoder to obtain the first latent variable z t . Input the first latent variable z t and the preset action a t into the mixture density recurrent neural network, process them using the mixture density recurrent neural network, and output the second latent variable z t+1 corresponding to the next state of the initial state.

[0117] Input the second latent variable z t+1 into the variational autoencoder, and perform reconstruction processing using the variational autoencoder to obtain the initial predicted state o t+1 .

[0118] In this embodiment, the physical attributes of the patch (such as size, position, rotation angle, etc.) will also be continuously adjusted during the optimization process. By optimizing these attributes of the patch, the patch can be more covertly integrated into the environment, thereby improving the concealment of the attack and reducing the risk of being detected and recognized. By finely adjusting the physical attributes of the patch, it is ensured that the patch has a strong attack effect during the optimization process and can maintain good concealment in the environment.

[0119] In this embodiment, the variational autoencoder and the mixture density recurrent neural network are pre-trained. The combination of mean squared error and Kullback–Leibler divergence is used as the loss function for training the VAE, and the Gaussian mixture loss is used as the loss function for training the MD-RNN.

[0120] Based on the same inventive concept, another embodiment of the present disclosure provides a method for generating physical patches based on a diffusion model, and the method specifically includes:

[0121] Step 201, define the attack scenario and target.

[0122] Determine the specific environment of the attack and the specific target state that the attacker hopes the agent to reach. This step requires clarifying the attack scenario, such as a specific section or specific traffic situation in autonomous driving, and the state that the attacker hopes the agent to reach through the attack, such as deviating from the lane, reaching the target position, etc.

[0123] Step 202, collect the initial dataset.

[0124] In the absence of an attack, use the pre-trained policy to perform multiple simulations in the environment, and collect the dataset of the initial state, action, and next state. The purpose of this step is to obtain the normal operation data of the environment and provide a basis for subsequent model training.

[0125] Step 203, add noise to enhance the dataset.

[0126] To increase the diversity of the dataset, add different intensities of noise to the collected actions and perform more simulations to explore different states of the environment. In this way, various possible environmental changes can be simulated, and the generalization ability of the model can be enhanced.

[0127] Step 204, construct a differentiable dynamic model.

[0128] Use the collected data to train the VAE so that it can encode and decode the environmental state. Utilize the state encoding output by the VAE to train the MD-RNN to predict the next state of the environment. Combine the VAE and the MD-RNN to form a complete differentiable dynamic model, which can predict the next state given the current state and action. The purpose of this step is to construct a dynamic model that can be used to optimize the attack perturbation, enabling the attacker to adjust the physical patch through optimization methods such as gradient descent.

[0129] In this embodiment, a variational autoencoder (VAE) and a mixture density recurrent neural network (MD-RNN) are used to construct a dynamic model, which is denoted as f(·,·;w). The model receives the environmental state and action as inputs and predicts the next environmental state. W represents the trainable parameters of the model. A combination of mean squared error and Kullback–Leibler divergence is used as the loss function for training the VAE, and Gaussian mixture loss is used as the loss function for training the MD-RNN.

[0130] The VAE is used to process the encoding and decoding of the environmental state, while the MD-RNN is used to handle the dynamic changes of state transitions. Through this dynamic model, an attacker can predict the changes in the environment when a certain action is selected, which is crucial for designing physical patches that can mislead the target agent towards a specific target state.

[0131] The main function of the prediction model F can be expressed as

[0132] F:o i ×a i →o i+1 (3-2)

[0133] where o i is the environmental state currently observed by the agent, a i is the action selected by the agent, and o i+1 is the output of the environmental dynamic model, corresponding to the state of the agent predicted by the model at the next moment.

[0134] Through the dynamic model, the physical patch attack algorithm can optimize the attack perturbation through gradient descent, thus effectively guiding the attacker to select the best physical patch content to achieve the attack goal. Through this method, the attacker can achieve targeted attacks through perturbations in the physical environment without directly accessing the perception module of the target agent.

[0135] Step 205, train the diffusion model.

[0136] Sample background images from the environmental background dataset, generate noisy data through the forward diffusion process, and then train the denoising network by minimizing the diffusion loss function. The purpose of this step is to generate physical patches that highly blend with the environmental background.

[0137] Step 206, initialize the parameters of the attack algorithm.

[0138] Set the initial parameters of the attack algorithm, including the time window, attack perturbation bound, learning rate, etc., and randomly initialize the perturbation image of the physical patch. The purpose of this step is to prepare the running environment of the attack algorithm to ensure the smooth progress of the algorithm.

[0139] Step 207, generate the initial physical patch.

[0140] Use the sampled background image as input and input it into a pre-trained diffusion model. The model generates a prototype of the physical patch that visually blends with the environmental background by learning the data distribution. The purpose of this step is to generate an initial physical patch to provide a starting point for the subsequent optimization process.

[0141] The steps to gradually generate the physical patch from noise through the reverse diffusion process are as follows:

[0142] Sample the initial noise from the standard Gaussian distribution As the starting point of the reverse process. To enhance the generation diversity, the latent space interpolation technique is adopted to linearly combine multiple noise samples.

[0143] Iteratively execute the reverse process to gradually generate the patch image Δo that fuses with the environmental background. At each step t, input the current noisy image x t and the time step embedding into the denoising network to predict the noise ∈ θ (x t , t). Calculate x t-1 according to the reverse update formula, and impose constraint conditions (such as truncating the pixel values to [0, 1]) to ensure the physical feasibility of the generated image. After generating the image x t-1 at each step, it needs to be input into the reverse process of the next time step until t = 0.

[0144] Step 208, optimization of the physical patch.

[0145] By adding the physical patch output by the diffusion model to the environment, continuously optimize the patch content using the differential dynamic model to achieve a balance between concealment and attack effect, and find the optimal physical patch. The purpose of this step is to further adjust the physical patch to achieve the best balance between attack effect and concealment.

[0146] Execute a multi-objective optimization process to synchronously adjust the content, position, and size of the physical patch, aiming to determine the optimal combination of physical properties to maximize the effectiveness of the attack. The purpose of this step is to find the physical patch that can achieve a balance between attack effect and concealment.

[0147] In this embodiment, the objective loss gradient is used for optimization. The determination process of the objective loss gradient specifically includes:

[0148] Calculate the attack loss function according to the initial prediction state and the preset target state. The attack loss function is used to measure the misleading ability of the patch to the agent's decision-making, and is defined as the cumulative difference between the target state and the true state. The attack loss function is expressed by the formula:

[0149]

[0150] where L attack is the attack loss function, o t (φ) is the initial prediction state, and o target is the predicted target state.

[0151] Determine the physical patch in the background image corresponding to the initial prediction state, and determine the concealment loss function based on the physical patch in the background image corresponding to the initial prediction state and the original background image. The concealment loss function is composed of the weighted sum of the L2 norm loss and the perceptual loss, taking into account both the pixel-level perturbation amplitude and the semantic-level visual consistency. The concealment loss function is expressed by the formula:

[0152]

[0153] where, L stealth is the concealment loss function, Δo is the physical patch in the background image corresponding to the initial prediction state, o ori is the original background image, φ l (I generated ) is the input of the l-th layer network corresponding to the generated physical patch, φ l (I target ) is the input of the l-th layer network corresponding to the original background image, λ L2 is the L2 norm.

[0154] The L2 norm loss directly constrains the difference amplitude between the patch pixel values and the background, suppressing overly conspicuous local perturbations. The perceptual loss can utilize the deep neural network to capture the high-level features of the image, making the generated image more in line with human visual perception. The concealment loss reflects the visual consistency and structural similarity between the physical patch and the covered environmental background.

[0155] Obtain the physical patch attributes in the background image corresponding to the prediction state, and determine the physical feasibility loss function according to the physical patch attributes and the preset attribute constraints, where the physical feasibility loss function is expressed by the formula:

[0156]

[0157] where, L feasible is the physical feasibility loss function, A represents the area size of the physical patch corresponding to the physical patch attributes, θ represents the rotation angle of the physical patch corresponding to the physical patch attributes, A0 represents the area size of the entire picture observed by the agent, P=(x,y) represents the coordinate position of the physical patch, is the influence function for calculating its impact on the concealment loss according to the area size of the physical patch. The larger the area of the physical patch, the greater the corresponding interference to the environment, is the contribution for calculating its impact on the concealment loss according to the rotation angle of the physical patch, is the influence function for calculating its impact on the concealment loss according to the position of the physical patch. Both the rotation angle and the position of the physical patch will affect its degree of integration with the environmental background, thereby affecting the concealment, λ ATo adjust the influence degree of the physical patch area on the final concealment loss function, λ P is the weight corresponding to the physical patch position influence function, λ θ is the weight corresponding to the physical patch rotation angle influence function.

[0158] The physical feasibility loss function reflects the impact of the physical patch on the environment, restricting the patch area A not to exceed the proportion threshold τ of the total area A0 of the observation screen A , calculates the patch position through the indicator function, and penalizes physical patches that exceed the specified range.

[0159] Determine the target loss gradient according to the attack loss function, the concealment loss function and the physical feasibility loss function, where the target loss gradient is expressed by the formula:

[0160]

[0161] Among them, is the target loss gradient, w1 is the weight value corresponding to the attack loss function, w2 is the weight value corresponding to the concealment loss function, w3 is the weight value corresponding to the physical feasibility loss function, and the physical patch attribute φ=(p x , p y , s, θ), s is the physical patch size, θ is the rotation angle of the physical patch, p x , p y is the coordinate position of the physical patch.

[0162] The attack loss function measures the impact of the patch on the agent's decision-making, the concealment loss function ensures the visual consistency of the patch with the environmental background, and the physical feasibility loss function restricts the physical attributes of the patch. By determining the target loss gradient according to the attack loss function, the concealment loss function and the physical feasibility loss function, the optimal solution of the comprehensive performance is achieved through dynamic weight allocation.

[0163] In this embodiment, the core of the physical patch attack method based on the diffusion model lies in using the characteristics of the diffusion model to optimize the generation of physical patches. This method first determines the best position and size of the physical patch through in-depth analysis of the target agent and the environment. This step utilizes the differentiable characteristics of the environmental dynamic model, and through calculation and simulation, accurately places the patch at the position that can most affect the target agent's decision-making. This optimization not only improves the pertinence of the attack, maximizes the attack effect, but also ensures that the size and position of the patch can minimize the interference to the environment, thereby reducing the risk of being detected by the target agent.

[0164] Diffusion models generate high-quality samples by simulating the gradual diffusion process from the data distribution to the noise distribution and reversing this process. Compared with traditional generative models, diffusion models can generate physical patches that blend more naturally with the environmental background, reducing the risk of being recognized by the target agent. The key advantage of this process is that it can generate physical patches that are visually closer to the natural state, thus significantly reducing the sensitivity of the target agent to anomalies and improving the concealment of the attack.

[0165] In addition, when generating adversarial samples using the physical patch attack method based on diffusion models, attention is paid to minimizing and concealing perturbations. Through the gradual reverse generation process of the diffusion model, the degree of perturbation in the patch can be precisely controlled to ensure that while guiding the behavior of the target agent, the patch does not deviate too much from the natural state of the environment. This fine-grained perturbation control not only improves the success rate of the attack but also provides more operating space for the attacker, enabling them to flexibly adjust the attack strategy without alerting the target agent.

[0166] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server, etc. The method of this embodiment can also be applied to a distributed scenario, and completed by multiple devices cooperating with each other. In such a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.

[0167] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0168] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a physical patch generation device based on a diffusion model.

[0169] Refer to Figure 2 , Figure 2 For the physical patch generation device based on the diffusion model of the embodiment, including:

[0170] A data acquisition module 201, configured to acquire an original background image and input the original background image into a pre-trained diffusion model;

[0171] The forward diffusion module 202 is configured to perform forward diffusion on the original background image by using the diffusion model to obtain a noisy background image;

[0172] The reverse diffusion module 203 is configured to perform reverse diffusion processing on the noisy background image by using the diffusion model to obtain an initial physical patch;

[0173] The target background image determination module 204 is configured to add the initial physical patch to the original background image to obtain a target background image;

[0174] The target physical patch determination module 205 is configured to obtain the initial state and a preset action of the agent in the target background image, input the initial state and the preset action into a pre-trained differential dynamics model, and output a target physical patch after being processed by the differential dynamics model.

[0175] In some embodiments, the device further includes a model training module, and the model training module is

[0176] specifically configured to:

[0177] Obtain a first training data set and a preset noise, where the first training data set contains training background images;

[0178] Input the training background image and the preset noise into an initial diffusion model, and perform forward diffusion processing on the training background image by using the initial diffusion model to obtain a training noisy background image;

[0179] Perform reverse diffusion processing on the training noisy background image by using the initial diffusion model to obtain training noise and a training physical patch;

[0180] Determine a diffusion loss function according to the preset noise and the training noise;

[0181] Train the initial diffusion model based on the diffusion loss function until the diffusion loss function converges to a first preset convergence threshold, and the initial diffusion model training is completed to obtain a diffusion model.

[0182] In some embodiments, the model training module is specifically configured to:

[0183] Obtain a preset number of time steps. Within the preset number of time steps, gradually add the preset noise to the training background image by using the initial diffusion model to obtain a training noisy background image, where the training noisy background image is represented by the formula:

[0184]

[0185] where, x tis the training noisy background image obtained after the t-th step, and α t is the noise scheduling coefficient corresponding to the preset noise, t is the time step, and t ∈ {1, 2,..., T}, x t-1 is the training noisy background image corresponding to the (t - 1)-th step, x t-1 relative to x t is the training background image, and ∈ is the preset noise.

[0186] In some embodiments, the initial diffusion model includes a denoising network, and the model training module is specifically configured to:

[0187] Within the preset time step, use the denoising network to perform reverse diffusion processing on the training noisy background image to obtain training noise, a denoised background image, and training physical patches, where the denoised background image is represented by the formula:

[0188]

[0189] where x t-1 ′ is the denoised background image, σ t is the noise variance parameter, z is a random noise variable that follows a standard normal distribution and is used to adjust the generated samples during the denoising process, so that, ∈ θ is the denoising network, is the standard multi-dimensional Gaussian distribution, which means that each dimension of z is independent and follows the standard normal distribution N(0, 1).

[0190] In some embodiments, the model training module is specifically configured to:

[0191] Determine a diffusion loss function according to the preset noise and the training noise, and the diffusion loss function is represented by the formula:

[0192]

[0193] where L diff is the diffusion loss function, x0 is the training background image, ∈ θ (x t , t) is the training noise.

[0194] In some embodiments, the target physical patch determination module 205 is specifically configured to:

[0195] Input the initial state and the preset action into a pre-trained differential dynamic model, and after being processed by the differential dynamic model, obtain an initial predicted state, and determine a target loss gradient according to the initial predicted state and the preset target state to complete the iterative process;

[0196] Determine whether the target loss gradient is less than a preset gradient threshold;

[0197] In response to the target loss gradient being greater than or equal to the preset gradient threshold, obtain the physical patch attributes in the background image corresponding to the predicted state, adjust the physical patch attributes according to the target loss gradient to obtain a new physical patch, add the new initial physical patch to the original background image, determine a new initial state, and repeat the iterative process;

[0198] In response to the target loss gradient being less than the preset gradient threshold, exit the iterative process, use the initial prediction state as the target prediction state, and use the physical patch in the background image corresponding to the target prediction state as the target physical patch.

[0199] In some embodiments, the differential dynamic model includes a variational autoencoder, a variational decoder, and a mixture density recurrent neural network; the target physical patch determination module 205 is specifically configured to:

[0200] Input the initial state into the variational autoencoder, and perform compression processing using the variational autoencoder to obtain a first latent variable;

[0201] Input the first latent variable and the preset action into the mixture density recurrent neural network, and process them using the mixture density recurrent neural network to output a second latent variable corresponding to the next state of the initial state;

[0202] Input the second latent variable into the variational decoder, and perform reconstruction processing using the variational decoder to obtain an initial prediction state.

[0203] In some embodiments, the target physical patch determination module 205 is specifically configured to:

[0204] Calculate an attack loss function according to the initial prediction state and a preset target state, and the attack loss function is expressed by the formula:

[0205]

[0206] where L attack is the attack loss function, o t (φ) is the initial prediction state, o target is the predicted target state;

[0207] Determine the physical patch in the background image corresponding to the initial prediction state, and determine a concealment loss function according to the physical patch in the background image corresponding to the initial prediction state and the original background image. The concealment loss function is expressed by the formula:

[0208]

[0209] Among them, L stealth is the concealment loss function, Δo is the physical patch in the background image corresponding to the initial prediction state, and o ori is the original background image, and φ l (I generated ) is the input of the l-th layer network corresponding to the generated physical patch, and φ l (I target ) is the input of the l-th layer network corresponding to the original background image, and λ L2 is the L2 norm;

[0210] Obtain the physical patch attributes in the background image corresponding to the prediction state, and determine the physical feasibility loss function according to the physical patch attributes and the preset attribute constraints, where the physical feasibility loss function is expressed by the formula:

[0211]

[0212] Among them, L feasible is the physical feasibility loss function, A represents the area size of the physical patch corresponding to the physical patch attributes, θ represents the rotation angle of the physical patch corresponding to the physical patch attributes, A0 represents the area size of the entire picture observed by the agent, and P=(x,y) represents the coordinate position of the physical patch, is the function to calculate the impact of the physical patch area on the concealment loss. The larger the area of the physical patch, the greater the corresponding interference to the environment, is the contribution to calculate the concealment loss according to the rotation angle of the physical patch, is the function to calculate the impact of the physical patch position on the concealment loss. The rotation angle and position of the physical patch will both affect its fusion degree with the environmental background, thus affecting the concealment, and λ A is to adjust the influence degree of the physical patch area on the final concealment loss function, and λ P is the weight corresponding to the physical patch position influence function, and λ θ is the weight corresponding to the physical patch rotation angle influence function;

[0213] Determine the target loss gradient according to the attack loss function, the concealment loss function, and the physical feasibility loss function, where the target loss gradient is expressed by the formula:

[0214]

[0215] Among them, is the target loss gradient, w1 is the weight value corresponding to the attack loss function, w2 is the weight value corresponding to the concealment loss function, and w3 is the weight value corresponding to the physical feasibility loss function.

[0216] For the convenience of description, when describing the above device, it is divided into various modules according to functions and described separately. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0217] The device in the above embodiment is used to implement the corresponding physical patch generation method based on the diffusion model in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0218] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the physical patch generation method based on the diffusion model described in any of the above embodiments.

[0219] Figure 3 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.

[0220] The processor 1010 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0221] The memory 1020 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and called and executed by the processor 1010.

[0222] The input / output interface 1030 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure), or can be externally connected to the device to provide corresponding functions. Among them, the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.

[0223] The communication interface 1040 is used to connect to the communication module (not shown in the figure) to achieve communication interaction between this device and other devices. Among them, the communication module can achieve communication through wired means (such as USB, network cable, etc.), or can achieve communication through wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0224] The bus 1050 includes a path to transmit information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0225] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and do not have to include all the components shown in the figure.

[0226] The electronic device of the above embodiment is used to implement the corresponding physical patch generation method based on the diffusion model in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0227] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium, and the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to make the computer execute the physical patch generation method based on the diffusion model as described in any of the above embodiments.

[0228] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0229] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the physical patch generation method based on the diffusion model described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0230] It can be understood that before using the technical solutions of the various embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.

[0231] For example, in response to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application program, server, or storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message.

[0232] As an optional but non-limiting implementation manner, the way of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0233] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure, and other ways that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0234] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; within the context of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above, and for the sake of brevity, they are not provided in detail.

[0235] In addition, for simplicity of explanation and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be entirely within the understanding of those skilled in the art). In cases where specific details (such as circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure may be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0236] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0237] The embodiments of the present disclosure are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A physical patch generation method based on a diffusion model, characterized in that, Including: Obtain an original background image, and input the original background image into a pre-trained diffusion model; Use the diffusion model to perform forward diffusion on the original background image to obtain a noisy background image; Use the diffusion model to perform reverse diffusion processing on the noisy background image to obtain an initial physical patch; Add the initial physical patch to the original background image to obtain a target background image; Obtain the initial state and a preset action of the agent in the target background image, input the initial state and the preset action into a pre-trained differential dynamics model, and after being processed by the differential dynamics model, output a target physical patch.

2. The method according to claim 1, characterized in that, The training process of the diffusion model includes: Obtain a first training dataset and a preset noise, where the first training dataset contains training background images; Input the training background image and the preset noise into an initial diffusion model, and use the initial diffusion model to perform forward diffusion processing on the training background image to obtain a training noisy background image; Use the initial diffusion model to perform reverse diffusion processing on the training noisy background image to obtain a training noise and a training physical patch; Determine a diffusion loss function according to the preset noise and the training noise; Based on the diffusion loss function, train the initial diffusion model until the diffusion loss function converges to a first preset convergence threshold, and the initial diffusion model training is completed to obtain a diffusion model.

3. The method according to claim 2, characterized in that, The step of using the initial diffusion model to perform forward diffusion processing on the training background image to obtain a training noisy background image includes: Obtain a preset time step. Within the preset time step, use the initial diffusion model to gradually add the preset noise to the training background image to obtain a training noisy background image, where the training noisy background image is represented by the formula: Among them, x t is the training noisy background image obtained after the t-th step, α t is the noise scheduling coefficient corresponding to the preset noise, t is the time step, and t ∈ {1, 2,..., T}, x t-1 is the training noisy background image corresponding to the (t - 1)-th step, x t-1 relative to x t is the training background image, ∈ is the preset noise, is the standard multi-dimensional Gaussian distribution, which means that each dimension of ∈ is independent and follows the standard normal distribution N(0, 1).

4. The method according to claim 3, characterized in that, The initial diffusion model includes a denoising network; The step of using the initial diffusion model to perform reverse diffusion processing on the training noisy background image to obtain a training noise and a training physical patch includes: Within the preset time step, use the denoising network to perform reverse diffusion processing on the training noisy background image to obtain a training noise, a denoised background image, and a training physical patch, where the denoised background image is represented by the formula: Among them, x t-1 ′ is the denoised background image, σ t is the noise variance parameter, z is a random noise variable subject to the standard normal distribution, which is used to adjust the generated samples during the denoising process to make them closer to the real distribution, ∈ θ is the denoising network, is the standard multi-dimensional Gaussian distribution, which means that each dimension of z is independent and subject to the standard normal distribution N(0, 1).

5. The method according to claim 3, wherein The step of determining a diffusion loss function according to the preset noise and the training noise includes: Determine a diffusion loss function according to the preset noise and the training noise, and the diffusion loss function is represented by the formula: Among them, L diff is the diffusion loss function, x0 is the training background image, and ∈ θ (x t , t) is the training noise.

6. The method according to claim 1, wherein The step of inputting the initial state and the preset action into a pre-trained differential dynamics model, and after being processed by the differential dynamics model, outputting a target physical patch includes: Input the initial state and the preset action into a pre-trained differential dynamics model, and after being processed by the differential dynamics model, obtain an initial predicted state. Determine a target loss gradient according to the initial predicted state and a preset target state to complete an iterative process; Judge whether the target loss gradient is less than a preset gradient threshold; In response to the target loss gradient being greater than or equal to a preset gradient threshold, obtain the physical patch attributes in the background image corresponding to the predicted state, adjust the physical patch attributes according to the target loss gradient to obtain a new physical patch, add the new initial physical patch to the original background image, determine a new initial state, and repeat the iterative process; In response to the target loss gradient being less than the preset gradient threshold, exit the iterative process, use the initial predicted state as the target predicted state, and use the physical patch in the background image corresponding to the target predicted state as the target physical patch.

7. The method according to claim 6, characterized in that, The differential dynamic model includes a variational autoencoder, a variational decoder, and a mixture density recurrent neural network; The step of inputting the initial state and the preset action into a pre-trained differential dynamic model and obtaining an initial predicted state through the processing of the differential dynamic model includes: Input the initial state into the variational autoencoder and perform compression processing using the variational autoencoder to obtain a first latent variable; Input the first latent variable and the preset action into the mixture density recurrent neural network and process them using the mixture density recurrent neural network to output a second latent variable corresponding to the next state of the initial state; Input the second latent variable into the variational decoder and perform reconstruction processing using the variational decoder to obtain an initial predicted state.

8. The method according to claim 6, wherein The step of determining the target loss gradient according to the initial predicted state and a preset target state includes: Calculate an attack loss function according to the initial predicted state and the preset target state. The attack loss function is expressed by the formula: Among them, L attack is the attack loss function, o t (φ) is the initial prediction state, o target is the predicted target state; Determine the physical patch in the background image corresponding to the initial predicted state, and determine a concealment loss function according to the physical patch in the background image corresponding to the initial predicted state and the original background image. The concealment loss function is expressed by the formula: Among them, L stealth is the concealment loss function, Δo is the physical patch in the background image corresponding to the initial prediction state, and o ori is the original background image, and φ l (I generated ) is the input of the l-th layer network corresponding to the generated physical patch, and φ l (I target ) is the input of the l-th layer network corresponding to the original background image, and λ L2 is the L2 norm; Obtain the physical patch attributes in the background image corresponding to the predicted state, and determine a physical feasibility loss function according to the physical patch attributes and a preset attribute constraint. The physical feasibility loss function is expressed by the formula: Among them, L feasible is the physical feasibility loss function, A represents the area size of the physical patch corresponding to the physical patch attribute, θ represents the rotation angle of the physical patch corresponding to the physical patch attribute, A0 represents the area size of the entire picture observed by the agent, P=(x, y) represents the coordinate position of the physical patch, is the influence function for calculating its impact on the concealment loss according to the area size of the physical patch. The larger the area of the physical patch, the greater the corresponding interference to the environment. is the contribution of the physical patch to the concealment loss calculated according to the rotation angle of the physical patch. is the influence function for calculating its impact on the concealment loss according to the position of the physical patch. The rotation angle and position of the physical patch will both affect its degree of integration with the environmental background, thus affecting the concealment. λ A is to adjust the influence degree of the physical patch area on the final concealment loss function, λ P is the weight corresponding to the physical patch position influence function, λ θ is the weight corresponding to the physical patch rotation angle influence function; Determine the target loss gradient according to the attack loss function, the concealment loss function, and the physical feasibility loss function. The target loss gradient is expressed by the formula: Among them, is the target loss gradient, w1 is the weight value corresponding to the attack loss function, w2 is the weight value corresponding to the concealment loss function, and w3 is the weight value corresponding to the physical feasibility loss function.

9. A physical patch generation device based on a diffusion model, characterized in that, including: A data acquisition module, configured to acquire an original background image and input the original background image into a pre-trained diffusion model; A forward diffusion module, configured to perform forward diffusion on the original background image using the diffusion model to obtain a noisy background image; A reverse diffusion module, configured to perform reverse diffusion processing on the noisy background image using the diffusion model to obtain an initial physical patch; A target background image determination module, configured to add the initial physical patch to the original background image to obtain a target background image; A target physical patch determination module, configured to acquire the initial state and a preset action of an agent in the target background image, input the initial state and the preset action into a pre-trained differential dynamic model, and output a target physical patch through the processing of the differential dynamic model.

10. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the method according to any one of claims 1 to 8.