A dynamic target optical invisibility method and device based on causal reasoning
By optimizing the adversarial texture generation method through causal reasoning and incremental learning, the robustness and adaptability issues of existing optical cloaking technologies in dynamic scenes are solved, achieving stable optical cloaking effects and efficient attack capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG UNIV OF TECH
- Filing Date
- 2026-01-30
- Publication Date
- 2026-06-19
Smart Images

Figure CN122244269A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of security analysis of intelligent monitoring, and specifically relates to a dynamic target optical stealth method and device based on causal reasoning. Background Technology
[0002] Deep neural networks are widely used in target classification, detection, and tracking tasks in the field of intelligent surveillance. However, the existence of adversarial examples reveals a security weakness in intelligent surveillance systems: by designing target optical camouflage technologies and devices, the prediction results of deep neural networks can be misled, causing the intelligent surveillance system to fail to correctly detect the presence of target objects. This means that target optical camouflage technologies and devices can disable the detection function of the surveillance system, allowing unidentified targets to penetrate security barriers, thereby posing a substantial risk to personnel and property safety.
[0003] Existing optical stealth technologies include: (1) Invasive optical stealth technology: By applying perturbation textures to the target surface through pasting, drawing, or wearing, it requires direct contact with the target object and has poor concealment. (2) Non-invasive optical stealth technology: By using projection equipment to change the lighting conditions and surface texture of the target object in real time, an attack can be achieved. Most optical stealth technologies are invasive and require direct contact with the target object through 2D / 3D printing, such as wearing clothes with colored anti-textures, printing and pasting anti-textures on the target surface. Intuitively, non-invasive optical stealth technology poses a more serious threat than invasive optical stealth technology because it can achieve stealth effects without physical access to the target object.
[0004] Existing non-invasive optical stealth technologies mostly employ simple image transformations (such as rotation, translation, and scaling) to simulate changes in the physical environment. Their simulation accuracy is low and cannot accurately reflect the real optical projection and imaging process of surveillance cameras. Furthermore, most existing technologies are offline and static, meaning they generate static adversarial textures from training data in the digital domain and then project these textures into the physical environment. This results in an inability to adapt to dynamic changes in the surveillance footage in real time, leading to unstable attack effectiveness in complex scenarios.
[0005] Furthermore, traditional threat detection technologies suffer from systemic flaws in dynamic threat perception and response. Existing solutions primarily rely on single-modal data analysis, which not only struggles to effectively integrate the spatiotemporal correlation characteristics of multi-source heterogeneous data, leading to temporal misalignments and semantic gaps during attack chain reconstruction, but also lacks the ability to quantify the scope and speed of threat propagation. This results in persistent core problems such as high false positive rates and ambiguous attack root cause localization. Moreover, most causal discovery tools are disconnected from real-world perception; they typically process pre-cleaned structured time-series data and lack the ability to perform end-to-end causal modeling from raw unstructured data, rendering the attack process unexplainable.
[0006] Therefore, the field of intelligent surveillance security urgently needs a target optical stealth technology with real-time dynamic adaptive capabilities, high robustness under varying physical conditions, and the ability to effectively bridge the gap between perception and reasoning, in order to support security assessment and defense research of intelligent surveillance systems. Summary of the Invention
[0007] To address the issues of poor robustness, lack of real-time adaptive capability, and unexplainable attack processes in existing non-invasive optical stealth technologies, this invention proposes a dynamic target optical stealth method and device based on causal reasoning for security analysis in the field of intelligent surveillance, aiming to achieve efficient and robust dynamic target optical stealth in dynamic and complex scenarios.
[0008] To achieve the above-mentioned objectives, this invention provides a dynamic target optical stealth method based on causal reasoning, comprising the following steps: Step 1: Collect baseline physics scene images, adversarial textures, and physics scene images with adversarial textures overlaid as training datasets, and train a neural network model based on a latent diffusion model. Step 2: Construct a causal graph model for the deep neural network prediction task. By causally intervening in the spurious bias variables in the causal graph model, causal guidance is provided for the generation of adversarial textures. In the same iteration, a comprehensive optimization including projectibility physical constraints and environmental robustness constraints is applied to the adversarial textures. Step 3: Continuously acquire video frames in dynamic scenes as real-time background information. Based on the real-time background information, use the trained neural network model to update the comprehensive and optimized adversarial texture in real time, and continuously project the updated optical projection texture onto the surface of the moving target to maintain a stable optical cloaking effect in dynamic environments.
[0009] Preferably, in step 1, the neural network model based on the latent diffusion model includes: an image fusion module, an encoder module, and a computation module, used to approximate the real projection acquisition process; The image fusion module receives a baseline physical scene image and adversarial texture from the training dataset as input, and performs feature fusion using a preset fusion function to obtain an intermediate fused image; wherein the preset fusion function is expressed as follows: , For intermediate fusion images, For fusion function, As a reference physical scene image, For adversarial textures; The encoder module is used to extract features and perform dimensionality reduction encoding on the intermediate fused image to obtain a latent feature vector; The computation module is configured as a denoising backbone network U-Net, which performs denoising prediction in the latent space using latent feature vectors as input, to obtain a physical scene image superimposed with adversarial textures. To simulate physical projection transformation.
[0010] By using a baseline physical scene image and adversarial texture as input data pairs, and a physical scene image with adversarial texture superimposed as output, a neural network model based on a latent diffusion model is trained to represent the interaction features between the physical scene and the texture, realizing the optical transformation of the adversarial texture from digital space to physical space; and by performing denoising prediction in the latent space, the complex transformations such as illumination attenuation, color mixing, and geometric distortion generated during physical projection are effectively simulated, thereby achieving high-fidelity representation and simulation of physical scene features.
[0011] More preferably, an incremental learning mechanism is adopted and combined with the elastic weight consolidation algorithm to establish a first joint loss including physical approximation loss and elastic weight merging regularization loss, which is used to optimize the neural network model parameters to obtain the trained neural network model; Physical approximation loss is expressed as ;in, For the sampled Gaussian noise, For the denoising backbone network U-Net, For the time step of the diffusion process, It is the square of the L2 norm, used to calculate the Euclidean distance between the predicted noise and the actual noise; Elastic weight pooling regularization loss is expressed as ,in It is the first of the Fisher information matrix. One element, These are pre-trained weights. The weights are after fine-tuning. It is a regularization hyperparameter.
[0012] To address the "catastrophic forgetting" problem that may occur when adapting to specific physical scenarios—a phenomenon where neural network models, during the learning of new tasks, experience a sharp decline or even complete loss of performance on previously learned tasks due to significant adjustments to model parameters to adapt to new data—an incremental learning mechanism based on Elastic Weight Consolidation (EWC) is introduced during the fine-tuning process. This mechanism, in addition to using a physical approximation loss to optimize neural network model parameters, penalizes changes in the neural network model's weight parameters through a first joint loss. This aims to preserve the original generalization ability when adapting to new physical scenarios and prevent the learning process of the new task (fine-tuning) from excessively influencing the memory of the old task (pre-training).
[0013] Optionally, when optimizing the neural network model parameters, a differentiation strategy is implemented, including: applying a high penalty to the regularization term for weights with large Fisher information values, and limiting the weights after fine-tuning. Deviation from pre-trained weights This allows for the locking in of general knowledge; for weights with smaller Fisher information values, adjustments are made based on the physical scene image.
[0014] Through this joint differential optimization, the neural network model can eventually converge to one that can accurately simulate the lighting and geometric characteristics of the current physical environment while maintaining strong generalization performance.
[0015] Preferably, the causal graph model described in step 2 includes causal factor variables, spurious bias variables, output variables, and confounding variables; The causal factor variables represent input features that have a direct causal relationship with the neural network model prediction task, corresponding to the essential structural features of the target object in the optical stealth task; The spurious bias variable represents an input feature that has no direct causal relationship with the neural network model's prediction task but can affect the prediction result, corresponding to the visual pattern of adversarial textures. The output variable represents the prediction result of the neural network model; The confounding variable refers to a factor that simultaneously affects both causal and spurious bias variables, corresponding to the current physical environment characteristics in optical stealth, including at least one of ambient lighting conditions and imaging ensemble location.
[0016] More preferably, the causal intervention in step 2 is to perform an intervention operation on the spurious bias variable to block the backdoor path formed by the confounding variable and the causal factor variable; Among them, the intervention operation The causal effect is calculated using a backdoor adjustment formula: In the formula, This indicates the presence of spurious bias variables. Causal factors variables under the conditions of intervention The distribution, This indicates a mixed variable.
[0017] Preferably, the causal effect under intervention is calculated using a backdoor adjustment formula, and the difference between the output distributions of causal factor variables under interventions in the baseline physical scene image and adversarial texture is measured based on the maximum average difference, in order to construct a causal intervention loss function, the calculation formula of which is as follows: , In the formula, For causal intervention loss function, As a reference physical scene image, It is the maximum mean difference measure, used to measure the difference between two probability distributions. Indicates that in a given and hour The probability distribution, Indicates that in a given adversarial texture and hour The probability distribution.
[0018] The average difference method is a nonparametric approach that offers flexibility, high computational efficiency, and a robust statistical foundation, making it well-suited for comparing high-dimensional, complex distributions. By minimizing the causal intervention loss function, deep neural networks can be guided to strengthen their dependence on non-causal visual cues during prediction. This provides causal guidance for the generation of adversarial textures, ensuring that the generated adversarial textures stably influence the neural network model's dependence on false biases while maintaining the target structural features. Consequently, consistent optical cloaking effects are achieved under various environmental conditions.
[0019] Preferably, the projectibility constraint is implemented by converting the adversarial texture from the RGB color space to the LAB color space and constraining the numerical ranges of the A and B channels in the LAB color space. The projection compensation loss function corresponding to the projectibility constraint is expressed as: , In the formula, The projection compensation loss function is... and To counteract the color channel components of textures in LAB space, and Preset thresholds for channels A and B, It is a function with maximum value. It is the square of the L2 norm.
[0020] Preferably, the environmental robustness constraint is achieved by minimizing the sharpness-perceived loss function, which is used to improve the robustness of the adversarial texture to natural transformations in physical space. The corresponding sharpness-perceived loss function is expressed as follows: , In the formula, Let be the sharpness-aware loss function, representing the loss function applied to the dataset. Above calculations, regarding adversarial textures SAM loss; This indicates that the constraints are satisfied. disturbance Find the maximum sharpness perception loss function in the process; For the Euclidean norm, The disturbance radius is... This is the loss function for causal intervention.
[0021] Preferably, the optimized adversarial texture is achieved by minimizing the total loss function, which is a weighted sum of the projection compensation loss function, the sharpness-perceived loss function, and the texture norm regularization term, expressed as: , In the formula, For the total loss function, , and These are the projection compensation loss functions. Sharpness-perceived loss function The weight parameters of the texture norm regularization term, It is the square of the L2 norm. For adversarial textures.
[0022] The present invention also provides a dynamic target optical stealth device based on causal reasoning, including a memory and a processor. The memory is used to store a computer program, and the processor is used to implement the dynamic target optical stealth method based on causal reasoning when the computer program is executed.
[0023] Compared with the prior art, the beneficial effects of the present invention include at least the following: The present invention provides a dynamic target optical cloaking method and device based on causal reasoning. It trains a projection neural network model based on a latent diffusion model to learn and approximate the complex physical transformations in the actual projection acquisition process. Based on this, a causal graph model is constructed to predict the task, guiding the neural network model to strengthen its dependence on non-causal visual cues during prediction, thus providing causal guidance for the generation of adversarial textures. Furthermore, in the same iteration, a comprehensive optimization, including projectability physical constraints and environmental robustness constraints, is applied to the adversarial texture to ensure that the texture can be implemented by a physical projection device and to improve the texture's robustness to natural environmental changes. The optimized adversarial texture is updated in real time by the projection neural network model and drives the projection device to act on the dynamic scene, achieving a stable, effective, and physically feasible optical cloaking effect in a variable physical environment. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.
[0025] Figure 1 This is a flowchart illustrating the dynamic target optical stealth method based on causal reasoning provided by the present invention.
[0026] Figure 2 This describes the inference process of a neural network model based on a potential diffusion model.
[0027] Figure 3 This is a schematic diagram of the cause-effect graph model structure.
[0028] Figure 4 This is a bar chart comparing the attack success rates of the embodiments of the present invention and existing technologies on different target detection models. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and given in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0030] The inventive concept of this invention addresses the problems of poor robustness, lack of real-time adaptive capability, and unexplainable attack processes in existing non-invasive optical cloaking technologies. This invention provides a dynamic target optical cloaking method based on causal reasoning, employing a camera, projector, and intelligent monitoring system to form a dynamic target optical cloaking hardware system. By training a ProjectNet neural network model based on a latent diffusion model to approximate the real projection acquisition process, a causal graph model is constructed, and backdoor adjustments are used for causal intervention. The maximum mean difference is used to measure the distribution differences of causal factors under different inputs, enhancing the dynamic target optical cloaking effect of false associations. During the optimization process, the model is guided to strengthen its dependence on visual cues. In dynamic scenes, adversarial textures are subjected to closed-loop iteration and continuous optimization to calculate optical cloaking textures in real time. These textures are then projected onto the surface of the moving target to achieve a continuous optical cloaking effect.
[0031] like Figure 1 As shown in the embodiment, a dynamic target optical stealth method based on causal reasoning is provided, including the following steps: S1. Collect baseline physical scene images, adversarial textures, and physical scene images with adversarial textures overlaid as training datasets, and train a neural network model based on a latent diffusion model.
[0032] In this embodiment, a latent diffusion-based neural network model, ProjectNet, fine-tuned through incremental learning, is designed to approximate the actual projection acquisition process π and generate an approximate function. Specific methods include: A baseline physical scene image, adversarial textures, and a physical scene image with adversarial textures overlaid are collected as training datasets. First, a hardware acquisition system consisting of a projector and a camera is constructed. In this embodiment, the hardware acquisition system maintains a preset distance from the target object, for example, approximately 1.5 meters, to ensure that the projection coverage and image clarity meet predetermined quality standards. Based on this, the acquisition process specifically includes acquiring three sets of key data: using the hardware acquisition system to acquire the original state of the physical scene without overlaid digital perturbation textures, and acquiring the baseline physical scene image. Using a hardware acquisition system to capture computer-generated adversarial textures The adversarial texture is projected onto the surface of a target object (such as a pedestrian or vehicle), causing it to undergo an optical transformation from digital space to physical space; a hardware acquisition system simultaneously acquires physical scene images overlaid with the adversarial texture. The physical scene image It incorporates realistic physical environment features, specifically including geometric distortions and photometric variations (such as ambient lighting and color decay). Finally, the baseline physical scene image is... and adversarial textures Marked as model input, the physically captured images of the scene will be used. These are labeled as training targets, thereby constructing an "input-output" mapping dataset that can characterize the properties of the current physical environment. This training dataset was configured for subsequent fine-tuning of the ProjectNet model using an incremental learning mechanism.
[0033] This embodiment employs a deep neural network architecture based on a latent diffusion model (i.e., ProjectNet) to characterize the interaction features between the physical scene and adversarial textures. Specifically, as... Figure 2 As shown, this step first configures an image fusion module to use the benchmark physical scene images in the training dataset. and adversarial textures And through a preset fusion function The two are then fused to generate an intermediate fused image that incorporates both scene geometry and texture semantics. Subsequently, a pre-trained encoder unit was used. For the intermediate fused image Feature extraction and dimensionality reduction encoding are performed to map the pixel space from a high-dimensional pixel space to a low-dimensional latent space, resulting in latent feature vectors. Finally, the latent feature vectors Input to a U-Net architecture configured for denoising backbone In the computational unit; the denoising backbone U-Net is combined with the time step index. Perform denoising prediction within the latent space The aim is to simulate the complex transformations such as light attenuation, color mixing, and geometric distortion that occur during physical projection, thereby achieving high-fidelity representation and simulation of physical scene characteristics.
[0034] Parameter fine-tuning and updating combined with incremental learning mechanisms: To address the "catastrophic forgetting" problem that may occur when adapting to specific physical scenarios, which refers to the phenomenon where, during the learning of a new task, the model parameters are drastically adjusted to adapt to the new data, leading to a sharp decline or even complete loss of performance on previously learned tasks, an incremental learning mechanism based on Elastic Weight Consolidation (EWC) is introduced in the fine-tuning process. This mechanism, in addition to using physical approximation loss to optimize the neural network model parameters, penalizes changes in the neural network model weight parameters through a first joint loss. This ensures that the original generalization ability is preserved when adapting to new physical scenarios, preventing the learning process of the new task (fine-tuning) from excessively affecting the memory of the old task (pre-training).
[0035] Specifically, this includes: firstly, calculating the Fisher information matrix of the pre-trained model parameters. Quantize each parameter Importance weights are assigned to the original general vision task. Then, a first joint loss is constructed, comprising physical approximation loss and elastic weight merging regularization loss. This is used to optimize the parameters of the neural network model and obtain the trained neural network model.
[0036] Physical approximation loss is expressed as ;in, For the sampled Gaussian noise, For the denoising backbone network U-Net, For the time step of the diffusion process, It is the square of the L2 norm, used to calculate the Euclidean distance (i.e., error) between the predicted noise and the actual noise.
[0037] Elastic weight pooling regularization loss is expressed as ,in It is the first of the Fisher information matrix. One element, These are pre-trained weights (for the old task). The weights are fine-tuned (new task). It is a regularization hyperparameter.
[0038] When updating parameters via backpropagation, a differentiation strategy is implemented: for key parameters with large Fisher information values, a high penalty is applied to the regularization term to limit the weights after fine-tuning. Deviation from pre-trained weights This allows for the locking in of general knowledge; for non-critical parameters with small Fisher information values, significant adjustments are allowed based on physical data. Through this joint optimization, ProjectNet ultimately converges to an approximate function that accurately simulates the lighting and geometric properties of the current physical environment while maintaining strong generalization performance. .
[0039] S2. Construct a causal graph model for deep neural network prediction tasks. By causally intervening in the spurious bias variables in the causal graph model, causal guidance is provided for the generation of adversarial textures. In the same iteration, a comprehensive optimization including projectibility physical constraints and environmental robustness constraints is applied to the adversarial textures.
[0040] In this embodiment, to facilitate the explanation of the optical cloaking perturbation generation process based on causal reasoning, the relevant symbols are defined as follows: S represents the output causal factor variable of the neural network model under given input conditions; V represents the spurious bias variable in the input image that has no direct causal relationship with the task but may affect the prediction result of the neural network model; Y represents the final prediction result output by the neural network model; and n represents the confounding variable that simultaneously affects the causal factor variable S and the spurious bias variable V.
[0041] Human causal reasoning ability can identify the true causal factors S (such as contour, shape, etc.) used for prediction tasks and ignore spurious biases V (such as background, color, etc.) irrelevant to the task. Conversely, deep neural networks often fit both the true causal relationship S→Y and the spurious correlation V→Y simultaneously during training, resulting in poor model robustness to input perturbations or scene changes. Therefore, based on the above variable definitions, this embodiment constructs a causal graph model of the neural network prediction process (…). Figure 3 This method formalizes the prediction process of deep neural networks and suppresses the unstable effects introduced by spurious bias variables by applying causal intervention operations to the input variables. The causal intervention mechanism blocks the "backdoor effect" of V on the S→Y path, achieving decoupling and control of the true causal factors. To eliminate the backdoor path formed between spurious bias variables and confounding variables, this embodiment employs a backdoor adjustment strategy, adjusting the input baseline physical scene image without altering the physical scene structure. A virtual intervention was performed to obtain the output distribution under causal intervention conditions, and the corresponding relationship is shown below: , In the formula, This indicates the presence of spurious bias variables. Causal factors variables under the conditions of intervention The distribution, This indicates a mixed variable. Can be quantified When intervening The changes are unaffected by the backdoor effect.
[0042] Based on this, this embodiment constructs a causal intervention loss function for generating optical stealth perturbations by comparing the response differences of output causal factor variables under different causal intervention conditions. Specifically, the causal effect under intervention is calculated through a backdoor adjustment formula, and the difference between the output distributions of causal factor variables under interventions of a baseline physical scene image and adversarial textures is measured based on the maximum average difference. The calculation formula is as follows: , in, Therefore, it can be further rewritten as: ,in, It is a measure of the distance between two distributions; Represents a given baseline physical scene image and Time-cause factor variables The probability distribution (denoted as) ); Indicates that in a given adversarial texture and Time-cause factor variables The probability distribution (denoted as) ).
[0043] By minimizing the causal intervention loss function, the generated adversarial texture can be guided to stably influence the neural network model's dependence on false biases while maintaining the target structural features, thereby achieving consistent optical stealth effects under different environmental conditions.
[0044] Since the output distribution under causal intervention conditions is difficult to obtain directly through analytical methods, this embodiment uses the maximum mean difference (MMD) as a measure of probability distribution difference. By sampling the causal factors of the output under both reference and perturbation input conditions, and calculating the similarity between samples based on the Gaussian kernel function, the MMD loss used to measure the difference between the two types of output distributions is obtained. Specifically, from the probability distribution... Extraction A sample, denoted as From the probability distribution Extraction A sample, denoted as And based on this, construct the kernel matrix: , in, , , The kernel matrix records the similarity of sample points within the probability distribution and between two probability distributions, respectively. It is a Gaussian kernel, used to map low-dimensional samples to high-dimensional space; , Denotes the specific set of samples drawn from the corresponding probability distribution, where and This is the index of the sample. Next, substituting the kernel matrix above into the definition of MMD, we calculate the square of MMD: , In the formula, For elements in the kernel matrix, For the kernel matrix rows, For the core matrix columns.
[0045] Finally, the square root of the equation is taken to measure the difference between the two probability distributions. Therefore, the causal intervention loss is expressed as: .
[0046] By incorporating the causal intervention loss into the optimization process of optical stealth perturbation, the generated perturbation texture can be guided to have a stable influence on the response introduced by non-causal factors in the network prediction without changing the target structural features, thereby improving the stability and consistency of optical stealth effect under different models and different physical environment conditions.
[0047] Furthermore, to prevent the generation of adversarial textures that cannot be projected, a projection compensation loss is introduced. Given that the brightness of the projector is constant, this loss focuses only on optimizing the color channels, which differs from existing methods that simultaneously optimize the brightness and color channels. By penalizing the extreme values within the color channels, projectable adversarial textures can be generated effectively.
[0048] Because the luminance and chrominance channels in the RGB color model are merged, it is impossible to calculate the chrominance channels separately. In contrast, the LAB color model consists of L (luminance channel), A (red / green channel), and B (yellow / blue channel), which allows for convenient calculation of each chrominance channel individually. Therefore, this invention uses the standardized LAB color space model to replace the traditional RGB model, achieving decoupled modeling of luminance and chrominance. Specifically, it first involves adversarial textures... Convert from the RGB color model to the LAB color model. The L / A / B channel components in the LAB color model can be obtained in the following ways: , In the formula, This represents the color model conversion function. , , These represent adversarial textures in the LAB color model. The output of the corresponding three channels.
[0049] Given the constant brightness of the projector and the uncertainty of the physical scene, this invention constrains extreme values to enable adversarial textures. (Perturbation) can be projected onto and There exist extrema in the equation. Therefore, the projection compensation loss function is expressed as: , In the formula, The projection compensation loss function is... and To counteract the color channel components of textures in LAB space, and Preset thresholds for channels A and B, It is a function with maximum value. It is the square of the L2 norm.
[0050] This invention addresses specific natural variations in non-invasive optical cloaking technology, including subtle changes in projector brightness and minor noise introduced by the camera. By introducing a loss of sharpness, robustness against attacks on these natural variations is enhanced.
[0051] Assuming that distribution 𝒟 contains all possible natural transformations in physical space, this invention extracts independent and identically distributed training datasets from distribution 𝒟. ,pass Choose robust adversarial textures. Combining the training dataset This is achieved by introducing a loss of sharpness. Upper limit of sharpness: , in, , , It's a hyperparameter. For disturbance, The disturbance radius is... It is an L2 norm.
[0052] The terms in square brackets in the equation are measured from... Calculated by the rate at which the loss increases when moving to a nearby value. Sharpness. Robustness is achieved through the following methods: , in, , , , Let be the sharpness-aware loss function, representing the loss function applied to the dataset. Above calculations, regarding adversarial textures SAM loss; through perturbation One around By finding the nearest neighbor value that maximizes the loss within the norm sphere, the sharpness of the loss can be calculated, thereby improving robustness.
[0053] S3. In dynamic scenes, continuously acquire video frames as real-time background information. Based on the real-time background information, use the trained neural network model to update the comprehensive and optimized adversarial texture in real time, and continuously project the updated optical projection texture onto the surface of the moving target to maintain a stable optical stealth effect in dynamic environments.
[0054] In this embodiment, to ensure that the training and test sets satisfy the independent and identically distributed condition, an attack strategy adaptable to arbitrary targets and backgrounds is proposed. (Collection Thread) It consists of multiple consecutive frames, with the content between adjacent frames remaining coherent. Utilizing the similarity between consecutive frames, this invention defines these frames as a training dataset and optimizes them to minimize the total loss function. Simultaneously, the generated adversarial textures are projected onto the attack scene, which is then captured by the acquisition thread. Continuous data acquisition. The total loss function is a weighted sum of the projection compensation loss function, the sharpness-perceived loss function, and the texture norm regularization term, expressed as: , In the formula, For the total loss function, , and These are the projection compensation loss functions. Sharpness-perceived loss function The weight parameters of the texture norm regularization term, It is the square of the L2 norm. For adversarial textures.
[0055] The optimized adversarial texture is mapped onto the surface of the target object in real time through a projection device. The acquisition thread continuously collects and verifies the attack effect, and iterates and continuously optimizes the adversarial texture to achieve dynamic adaptability to any target and background.
[0056] Based on the same inventive concept, the embodiment also provides a dynamic target optical stealth device based on causal reasoning, including a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to implement the dynamic target optical stealth method based on causal reasoning when the computer program is executed.
[0057] To verify the effectiveness, stability, and model generalization ability of the dynamic target optical stealth method and device based on causal reasoning proposed in this invention in real physical environments, this embodiment selects several representative deep neural network target detection models as attack targets and compares the method of this invention with existing non-invasive optical stealth technologies. The comparison methods include OPAD, SPAA, and TIAPT, all of which are typical optical-physical adversarial attack schemes in the prior art.
[0058] Experimental results are as follows Figure 4 As shown, the attack success rates on three target detection models with significantly different structures—RetinaNet, YOLO V3, and Mask R-CNN—are compared. The experimental data demonstrates that the method of this invention achieves significantly higher attack success rates than existing techniques in all three models, and there is no significant performance degradation with changes in model structure.
[0059] Specifically, on the RetinaNet detection model, the attack success rate of existing techniques is concentrated in the 40%–45% range, indicating a clear upper limit to their attack capability against this type of single-stage detection model. However, after adopting the technical solution of this invention, the attack success rate increases to over 80%, verifying the technical effectiveness of this invention in stably destroying the discrimination capability of the detection model even under complex physical conditions. In the Yolo V3 detection model, the attack effectiveness of existing methods fluctuates significantly with changes in model structure, while the method of this invention maintains a high and stable attack success rate, indicating its good adaptability to different network structures. Furthermore, on the Mask R-CNN detection model, which has a more complex structure and higher requirements for feature robustness, the attack success rate of some existing techniques decreases significantly, while the method of this invention remains at a high level, without showing obvious failure.
[0060] Comparative analysis confirms that the method of this invention exhibits highly consistent attack effects across different detection models, and its attack success rate is significantly higher than that of existing technologies in multiple models. This indicates that the present invention does not perform overfitting optimization for specific network structures, but rather weakens the dependence of deep neural networks on false associations through a causal intervention mechanism, fundamentally improving the stability and generalization ability of optical stealth attacks in cross-model scenarios.
[0061] Further analysis of the technical solution of this invention reveals that the above experimental results mainly stem from the synergistic effect of the following technical features: First, by constructing an intervention loss function based on causal reasoning, the backdoor path caused by non-causal factors such as background and color during the prediction process of deep neural networks is explicitly blocked, enabling the generated perturbation texture to produce a consistent misleading effect in different models; Second, by introducing ProjectNet based on a latent diffusion model and combining incremental learning and elastic weight merging strategies, a high-precision approximation of the real projection-acquisition process is achieved, thereby ensuring the transferability of the perturbation in physical space; Third, through the joint optimization of the real-time attack strategy and robust loss, the perturbation can still maintain a stable attack effect under dynamic scenes, lighting changes, and noise interference conditions.
[0062] In summary, the experimental results fully demonstrate that the present invention has significant technical advantages in the field of non-invasive dynamic target optical stealth. It can achieve stable, efficient and well-generalized optical stealth effects in various target detection models and complex physical environments, providing a reliable technical means for security assessment and defense mechanism design of intelligent monitoring systems.
[0063] Thus, the optical-physical adversarial attack technology and device based on causal reasoning has been completed. First, a lightweight projection-acquisition approximation model is constructed, combining incremental learning and elastic weight merging strategies to accurately simulate the optical processes of actual projection and camera imaging while preserving the model's original task generalization performance. Second, a multi-loss collaborative optimization framework is introduced into the perturbation optimization, including causal intervention loss, projection realizability loss (based on LAB color space gamut constraints), and robustness loss formed by expectation transformation and loss sharpness optimization. This simulates various natural physical transformations such as brightness changes, perspective shifts, complex backgrounds, and noise, while maintaining the stability of the perturbation under multiple environments. Finally, through a real-time attack strategy, the optimized perturbation texture is projected onto the surface of a target object in the monitoring scene, continuously acquired and verified by the camera. Feedback loops are used to dynamically adjust the optimization parameters, achieving non-contact optical-physical domain adversarial attacks against arbitrary targets and backgrounds. This concept significantly outperforms existing technologies in terms of attack effectiveness, real-time performance, robustness, and interpretability, and provides practical technical support for the security assessment and defense design of intelligent monitoring systems.
[0064] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A dynamic target optical stealth method based on causal reasoning, characterized in that, Includes the following steps: Step 1: Collect baseline physics scene images, adversarial textures, and physics scene images with adversarial textures overlaid as training datasets, and train a neural network model based on a latent diffusion model. Step 2: Construct a causal graph model for the deep neural network prediction task. By causally intervening in the spurious bias variables in the causal graph model, causal guidance is provided for the generation of adversarial textures. In the same iteration, a comprehensive optimization including projectibility physical constraints and environmental robustness constraints is applied to the adversarial textures. Step 3: Continuously acquire video frames in dynamic scenes as real-time background information. Based on the real-time background information, use the trained neural network model to update the comprehensive and optimized adversarial texture in real time, and continuously project the updated optical projection texture onto the surface of the moving target to maintain a stable optical cloaking effect in dynamic environments.
2. The dynamic target optical stealth method based on causal reasoning according to claim 1, characterized in that, In step 1, the neural network model based on the latent diffusion model includes: an image fusion module, an encoder module, and a computation module, which are used to approximate the real projection acquisition process; The image fusion module receives a baseline physical scene image and adversarial texture from the training dataset as input, and performs feature fusion using a preset fusion function to obtain an intermediate fused image; wherein the preset fusion function is expressed as follows: , For intermediate fusion images, For fusion function, As a reference physical scene image, For adversarial textures; The encoder module is used to extract features and perform dimensionality reduction encoding on the intermediate fused image to obtain a latent feature vector; The computation module is configured as a denoising backbone network U-Net, which performs denoising prediction in the latent space using latent feature vectors as input, to obtain a physical scene image superimposed with adversarial textures. To simulate physical projection transformation.
3. The dynamic target optical stealth method based on causal reasoning according to claim 2, characterized in that, An incremental learning mechanism is adopted and combined with the elastic weight consolidation algorithm to establish a first joint loss including physical approximation loss and elastic weight merging regularization loss, which is used to optimize the neural network model parameters and obtain the trained neural network model. Physical approximation loss is expressed as ;in, For the sampled Gaussian noise, For the denoising backbone network U-Net, For the time step of the diffusion process, It is the square of the L2 norm, used to calculate the Euclidean distance between the predicted noise and the actual noise; Elastic weight pooling regularization loss is expressed as ,in It is the first of the Fisher information matrix. One element, These are pre-trained weights. The weights are after fine-tuning. It is a regularization hyperparameter.
4. The dynamic target optical stealth method based on causal reasoning according to claim 1, characterized in that, The causal graph model described in step 2 includes causal factor variables, spurious bias variables, output variables, and confounding variables; The causal factor variables represent input features that have a direct causal relationship with the neural network model prediction task, corresponding to the essential structural features of the target object in the optical stealth task; The spurious bias variable represents an input feature that has no direct causal relationship with the neural network model's prediction task but can affect the prediction result, corresponding to the visual pattern of adversarial textures. The output variable represents the prediction result of the neural network model; The confounding variable refers to a factor that simultaneously affects both causal and spurious bias variables, corresponding to the current physical environment characteristics in optical stealth, including at least one of ambient lighting conditions and imaging ensemble location.
5. The dynamic target optical stealth method based on causal reasoning according to claim 4, characterized in that, The causal intervention described in step 2 involves performing an intervention operation on the spurious bias variable to block the backdoor path formed by the confounding variable and the causal factor variable; Among them, the intervention operation The causal effect is calculated using a backdoor adjustment formula: In the formula, This indicates the presence of spurious bias variables. Causal factors variables under the conditions of intervention The distribution, This indicates a mixed variable.
6. The dynamic target optical stealth method based on causal reasoning according to claim 5, characterized in that, The causal effect under intervention is calculated by adjusting the formula through a backdoor, and the difference between the output distribution of causal factor variables under interventions in the baseline physical scene image and adversarial texture is measured based on the maximum mean difference, so as to construct a causal intervention loss function. The calculation formula is as follows: , In the formula, For causal intervention loss function, As a reference physical scene image, It is the maximum mean difference measure, used to measure the difference between two probability distributions. Indicates that in a given and hour The probability distribution, Indicates that in a given adversarial texture and hour The probability distribution.
7. The dynamic target optical stealth method based on causal reasoning according to claim 1, characterized in that, The projectibility constraint is achieved by converting the adversarial texture from the RGB color space to the LAB color space and constraining the numerical ranges of the A and B channels in the LAB color space. The projection compensation loss function corresponding to the projectibility constraint is expressed as follows: , In the formula, The projection compensation loss function is... and To counteract the color channel components of textures in LAB space, and Preset thresholds for channels A and B, It is a function with maximum value. It is the square of the L2 norm.
8. The dynamic target optical stealth method based on causal reasoning according to claim 1, characterized in that, The environmental robustness constraint is achieved by minimizing the sharpness-perceived loss function, which is used to improve the robustness of the adversarial texture to natural transformations in physical space. The corresponding sharpness-perceived loss function is expressed as follows: , In the formula, Let be the sharpness-aware loss function, representing the loss function applied to the dataset. Above calculations, regarding adversarial textures SAM loss; This indicates that the constraints are satisfied. disturbance Find the maximum sharpness perception loss function in the process; For the Euclidean norm, The disturbance radius is... This is the loss function for causal intervention.
9. The dynamic target optical stealth method based on causal reasoning according to claim 8, characterized in that, The optimized adversarial texture is achieved by minimizing the total loss function, which is a weighted sum of the projection compensation loss function, the sharpness-perceived loss function, and the texture norm regularization term, expressed as: , In the formula, For the total loss function, , and These are the projection compensation loss functions. Sharpness-perceived loss function The weight parameters of the texture norm regularization term, It is the square of the L2 norm. For adversarial textures.
10. A dynamic target optical stealth device based on causal reasoning, comprising a memory and a processor, wherein the memory is used to store a computer program, characterized in that, The processor is configured to implement the dynamic target optical stealth method based on causal reasoning as described in any one of claims 1 to 9 when executing the computer program.