A method and apparatus for detecting occlusion in dynamic degradation decomposition combined with physical restoration

By employing dynamic degradation decomposition and physical repair methods, the problem of low PPE detection accuracy in high-risk industrial scenarios of traditional CNNs has been solved, achieving higher detection accuracy and safety.

CN121837642BActive Publication Date: 2026-05-26YANAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
YANAN UNIV
Filing Date
2026-03-12
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Traditional CNNs are prone to low detection accuracy and safety risks when detecting industrial personal protective equipment (PPE) in high-risk industrial scenarios such as chemical and construction industries, due to background noise drowning out features and occlusion.

Method used

A method combining dynamic degradation decomposition and physical restoration is adopted. Multi-scale features are extracted through the dynamic degradation decomposition backbone network, decoupling and restoration are performed using an atmospheric scattering model, and local features of the feature map are enhanced through a context relationship module. Finally, a decoupled detection head is used for accurate prediction.

Benefits of technology

It effectively solves the problem of feature blurring caused by background noise and occlusion, improves the accuracy and safety of PPE detection, and provides detection assurance in complex industrial environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121837642B_ABST
    Figure CN121837642B_ABST
Patent Text Reader

Abstract

This application discloses an occlusion detection method and apparatus combining dynamic degradation decomposition and physical restoration, relating to the fields of computer vision and industrial safety monitoring technology. The method includes: extracting multi-scale features from industrial scene images using a dynamic degradation decomposition backbone network; explicitly decomposing and suppressing the dynamic degradation patterns of the industrial scene images in manifold space; generating degradation correction prompts and strategy prompts; inputting these prompts into a physically guided dehazing neck network; and using the physical inductive bias of an atmospheric scattering model to decouple and restore the multi-scale features, generating a multi-scale restored feature map; and using a context relation module to perform semantic reconnection operations on the multi-scale restored feature map to generate a context-enhanced feature map, which is then input into a decoupled detection head to predict the location and category of industrial personal protective equipment (PPE) in the industrial scene image. This solves the problem of low accuracy in detecting the wearing of industrial PPE in high-risk industrial scenes, which poses safety risks in industrial production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer vision and industrial safety monitoring technology, and in particular to an occlusion detection method and device that combines dynamic degradation decomposition with physical repair. Background Technology

[0002] With the increasing safety standards in industrial production, the detection of industrial personal protective equipment (PPE) wearing is crucial in high-risk environments such as chemical plants and construction sites. However, actual industrial scenarios are often unstructured, facing severe challenges such as pervasive smoke and dust, drastic changes in lighting, and obstruction from machinery. In these environments, PPE targets (such as safety helmets and reflective vests) often exhibit blurred outlines, lack of texture, and severe scale differences, leading to a sharp decline in the performance of traditional convolutional neural network (CNN) detectors.

[0003] Existing technologies suffer from two fundamental drawbacks: First, features are easily overwhelmed by background noise. Traditional CNNs are essentially low-pass filters; when processing industrial images filled with smoke and low contrast, crucial high-frequency edges and texture details are treated as noise and filtered out, causing small target features to completely disappear in deep networks. Second, occlusion causes semantic breaks in targets. In partially occluded or smoky environments, the natural connection between an object and its context (e.g., the spatial relationship between a worker's head and a safety helmet) is severed, leading to blurred localization. This results in low accuracy in detecting the wearing of industrial personal protective equipment (PPE) in high-risk industrial settings such as chemical and construction industries, posing safety risks in industrial production. Summary of the Invention

[0004] In this embodiment of the application, by providing an occlusion detection method that combines dynamic degradation decomposition with physical repair, the problem of low accuracy in detecting the wearing of industrial personal protective equipment (PPE) in high-risk industrial scenarios such as chemical industry and construction is solved. This is because the features of small targets are easily submerged by background noise and disappear, and the semantic breaks of targets caused by occlusion lead to fuzzy positioning.

[0005] In a first aspect, embodiments of this application provide an occlusion detection method combining dynamic degradation decomposition and physical restoration. The method includes: receiving an industrial scene image; extracting multi-scale features of the industrial scene image using a dynamic degradation decomposition backbone network; explicitly decomposing and suppressing the dynamic degradation pattern of the industrial scene image in the manifold space; generating degradation correction prompts and strategy prompts; inputting the degradation correction prompts and strategy prompts into a physically guided dehazing neck network; using the physical inductive bias of an atmospheric scattering model to decouple and restore the multi-scale features, generating a multi-scale restored feature map; and performing semantic reconnection operations on the multi-scale restored feature map using a context relation module. The local features of the multi-scale repair feature map are enhanced to generate a context-enhanced feature map. The semantic reconnection operation includes: constructing a query, key, and value mapping for the multi-scale repair feature map to obtain appearance relation compatibility and spatial geometric location compatibility; generating attention weights based on appearance relation compatibility and spatial geometric location compatibility through a dual relation attention mechanism to enhance the local features of the multi-scale repair feature map and generate a context-enhanced feature map; the context-enhanced feature map is input into a decoupled detection head to predict the location and category of industrial personal protective equipment in industrial scene images; wherein, the location of industrial personal protective equipment is presented in the form of a prediction box.

[0006] In one possible implementation, the receiving of industrial scene images, extraction of multi-scale features from the industrial scene images using a dynamic degradation decomposition backbone network, and explicit decomposition and suppression of dynamic degradation patterns of the industrial scene images in the manifold space, generating degradation correction prompts and strategy prompts, includes: the dynamic degradation decomposition backbone network processes the industrial scene images through a cross-domain degradation analyzer, explicitly decomposing and suppressing the dynamic degradation patterns of the industrial scene images in the manifold space using the cross-domain degradation analyzer, generating degradation correction prompts and strategy prompts; each level of the dynamic degradation decomposition backbone network processes the features of the previous level through a dynamic decomposition mechanism, and the dynamic degradation decomposition backbone network... The calculation process for each level is as follows: ;in, For the first Features of each level of output For the first The gating weights of each level of output, Represents the dynamic degradation decomposition backbone network. The dynamic decomposition mechanism operates at each level. This is a feature of the next higher level. For degradation correction prompts, This is a strategy suggestion.

[0007] In one possible implementation, the degradation correction prompts and strategy prompts are input into a physically guided dehazing neck network. Using the physical inductive bias of the atmospheric scattering model, multi-scale features are decoupled and repaired to generate a multi-scale repaired feature map. This includes: the physically guided dehazing neck network comprising a top-down path for physically guided repair and a bottom-up path for spatial detail injection; the physically guided repair process includes: using a physically guided decoupling module, based on the atmospheric scattering model, decoupling the multi-scale features into haze-related features and content features, and then recombining them. ;in, For the first The repaired features output by the layer physical guidance decoupling module. Indicates the first Operation functions of the layer physical guidance decoupling module. Indicating degenerate trunk features, Indicates the first Repair features output by the layer physical guidance decoupling module. Indicates an upsampling operation. For degradation correction prompts, As a strategy suggestion, the repaired features output by the physical guidance decoupling module are decoupled into clear content features, transmission rate map features, and atmospheric light features using the physical inductive bias of the atmospheric scattering model. Feature recombination is achieved through the inverse transformation of the atmospheric scattering model to realize the spatial detail injection process. ;in, For multi-scale repair feature maps, To clearly define the content characteristics, For transmission rate map features, It is an atmospheric light characteristic.

[0008] In one possible implementation, the expression for appearance relationship compatibility is: ;in, To query pixels Corresponding query features With key pixels Corresponding context features Compatibility of appearance relationships between them This is the first learnable weight matrix. This is the second learnable weight matrix. For query features The function representation, Contextual features The function representation of ; the expression for spatial geometric location compatibility is: ;in, To query pixels Corresponding query features With key pixels Corresponding context features Spatial geometric position compatibility between them This is the third learnable weight matrix. For use in calculation and A function of the spatial geometric position between them This is the activation function.

[0009] In one possible implementation, the step of generating attention weights based on appearance relationship compatibility and spatial geometric location compatibility through a dual-relationship attention mechanism to enhance local features of the multi-scale repair feature map and generate a context-enhanced feature map includes: a context relationship module fusing appearance relationship compatibility and spatial geometric location compatibility through a Softmax function to generate attention weights; the expression for the attention weights is: ;in, For attention weights, It is an exponential function. Key pixels Corresponding context features To query pixels Corresponding query features With key pixels Corresponding context features Spatial geometric position compatibility between them To query pixels Corresponding query features With key pixels Corresponding context features The appearance relationship compatibility between them is assessed; attention weights are used to weight and aggregate the value features to extract the aggregated global context information; the aggregated global context information is then residually connected with the original features of the multi-scale insulated feature map to enhance the local features of the multi-scale insulated feature map, thereby generating a context-enhanced feature map, expressed as: ;in, For context-enhanced feature maps, The original features of the multi-scale repair feature map, This is the fourth learnable weight matrix. This is the aggregated global context information. The spatial dimensions of the multi-scale repair feature map are given by a height of [value missing]. Width is , For attention weights, For key pixels, Key pixels The value characteristics at that location.

[0010] One possible implementation also includes: using a quality-aware dynamic focusing loss function to perform detection optimization on the overall model framework during end-to-end training. The quality-aware dynamic focusing loss function uses the intersection-union ratio of the predicted boxes as a quality indicator and dynamically adjusts the gradient weights of easy and difficult samples. The overall model framework includes a dynamically degenerate decomposition backbone network, a physically guided dehazing neck network, a context relation module, and a decoupled detection head.

[0011] In one possible implementation, the quality-aware dynamic focus loss function includes zoom loss, Wise-IoU loss, and distributed focus loss; the expression for the quality-aware dynamic focus loss function is: ;in, For quality-perceived dynamic focusing loss function, For the first Index of positive samples, For classification confidence, For zoom loss, To predict soft tags, To dynamically focus weights, For Wise-IoU loss, Predict bounding boxes, True bounding box, This is the distributed loss balance coefficient. For distributed focusing loss, To predict the bounding box coordinate distribution, The coordinate distribution of the true bounding box; , ;in, These are dynamically adjusted gradient weight coefficients. For outlier degree, As the first hyperparameter, This is the second hyperparameter. For the current crossover loss, The moving average is used; the zoom loss uses the intersection-union ratio of the predicted bounding boxes as the predicted soft label. To weighted classification confidence The expression is: ;in, For zoom loss, For balance coefficient, This is the adjustment coefficient.

[0012] Secondly, embodiments of this application provide an occlusion detection device combining dynamic degradation decomposition and physical restoration. The device includes: a prompt generation module for receiving an industrial scene image, extracting multi-scale features of the industrial scene image using a dynamic degradation decomposition backbone network, explicitly decomposing and suppressing the dynamic degradation mode of the industrial scene image in the manifold space, and generating degradation correction prompts and strategy prompts; a multi-scale restoration feature map generation module for inputting the degradation correction prompts and strategy prompts into a physically guided dehazing neck network, decoupling and restoring the multi-scale features using the physical inductive bias of an atmospheric scattering model, and generating a multi-scale restoration feature map; a semantic reconnection operation including: constructing a query, key, and value mapping for the multi-scale restoration feature map, obtaining appearance relation compatibility and spatial geometric location compatibility; generating attention weights based on appearance relation compatibility and spatial geometric location compatibility through a dual relational attention mechanism to enhance the local features of the multi-scale restoration feature map and generate a context-enhanced feature map; and a prediction module for inputting the context-enhanced feature map into a decoupled detection head to predict the location and category of industrial personal protective equipment in the industrial scene image; wherein the location of the industrial personal protective equipment is presented in the form of a prediction bounding box.

[0013] Thirdly, embodiments of this application provide an occlusion detection server that combines dynamic degradation decomposition with physical repair, including a memory and a processor; the memory is used to store computer-executable instructions; the processor is used to execute the computer-executable instructions to implement the method described in the first aspect or any possible implementation of the first aspect.

[0014] Fourthly, embodiments of this application provide a computer-readable storage medium storing executable instructions, which, when executed by a computer, enable the method described in the first aspect or any possible implementation thereof.

[0015] One or more technical solutions provided in this application embodiment have at least the following technical effects: This application embodiment provides an occlusion detection method combining dynamic degradation decomposition and physical restoration. It receives industrial scene images, extracts multi-scale features of the industrial scene images using a dynamic degradation decomposition backbone network, and explicitly decomposes and suppresses the dynamic degradation patterns of the industrial scene images in the manifold space, generating degradation correction prompts and strategy prompts. This process effectively solves the image degradation problem caused by factors such as smoke and dust and drastic changes in lighting in industrial scenes, providing a high-quality foundation for subsequent feature processing. The physically guided dehazing neck network uses the physical inductive bias of the atmospheric scattering model to deeply decouple and restore multi-scale features, generating multi-scale restored feature maps. This network fully utilizes physical laws to effectively remove interference factors such as fog and blur in the image, restoring the clarity and integrity of the features, and further improving feature quality. The context relationship module constructs query, key, and value mappings to obtain appearance relationship compatibility and spatial geometric position compatibility, and uses a dual relationship attention mechanism to generate attention weights, performing semantic reconnection operations on the multi-scale restored feature maps to generate context-enhanced feature maps. This operation enhances the local features of the feature map, effectively solving the problem of semantic fragmentation caused by occlusion, enabling the model to better understand the relationship between the target and the context, and improving the expressive power and semantic integrity of the features. The decoupled detection head, based on context-enhanced feature maps, can accurately predict the location and category of industrial personal protective equipment (PPE) in industrial scene images, with precise bounding box localization and high category recognition accuracy. This application demonstrates superior detection performance in complex industrial environments, providing strong technical support for industrial production safety. It solves the problems in existing technologies for PPE detection, such as features being easily submerged by background noise leading to the disappearance of small target features, and semantic fragmentation caused by occlusion resulting in blurred localization. These issues lead to reduced accuracy in detecting the wearing of industrial PPE in high-risk industrial scenarios such as chemical and construction industries, posing safety risks in industrial production. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments of this application or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart of an occlusion detection method combining dynamic degradation decomposition and physical repair, provided as an embodiment of this application.

[0018] Figure 2 This is a schematic diagram of the overall architecture of this application provided for an embodiment of this application.

[0019] Figure 3 This is a schematic diagram of the architecture of the dynamic decomposition mechanism module provided in the embodiments of this application.

[0020] Figure 4 A visualization comparison of the detection results of DDP-YOLO on the SFCHD smoke dataset provided in the embodiments of this application.

[0021] Figure 5 A visualization comparison of the detection results of DDP-YOLO on the PPEs strongly occluded dataset provided in this application embodiment.

[0022] Figure 6 A visualization comparison of the detection results of DDP-YOLO on the PP02 complex industrial background dataset provided in this application embodiment.

[0023] Figure 7 This is a schematic diagram of an occlusion detection device that combines dynamic degradation decomposition with physical repair, provided in an embodiment of this application.

[0024] Figure 8 This is a schematic diagram of an occlusion detection server that combines dynamic degradation decomposition with physical repair, provided as an embodiment of this application. Detailed Implementation

[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0026] The following description of some technologies involved in the embodiments of this application is provided to aid understanding and should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, some descriptions of well-known functions and structures are omitted in the following description.

[0027] This application provides an occlusion detection method that combines dynamic degradation decomposition with physical restoration, such as... Figure 1 As shown, the method includes steps S101 to S104. Wherein, Figure 1 This is merely one execution order shown in the embodiments of this application, and does not represent the only execution order of an occlusion detection method that combines dynamic degradation decomposition with physical restoration. Where the final result can be achieved, [the following is possible]. Figure 1 The steps shown can be performed in parallel or in reverse order.

[0028] S101: Receives industrial scene images, extracts multi-scale features of the industrial scene images using a dynamic degradation decomposition backbone network, explicitly decomposes and suppresses the dynamic degradation patterns of the industrial scene images in the manifold space, and generates degradation correction prompts and strategy prompts.

[0029] The system receives industrial scene images, extracts multi-scale features of the industrial scene images using a dynamic degradation decomposition backbone network, explicitly decomposes and suppresses the dynamic degradation patterns of the industrial scene images in the manifold space, and generates degradation correction prompts and policy prompts, including the following:

[0030] The dynamic degradation decomposition backbone network processes industrial scene images through a cross-domain degradation analyzer. The cross-domain degradation analyzer explicitly decomposes and suppresses the dynamic degradation patterns of industrial scene images in the manifold space, generating degradation correction prompts and strategy prompts.

[0031] Specifically, the Dynamic Degradation Decomposition backbone network first processes the input industrial scene image using a Cross-Domain Degradation Analyzer (CDDA). CDDA does not analyze the image in a single domain, but rather deconstructs image features in depth from the perspective of frequency and spatial domain interaction on the image manifold. Through this cross-domain analysis, CDDA can accurately capture various dynamic degradation patterns in the image, such as image degradation caused by factors like smoke and dust, drastic changes in lighting, and occlusion by mechanical equipment. Based on these analysis results, CDDA generates two types of key prompts: degradation correction prompts and strategy prompts. Degradation correction prompts mainly guide the subsequent network's image feature restoration work, helping the network understand which parts of the image are degraded and the degree of degradation, thus enabling targeted correction. Strategy prompts are used to adjust the path and method of subsequent feature extraction, dynamically adjusting the feature extraction strategy according to different degradation conditions to better adapt to the complex and ever-changing industrial scene.

[0032] Each level of the dynamically degenerate decomposition backbone network processes the features of the previous level through a dynamic decomposition mechanism. The calculation process for each level is as follows: .in, For the first Features of each level of output For the first The gating weights of each level of output, Represents the dynamic degradation decomposition backbone network. The dynamic decomposition mechanism operates at each level. This is a feature of the next higher level. For degradation correction prompts, This is a strategy suggestion.

[0033] Previous level features This includes the image feature extraction and processing results from the previous layer, serving as the input information for the current layer. In the above operations, the previously generated degradation correction prompts are used to perform preliminary denoising on the input features, removing noise components introduced by degradation patterns and retaining more valuable feature information. Simultaneously, gating weights are generated based on strategy prompts. These gating weights can dynamically fuse features from different receptive fields. Features from different receptive fields contain information at different image scales. By adjusting the gating weights, the network can flexibly select and combine features of different scales during feature extraction, based on the actual image conditions, thus completing the initial removal of degradation factors such as smoke and noise while extracting features. Finally, the... The features and gating weights output from each level will be used as input for the next level of feature processing, and will continue to be passed and processed in the network.

[0034] Figure 2 This is a schematic diagram of the overall architecture of this application, provided for embodiments of this application. It covers the complete process from input RGB image (in this application, an industrial scene image) to final output, including input and analyzer, dynamic degradation decomposition backbone network (… The backbone, the physically-guided haze removal neck (PHATNeck), and the head section.

[0035] Dynamic degradation decomposition backbone network ( The Backbone integrates a Cross-Domain Degradation Analyzer (CDDA) and a Dynamic Degradation Mechanism (DDM). The input to the DDM backbone network is a convolutional layer, followed by multiple processing modules at different levels. Each level includes downsampling and DDM module operations, with the input image sequentially processed through multiple levels (e.g., L1 to L4). Each level first adjusts the size and number of channels of the input feature map, such as from... After downsampling, it becomes from Form. Among them, The height of the feature map, The width of the feature map. , This represents the number of channels in the feature map at different levels.

[0036] The main function of this Dynamic Decomposition Mechanism (DDM) is to decouple degenerate features, separating clean features from degenerate components. The entire process can be described as: input degenerate features... (dimension is) Enter the cross-chain block structure from the left. The structure contains degradation correction hints, representing the number of channels in the feature map. The two branches are "and policy hints," with the degradation correction hint branch proceeding through a series of g... n Layered operations such as Conv, LayerNorm, 1×1Conv, and LeakyReLU extract degradation pattern information, while the policy suggestion branch generates guiding policies through similar convolution and normalization operations. The outputs of these two branches are adaptively fused in the ADB module. The fused features then enter the DUs module for further processing. This module extracts multi-scale degradation information through operations such as Fusion fusion layers, multi-layer convolutional ReLU activation, MaxPool max pooling, and AdaptiveAvgPool adaptive average pooling, ultimately outputting a refined degradation representation. (Adaptive weights). In the feature refinement path on the right, the original degenerate features... First, a 2C-dimensional feature is obtained by concatenating the intermediate features using a Concat operation. Then, a depthwise separable convolution (DWConv) is used for efficient spatial feature extraction. Next, the feature is reduced in dimensionality by a Proj 2 / C projection layer, and then processed separately. (Semi-channel element-wise modulation) and The channel modulation operation (full-channel element-wise modulation) passes through the Proj 2 / C projection layer again and finally through... The Conv convolutional layer outputs clean features after decomposition, and the entire process decouples degraded features into clean scene features and separable degraded components. This provides a more accurate representation foundation for subsequent feature recovery and enhancement.

[0037] Figure 3 This is a schematic diagram of the architecture of the dynamic decomposition mechanism module provided in an embodiment of this application. Degradation features. The input to the module represents image features with degradation effects (such as blurring, noise, etc.), and the shape is... ,in This refers to the number of channels in an image exhibiting degradation characteristics. Degradation correction tips. Prompt messages are provided to guide degradation correction, and the auxiliary module processes degradation features. Strategy prompts. Provides policy guidance and prompts, influencing how the module operates. Degeneracy characteristics. The output of this module has been processed to remove some features affected by degradation. `gnConv` is a recursively gated convolution used to adjust the number of channels or perform simple feature transformations. `LayerNorm` is a layer normalization operation that normalizes features, accelerating model convergence and improving stability. `LeakyReLU` is an activation function. `1×1Conv` is a 1×1 convolution operation. The `ADB` module is an adaptive decomposition block. The `DUs` module represents a decomposition unit. `Fusion` is a feature fusion operation. `Concat` is a concatenation operation. `AdaptiveAvgPool` is an adaptive average pooling operation. `MaxPool` is a max pooling operation. `ReLU` is an activation function. `Conv` is a convolution operation. For adaptive weights, Proj 2 / C performs a projection operation, projecting features onto a specified number of channels. DWConv performs a depthwise convolution. This is element-wise multiplication. This is element-wise addition.

[0038] S102: Input the degradation correction prompts and strategy prompts into the physics-guided dehazing neck network, and use the physical inductive bias of the atmospheric scattering model to decouple and repair multi-scale features, generating multi-scale repaired feature maps.

[0039] The degradation correction and strategy prompts are input into the physics-guided dehazing neck network. The physical inductive bias of the atmospheric scattering model is used to decouple and repair multi-scale features, generating a multi-scale repaired feature map, including the following:

[0040] The physical-guided defogging neck network includes a top-down path for physical-guided repair and a bottom-up path for spatial detail injection.

[0041] The dual-path design of the Physically Guided Dehazing Neck Network (PHAT Neck) enables the network to process features from different angles. The top-down path focuses on using physical models for semantic repair, while the bottom-up path focuses on integrating shallow details into deep features to improve the integrity and accuracy of the features.

[0042] like Figure 2 As shown, the physically guided dehazing neck network receives feature maps from different levels. , , Output feature map ( , , O3, O4, and O5 are enhanced feature maps output at three different levels (scales), corresponding to the outputs of the detection heads for small, medium, and large targets, respectively. Feature processing is performed from top to bottom using a Physically Guided Decoupling Module (PHDT). As shown in the figure, the high-level feature maps first enter the PHDT module, while simultaneously incorporating degradation correction prompts (…). The PHDT module performs physical-guided dehazing. Features processed by the PHDT module are then concatenated with higher-level features that have undergone upsampling (Up) operations. This top-down approach allows the network to perform semantic-level repair on features based on a physical model (such as the physical inductive bias of an atmospheric scattering model), gradually passing high-level semantic information downwards and correcting feature blurring caused by degradation factors such as fog. The bottom-up path focuses on injecting spatial details. Low-level feature maps contain rich spatial details. The bottom-up path uses downsampling (Dn) to gradually integrate the details from shallow features into the deeper features repaired by the top-down path. For example, as shown in the figure, features processed and concatenated by the PHDT module are concatenated with other downsampled features during downward propagation, effectively injecting shallow details into deeper features and improving feature integrity and accuracy. The PHDT module, as the core module of the top-down path, performs decoupling and repair operations on input features based on a physical model, and is a key component for achieving physical-guided dehazing. Upsampling (Up) and downsampling (Dn) operations: Upsampling is used to adjust the size of the feature map so that it can be concatenated with features of different levels; downsampling is used to extract and pass on detailed information in shallow features to adapt to the processing needs of features of different levels.

[0043] The physical-guided repair process includes: using a physical-guided decoupling module, based on an atmospheric scattering model, decoupling multi-scale features into haze-related features and content features, and then recombining them. .in, For the first The repaired features output by the layer physical guidance decoupling module. Indicates the first Operation functions of the layer physical guidance decoupling module. Indicating degenerate trunk features, Indicates the first Repair features output by the layer physical guidance decoupling module. Indicates an upsampling operation. For degradation correction prompts, This is a strategy suggestion.

[0044] Using the physical inductive bias of the Atmospheric Scattering Model (ASM), the repaired features output by the physical guidance decoupling module are decoupled into clear content features, transmission rate map features, and atmospheric light features.

[0045] Feature reconstruction is achieved through the inverse transformation of the atmospheric scattering model to realize the spatial detail injection process: .in, For multi-scale repair feature maps, To clearly define the content characteristics, For transmission rate map features, It is an atmospheric light characteristic.

[0046] This inverse transformation process simulates the propagation of light in a haze-free environment, recombining the separated clear content features, transmission rate map features, and atmospheric light features to remove the effects of degradation factors such as haze and restore image clarity. Simultaneously, by incorporating a bottom-up path spatial detail injection operation, detailed information from shallow features is integrated into the recombined features. This results in a multi-scale repaired feature map that possesses both clear semantic information and rich spatial details, providing high-quality feature input for subsequent semantic reconnection and object detection.

[0047] The core function of this Physically Guided Decoupling (PHDT) Block is to decouple and refine features using prior physical knowledge. The entire process can be described as follows: Input features First, we enter the physical guidance decoupler section, which contains two parallel encoder branches: the Atmospheric Light Encoder (ALE) and the Transmission Map Encoder (TME). These two encoders extract correction prompts and corresponding physical parameter information, respectively. In the TCSC module, through... Feature processing, ALE encoder output atmospheric light parameters TME encoder output transmission parameters Then, these two parameters are fused and transformed through a 1×1 convolutional layer. The entire decoupling process follows the mathematical expression of the physical scattering model. ,in The transmission map is used to adjust the degree of feature attenuation. Atmospheric light is used to remove the influence of ambient light through a subtraction operation. After removing the atmospheric light component, the transmission map is then processed. The weighted average is then compared with the original input. Residual connections are performed to obtain refined output features. This design decouples features into scene information and environmental interference information by introducing prior knowledge from a physical scattering model, thereby achieving more robust feature representation and enhancement effects.

[0048] The Physically Guided Decoupling Module (PHDT Block) includes the Atmospheric Light Encoder (ALE) and the Transmission Map Encoder (TME). The ALE receives degradation correction cues as input and outputs atmospheric light-related features for feature decoupling and reconstruction. The TME also takes degradation correction cues as input and outputs transmission map features to help the module understand how light travels in the image.

[0049] S103: Utilize the context relation module to perform semantic reconnection on the multi-scale repair feature map, enhance the local features of the multi-scale repair feature map, and generate a context-enhanced feature map.

[0050] The main function of this CRM context association module is to integrate features of the target focus area. and global context features This enhances feature representation through an attention mechanism. The entire process can be described as follows: In the target focusing path of the upper branch, the input target focusing region features... First, through two parallel... The convolutional layer performs feature transformation to generate query features for subsequent attention calculations; simultaneously, in the global context path of the next branch, global context features are... The features are fused through element-wise multiplication of the target and contextual information after extraction of highly photosensitive regions. Next, the relational attention unit calculates the relationship matrix between the target and contextual features, and normalizes it using the Softmax function to generate an attention map to guide feature selection. In the bottom fusion stage, the module provides three element-wise operation methods: element-wise multiplication to emphasize relevance, element-wise addition for feature accumulation, and element-wise subtraction for difference modeling. Finally, the fused features are mapped to weights to generate the final output features. The entire process effectively combines local target information with global contextual information through the attention mechanism, thereby improving the discriminative power of the features.

[0051] The semantic reconnection operation includes: constructing query, key, and value mappings for multi-scale repair feature maps, and obtaining appearance relation compatibility and spatial geometric location compatibility. Through a dual relational attention mechanism, attention weights are generated based on appearance relation compatibility and spatial geometric location compatibility to enhance the local features of the multi-scale repair feature maps and generate context-enhanced feature maps.

[0052] like Figure 2 As shown, the dual relational attention mechanism in this application is a cross-attention mechanism. Q is the query, K is the key, and V is the value. I represents the input feature map, i.e., the multi-scale repaired feature map. For frequency domain features, Conv is the convolution operation, and MaxPooling is the maximum pooling operation.

[0053] The expression for appearance compatibility is: ;in, To query pixels Corresponding query features With key pixels Corresponding context features Compatibility of appearance relationships between them This is the first learnable weight matrix. This is the second learnable weight matrix. For query features The function representation, Contextual features The function representation of .

[0054] The expression for spatial geometric location compatibility is: ;in, To query pixels Corresponding query features With key pixels Corresponding context features Spatial geometric position compatibility between them This is the third learnable weight matrix. For use in calculation and A function of the spatial geometric position between them This is the activation function.

[0055] By employing a dual relational attention mechanism, attention weights are generated based on appearance relation compatibility and spatial geometric location compatibility to enhance local features of multi-scale insulated feature maps and generate context-enhanced feature maps, including the following:

[0056] The Contextual Relation Module (CRM) uses the Softmax function to combine appearance relationship compatibility and spatial geometric location compatibility to generate attention weights.

[0057] The expression for attention weights is: ;in, For attention weights, It is an exponential function. Key pixels Corresponding context features To query pixels Corresponding query features With key pixels Corresponding context features Spatial geometric position compatibility between them To query pixels Corresponding query features With key pixels Corresponding context features Appearance compatibility between them.

[0058] Attention weights are used to perform weighted aggregation of value features, and the aggregated global context information is extracted.

[0059] The aggregated global context information is residually concatenated with the original features of the multi-scale insulated feature map to enhance the local features of the multi-scale insulated feature map, thereby generating a context-enhanced feature map. The expression is as follows: .in, For context-enhanced feature maps, The original features of the multi-scale repair feature map, This is the fourth learnable weight matrix. This is the aggregated global context information. The spatial dimensions of the multi-scale repair feature map are given by a height of [value missing]. Width is , For attention weights, For key pixels, Key pixels The value characteristics at that location.

[0060] S104: Input the context-enhanced feature map into the decoupled detection head to predict the location and category of industrial personal protective equipment (PPE) in the industrial scene image. The location of the PPE is presented as a predicted bounding box.

[0061] Specifically, the decoupled detection head includes a classification branch and a regression branch. The industrial personal protective equipment (PPE) in this application is the target.

[0062] This application also includes the following.

[0063] A quality-aware dynamic focusing loss function is employed to optimize the overall model framework during end-to-end training. This loss function uses the intersection-union ratio (IUU) of predicted bounding boxes as a quality metric to dynamically adjust the gradient weights for easy and difficult samples. The overall model framework comprises a dynamically degenerate backbone network, a physically guided dehazing neck network, a contextual relationship module, and a decoupled detection head.

[0064] Specifically, the overall model framework of this application is the DDP-YOLO framework.

[0065] The quality-perceived dynamic focus loss function includes zoom loss, Wise-IoU loss, and distributed focus loss.

[0066] Specifically, this application designs a novel multi-task joint loss function, the Quality-Aware Dynamic Focusing Loss (QDFL). This loss function not only includes standard classification loss and bounding box regression loss, but also introduces an adaptive weighting mechanism to enhance the model's learning ability for blurred, occluded, or small target samples. Specifically, for each positive sample... The expression for the quality-perceived dynamic focusing loss function is as follows. Wherein, This represents the set of indices for all positive samples in the current batch.

[0067] The expression for the quality-perceived dynamic focusing loss function is: .in, For quality-perceived dynamic focusing loss function, For the first Index of positive samples, For classification confidence, For zoom loss, To predict soft tags, To dynamically focus weights, For Wise-IoU loss, Predict bounding boxes, True bounding box, This is the distributed loss balance coefficient. For distributed focusing loss, To predict the bounding box coordinate distribution, This represents the true bounding box coordinate distribution.

[0068] Specifically, the Wise-IoU loss (WIoU) introduces a non-monotonic focusing mechanism. It calculates outlier. This value reflects the quality of the match between the predicted bounding box and the ground truth bounding box.

[0069] , .in, These are dynamically adjusted gradient weight coefficients. For outlier degree, As the first hyperparameter, This is the second hyperparameter. For the current crossover loss, It is a moving average.

[0070] Specifically, for high-quality samples ( Small), Reduce to prevent overfitting; for low-quality samples ( Extremely high (e.g., due to noise or severe occlusion). Similarly, reduce to suppress harmful gradients; while for moderately high-quality "difficult but learnable" samples, Maintain a high value to enhance learning.

[0071] Zoom loss uses the intersection-over-union ratio of the predicted bounding boxes as the predicted soft label. To weighted classification confidence The expression is: ;in, For zoom loss, For balance coefficient, This is the adjustment coefficient.

[0072] Specifically, the zoom loss uses the intersection-over-union ratio (IoU) as a soft label to train the classification branch, ensuring that high-confidence detection boxes also have high localization accuracy. The distributed focusing loss is used to handle blurred boundaries, transforming the regression problem into general distribution learning, adapting to the edge uncertainty caused by smoke.

[0073] To verify the effectiveness of the proposed occlusion detection method (DDP-YOLO) based on dynamic degradation decomposition combined with physical restoration, three representative benchmark datasets, SFCHD (smoky environment), PPEs (severe occlusion), and PP02 (complex background), were selected for testing.

[0074] Figure 4 This is a visualization comparison of the detection results of DDP-YOLO on the SFCHD smoke dataset provided in this application embodiment. Ours represents the method of this application, while YOLOV7-tiny, Faster R-CNN, and RetinaNet are existing models. DDP-YOLO represents the overall model framework of this application. In the SFCHD scene test with dense smoke, from... Figure 4 It is clear that traditional models such as YOLOv11 suffer from numerous missed detections due to their inability to handle smoke interference. In contrast, the DDP-YOLO proposed in this application utilizes a dynamically degenerate decomposition backbone network (…). The synergistic effect of Backbone and Physically Guided Defogging Neck Network (PHAT) successfully detected blurred worker targets in images, demonstrating significant advantages in smoky environments.

[0075] Figure 5This is a visualization comparison of the detection results of DDP-YOLO on a heavily occluded PPEs dataset, provided in this embodiment of the application. The PPEs dataset test focuses on partially occluded PPE targets such as gloves and helmets. The figure shows that other models fail when faced with occlusion, while DDP-YOLO, thanks to the powerful contextual reasoning capabilities of its Context Relationship Module (CRM), maintains accurate localization, effectively solving the target detection challenge caused by occlusion.

[0076] Figure 6 The visualization comparison of DDP-YOLO detection results on the PP02 complex industrial background dataset, provided for embodiments of this application, visually demonstrates the superior detection performance of this application in complex environments. These experimental results fully demonstrate that the method of this application possesses superior detection performance and computational efficiency in complex industrial environments, providing reliable technical support for industrial safety monitoring.

[0077] This application also provides an occlusion detection device 700 that combines dynamic degradation decomposition with physical repair, such as... Figure 7 As shown, the device includes: a prompt generation module 701, a multi-scale repair feature map generation module 702, and a prediction module 703.

[0078] The prompt generation module 701 is used to receive industrial scene images, extract multi-scale features of industrial scene images using a dynamic degradation decomposition backbone network, explicitly decompose and suppress the dynamic degradation mode of industrial scene images in manifold space, and generate degradation correction prompts and strategy prompts.

[0079] The multi-scale restoration feature map generation module 702 is used to input degradation correction prompts and strategy prompts into the physically guided dehazing neck network. Utilizing the physical inductive bias of the atmospheric scattering model, it decouples and restores multi-scale features, generating multi-scale restoration feature maps. Semantic reconnection operations include: constructing query, key, and value mappings for the multi-scale restoration feature maps, and obtaining appearance relation compatibility and spatial geometric location compatibility. Through a dual relational attention mechanism, attention weights are generated based on appearance relation compatibility and spatial geometric location compatibility to enhance local features of the multi-scale restoration feature maps and generate context-enhanced feature maps.

[0080] The prediction module 703 is used to input the context-enhanced feature map into the decoupled detection head to predict the location and category of industrial personal protective equipment in industrial scene images. The location of the industrial personal protective equipment is presented in the form of a predicted bounding box.

[0081] Some modules in the apparatus described in this application can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0082] The apparatus or module described in the above embodiments can be implemented by a computer chip or physical entity, or by a product with a certain function. For ease of description, the above apparatus is described by dividing it into various modules according to their functions. When implementing the embodiments of this application, the functions of each module can be implemented in one or more software and / or hardware. Of course, a module that implements a certain function can also be implemented by combining multiple sub-modules or sub-units.

[0083] The methods, apparatus, or modules described in this application can be implemented in a computer-readable program code manner. The controller can be implemented in any suitable manner, for example, as a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of a memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code manner, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included within it for implementing various functions can also be considered as structures within the hardware component. Alternatively, the device used to implement various functions can be viewed as either a software module that implements the method or a structure within a hardware component.

[0084] like Figure 8As shown in the figure, this application embodiment also provides an occlusion detection server that combines dynamic degradation decomposition with physical restoration, including a memory 801 and a processor 802; the memory 801 is used to store computer-executable instructions; the processor 802 is used to execute the computer-executable instructions to implement the occlusion detection method that combines dynamic degradation decomposition with physical restoration described above in this application embodiment.

[0085] This application also provides a computer-readable storage medium storing executable instructions. When a computer executes the executable instructions, it can implement the occlusion detection method combining dynamic degradation decomposition and physical repair described above in this application.

[0086] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary hardware. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, or it can be embodied in the process of data migration. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, mobile terminal, server, or network device, etc.) to execute the methods described in the embodiments of this application.

[0087] The various embodiments described in this specification are presented in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. All or part of this application can be used in numerous general-purpose or special-purpose computer system environments or configurations.

[0088] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of this application.

Claims

1. An occlusion detection method combining dynamic degradation decomposition and physical restoration, characterized in that, include: The system receives industrial scene images, extracts multi-scale features of the industrial scene images using a dynamic degradation decomposition backbone network, and explicitly decomposes and suppresses the dynamic degradation mode of the industrial scene images in the manifold space, generating degradation correction prompts and strategy prompts. The degradation correction prompts and strategy prompts are input into the physics-guided dehazing neck network. The physical inductive bias of the atmospheric scattering model is used to decouple and repair multi-scale features, generating multi-scale repaired feature maps. The context relation module is used to perform semantic reconnection on the multi-scale repair feature map, which enhances the local features of the multi-scale repair feature map and generates a context-enhanced feature map. The semantic reconnection operation includes: constructing query, key, and value mappings for multi-scale repair feature maps to obtain appearance relation compatibility and spatial geometric location compatibility; generating attention weights based on appearance relation compatibility and spatial geometric location compatibility through a dual relation attention mechanism to enhance the local features of the multi-scale repair feature maps and generate context-enhanced feature maps. The context-enhanced feature map is input into the decoupled detection head to predict the location and category of industrial personal protective equipment in industrial scene images; the location of the industrial personal protective equipment is presented in the form of a predicted bounding box. The degradation correction prompts and strategy prompts are input into the physics-guided dehazing neck network. Using the physical inductive bias of the atmospheric scattering model, multi-scale features are decoupled and repaired to generate a multi-scale repaired feature map, including: The physical-guided defogging neck network includes a top-down path for physical-guided repair and a bottom-up path for spatial detail injection; The physical-guided repair process includes: using a physical-guided decoupling module, based on an atmospheric scattering model, decoupling multi-scale features into haze-related features and content features, and then recombining them. ;in, For the first The repaired features output by the layer physical guidance decoupling module. Indicates the first Operation functions of the layer physical guidance decoupling module. Indicating degenerate trunk features, Indicates the first Repair features output by the layer physical guidance decoupling module. Indicates an upsampling operation. For degradation correction prompts, Provide strategy hints; Using the physical inductive bias of the atmospheric scattering model, the repaired features output by the physical guidance decoupling module are decoupled into clear content features, transmission rate map features, and atmospheric light features. Feature reconstruction is achieved through the inverse transformation of the atmospheric scattering model to realize the spatial detail injection process: ;in, For multi-scale repair feature maps, To clearly define the content characteristics, For transmission rate map features, Atmospheric light characteristics; The expression for appearance compatibility is: ;in, To query pixels Corresponding query features With key pixels Corresponding context features Compatibility of appearance relationships between them This is the first learnable weight matrix. This is the second learnable weight matrix. For query features The function representation, Contextual features The function representation; The expression for spatial geometric location compatibility is: ;in, To query pixels Corresponding query features With key pixels Corresponding context features Spatial geometric position compatibility between them This is the third learnable weight matrix. For use in calculation and A function of the spatial geometric position between them This is the activation function.

2. The occlusion detection method combining dynamic degradation decomposition and physical restoration according to claim 1, characterized in that, The received industrial scene image is processed by extracting multi-scale features of the industrial scene image using a dynamic degradation decomposition backbone network, and explicitly decomposing and suppressing the dynamic degradation mode of the industrial scene image in the manifold space, generating degradation correction prompts and strategy prompts, including: The dynamic degradation decomposition backbone network processes industrial scene images through a cross-domain degradation analyzer. The cross-domain degradation analyzer explicitly decomposes and suppresses the dynamic degradation patterns of industrial scene images in the manifold space, generating degradation correction prompts and strategy prompts. Each level of the dynamically degenerate decomposition backbone network processes the features of the previous level through a dynamic decomposition mechanism. The calculation process for each level is as follows: ;in, For the first Features of each level of output For the first The gating weights of each level of output, Represents the dynamic degradation decomposition backbone network. The dynamic decomposition mechanism operates at each level. This is a feature of the next higher level. For degradation correction prompts, This is a strategy suggestion.

3. The occlusion detection method combining dynamic degradation decomposition and physical restoration according to claim 1, characterized in that, The method employs a dual-relational attention mechanism, generating attention weights based on appearance relationship compatibility and spatial geometric position compatibility to enhance local features of multi-scale repair feature maps and generate context-enhanced feature maps, including: The context module uses the Softmax function to fuse appearance relationship compatibility and spatial geometric position compatibility to generate attention weights. The expression for attention weights is: ;in, For attention weights, It is an exponential function. Key pixels Corresponding context features To query pixels Corresponding query features With key pixels Corresponding context features Spatial geometric position compatibility between them To query pixels Corresponding query features With key pixels Corresponding context features Compatibility of appearance relationships between them; We use attention weights to perform weighted aggregation of value features and extract the aggregated global context information. The aggregated global context information is residually concatenated with the original features of the multi-scale insulated feature map to enhance the local features of the multi-scale insulated feature map, thereby generating a context-enhanced feature map. The expression is as follows: ;in, For context-enhanced feature maps, The original features of the multi-scale repair feature map, This is the fourth learnable weight matrix. This is the aggregated global context information. The spatial dimensions of the multi-scale repair feature map are given by a height of [value missing]. Width is , For attention weights, For key pixels, Key pixels The value characteristics at that location.

4. The occlusion detection method combining dynamic degradation decomposition and physical restoration according to claim 3, characterized in that, Also includes: A quality-aware dynamic focusing loss function is used to optimize the overall model framework during end-to-end training. The quality-aware dynamic focusing loss function uses the intersection-union ratio of the predicted boxes as a quality indicator and dynamically adjusts the gradient weights of easy and difficult samples. The overall model framework includes a dynamically degenerate backbone network, a physically guided dehazing neck network, a context relation module, and a decoupled detection head.

5. The occlusion detection method combining dynamic degradation decomposition and physical restoration according to claim 4, characterized in that, The quality-perceived dynamic focus loss function includes zoom loss, Wise-IoU loss, and distributed focus loss; The expression for the quality-perceived dynamic focusing loss function is: ;in, For quality-perceived dynamic focusing loss function, For the first Index of positive samples, For classification confidence, For zoom loss, To predict soft tags, To dynamically focus weights, For Wise-IoU loss, Predict bounding boxes, True bounding box, This is the distributed loss balance coefficient. For distributed focusing loss, To predict the bounding box coordinate distribution, The coordinate distribution of the true bounding box; , ;in, These are dynamically adjusted gradient weight coefficients. For outlier degree, As the first hyperparameter, This is the second hyperparameter. For the current crossover loss, It is a moving average; Zoom loss uses the intersection-over-union ratio of the predicted bounding boxes as the predicted soft label. To weighted classification confidence The expression is: ;in, For zoom loss, For balance coefficient, This is the adjustment coefficient.

6. An occlusion detection device combining dynamic degradation decomposition and physical repair, characterized in that, include: The prompt generation module receives industrial scene images, extracts multi-scale features of the industrial scene images using a dynamic degradation decomposition backbone network, explicitly decomposes and suppresses the dynamic degradation mode of the industrial scene images in the manifold space, and generates degradation correction prompts and strategy prompts. A module for generating multi-scale repair feature maps is used to input degradation correction prompts and strategy prompts into the physics-guided dehazing neck network. It utilizes the physical inductive bias of the atmospheric scattering model to decouple and repair multi-scale features, generating multi-scale repair feature maps. The semantic reconnection operation includes: constructing query, key, and value mappings for multi-scale repair feature maps to obtain appearance relation compatibility and spatial geometric location compatibility; generating attention weights based on appearance relation compatibility and spatial geometric location compatibility through a dual relation attention mechanism to enhance the local features of the multi-scale repair feature maps and generate context-enhanced feature maps. The prediction module is used to input the context-enhanced feature map into the decoupled detection head to predict the location and category of industrial personal protective equipment in industrial scene images; the location of industrial personal protective equipment is presented in the form of a prediction bounding box. The degradation correction prompts and strategy prompts are input into the physics-guided dehazing neck network. Using the physical inductive bias of the atmospheric scattering model, multi-scale features are decoupled and repaired to generate a multi-scale repaired feature map, including: The physical-guided defogging neck network includes a top-down path for physical-guided repair and a bottom-up path for spatial detail injection; The physical-guided repair process includes: using a physical-guided decoupling module, based on an atmospheric scattering model, decoupling multi-scale features into haze-related features and content features, and then recombining them. ;in, For the first The repaired features output by the layer physical guidance decoupling module. Indicates the first Operation functions of the layer physical guidance decoupling module. Indicating degenerate trunk features, Indicates the first Repair features output by the layer physical guidance decoupling module. Indicates an upsampling operation. For degradation correction prompts, Provide strategy hints; Using the physical inductive bias of the atmospheric scattering model, the repaired features output by the physical guidance decoupling module are decoupled into clear content features, transmission rate map features, and atmospheric light features. Feature reconstruction is achieved through the inverse transformation of the atmospheric scattering model to realize the spatial detail injection process: ;in, For multi-scale repair feature maps, To clearly define the content characteristics, For transmission rate map features, Atmospheric light characteristics; The expression for appearance compatibility is: ;in, To query pixels Corresponding query features With key pixels Corresponding context features Compatibility of appearance relationships between them This is the first learnable weight matrix. This is the second learnable weight matrix. For query features The function representation, Contextual features The function representation; The expression for spatial geometric location compatibility is: ;in, To query pixels Corresponding query features With key pixels Corresponding context features Spatial geometric position compatibility between them This is the third learnable weight matrix. For use in calculation and A function of the spatial geometric position between them This is the activation function.

7. An occlusion detection server combining dynamic degradation decomposition and physical repair, characterized in that, Including memory and processor; The memory is used to store computer-executable instructions; The processor is configured to execute the computer-executable instructions to implement the method according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores executable instructions, which, when executed by a computer, enable the implementation of the method as described in any one of claims 1-5.