A method, system, electronic device and medium for human body recognition in a thick smoke environment

Through the combination of DyTransUNet and DyWinFormer networks, combined with the improved YOLOv8 network, the problems of poor image quality and low recognition accuracy of human body recognition in thick smoke environments are solved, and high-quality smoke removal and accurate human body recognition are achieved.

CN120126185BActive Publication Date: 2025-07-25EAST CHINA JIAOTONG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510607536.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-07-25
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

In the prior art, when human body recognition in thick smoke environments, the image quality is poor, the contour is blurred, and the recognition accuracy is low, making it difficult to meet the real-time and robust needs of emergency rescue and security monitoring.

Method used

The DyTransUNet network is used to accurately segment the smoke area, the DyWinFormer network is used to target smoke repair, and combined with the improved YOLOv8 network for human object recognition. The identification accuracy is improved through technical means such as dynamic patch division, smoke perception position coding, and dynamic mask guidance attention and area perception style modulation.

Benefits of technology

It realizes accurate positioning and repair of smoke areas in thick smoke environments, generates high-quality smoke removal images, improves the accuracy and robustness of human body recognition, and reduces the computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126185B_ABST
    Figure CN120126185B_ABST
Patent Text Reader

Abstract

This application belongs to the field of computer vision technology and discloses a method, system, electronic device, and medium for human body recognition in a thick smoke environment. The method includes: obtaining an image of a thick smoke environment to be processed, using the DyTransUNet network to process the image of the thick smoke environment, segmenting the smoke area in the image, and generating a smoke mask containing pixel-level smoke position and concentration information. The DyTransUNet network is a neural network based on the convolutional neural network (CNN) and Transformer; using the DyWinFormer network, based on the smoke mask and the image of the thick smoke environment, to repair the smoke area and generate a clear image after removing the smoke. The DyWinFormer network is a neural network based on a moving window and a dynamic mask; using the YOLOv8 network based on the attention mechanism to perform human target recognition on the clear image and output the position information of the human target. This method improves the accuracy of human body recognition in a thick smoke environment through accurate segmentation of the smoke area, targeted repair, and an optimized human body recognition network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a method, system, electronic device and medium for human body recognition in a thick smoke environment. Background Art

[0002] In fires, industrial accidents or other emergency scenarios, the thick smoke environment severely hinders the detection and recognition of human targets by traditional visual sensors (such as RGB cameras). The scattering and absorption of light by smoke particles lead to a decrease in image contrast, loss of details, and the generation of a large amount of noise, causing the performance of visible light-based human body recognition algorithms to drop sharply. In addition, problems such as high temperature, low visibility, and dynamic occlusion often accompany the thick smoke environment, further increasing the difficulty of recognition. Traditional object detection methods (such as YOLO, Faster R-CNN, etc.) rely on clear visual features and are prone to false detection or missed detection under the interference of thick smoke, making it difficult to meet the real-time and robustness requirements of scenarios such as emergency rescue and security monitoring.

[0003] In order to improve the accuracy of human target recognition in thick smoke, a large number of image de-smoking algorithms have been proposed from different perspectives such as image enhancement, image restoration, and image fusion. In terms of image enhancement, methods such as histogram equalization and Retinex theory are used to improve the image quality, but these methods can only limitedly enhance the contrast and are difficult to fundamentally solve the problem of detail loss caused by smoke; in the field of image restoration, the dark channel prior algorithm based on the atmospheric scattering model and its improved methods are widely used for smoke removal. However, such methods rely strongly on the accuracy of smoke concentration estimation and have poor effects when dealing with non-uniform smoke or dynamic smoke scenarios; in the direction of image fusion, data fusion technologies of multi-modal sensors (such as visible light and infrared imaging) significantly improve the recognition robustness, but there are practical application bottlenecks such as high equipment costs and difficult data registration.

[0004] Therefore, the existing de-smoking algorithms still have problems such as large computational overhead, poor processing effect on non-uniform smoke, easy loss of details or introduction of artifacts in the application of real and complex thick smoke scenarios, resulting in the inability to accurately segment the smoke area, perform targeted repair, and optimize the human body recognition network, making it difficult to improve the accuracy of human body recognition in a thick smoke environment. Summary of the Invention

[0005] Aiming at the problems of poor image quality, blurred contours, and low recognition accuracy faced by the prior art in human body recognition in a thick smoke environment, the present invention aims to provide a deep learning-based human body recognition method after thick smoke removal, which improves the accuracy of human body recognition in a thick smoke environment through accurate segmentation of the smoke area, targeted repair, and optimized human body recognition network.

[0006] To achieve the above object, a first aspect of the present invention provides a method for human body recognition in a thick smoke environment, the method comprising:

[0007] Obtain a thick smoke environment image to be processed;

[0008] Process the thick smoke environment image using the DyTransUNet network to segment the smoke area in the image, generate a smoke mask containing pixel-level smoke position and concentration information, and the DyTransUNet network is a neural network based on the convolutional neural network CNN and Transformer;

[0009] Use the DyWinFormer network to repair the smoke area based on the smoke mask and the thick smoke environment image to generate a clear image after removing the smoke, and the DyWinFormer network is a neural network based on a moving window and a dynamic mask;

[0010] Use the YOLOv8 network based on the attention mechanism to perform human target recognition on the clear image and output the position information of the human target.

[0011] As an alternative implementation manner of the first aspect of the present application, the DyTransUNet network includes: a CNN-DyTransformer hybrid encoder for extracting features of the thick smoke environment image, and the CNN-DyTransformer hybrid encoder adopts a dynamic patch division strategy to adaptively determine the patch size according to the local gradient information of the feature map, and add relative position encoding of smoke diffusion perception after serializing the patches; a DyTransformer encoder, including multiple layers of multi-head self-attention modules and multi-layer perceptron modules, and adopting a dynamic hyperbolic tangent activation function; and a cascaded upsampler for restoring the features output by the DyTransformer encoder to the original resolution and fusing them with the high-resolution features in the CNN-DyTransformer hybrid encoder through skip connections to output the smoke mask.

[0012] As an alternative implementation manner of the first aspect of the present application, the dynamic patch division strategy is based on the gradient magnitude of the feature map and includes: using small-sized patches in areas where the gradient magnitude is greater than a preset threshold; using large-sized patches in areas where the gradient magnitude is less than or equal to the preset threshold.

[0013] As an alternative implementation of the first aspect of the present application, the DyWinFormer network includes: a convolutional head for performing preliminary feature extraction and downsampling on the fused thick smoke environment image and the smoke mask; a DyWinFormer body containing multiple DyWinFormer blocks, where each DyWinFormer block includes a multi-head context attention module based on a moving window and a multi-layer perceptron module, and uses a dynamic hyperbolic tangent activation function; and a region-aware style modulation module for spatially adaptively modulating the weights of convolutional layers in the network according to the smoke concentration information and image local gradient information provided by the smoke mask.

[0014] As an alternative implementation of the first aspect of the present application, in the multi-head context attention module based on a moving window: the calculation of multi-head context attention is guided by an improved in-window mask; the improved in-window mask is generated according to the binary mask information and smoke concentration information provided by the smoke mask, and combines a learnable scaling factor and local smoke concentration features to adjust the attention weights from different pixel regions.

[0015] As an alternative implementation of the first aspect of the present application, in the region-aware style modulation module: a dual-path style representation based on smoke-concentration image-conditioned style and noise-unconditioned style is constructed, and the final style representation is obtained through a smoke-aware fusion function to perform region-aware modulation on the convolutional weights, and the modulation process includes a regulation coefficient related to the smoke concentration.

[0016] As an alternative implementation of the first aspect of the present application, the YOLOv8 network based on the attention mechanism includes: a ConvNeXt backbone network for extracting features of the clear image after smoke removal, where the ConvNeXt backbone network contains multiple ConvNeXt blocks optimized for the characteristics of the smoke-removed image; and a neck network and a detection head, where at least one SE-C2f module is integrated in the neck network, and the SE-C2f module includes an improved squeeze-and-excitation attention mechanism that uses average pooling and max pooling dual pooling operations and introduces a regulation coefficient related to the smoke concentration.

[0017] In a second aspect, an embodiment of the present application provides a system for human body recognition in a thick smoke environment, and the system includes:

[0018] An image acquisition module for acquiring a thick smoke environment image to be processed;

[0019] A smoke segmentation module, configured to process the thick smoke environment image by using a DyTransUNet network, segment the smoke area in the image, and generate a smoke mask containing pixel-level smoke position and concentration information, where the DyTransUNet network is a neural network based on a convolutional neural network (CNN) and a Transformer;

[0020] An image restoration module, configured to use a DyWinFormer network to restore the smoke area based on the smoke mask and the thick smoke environment image, and generate a clear image after smoke removal, where the DyWinFormer network is a neural network based on a mobile window and a dynamic mask;

[0021] A human body recognition module, configured to perform human target recognition on the clear image by using a YOLOv8 network based on an attention mechanism, and output the position information of the human target.

[0022] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0023] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0024] Compared with the prior art, the present invention has the following beneficial effects:

[0025] (1) Precise positioning and restoration: By adopting a strategy of segmentation first and then restoration, unnecessary processing of the smoke-free area is avoided, the computational complexity is reduced, and at the same time, different concentration smoke areas can be processed differently, improving the processing effect on non-uniform smoke.

[0026] (2) High-quality smoke removal: DyTransUNet uses dynamic patches, smoke-aware positional encoding, and DyT activation to more accurately segment fuzzy and dynamic smoke edges; DyWinFormer can better retain the details of the smoke-free area and ensure the style consistency between the restored area and the surrounding environment while restoring the smoke area through dynamic mask-guided attention and region-aware style modulation, generating high-quality smoke-removed images.

[0027] (3)Accurate human body recognition: Aiming at problems such as artifacts, blurring, and color imbalance that may exist in the images after smoke removal, the improved YOLOv8 adopts the ConvNeXt backbone and conducts targeted optimizations (large kernels, grouped convolutions, spatial adaptive LN, smoke perception preprocessing, etc.). Combining with the SE-C2f module to enhance the key feature channels, it improves the accuracy and robustness of human body recognition on smoke-removed images. Description of the Drawings

[0028] Figure 1 It is a flowchart of a method for human body recognition in a thick smoke environment provided by an embodiment of the present invention;

[0029] Figure 2 It is a schematic diagram of the DyTransUNet smoke segmentation network structure in an embodiment of the present invention;

[0030] Figure 3 It is a schematic diagram of the DyWinFormer smoke restoration network structure in an embodiment of the present invention;

[0031] Figure 4 It is a schematic diagram of the DyWinFormer block structure in an embodiment of the present invention;

[0032] Figure 5 It is a schematic diagram of the SE-C2f module structure in an embodiment of the present invention;

[0033] Figure 6 It is a schematic diagram of the structure of a system for human body recognition in a thick smoke environment provided by an embodiment of the present invention. Detailed Embodiments

[0034] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.

[0035] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order different from those illustrated or described herein. In addition, "and / or" in the specification and claims means at least one of the connected objects. The character " / ", generally represents an "or" relationship between the associated objects before and after. In the description of the present invention, "a plurality" means two or more, unless otherwise specifically defined.

[0036] Embodiment 1

[0037] Please refer to Figure 1 , which is a flowchart of a method for human body recognition in a thick smoke environment proposed in the first embodiment of the present invention. Its core process includes three main steps: smoke segmentation (Step 1), smoke restoration (Step 2), and human body recognition (Step 3). Each step uses a specially designed and trained deep neural network.

[0038] Step 1: Obtain the thick smoke environment image to be processed, and use the DyTransUNet network to process the thick smoke environment image to segment the smoke area in the image and generate a smoke mask containing pixel-level smoke position and concentration information. The DyTransUNet network is a neural network based on the convolutional neural network CNN and Transformer.

[0039] First, construct a smoke image segmentation dataset, including:

[0040] S11. Based on the fog-free images in the public dataset (such as Dense-Haze), a synthetic smoke image with different concentrations, morphologies, and distribution characteristics and its corresponding fog-free original image are generated by combining physical simulation (such as a smoke diffusion model based on hydrodynamics) and data augmentation, and a training dataset for the smoke removal (restoration) task is constructed.

[0041] S12. Based on the real smoke images in the public dataset (such as Dense-Haze), use professional annotation tools to accurately annotate the smoke area in each image at the pixel level. Preferably, the polygon annotation method is used to ensure the accuracy of the smoke edge (especially the transition between the semi-transparent area and the background), and a smoke segmentation mask is generated.

[0042] S13. Data augmentation is performed on the real smoke images and their segmentation masks annotated in step S12, including: random rotation (such as ±30°), mirror flipping, and elastic deformation in the spatial dimension; channel separation and recombination, Gamma correction (such as γ∈[0.8, 1.2]), and color offset simulation based on the atmospheric scattering model in the color space to enhance the robustness of smoke features under different lighting conditions.

[0043] Secondly, construct a smoke image segmentation neural network (DyTransUNet network), including:

[0044] In this embodiment, the DyTransUNet network includes:

[0045] The CNN-DyTransformer hybrid encoder is used to extract the features of images in a thick smoke environment. The CNN-DyTransformer hybrid encoder adopts a dynamic patch division strategy, adaptively determines the patch size according to the local gradient information of the feature map, and adds a relative position encoding with smoke diffusion perception after serializing the patches; among them, the dynamic patch division strategy is based on the gradient magnitude of the feature map, including: using small-sized patches in areas where the gradient magnitude is greater than the preset threshold, and using large-sized patches in areas where the gradient magnitude is less than or equal to the preset threshold.

[0046] The DyTransformer encoder contains multiple layers of multi-head self-attention modules and multi-layer perceptron modules, and uses a dynamic hyperbolic tangent activation function.

[0047] The cascaded upsampler is used to restore the features output by the DyTransformer encoder to the original resolution, and fuse them with the high-resolution features in the CNN-DyTransformer hybrid encoder through skip connections to output a smoke mask.

[0048] Specifically, as Figure 2 shown, the way to construct the DyTransUNet network is as follows:

[0049] First, construct the CNN-DyTransformer hybrid encoder. DyTransUNet uses a convolutional neural network as a feature extractor to extract the features in the image to obtain a feature map, and converts the feature map into a sequence as the input of the DyTransformer. The specific way to convert the sequence is as follows:

[0050] The feature map generated by the CNN is segmented and flattened into a 2D patch sequence, and a trainable linear projection is used to map the vectorized sequence to a D-dimensional embedding space, and position encoding is added to retain spatial information. In a thick smoke scene, the diffusion of smoke has obvious spatial continuity characteristics. Although the self-attention mechanism of the traditional Transformer can capture long-range dependencies, it will lose precise spatial position information. And the accurate segmentation of the smoke edge requires maintaining the sensitivity of local spatial relationships, especially for the gradual transition area between semi-transparent smoke and the background. In addition, the dynamic characteristics of smoke diffusion make the modeling of position information more critical, and the encoder needs to be able to distinguish the spatial differences between the smoke core area and the diffusion edge. Therefore, an additional term with smoke diffusion perception is introduced into the position encoding to better model the dynamic characteristics of smoke. And a dynamic patch division strategy is adopted to adaptively select the patch size according to the local gradient features. Small patches are used in high-gradient areas (smoke edges) to retain the details of the smoke edge; large patches are used in low-gradient areas to reduce the amount of calculation. The serialization formula is as follows:

[0051]

[0052] wherein represents the initial input sequence after embedding and positional encoding, represents the dynamically partitioned image patches, represents represents the dynamically determined patch size ( ), which is determined by the local gradient magnitude. is the number of image patches, represents the height of the image, represents the width of the image. is the size-related block embedding projection matrix (when ), ; when ), ), which maps the pixel values of the image patches to the feature space required by the Transformer. is the positional embedding, adopting a relative positional encoding scheme to ensure the consistency of the positional relationship of patches of different sizes. The patch partitioning process first calculates the gradient magnitude of the feature map, and uses 8×8 small patches in the regions where the gradient magnitude is greater than the threshold , and 16×16 large patches in the remaining regions.

[0053] Construct the DyTransformer encoder. Smoke usually has characteristics such as semi-transparency, dynamic diffusion, and blurred boundaries, and its shape and density change rapidly with the environment, which makes it possible that traditional Transformer normalization methods may not fully adapt to the non-uniform distribution of smoke when normalizing features. Dynamic Tanh (DyT) can more flexibly adapt to local changes in smoke features by dynamically adjusting the slope and offset of the non-linear transformation, enhancing the sensitivity of the model to low-contrast regions. The DyTransformer encoder consists of L layers of multi-head self-attention (MSA) and multi-layer perceptron blocks (MLP). Therefore, the output of the th layer can be written as:

[0054]

[0055]

[0056] wherein, is the output of the th layer, is the intermediate output of the th layer. is the dynamic hyperbolic tangent, and the formula is as follows:

[0057]

[0058] wherein, and are learnable per-channel vector parameters is a learnable scalar parameter that allows for different scalings based on the input range. is the hyperbolic sine function, and the specific formula is:

[0059]

[0060] The dynamic characteristics of DyT help the model distinguish smoke from similar noises. Traditional LN adopts a fixed normalization strategy for all feature channels, making it difficult to highlight the weak pixel value differences and feature differences between the smoke area and the background. DyT endows different channels with differential adjustment capabilities through learnable parameters, enabling the Transformer to more effectively suppress irrelevant backgrounds during global self-attention calculation and highlighting the long-range dependencies related to smoke.

[0061] Construct a Cascaded Upsampler (CUP). After reshaping the sequence of hidden features into the shape of (D is the output dimension), build the CUP by cascading multiple upsampling blocks. Each upsampling block consists of a 2x upsampling operation, a 3×3 convolutional layer, and a ReLU layer. An edge-aware upsampling kernel is specially designed to maintain the clarity of the smoke boundary. Finally, restore from to .

[0062] CUP and CNN-DyTransformer form a U-shaped architecture, which combines the high-resolution CNN feature maps in the encoder with the global context features encoded by DyTransformer through skip connections (modeling the smoke diffusion law) to achieve accurate smoke boundary localization and region segmentation.

[0063] Finally, use the smoke image and segmentation mask dataset constructed in the previous step to train the DyTransUNet network so that it can accurately segment the smoke area in the input image and output a pixel-level smoke mask.

[0064] Step 2: Use the DyWinFormer network to repair the smoke area based on the smoke mask and the thick smoke environment image to generate a clear image after smoke removal. The DyWinFormer network is a neural network based on moving windows and dynamic masks.

[0065] First, construct a smoke image repair (smoke removal) dataset, including:

[0066] S21. Based on the fog-free images in the public dataset (such as Dense-Haze), in a specific area of the image (such as the center of the picture or a random area), use a smoke diffusion model based on hydrodynamics to generate non-uniformly distributed thick smoke. Simulate different local smoke concentrations by adjusting the particle density (such as 0.1 - 0.5 g / m³) and the diffusion speed (such as 0.1 - 0.3 m / s), and use the alpha blending technique to ensure a natural transition at the smoke edges.

[0067] S22. Perform data augmentation on the generated local smoke images: randomly crop (such as from 256x256 to 512x512), rotate (such as ±30°), adjust the brightness (such as ±20%); apply the CutMix strategy to combine the smoke areas of different images; add Gaussian noise (such as σ = 0.01 - 0.05) and motion blur (such as kernel size from 3x3 to 7x7).

[0068] S23. Pair the fog-free original image, the generated image with local smoke, and the corresponding binary smoke area mask (indicating the smoke position) to form a dataset for training the smoke restoration network.

[0069] Secondly, construct a de-smoke neural network (DyWinFormer network), including:

[0070] In this embodiment, the DyWinFormer network includes:

[0071] A convolutional head for performing preliminary feature extraction and downsampling on the fused thick smoke environment image and the smoke mask;

[0072] The DyWinFormer main body contains multiple DyWinFormer blocks. The DyWinFormer block contains a multi-head context attention module based on a moving window and a multi-layer perceptron module, and uses a dynamic hyperbolic tangent activation function; the smoke concentration shows a gradient distribution that is light at the edges and thick in the center in space. The traditional attention mechanism equally weights all regions and cannot distinguish between slightly blurred regions that need fine-tuning and severely occluded regions that need reconstruction. Among them, in the multi-head context attention module based on a moving window: the calculation of multi-head context attention is guided by an improved in-window mask; the improved in-window mask is generated according to the binary mask information and smoke concentration information provided by the smoke mask, and combines a learnable scaling factor and local smoke concentration features to adjust the attention weights from different pixel regions; in the region-aware style modulation module: construct a dual-path style representation based on the image condition style of smoke concentration and the noise unconditional style, and obtain the final style representation through a smoke-aware fusion function to perform region-aware modulation on the convolutional weights. The modulation process includes a modulation coefficient related to the smoke concentration;

[0073] The region-aware style modulation module is used to perform spatial adaptive modulation on the weights of the convolutional layers in the network according to the smoke concentration information provided by the smoke mask and the local gradient information of the image.

[0074] Specifically, as Figure 3 shown, the DyWinFormer network is constructed as follows:

[0075] Construct the convolutional head. Use a convolutional neural network to accept the image and the mask , splice their channels and input them into the first convolutional layer to expand the number of channels to 180, and then reduce the resolution to through 3 convolutional layers with a stride of 2. Finally, flatten the downsampled feature map into a sequence through a flattening layer as the input of DyWinformer.

[0076] Construct the DyWinFormer body. The body includes 5 multi-head context attention DyWimFormer blocks that fuse moving windows (as Figure 4 shown), and there is an additional mask-guided efficient attention mechanism. The output of the layer can be written as:

[0077]

[0078]

[0079] where, is the output of the layer, is the intermediate output of the layer. is the dynamic hyperbolic tangent, is the multi-head context attention module based on the moving window.

[0080] Construct the multi-head context attention module based on the moving window (MCA). In the heavily smoke-polluted area, there is often an uneven degree of damage, and it is necessary to dynamically adjust the repair weights of different regions. The traditional attention mechanism treats all regions equally and cannot effectively distinguish the feature interactions between the intact regions (which should be retained) and the smoke-damaged regions (which need to be repaired). By introducing a dynamic mask-guided attention mechanism, it can intelligently enhance the feature propagation between effective regions while suppressing the noise interference brought by the damaged regions. Using shifted windows and dynamic masks, non-local interactions are performed with several feasible tokens, and the output is calculated as the weighted sum of the effective tokens. The specific formula is as follows:

[0081]

[0082] where, is the query matrix, is the key matrix, is the value matrix, both obtained by linearly projecting the input features to get. is a factor for scaling the dot-product attention scores, is a learnable scaling factor that controls the guiding strength of the mask, introducing a learnable , and the model will automatically adjust the balance between the mask and content attention. is the local smoke concentration feature. When is the case, the term dominates, forcing the model to refer to the context far from the damaged area. When is the case, content attention dominates and the original features are retained. Introducing a learnable and , the model can automatically adjust the dynamic balance among the mask, smoke concentration, and content attention. Adjust the guiding strength of the mask: for scenes with clear structures, increase to strictly follow the mask; for scenes with complex textures, decrease to rely on content attention. This improvement enables the smoke restoration to achieve an optimal balance among retaining details, reconstructing structures, and suppressing noise, especially suitable for the restoration of non-uniform contamination in complex natural scenes. is the improved in-window mask, which is generated by fusing the binary mask (0 represents the damaged area, 1 represents the valid area) and the smoke concentration map , flattened and mapped to a matrix with the same dimension as . For the valid area ( ), a positive bias is given to enhance its attention weight; for the damaged area ( ), a negative bias (or zero) is given to suppress the interference of irrelevant features. The specific formula is as follows:

[0083]

[0084] where, is a hyperparameter. The + sign indicates enhancing the attention weight, and the - sign indicates masking the attention weight to suppress the participation of invalid area features in attention aggregation. represents the smoke concentration at position (0 means no smoke, 1 means thick smoke), represents the value of the binary mask at position , represents the value of the binary mask at position The values are designed in such a way that the characteristics of the smoke-free area are maximally preserved, the mildly smoky area is moderately restored, and the severely smoky area is reconstructed with emphasis, while avoiding over-smoothing.

[0085] Construct a region-aware style modulation module. The restoration of the smoke-polluted area needs to consider both global consistency and local adaptability simultaneously. Traditional style transfer methods often struggle to balance these two requirements: global style transfer can lead to a lack of harmony between the restored area and the surrounding environment, while fully local processing is prone to obvious restoration traces. By introducing region-aware dual-path style modulation, the feature expression between the restored area and the preserved area can be intelligently balanced. Precise control of the output is achieved by implementing region-adaptive convolutional layer weight modulation during the reconstruction process with smoke concentration-aware noise input. To enhance the representation ability of the noise input and achieve local style control, construct a dual-path style representation based on smoke concentration: image-conditioned style (maintaining the consistency of the smoke-free area) and noise-unconditioned style (restoring the smoky area) are expressed as follows:

[0086]

[0087]

[0088]

[0089] Among them, is the smoke concentration map, represents the mapping function for extracting unconditional style features, represents random noise, represents the image features that have fused the region-aware style of, represents the size adjustment operation to ensure that the style vector can be weighted and fused pixel by pixel with the image features , represents the mapping function for generating the conditional style, is the local gradient feature. is an improved spatially adaptive region mask, and its probability of taking values is positively correlated with the local smoke concentration , is the Sigmoid function. By fusing the two style representations through smoke perception, the final style representation is obtained:

[0090]

[0091] Among them is the improved smoke perception mapping function, is the dynamic style weight map combined with the smoke diffusion characteristics, are the statistical quantities of the original image features. The post-convolution weights are modulated for region perception:

[0092]

[0093]

[0094] Among them, represents the weight after the first modulation, introducing position-aware style modulation, represents the convolutional kernel weight after region-aware normalization, represents the final style vector, represents the smoke concentration map, represents the final style vector of the th channel, represents the spatial position, are the input channel, output channel, and convolutional kernel size respectively, is a very small constant to prevent the denominator from being zero. and are the adjustment coefficients related to smoke concentration. In high-smoke-concentration areas ( >0), stronger style transformation capabilities are obtained; in smoke-free areas ( ≈0), the original features are maintained; in the transition area, natural gradients are achieved. Ensure that the repaired area has a natural transition in lighting and color with the surrounding environment while maintaining the authenticity of texture details.

[0095] Finally, the DyWinFormer network is trained using the constructed smoke image restoration (de-smoking) dataset so that it can repair the smoke area according to the input smoke image and smoke mask and output a clear de-smoked image.

[0096] Step 3: Use the YOLOv8 network based on the attention mechanism to perform human target recognition on the clear image and output the position information of the human target.

[0097] First, construct a human recognition dataset, including:

[0098] S31. Use a high-resolution RGB camera to collect image sequences (e.g., 30fps, 1920x1080 pixels, PNG format) in diverse scenarios (indoor, outdoor, different lighting, different simulated or real smoke concentrations).

[0099] S32. Use an annotation tool (such as LabelImg) to annotate the human targets in the images with rectangular boxes, requiring the bounding boxes to be close to the targets and have a high annotation confidence (e.g., ≥0.95), and generate label files (including class IDs and normalized coordinates) that meet the requirements of the object detection model (such as YOLOv8).

[0100] S33. Perform multi-stage data augmentation on the labeled RGB images:

[0101] Color jittering: brightness adjustment (e.g., ±20%), saturation variation (e.g., ±15%), hue shift (e.g., ±10%).

[0102] Add noise: Gaussian random noise (e.g., N(0, 0.01)).

[0103] Structure enhancement: Apply CutMix (e.g., mixing ratio α = 0.4) and Mosaic (e.g., 4-image stitching) enhancements.

[0104] S34. Implement multi-level annotation verification: automated verification (such as aspect ratio of bounding boxes), manual cross-checking, pre-trained model-assisted verification (confidence threshold), and clean unqualified samples.

[0105] Secondly, construct a human body recognition network (YOLOv8 network based on the attention mechanism), including:

[0106] In this embodiment, the YOLOv8 network based on the attention mechanism includes:

[0107] The ConvNeXt backbone network is used to extract features of the clear image after de-smoking. The ConvNeXt backbone network contains multiple ConvNeXt blocks optimized for the characteristics of de-smoked images; and

[0108] The neck network and the detection head, where at least one SE-C2f module is integrated in the neck network. The SE-C2f module contains an improved squeeze-and-excitation attention mechanism that adopts dual pooling operations of average pooling and max pooling and introduces a regulation coefficient related to the smoke concentration.

[0109] Specifically, as Figure 5 shown, the method for constructing the improved YOLOv8 network (YOLOv8 network based on the attention mechanism) is as follows:

[0110] Construct the ConvNeXt backbone network. Although the visibility of the image after smoke removal is improved, there are still three typical feature problems: local details are blurred, especially the loss of low-frequency texture information; the smoke removal algorithm may cause over-enhancement of some color channels, resulting in color channel imbalance; the high-frequency noise introduced during the smoke removal process will leave noise artifacts. The CSPDarknet53 backbone network adopted by the traditional YOLOv8 has obvious limitations: its dense cross-stage connection structure is prone to spreading the noise artifacts remaining after smoke removal; the fixed-size convolutional kernels are difficult to adapt to the regional feature differences of the image after smoke removal; the single ReLU activation function will exacerbate the color channel imbalance problem. In contrast, the modern design of ConvNeXt targets these problems in the following ways: its depthwise separable convolution structure can effectively separate channel features and alleviate color imbalance; the staged downsampling strategy retains multi-scale features to compensate for detail loss; layer normalization stabilizes the feature distribution and suppresses noise interference; the smooth characteristic of the GELU activation function avoids artificially introducing non-linear distortion; the large kernel convolution operation (7×7) provides a wider receptive field and effectively captures the global context relationship of the image after smoke removal. These characteristics make ConvNeXt more suitable for processing the complex feature distribution of the image after smoke removal than the original YOLOv8 backbone network.

[0111] The input part is a 3-channel visible light image input by the RGB branch , and first, perform preprocessing for smoke perception:

[0112]

[0113] where is the smoke residue intensity map, is the compensated noise, are learnable parameters.

[0114] Extract features through the 4-stage improved ConvNeXt block, and each stage is optimized according to the characteristics of the image after smoke removal:

[0115] Stage 1 ( resolution): Use 7×7 large kernel depth convolution, combined with the GELU activation function, to focus on restoring low-frequency contour features. Design a special edge enhancement convolutional kernel:

[0116]

[0117] where represents the edge enhancement convolutional kernel, which strengthens the edge information in the image to make up for the contour weakening problem caused by the smoke removal process.

[0118] Stage 2 ( resolution): Adopt grouped point convolution (group = 4) to process color channel imbalance and introduce channel attention:

[0119]

[0120] where is the smoke influence coefficient of the th channel, the original th channel's weight representation at the th position in the convolutional kernel, represents the weight of the th channel after channel attention adjustment at the th position in the convolutional kernel.

[0121] Stage 3 ( resolution): Use dynamic sparse convolution to adjust the sparse pattern of the convolutional kernel according to the smoke concentration map :

[0122]

[0123] where indicates whether to retain the binary judgment for participating in the convolutional calculation (if the value is greater than 0.5, it is retained; if it is less than 0.5, it is discarded), represents the smoke concentration of the central pixel at position , and represents the smoke concentration of the current neighborhood pixel at position .

[0124] Stage 4 ( resolution): Adopt hybrid dilated convolution (dilation = [1, 2, 3]) to expand the receptive field and compensate for the loss of deep features. The feature extraction formula is extended to:

[0125]

[0126]

[0127] where represents the finally extracted deep RGB image features, represents the depthwise separable convolution of the th stage, including the following improvements:

[0128]

[0129]

[0130] where is the attention mask for smoke perception, is the number of channels of the input feature map, and is a local statistic based on the smoke area calculation. In particular, the layer normalization parameter is extended to be spatially adaptive:

[0131]

[0132]

[0133] Construct the SE-C2f module. Although the visibility of the de-smoked image is improved, there may still be problems such as blurred local details, unbalanced color channels, and residual high-frequency noise. These problems will lead to the insufficient prominence of the features of the human target, especially in areas with heavy smoke occlusion. Although the traditional C2f module enhances the representation ability through multi-path feature extraction, it lacks targeted enhancement of key channels. The channel attention mechanism of the SE module can adaptively recalibrate the channel feature responses, effectively enhancing the saliency of the weak human features in the de-smoked image. The structure of the original C2f module is as follows:

[0134]

[0135] where represents the final output feature map of the module, which is used as the input of the downstream network, contains two 3×3 convolutions and a residual connection, is the smoke concentration map, represents feature modulation. represents concatenation along the channel dimension.

[0136] Insert the improved SE module after C2f:

[0137]

[0138]

[0139]

[0140] where represents the intermediate vector for feature statistical compression in the channel dimension, represents the channel attention map, which determines the degree to which each channel is amplified in the output, represents the output feature map after attention weighting and fusion with smoke residual modulation, is the weight matrix of the first fully connected layer, and the compression ratio is dynamically adjustable reduces the number of channels from to . and is the adjustment coefficient for smoke perception, which enables the characteristic channels in high-smoke-concentration areas to obtain greater enhancement weights, preserves the original feature distribution in the smoke-free area, and captures feature statistics more comprehensively through dual pooling (average pooling + max pooling). is the Sigmoid activation function, is the ReLU activation function. is the weight matrix of the second fully connected layer, which restores the number of channels to .

[0141] Add an SE module after C2f. Utilize its channel attention mechanism to adaptively recalibrate the channel feature responses, which can effectively enhance the saliency of weak human features in the de-smoked image. The SE module combines average pooling and max pooling operations to comprehensively capture feature statistics, and introduces an adjustment coefficient related to the smoke concentration to dynamically adjust the feature enhancement strategy according to the smoke concentration in different regions. It improves the prominence of the human target in the feature map and enhances the robustness of the human recognition network to the noise and artifacts in the de-smoked image.

[0142] Finally, use the constructed human recognition dataset (obtain clear images by processing the images through the trained smoke removal network and then pair them with labels) to train the improved YOLOv8 neural network so that it can accurately identify human targets in the de-smoked images.

[0143] Embodiment 2

[0144] Please refer to Figure 6 , which shows the structural schematic diagram of a human recognition system proposed in the second embodiment of this application. The system includes the following key modules:

[0145] The smoke segmentation module 100 is configured to obtain the image of the thick smoke environment to be processed, process the image of the thick smoke environment by using the DyTransUNet network, segment the smoke area in the image, and generate a smoke mask containing pixel-level smoke position and concentration information. The DyTransUNet network is a neural network based on the convolutional neural network CNN and Transformer;

[0146] The image restoration module 200 is configured to use the DyWinFormer network to restore the smoke area based on the smoke mask and the image of the thick smoke environment to generate a clear image after smoke removal. The DyWinFormer network is a neural network based on moving windows and dynamic masks;

[0147] The human recognition module 300 is configured to use the YOLOv8 network based on the attention mechanism to identify human targets in the clear image and output the position information of the human targets.

[0148] A system for human body recognition in a thick smoke environment in an embodiment of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a Network Attached Storage (NAS), a personal computer (PC), etc. The embodiments of the present application do not make specific limitations.

[0149] A system for human body recognition in a thick smoke environment in an embodiment of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0150] A system for human body recognition in a thick smoke environment provided by the embodiments of the present application can implement Figure 1 each process implemented in a method embodiment of a method for human body recognition in a thick smoke environment. To avoid repetition, it will not be elaborated here.

[0151] Optionally, an embodiment of the present application further provides an electronic device, including a processor, a memory, a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above method embodiment of a method for human body recognition in a thick smoke environment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0152] An embodiment of the present application further provides a readable storage medium. A program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, it implements each process of the above method embodiment of a method for human body recognition in a thick smoke environment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0153] Wherein, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk, or an optical disc, etc.

[0154] It should be noted that in this text, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0155] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0156] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of the present application and without departing from the spirit of the present application and the scope protected by the claims, can also make many forms, all of which fall within the protection scope of the present application.

Claims

1. A method for human body recognition in a thick smoke environment, characterized in that, The method includes: Obtaining a smoky environment image to be processed; Processing the smoky environment image using the DyTransUNet network to segment the smoke area in the image and generate a smoke mask containing pixel-level smoke position and concentration information. The DyTransUNet network is a neural network based on the convolutional neural network (CNN) and Transformer. The DyTransUNet network includes: a CNN-DyTransformer hybrid encoder for extracting features of the smoky environment image. The CNN-DyTransformer hybrid encoder adopts a dynamic patch division strategy, adaptively determines the patch size according to the local gradient information of the feature map, and adds relative position encoding with smoke diffusion perception after serializing the patches; a DyTransformer encoder containing multiple multi-head self-attention modules and multi-layer perceptron modules, and adopting a dynamic hyperbolic tangent activation function; and a cascaded upsampler for restoring the features output by the DyTransformer encoder to the original resolution and fusing them with the high-resolution features in the CNN-DyTransformer hybrid encoder through skip connections to output the smoke mask; Using the DyWinFormer network, based on the smoke mask and the smoky environment image, repairing the smoke area to generate a clear image after removing smoke. The DyWinFormer network is a neural network based on moving windows and dynamic masks. The DyWinFormer network includes: a convolutional head for performing preliminary feature extraction and downsampling on the fused smoky environment image and the smoke mask; a DyWinFormer main body containing multiple DyWinFormer blocks. The DyWinFormer block contains a multi-head context attention module based on moving windows and a multi-layer perceptron module, and adopts a dynamic hyperbolic tangent activation function; and a region-aware style modulation module for spatially adapting the modulation of the weights of the convolutional layers in the network according to the smoke concentration information and image local gradient information provided by the smoke mask; Using the YOLOv8 network based on the attention mechanism to perform human target recognition on the clear image and output the position information of the human target.

2. The method according to claim 1, wherein The dynamic patch division strategy is based on the gradient magnitude of the feature map and includes: In areas where the gradient magnitude is greater than a preset threshold, small-sized patches are used; In areas where the gradient magnitude is less than or equal to the preset threshold, large-sized patches are used.

3. The method for human body recognition in a thick smoke environment according to claim 1, characterized in that, In the multi-head context attention module based on moving windows: The calculation of multi-head context attention is guided by an improved in-window mask; The improved in-window mask is generated according to the binary mask information and smoke concentration information provided by the smoke mask, and combines a learnable scaling factor and local smoke concentration features to adjust the attention weights from different pixel regions.

4. A method for human body recognition in a thick smoke environment according to claim 1, characterized in that, In the region-aware style modulation module: A dual-path style representation of image conditional style based on smoke concentration and noise unconditional style is constructed, and the final style representation is obtained through a smoke-aware fusion function to perform region-aware modulation on the convolution weights. The modulation process includes an adjustment coefficient related to the smoke concentration.

5. A method for human body recognition in a thick smoke environment according to claim 1, characterized in that, The YOLOv8 network based on the attention mechanism includes: A ConvNeXt backbone network, used for extracting features of the clear image after smoke removal, wherein the ConvNeXt backbone network comprises a plurality of ConvNeXt blocks optimized for the characteristics of the smoke-removed image; and A neck network and a detection head, wherein at least one SE-C2f module is integrated in the neck network, and the SE-C2f module includes an improved squeezing and incentive attention mechanism that uses average pooling and maximum pooling double pooling operations and introduces an adjustment coefficient related to smoke concentration.

6. A system for human body recognition in a thick smoke environment, characterized in that, The system comprises: An image acquisition module is used to acquire the smoke environment image to be processed; A smoke segmentation module is configured to process the dense smoke environment image using a DyTransUNet network, segment the smoke area in the image, and generate a smoke mask containing pixel-level smoke position and concentration information. The DyTransUNet network is a neural network based on a convolutional neural network (CNN) and a Transformer; the DyTransUNet network includes: a CNN-DyTransformer hybrid encoder for extracting features of the dense smoke environment image, the CNN-DyTransformer hybrid encoder adopts a dynamic patch partitioning strategy, adaptively determines the patch size according to the local gradient information of the feature map, and serializes the patch and then adds a relative position code for smoke diffusion perception; a DyTransformer encoder, including a multi-layer multi-head self-attention module and a multi-layer perceptron module, and adopts a dynamic hyperbolic tangent activation function; and a cascade upsampler, for restoring the features output by the DyTransformer encoder to the original resolution, and fusing them with the high-resolution features in the CNN-DyTransformer hybrid encoder through a jump connection to output the smoke mask; An image restoration module, configured to utilize the DyWinFormer network to restore the smoke area based on the smoke mask and the thick smoke environment image, generating a clear image without smoke. The DyWinFormer network is a neural network based on mobile windows and dynamic masks; the DyWinFormer network includes: a convolutional head for performing preliminary feature extraction and downsampling on the fused thick smoke environment image and the smoke mask; a DyWinFormer main body containing multiple DyWinFormer blocks, where each DyWinFormer block includes a multi-head context attention module based on mobile windows and a multi-layer perceptron module, and uses a dynamic hyperbolic tangent activation function; and a region-aware style modulation module for performing spatial adaptive modulation on the weights of the convolutional layers in the network according to the smoke concentration information and image local gradient information provided by the smoke mask. A human body recognition module, configured to perform human body target recognition on the clear image using the YOLOv8 network based on the attention mechanism, and output the position information of the human body target.

7. An electronic device, characterized in that, It includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a method for human body recognition in a thick smoke environment as described in any one of claims 1-5 are implemented.

8. A readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, the steps of a method for human body recognition in a thick smoke environment as described in any one of claims 1-5 are implemented.

Citation Information

Patent Citations

  • Transform-based efficient defogging semantic segmentation method and application thereof

    CN117058024A

  • Human body recognition method in indoor fire smoke scene

    CN118968557A