Abnormal weather identification method, device, electronic equipment and program product based on deep learning

By introducing the EGMA and C2fPRFF modules to optimize the YOLOv8s model, the accuracy and robustness issues of image recognition under severe weather conditions are solved, and the recognition accuracy and stability of abnormal weather conditions are improved.

CN120182788BActive Publication Date: 2025-09-12STREAMAP TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510549910.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-09-12
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

Abnormal weather interferes with sensors, resulting in degraded image quality and difficulty in distinguishing between foreground and background. Existing technologies are insufficient in recognition accuracy and robustness, especially in severe weather conditions such as sandstorms, fog, rain and heavy snow, where the false detection and missed detection rates are high.

Method used

A weather recognition model based on deep learning is adopted, and the enhanced global hybrid attention mechanism (EGMA) module and the progressive receptive field fusion (C2fPRFF) module are introduced. The EGMA module is used to enhance the feature expression ability and global correlation strength, and the C2fPRFF module is used to capture multi-scale feature information. It is combined with the YOLOv8s model for optimization.

Benefits of technology

It significantly improves the accuracy and robustness of image recognition under adverse weather conditions, reduces the probability of false detection and missed detection, and improves the model's ability to understand and process complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182788B_ABST
    Figure CN120182788B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, electronic device and program product for abnormal weather recognition based on deep learning. The recognition method sets a C2fPRFF module in the backbone network through the weather recognition model, extracts multi-scale receptive field features through the progressive fusion of stacked PRFFBlocks, and enhances the model's understanding of contextual information. The neck network is provided with an EGMA module, which can combine the channel attention and spatial attention sub-modules, and add a channel shuffling operation to effectively improve the expressive ability of features and help the model better deal with the problems of poor image quality and difficulty in distinguishing foreground and background. In addition, by introducing a newly designed focused fusion intersection-over-union loss, the positioning accuracy of target detection and the convergence speed of the model can be optimized, especially under complex background and interference conditions, which improves the robustness and accuracy of the model. The comprehensive improvements in the three aspects can significantly improve the detection performance and reliability of the weather recognition model in harsh environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of image processing technology, and in particular relates to an abnormal weather recognition method, abnormal weather recognition device, electronic device and computer program product based on deep learning. Background Art

[0002] Abnormal weather recognition technology can be used to automatically detect and classify extreme weather conditions, such as sandstorms, fog, rainfall, and snow. This technology has important applications in areas such as traffic safety, agricultural production, and urban management. Real-time identification of abnormal weather can provide early warnings, reduce accident risks, optimize resource scheduling, and enhance society's ability to respond to abnormal weather.

[0003] However, abnormal weather conditions often interfere with sensors, such as lens obstruction or image blur, affecting recognition accuracy. Furthermore, phenomena such as dust, rain, and fog can cause dramatic background changes, making it difficult to distinguish between foreground and background, which can easily lead to false detections and missed detections. Summary of the Invention

[0004] This application provides an abnormal weather identification method, abnormal weather identification device, electronic device and computer program product based on deep learning, which can effectively improve the expression ability of features and help the model better deal with the problems of poor image quality and difficulty in distinguishing foreground and background.

[0005] In a first aspect, the present application provides a method for identifying abnormal weather based on deep learning, comprising:

[0006] The backbone network of the pre-trained weather recognition model extracts features from the image to be detected to obtain image features; the image to be detected includes meteorological information;

[0007] The neck network based on the weather recognition model fuses the image features to obtain the target fusion features;

[0008] The detection network based on the weather recognition model detects the target fusion features and obtains the detection results of abnormal weather in the image to be detected;

[0009] Among them, the neck network is equipped with an EGMA module, which enhances the expressiveness and global correlation strength of weather features based on at least two attention mechanisms.

[0010] Furthermore, the EGMA module includes a channel-aware modulation submodule, an ECA submodule, and a cross-channel spatial attention submodule ESA submodule. For the first input feature of the EGMA module:

[0011] Performing a channel feature enhancement operation on the first input feature through a channel-aware modulation submodule to obtain a first enhanced feature;

[0012] Perform a channel attention operation on the first enhanced feature through the ECA submodule to obtain a second enhanced feature;

[0013] Performing a cross-channel spatial attention operation on the second enhanced feature through a cross-channel spatial attention submodule to obtain a multi-attention enhanced feature;

[0014] The ESA submodule performs a spatial attention operation on the multiple attention-enhanced features to obtain a first output feature corresponding to the first input feature.

[0015] Furthermore, the channel-aware modulation submodule includes a dimension permutation layer, a double-layer multi-perceptron layer, an inverse dimension permutation layer, an activation function layer, and a first weighted layer; the channel-aware modulation submodule performs a channel feature enhancement operation on the first input feature to obtain a first enhanced feature, including:

[0016] The dimensional parameters of the first input feature are permuted through the dimensional permutation layer to obtain the flattened spatial feature;

[0017] The first multi-perceptron layer performs a channel compression operation on the flattened spatial features and a nonlinear activation operation to obtain compressed features;

[0018] Performing a channel expansion operation on the compressed features through a second multi-perceptron layer to obtain a first expanded feature that matches the dimension of the flattened spatial feature;

[0019] Perform inverse permutation on each dimensional parameter of the first extended feature through an inverse dimensional permutation layer to obtain a reconstructed feature whose dimensional parameters match the first input feature;

[0020] Perform linear activation operations on the reconstructed features through the activation function layer to obtain channel modulation weights;

[0021] A first enhanced feature is obtained by performing a weighted operation on the first input feature based on the channel modulation weight through the first weighted layer.

[0022] Furthermore, the cross-channel spatial attention submodule includes a channel shuffling layer, a CBR layer, a CBS layer and a second weighted layer; the cross-channel spatial attention submodule performs a cross-channel spatial attention operation on the second enhanced feature to obtain a multi-attention enhanced feature, including:

[0023] performing a channel shuffling operation on the second enhanced feature through a channel shuffling layer to obtain a recombined feature;

[0024] Through the CBR layer, channel compression operation, normalization operation and nonlinear transformation operation are performed on the recombined features in sequence to obtain compressed activation features;

[0025] The CBS layer performs channel expansion, normalization, and linear activation operations on the compressed activation features to obtain spatial weights.

[0026] The second weighted layer performs a weighted operation on the reorganized features based on the spatial weights to obtain multi-attention enhanced features.

[0027] Furthermore, the backbone network is provided with a C2fPRFF module, including a first convolutional layer, a progressive fusion structure including at least two PRFFBlocks, a first concatenation layer, and a second convolutional layer. For each second input feature of the C2fPRFF module:

[0028] Perform channel compression on the second input feature through the first convolution layer to obtain convolution features;

[0029] Through the progressive fusion structure, the convolution features are sequentially fused through the series of PRFFBlocks to obtain the hierarchical fusion features corresponding to each level;

[0030] The convolution feature and the fusion feature of each level are concatenated through the first concatenation layer to obtain the first concatenation feature;

[0031] A channel expansion operation is performed on the first concatenated feature through a second convolutional layer to obtain a second output feature corresponding to the second input feature; and the image feature is obtained based on the second output feature.

[0032] Furthermore, PRFFBlock includes a progressive convolution structure, a second convolution layer, a third convolution layer, and an ESE layer third convolution layer. The progressive convolution structure includes at least two convolution layers connected in series. For the third input feature of each PRFFBlock:

[0033] Performing a progressive convolution operation on the third input feature through at least two convolutional layers connected in series in a progressive convolution structure to obtain hierarchical convolution features corresponding to each level;

[0034] Performing a splicing operation on the third input feature and the convolution features of each level through the second splicing layer to obtain a second splicing feature;

[0035] Performing a channel expansion operation on the second concatenated feature through the third convolutional layer to obtain a second expanded feature;

[0036] Perform global average pooling and convolution operations on the second extended feature through the ESE layer, and weight the calculated channel attention weight on the second extended feature to obtain the initial fusion feature;

[0037] A splicing operation is performed on the third input feature and the initial fusion feature through the third splicing layer to obtain a third output feature corresponding to the third input feature.

[0038] Furthermore, the weather recognition model is trained based on regression loss and classification loss; the regression loss includes focused fusion intersection-over-union loss, and the formula is as follows:

[0039]

[0040]

[0041]

[0042]

[0043]

[0044]

[0045]

[0046] in, To focus on the fusion intersection-over-union loss, IoU is the intersection-over-union ratio between the predicted box and the true box; γ is a parameter that controls the degree of outlier suppression; To blend and compare losses; L IoU is the intersection and comparison loss; L dis is the center distance loss; L shp is the width and height loss; L ang is the angle loss; Δ is 2 times the center distance loss; b and Represent the center point of the real box and the center point of the predicted box respectively; Represents the square of the Euclidean distance between the center point of the real box and the predicted box; and Represent the width and height of the minimum bounding rectangle respectively; Ω is 2 times the width and height loss; w and are the width of the predicted border and the true border respectively; Represents the square of the Euclidean distance between the true box and the predicted box width; h and are the heights of the predicted bounding box and the true bounding box respectively; Represents the square of the Euclidean distance between the true box and the predicted box height; The angle loss is 2 times; and Represent the width and height of the rectangular box constructed by the center points of the true box and the predicted box, respectively.

[0047] In a second aspect, the present application provides an abnormal weather identification device, comprising:

[0048] An extraction module is used to extract features from the image to be detected based on the backbone network of the pre-trained weather recognition model to obtain image features; the image to be detected includes meteorological information;

[0049] The fusion module is used to fuse image features based on the neck network of the weather recognition model to obtain target fusion features;

[0050] The detection module is used to detect the target fusion features based on the detection network of the weather recognition model to obtain the detection results of abnormal weather in the image to be detected;

[0051] Among them, the neck network is equipped with an EGMA module, which enhances the expressiveness and global correlation strength of weather features based on at least two attention mechanisms.

[0052] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method according to the first aspect are implemented.

[0053] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method of the first aspect are implemented.

[0054] In a fifth aspect, the present application provides a computer program product, which includes a computer program. When the computer program is executed by one or more processors, it implements the steps of the method of the first aspect.

[0055] Compared with the prior art, the present application has the following advantages: the weather recognition model introduces an Enhanced Global Mixed Attention (EGMA) module into the neck network. The EGMA module can enhance the expressive power of key features by fusing at least two attention mechanisms and improve the global correlation strength between features, which helps the model better handle images with poor quality and difficult-to-distinguish panoramic backgrounds, such as occluded or blurred images. The target fusion features extracted and fused by the neck network will contain meteorological features with strong global correlation and high discrimination, significantly improving the expressive power of the target fusion features, thereby improving the recognition accuracy of images with poor image quality and difficult-to-distinguish foreground and background.

[0056] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0058] Figure 1 Schematic diagram of the network structure of the weather recognition model provided in the embodiment of the present application;

[0059] Figure 2 Schematic diagram of the positional relationship between the detection bounding box and the predicted bounding box provided in an embodiment of the present application;

[0060] Figure 3 This is a flowchart of an abnormal weather identification method based on deep learning provided by an embodiment of the present application;

[0061] Figure 4 Schematic diagram of the network structure of the EGMA module provided in the embodiment of the present application;

[0062] Figure 5 Schematic diagram of the network structure of the C2fPRFF module provided in an embodiment of the present application;

[0063] Figure 6 This is a schematic diagram of the network structure of the PRFFBlock provided in an embodiment of the present application;

[0064] Figure 7 Schematic diagram of the structure of the abnormal weather identification device provided in an embodiment of the present application;

[0065] Figure 8 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0066] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it should be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuit methods, and other related information are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0067] Inclement weather often physically impacts sensors, such as camera lenses being blocked or covered in snow, significantly degrading input image quality. Furthermore, weather phenomena such as sandstorms, rainfall, and dense fog can easily cause dynamic background changes, further complicating the distinction between foreground objects and their background. The combined effects of these factors significantly interfere with the recognition of abnormal weather conditions, leading to a significant increase in false positives.

[0068] To address this issue, this application proposes a weather recognition model comprising a backbone network and a neck network detection head. Specifically, the backbone network extracts features from an input image to be detected, i.e., an image containing meteorological information, to obtain image features; the neck network fuses these image features to obtain target fused features; and the detection head detects the target fused features to obtain a detection result for abnormal weather damage in the image to be processed.

[0069] To help the model better handle images with poor quality and indistinguishable foreground and background, the EGMA module enhances the expressiveness of key features and improves the global correlation strength between features through the fusion of at least two attention mechanisms. This helps the model better handle images with poor quality and indistinguishable panoramic backgrounds, such as occluded or blurred images. The target fusion features extracted and fused through this neck network will include meteorological features with strong global correlation and high discrimination, significantly improving the expressiveness of the target fusion features and, in turn, improving the recognition accuracy of images with poor quality and indistinguishable foreground and background.

[0070] In some embodiments, the EGMA module includes a channel-aware modulation submodule, an ECA submodule, and a cross-channel spatial attention submodule (ESA submodule). This hybrid of multiple attention mechanisms not only enhances the representation of key information within features, but also effectively facilitates cross-channel information exchange and accurately extracts important features in the spatial dimension. This allows the model to effectively handle abnormal weather conditions, even in scenarios where image quality degrades due to inclement weather. This structural improvement comprehensively enhances the model's ability to understand and process complex patterns.

[0071] In some embodiments, the channel-aware modulation submodule includes a dimensionality permutation layer, a two-layer multi-sensor layer, an inverse dimensionality permutation layer, and the first weighted activation function layer. Through the collaborative design of each network layer, the channel-aware modulation submodule enables the network to learn the importance and relevance between channels, thereby selectively enhancing key channels and suppressing redundant channels, highlighting effective feature responses, and improving the discriminability and robustness of feature expression. This makes the target fusion feature more discriminative, and overall enhances the model's ability to perceive abnormal weather in complex scenarios.

[0072] In some embodiments, the cross-channel spatial attention submodule includes a channel shuffling layer, a CBR layer, a CBS layer and a second weighted layer. Among them, the introduction of the channel shuffling layer promotes cross-channel information interaction, and the CBR and CBS layers extract spatial attention features in turn and perform nonlinear transformation and normalization processing, which helps to enhance the spatial response of salient areas. Therefore, the overall design of the cross-channel spatial attention submodule can enhance the model's ability to focus on key spatial areas, strengthen the expression of spatial features, and provide more recognizable spatial information for the final target fusion features, thereby improving the perception accuracy and robustness in complex scenes.

[0073] Abnormal weather conditions, especially dust storms, fog, rain, and heavy snow, not only reduce visibility, blurring images captured by cameras and affecting input image quality, but also significantly alter image contrast. For example, white backgrounds in snowy weather can easily become overexposed, while dust storms and fog can darken the image overall. These changes in image contrast not only diminish the visibility of abnormal weather features but also complicate feature extraction, interfering with the accuracy of recognition algorithms and significantly increasing the probability of false and missed detections.

[0074] In some embodiments, to reduce the probability of false detection and missed detection, the backbone network is equipped with a C2f Progressive Receptive Field Fusion (C2fPRFF) module, which includes a first convolutional layer, a progressive fusion structure consisting of at least two PRFFBlocks, and a first concatenated layer and a second convolutional layer. The progressive fusion structure gradually stacks and fuses features using PRFFBlocks with multiple levels of receptive fields, effectively capturing multi-scale feature information and improving the perception of objects of varying sizes and contextual structures. This module helps enhance the richness and hierarchy of feature representations, extracting more robust and discriminative feature representations.

[0075] In some embodiments, PRFFBlock includes a progressive convolution structure, a second splicing layer, a third convolution layer, and an ESE layer third splicing layer, where the progressive convolution structure includes at least two convolution layers in series. The progressive convolution structure enables the model to learn more refined contextual information and capture richer spatial hierarchical features by continuously and progressively stacking multiple convolution operations and performing feature fusion. This improvement significantly enhances the performance of the model in complex environments, especially in severe weather conditions, effectively improving the accuracy and robustness of detection and reducing the probability of false detection and missed detection.

[0076] In some embodiments, the weather recognition model can be improved based on the YOLO series models, given the following advantages of the YOLO series models:

[0077] The YOLO series of models are end-to-end, single-stage detection frameworks that achieve efficient object detection. Compared to two-stage detectors such as Faster R-CNN, YOLO directly regresses bounding boxes and categories through a single forward pass, improving detection speed and making it suitable for real-time applications. Its architecture has been continuously optimized, including the introduction of anchor-free mechanisms and feature pyramids (FPN and PAN), enhancing its detection capabilities for small and multi-scale objects. Furthermore, the YOLO series of models has been continuously optimized for lightweightness, computational efficiency, and robustness, making it highly adaptable to both embedded devices and cloud-based inference scenarios. YOLOv8 is a major upgrade to the YOLO series, supporting tasks such as object detection, image classification, and instance segmentation. Its architecture consists of a backbone network, a neck network, and a detection head. The backbone network uses a C2f module to improve feature extraction efficiency, while the neck network uses a PANet architecture to enhance multi-scale feature fusion. The detection head incorporates an anchor-free design and a Dependent Fluent (DFL) loss to improve detection accuracy and flexibility. YOLOv8 also offers five model variants: n / s / m / l / x, adapting to different scenarios. For the abnormal weather recognition task, YOLOv8s was optimized to balance accuracy and real-time performance, improving detection results.

[0078] Based on this, YOLOv8s can be preferred when building a weather recognition model. For example, if YOLOv8s is used as the basic network and improved by the EGMA module and C2fPRFF module, the network structure of the weather recognition model can be found in Figure 1 .

[0079] based on Figure 1In this weather recognition model, the backbone network extracts multi-scale features from the input image through a sequentially connected sequence of GroupConv-GroupConv-C2fPRFF–3×[GroupConv- C2fPRFF×2]-SPPF layers. GroupConv is a special form of convolution that groups input features by channel, performs independent convolution on each group, and then concatenates the outputs of all groups. Each [GroupConv- C2fPRFF×2] layer can be considered a feature extraction module for one scale, and three [GroupConv- C2fPRFF×2] layers represent feature extraction modules for three scales connected in sequence. The fast spatial pyramid pooling (SPPF) layer rapidly extracts multi-scale features without changing the size of the input feature map, improving the model's ability to perceive objects of varying sizes while reducing computational overhead. SPPF uses a continuous process of pooling, concatenation, and convolution to rapidly extract multi-scale information, enhancing feature representation while minimizing computational overhead. Among them, the output of the feature extraction module of the last scale will pass through SPPF, and the output of SPPF will be used as the input of the neck network of this scale for feature extraction.

[0080] Based on the feature extraction at each scale in the backbone network, the neck network implements feature fusion operations at the corresponding scale. For each scale, the output features of the other two scales are aligned and then concatenated. These features are then output to the corresponding decoupled detection head via the c2f module and the EGMA module. Scale alignment uses convolution to align large scales with small ones, while upsampling is used to align small scales with large ones.

[0081] The corresponding detection head detects the target fusion features at each scale and outputs the detection results.

[0082] During training, in addition to using a 4 × reg_max distributed regression strategy to make the detection box regression more detailed and stable, thereby improving positioning accuracy, the total number of categories (nc) is set to guide the output layer to generate a corresponding number of category prediction results, ensuring that the detection task can cover all target categories.

[0083] When the trained weather recognition model is put into application, redundant detection frame removal (such as non-maximum suppression, NMS) and confidence filtering operations can be further introduced to eliminate overlapping detection frames and screen out high-confidence prediction results, thereby improving the accuracy and reliability of the final detection results.

[0084] To enhance the perception capability of the weather recognition model, the backbone network introduces the C2fPRFF module, which realizes the progressive fusion of multi-scale receptive fields through stacked PRFFBlocks, enhances the understanding of contextual information, and improves its ability to process complex content, thereby reducing the probability of false detection and missed detection in abnormal weather recognition. The neck network sets up the EGMA module, which integrates the channel attention and spatial attention mechanisms, and introduces the channel shuffling operation, effectively improving the feature expression capability and enhancing the adaptability and robustness of the model in scenarios with poor image quality or where the foreground and background are difficult to distinguish.

[0085] In some embodiments, to more accurately address the challenge of abnormal weather identification, a specialized dataset can be created for training weather recognition models. Specifically, a dataset called Abnormal Weather Detect (AWD) is created. This dataset focuses on collecting and analyzing instances of abnormal weather. The images in the dataset can be sourced from the internet or from samples collected on-site in real-world vehicle environments to ensure high data authenticity and broad diversity. The AWD dataset is unique in that it covers a variety of abnormal weather scenarios, providing a valuable and representative resource for model training and evaluation.

[0086] This application focuses on identifying and classifying four common abnormal weather conditions, aiming to accurately address the multiple challenges of representative weather identification. These four abnormal weather categories can include sandstorms, fog, rain, and snow, each of which is carefully classified and defined. The AWD dataset, created based on these four abnormal weather conditions, contains 3,000 high-quality images, providing a rich visual resource for research. Each image is manually annotated.

[0087] For example, each scene can include multiple time periods such as day and night to increase the complexity and comprehensiveness of the scene. Each image can be configured with a corresponding txt label file that details the location and category of abnormal weather.

[0088] In order to ensure the effectiveness of model training and the reliability of evaluation results, each image in the dataset of this application is divided into a training set and a validation set in a ratio of 8:2.

[0089] In some embodiments, in order to improve the positioning accuracy of abnormal weather and the convergence speed of the model, the loss function includes classification loss during the training of the weather recognition model. Specifically, the classification loss adopts Binary Cross-Entropy Loss (BCE Loss), which is a loss function commonly used in binary classification problems. It measures the difference between the model's predicted category and the true category and is used to determine the specific category in the anchor box. Specifically, for each sample, if its true category is yi ∈[0,1], the predicted category is ∈[0,1], the binary cross entropy loss can be expressed as:

[0090]

[0091] in, L BCE is BCELoss, n is the number of image samples of all urban facilities anomalies, y i is the true category of the i-th image sample, is the predicted category of the i-th image sample.

[0092] Bounding box regression loss usually includes Complete Intersection over Union loss (CIoU Loss) and Dynamic Focal Loss (DFL loss), which are used to measure the error between the predicted bounding box and the true bounding box. CIoU Loss is improved based on commonly used loss functions such as Intersection over Union (IoU), Generalized Intersection over Union (GIoU), and Distance Intersection over Union (DIoU). Compared with previous loss functions, it adds a penalty term for aspect ratio. When the center points of the predicted bounding box and the true bounding box coincide, it can better distinguish the errors in different situations and has scale invariance. Its formula is as follows:

[0093]

[0094] in, L CIoU is CIoU Loss, b and Represent the center point of the real box and the center point of the predicted box respectively; ρ is the Euclidean distance between the predicted bounding box and the true bounding box; c is the diagonal distance between the predicted bounding box and the true bounding box; v To ensure the consistency of the relative proportions between the predicted bounding box and the true bounding box, IoU is the intersection-over-union ratio of the predicted bounding box and the true bounding box; α is the weight coefficient; w and are the width of the predicted border and the true border respectively; h and are the heights of the predicted bounding box and the true bounding box, respectively.

[0095] Based on CIoU Loss, the introduction of DFL loss can further improve the bounding box regression accuracy. DFL loss discretizes the uncertainty in the bounding box coordinate prediction, making the regression process more refined and robust, effectively reducing the prediction error. The formula of DFL loss is as follows:

[0096]

[0097] in, S i is the cross entropy loss between the real border and the predicted border on the left, S i+1 is the cross entropy loss between the true box and the predicted box on the right.

[0098] However, although the CIoU loss takes into account the three important factors of positioning loss: overlapping area, center point distance and aspect ratio, the α in the CIoU loss formula is v ,There are still problems with the design of this item, which slows down the convergence speed.

[0099] In order to improve the convergence speed, the aspect ratio penalty term is split based on the penalty term of the original CIoU loss, processing the width and height penalties separately, and introducing the angle penalty term, thus proposing a new IoU loss that integrates the Kombinance Intersection over Union (KIoU) loss. This loss function includes overlap loss, center distance loss, width and height loss, and angle loss. Figure 2 The core idea of ​​KIoU loss includes: first, guiding the predicted box to quickly fit to the closest axis position, that is, making the center point of the predicted box parallel to the center point of the true box in the horizontal or vertical direction. After that, the predicted box only needs to be further adjusted in one direction. By introducing the angle penalty term, the total number of degrees of freedom can be effectively reduced and the convergence speed can be accelerated. At the same time, splitting the aspect ratio penalty term into a high penalty term and a wide penalty term can more accurately adjust the model's prediction error in different directions. The splitting of the aspect ratio penalty term avoids treating the aspect ratio as a whole, allowing the error of each dimension (width and height) to be optimized more independently and specifically, thereby making the predicted positioning results more accurate.

[0100] Specifically, the formula for KIoU loss is as follows:

[0101]

[0102]

[0103]

[0104]

[0105]

[0106]

[0107] in, is the intersection over union loss; IoU is the intersection over union ratio between the predicted box and the true box; L IoU is the intersection and comparison loss; L dis is the center distance loss; L shp is the width and height loss; L ang is the angle loss; Δ is 2 times the center distance loss; b and Represent the center point of the real box and the center point of the predicted box respectively; Represents the square of the Euclidean distance between the center point of the real box and the predicted box; and Represent the width and height of the minimum bounding rectangle respectively; Ω is 2 times the width and height loss; w and are the width of the predicted border and the true border respectively; Represents the square of the Euclidean distance between the true box and the predicted box width; h and are the heights of the predicted bounding box and the true bounding box respectively; Represents the square of the Euclidean distance between the true box and the predicted box height; The angle loss is 2 times; and Represent the width and height of the rectangular box constructed by the center points of the true box and the predicted box, respectively.

[0108] Considering the sample imbalance problem in the bounding box regression process, that is, the number of high-quality prediction boxes (small errors) in an image is far less than the number of low-quality prediction boxes (large errors). These low-quality prediction boxes may produce excessively large gradients, which will adversely affect the training process. To address this problem, this application proposes a new loss function that combines Focal Loss and KIoU loss, called Focal Kombine Intersection over Union (Focal KIoU Loss). The formula is as follows:

[0109]

[0110] in, To focus on the fusion intersection-over-union loss, γ is a parameter that controls the degree of outlier suppression.

[0111] Combining the above formulas, the formula for the total loss function Loss is as follows:

[0112]

[0113] Among them, λ1 and λ2 are balance coefficients.

[0114] Performing weather recognition as described in any of the aforementioned embodiments based on this total loss function can improve the model convergence speed and thus the efficiency of model training. By continuously iterating the weather recognition model, the model can better and more stably recognize abnormal weather, and ultimately a robust version can be selected from multiple converged versions as the trained weather recognition model.

[0115] In some embodiments, weather recognition model training combines classification and regression loss optimization. Classification loss uses binary cross entropy loss (BCE Loss) to determine anchor box categories; regression loss consists of focal KIoU loss and Dependency Loss (DFL Loss) to measure the error between the predicted box and the ground-truth box. The positive and negative sample matching strategy uses TAL dynamic matching to optimize target allocation and improve detection accuracy.

[0116] In some embodiments, in order to comprehensively and accurately measure the performance of each version of the weather recognition model, after at least one version of the weather recognition model is converged based on the training set, the converged model can be evaluated through the validation set to avoid model overfitting, verify the generalization of the model, and ensure that the model can run stably after deployment.

[0117] Specifically, the performance of the weather recognition model on the validation set can be evaluated according to preset conditions.

[0118] For example, these preset conditions may include performance indicators such as the intersection-over-union ratio, detection accuracy, and recall rate of the abnormal weather recognition device for abnormal weather.

[0119] IoU is an indicator that evaluates the degree of overlap between the predicted bounding box and the true bounding box. It calculates the ratio of the intersection and union of the predicted bounding box and the true bounding box. It plays a key role in determining whether it is a correct detection. The calculation formula of IoU can be written as:

[0120]

[0121] Precision (p), also known as the precision rate, refers to the ratio of correct positive predictions to all positive predictions, as shown in the formula:

[0122]

[0123] Recall (R), also known as the recall rate, refers to the ratio of correctly predicted positive results to all actual positive results, as shown in the formula:

[0124]

[0125] F1-Score is a comprehensive indicator for measuring model performance. It combines the dual advantages of precision and recall to comprehensively evaluate the balanced performance of the model in abnormal weather detection tasks, as shown in the formula:

[0126]

[0127] Average Precision (AP) is calculated from precision and recall. A line graph of precision is drawn based on the recall value, and the area under the line is calculated, as shown in the formula:

[0128]

[0129] The mean average precision (mAP) refers to the average of the average precision AP of C different abnormal weather categories, as shown in the formula:

[0130]

[0131] That is to say, after verifying each version of the weather recognition model through the validation set, the weather recognition model of each version can be comprehensively evaluated based on the above indicators, so as to determine the weather recognition model with the best performance from each version as the trained weather recognition model.

[0132] In some embodiments, the model runtime environment includes an Intel Xeon Platinum 8255C processor, 314 GB of memory, an NVIDIA Tesla V100 32 GB graphics card, and a CentOS 8.5.2 (64-bit) operating system. This weather recognition model is built based on the PyTorch framework, with an input image size of [640, 640] and a multi-scale training strategy. The experiment set the batch size to 64 and trained for 200 epochs. During model training, the SGD optimizer was used with an initial learning rate of 0.01, combined with a cosine decay strategy for dynamic adjustment of the learning rate. The SGD optimizer can set the momentum factor to 0.937 to help accelerate the gradient descent process and suppress oscillations; the weight decay coefficient is set to 0.0005 to regularize the model and prevent overfitting.

[0133] Based on the network structure of the weather recognition model in the previous embodiment, this application proposes an abnormal weather recognition method based on deep learning.

[0134] The abnormal weather identification method based on deep learning provided in the embodiments of the present application can be applied to electronic devices such as mobile phones, tablet computers, vehicle-mounted equipment, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). The embodiments of the present application do not impose any restrictions on the specific types of electronic devices.

[0135] In order to illustrate the technical solution proposed in this application, each embodiment will be described below using an electronic device as the execution entity.

[0136] Figure 3 A schematic flow chart of the abnormal weather identification method based on deep learning provided by the present application is shown. The abnormal weather identification method based on deep learning includes:

[0137] Step 310: The electronic device extracts features from the image to be detected based on the backbone network of the pre-trained weather recognition model.

[0138] Step 320: The electronic device fuses the image features based on the neck network of the weather recognition model to obtain target fusion features.

[0139] Step 330: The electronic device detects the target fusion features based on the detection network of the weather recognition model to obtain a detection result of abnormal weather in the image to be detected.

[0140] In this embodiment, the EGMA module introduced into the neck network enhances the expressive power of key features by fusing at least two attention mechanisms and increases the global correlation strength between features. This helps the model better handle images with poor quality and difficult-to-distinguish panoramic backgrounds, such as occluded or blurred images. The target fusion features extracted and fused through this neck network include meteorological features with strong global correlation and high discrimination, significantly improving the expressive power of the target fusion features and, in turn, improving the recognition accuracy of images with poor quality and difficult-to-distinguish foreground and background.

[0141] In some embodiments, the EGMA module includes a channel-aware modulation submodule, an ECA submodule, and a cross-channel spatial attention submodule (ESA submodule). For a first input feature input to the EGMA module, the electronic device may perform the following operations:

[0142] Step A1: The electronic device performs a channel feature enhancement operation on the first input feature through the channel-aware modulation submodule to obtain a first enhanced feature.

[0143] The electronic device processes the first input feature through a channel-aware modulation submodule, which can model the dependency between channels, selectively enhance key channels, and suppress invalid channels, thereby obtaining a first enhanced feature with stronger discrimination ability.

[0144] Step A2: The electronic device performs a channel attention operation on the first enhanced feature through the ECA submodule to obtain a second enhanced feature.

[0145] Based on the first enhanced feature, the electronic device further performs a lightweight channel attention operation through the ECA (Efficient Channel Attention) sub-module to refine and strengthen important channel information, suppress irrelevant redundancy, and output a more discriminative second enhanced feature.

[0146] Step A3: The electronic device performs a cross-channel spatial attention operation on the second enhanced feature through a cross-channel spatial attention submodule to obtain a multi-attention enhanced feature.

[0147] The electronic device applies a cross-channel spatial attention sub-module to the second enhanced feature, enhances the perception of key areas and positions through information interaction between channels and spatial feature modeling, and obtains a multi-attention enhanced feature that integrates channel and spatial attention.

[0148] Step A43: The electronic device performs a spatial attention operation on the multiple attention enhancement features through the ESA sub-module to obtain a first output feature corresponding to the first input feature.

[0149] Finally, the electronic device performs spatial attention operations on the multiple attention enhancement features through the ESA (Enhanced Spatial Attention) sub-module to further strengthen the response to significant spatial information, and finally outputs the first output feature corresponding to the first input feature to improve the overall perception effect and downstream task performance.

[0150] This embodiment constructs a multi-layered, multi-dimensional attention enhancement mechanism by gradually introducing channel-aware modulation, ECA channel attention, and cross-channel spatial attention (ESA) modules. This process effectively mines and enhances key information (especially channel and spatial information) in input features, suppresses invalid or interfering features, and improves the discriminability and robustness of feature representation by fully modeling inter-channel relationships and spatially salient regions, thereby providing higher-quality feature input for subsequent tasks.

[0151] In some embodiments, the channel-aware modulation submodule includes a dimension permutation layer, a two-layer multi-perceptron layer, an inverse dimension permutation layer, an activation function layer, and a first weighted layer; the channel-aware modulation submodule performs a channel feature enhancement operation on the first input feature to obtain a first enhanced feature, including:

[0152] Step A11: The electronic device permutes the dimensional parameters of the first input feature through a dimensional permutation layer to obtain a flattened spatial feature.

[0153] The electronic device uses a dimensionality permutation layer to rearrange the spatial and channel dimensions of the first input feature. For example, it transforms the feature shape from B×C×H×W (batch×number of channels×height×width) to B×(H×W)×C, flattening the spatial information while retaining the channel feature vector corresponding to each spatial location. This operation helps to model the dependencies between features in the channel dimension, thereby improving the expressiveness and discriminability of channel features.

[0154] Step A12: The electronic device performs a channel compression operation and a nonlinear activation operation on the flattened spatial features through the first multi-sensor layer to obtain compressed features.

[0155] Step A13: The electronic device performs a channel expansion operation on the compressed feature through the second multi-sensor layer to obtain a first expanded feature that matches the dimension of the flattened spatial feature.

[0156] The two-layer multi-perceptron layer effectively models inter-channel dependencies. Specifically, the electronic device uses the first multi-perceptron layer to perform channel compression and nonlinear activation on the flattened spatial features, extracting key channel information and obtaining compact compressed features. Subsequently, the second multi-perceptron layer performs channel expansion on the compressed features, restoring them to the original channel dimension to obtain the first expanded features, thus ensuring feature structure alignment and maintaining information integrity.

[0157] Step A14: The electronic device performs inverse permutation on the dimensional parameters of the first extended feature through an inverse dimensional permutation layer to obtain a reconstructed feature whose dimensional parameters match the first input feature.

[0158] The electronic device restores the first extended feature from B×(H×W)×C to a spatial-channel arrangement of B×C×H×W through an inverse dimensional permutation operation, and obtains a reconstructed feature consistent with the dimension of the first input feature.

[0159] Step A15: The electronic device performs a linear activation operation on the reconstructed features through an activation function layer to obtain a channel modulation weight.

[0160] Step A16: The electronic device performs a weighted operation on the first input feature based on the channel modulation weight through the first weighted layer to obtain a first enhanced feature.

[0161] The electronic device normalizes the reconstructed features using an activation function (such as Sigmoid) to generate channel modulation weights. The first input features are then weighted channel by channel through the first weighting layer to highlight key features and suppress irrelevant interference, resulting in the first enhanced features. Specifically, the weighting operation essentially represents the channel modulation weights as channel attention features and multiplies them element-wise with the first input features.

[0162] In some embodiments, the cross-channel spatial attention submodule includes a channel shuffling layer, a CBR layer, a CBS layer, and a second weighted layer; the cross-channel spatial attention submodule performs a cross-channel spatial attention operation on the second enhanced feature to obtain a multi-attention enhanced feature, including:

[0163] Step A31: The electronic device performs a channel shuffling operation on the second enhanced feature through the channel shuffling layer to obtain a recombined feature.

[0164] The electronic device rearranges the channels of the second enhanced features through a channel shuffling layer, promotes information interaction between channels, more effectively mixes and enriches feature information, and obtains more expressive recombined features.

[0165] Step A32: The electronic device sequentially performs channel compression, normalization, and nonlinear transformation operations on the reconstructed features through the CBR layer to obtain compressed activation features.

[0166] The electronic device performs channel compression, normalization and nonlinear transformation operations on the reconstructed features in sequence through the CBR layer (Convolution, Batch NormalizationReLU activation function, which is a combination of convolution, batch normalization and ReLU activation) to extract compact and discriminative compressed activation features.

[0167] Step A33: The electronic device sequentially performs channel expansion operation, normalization operation, and linear activation operation on the compressed activation features through the CBS layer to obtain spatial weights.

[0168] The electronic device performs channel expansion, normalization and linear activation operations on the compressed activation features in sequence through the CBS layer (Convolution, Batch Normalization Sigmoid activation function, that is, a combination of convolution, batch normalization and Sigmoid activation) to generate normalized spatial attention weights. The spatial attention weights can be used to highlight the discriminative regional features in the image and suppress irrelevant interference.

[0169] Step A34: The electronic device performs a weighted operation on the reorganized features based on the spatial weights through the second weighted layer to obtain multi-attention enhanced features.

[0170] Finally, the electronic device uses the second weighting layer to perform a weighted operation on the reorganized features based on the spatial weights to obtain the multi-attention enhanced features. Specifically, the weighting operation essentially multiplies the spatial weights as spatial attention features by the reorganized features element-by-element.

[0171] In this embodiment, the electronic device introduces a channel shuffling mechanism and a layer-by-layer compression-expansion structure, combined with spatial attention weights to guide feature enhancement, effectively improving the spatial differentiation ability of features and the collaborative perception ability between channels, thereby enhancing the model's perception of key areas.

[0172] In some embodiments, Figure 4 The network structure diagram of the EGMA module is shown. Figure 4 In this network structure, for the first input feature, the electronic device first extracts attention features in the channel dimension through the channel attention submodule. It then performs a channel shuffling operation to promote the interaction and fusion of cross-channel information. The processed features are then input into the spatial attention submodule to further explore significant features in the spatial dimension. Through this step-by-step processing, the model can more effectively focus on key areas and important channels, optimize feature representation, and improve overall task performance.

[0173] Specifically, if the first input feature is expressed as In the channel attention submodule, first, the first input feature is converted from B×C×H×W to B×(H×W)×C through a dimensionality permutation operation. Then, a two-layer multilayer perceptron (MLP) is used to capture the dependencies between channels. The first layer of MLP reduces the number of channels to 1 / 4 of the original and introduces nonlinearity through the ReLU activation function; then, the second layer of MLP restores the number of channels to the original dimension to obtain the first extended feature. Next, the dimensional parameters of the first extended feature are restored to B×C×H×W through inverse dimensionality permutation to obtain the reconstructed feature; the reconstructed feature generates the channel attention feature through the Sigmoid activation function. Finally, the channel attention feature is multiplied element-by-element with the first input feature to obtain the first enhanced feature; the first enhanced feature is further processed by the ECA module to obtain the second enhanced feature. Y GCA .

[0174]

[0175] in, Permute is the dimension permutation operation, X Permute To flatten the spatial features, W 1 is the channel compression operation performed by the first layer MLP, and ReLU is the nonlinear operation performed by the activation function; W 2 is the channel expansion operation performed by the second layer MLP; X GCA is the channel attention map, Reverse Permute is the inverse dimension permutation operation.

[0176] In order to promote further mixing and sharing of feature information, a channel shuffling operation is introduced. Specifically, first, Y GCA The features are divided into 4 groups, each containing C / 4 channels. Then, the channel order within each group is transposed to randomize the channel order within each group. Then, the shuffled features are reassembled to restore the original B×C×H×W shape to obtain the recombined features. Y CS . Y CS It can promote the effective mixing and sharing of features from subsequent channels, thereby improving the ability of feature expression.

[0177] ChannelShuffle Performs a channel shuffle operation.

[0178] Assume that the convolution operation in the CBR layer is 5×5 convolution to achieve channel compression, and the convolution in the CBS layer is 7×7 convolution to achieve channel restoration. Then after the channel shuffling operation, in the cross-channel spatial attention submodule, first, YCS After a 5×5 convolution layer, the number of channels is compressed to 1 / 4 of the original; then, batch normalization and ReLU activation function are used for nonlinear transformation to obtain compressed activation features. Then, a 7×7 convolution layer is used to restore the number of channels to the original dimension C, and batch normalization is performed again. Subsequently, the Sigmoid activation function is used to generate spatial weights, namely spatial attention features. Finally, the spatial attention features are X GSA and Y CS Multiply element by element to get the multi-attention enhanced features. Input the multi-attention enhanced features into the ESA module to further optimize the spatial features and get the first output feature O . O It integrates channel and spatial saliency information and has stronger feature expression capabilities.

[0179]

[0180] Among them, the mathematical description of the ECA module can be expressed as follows:

[0181] Given an input tensor , then:

[0182]

[0183] The mathematical description of the ESA module can be expressed as follows:

[0184] Given an input tensor , then:

[0185]

[0186] In this embodiment, the EGMA module significantly enhances the expressive power of the input features. Even when the image quality is degraded due to adverse weather conditions such as lens occlusion and snow cover, it can still effectively extract and enhance key visual information, thereby reducing the interference of the physical environment on sensor perception and improving the robustness of the detection system. In the face of dynamic background changes caused by sandstorms, rain, and dense fog, EGMA relies on the spatial attention sub-module to accurately capture the significant features of the spatial dimension, helping the model to accurately distinguish between foreground and complex background, reduce false detection rate and improve detection accuracy. At the same time, EGMA also promotes cross-channel information interaction through a channel shuffling mechanism, and combines the spatial attention reinforcement structure to effectively improve the perception ability and stability of the model in complex environments.

[0187] In some embodiments, the backbone network is provided with a C2fPRFF module. Figure 5A schematic diagram of the network structure of the C2fPRFF module is shown, including a first convolutional layer, a progressive fusion structure including N PRFFBlocks, a first splicing layer, and a second convolutional layer, where N ≥ 2. For each second input feature input to the C2fPRFF module, the electronic device may perform the following steps:

[0188] Step B1: The electronic device performs a channel compression operation on the second input feature through the first convolution layer to obtain a convolution feature.

[0189] The electronic device performs a channel compression operation on the second input feature through the first convolution layer to extract compact low-dimensional convolution features, reduce computational complexity and highlight key semantic information.

[0190] Step B2: The electronic device sequentially performs progressive fusion operations on the convolution features through the serially connected PRFFBlocks via the progressive fusion structure to obtain hierarchical fusion features corresponding to each level.

[0191] The electronic device sends the convolutional features into the progressive fusion structure. The progressive fusion structure gradually extracts multi-scale contextual information through the gradual superposition and fusion of multi-level receptive fields, and generates fusion features corresponding to each level.

[0192] Step B3: The electronic device performs a splicing operation on the convolutional features and the fusion features of each level through the first splicing layer to obtain a first splicing feature.

[0193] The electronic device splices the original convolution features with the fusion features of each level in the channel dimension through the first splicing layer to form a first splicing feature containing multi-scale information, thereby enhancing the richness of feature representation.

[0194] Step B4: The electronic device performs a channel expansion operation on the first splicing feature through a second convolutional layer to obtain a second output feature corresponding to the second input feature; and the image feature is obtained based on the second output feature.

[0195] The electronic device performs a channel expansion operation on the first spliced ​​feature through the second convolution layer, restores the feature to the target dimension, and generates a second output feature corresponding to the second input feature as the basis for subsequent image feature extraction.

[0196] In this embodiment, the electronic device performs multi-scale semantic fusion and feature enhancement on the second input features. The progressive PRFFBlock structure refines contextual information layer by layer, effectively improving the model's ability to recognize complex structures. Multi-level feature concatenation and channel expansion operations further enrich feature expression, significantly enhancing the model's ability to capture details and semantic information in tasks such as image understanding and object detection.

[0197] In some embodiments, PRFFBlock includes a progressive convolution structure, a second convolution layer, a third convolution layer, an ESE layer, and a third convolution layer, wherein the progressive convolution structure includes at least two convolution layers connected in series; for the third input feature of each PRFFBlock:

[0198] Step C1: The electronic device performs a progressive convolution operation on the third input feature through at least two convolution layers connected in series in a progressive convolution structure to obtain hierarchical convolution features corresponding to each level.

[0199] The electronic device performs layer-by-layer convolution processing on the third input feature through a progressive convolution structure (consisting of at least two convolution layers in series), extracts multi-scale semantic information from shallow to deep layers, obtains hierarchical convolution features corresponding to each level, and gradually enhances local details and global structural information.

[0200] Step C2: The electronic device performs a splicing operation on the third input feature and the convolution features of each level through the second splicing layer to obtain a second splicing feature.

[0201] The electronic device splices the original third input feature with the convolution features of different levels in the channel dimension through the second splicing layer to obtain the second splicing feature, thereby realizing multi-scale fusion of features and improving the representation capability.

[0202] Step C3: The electronic device performs a channel expansion operation on the second concatenated feature through the third convolutional layer to obtain a second expanded feature.

[0203] The electronic device performs a channel expansion operation on the second concatenated feature through the third convolutional layer, restores the channel dimension to the expected scale, and generates a second expanded feature with complete structure and rich semantics.

[0204] Step C4: The electronic device performs a global average pooling operation and a convolution operation on the second extended feature through the ESE layer, and weights the second extended feature with the calculated channel attention weight to obtain an initial fusion feature.

[0205] The electronic device performs a global average pooling operation on the second extended features through the ESE layer (Efficient Squeeze-and-Excitation layer), calculates the channel attention weights through the convolution operation, and weights the second extended features channel by channel to obtain an initial fused feature with key enhancement to highlight the key semantic areas.

[0206] Step C5: The electronic device performs a splicing operation on the third input feature and the initial fusion feature through a third splicing layer to obtain a third output feature corresponding to the third input feature.

[0207] The electronic device splices the third input feature with the initial fusion feature through the third splicing layer to form a third output feature that contains both original information and enhanced attention information, which is used for subsequent higher-level feature extraction and task processing.

[0208] In this embodiment, the electronic device performs multi-level convolution extraction, feature fusion, and channel attention enhancement on the third input feature. The progressive convolution structure ensures the depth and continuity of feature extraction, the splicing mechanism integrates multi-scale information, and the attention mechanism introduced by the ESE layer further enhances the model's perception of key channel features. The final output third output feature significantly enhances the discriminability and expressiveness of the feature while maintaining the integrity of the original information, providing strong feature support for subsequent visual tasks.

[0209] In some embodiments, Figure 6 The network structure diagram of PRFFBlock is shown. Figure 6 The network structure, for the third input feature , PRFFBlock performs the following steps:

[0210] First, in order to facilitate the splicing of scale features, a list can be initialized Y , including the third input feature X :

[0211]

[0212] Then, for the i-th (i∈{0, 1, ..., n-1}) convolution layer, the level convolution features output by the i-1 convolution layers are used as the input of the i convolution layer, and a 3×3 convolution operation is performed. It is worth noting that, except for the first convolution layer, the input is the third input feature. X Except for the above, the inputs of other convolutional layers are all the hierarchical convolution features output by the previous layer.

[0213]

[0214] in, Y 0= X After each convolution operation, the result is added to the list Y middle.

[0215] The list Y The output of all convolutional layers and the third input feature X The second splicing feature is obtained by splicing, and the second extended feature is obtained by channel fusion through 1×1.

[0216]

[0217] Finally, the ESE attention mechanism is used to adjust the importance of the second extended feature with different receptive field sizes to obtain the initial fusion feature Y Agg ; Use residual connection to connect the third input feature X With the initial fusion feature Y Agg Splicing to get the third output feature O .

[0218]

[0219] In this embodiment, PRFFBlock achieves progressive feature extraction and refined modeling by continuously stacking multiple 3×3 convolution operations and fusing multi-scale features. This structure helps the model learn richer contextual relationships and spatial hierarchical information, significantly improving the accuracy and robustness of object detection in complex environments, especially in adverse weather conditions such as heavy rain, dense fog, or low light.

[0220] In some embodiments, the recognition method introduces a C2fPRFF module into the backbone network, achieves progressive fusion through stacked PRFFBlocks, extracts multi-scale receptive field features, and enhances the model's ability to understand contextual information. The EGMA module is integrated into the neck network, combining channel attention and spatial attention mechanisms, and introducing channel shuffling operations to effectively improve feature expression capabilities and enhance the model's performance in scenes with degraded image quality and in which foreground and background are difficult to distinguish. Furthermore, a newly designed focused fusion intersection-over-union loss function is used to accelerate model convergence while optimizing target positioning accuracy, significantly improving robustness and accuracy under complex background and interference conditions. Through the collaborative optimization of the above three aspects, the detection performance and reliability of the weather recognition model in harsh environments are significantly enhanced.

[0221] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0222] Corresponding to the abnormal weather identification method based on deep learning in the above embodiment, Figure 7 A structural block diagram of the abnormal weather identification device 7 provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.

[0223] Reference Figure 7 , the abnormal weather identification device 7 includes:

[0224] Extraction module 71, used for extracting features from the image to be detected based on the backbone network of the pre-trained weather recognition model to obtain image features; the image to be detected includes meteorological information;

[0225] A fusion module 72 is used to fuse image features based on the neck network of the weather recognition model to obtain target fusion features;

[0226] A detection module 73 is configured to detect target fusion features based on a detection network of a weather recognition model to obtain a detection result of abnormal weather in the image to be detected;

[0227] Among them, the neck network is equipped with an EGMA module, which enhances the expressiveness and global correlation strength of weather features based on at least two attention mechanisms.

[0228] Optionally, the EGMA module includes a channel-aware modulation submodule, an ECA submodule, and a cross-channel spatial attention submodule ESA submodule, and the fusion module 72 includes a fusion unit for:

[0229] Performing a channel feature enhancement operation on the first input feature through a channel-aware modulation submodule to obtain a first enhanced feature;

[0230] Perform a channel attention operation on the first enhanced feature through the ECA submodule to obtain a second enhanced feature;

[0231] Performing a cross-channel spatial attention operation on the second enhanced feature through a cross-channel spatial attention submodule to obtain a multi-attention enhanced feature;

[0232] The ESA submodule performs a spatial attention operation on the multiple attention enhancement features to obtain the first output feature corresponding to the first input feature.

[0233] Optionally, the channel-aware modulation submodule includes a dimension permutation layer, a double-layer multi-sensor layer, an inverse dimension permutation layer, an activation function layer, and a first weighted layer; the fusion unit is specifically configured to:

[0234] The dimensional parameters of the first input feature are permuted through the dimensional permutation layer to obtain the flattened spatial feature;

[0235] The first multi-perceptron layer performs a channel compression operation on the flattened spatial features and a nonlinear activation operation to obtain compressed features;

[0236] Performing a channel expansion operation on the compressed features through a second multi-perceptron layer to obtain a first expanded feature that matches the dimension of the flattened spatial feature;

[0237] Perform inverse permutation on each dimensional parameter of the first extended feature through an inverse dimensional permutation layer to obtain a reconstructed feature whose dimensional parameters match the first input feature;

[0238] Perform linear activation operations on the reconstructed features through the activation function layer to obtain channel modulation weights;

[0239] A first enhanced feature is obtained by performing a weighted operation on the first input feature based on the channel modulation weight through the first weighted layer.

[0240] Optionally, the cross-channel spatial attention submodule includes a channel shuffling layer, a CBR layer, a CBS layer, and a second weighted layer; the fusion unit is specifically used to:

[0241] Performing a channel shuffling operation on the second enhanced feature through a channel shuffling layer to obtain a recombined feature;

[0242] Through the CBR layer, channel compression operation, normalization operation and nonlinear transformation operation are performed on the recombined features in sequence to obtain compressed activation features;

[0243] The CBS layer performs channel expansion, normalization, and linear activation operations on the compressed activation features to obtain spatial weights.

[0244] The second weighted layer performs a weighted operation on the reorganized features based on the spatial weights to obtain multi-attention enhanced features.

[0245] Optionally, the backbone network is provided with a C2fPRFF module, including a first convolutional layer, a progressive fusion structure including at least two PRFFBlocks, a first splicing layer and a second convolutional layer. For each second input feature input to the C2fPRFF module, the extraction module includes an extraction unit, and the extraction unit is used to:

[0246] Perform channel compression on the second input feature through the first convolution layer to obtain convolution features;

[0247] Through the progressive fusion structure, the convolution features are sequentially fused through the series of PRFFBlocks to obtain the hierarchical fusion features corresponding to each level;

[0248] The convolution feature and the fusion feature of each level are concatenated through the first concatenation layer to obtain the first concatenation feature;

[0249] A channel expansion operation is performed on the first concatenated feature through a second convolutional layer to obtain a second output feature corresponding to the second input feature; and the image feature is obtained based on the second output feature.

[0250] Optionally, the PRFFBlock includes a progressive convolution structure, a second convolution layer, a third convolution layer, and an ESE layer third convolution layer, where the progressive convolution structure includes at least two convolution layers connected in series; for the third input feature input to each PRFFBlock: the extraction unit is specifically used to:

[0251] Performing a progressive convolution operation on the third input feature through at least two convolutional layers connected in series in a progressive convolution structure to obtain hierarchical convolution features corresponding to each level;

[0252] Performing a splicing operation on the third input feature and the convolution features of each level through the second splicing layer to obtain a second splicing feature;

[0253] Performing a channel expansion operation on the second concatenated feature through the third convolutional layer to obtain a second expanded feature;

[0254] Perform global average pooling and convolution operations on the second extended feature through the ESE layer, and weight the calculated channel attention weight on the second extended feature to obtain the initial fusion feature;

[0255] A splicing operation is performed on the third input feature and the initial fusion feature through the third splicing layer to obtain a third output feature corresponding to the third input feature.

[0256] Optionally, the weather recognition model is trained based on regression loss and classification loss; the regression loss includes focused fusion intersection-over-union loss, and the formula is as follows:

[0257]

[0258]

[0259]

[0260]

[0261]

[0262]

[0263]

[0264] in, To focus on the fusion intersection-over-union loss, IoU is the intersection-over-union ratio between the predicted box and the true box; γ is a parameter that controls the degree of outlier suppression; To blend and compare losses; L IoU is the intersection and comparison loss; L dis is the center distance loss; L shp is the width and height loss; L ang is the angle loss; Δ is 2 times the center distance loss; b and Represent the center point of the real box and the center point of the predicted box respectively; Represents the square of the Euclidean distance between the center point of the real box and the predicted box; and Represent the width and height of the minimum bounding rectangle respectively; Ω is 2 times the width and height loss; w and are the width of the predicted border and the true border respectively; Represents the square of the Euclidean distance between the true box and the predicted box width; h and are the heights of the predicted bounding box and the true bounding box respectively; Represents the square of the Euclidean distance between the true box and the predicted box height; The angle loss is 2 times; and Represent the width and height of the rectangular box constructed by the center points of the true box and the predicted box, respectively.

[0265] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0266] Figure 8 This is a schematic diagram of the physical structure of an electronic device provided in one embodiment of the present application. Figure 8 As shown, the electronic device 8 of this embodiment includes: at least one processor 80 ( Figure 8 Only one processor is shown), a memory 81 stores a computer program 82 in the memory 81 and can be run on at least one processor 80, and when the processor 80 executes the computer program 82, the steps in any of the above-mentioned abnormal weather identification method embodiments based on deep learning are implemented, such as Figure 3 Steps 310-330 are shown.

[0267] The processor 80 may be a central processing unit (CPU), or other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor.

[0268] In some embodiments, the memory 81 may be an internal storage unit of the electronic device 8, such as a hard disk or memory of the electronic device 8. In other embodiments, the memory 81 may also be an external storage device of the electronic device 8, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the electronic device 8.

[0269] Furthermore, the memory 81 may include both an internal storage unit of the electronic device 8 and an external storage device. The memory 81 is used to store operating devices, applications, boot loaders, data, and other programs, such as computer program code. The memory 81 may also be used to temporarily store data that has been output or is about to be output.

[0270] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the above-mentioned device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0271] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the steps in the above-mentioned various method embodiments.

[0272] An embodiment of the present application provides a computer program product. When the computer program product is run on an electronic device, the electronic device can implement the steps in the above-mentioned method embodiments when executing the computer program product.

[0273] If this integrated unit is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application can implement all or part of the process steps in the above-mentioned method embodiments by using a computer program to instruct the relevant hardware. The above-mentioned computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The above-mentioned computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The above-mentioned computer-readable medium can at least include: any entity or device capable of carrying computer program code to a camera / electronic device, a recording medium, computer memory, read-only memory (ROM), random access memory (RAM), an electrical carrier signal, a telecommunications signal, or a software distribution medium. Examples include a USB flash drive, a removable hard drive, a magnetic disk, or an optical disk.

[0274] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0275] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0276] In the embodiments provided in this application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely illustrative. For example, the division of the above modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0277] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0278] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for identifying abnormal weather based on deep learning, characterized in that: include: The backbone network based on the pre-trained weather recognition model extracts features from the image to be detected to obtain image features; The image to be detected includes meteorological information; fusing the image features based on the neck network of the weather recognition model to obtain target fusion features; Detecting the target fusion features based on the detection network of the weather recognition model to obtain a detection result of abnormal weather in the image to be detected; The neck network is provided with an EGMA module, and the EGMA module enhances the global correlation strength of the expression ability of weather features based on at least two attention mechanisms; The EGMA module includes a channel-aware modulation submodule, an ECA submodule, a cross-channel spatial attention submodule, and an ESA submodule. For the first input feature of the EGMA module: Performing a channel feature enhancement operation on the first input feature by the channel-aware modulation submodule to obtain a first enhanced feature; Performing a channel attention operation on the first enhanced feature through the ECA submodule to obtain a second enhanced feature; Performing a cross-channel spatial attention operation on the second enhanced feature through the cross-channel spatial attention submodule to obtain a multi-attention enhanced feature; The ESA submodule performs a spatial attention operation on the multiple attention enhancement features to obtain a first output feature corresponding to the first input feature; The channel-aware modulation submodule includes a dimension permutation layer, a double-layer multi-perceptron layer, an inverse dimension permutation layer, an activation function layer, and a first weighting layer; performing a channel feature enhancement operation on the first input feature by the channel-aware modulation submodule to obtain a first enhanced feature includes: Permuting the dimensional parameters of the first input feature through the dimensional permutation layer to obtain a flattened spatial feature; Performing a channel compression operation and a nonlinear activation operation on the flattened spatial features through the first multi-perceptron layer to obtain compressed features; Performing a channel expansion operation on the compressed feature through the second multi-perceptron layer to obtain a first expanded feature that matches the dimension of the flattened spatial feature; Performing inverse permutation on each dimensional parameter of the first extended feature through the inverse dimensional permutation layer to obtain a reconstructed feature whose dimensional parameters match the first input feature; Performing a linear activation operation on the reconstructed features through the activation function layer to obtain channel modulation weights; Performing a weighted operation on the first input feature based on the channel modulation weight by the first weighted layer to obtain the first enhanced feature; The cross-channel spatial attention submodule includes a channel shuffling layer, a CBR layer, a CBS layer and a second weighted layer; the cross-channel spatial attention submodule performs a cross-channel spatial attention operation on the second enhanced feature to obtain a multi-attention enhanced feature, including: performing a channel shuffling operation on the second enhanced feature through the channel shuffling layer to obtain a recombined feature; Performing channel compression operation, normalization operation and nonlinear transformation operation on the recombined features in sequence through the CBR layer to obtain compressed activation features; Performing channel expansion operation, normalization operation and linear activation operation on the compressed activation feature in sequence through the CBS layer to obtain a spatial weight; The second weighted layer performs a weighted operation on the reorganized features based on the spatial weights to obtain the multi-attention enhanced features.

2. The abnormal weather identification method according to claim 1, characterized in that: The backbone network is provided with a C2fPRFF module, which includes a first convolutional layer, a progressive fusion structure including at least two PRFFBlocks, a first splicing layer, and a second convolutional layer. For each second input feature input to the C2fPRFF module: Performing a channel compression operation on the second input feature through the first convolution layer to obtain a convolution feature; The progressive fusion structure sequentially performs progressive fusion operations on the convolution features through the series-connected PRFFBlocks to obtain hierarchical fusion features corresponding to each level; Performing a splicing operation on the convolutional features and the fusion features of each level through the first splicing layer to obtain a first splicing feature; A channel expansion operation is performed on the first concatenated feature through the second convolutional layer to obtain a second output feature corresponding to the second input feature; and the image feature is obtained based on the second output feature.

3. The abnormal weather identification method according to claim 2, characterized in that: The PRFFBlock includes a progressive convolution structure, a second convolution layer, a third convolution layer, an ESE layer, and a third convolution layer. The progressive convolution structure includes at least two convolution layers connected in series. For the third input feature of each PRFFBlock: Performing a progressive convolution operation on the third input feature through at least two convolutional layers connected in series in the progressive convolution structure to obtain hierarchical convolution features corresponding to each level; Performing a splicing operation on the third input feature and each of the level convolution features through the second splicing layer to obtain a second splicing feature; Performing a channel expansion operation on the second concatenated feature through the third convolutional layer to obtain a second expanded feature; Performing a global average pooling operation and a convolution operation on the second extended feature through the ESE layer, and weighting the second extended feature with the calculated channel attention weight to obtain an initial fused feature; A splicing operation is performed on the third input feature and the initial fusion feature through the third splicing layer to obtain a third output feature corresponding to the third input feature.

4. The abnormal weather identification method according to any one of claims 1 to 3, characterized in that: The weather recognition model is trained based on regression loss and classification loss; the regression loss includes focused fusion intersection-over-union loss, and the formula is as follows: Among them, the To focus on the fusion intersection-over-union loss, the IoU is the intersection-over-union ratio between the predicted box and the real box; the γ is a parameter that controls the degree of outlier suppression; the For the purpose of integration and comparison of losses; L IoU is the intersection and union loss; L dis is the center distance loss; L shp is the width and height loss; L ang is the angle loss; the Δ is 2 times the center distance loss; the b and stated Represent the center point of the real box and the center point of the predicted box respectively; Represents the square of the Euclidean distance between the center point of the real box and the predicted box respectively; and stated Respectively represent the width and height of the minimum circumscribed rectangular frame; the Ω is 2 times the width and height loss; the w and are the widths of the predicted border and the true border respectively; Represents the square of the Euclidean distance between the true box and the predicted box width; h and are the heights of the predicted border and the true border respectively; Represents the square of the Euclidean distance between the true box and the predicted box height; The angle loss is 2 times; and stated Represent the width and height of the rectangular box constructed by the center points of the true box and the predicted box, respectively.

5. An abnormal weather identification device, characterized in that: include: The extraction module is used to extract features from the image to be detected based on the backbone network of the pre-trained weather recognition model to obtain image features; The image to be detected includes meteorological information; A fusion module, configured to fuse the image features based on the neck network of the weather recognition model to obtain target fusion features; a detection module, configured to detect the target fusion features based on the detection network of the weather recognition model, and obtain a detection result of abnormal weather in the image to be detected; The neck network is provided with an EGMA module, and the EGMA module enhances the global correlation strength of the expression ability of weather features based on at least two attention mechanisms; The EGMA module includes a channel-aware modulation submodule, an ECA submodule, a cross-channel spatial attention submodule, and an ESA submodule. The fusion module includes a fusion unit, which is used to: For the first input feature of the EGMA module: Performing a channel feature enhancement operation on the first input feature by the channel-aware modulation submodule to obtain a first enhanced feature; Performing a channel attention operation on the first enhanced feature through the ECA submodule to obtain a second enhanced feature; Performing a cross-channel spatial attention operation on the second enhanced feature through the cross-channel spatial attention submodule to obtain a multi-attention enhanced feature; The ESA submodule performs a spatial attention operation on the multiple attention enhancement features to obtain a first output feature corresponding to the first input feature; The channel-aware modulation submodule includes a dimension permutation layer, a double-layer multi-sensor layer, an inverse dimension permutation layer, an activation function layer, and a first weighting layer; the fusion unit is specifically configured to: Permuting the dimensional parameters of the first input feature through the dimensional permutation layer to obtain a flattened spatial feature; Performing a channel compression operation and a nonlinear activation operation on the flattened spatial features through the first multi-perceptron layer to obtain compressed features; Performing a channel expansion operation on the compressed feature through the second multi-perceptron layer to obtain a first expanded feature that matches the dimension of the flattened spatial feature; Performing inverse permutation on each dimensional parameter of the first extended feature through the inverse dimensional permutation layer to obtain a reconstructed feature whose dimensional parameters match the first input feature; Performing a linear activation operation on the reconstructed features through the activation function layer to obtain channel modulation weights; Performing a weighted operation on the first input feature based on the channel modulation weight by the first weighted layer to obtain the first enhanced feature; The cross-channel spatial attention submodule includes a channel shuffling layer, a CBR layer, a CBS layer and a second weighted layer; the fusion unit is specifically used to: performing a channel shuffling operation on the second enhanced feature through the channel shuffling layer to obtain a recombined feature; Performing channel compression operation, normalization operation and nonlinear transformation operation on the recombined features in sequence through the CBR layer to obtain compressed activation features; Performing channel expansion operation, normalization operation and linear activation operation on the compressed activation feature in sequence through the CBS layer to obtain a spatial weight; The second weighted layer performs a weighted operation on the reorganized features based on the spatial weights to obtain the multi-attention enhanced features.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the abnormal weather identification method based on deep learning as described in any one of claims 1 to 4 is implemented.

7. A computer program product, wherein the computer program product stores a computer program, characterized in that: When the computer program is executed by a processor, the abnormal weather identification method based on deep learning as described in any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Marine weather prediction method and system based on neural network

    CN118962852A

  • Remote sensing image change detection method and device based on iterative Mama architecture

    CN119205638A