A method and system for detecting ice accretion, an electronic device and a medium

By designing the encoding layer and multi-head attention layer of the ice jam detection model, the problems of resolution and accuracy in ice jam detection are solved, and more efficient ice jam detection is achieved.

CN116977778BActive Publication Date: 2025-11-25NANHU LAB
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310738639.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-20
Publication Date
2025-11-25
Estimated Expiration
2043-06-20

AI Technical Summary

Technical Problem

Existing technologies for ice jam detection suffer from low spatial resolution, limited temporal resolution, and low accuracy in ice jam edge detection, making it difficult to achieve intelligent detection of ice jams.

Method used

An ice slab detection model is adopted, which includes an encoding layer, a multi-head attention layer, and a decoding layer. The encoding layer improves feature extraction capability through depth-parameterized convolution and inverted residual structure. The multi-head attention layer takes into account both channel features and spatial features and is used for remote information transmission to improve the accuracy of ice slab edge detection.

Benefits of technology

It improves the accuracy of ice edge detection and enhances the accuracy of ice jam detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116977778B_ABST
    Figure CN116977778B_ABST
Patent Text Reader

Abstract

The application discloses an ice slurry detection method and system, electronic equipment and medium, and relates to the technical field of remote sensing semantic segmentation. The method comprises the following steps: inputting a to-be-predicted river remote sensing image into a trained ice slurry detection model to obtain ice slurry in the to-be-predicted river remote sensing image; the trained ice slurry detection model is trained by taking a sample river remote sensing image as input and taking ice slurry in the sample river remote sensing image as output; the ice slurry detection model comprises an encoding layer, a multi-head attention layer and a decoding layer; the output end of the encoding layer is connected with the input end of the multi-head attention layer and the input end of the decoding layer, and the output end of the multi-head attention layer is connected with the input end of the decoding layer; the encoding layer comprises a plurality of feature encoding modules; the feature encoding module is an inverted residual structure and comprises a deep parameterized convolution layer; the multi-head attention layer comprises a spatial attention feature mechanism and a channel attention feature mechanism. The application can improve the accuracy of ice slurry edge detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing semantic segmentation technology, and in particular to an ice shard detection method, system, electronic device, and medium. Background Technology

[0002] During the early winter freezing and early spring thawing periods, due to the warmer temperatures at higher latitudes in low-latitude regions, rivers freeze earlier and thaw later as they move from low to high latitudes. When the glacial water from upstream travels downstream, it easily forms ice dams in narrower sections and at bends, blocking some of the upstream water, increasing the river's storage capacity, and raising upstream water levels—a phenomenon known as ice jam flooding. When the ice melts and the river thaws, this stored water is rapidly released and carried downstream, forming ice floods that pose a significant threat to people's lives and property. Therefore, there is a need to find a method for intelligent detection of ice jam floods.

[0003] Traditional monitoring methods rely primarily on on-site personnel patrols, which struggle to provide a comprehensive, large-scale spatial overview, resulting in lagging supervision. To address these shortcomings, patent application CN105913023A proposes a detection method utilizing multispectral and SAR images. This method iteratively performs superpixel segmentation using a combination of NDSI detection and clustering to obtain optimal ice jam detection results, enabling intelligent detection of ice jams. However, the satellite remote sensing method employed is limited by temporal resolution, hindering real-time monitoring of target areas, and its low spatial resolution also fails to meet the detection requirements for small ice floes. Patent application CN111160311A proposes a semantic segmentation method using a multi-attention mechanism and dual-stream fusion. This method extracts object information using positional attention and channel attention separately, then fuses the results to output the detection findings. While this method improves detection performance to some extent, the application of the two attention mechanisms on different branch features makes its effectiveness highly dependent on feature encoding capabilities. Furthermore, the lack of a targeted feature extraction encoder results in low accuracy for ice jam edge detection, consequently leading to inaccurate ice jam detection results. Summary of the Invention

[0004] The purpose of this invention is to provide an ice shard detection method, system, electronic device, and medium that can improve the accuracy of ice shard edge detection.

[0005] To achieve the above objectives, the present invention provides the following solution:

[0006] A method for detecting ice crystals, comprising:

[0007] Acquire remote sensing images of the river to be predicted;

[0008] The ice crystals in the remote sensing image of the river to be predicted are obtained by inputting the image of the river to be predicted into the trained ice crystal detection model. The trained ice crystal detection model is obtained by training the sample remote sensing image of the river as input and the ice crystals in the sample remote sensing image as output. The ice crystal detection model includes an encoding layer, a multi-head attention layer, and a decoding layer. The output of the encoding layer is connected to the input of the multi-head attention layer and the input of the decoding layer, and the output of the multi-head attention layer is connected to the input of the decoding layer. The encoding layer includes multiple feature encoding modules. The feature encoding modules are inverted residual structures and include depth parameterized convolutional layers. The multi-head attention layer includes spatial attention feature mechanism and channel attention feature mechanism.

[0009] Optionally, the coding layer includes: a first Conv2d convolutional layer, a first feature coding convolutional layer, a second feature coding convolutional layer, a third feature coding convolutional layer, a fourth feature coding convolutional layer, and a fifth feature coding convolutional layer connected in sequence;

[0010] The first feature encoding convolutional layer includes a first encoding input layer, a first feature encoding layer, and a second Conv2d convolutional layer connected in sequence; the second feature encoding convolutional layer includes a second encoding input layer, a second feature encoding layer, and a third Conv2d convolutional layer connected in sequence; the third feature encoding convolutional layer includes a third encoding input layer, a third feature encoding layer, and a fourth Conv2d convolutional layer connected in sequence; the fourth feature encoding convolutional layer includes a fourth encoding input layer, a fourth feature encoding layer, and a fifth Conv2d convolutional layer connected in sequence; the fifth feature encoding convolutional layer includes a fifth encoding input layer, a fifth feature encoding layer, and a sixth Conv2d convolutional layer connected in sequence.

[0011] The first feature encoding layer, the second feature encoding layer, and the fourth feature encoding layer each include three sequentially connected feature encoding modules; the third feature encoding layer includes four sequentially connected feature encoding modules; the fifth feature encoding layer includes one feature encoding module; the feature encoding module includes a first stacking operation layer and a first Conv convolutional layer, a first Do-Conv layer, a second Conv convolutional layer, and a first activation layer connected in sequence; the input of the first Conv convolutional layer is connected to the input of the first activation layer through the first stacking operation layer;

[0012] The outputs of the first, second, third, fourth, and fifth Conv2d convolutional layers are all connected to the input of the multi-head attention layer, and the output of the sixth Conv2d convolutional layer is connected to the decoding layer.

[0013] Optionally, the multi-head attention layer includes five multi-head attention modules; each of the five multi-head attention modules includes an input feature vector branch, a second overlay operation layer, a channel attention branch, and a spatial attention branch; the input end of the input feature vector branch is connected to the first input end of the channel attention branch and the first input end of the spatial attention branch, respectively, serving as the input end of the multi-head attention module; the first output end of the input feature vector branch is connected to the second input end of the channel attention branch and the second input end of the spatial attention branch, respectively; the second output end of the input feature vector branch, the output end of the channel attention branch, and the output end of the spatial attention branch are all connected to the input end of the second overlay operation layer; the output end of the second overlay operation layer is the output layer of the multi-head attention module.

[0014] The input of the first multi-head attention module is connected to the output of the first Conv2d convolutional layer; the input of the second multi-head attention module is connected to the output of the second Conv2d convolutional layer; the input of the third multi-head attention module is connected to the output of the third Conv2d convolutional layer; the input of the fourth multi-head attention module is connected to the output of the fourth Conv2d convolutional layer; and the input of the fifth multi-head attention module is connected to the output of the fifth Conv2d convolutional layer. The output layers of the first, second, third, fourth, and fifth multi-head attention modules are all connected to the input of the decoding layer.

[0015] Optionally, the decoding layer includes a first concatenated feature encoding deconvolution layer, a second concatenated feature encoding deconvolution layer, a third concatenated feature encoding deconvolution layer, a fourth concatenated feature encoding deconvolution layer, a fifth concatenated feature encoding deconvolution layer, and a second activation layer;

[0016] The first concatenated feature encoding deconvolution layer comprises a first decoding input layer, a first concatenation layer, a sixth feature encoding layer, and a first ConvT2d convolution layer connected in sequence; the second concatenated feature encoding deconvolution layer comprises a second decoding input layer, a second concatenation layer, a seventh feature encoding layer, and a second ConvT2d convolution layer connected in sequence; the third concatenated feature encoding deconvolution layer comprises a third decoding input layer, a third concatenation layer, an eighth feature encoding layer, and a third ConvT2d convolution layer connected in sequence; the fourth concatenated feature encoding deconvolution layer comprises a fourth decoding input layer, a fourth concatenation layer, a ninth feature encoding layer, and a fourth ConvT2d convolution layer connected in sequence; and the fifth concatenated feature encoding deconvolution layer comprises a fifth decoding input layer, a fifth concatenation layer, a tenth feature encoding layer, and a fifth ConvT2d convolution layer connected in sequence.

[0017] The sixth, seventh, eighth, and ninth feature coding layers each include three sequentially connected feature coding modules; the tenth feature coding layer includes two sequentially connected feature coding modules.

[0018] The input of the first decoding input layer is connected to the output of the multi-head attention module output layer of the fifth multi-head attention module; the input of the sixth feature encoding layer is connected to the output of the sixth Conv2d convolutional layer; the input of the second decoding input layer is connected to the output of the fourth multi-head attention module; the input of the seventh feature encoding layer is connected to the output of the first Conv2d convolutional layer; the input of the third decoding input layer is connected to the output of the third multi-head attention module; the input of the eighth feature encoding layer is connected to the output of the second Conv2d convolutional layer; the input of the fourth decoding input layer is connected to the output of the second multi-head attention module; the input of the ninth feature encoding layer is connected to the output of the third Conv2d convolutional layer; the input of the fifth decoding input layer is connected to the output of the first multi-head attention module; the input of the tenth feature encoding layer is connected to the output of the fourth Conv2d convolutional layer; and the input of the second activation layer is connected to the output of the fifth Conv2d convolutional layer.

[0019] Optionally, the spatial attention branch includes: a fourth Conv convolutional layer, a global average pooling layer, a max pooling layer, a fifth Conv convolutional layer, a third activation layer, and a first convolutional operation layer connected in sequence; the input of the global average pooling layer is connected to the input of the input feature vector branch and the input of the channel attention branch, respectively; the input of the first convolutional operation layer is connected to the first output of the input feature vector branch; and the output of the first convolutional operation layer is connected to the input of the second stacking operation layer.

[0020] Optionally, the channel attention branch includes: a sixth Conv convolutional layer, a first reshape layer, a second reshape layer, a seventh Conv convolutional layer, a product operation layer, and a second convolutional operation layer; the input of the sixth Conv convolutional layer is connected to the input of the global average pooling layer and the input of the input feature vector branch; the output of the sixth Conv convolutional layer is connected to the inputs of the first reshape layer and the second reshape layer, respectively; the output of the first reshape layer is connected to the input of the seventh Conv convolutional layer; the outputs of the seventh Conv convolutional layer and the second reshape layer are respectively connected to the input of the product operation layer, the output of the product operation layer is connected to the first input of the second convolutional operation layer, the second input of the second convolutional operation layer is connected to the first output of the input feature vector branch; and the output of the convolutional operation layer is connected to the input of the second stacking operation layer.

[0021] Optionally, the process of determining the trained ice crystal detection model includes:

[0022] Construct an ice crystal detection model;

[0023] Images of sample rivers flowing from low latitudes to high latitudes were captured during the early stages of freezing in early winter and the early stages of ice thawing in early spring.

[0024] Sample river remote sensing images are obtained from the river section images; the sample river remote sensing images are all images including ice floes in the river section images;

[0025] The ice floes in the sample river remote sensing image are obtained from the sample river remote sensing image;

[0026] The ice detection model is trained by using the sample river remote sensing image as input and the ice floes in the sample river remote sensing image as output.

[0027] An ice detection system, comprising:

[0028] The acquisition module is used to acquire remote sensing images of the river to be predicted.

[0029] An ice jam detection module is used to input the remote sensing image of the river to be predicted into a trained ice jam detection model to obtain ice jams in the remote sensing image of the river to be predicted. The trained ice jam detection model is obtained by training the ice jam detection model with the sample river remote sensing image as input and the ice jams in the sample river remote sensing image as output. The ice jam detection model includes an encoding layer, a multi-head attention layer, and a decoding layer. The output of the encoding layer is connected to the input of the multi-head attention layer and the input of the decoding layer, and the output of the multi-head attention layer is connected to the input of the decoding layer. The encoding layer includes multiple feature encoding modules. The feature encoding modules are inverted residual structures and include depth parameterized convolutional layers. The multi-head attention layer includes spatial attention feature mechanisms and channel attention feature mechanisms.

[0030] An electronic device, comprising:

[0031] A memory and a processor, the memory being used to store a computer program, the processor running the computer program to cause the electronic device to perform the ice detection method according to the above description.

[0032] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the ice detection method as described above.

[0033] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects: The ice jam detection model of the present invention includes an encoding layer, a multi-head attention layer, and a decoding layer. The encoding layer includes a feature encoding module, which is composed of a combination of depth parameterized convolution and inverted residual structure, which can effectively improve the feature extraction capability. The multi-head attention layer of the present invention takes into account both channel features and spatial features for remote information transmission, improves the ability of extracted features to represent the difference between the target and the background, and can improve the accuracy of ice jam edge detection, thereby making the detection results of ice jams more accurate. Attached Figure Description

[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0035] Figure 1 This is a flowchart of the semantic segmentation-based ice floe detection method for the Yellow River according to the present invention;

[0036] Figure 2 This is a structural diagram of the ice crystal detection model of the present invention;

[0037] Figure 3 This is a structural diagram of the feature coding convolutional layer of the present invention, which includes three feature coding modules;

[0038] Figure 4 This is a diagram of the depth hyperparameterized convolution structure of the present invention;

[0039] Figure 5 This is a structural diagram of the multi-head attention module of the present invention;

[0040] Figure 6 This is a structural diagram of the concatenated feature encoding deconvolution layer of the present invention, which includes three feature encoding modules;

[0041] Figure 7 This is a comparison chart showing the results of processing region 1 using the ice detection model of this invention and existing models;

[0042] Figure 8 This is a comparison chart showing the results of processing Region 2 using the ice detection model of this invention and existing models;

[0043] Figure 9 This is a comparison chart showing the results of processing region 3 using the ice detection model of the present invention and existing models. Detailed Implementation

[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0045] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0046] This invention provides a method for detecting ice crystals, comprising:

[0047] Obtain remote sensing images of the river to be predicted.

[0048] The ice crystals in the remote sensing image of the river to be predicted are obtained by inputting the image of the river to be predicted into the trained ice crystal detection model. The trained ice crystal detection model is obtained by training the sample remote sensing image of the river as input and the ice crystals in the sample remote sensing image as output. The ice crystal detection model includes an encoding layer, a multi-head attention layer, and a decoding layer. The output of the encoding layer is connected to the input of the multi-head attention layer and the input of the decoding layer, and the output of the multi-head attention layer is connected to the input of the decoding layer. The encoding layer includes multiple feature encoding modules. The feature encoding modules are inverted residual structures and include depth parameterized convolutional layers. The multi-head attention layer includes spatial attention feature mechanism and channel attention feature mechanism.

[0049] In practical applications, such as Figure 2 As shown, the encoding layer includes: a first Conv2d convolutional layer, a first feature encoding convolutional layer, a second feature encoding convolutional layer, a third feature encoding convolutional layer, a fourth feature encoding convolutional layer, and a fifth feature encoding convolutional layer connected in sequence.

[0050] like Figure 3 As shown, the first feature encoding convolutional layer includes a first encoding input layer, a first feature encoding layer, and a second Conv2d convolutional layer connected in sequence; the second feature encoding convolutional layer includes a second encoding input layer, a second feature encoding layer, and a third Conv2d convolutional layer connected in sequence; the third feature encoding convolutional layer includes a third encoding input layer, a third feature encoding layer, and a fourth Conv2d convolutional layer connected in sequence; the fourth feature encoding convolutional layer includes a fourth encoding input layer, a fourth feature encoding layer, and a fifth Conv2d convolutional layer connected in sequence; and the fifth feature encoding convolutional layer includes a fifth encoding input layer, a fifth feature encoding layer, and a sixth Conv2d convolutional layer connected in sequence.

[0051] The first feature coding layer, the second feature coding layer, and the fourth feature coding layer each include three sequentially connected feature coding modules; the third feature coding layer includes four sequentially connected feature coding modules; and the fifth feature coding layer includes one feature coding module. Figure 3 As shown, the feature encoding module includes a first stacking operation layer and a first Conv convolutional layer, a first Do-Conv layer, a second Conv convolutional layer, and a first activation layer connected in sequence; the input of the first Conv convolutional layer is connected to the input of the first activation layer through the first stacking operation layer.

[0052] The outputs of the first, second, third, fourth, and fifth Conv2d convolutional layers are all connected to the input of the multi-head attention layer, and the output of the sixth Conv2d convolutional layer is connected to the decoding layer.

[0053] In practical applications, such as Figure 2 As shown, the multi-head attention layer includes five multi-head attention modules; as Figure 5 As shown, each of the five multi-head attention modules includes an input feature vector branch, a second overlay operation layer, a channel attention branch, and a spatial attention branch. The input end of the input feature vector branch is connected to the first input end of the channel attention branch and the first input end of the spatial attention branch, respectively, serving as the input end of the multi-head attention module. The first output end of the input feature vector branch is connected to the second input end of the channel attention branch and the second input end of the spatial attention branch, respectively. The second output end of the input feature vector branch, the output end of the channel attention branch, and the output end of the spatial attention branch are all connected to the input end of the second overlay operation layer. The output end of the second overlay operation layer is the output layer of the multi-head attention module.

[0054] The input of the first multi-head attention module is connected to the output of the first Conv2d convolutional layer; the input of the second multi-head attention module is connected to the output of the second Conv2d convolutional layer; the input of the third multi-head attention module is connected to the output of the third Conv2d convolutional layer; the input of the fourth multi-head attention module is connected to the output of the fourth Conv2d convolutional layer; and the input of the fifth multi-head attention module is connected to the output of the fifth Conv2d convolutional layer. The output layers of the first, second, third, fourth, and fifth multi-head attention modules are all connected to the input of the decoding layer.

[0055] In practical applications, such as Figure 2 As shown, the decoding layer includes a first concatenated feature encoding deconvolution layer, a second concatenated feature encoding deconvolution layer, a third concatenated feature encoding deconvolution layer, a fourth concatenated feature encoding deconvolution layer, a fifth concatenated feature encoding deconvolution layer, and a second activation layer.

[0056] The first concatenated feature encoding deconvolution layer includes a first decoding input layer, a first concatenation layer, a sixth feature encoding layer, and a first ConvT2d convolution layer connected in sequence; the second concatenated feature encoding deconvolution layer includes a second decoding input layer, a second concatenation layer, a seventh feature encoding layer, and a second ConvT2d convolution layer connected in sequence; the third concatenated feature encoding deconvolution layer includes a third decoding input layer, a third concatenation layer, an eighth feature encoding layer, and a third ConvT2d convolution layer connected in sequence; the fourth concatenated feature encoding deconvolution layer includes a fourth decoding input layer, a fourth concatenation layer, a ninth feature encoding layer, and a fourth ConvT2d convolution layer connected in sequence; and the fifth concatenated feature encoding deconvolution layer includes a fifth decoding input layer, a fifth concatenation layer, a tenth feature encoding layer, and a fifth ConvT2d convolution layer connected in sequence.

[0057] like Figure 6 As shown, the sixth feature coding layer, the seventh feature coding layer, the eighth feature coding layer and the ninth feature coding layer each include three feature coding modules connected in sequence; the tenth feature coding layer includes two feature coding modules connected in sequence.

[0058] The input of the first decoding input layer is connected to the output of the multi-head attention module output layer of the fifth multi-head attention module; the input of the sixth feature encoding layer is connected to the output of the sixth Conv2d convolutional layer; the input of the second decoding input layer is connected to the output of the fourth multi-head attention module; the input of the seventh feature encoding layer is connected to the output of the first Conv2d convolutional layer; the input of the third decoding input layer is connected to the output of the third multi-head attention module; the input of the eighth feature encoding layer is connected to the output of the second Conv2d convolutional layer; the input of the fourth decoding input layer is connected to the output of the second multi-head attention module; the input of the ninth feature encoding layer is connected to the output of the third Conv2d convolutional layer; the input of the fifth decoding input layer is connected to the output of the first multi-head attention module; the input of the tenth feature encoding layer is connected to the output of the fourth Conv2d convolutional layer; and the input of the second activation layer is connected to the output of the fifth Conv2d convolutional layer.

[0059] In practical applications, the structure of the first Do-Conv layer is as follows: Figure 4 As shown.

[0060] In practical applications, such as Figure 5As shown, the spatial attention branch includes: a fourth Conv convolutional layer, a global average pooling layer, a max pooling layer, a fifth Conv convolutional layer, a third activation layer, and a first convolutional operation layer connected in sequence; the input of the global average pooling layer is connected to the input of the input feature vector branch and the input of the channel attention branch, respectively; the input of the first convolutional operation layer is connected to the first output of the input feature vector branch; and the output of the first convolutional operation layer is connected to the input of the second stacking operation layer.

[0061] In practical applications, such as Figure 5 As shown, the channel attention branch includes: a sixth Conv convolutional layer, a first reshape layer, a second reshape layer, a seventh Conv convolutional layer, a product operation layer, and a second convolutional operation layer; the input of the sixth Conv convolutional layer is connected to the input of the global average pooling layer and the input of the input feature vector branch; the output of the sixth Conv convolutional layer is connected to the inputs of the first reshape layer and the second reshape layer, respectively; the output of the first reshape layer is connected to the input of the seventh Conv convolutional layer; the outputs of the seventh Conv convolutional layer and the second reshape layer are respectively connected to the input of the product operation layer, the output of the product operation layer is connected to the first input of the second convolutional operation layer, the second input of the second convolutional operation layer is connected to the first output of the input feature vector branch; and the output of the convolutional operation layer is connected to the input of the second stacking operation layer.

[0062] In practical applications, the process of determining the trained ice crystal detection model includes:

[0063] Construct an ice crystal detection model.

[0064] Images of sample rivers flowing from low latitudes to high latitudes were captured during the early stages of freezing in early winter and the early stages of ice thawing in early spring.

[0065] Sample river remote sensing images are obtained from the river section images; the sample river remote sensing images are all images including ice floes in the river section images.

[0066] The ice floes in the sample river remote sensing image are obtained from the sample river remote sensing image.

[0067] The ice detection model is trained by using the sample river remote sensing image as input and the ice floes in the sample river remote sensing image as output.

[0068] The advantages of this invention are: (1) A new feature extraction block (FEB) is designed, which is composed of a combination of deep parameterized convolution and inverted residual structure, which can effectively improve the feature extraction capability; (2) A multi-head attention block (MAB) is designed to take into account both channel features and spatial features for remote information transmission, thereby improving the ability of extracted features to represent the differences between the target and the background.

[0069] This invention also provides an embodiment applying the above method to the Yellow River, and the method will be described in more detail, such as... Figure 1 The following are included:

[0070] Step 101: Collect images of the Yellow River flowing from low latitudes to high latitudes during early winter and early spring using drones. Images of the Yellow River flowing from low latitudes to high latitudes should be captured during the initial stages of freezing in early winter and the initial stages of ice thawing in early spring.

[0071] Step 102: Perform preprocessing on the UAV images, including stitching, cropping, radiometric calibration, and geometric correction. Preprocess the images by stitching, cropping, radiometric calibration, and geometric correction, and then select regions of interest containing ice spikes.

[0072] Step 103: Label the ice crystals and background based on expert knowledge to create a deep learning training sample set. Manually label the ice crystal targets and background based on expert knowledge to create a deep learning training sample set.

[0073] Step 104: Design an ice crystal detection model based on semantic segmentation.

[0074] like Figure 2 As shown, the Ice Detection Model (IDM) mainly consists of three parts: feature encoding, feature propagation, and feature decoding.

[0075] In the feature encoding part, a feature encoding module is constructed using deep hyperparameterized convolution and inverted residual structure to extract high-dimensional feature representation information of the target object; in the feature transmission part, features of different resolutions are transmitted remotely through a designed multi-head attention block (MAB); in the feature decoding part, remote information and high-dimensional information are superimposed, and the resolution is restored to be consistent with the original input layer by layer through a combination of deconvolution, Concat and FEB, and finally the pixel-by-pixel ice crystal recognition probability is output.

[0076] Step 1: Combine deep hyperparameterized convolution and inverted residual structure to construct a feature encoding module to achieve high-dimensional feature representation of the target object.

[0077] The specific process is as follows: network structure parameters are shown in Table 1, from the input graph to d5. InputSize is the scale of the network input feature vector map: C×M. 2 Where C is the number of channels, M is the length and width of the feature map; Operator is the feature extraction network structure at each resolution; h is the inverted residual channel expansion factor; n is the number of repetitions of the FEB module; c is the number of output channels of the Conv2d convolution; s is the convolution stride, which shrinks the feature map during the encoding stage to extract high-dimensional features.

[0078] Table 1 Model Parameter Table

[0079]

[0080] Specifically, the FEB structure is as follows: Figure 3 As shown within the dashed box, the input feature vector's channel count is first expanded to h times its original size using a 1×1 convolution (first Conv convolutional layer). Then, feature learning is performed using a depthwise hyperparameterized convolution (first Do-Conv layer). Finally, a 1×1 convolution (second Conv convolutional layer) restores the feature channel count to match the input. The first activation layer in this module uses the Gelu function to enhance the model's non-linearity. This module utilizes a first stacking operation layer. By superimposing the feature values, the network learns the residual between the input feature vector and the true feature.

[0081] With 16×160 2 32×80 2 Taking the stage operation (e1 to e2) as an example, the input feature vector map has a length and width of M=160 and a number of channels C=16. This stage operation consists of n=3 FEB operations and one Conv2d convolution operation in series. The Conv2d convolution stride s=2 and the number of output channels c=32. When operating inside the FEB structure, the number of feature channels is expanded to be 3 times the input value h.

[0082] In detail, FEB is used to explicitly model information, learning the difference between the input feature vector and the true value: first, the input scale is 16×160. 2 The feature map is expanded by 1×1 convolution to have 3 times the number of channels as h = 3 times the input value, resulting in a scale of 48×160. 2 The feature map is then processed by a Do-Conv operation without changing its scale. Next, a 1×1 convolution is performed to reduce the number of channels by a factor of 1 / h, maintaining consistency with the input value. This channel number transformation within the structure ensures both the number of features and keeps the overall model's channel count low, effectively reducing the overall model parameter count and memory usage. The 1×1 convolution result is then activated using the Gelu activation function to improve the model's non-linear expressive power. Finally, the result is summed with the input feature vector. This residual mechanism allows the model to learn the difference between the input feature vector and the true value within the FEB structure.

[0083] The Conv2d operation in this stage is used to complete the stage feature output. The FEB operation does not change the feature scale; after three operations, it remains 16×160. 2 Conv2d convolution uses a stride of s=2, scaling the feature map's width and height to half the size of the input feature vector. It also outputs 32 channels (c=32), expanding the feature map's channels to 32, resulting in an output feature map of 32×80 pixels. 2 This feature map is called the stage modeling feature map of the feature encoding process, serving as a basis for the next stage of 32×80. 2 On the one hand, it serves as the input feature vector for scale, and on the other hand, it serves as the input feature vector for remote information transmission.

[0084] The first Do-Conv layer in this invention, such as Figure 4 As shown, D is the depthwise separable convolution kernel, W is the regular convolution kernel, P is the input feature vector, M and N are the length and width of the input feature vector, Dmul is the product of the two, and in this embodiment, M = N. Cin and Cout are the number of channels of the input feature vector and the output feature vector, respectively. In this invention, Cin = Cout. The first Do-Conv layer first uses the depthwise convolution operator ° to calculate the depthwise convolution D°P on the input feature vector, and then obtains the output out of the structure by multiplying the output result (D°P) of the regular convolution operator w. The combination of depthwise convolution and regular convolution increases the learnable parameters of the model, enabling the convolutional layer to learn channel information on the basis of learning neighborhood features. This hyperparameterization increases the computational cost slightly and does not lead to an increase in the computational complexity of inference, but can effectively improve the model performance.

[0085] Step 2: Based on the features obtained above at different resolutions, remote information transmission is performed through the designed multi-head attention module. In Table 1, column f in step represents the encoded features e with the same number, which are obtained after optimization by the attention module. This process only optimizes the features to achieve remote information transmission and skip connections without changing the scale.

[0086] The multi-head attention module network structure designed in this invention is as follows: Figure 5 As shown, the input feature vector for feature propagation is the feature output from the feature encoding process. The dashed boxes containing the channel weight coefficients and spatial weight coefficients represent the channel attention branch and the spatial attention branch, respectively. Not all information in the input feature vector (input through the input feature vector branch) is helpful for ice crystal extraction. The core role of the attention mechanism is to highlight useful information and suppress useless information. Specifically, the channel attention weight (CW) is used to select channel information that is highly relevant to the detection task and integrate multi-channel data; the spatial attention weight (GW) is used to characterize the correlation of the input feature vector in the spatial neighborhood, automatically capturing features of important pixel regions.

[0087] Specifically, the channel attention branch completes explicit modeling and feature interaction between different channels. The input feature vector (matrix shape H×W×C) is compressed into a 1×H×W compressed vector by 1×1 convolution through the sixth Conv convolutional layer. The compressed vector and the input feature vector are then reshaped by the first and second reshape layers to obtain vectors of dimensions HW×1×1 and HW×C, respectively. The vector HW×1×1 is then compressed into C×1×1 by convolution through the seventh Conv convolutional layer. The channel feature weights CW (Equation 1) are obtained by multiplying the vector HW×C by the convolutional layer. Finally, the channel enhancement features are obtained by multiplying the vector HW×C by the second convolutional layer.

[0088] CW=F conv {F reshape [F conv (Input)]}×F reshape (Input)(Equation 1), where,

[0089] F reshape (Input) represents the reshape operation performed on the input feature vector Input, which is the input feature vector obtained by processing the input feature vector through the second reshape layer. F conv (Input) represents the convolution operation performed on the input feature vector Input, which is the input feature vector passed through the input feature vector branch. In other words, it's the compressed vector obtained by processing the input feature vector through the sixth Conv convolutional layer. F reshape [F conv [Input] indicates that F conv (Input) performs a reshape operation, that is, the compressed vector is processed by the first reshape layer, F conv {F reshape [F conv ( / nput)]} indicates that F reshape [F conv (Input)] performs a convolution operation, i.e., F reshape [F conv [Input] The result obtained by processing through the seventh Conv convolutional layer.

[0090] Specifically, the spatial attention branch performs global average pooling and max pooling operations on the input feature vector to obtain the channel description of the feature. It then performs a 3×3 convolutional layer and Gelu activation function through the fifth Conv convolutional layer to obtain the spatial weight coefficient GW (Equation 2). Finally, it multiplies the weight coefficient with the input feature vector through the first convolutional operation layer to obtain the spatial attention feature.

[0091] GW = Fconv {F pool [F conv (Input)]}(Equation 2), where F conv (Input) represents the convolution operation performed on the input feature vector Input, which is the input feature vector passed through the input feature vector branch. That is, the result obtained by processing the input feature vector through the fourth Conv convolutional layer, F. pool [F conv [Input] indicates that F conv (Input) performs global average pooling and max pooling operations, F conv {F pool [F conv [Input]} represents the input to F. pool [F conv (Input)] performs a convolution operation, i.e., F pool [F conv [Input] The result obtained by processing through the fifth Conv convolutional layer.

[0092] The values ​​of spatial attention features, channel attention features, and input feature vector are summed through the second superposition operation layer (Equation 3) to obtain the remote transfer feature Output. By comprehensively using spatial dimension and channel dimension attention, reinforcement learning can be achieved on different dimensions of extracted features to obtain better feature representation.

[0093] Output=Input+Input×CW+Input×GW (Formula 3).

[0094] Step 3: Process the deepest encoded features obtained in Step 1 (scale 256×10). 2 Decoding is performed to recover the resolution layer by layer from the high-dimensional features. The operations at each scale stage of the decoding process include channel concat, FEB, and deconvolution operations. The specific network parameters for each layer are shown in the d5+f5 to probability section of Table 1. h is the inverted residual channel number expansion factor, n is the number of FEB module repetitions, c is the number of ConvT2d deconvolution output channels, and s is the scale expansion factor of the deconvolution process. At the end of the network layer, it is transformed into the output category through two sigmoid functions.

[0095] Stage operations at a single scale, such as Figure 6 As shown, the decoding features from the high level and the remotely transmitted features of the same scale obtained in step two are fused to realize remote information transmission between the encoding and decoding processes. Then, the model's learning ability is enhanced by the FEB module, and deconvolution is used to upsample the feature map and restore the resolution of the feature map.

[0096] Based on the network parameters in Table 1, with a resolution of 64×40 2 We get 32×802 Taking the stage operation (d3+f3 to d2) as an example, the input feature vector map has a length and width M=40 and a number of channels C=64. This stage operation consists of a Concat channel superposition operation, n=3 FEB operations and a ConvT2d deconvolution operation in series.

[0097] In detail, the Concat channel concatenation operation has two feature inputs. The input to the seventh feature encoding layer is the high-level feature decoding feature d3, and the long pass information f3 is the long pass information from the encoded features of the same scale after passing through the multi-head attention module. Both sets of features (scale 64×40) 2 The images were overlaid on the channel to obtain a scale of 128×40. 2 The features are fused, and after three FEB operations, only the features are further optimized and fitted to the output without changing the scale. The FEB operation parameter h changes here, but the structure is the same as before and will not be described in detail. ConvT2d is a deconvolution operation, and s=2 determines that the feature map window width is expanded to twice the input, thus condensing the information from 128 channels into c=32 output channels. The final output scale is 32×80. 2 The feature map is then used to proceed to the next level of computation.

[0098] Finally, the model performs two sigmoid operations through the second activation layer (Equation 4), where t represents the input. The feature encoding output is an unnormalized predicted probability value. The first sigmoid operation normalizes the probability value distribution range from (-∞, +∞) to [0,1], where t represents the probability value. The second sigmoid operation converts the probability value into an output detection result type of 0 or 1, where t represents the normalized probability value. The result of the operation is 0 for background and 1 for icicles.

[0099]

[0100] Step 105: Train the ice floe detection model based on labeled samples and obtain the model weights. This embodiment uses the Yellow River ice floe UAV remote sensing image dataset, which contains a total of 21,293 image pairs. The image size is 320×320 pixels. After random allocation, the training set, test set, and validation set contain 15,966, 2,661, and 2,661 image pairs, respectively.

[0101] All networks trained in this application were implemented on the Python-based deep learning framework PyTorch 1.11, using an NVIDIA A40 device. To ensure fairness in the method comparison, all network model hyperparameters were uniformly set as follows: training epochs of 100, learning rate of 0.01, batch size of 16 training samples, and loss function BCE loss.

[0102] Step 106: Input the target image into the ice crystal detection model to identify ice crystals in the target image. The target image is acquired by a drone and is a two-dimensional visible light color image.

[0103] This invention employs a novel feature encoding module during the encoding process, which enhances feature extraction capabilities through a combination of deep parameterized convolution and inverted residual structures. It also designs an attention mechanism that takes into account both channel and spatial features to improve the quality of long-distance information transmission and highlight key information. Finally, it outputs the detection result of whether it is an ice crystal pixel by pixel.

[0104] This invention also provides an experiment to verify the accuracy of the ice detection model of this invention. UNet, DANet, DDRNet, ENet, and SFNet, which have similar network structures to the ice detection model of this invention, are used for comparison. The F1 score, precision, recall, and overall accuracy of the ice detection results are selected as the performance evaluation indicators of the method, as shown below:

[0105]

[0106]

[0107] Using the pixel as the smallest evaluation unit, TP represents the number of pixels that are icicles but are detected as icicles, TN represents the number of pixels that are icicles but are detected as background, FP represents the number of pixels that are background but are detected as icicles, and FN represents the number of pixels that are background but are detected as background.

[0108] The quantified experimental results are shown in Table 2. SFNet achieved the best overall performance index (F1) at 81.897%, while UNet performed the worst at 81.155%. The compared models performed similarly on the dataset, with little performance difference, and an average detection accuracy of approximately 81.6%. The proposed method outperforms SFNet by 0.924%, 1.038%, 0.872%, and 0.975% on the four metrics, respectively, demonstrating comprehensive superiority over the compared models. Overall, it outperforms the average model level by approximately 1.2%.

[0109] Table 2 Comparison of accuracy of UAV ice datasets (%)

[0110] Model F1 value Accuracy Recall rate Total accuracy UNet 81.155 79.139 81.136 81.175 DANet 81.614 79.553 81.971 81.26 DDRNet 81.584 79.604 81.598 81.569 ENet 81.589 79.574 81.748 81.431 SFNet 81.897 79.912 82.07 81.724 IDM 82.821 80.95 82.942 82.701

[0111] Figure 7 , Figure 8 and Figure 9 For a qualitative comparison of typical detection results, Region 1 is an area with a large number of floating ice floes and unfrozen water puddles in the middle. The main focus is on comparing the edge detection accuracy of each model. Figure 7 (a) is the original image of region one. Figure 7(b) is Figure 7 (a) The corresponding image of icicles. Figure 7 (c) To use UNet Figure 7 (a) Graph showing the results of the processing. Figure 7 (d) Using DANet Figure 7 (a) Graph showing the results of the processing. Figure 7 (e) is for using DDRNet Figure 7 (a) Graph showing the results of the processing. Figure 7 (f) is to use ENet for Figure 7 (a) Graph showing the results of the processing. Figure 7 (g) Using SFNet Figure 7 (a) Graph showing the results of the processing. Figure 7 (h) refers to the use of IDM for Figure 7 (a) The result of the processing, such as Figure 7 As shown, the edges of puddles are complex and irregular. DANet identifies them as approximately elliptical, while DDRNet and SFNet have weaker ability to identify subtle edges. The method proposed in this invention produces detection results that are closer to the label edges. Region 2 is a large background area with only a few ice crystals. The main comparison focuses on the detection capabilities of each model in the background and in similar scenarios. Figure 8 (a) is the original image of region two. Figure 8 (b) is Figure 8 (a) The corresponding image of icicles. Figure 8 (c) To use UNet Figure 8 (a) Graph showing the results of the processing. Figure 8 (d) Using DANet Figure 8 (a) Graph showing the results of the processing. Figure 8 (e) is for using DDRNet Figure 8 (a) Graph showing the results of the processing. Figure 8 (f) is to use ENet for Figure 8 (a) Graph showing the results of the processing. Figure 8 (g) Using SFNet Figure 8 (a) Graph showing the results of the processing. Figure 8 (h) refers to the use of IDM for Figure 8 (a) The result of the processing, such as Figure 8 As shown, the DANet and DDRNet models are basically unable to identify ice floes, and ENet shows significant missed detections. Our method and UNet perform slightly better than other comparative methods in this type of scenario. Region 3 represents ice floe detection areas of varying sizes, primarily examining the model's adaptability to objects of different scales and its overall detection capabilities. Figure 9 (a) is the original image of region three. Figure 9 (b) is Figure 9(a) The corresponding image of icicles. Figure 9 (c) To use UNet Figure 9 (a) Graph showing the results of the processing. Figure 9 (d) Using DANet Figure 9 (a) Graph showing the results of the processing. Figure 9 (e) is for using DDRNet Figure 9 (a) Graph showing the results of the processing. Figure 9 (f) is to use ENet for Figure 9 (a) Graph showing the results of the processing. Figure 9 (g) Using SFNet Figure 9 (a) Graph showing the results of the processing. Figure 9 (h) refers to the use of IDM for Figure 9 (a) The result of the processing, such as Figure 9 As shown, DANet and DDRNet are weak in detecting small ice particles, ENet has gaps in detecting large ice floes, and the other three methods show little difference. Overall, the ice detection model of this invention is superior to the comparison models.

[0112] This invention also provides an ice detection system corresponding to the above method, comprising:

[0113] The acquisition module is used to acquire remote sensing images of the river to be predicted.

[0114] An ice jam detection module is used to input the remote sensing image of the river to be predicted into a trained ice jam detection model to obtain ice jams in the remote sensing image of the river to be predicted. The trained ice jam detection model is obtained by training the ice jam detection model with the sample river remote sensing image as input and the ice jams in the sample river remote sensing image as output. The ice jam detection model includes an encoding layer, a multi-head attention layer, and a decoding layer. The output of the encoding layer is connected to the input of the multi-head attention layer and the input of the decoding layer, and the output of the multi-head attention layer is connected to the input of the decoding layer. The encoding layer includes multiple feature encoding modules. The feature encoding modules are inverted residual structures and include depth parameterized convolutional layers. The multi-head attention layer includes spatial attention feature mechanisms and channel attention feature mechanisms.

[0115] This invention also provides an electronic device, comprising:

[0116] A memory and a processor, the memory being used to store a computer program, the processor running the computer program to cause the electronic device to perform the ice detection method according to the above embodiments.

[0117] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the ice detection method described in the above embodiments.

[0118] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.

[0119] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for detecting ice crystals, characterized in that, include: Acquire remote sensing images of the river to be predicted; The ice crystals in the remote sensing image of the river to be predicted are obtained by inputting the trained ice crystal detection model into the remote sensing image of the river to be predicted. The trained ice jam detection model is obtained by training a sample river remote sensing image as input and ice jams in the sample river remote sensing image as output. The ice jam detection model includes an encoding layer, a multi-head attention layer, and a decoding layer. The output of the encoding layer is connected to the input of the multi-head attention layer and the input of the decoding layer, and the output of the multi-head attention layer is connected to the input of the decoding layer. The encoding layer includes multiple feature encoding modules. The feature encoding module is an inverted residual structure and includes a depth parameterized convolutional layer. The feature encoding module includes a first stacking operation layer and a first Conv convolutional layer, a first Do-Conv layer, a second Conv convolutional layer, and a first activation layer connected in sequence. The input of the first Conv convolutional layer is connected to the input of the first activation layer through the first stacking operation layer. The multi-head attention layer includes a spatial attention feature mechanism and a channel attention feature mechanism.

2. The ice crystal detection method according to claim 1, characterized in that, The coding layer includes: a first Conv2d convolutional layer, a first feature coding convolutional layer, a second feature coding convolutional layer, a third feature coding convolutional layer, a fourth feature coding convolutional layer, and a fifth feature coding convolutional layer connected in sequence; The first feature encoding convolutional layer includes a first encoding input layer, a first feature encoding layer, and a second Conv2d convolutional layer connected in sequence; the second feature encoding convolutional layer includes a second encoding input layer, a second feature encoding layer, and a third Conv2d convolutional layer connected in sequence; the third feature encoding convolutional layer includes a third encoding input layer, a third feature encoding layer, and a fourth Conv2d convolutional layer connected in sequence; the fourth feature encoding convolutional layer includes a fourth encoding input layer, a fourth feature encoding layer, and a fifth Conv2d convolutional layer connected in sequence; the fifth feature encoding convolutional layer includes a fifth encoding input layer, a fifth feature encoding layer, and a sixth Conv2d convolutional layer connected in sequence. The first feature coding layer, the second feature coding layer, and the fourth feature coding layer each include three sequentially connected feature coding modules; the third feature coding layer includes four sequentially connected feature coding modules; and the fifth feature coding layer includes one feature coding module. The outputs of the first, second, third, fourth, and fifth Conv2d convolutional layers are all connected to the input of the multi-head attention layer, and the output of the sixth Conv2d convolutional layer is connected to the decoding layer.

3. The ice crystal detection method according to claim 2, characterized in that, The multi-head attention layer includes five multi-head attention modules; each of the five multi-head attention modules includes an input feature vector branch, a second superposition operation layer, a channel attention branch, and a spatial attention branch. The input terminals of the input feature vector branch are connected to the first input terminals of the channel attention branch and the spatial attention branch, respectively, serving as the input terminals of the multi-head attention module; the first output terminal of the input feature vector branch is connected to the second input terminals of the channel attention branch and the spatial attention branch, respectively; the second output terminals of the input feature vector branch, the channel attention branch, and the spatial attention branch are all connected to the input terminals of the second superposition operation layer; the output terminal of the second superposition operation layer is the output layer of the multi-head attention module. The input of the first multi-head attention module is connected to the output of the first Conv2d convolutional layer; the input of the second multi-head attention module is connected to the output of the second Conv2d convolutional layer; the input of the third multi-head attention module is connected to the output of the third Conv2d convolutional layer; the input of the fourth multi-head attention module is connected to the output of the fourth Conv2d convolutional layer; and the input of the fifth multi-head attention module is connected to the output of the fifth Conv2d convolutional layer. The output layers of the first, second, third, fourth, and fifth multi-head attention modules are all connected to the input of the decoding layer.

4. The ice crystal detection method according to claim 3, characterized in that, The decoding layer includes a first concatenated feature encoding deconvolution layer, a second concatenated feature encoding deconvolution layer, a third concatenated feature encoding deconvolution layer, a fourth concatenated feature encoding deconvolution layer, a fifth concatenated feature encoding deconvolution layer, and a second activation layer; The first concatenated feature encoding deconvolution layer comprises a first decoding input layer, a first concatenation layer, a sixth feature encoding layer, and a first ConvT2d convolution layer connected in sequence; the second concatenated feature encoding deconvolution layer comprises a second decoding input layer, a second concatenation layer, a seventh feature encoding layer, and a second ConvT2d convolution layer connected in sequence; the third concatenated feature encoding deconvolution layer comprises a third decoding input layer, a third concatenation layer, an eighth feature encoding layer, and a third ConvT2d convolution layer connected in sequence; the fourth concatenated feature encoding deconvolution layer comprises a fourth decoding input layer, a fourth concatenation layer, a ninth feature encoding layer, and a fourth ConvT2d convolution layer connected in sequence; and the fifth concatenated feature encoding deconvolution layer comprises a fifth decoding input layer, a fifth concatenation layer, a tenth feature encoding layer, and a fifth ConvT2d convolution layer connected in sequence. The sixth, seventh, eighth, and ninth feature coding layers each include three sequentially connected feature coding modules; the tenth feature coding layer includes two sequentially connected feature coding modules. The input of the first decoding input layer is connected to the output of the multi-head attention module output layer of the fifth multi-head attention module; the input of the sixth feature encoding layer is connected to the output of the sixth Conv2d convolutional layer; the input of the second decoding input layer is connected to the output of the fourth multi-head attention module; the input of the seventh feature encoding layer is connected to the output of the first Conv2d convolutional layer; the input of the third decoding input layer is connected to the output of the third multi-head attention module; the input of the eighth feature encoding layer is connected to the output of the second Conv2d convolutional layer; the input of the fourth decoding input layer is connected to the output of the second multi-head attention module; the input of the ninth feature encoding layer is connected to the output of the third Conv2d convolutional layer; the input of the fifth decoding input layer is connected to the output of the first multi-head attention module; the input of the tenth feature encoding layer is connected to the output of the fourth Conv2d convolutional layer; and the input of the second activation layer is connected to the output of the fifth Conv2d convolutional layer.

5. The ice crystal detection method according to claim 3, characterized in that, The spatial attention branch includes: a fourth Conv convolutional layer, a global average pooling layer, a max pooling layer, a fifth Conv convolutional layer, a third activation layer, and a first convolutional operation layer connected in sequence; the input of the global average pooling layer is connected to the input of the input feature vector branch and the input of the channel attention branch, respectively; the input of the first convolutional operation layer is connected to the first output of the input feature vector branch; and the output of the first convolutional operation layer is connected to the input of the second stacking operation layer.

6. The ice crystal detection method according to claim 5, characterized in that, The channel attention branch includes: a sixth Conv convolutional layer, a first reshape layer, a second reshape layer, a seventh Conv convolutional layer, a product operation layer, and a second convolutional operation layer; the input of the sixth Conv convolutional layer is connected to the input of the global average pooling layer and the input of the input feature vector branch; the output of the sixth Conv convolutional layer is connected to the inputs of the first reshape layer and the second reshape layer, respectively; the output of the first reshape layer is connected to the input of the seventh Conv convolutional layer; the outputs of the seventh Conv convolutional layer and the second reshape layer are respectively connected to the input of the product operation layer, the output of the product operation layer is connected to the first input of the second convolutional operation layer, the second input of the second convolutional operation layer is connected to the first output of the input feature vector branch; and the output of the convolutional operation layer is connected to the input of the second stacking operation layer.

7. The ice crystal detection method according to claim 1, characterized in that, The process of determining the trained ice crystal detection model includes: Construct an ice crystal detection model; Images of sample rivers flowing from low latitudes to high latitudes were captured during the early stages of freezing in early winter and the early stages of ice thawing in early spring. Sample river remote sensing images are obtained from the river section images; the sample river remote sensing images are all images including ice floes in the river section images; The ice floes in the sample river remote sensing image are obtained from the sample river remote sensing image; The ice detection model is trained by using the sample river remote sensing image as input and the ice floes in the sample river remote sensing image as output.

8. An ice crystal detection system, characterized in that, include: The acquisition module is used to acquire remote sensing images of the river to be predicted. An ice jam detection module is used to input the remote sensing image of the river to be predicted into a trained ice jam detection model to obtain ice jams in the remote sensing image of the river to be predicted. The trained ice jam detection model is obtained by training the ice jam detection model with the sample river remote sensing image as input and the ice jams in the sample river remote sensing image as output. The ice jam detection model includes an encoding layer, a multi-head attention layer, and a decoding layer. The output of the encoding layer is connected to the input of the multi-head attention layer and the input of the decoding layer, and the output of the multi-head attention layer is connected to the input of the decoding layer. The encoding layer includes multiple feature encoding modules. The feature encoding module is an inverted residual structure and includes a depth parameterized convolutional layer. The feature encoding module includes a first stacking operation layer and a first Conv convolutional layer, a first Do-Conv layer, a second Conv convolutional layer, and a first activation layer connected in sequence. The input of the first Conv convolutional layer is connected to the input of the first activation layer through the first stacking operation layer. The multi-head attention layer includes a spatial attention feature mechanism and a channel attention feature mechanism.

9. An electronic device, characterized in that, include: A memory and a processor, the memory for storing a computer program, the processor for running the computer program to cause the electronic device to perform the ice detection method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the ice detection method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Cooperated detecting method for ice of The Yellow River based on multispectral image and SAR image

    CN105913023A

  • Yellow River ice semantic segmentation method based on multi-attention mechanism double-flow fusion network

    CN111160311A

  • River ice distribution intelligent extraction method, device and equipment and storage medium

    CN116012738A