River Ice Semantic Segmentation Method Based on YOLOv8 Instance Segmentation Framework

By introducing an improved channel attention mechanism and maximum value suppression algorithm in the YOLOv8 instance segmentation framework, the balance problem of speed and accuracy in Heling semantic segmentation is solved, and efficient segmentation of complex edge targets is achieved.

CN120182609BActive Publication Date: 2025-07-22NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510659827.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-07-22
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

The existing Heling semantic segmentation method is difficult to improve segmentation accuracy while maintaining the segmentation speed, especially the segmentation effect of complex edge targets is poor.

Method used

The Heling semantic segmentation method based on the YOLOv8 instance segmentation framework is adopted, and hierarchical feature extraction is performed through the feature encoder. Combined with the improved channel attention mechanism in the channel attention network, the bounding box and segmentation results are generated by the detection head and segmentation head after the characteristics are fused, and maximum value suppression is performed through the output network to optimize the segmentation effect.

Benefits of technology

The effective balance of the speed and accuracy of Heling semantic segmentation is achieved, the segmentation effect of complex edge targets is improved, and the model detection accuracy of dense targets is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182609B_ABST
    Figure CN120182609B_ABST
Patent Text Reader

Abstract

This application relates to the field of deep learning technology, and particularly to a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework, including: inputting the river ice image to be segmented into a feature encoder for hierarchical feature extraction to obtain three features of different scales; respectively inputting the three features of different scales into three channel attention modules in a channel attention network for adaptive calibration and fusion to obtain three fused features of different scales; respectively inputting the three fused features of different scales into three different detection heads to obtain the bounding box detection results of each instance; inputting the three fused features of different scales into a segmentation head to obtain the full-image segmentation result; performing non-maximum suppression on the bounding box detection results by an output network to obtain the final bounding box, and cropping the full-image segmentation result based on the final bounding box to obtain the final segmentation result. This method realizes an effective balance between the speed and accuracy of river ice semantic segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical fields of artificial intelligence and deep learning, and particularly to a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework. Background Art

[0002] River ice, also known as ice flood, commonly known as ice raft, is a hydrological phenomenon in which the water level of rivers rises significantly due to the resistance of ice jams to the flow of water. It mostly occurs in river sections flowing from lower latitudes to higher latitudes and with an obvious north-south flow direction. For such river sections, during the three stages of freezing, ice cover, and thawing, the downstream freezes earlier than the upstream during freezing, and the upstream thaws earlier than the downstream during thawing, thus causing ice floods. When the ice accumulates and gradually evolves into ice jams, ice dams, and ice pressures, it is likely to cause serious damage to buildings, cultivated land, etc. near the river channel.

[0003] Today, with the development and maturity of deep learning technology, most tasks in the field of computer vision rely on deep learning technology to complete. Semantic segmentation is one of the basic tasks in the field of computer vision, and its goal is to perform pixel-level segmentation on a given image. River ice segmentation is one of the most important technologies in ice regime monitoring research. If the river ice image can be segmented quickly and accurately, the situation of ice in the river can be monitored in real time and early warnings can be given in a timely manner.

[0004] However, the currently proposed river ice semantic segmentation methods are difficult to balance segmentation speed and segmentation accuracy. For example, the river ice semantic segmentation technology based on deep learning proposed by Abhineet et al., and the technology proposed by Zhang et al. that first extracts features from river ice images at multiple scales, then fuses them into a unified representation, and finally performs semantic segmentation on the unified representation, etc. Among these technologies, the Transformer structure represented by Segformer and the fully convolutional network structure represented by Unet++ have high segmentation accuracy, but high computational complexity and slow inference speed. While the segmentation model with instance objects as the recognition unit, such as the YOLOV8 model, has a speed advantage, but for river ice targets with complex edge shapes, the detail segmentation accuracy is low, and there will be missegmentation phenomena such as adhesion and merging. Summary of the Invention

[0005] In view of this, the embodiments of the present application propose a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework, aiming to achieve fast river ice semantic segmentation by utilizing the high inference efficiency of the YOLOv8 framework, and adding an improved lightweight channel attention mechanism in the feature fusion stage, effectively overcoming the problem of poor segmentation effect of the YOLOv8 framework on complex object edges, and achieving an effective balance between the speed and accuracy of river ice semantic segmentation.

[0006] To achieve the above object, an embodiment of the present application proposes a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework. The method includes the following steps: Input the river ice image to be segmented into a feature encoder for hierarchical feature extraction. The feature encoder contains a serialized combination of multiple composite convolution modules, and finally three different scales of features are obtained; Use the three different scales of features as the original features and input them into three independent channel attention modules in the channel attention network respectively. Each channel attention module consists of two parallel branches, namely the average pooling branch and the max pooling branch. The average pooling branch is used to perform adaptive average pooling on the input features to extract channel features, and then perform non-linear mapping through two KAN linear layers. The max pooling branch is used to perform adaptive max pooling on the input features to extract channel features, and then perform non-linear mapping through two KAN linear layers. The outputs of the two branches are activated through the Sigmoid activation function layer after addition, and then multiplied by the corresponding original features to achieve adaptive calibration and fusion, and finally three different scales of fused features are obtained; Input the three different scales of fused features into three different detection heads to detect the bounding boxes and categories of instances, and obtain the bounding box detection results and classification results of each instance; Input the three different scales of fused features into the segmentation head. The segmentation head first generates a prototype segmentation result based on the fused feature with the largest scale, and then generates the segmentation results of each channel based on the three different scales of fused features. Finally, the prototype segmentation result and the segmentation results of each channel are connected together to obtain the full-image segmentation result; The output network performs non-maximum suppression on the bounding box detection results to obtain the final bounding boxes, and then crops the full-image segmentation result based on the final bounding boxes to obtain the final segmentation results of each instance.

[0007] To achieve the above object, an embodiment of the present application also proposes an electronic device. The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework as described above.

[0008] To achieve the above object, an embodiment of the present application also provides a computer-readable storage medium storing a computer program, which when executed by a processor, can implement a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework as described above.

[0009] A method for ice floe semantic segmentation based on the YOLOv8 instance segmentation framework proposed in this application realizes ice floe semantic segmentation by constructing, training, and using an ice floe semantic segmentation model composed of a feature encoder, a channel attention network, three detection heads, a segmentation head, and an output network. The feature encoder performs hierarchical feature extraction on the ice floe image to be segmented through a serial combination of multiple composite convolution modules, and can extract three different scales of features. Three independent channel attention modules in the channel attention network further process the three different scales of features respectively, and an improved channel attention mechanism is added in the feature fusion stage to fully encode the feature information to achieve adaptive calibration and fusion. Finally, three different scales of high-quality fused features can be obtained, which helps to improve the accuracy and speed of ice floe semantic segmentation. The detection head and the segmentation head respectively output the bounding box detection results of each instance and the full-image segmentation result. Finally, the output network performs non-maximum suppression on the bounding box detection results to obtain the final bounding box, and then crops the full-image segmentation result based on the final bounding box to obtain the final segmentation result of each instance. Performing non-maximum suppression on the bounding box detection results can effectively improve the detection accuracy of the model for dense targets. All in all, this method effectively overcomes the problem that the YOLOv8 framework has a poor segmentation effect on complex object edges, and realizes an effective balance between the speed and accuracy of ice floe semantic segmentation.

[0010] Optionally, the feature encoder consists of five cascaded composite convolution modules. The input of the first composite convolution module is the ice floe image to be segmented, and the input of the subsequent composite convolution module is the output of the previous composite convolution module. Each composite convolution module consists of an initial convolution unit and a C2f unit. The initial convolution unit adopts a 3×3 convolution with a downsampling strategy of stride 2, and reduces the input scale to 1 / 2 of the original scale by dynamically setting the convolution kernel size and padding size to obtain the preliminary dimensionality-reduced feature. The C2f unit expands the number of channels of the preliminary dimensionality-reduced feature to 2×e×C through a 1×1 convolution, where C is the target output channel number and e is the channel expansion coefficient. Then, through a dimension splitting operation, the expanded feature is evenly decomposed into a main branch feature and a secondary branch feature. The secondary branch feature is non-linearly transformed through a Bottleneck layer, and the main branch feature and the non-linearly transformed secondary branch feature are concatenated in the channel dimension, and then feature fusion is performed through a 1×1 convolution and output. The feature encoder uses the features output by the last three composite convolution modules as three different scales of features, and the scales of the three different scales of features are 1 / 8, 1 / 16, and 1 / 32 of the ice floe image to be segmented respectively.

[0011] Optionally, the Bottleneck layer adopts a compression-expansion structure. First, it uses a 3×3 convolution to compress the number of channels to e×C, then uses a 3×3 convolution to restore the number of channels to the original number of channels, and realizes feature reuse through a residual connection, finally realizing the non-linear transformation of the sub-branch features; the composite convolution module ensures the size consistency of the convolution operation through a dynamic kernel padding algorithm. For a convolution kernel with a dilation rate d greater than 1, its equivalent kernel size is automatically extended to d×(k - 1)+1, where k is the size of the original convolution kernel. In the Bottleneck layer, the residual connection is enabled if and only if the number of input channels is equal to the number of output channels. Each convolution in the composite convolution module adopts a grouped convolution strategy, and the number of groups g is adjustable.

[0012] Optionally, for the channel attention module, its average pooling branch consists of an average pooling unit, two average pooling KAN linear layers, and an average pooling multi-layer perceptron, and its max pooling branch consists of a max pooling unit, two max pooling KAN linear layers, and a max pooling multi-layer perceptron; in the average pooling branch, the average pooling unit is used to perform adaptive average pooling on the input features to extract channel features, the first average pooling KAN linear layer is used for channel compression, the second average pooling KAN linear layer is used to restore the channel dimension, and the average pooling multi-layer perceptron is used to generate an average pooling channel weight matrix based on the output of the second average pooling KAN linear layer; in the max pooling branch, the max pooling unit is used to perform adaptive max pooling on the input features to extract channel features, the first max pooling KAN linear layer is used for channel compression, the second max pooling KAN linear layer is used to restore the channel dimension, and the max pooling multi-layer perceptron is used to generate a max pooling channel weight matrix based on the output of the second max pooling KAN linear layer; the average pooling channel weight matrix and the max pooling channel weight matrix are activated through a Sigmoid activation function layer after addition to obtain the final channel weight matrix, and finally the final channel weight matrix is multiplied element-wise with the corresponding original features to obtain fused features of the same scale as the original features.

[0013] Optionally, the structures of the three detection heads are the same, and each consists of a bounding box detection branch and a class detection branch. The structures of the bounding box detection branch and the class detection branch are the same, and each consists of two convolutional layers and a standard two-dimensional convolutional layer. The bounding box detection branch and the class detection branch respectively output the bounding box detection results and classification results of each instance.

[0014] Optionally, the segmentation head is composed of a ProtoNet network, a segmentation coefficient generation network and an output layer; the ProtoNet network transforms the channel dimension of the largest scale fusion feature through a 3×3 convolution layer, thereby improving the nonlinear expression capability while maintaining the integrity of the feature space, and then performs feature upsampling through a learnable transposed convolution layer, and finally completes the recognition of the prototype through two convolution layers again to generate a prototype segmentation result; the segmentation coefficient generation network is composed of three scale branches and a 1×1 convolution layer, each scale branch is provided with a feature refinement module including two cascaded 3×3 convolution layers, the first 3×3 convolution layer is used to compress the number of channels to 1 / 4 of the original number of channels, the second 3×3 convolution layer is used to optimize the feature representation through learnable parameters, and finally the output of each scale branch is uniformly mapped to a 32-dimensional segmentation coefficient space through a 1×1 convolution layer, and the multi-scale segmentation coefficients are spliced along the spatial dimension through a dimensional reorganization operation to form a complete mask coefficient matrix; the output layer is used to perform a matrix multiplication operation on the prototype segmentation result and the mask coefficient matrix, dynamically generate an accurate segmentation mask for each instance, and thus obtain a full-image segmentation result.

[0015] Optionally, the output network performs maximum suppression on the bounding box detection result using the SoftNMS maximum suppression algorithm to obtain a final bounding box. When the SoftNMS maximum suppression algorithm detects that the IoU of a bounding box is greater than a preset threshold, it reduces the confidence of the bounding box, but does not delete the bounding box whose IoU is greater than the preset threshold.

[0016] Optionally, the feature encoder, the channel attention network, the three different detection heads, the segmentation head and the output network together constitute the Heling semantic segmentation model. When training the Heling semantic segmentation model, the polygon key point alignment algorithm is used to align the true value data corresponding to the training samples, and harmless interpolation is performed on the basis of retaining the original points as much as possible to calculate the loss value. The execution process of the polygon key point alignment algorithm includes:

[0017] Initialize the data structure. First, build a maximum heap structure, traverse the original boundary segment set, encapsulate each segment into an object containing four attributes: segment index, start and end point coordinates, initial key point number, and segment length. Then store them in the maximum heap in descending order of weight value to form a priority processing queue.

[0018] Dynamic key point allocation, calculate the current total number of key points, extract the longest weighted line segment at the top of the heap, increase the number of key points of the longest weighted line segment at the top of the heap by 1, update the weight of the longest weighted line segment at the top of the heap and put it back into the heap, and simultaneously increase the total number of key points, and execute the dynamic key point allocation operation in a loop to ensure that the long edge gets more key points first, so as to achieve adaptive maintenance of the boundary shape;

[0019] For ordered key point reconstruction, after sorting the line segments in the heap according to the original index, key points are generated. First, the starting points of the original line segments are retained, and then new points are evenly inserted on each line segment. The interpolation ratio factor for each point is t = i / (pc - 1), where i is the index value of the point, pc is the total number of points, and t is the interpolation ratio factor. The coordinates are generated by linear interpolation. After interpolation is completed, the coordinates of the original boundary end points are finally appended.

[0020] The process of generating coordinates by linear interpolation is expressed by the formula:

[0021] x' = x_start + t×(x_end - x_start);

[0022] y' = y_start + t×(y_end - y_start);

[0023] Among them, x_start and y_start respectively represent the starting point positions of the line segment in terms of length and width, x_end and y_end respectively represent the ending point positions of the line segment in terms of length and width, and x' and y' represent the new coordinates obtained after interpolation. After alignment is completed, the bounding box detection results of the training samples are used to crop the full-image segmentation results of the training samples output by the segmentation head to obtain the final segmentation results of each instance of the training samples. Then, cross-entropy is calculated with the aligned ground truth data to obtain the loss value. Based on the loss value, the backpropagation algorithm is used to update the parameters until the preset convergence condition is met, and then the trained river ice semantic segmentation model is obtained. Description of the Drawings

[0024] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the related art, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application or the related art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] Figure 1 It is a flowchart of a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework provided in an embodiment of the present application;

[0026] Figure 2 It is a structural schematic diagram of a river ice semantic segmentation model provided in an embodiment of the present application;

[0027] Figure 3 It is a structural schematic diagram of a channel attention module provided in an embodiment of the present application;

[0028] Figure 4It is a schematic structural diagram of a splitting head provided in an embodiment of the present application;

[0029] Figure 5 It is a comparison diagram of the final segmentation result and the river ice image to be segmented provided in an embodiment of the present application;

[0030] Figure 6 It is a schematic diagram of the experimental results of a comparative experiment conducted on the NWPU_YRCC2 dataset provided in an embodiment of the present application;

[0031] Figure 7 It is a schematic diagram of the experimental results of a comparative experiment conducted on the NWPU_YRCC_MS dataset provided in an embodiment of the present application;

[0032] Figure 8 It is a schematic diagram of the experimental results of a comparative experiment conducted on the Albert River Ice Segmentation dataset provided in an embodiment of the present application;

[0033] Figure 9 It is a schematic structural diagram of an electronic device provided in another embodiment of the present application. Detailed implementation manners

[0034] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in the embodiments of the present application, many technical details are provided for readers to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions required to be protected by the present application can still be implemented. The following division of each embodiment is for convenience of description and should not constitute any limitation on the specific implementation manner of the present application. Each embodiment can be combined and cross-referenced with each other on the premise of not being contradictory.

[0035] An embodiment of the present application proposes a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework, which is applied to an electronic device. Among them, the electronic device can be a terminal or a server. In this embodiment and the following embodiments, the server is taken as an example for illustration. The implementation details of a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework proposed in this embodiment will be specifically described below. The following content is only implementation details provided for convenient understanding and is not necessary for implementing this solution.

[0036] The specific process of a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework proposed in this embodiment can be as Figure 1 shown and includes:

[0037] Step 11: Input the river ice image to be segmented into the feature encoder for hierarchical feature extraction. The feature encoder contains a serialized combination of multiple composite convolution modules, and finally three different scales of features are obtained.

[0038] In a specific implementation, the river ice semantic segmentation task is implemented by a river ice semantic segmentation model, which consists of a feature encoder, a channel attention network, three detection heads, a segmentation head, and an output network. The river ice image to be segmented is used as the input of the model and enters the feature encoder. The feature encoder contains a serialized combination of multiple composite convolution modules. With the serialized combination of these composite convolution modules, hierarchical feature extraction is performed on the river ice image to be segmented, and finally three different scales of features can be obtained.

[0039] In an example, the river ice image to be segmented is an aerial image taken by a drone, or it can also be a multispectral remote sensing image taken by a remote sensing satellite.

[0040] In an example, the specific structure of the river ice semantic segmentation model can be as Figure 2 shown. The river ice semantic segmentation model is based on YOLOv8 and uses the YOLO Backbone network as the backbone. However, in practical applications, other architectures can also be used to build the model.

[0041] In an example, as Figure 2 shown, the feature encoder consists of five cascaded composite convolution modules (i.e., Figure 2 P1 to P5 in it). The input of the first composite convolution module is the river ice image to be segmented, and the input of the subsequent composite convolution module is the output of the previous composite convolution module.

[0042] Each composite convolution module consists of an initial convolution unit and a C2f unit.

[0043] The initial convolution unit uses a 3×3 convolution with a downsampling strategy of stride 2. By dynamically setting the convolution kernel size and padding size, it is ensured that the input scale is strictly reduced to 1 / 2 of the original scale to obtain the preliminary dimensionality-reduced features. Among them, the convolution layer and the batch normalization layer adopt a zero-bias design, followed by a SiLU activation function to enhance the non-linear expression ability.

[0044] The C2f unit expands the number of channels of the preliminary dimensionality-reduced features to 2×e×C through a 1×1 convolution, where C is the target output channel number and e is the channel expansion coefficient. Then, through a dimension splitting operation, the expanded features are evenly decomposed into the main branch features and the sub-branch features. The main branch features retain the original feature information, and the sub-branch features need to undergo a non-linear transformation through a Bottleneck layer. Subsequently, the main branch features and the non-linearly transformed sub-branch features are concatenated in the channel dimension, and then feature fusion is performed through a 1×1 convolution and output.

[0045] The feature encoder uses the features output by the last three composite convolution modules (P3, P4, P5) as features of three different scales, and the scales of the features of the three different scales are 1 / 8, 1 / 16, and 1 / 32 of the river ice image to be segmented respectively.

[0046] In one example, the Bottleneck layer adopts a squeeze-and-excitation structure. First, it uses a 3×3 convolution to compress the number of channels to e×C, then uses a 3×3 convolution to restore the number of channels to the original number of channels, and realizes feature reuse through residual connection, and finally realizes the non-linear transformation of the sub-branch features.

[0047] It should be noted that the composite convolution module needs to ensure the size consistency of the convolution operation through the dynamic kernel padding algorithm. For a convolution kernel with a dilation rate d greater than 1, its equivalent kernel size is automatically extended to d×(k−1)+1, where k is the original convolution kernel size, and the symmetric padding value is calculated accordingly. In the Bottleneck layer, the residual connection is enabled if and only if the number of input channels is equal to the number of output channels. This conditional judgment mechanism effectively suppresses the gradient dissipation problem while maintaining the feature expression ability. Each convolution in the composite convolution module adopts a grouped convolution strategy, and the number of groups g is adjustable. By reusing parameters, the computational complexity is reduced, enabling the model to be deployed and processed in real time on embedded devices such as drones.

[0048] Step 12: Use the features of the three different scales as the original features and input them into three independent channel attention modules in the channel attention network respectively. Each channel attention module consists of two parallel branches, namely the average pooling branch and the max pooling branch. The average pooling branch is used to perform adaptive average pooling on the input features to extract channel features, and then perform non-linear mapping through two KAN linear layers. The max pooling branch is used to perform adaptive max pooling on the input features to extract channel features, and then perform non-linear mapping through two KAN linear layers. The outputs of the two branches are activated through the Sigmoid activation function layer after being added, and then multiplied by the corresponding original features to achieve adaptive calibration and fusion, finally obtaining the fused features of the three different scales.

[0049] In a specific implementation, the channel attention network is the most core part of the river ice semantic segmentation model. Features of three different scales will be used as the original features and are respectively input into three independent channel attention modules in the channel attention network. Each channel attention module consists of two parallel branches, namely the average pooling branch and the max pooling branch. The average pooling branch is used to perform adaptive average pooling on the input features to extract channel features, and then perform non-linear mapping through two KAN linear layers. The max pooling branch is used to perform adaptive max pooling on the input features to extract channel features, and then perform non-linear mapping through two KAN linear layers. The outputs of the two branches are activated through a Sigmoid activation function layer after addition, and then multiplied by the corresponding original features to achieve adaptive calibration and fusion, finally obtaining fused features of three different scales.

[0050] In one example, as Figure 2 shown, IceChannelAttn1, IceChannelAttn2, and IceChannelAtt3 are three independent channel attention modules. The features output by P3 of the feature extractor will enter IceChannelAttn1, the features output by P4 will enter IceChannelAttn2, and the features output by P5 will enter IceChannelAtt3.

[0051] In one example, the specific structure of the channel attention module can be as Figure 3 shown. For any channel attention module, its average pooling branch consists of an average pooling unit, two average pooling KAN linear layers, and an average pooling multi-layer perceptron. Its max pooling branch consists of a max pooling unit, two max pooling KAN linear layers, and a max pooling multi-layer perceptron.

[0052] In the average pooling branch, the average pooling unit is used to perform adaptive average pooling on the input features to extract channel features. The first average pooling KAN linear layer is used for channel compression, the second average pooling KAN linear layer is used to restore the channel dimension, and the average pooling multi-layer perceptron is used to generate an average pooling channel weight matrix based on the output of the second average pooling KAN linear layer.

[0053] In the max pooling branch, the max pooling unit is used to perform adaptive max pooling on the input features to extract channel features. The first max pooling KAN linear layer is used for channel compression, the second max pooling KAN linear layer is used to restore the channel dimension, and the max pooling multi-layer perceptron is used to generate a max pooling channel weight matrix based on the output of the second max pooling KAN linear layer.

[0054] The average pooling channel weight matrix and the max pooling channel weight matrix are activated through the Sigmoid activation function layer after addition to obtain the final channel weight matrix. Finally, the final channel weight matrix is multiplied by the corresponding original features channel by channel to obtain the fused features of the same scale as the original features.

[0055] Step 13: Input the fused features of three different scales into three different detection heads respectively to detect the bounding boxes and categories of instances, and obtain the bounding box detection results and classification results of each instance.

[0056] In specific implementation, the fused features of three different scales will be input into three different detection heads respectively to detect the bounding boxes and categories of instances, and finally the bounding box detection results and classification results of each instance can be obtained.

[0057] In one example, the structures of the three detection heads are the same, and each consists of a bounding box detection branch and a category detection branch. The structures of the bounding box detection branch and the category detection branch are also the same, and each consists of two convolutional layers and a standard two-dimensional convolutional layer. The bounding box detection branch and the category detection branch will output the bounding box detection results and classification results of each instance respectively.

[0058] Step 14: Input the fused features of three different scales into the segmentation head. The segmentation head first generates the prototype segmentation result according to the fused feature with the largest scale, then generates the segmentation results of each channel according to the fused features of three different scales, and finally connects the prototype segmentation result and the segmentation results of each channel together to obtain the full-image segmentation result.

[0059] In specific implementation, while the detection head is working, the fused features of three different scales will also be input into the segmentation head to implement the work of the segmentation head. The segmentation head first generates the prototype segmentation result according to the fused feature with the largest scale, then generates the segmentation results of each channel according to the fused features of three different scales, and finally connects the prototype segmentation result and the segmentation results of each channel together to obtain the full-image segmentation result.

[0060] In one example, the specific structure of the segmentation head can be as Figure 4 shown. The specific segmentation head consists of a ProtoNet network, a segmentation coefficient generation network, and an output layer.

[0061] The ProtoNet network transforms the fused feature with the largest scale (i.e., the fused feature output by IceChannelAttn1 in Figure 2 ) through a 3×3 convolutional layer for channel dimension transformation, improves the non-linear expression ability while maintaining the integrity of the feature space, then performs feature upsampling through a learnable transposed convolutional layer, and finally completes the recognition of the prototype through two convolutional layers again to generate the prototype segmentation result.

[0062] The segmentation coefficient generation network consists of three scale branches and a 1×1 convolutional layer. Each scale branch is equipped with a feature refinement module containing two cascaded 3×3 convolutional layers. The first 3×3 convolutional layer is used to compress the number of channels to 1 / 4 of the original number of channels, and the second 3×3 convolutional layer is used to optimize the feature representation through learnable parameters. Finally, the output of each scale branch is uniformly mapped to a 32-dimensional segmentation coefficient space through a 1×1 convolutional layer, and the multi-scale segmentation coefficients are concatenated along the spatial dimension through a dimension reorganization operation to form a complete mask coefficient matrix.

[0063] The output layer is used to perform a matrix multiplication operation on the prototype segmentation result and the mask coefficient matrix to dynamically generate an accurate segmentation mask for each instance, thereby obtaining the full-image segmentation result.

[0064] The segmentation head effectively reduces the computational complexity through the segmentation prototype sharing mechanism, and at the same time maintains the sensitivity to targets of different scales by combining the multi-scale coefficient fusion strategy. In particular, a parametric deconvolution upsampling is adopted in the prototype generation network to replace the traditional interpolation method, enhancing the learnability of the feature space transformation. The segmentation coefficient generation network adopts a double-convolution layer cascade structure, suppressing feature redundancy through channel compression while ensuring the receptive field, and finally achieving a balanced optimization of segmentation accuracy and computational efficiency.

[0065] Step 15: The output network performs non-maximum suppression on the bounding box detection results to obtain the final bounding box, and then crops the full-image segmentation result based on the final bounding box to obtain the final segmentation result of each instance.

[0066] In a specific implementation, after the output network obtains the bounding box detection results, classification results, and full-image segmentation results, it can perform non-maximum suppression on the bounding box detection results to obtain the final bounding box, and then crop the full-image segmentation result based on the final bounding box to obtain the final segmentation result of each instance.

[0067] It should be noted that the principle of the standard non-maximum suppression algorithm is to calculate the IoU between other bounding boxes and the current bounding box with the maximum confidence. If the IoU is greater than a certain threshold, the bounding boxes that meet the conditions around the current bounding box with the maximum confidence will be deleted, which may cause dense targets to be recognized as a single target or lost. Therefore, in this embodiment, the output network uses the SoftNMS non-maximum suppression algorithm to perform non-maximum suppression on the bounding box detection results to obtain the final bounding box. When the SoftNMS non-maximum suppression algorithm detects that the IoU of a bounding box is greater than the preset threshold, it reduces the confidence of the bounding box, but does not delete the bounding box whose IoU is greater than the preset threshold. With such a design, dense targets can be better recognized after selecting an appropriate confidence threshold.

[0068] In one embodiment, the comparison between the final segmentation result of each instance and the original river image to be segmented can be as follows: Figure 5 shown.

[0069] In an example, the feature encoder, channel attention network, three different detection heads, segmentation head and output network together constitute the Heling semantic segmentation model. When training the Heling semantic segmentation model, it is necessary to use the polygon key point alignment algorithm to align the true value data corresponding to the training samples, and perform harmless interpolation on the basis of retaining the original points as much as possible to calculate the loss value.

[0070] The traditional alignment using linear interpolation will cause the original points to change, which will lead to large deviations in image edge targets and non-convex polygonal targets. The polygon key point alignment algorithm performs harmless interpolation while retaining the original points as much as possible, thus avoiding this problem.

[0071] The execution process of the polygon key point alignment algorithm is as follows:

[0072] First, the data structure is initialized. The maximum heap structure is constructed first, and the original boundary segment set is traversed. Each segment is encapsulated as an object containing four attributes: segment index, start and end point coordinates, initial number of key points (default is 1), and segment length (calculated by Euclidean distance). Then, it is stored in the maximum heap in descending order according to the weight value (length divided by the current number of key points) to form a priority processing queue.

[0073] Next is the dynamic key point allocation, which calculates the current total number of key points (initially the number of original segments), extracts the longest weighted segment at the top of the stack, increases the number of key points of the longest weighted segment at the top of the stack by 1, updates the weight of the longest weighted segment at the top of the stack and puts it back into the stack, and simultaneously increases the total number of key points. The dynamic key point allocation operation is executed in a loop to ensure that the long edges get more key points first, so as to achieve adaptive maintenance of the boundary shape.

[0074] Next is the ordered key point reconstruction. After sorting the line segments in the heap according to the original index, the key points are generated. First, the starting point of the original line segment is retained, and then new points are evenly inserted on each line segment. The interpolation scale factor of each point is t. The calculation formula of t is t= i / (pc-1), where i is the index value of the point, pc is the total number of points, and t is the interpolation scale factor. The coordinates are generated by linear interpolation. After the interpolation is completed, the original boundary end point coordinates are finally appended.

[0075] The process of generating coordinates through linear interpolation can be expressed by the formula:

[0076] x'= x_start + t×(x_end-x_start);

[0077] y' = y_start + t×(y_end - y_start);

[0078] Among them, x_start and y_start respectively represent the starting point positions of the line segment in terms of length and width, x_end and y_end respectively represent the ending point positions of the line segment in terms of length and width, and x' and y' represent the new coordinates obtained after interpolation.

[0079] After alignment, the bounding box detection results of the training samples are used to crop the full-image segmentation results of the training samples output by the segmentation head to obtain the final segmentation results of each instance of the training samples. Then, cross-entropy is calculated with the aligned ground truth data to obtain the loss value. Based on the loss value, the parameter update is performed using the backpropagation algorithm until the preset convergence condition is met, and finally, the trained ice jam semantic segmentation model is obtained.

[0080] An ice jam semantic segmentation method based on the YOLOv8 instance segmentation framework proposed in this embodiment constructs, trains, and uses an ice jam semantic segmentation model composed of a feature encoder, a channel attention network, three detection heads, a segmentation head, and an output network to achieve ice jam semantic segmentation. The feature encoder performs hierarchical feature extraction on the ice jam image to be segmented through the serial combination of multiple composite convolution modules, and can extract three different scales of features. Three independent channel attention modules in the channel attention network further process the three different scales of features, and an improved channel attention mechanism is added in the feature fusion stage to fully encode the feature information to achieve adaptive calibration and fusion. Finally, high-quality fused features of three different scales can be obtained, which helps to improve the accuracy and speed of ice jam semantic segmentation. The detection head and the segmentation head respectively output the bounding box detection results and the full-image segmentation results of each instance. Finally, the output network performs non-maximum suppression on the bounding box detection results to obtain the final bounding box, and then crops the full-image segmentation results based on the final bounding box to obtain the final segmentation results of each instance. Performing non-maximum suppression on the bounding box detection results can effectively improve the detection accuracy of the model for dense targets. In summary, this method effectively overcomes the problem of poor segmentation effect of the YOLOv8 framework on complex object edges and achieves an effective balance between the speed and accuracy of ice jam semantic segmentation.

[0081] The step divisions of the above various methods are only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, they are all within the protection scope of this application; adding insignificant modifications or introducing insignificant designs to the algorithm or process, but not changing the core design of its algorithm and process, are all within the protection scope of this application.

[0082] In one embodiment, to verify the effectiveness of a river ice semantic segmentation method (river ice semantic segmentation model, hereinafter referred to as UiceNet or OURS for short) based on the YOLOv8 instance segmentation framework proposed in this application, we used the datasets NWPU_YRCC2, NWPU_YRCC_MS, and Albert River Ice Segmentation for training and testing, and compared with other methods.

[0083] NWPU_YRCC2 is an aerial photography dataset of the Yellow River ice floes, which contains 1,220 training images and 305 test images. NWPU_YRCC_MS is a hyperspectral dataset of satellite remote sensing of the Yellow River, which contains 960 training images and 241 test images. Albert River Ice Segmentation is an aerial photography dataset of the Albert River, which contains 554 training images and 139 test images after processing.

[0084] The experimental results are as Figure 6 、 Figure 7 、 Figure 8 shown. Figure 6 This is a schematic diagram of the experimental results of the comparative experiment conducted on the NWPU_YRCC2 dataset. Figure 7 This is a schematic diagram of the experimental results of the comparative experiment conducted on the NWPU_YRCC_MS dataset. Figure 8 This is a schematic diagram of the experimental results of the comparative experiment conducted on the Albert River Ice Segmentation dataset. Among many commonly used algorithms or models, including convolutional and Transformer-based networks, this method outperforms traditional commonly used algorithms on multiple datasets. The evaluation metric for the semantic segmentation task is mIoU, which is the mean of the intersection over union (IoU) and is used to measure the performance of the semantic segmentation model. In the comparison with convolutional and Transformer-based semantic segmentation models, UiceNet shows very competitive results, achieving 87.29% mIoU on NWPU_YRCC_MS, 92.69% mIoU on NWPU_YRCC2, and even 94.00% mIoU on Albert River Ice Segmentation.

[0085] At the same time, when UiceNet infers 768px×768px images using a single NVIDIA Tesla V100 GPU, the running speed can reach 14ms per input image, that is, it can smoothly process a video stream of 71FPS.

[0086] The above experimental results show that UiceNet can balance accuracy and inference speed.

[0087] Another embodiment of the present application proposes an electronic device, and its specific structure can be as Figure 9 shown, including: at least one processor 21; and a memory 22 communicatively connected to the at least one processor 21; wherein, the memory 22 stores instructions executable by the at least one processor 21, and the instructions are executed by the at least one processor 21 so that the at least one processor 21 can execute a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework as described in the above method embodiment.

[0088] Among them, the memory and the processor are connected by a bus. The bus can include any number of interconnected buses and bridges, and the bus can connect various circuits of one or more processors and memories together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be an element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium.

[0089] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory can be used to store data used by the processor when executing operations.

[0090] Another embodiment of the present application proposes a computer-readable storage medium storing a computer program, which when executed by a processor, can implement a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework as described in the above method embodiment.

[0091] That is, those skilled in the art can understand that all or part of the steps in the above method embodiments can be completed by instructing relevant hardware through a program. This program is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0092] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application, and in practical applications, various changes can be made in form and details without departing from the spirit and scope of the present application.

Claims

1. A river ice semantic segmentation method based on the YOLOv8 instance segmentation framework, characterized in that, Including: Input the river ice image to be segmented into the feature encoder for hierarchical feature extraction. The feature encoder contains a serialized combination of multiple composite convolution modules, and finally three different scales of features are obtained; Take the three different scales of features as the original features and input them into three independent channel attention modules in the channel attention network respectively. Each channel attention module consists of two parallel branches, namely the average pooling branch and the max pooling branch. The average pooling branch is used to perform adaptive average pooling on the input features to extract channel features, and then perform non-linear mapping through two KAN linear layers. The max pooling branch is used to perform adaptive max pooling on the input features to extract channel features, and then perform non-linear mapping through two KAN linear layers. The outputs of the two branches are activated through the Sigmoid activation function layer after addition, and then multiplied by the corresponding original features to achieve adaptive calibration and fusion, and finally three different scales of fused features are obtained; Input the three different scales of fused features into three different detection heads to detect the bounding boxes and categories of the instances, and obtain the bounding box detection results and classification results of each instance; Input the three different scales of fused features into the segmentation head. The segmentation head first generates the prototype segmentation result according to the fused feature with the largest scale, and then generates the segmentation results of each channel according to the three different scales of fused features. Finally, the prototype segmentation result and the segmentation results of each channel are connected together to obtain the full-image segmentation result; The output network performs non-maximum suppression on the bounding box detection results to obtain the final bounding boxes, and then crops the full-image segmentation results based on the final bounding boxes to obtain the final segmentation results of each instance.

2. The river ice semantic segmentation method based on the YOLOv8 instance segmentation framework according to claim 1, characterized in that The feature encoder consists of five cascaded composite convolution modules. The input of the first composite convolution module is the river ice image to be segmented, and the input of the subsequent composite convolution module is the output of the previous composite convolution module; Each composite convolution module consists of an initial convolution unit and a C2f unit; The initial convolution unit uses a 3×3 convolution with a downsampling strategy of stride 2. By dynamically setting the convolution kernel size and padding size, the input scale is reduced to 1 / 2 of the original scale to obtain the preliminary dimensionality reduction features; The C2f unit expands the number of channels of the preliminary dimensionality reduction features to 2×e×C through a 1×1 convolution, where C is the target output channel number and e is the channel expansion coefficient. Then, through the dimension splitting operation, the expanded features are evenly decomposed into the main branch features and the sub-branch features. The sub-branch features are non-linearly transformed through the Bottleneck layer, and the main branch features and the non-linearly transformed sub-branch features are concatenated in the channel dimension, and then feature fusion is performed through a 1×1 convolution and then output; The feature encoder takes the features output by the last three composite convolution modules as three different scales of features, and the scales of the three different scales of features are 1 / 8, 1 / 16, and 1 / 32 of the river ice image to be segmented respectively.

3. A method for river ice semantic segmentation based on the YOLOv8 instance segmentation framework according to claim 2, characterized in that, The Bottleneck layer adopts a compression-expansion structure. First, it uses a 3×3 convolution to compress the number of channels to e×C, then uses a 3×3 convolution to restore the number of channels to the original number of channels, and realizes feature reuse through residual connection, finally realizing the non-linear transformation of the secondary branch features; The composite convolution module ensures the size consistency of the convolution operation through the dynamic kernel padding algorithm. For a convolution kernel with a dilation rate d greater than 1, its equivalent kernel size is automatically extended to d×(k - 1)+1, where k is the size of the original convolution kernel. In the Bottleneck layer, the residual connection is enabled if and only if the number of input channels is equal to the number of output channels. Each convolution in the composite convolution module adopts a grouped convolution strategy, and the number of groups g is adjustable.

4. A river ice semantic segmentation method based on the YOLOv8 instance segmentation framework according to claim 3, characterized in that, For the channel attention module, its average pooling branch consists of an average pooling unit, two average pooling KAN linear layers, and an average pooling multi-layer perceptron, and its max pooling branch consists of a max pooling unit, two max pooling KAN linear layers, and a max pooling multi-layer perceptron; In the average pooling branch, the average pooling unit is used to perform adaptive average pooling on the input features to extract channel features. The first average pooling KAN linear layer is used for channel compression, the second average pooling KAN linear layer is used to restore the channel dimension, and the average pooling multi-layer perceptron is used to generate the average pooling channel weight matrix based on the output of the second average pooling KAN linear layer; In the max pooling branch, the max pooling unit is used to perform adaptive max pooling on the input features to extract channel features. The first max pooling KAN linear layer is used for channel compression, the second max pooling KAN linear layer is used to restore the channel dimension, and the max pooling multi-layer perceptron is used to generate the max pooling channel weight matrix based on the output of the second max pooling KAN linear layer; The average pooling channel weight matrix and the max pooling channel weight matrix are activated through the Sigmoid activation function layer after addition to obtain the final channel weight matrix. Finally, the final channel weight matrix is multiplied element-wise with the corresponding original features to obtain the fused features of the same scale as the original features.

5. A method for river ice semantic segmentation based on the YOLOv8 instance segmentation framework according to claim 4, characterized in that, The structures of the three detection heads are the same, and each consists of a bounding box detection branch and a class detection branch. The structures of the bounding box detection branch and the class detection branch are the same, and each consists of two convolutional layers and a standard two-dimensional convolutional layer. The bounding box detection branch and the class detection branch respectively output the bounding box detection results and classification results of each instance.

6. A river ice semantic segmentation method based on the YOLOv8 instance segmentation framework according to claim 5, characterized in that, The segmentation head consists of a ProtoNet network, a segmentation coefficient generation network, and an output layer; The ProtoNet network performs channel dimension transformation on the fused feature with the largest scale through a 3×3 convolutional layer, improving the non-linear expression ability while maintaining the integrity of the feature space. Then, it performs feature upsampling through a learnable transposed convolutional layer, and finally completes the prototype recognition through two convolutional layers again to generate the prototype segmentation result; The segmentation coefficient generation network consists of three scale branches and a 1×1 convolutional layer. Each scale branch is equipped with a feature refinement module containing two cascaded 3×3 convolutional layers. The first 3×3 convolutional layer is used to compress the number of channels to 1 / 4 of the original number of channels, and the second 3×3 convolutional layer is used to optimize the feature representation through learnable parameters. Finally, the 1×1 convolutional layer maps the outputs of each scale branch to a 32-dimensional segmentation coefficient space uniformly, and the multi-scale segmentation coefficients are concatenated along the spatial dimension through a dimension reorganization operation to form a complete mask coefficient matrix. The output layer is used to perform a matrix multiplication operation on the prototype segmentation result and the mask coefficient matrix to dynamically generate an accurate segmentation mask for each instance, thereby obtaining the full-image segmentation result.

7. A method for river ice semantic segmentation based on the YOLOv8 instance segmentation framework according to claim 6, characterized in that, The output network performs non-maximum suppression on the bounding box detection results using the SoftNMS non-maximum suppression algorithm to obtain the final bounding box. When the IoU of the detected bounding box is greater than the preset threshold, the SoftNMS non-maximum suppression algorithm reduces the confidence of the bounding box, but does not delete the bounding box with an IoU greater than the preset threshold.

8. A method for river ice semantic segmentation based on the YOLOv8 instance segmentation framework according to any one of claims 1 to 7, characterized in that, The feature encoder, channel attention network, three different detection heads, segmentation head, and output network together constitute the ice jam semantic segmentation model. When training the ice jam semantic segmentation model, the polygon key point alignment algorithm is used to align the ground truth data corresponding to the training samples, and harmless interpolation is performed on the basis of retaining the original points to calculate the loss value. The execution process of the polygon key point alignment algorithm includes: Data structure initialization: First, construct a max heap structure. Traverse the set of original boundary line segments, encapsulate each line segment as an object containing four attributes: line segment index, start and end coordinates, initial number of key points, and line segment length, and then store them in the max heap in descending order of weight value to form a priority processing queue. Dynamic key point allocation: Calculate the current total number of key points, extract the longest weighted line segment at the top of the heap, increase the number of key points of the longest weighted line segment at the top of the heap by 1, update the weight of the longest weighted line segment at the top of the heap and then re-enter the heap, and synchronously increase the total number of key points. Loop to execute the dynamic key point allocation operation to ensure that the long sides obtain more key points first, and achieve adaptive preservation of the boundary shape. Ordered key point reconstruction: After sorting the line segments in the heap by the original index, generate key points. First, retain the starting point of the original line segment, and then uniformly insert new points on each line segment. The interpolation ratio factor for each point is t = i / (pc - 1), where i is the index value of the point, pc is the total number of points, and t is the interpolation ratio factor. The coordinates are generated by linear interpolation, and finally the original boundary end coordinates are appended after the interpolation is completed. The process of generating coordinates by linear interpolation is expressed by the formula: x' = x_start + t×(x_end - x_start); y' = y_start + t×(y_end - y_start); Among them, x_start and y_start respectively represent the starting point positions of the line segment in terms of length and width, x_end and y_end respectively represent the ending point positions of the line segment in terms of length and width, and x' and y' represent the new coordinates obtained after interpolation; After alignment, the full-image segmentation result of the training sample output by the segmentation head is cropped using the bounding box detection result of the training sample output by the detection head to obtain the final segmentation result of each instance of the training sample. Then, cross-entropy is calculated with the aligned ground truth data to obtain the loss value. Based on the loss value, parameter update is performed using the backpropagation algorithm until the preset convergence condition is met, and then the trained ice jam semantic segmentation model is obtained.

9. An electronic device, characterized in that, Including: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a method for ice jam semantic segmentation based on the YOLOv8 instance segmentation framework as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it can implement a method for ice jam semantic segmentation based on the YOLOv8 instance segmentation framework as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Real-time semantic segmentation method for unmanned aerial vehicle aerial image of Yellow River ice

    CN114943835A

  • Ice detection method and system, electronic equipment and medium

    CN116977778A