YOLOv8 instance segmentation framework-based River-River semantic segmentation method
By introducing a lightweight channel attention mechanism into the YOLOv8 instance segmentation framework, the existing Heling semantic segmentation method has solved the shortcomings in taking into account speed and accuracy, and the efficient and precise segmentation effect of Heling semantic segmentation is achieved.
Patent Information
- Application Number
- CN202510659827.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-21
AI Technical Summary
The existing Heling semantic segmentation method is difficult to take into account both the segmentation speed and the segmentation accuracy. The YOLOV8 model has low detail segmentation accuracy when segmenting Heling targets with complex edge shapes, and is prone to incorrect segmentation such as sticky and merge.
The Heling semantic segmentation method based on the YOLOv8 instance segmentation framework improves the segmentation effect of target complex edges by adding an improved lightweight channel attention mechanism in the feature fusion stage. This method includes a feature encoder, a channel attention network, a detection head, a segmentation head and an output network. Through hierarchical feature extraction, a channel attention mechanism and a multi-scale segmentation strategy, the speed and accuracy of Heling semantic segmentation are balanced.
It effectively overcomes the problem of poor segmentation of complex edges of targets by YOLOV8 framework, realizes an effective balance between the speed and accuracy of Heling semantic segmentation, improves segmentation accuracy and maintains a faster inference speed.
Smart Images

Figure CN120182609A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the technical fields of artificial intelligence and deep learning, and particularly to a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework. Background Art
[0002] River ice, also known as ice flood and commonly called ice raft, is a hydrological phenomenon in which the water level of rivers rises significantly due to the resistance of ice jams to the water flow. It mostly occurs in river sections flowing from lower latitudes to higher latitudes and with an obvious north-south flow direction. For such river sections, during the three stages of freezing period, ice cover period and thawing period, the downstream freezes earlier than the upstream when the river freezes, and the upstream thaws earlier than the downstream during thawing, thus causing ice floods. When the ice accumulates more and more and gradually evolves into ice jams, ice dams and ice pressure, it is easy to cause serious damage to buildings, cultivated land, etc. near the river channel.
[0003] Today, with the development and maturity of deep learning technology, most tasks in the field of computer vision rely on deep learning technology. Semantic segmentation is one of the basic tasks in the field of computer vision, and its goal is to perform pixel-level segmentation on a given image. River ice segmentation is one of the most important technologies in ice regime monitoring research. If the river ice image can be segmented quickly and accurately, the situation of ice in the river can be monitored in real time and early warnings can be given in time.
[0004] However, the currently proposed river ice semantic segmentation methods are difficult to balance segmentation speed and segmentation accuracy. For example, the river ice semantic segmentation technology based on deep learning proposed by Abhineet et al., and the technology proposed by Zhang et al. that first extracts features of river ice images at multiple scales respectively, then fuses them into a unified representation, and finally performs semantic segmentation on the unified representation, etc. Among these technologies, the Transformer structure represented by Segformer and the fully convolutional network structure represented by Unet++ have high segmentation accuracy, but high computational complexity and slow inference speed. While the segmentation models with instance objects as the recognition unit, such as the YOLOV8 model, have speed advantages, but for river ice targets with complex edge shapes, the detail segmentation accuracy is low, and there will be missegmentation phenomena such as adhesion and merging. Summary of the Invention
[0005] In view of this, embodiments of the present application propose a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework, aiming to achieve fast river ice semantic segmentation by utilizing the high inference efficiency of the YOLOv8 framework, and adding an improved lightweight channel attention mechanism in the feature fusion stage, effectively overcoming the problem of poor segmentation effect of the YOLOv8 framework on complex target edges, and achieving an effective balance between the speed and accuracy of river ice semantic segmentation.
[0006] To achieve the above object, an embodiment of the present application proposes a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework. The method includes the following steps: input the river ice image to be segmented into a feature encoder for hierarchical feature extraction. The feature encoder contains a serialized combination of multiple composite convolution modules, and finally obtains three different scales of features; use the three different scales of features as the original features and input them into three independent channel attention modules in the channel attention network respectively. Each channel attention module consists of two parallel branches, namely the average pooling branch and the max pooling branch. The average pooling branch is used to perform adaptive average pooling on the input features to extract channel features, and then perform non-linear mapping through two KAN linear layers. The max pooling branch is used to perform adaptive max pooling on the input features to extract channel features, and then perform non-linear mapping through two KAN linear layers. The outputs of the two branches are activated through the Sigmoid activation function layer after addition, and then multiplied by the corresponding original features to achieve adaptive calibration and fusion, and finally obtain three different scales of fused features; input the three different scales of fused features into three different detection heads to detect the bounding boxes and categories of instances, and obtain the bounding box detection results and classification results of each instance; input the three different scales of fused features into the segmentation head. The segmentation head first generates a prototype segmentation result according to the fused feature with the largest scale, and then generates the segmentation results of each channel according to the three different scales of fused features. Finally, the prototype segmentation result and the segmentation results of each channel are connected together to obtain the full-image segmentation result; the output network performs non-maximum suppression on the bounding box detection results to obtain the final bounding boxes, and then crops the full-image segmentation result based on the final bounding boxes to obtain the final segmentation results of each instance.
[0007] To achieve the above object, an embodiment of the present application also proposes an electronic device. The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework as described above.
[0008] To achieve the above object, an embodiment of the present application also provides a computer-readable storage medium storing a computer program, which when executed by a processor, can implement a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework as described above.
[0009] A method for ice jam semantic segmentation based on the YOLOv8 instance segmentation framework proposed in this application constructs, trains, and uses an ice jam semantic segmentation model composed of a feature encoder, a channel attention network, three detection heads, a segmentation head, and an output network to achieve ice jam semantic segmentation. The feature encoder performs hierarchical feature extraction on the ice jam image to be segmented through the serial combination of multiple compound convolution modules, and can extract three different scales of features. Three independent channel attention modules in the channel attention network further process the three different scales of features respectively, and an improved channel attention mechanism is added in the feature fusion stage to fully encode the feature information to achieve adaptive calibration and fusion. Finally, high-quality fused features of three different scales can be obtained, which helps to improve the accuracy and speed of ice jam semantic segmentation. The detection head and the segmentation head respectively output the bounding box detection results of each instance and the full-image segmentation results. Finally, the output network performs non-maximum suppression on the bounding box detection results to obtain the final bounding boxes, and then crops the full-image segmentation results based on the final bounding boxes to obtain the final segmentation results of each instance. Performing non-maximum suppression on the bounding box detection results can effectively improve the detection accuracy of the model for dense targets. All in all, this method effectively overcomes the problem that the YOLOv8 framework has a poor segmentation effect on complex object edges, and achieves an effective balance between the speed and accuracy of ice jam semantic segmentation.
[0010] Optionally, the feature encoder is composed of five cascaded compound convolution modules. The input of the first compound convolution module is the ice jam image to be segmented, and the input of the subsequent compound convolution module is the output of the previous compound convolution module. Each compound convolution module is composed of an initial convolution unit and a C2f unit. The initial convolution unit adopts a 3×3 convolution with a downsampling strategy of stride 2, and reduces the input scale to 1 / 2 of the original scale by dynamically setting the convolution kernel size and padding size to obtain preliminary dimension-reduced features. The C2f unit expands the number of channels of the preliminary dimension-reduced features to 2×e×C through 1×1 convolution, where C is the target output channel number and e is the channel expansion coefficient. Then, the expanded features are evenly decomposed into a main branch feature and a secondary branch feature through a dimension splitting operation. The secondary branch feature is non-linearly transformed through a Bottleneck layer, and the main branch feature and the non-linearly transformed secondary branch feature are concatenated in the channel dimension, and then feature fusion is performed through 1×1 convolution and output. The feature encoder uses the features output by the last three compound convolution modules as three different scales of features, and the scales of the three different scales of features are 1 / 8, 1 / 16, and 1 / 32 of the ice jam image to be segmented respectively.
[0011] Optionally, the Bottleneck layer adopts a compression-expansion structure. First, it uses a 3×3 convolution to compress the number of channels to e×C, then uses a 3×3 convolution to restore the number of channels to the original number of channels, and realizes feature reuse through a residual connection, finally realizing a non-linear transformation of the sub-branch features; the composite convolution module ensures the size consistency of the convolution operation through a dynamic kernel padding algorithm. For a convolution kernel with a dilation rate d greater than 1, its equivalent kernel size is automatically expanded to d×(k - 1)+1, where k is the size of the original convolution kernel. In the Bottleneck layer, the residual connection is enabled if and only if the number of input channels is equal to the number of output channels. Each convolution in the composite convolution module adopts a grouped convolution strategy, and the number of groups g is adjustable.
[0012] Optionally, for the channel attention module, its average pooling branch consists of an average pooling unit, two average pooling KAN linear layers, and an average pooling multi-layer perceptron; its max pooling branch consists of a max pooling unit, two max pooling KAN linear layers, and a max pooling multi-layer perceptron; in the average pooling branch, the average pooling unit is used to perform adaptive average pooling on the input features to extract channel features, the first average pooling KAN linear layer is used for channel compression, the second average pooling KAN linear layer is used to restore the channel dimension, and the average pooling multi-layer perceptron is used to generate an average pooling channel weight matrix based on the output of the second average pooling KAN linear layer; in the max pooling branch, the max pooling unit is used to perform adaptive max pooling on the input features to extract channel features, the first max pooling KAN linear layer is used for channel compression, the second max pooling KAN linear layer is used to restore the channel dimension, and the max pooling multi-layer perceptron is used to generate a max pooling channel weight matrix based on the output of the second max pooling KAN linear layer; the average pooling channel weight matrix and the max pooling channel weight matrix are activated through a Sigmoid activation function layer after addition to obtain the final channel weight matrix, and finally the final channel weight matrix is multiplied with the corresponding original features channel by channel to obtain fused features of the same scale as the original features.
[0013] Optionally, the structures of the three detection heads are the same, and each consists of a bounding box detection branch and a class detection branch. The structures of the bounding box detection branch and the class detection branch are the same, and each consists of two convolutional layers and a standard two-dimensional convolutional layer. The bounding box detection branch and the class detection branch respectively output the bounding box detection results and classification results of each instance.
[0014] Optionally, the segmentation head is composed of a ProtoNet network, a segmentation coefficient generation network and an output layer; the ProtoNet network transforms the channel dimension of the largest scale fusion feature through a 3×3 convolution layer, thereby improving the nonlinear expression capability while maintaining the integrity of the feature space, and then performs feature upsampling through a learnable transposed convolution layer, and finally completes the recognition of the prototype through two convolution layers again to generate a prototype segmentation result; the segmentation coefficient generation network is composed of three scale branches and a 1×1 convolution layer, each scale branch is provided with a feature refinement module including two cascaded 3×3 convolution layers, the first 3×3 convolution layer is used to compress the number of channels to 1 / 4 of the original number of channels, the second 3×3 convolution layer is used to optimize the feature representation through learnable parameters, and finally the output of each scale branch is uniformly mapped to a 32-dimensional segmentation coefficient space through a 1×1 convolution layer, and the multi-scale segmentation coefficients are spliced along the spatial dimension through a dimensional reorganization operation to form a complete mask coefficient matrix; the output layer is used to perform a matrix multiplication operation on the prototype segmentation result and the mask coefficient matrix, dynamically generate an accurate segmentation mask for each instance, and thus obtain a full-image segmentation result.
[0015] Optionally, the output network performs maximum suppression on the bounding box detection result using the SoftNMS maximum suppression algorithm to obtain a final bounding box. When the SoftNMS maximum suppression algorithm detects that the IoU of a bounding box is greater than a preset threshold, it reduces the confidence of the bounding box, but does not delete the bounding box whose IoU is greater than the preset threshold.
[0016] Optionally, the feature encoder, the channel attention network, the three different detection heads, the segmentation head and the output network together constitute the Heling semantic segmentation model. When training the Heling semantic segmentation model, the polygon key point alignment algorithm is used to align the true value data corresponding to the training samples, and harmless interpolation is performed on the basis of retaining the original points as much as possible to calculate the loss value. The execution process of the polygon key point alignment algorithm includes: Initialize the data structure. First, build a maximum heap structure, traverse the original boundary segment set, encapsulate each segment into an object containing four attributes: segment index, start and end point coordinates, initial key point number, and segment length. Then store them in the maximum heap in descending order of weight value to form a priority processing queue. Dynamic key point allocation, calculate the current total number of key points, extract the longest weighted line segment at the top of the heap, increase the number of key points of the longest weighted line segment at the top of the heap by 1, update the weight of the longest weighted line segment at the top of the heap and put it back into the heap, and simultaneously increase the total number of key points, and execute the dynamic key point allocation operation in a loop to ensure that the long edge gets more key points first, so as to achieve adaptive maintenance of the boundary shape; For the reconstruction of ordered key points, after sorting the line segments in the heap according to the original index, key points are generated. First, the starting points of the original line segments are retained, and then new points are evenly inserted on each line segment. The interpolation ratio factor for each point is t = i / (pc - 1), where i is the index value of the point, pc is the total number of points, and t is the interpolation ratio factor. The coordinates are generated through linear interpolation. After the interpolation is completed, the coordinates of the ending points of the original boundary are finally appended. The process of generating coordinates through linear interpolation is expressed by the formula: x' = x_start + t×(x_end - x_start); y' = y_start + t×(y_end - y_start); Among them, x_start and y_start respectively represent the starting point positions of the line segment in terms of length and width, x_end and y_end respectively represent the ending point positions of the line segment in terms of length and width, and x' and y' represent the new coordinates obtained after interpolation. After the alignment is completed, the bounding box detection results of the training samples are used to crop the full-image segmentation results of the training samples output by the segmentation head, obtaining the final segmentation results of each instance of the training samples. Then, cross-entropy is calculated with the aligned ground truth data to obtain the loss value. Based on the loss value, the parameter update is performed using the backpropagation algorithm until the preset convergence condition is met, and the trained river ice semantic segmentation model is obtained. Description of the Drawings
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the related art, the following will briefly introduce the drawings required for the description of the embodiments of the present application or the related art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 It is a flowchart of a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework provided in an embodiment of the present application; Figure 2 It is a structural schematic diagram of a river ice semantic segmentation model provided in an embodiment of the present application; Figure 3 It is a structural schematic diagram of a channel attention module provided in an embodiment of the present application; Figure 4 It is a structural schematic diagram of a segmentation head provided in an embodiment of the present application; Figure 5 It is a comparison diagram of the final segmentation result and the river ice image to be segmented provided in an embodiment of the present application; Figure 6It is a schematic diagram of the experimental results of a comparative experiment conducted on the NWPU_YRCC2 dataset provided in an embodiment of the present application; Figure 7 It is a schematic diagram of the experimental results of a comparative experiment conducted on the NWPU_YRCC_MS dataset provided in an embodiment of the present application; Figure 8 It is a schematic diagram of the experimental results of a comparative experiment conducted on the Albert River Ice Segmentation dataset provided in an embodiment of the present application; Figure 9 It is a schematic diagram of the structure of an electronic device provided in another embodiment of the present application. Detailed implementation manners
[0019] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in the embodiments of the present application, many technical details are proposed to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions required to be protected by the present application can still be implemented. The following division of each embodiment is for convenience of description and should not constitute any limitation to the specific implementation manner of the present application. Each embodiment can be combined and cross-referenced with each other on the premise of not conflicting with each other.
[0020] An embodiment of the present application proposes a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework, which is applied to an electronic device. Herein, the electronic device can be a terminal or a server. In this embodiment and the following embodiments, the server is taken as an example for illustration. The implementation details of a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework proposed in this embodiment will be specifically described below. The following content is only implementation details provided for convenient understanding and is not necessary for implementing this solution.
[0021] The specific process of a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework proposed in this embodiment can be as Figure 1 shown and includes: Step 11: Input the river ice image to be segmented into a feature encoder for hierarchical feature extraction. The feature encoder includes a serialized combination of multiple composite convolution modules, and finally three different scales of features are obtained.
[0022] In a specific implementation, the river ice semantic segmentation task is implemented by a river ice semantic segmentation model, which consists of a feature encoder, a channel attention network, three detection heads, a segmentation head, and an output network. The river ice image to be segmented is used as the input of the model and enters the feature encoder. The feature encoder contains a serialized combination of multiple composite convolution modules. With the serialized combination of these composite convolution modules, hierarchical feature extraction is performed on the river ice image to be segmented, and finally, features of three different scales can be obtained.
[0023] In one example, the river ice image to be segmented is an aerial image taken by a drone, or it can also be a multispectral remote sensing image taken by a remote sensing satellite.
[0024] In one example, the specific structure of the river ice semantic segmentation model can be as Figure 2 shown. The river ice semantic segmentation model is based on YOLOv8 and uses the YOLO Backbone network as the backbone. However, in actual applications, other architectures can also be used to build the model.
[0025] In one example, as Figure 2 shown, the feature encoder consists of five cascaded composite convolution modules (i.e., Figure 2 P1 to P5 in
[0026] ). The input of the first composite convolution module is the river ice image to be segmented, and the input of the subsequent composite convolution module is the output of the previous composite convolution module.
[0027] Each composite convolution module consists of an initial convolution unit and a C2f unit.
[0028] The initial convolution unit uses a 3×3 convolution with a downsampling strategy of stride 2. By dynamically setting the convolution kernel size and padding size, it ensures that the input scale is strictly reduced to 1 / 2 of the original scale, obtaining preliminary dimension-reduced features. Among them, the convolution layer and the batch normalization layer adopt a zero-bias design, followed by a SiLU activation function to enhance the non-linear expression ability.
[0029] The feature encoder uses the features output by the last three composite convolution modules (P3, P4, P5) as features of three different scales, and the scales of the features of the three different scales are 1 / 8, 1 / 16, and 1 / 32 of the river ice image to be segmented, respectively.
[0030] In one example, the Bottleneck layer adopts a compression-expansion structure. First, it uses a 3×3 convolution to compress the number of channels to e×C, then uses a 3×3 convolution to restore the number of channels to the original number of channels, and realizes feature reuse through residual connection, and finally realizes the non-linear transformation of the sub-branch features.
[0031] It should be noted that the composite convolution module needs to ensure the size consistency of the convolution operation through the dynamic kernel padding algorithm. For a convolution kernel with a dilation rate d greater than 1, its equivalent kernel size is automatically expanded to d×(k−1)+1, where k is the original convolution kernel size, and the symmetric padding value is calculated accordingly. In the Bottleneck layer, the residual connection is enabled if and only if the number of input channels is equal to the number of output channels. This condition judgment mechanism effectively suppresses the gradient dissipation problem while maintaining the feature expression ability. Each convolution in the composite convolution module adopts a grouped convolution strategy, and the number of groups g is adjustable. By reusing parameters, the computational complexity is reduced, enabling the model to be deployed and processed in real time on embedded devices such as drones.
[0032] Step 12: Use the features of the three different scales as the original features and input them into three independent channel attention modules in the channel attention network respectively. Each channel attention module consists of two parallel branches, namely the average pooling branch and the max pooling branch. The average pooling branch is used to perform adaptive average pooling on the input features to extract channel features, and then perform non-linear mapping through two KAN linear layers. The max pooling branch is used to perform adaptive max pooling on the input features to extract channel features, and then perform non-linear mapping through two KAN linear layers. The outputs of the two branches are activated through the Sigmoid activation function layer after addition, and then multiplied by the corresponding original features to achieve adaptive calibration and fusion, and finally obtain the fused features of the three different scales.
[0033] In a specific implementation, the channel attention network is the most core part of the river ice semantic segmentation model. Features of three different scales will be used as the original features and are respectively input into three independent channel attention modules in the channel attention network. Each channel attention module consists of two parallel branches, namely the average pooling branch and the max pooling branch. The average pooling branch is used to perform adaptive average pooling on the input features to extract channel features, and then perform non-linear mapping through two KAN linear layers. The max pooling branch is used to perform adaptive max pooling on the input features to extract channel features, and then perform non-linear mapping through two KAN linear layers. The outputs of the two branches are added and then activated through a Sigmoid activation function layer, and then multiplied by the corresponding original features to achieve adaptive calibration and fusion, finally obtaining fused features of three different scales.
[0034] In one example, as Figure 2 shown, IceChannelAttn1, IceChannelAttn2, and IceChannelAtt3 are three independent channel attention modules. The features output by P3 of the feature extractor will enter IceChannelAttn1, the features output by P4 will enter IceChannelAttn2, and the features output by P5 will enter IceChannelAtt3.
[0035] In one example, the specific structure of the channel attention module can be as Figure 3 shown. For any channel attention module, its average pooling branch consists of an average pooling unit, two average pooling KAN linear layers, and an average pooling multi-layer perceptron. Its max pooling branch consists of a max pooling unit, two max pooling KAN linear layers, and a max pooling multi-layer perceptron.
[0036] In the average pooling branch, the average pooling unit is used to perform adaptive average pooling on the input features to extract channel features. The first average pooling KAN linear layer is used for channel compression, the second average pooling KAN linear layer is used to restore the channel dimension, and the average pooling multi-layer perceptron is used to generate an average pooling channel weight matrix based on the output of the second average pooling KAN linear layer.
[0037] In the max pooling branch, the max pooling unit is used to perform adaptive max pooling on the input features to extract channel features. The first max pooling KAN linear layer is used for channel compression, the second max pooling KAN linear layer is used to restore the channel dimension, and the max pooling multi-layer perceptron is used to generate a max pooling channel weight matrix based on the output of the second max pooling KAN linear layer.
[0038] The average pooling channel weight matrix and the max pooling channel weight matrix are activated through a Sigmoid activation function layer after addition to obtain the final channel weight matrix. Finally, the final channel weight matrix is multiplied by the corresponding original features channel by channel to obtain fused features of the same scale as the original features.
[0039] Step 13: Input the fused features of three different scales into three different detection heads respectively to detect the bounding boxes and categories of instances, and obtain the bounding box detection results and classification results of each instance.
[0040] In specific implementation, the fused features of three different scales will be input into three different detection heads respectively to detect the bounding boxes and categories of instances, and finally the bounding box detection results and classification results of each instance can be obtained.
[0041] In an example, the structures of the three detection heads are the same, and each consists of a bounding box detection branch and a category detection branch. The structures of the bounding box detection branch and the category detection branch are also the same, and each consists of two convolutional layers and a standard two-dimensional convolutional layer. The bounding box detection branch and the category detection branch will output the bounding box detection results and classification results of each instance respectively.
[0042] Step 14: Input the fused features of three different scales into the segmentation head. The segmentation head first generates a prototype segmentation result based on the fused feature with the largest scale, then generates the segmentation results of each channel based on the fused features of three different scales, and finally connects the prototype segmentation result and the segmentation results of each channel together to obtain the full-image segmentation result.
[0043] In specific implementation, while the detection head is working, the fused features of three different scales will also be input into the segmentation head to implement the work of the segmentation head. The segmentation head first generates a prototype segmentation result based on the fused feature with the largest scale, then generates the segmentation results of each channel based on the fused features of three different scales, and finally connects the prototype segmentation result and the segmentation results of each channel together to obtain the full-image segmentation result.
[0044] In an example, the specific structure of the segmentation head can be as Figure 4 shown. The specific segmentation head consists of a ProtoNet network, a segmentation coefficient generation network, and an output layer.
[0045] The ProtoNet network performs channel dimension transformation on the fused feature with the largest scale (i.e., the fused feature output by IceChannelAttn1 in Figure 2 ) through a 3×3 convolutional layer to enhance the non-linear expression ability while maintaining the integrity of the feature space. Then, it performs feature upsampling through a learnable transposed convolutional layer, and finally completes the recognition of the prototype through two convolutional layers again to generate the prototype segmentation result.
[0046] The segmentation coefficient generation network consists of three scale branches and a 1×1 convolutional layer. Each scale branch is equipped with a feature refinement module containing two cascaded 3×3 convolutional layers. The first 3×3 convolutional layer is used to compress the number of channels to 1 / 4 of the original number of channels, and the second 3×3 convolutional layer is used to optimize the feature representation through learnable parameters. Finally, the output of each scale branch is uniformly mapped to a 32-dimensional segmentation coefficient space through a 1×1 convolutional layer, and the multi-scale segmentation coefficients are concatenated along the spatial dimension through a dimension reorganization operation to form a complete mask coefficient matrix.
[0047] The output layer is used to perform matrix multiplication on the prototype segmentation result and the mask coefficient matrix to dynamically generate the accurate segmentation mask for each instance, thereby obtaining the full-image segmentation result.
[0048] The segmentation head effectively reduces the computational complexity through the segmentation prototype sharing mechanism, and at the same time maintains the sensitivity to targets of different scales by combining the multi-scale coefficient fusion strategy. In particular, the parametric deconvolution upsampling is adopted in the prototype generation network to replace the traditional interpolation method, enhancing the learnability of the feature space transformation. The segmentation coefficient generation network adopts a double-convolution layer cascade structure, suppressing feature redundancy through channel compression while ensuring the receptive field, and finally achieving a balanced optimization of segmentation accuracy and computational efficiency.
[0049] Step 15: The output network performs non-maximum suppression on the bounding box detection results to obtain the final bounding box, and then crops the full-image segmentation result based on the final bounding box to obtain the final segmentation result of each instance.
[0050] In a specific implementation, after the output network obtains the bounding box detection results, classification results, and full-image segmentation results, it can perform non-maximum suppression on the bounding box detection results to obtain the final bounding box, and then crop the full-image segmentation result based on the final bounding box to obtain the final segmentation result of each instance.
[0051] It should be noted that the principle of the standard non-maximum suppression algorithm is to calculate the IoU between other bounding boxes and the current bounding box with the maximum confidence. If the IoU is greater than a certain threshold, the bounding boxes that meet the conditions around the current bounding box with the maximum confidence will be deleted, which may cause dense targets to be recognized as a single target or lost. Therefore, in this embodiment, the output network uses the SoftNMS non-maximum suppression algorithm to perform non-maximum suppression on the bounding box detection results to obtain the final bounding box. When the SoftNMS non-maximum suppression algorithm detects that the IoU of a bounding box is greater than the preset threshold, it reduces the confidence of the bounding box, but does not delete the bounding box whose IoU is greater than the preset threshold. With such a design, dense targets can be better recognized after selecting an appropriate confidence threshold.
[0052] In one embodiment, the comparison between the final segmentation result of each instance and the original river image to be segmented can be as follows: Figure 5 shown.
[0053] In an example, the feature encoder, channel attention network, three different detection heads, segmentation head and output network together constitute the Heling semantic segmentation model. When training the Heling semantic segmentation model, it is necessary to use the polygon key point alignment algorithm to align the true value data corresponding to the training samples, and perform harmless interpolation on the basis of retaining the original points as much as possible to calculate the loss value.
[0054] The traditional alignment using linear interpolation will cause the original points to change, which will lead to large deviations in image edge targets and non-convex polygonal targets. The polygon key point alignment algorithm performs harmless interpolation while retaining the original points as much as possible, thus avoiding this problem.
[0055] The execution process of the polygon key point alignment algorithm is as follows: First, the data structure is initialized. The maximum heap structure is constructed first, and the original boundary segment set is traversed. Each segment is encapsulated as an object containing four attributes: segment index, start and end point coordinates, initial number of key points (default is 1), and segment length (calculated by Euclidean distance). Then, it is stored in the maximum heap in descending order according to the weight value (length divided by the current number of key points) to form a priority processing queue.
[0056] Next is the dynamic key point allocation, which calculates the current total number of key points (initially the number of original segments), extracts the longest weighted segment at the top of the stack, increases the number of key points of the longest weighted segment at the top of the stack by 1, updates the weight of the longest weighted segment at the top of the stack and puts it back into the stack, and simultaneously increases the total number of key points. The dynamic key point allocation operation is executed in a loop to ensure that long edges get more key points first, so as to achieve adaptive maintenance of the boundary shape.
[0057] Next is the ordered key point reconstruction. After sorting the line segments in the heap according to the original index, the key points are generated. First, the starting point of the original line segment is retained, and then new points are evenly inserted on each line segment. The interpolation scale factor of each point is t. The calculation formula of t is t= i / (pc-1), where i is the index value of the point, pc is the total number of points, and t is the interpolation scale factor. The coordinates are generated by linear interpolation. After the interpolation is completed, the original boundary end point coordinates are finally appended.
[0058] The process of generating coordinates through linear interpolation can be expressed by the formula: x'= x_start + t×(x_end-x_start); y'= y_start + t×(y_end-y_start); Among them, x_start and y_start respectively represent the starting point positions of the line segment in terms of length and width, x_end and y_end respectively represent the ending point positions of the line segment in terms of length and width, and x' and y' represent the new coordinates obtained after interpolation.
[0059] After alignment, the bounding box detection results of the training samples are used to crop the full-image segmentation results of the training samples output by the segmentation head, obtaining the final segmentation results of each instance of the training samples. Then, cross-entropy is calculated with the aligned ground truth data to obtain the loss value. Based on the loss value, the parameter update is performed using the backpropagation algorithm until the preset convergence condition is met, and then the trained ice jam semantic segmentation model is obtained.
[0060] An ice jam semantic segmentation method based on the YOLOv8 instance segmentation framework proposed in this embodiment constructs, trains, and uses an ice jam semantic segmentation model composed of a feature encoder, a channel attention network, three detection heads, a segmentation head, and an output network to implement ice jam semantic segmentation. The feature encoder performs hierarchical feature extraction on the ice jam image to be segmented through the serial combination of multiple composite convolution modules, and can extract three different scales of features. The three independent channel attention modules in the channel attention network further process the three different scales of features respectively, and an improved channel attention mechanism is added in the feature fusion stage to fully encode the feature information to achieve adaptive calibration and fusion. Finally, high-quality fused features of three different scales can be obtained, which helps to improve the accuracy and speed of ice jam semantic segmentation. The detection head and the segmentation head respectively output the bounding box detection results and the full-image segmentation results of each instance. Finally, the output network performs non-maximum suppression on the bounding box detection results to obtain the final bounding box, and then crops the full-image segmentation results based on the final bounding box to obtain the final segmentation results of each instance. Performing non-maximum suppression on the bounding box detection results can effectively improve the detection accuracy of the model for dense targets. In summary, this method effectively overcomes the problem of poor segmentation effect of the YOLOv8 framework on complex object edges and achieves an effective balance between the speed and accuracy of ice jam semantic segmentation.
[0061] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, they are all within the protection scope of this application; adding insignificant modifications or introducing insignificant designs to the algorithm or process, but without changing the core design of its algorithm and process, are all within the protection scope of this application.
[0062] In one embodiment, to verify the effectiveness of a river ice semantic segmentation method (river ice semantic segmentation model, hereinafter referred to as UiceNet or OURS for short) based on the YOLOv8 instance segmentation framework proposed in this application, we used the datasets NWPU_YRCC2, NWPU_YRCC_MS, and Albert River Ice Segmentation for training and testing, and compared with other methods.
[0063] NWPU_YRCC2 is an aerial photography dataset of the Yellow River ice floes, which contains 1,220 training images and 305 test images. NWPU_YRCC_MS is a hyperspectral dataset of satellite remote sensing of the Yellow River, which contains 960 training images and 241 test images. Albert River Ice Segmentation is an aerial photography dataset of the Albert River, which contains 554 training images and 139 test images after processing.
[0064] The experimental results are as Figure 6 、 Figure 7 、 Figure 8 shown. Figure 6 is a schematic diagram of the experimental results of the comparative experiment conducted on the NWPU_YRCC2 dataset. Figure 7 is a schematic diagram of the experimental results of the comparative experiment conducted on the NWPU_YRCC_MS dataset. Figure 8 is a schematic diagram of the experimental results of the comparative experiment conducted on the Albert River Ice Segmentation dataset. Among many commonly used algorithms or models, including convolutional and Transformer-based networks, this method outperforms traditional common algorithms on multiple datasets. The evaluation metric for the semantic segmentation task is mIoU, which is the mean of the intersection over union (IoU) and is used to measure the performance of the semantic segmentation model. In the comparison with convolutional and Transformer-based semantic segmentation models, UiceNet shows very competitive results, achieving 87.29% mIoU on NWPU_YRCC_MS, 92.69% mIoU on NWPU_YRCC2, and even 94.00% mIoU on Albert River Ice Segmentation.
[0065] Meanwhile, when UiceNet performs inference on 768px×768px images using a single NVIDIA Tesla V100 GPU, the running speed can reach 14 ms per input image, that is, it can smoothly process a video stream of 71 FPS.
[0066] The above experimental results show that UiceNet can balance accuracy and inference speed.
[0067] Another embodiment of the present application proposes an electronic device, and its specific structure can be as Figure 9 shown, including: at least one processor 21; and a memory 22 communicatively connected to the at least one processor 21; wherein, the memory 22 stores instructions executable by the at least one processor 21, and the instructions are executed by the at least one processor 21 to enable the at least one processor 21 to execute a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework as described in the above method embodiment.
[0068] Among them, the memory and the processor are connected by a bus. The bus can include any number of interconnected buses and bridges, and the bus can connect various circuits of one or more processors and memories together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be an element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium.
[0069] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory can be used to store the data used by the processor when executing operations.
[0070] Another embodiment of the present application proposes a computer-readable storage medium storing a computer program, which when executed by a processor, can implement a river ice semantic segmentation method based on the YOLOv8 instance segmentation framework as described in the above method embodiment.
[0071] That is, those skilled in the art can understand that all or part of the steps in the above method embodiments can be completed by instructing relevant hardware through a program. This program is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROMs, RAMs, magnetic disks, or optical discs that can store program codes.
[0072] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application, and in practical applications, various changes can be made to them in form and details without departing from the spirit and scope of the present application.
Claims
1. A HeLing semantic segmentation method based on the YOLOv8 instance segmentation framework, characterized in that: include: The river ridge image to be segmented is input into the feature encoder for hierarchical feature extraction. The feature encoder contains a serialized combination of multiple composite convolution modules, and finally obtains features of three different scales; The features of three different scales are taken as original features and input into three independent channel attention modules in the channel attention network. Each channel attention module consists of two parallel branches, namely the average pooling branch and the maximum pooling branch. The average pooling branch is used to perform adaptive average pooling on the input features to extract channel features, and then perform nonlinear mapping through two KAN linear layers. The maximum pooling branch is used to perform adaptive maximum pooling on the input features to extract channel features, and then perform nonlinear mapping through two KAN linear layers. The outputs of the two branches are added and activated by the Sigmoid activation function layer, and then multiplied with the corresponding original features to achieve adaptive calibration and fusion, and finally obtain three fusion features of different scales; The fused features of three different scales are input into three different detection heads to detect the bounding box and category of the instance, and the bounding box detection result and classification result of each instance are obtained; The fusion features of three different scales are input into the segmentation head. The segmentation head first generates the prototype segmentation result according to the fusion feature with the largest scale, and then generates the segmentation result of each channel according to the fusion features of three different scales. Finally, the prototype segmentation result and the segmentation result of each channel are connected together to obtain the segmentation result of the whole image. The output network performs maximum suppression on the bounding box detection results to obtain the final bounding box, and then crops the full image segmentation result based on the final bounding box to obtain the final segmentation result of each instance.
2. The HeLing semantic segmentation method based on the YOLOv8 instance segmentation framework according to claim 1, characterized in that: The feature encoder consists of five serially connected composite convolution modules. The input of the first composite convolution module is the river image to be segmented, and the input of the next composite convolution module is the output of the previous composite convolution module. Each composite convolution module consists of an initial convolution unit and a C2f unit; The initial convolution unit uses a 3×3 convolution with a downsampling strategy of step size 2. By dynamically setting the convolution kernel size and padding size, the input scale is reduced to 1 / 2 of the original scale to obtain preliminary dimensionality reduction features. The C2f unit expands the number of channels of the initial dimensionality reduction features to 2×e×C through 1×1 convolution, where C is the target output channel number and e is the channel expansion coefficient. The expanded features are then evenly decomposed into main branch features and secondary branch features through dimensional segmentation. The secondary branch features are nonlinearly transformed through the Bottleneck layer, and the main branch features and the secondary branch features after nonlinear transformation are concatenated in the channel dimension, and then output after feature fusion through 1×1 convolution; The feature encoder uses the features output by the last three composite convolution modules as features of three different scales, and the scales of the three different scales of features are 1 / 8, 1 / 16, and 1 / 32 of the river ice image to be segmented, respectively.
3. The HeLing semantic segmentation method based on the YOLOv8 instance segmentation framework according to claim 2, characterized in that: The Bottleneck layer adopts a compression-expansion structure. It first uses 3×3 convolution to compress the number of channels to e×C, then uses 3×3 convolution to restore the number of channels to the original number of channels, and realizes feature reuse through residual connection, and finally realizes nonlinear transformation of side branch features. The composite convolution module ensures the size consistency of the convolution operation through a dynamic kernel filling algorithm. For convolution kernels with a dilation rate d greater than 1, their equivalent kernel size is automatically expanded to d×(k-1)+1, where k is the size of the original convolution kernel. In the Bottleneck layer, the residual connection is enabled when and only when the number of input channels is equal to the number of output channels. Each convolution in the composite convolution module adopts a grouped convolution strategy, and the number of groups g is adjustable.
4. The HeLing semantic segmentation method based on the YOLOv8 instance segmentation framework according to claim 3, characterized in that: For the channel attention module, its average pooling branch consists of an average pooling unit, two average pooling KAN linear layers and an average pooling multilayer perceptron, and its maximum pooling branch consists of a maximum pooling unit, two maximum pooling KAN linear layers and a maximum pooling multilayer perceptron; In the average pooling branch, the average pooling unit is used to perform adaptive average pooling on the input features to extract channel features, the first average pooling KAN linear layer is used to perform channel compression, the second average pooling KAN linear layer is used to restore the channel dimension, and the average pooling multilayer perceptron is used to generate the average pooling channel weight matrix based on the output of the second average pooling KAN linear layer; In the maximum pooling branch, the maximum pooling unit is used to perform adaptive maximum pooling on the input features to extract channel features, the first maximum pooling KAN linear layer is used to perform channel compression, the second maximum pooling KAN linear layer is used to restore the channel dimension, and the maximum pooling multilayer perceptron is used to generate the maximum pooling channel weight matrix based on the output of the second maximum pooling KAN linear layer; The average pooling channel weight matrix and the maximum pooling channel weight matrix are added and activated through the Sigmoid activation function layer to obtain the final channel weight matrix. Finally, the final channel weight matrix is multiplied channel by channel with the corresponding original features to obtain a fusion feature of the same scale as the original feature.
5. The HeLing semantic segmentation method based on the YOLOv8 instance segmentation framework according to claim 4, characterized in that: The three detection heads have the same structure, which is composed of a bounding box detection branch and a category detection branch. The structures of the bounding box detection branch and the category detection branch are the same, which are composed of two convolutional layers and a standard two-dimensional convolutional layer. The bounding box detection branch and the category detection branch output the bounding box detection results and classification results of each instance respectively.
6. The method for semantic segmentation based on the YOLOv8 instance segmentation framework according to claim 5, characterized in that: The segmentation head consists of a ProtoNet network, a segmentation coefficient generation network, and an output layer; The ProtoNet network transforms the channel dimension of the largest fusion feature through a 3×3 convolution layer, improving the nonlinear expression capability while maintaining the integrity of the feature space. It then performs feature upsampling through a learnable transposed convolution layer, and finally completes the prototype recognition through two convolution layers again to generate the prototype segmentation result. The segmentation coefficient generation network consists of three scale branches and a 1×1 convolutional layer. Each scale branch is equipped with a feature refinement module consisting of two cascaded 3×3 convolutional layers. The first 3×3 convolutional layer is used to compress the number of channels to 1 / 4 of the original number of channels. The second 3×3 convolutional layer is used to optimize the feature representation through learnable parameters. Finally, the output of each scale branch is uniformly mapped to a 32-dimensional segmentation coefficient space through a 1×1 convolutional layer, and the multi-scale segmentation coefficients are spliced along the spatial dimension through a dimensional reorganization operation to form a complete mask coefficient matrix. The output layer is used to perform matrix multiplication operation on the prototype segmentation result and the mask coefficient matrix to dynamically generate the precise segmentation mask of each instance, thereby obtaining the full image segmentation result.
7. The method for semantic segmentation based on the YOLOv8 instance segmentation framework according to claim 6, characterized in that: The output network uses the SoftNMS maximum suppression algorithm to perform maximum suppression on the bounding box detection results to obtain the final bounding box. When the SoftNMS maximum suppression algorithm detects that the IoU of the bounding box is greater than the preset threshold, it reduces the confidence of the bounding box, but does not delete the bounding box whose IoU is greater than the preset threshold.
8. A HeLing semantic segmentation method based on the YOLOv8 instance segmentation framework according to any one of claims 1 to 7, characterized in that: The feature encoder, channel attention network, three different detection heads, segmentation head and output network together constitute the Heling semantic segmentation model. When training the Heling semantic segmentation model, the polygon key point alignment algorithm is used to align the true value data corresponding to the training samples, and harmless interpolation is performed on the basis of retaining the original points as much as possible to calculate the loss value. The execution process of the polygon key point alignment algorithm includes: Initialize the data structure. First, build a maximum heap structure, traverse the original boundary segment set, encapsulate each segment into an object containing four attributes: segment index, start and end point coordinates, initial key point number, and segment length. Then store them in the maximum heap in descending order of weight value to form a priority processing queue. Dynamic key point allocation, calculate the current total number of key points, extract the longest weighted line segment at the top of the heap, increase the number of key points of the longest weighted line segment at the top of the heap by 1, update the weight of the longest weighted line segment at the top of the heap and put it back into the heap, and simultaneously increase the total number of key points, and execute the dynamic key point allocation operation in a loop to ensure that the long edge gets more key points first, so as to achieve adaptive maintenance of the boundary shape; Ordered key point reconstruction: After sorting the line segments in the heap according to the original index, key points are generated. First, the starting point of the original line segment is retained, and then new points are evenly inserted on each line segment. The interpolation scale factor of each point is t = i / (pc-1), where i is the index value of the point, pc is the total number of points, and t is the interpolation scale factor. The coordinates are generated by linear interpolation. After the interpolation is completed, the original boundary end point coordinates are finally appended; The process of generating coordinates through linear interpolation is expressed by the formula: x'= x_start + t×(x_end-x_start); y'= y_start + t×(y_end-y_start); Among them, x_start and y_start represent the starting point of the line segment in length and width respectively, x_end and y_end represent the ending point of the line segment in length and width respectively, and x' and y' represent the new coordinates obtained after interpolation; After the alignment is completed, the bounding box detection results of the training samples output by the detection head are used to crop the full-image segmentation results of the training samples output by the segmentation head to obtain the final segmentation results of each instance of the training sample, and then cross entropy is performed with the aligned true value data to calculate the loss value. Based on the loss value, the back propagation algorithm is used to update the parameters until the preset convergence conditions are met, and the trained Heling semantic segmentation model is obtained.
9. An electronic device, characterized in that: include: at least one processor; and, a memory communicatively coupled to the at least one processor; Wherein, the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the Heling semantic segmentation method based on the YOLOv8 instance segmentation framework as described in any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it can implement the Heling semantic segmentation method based on the YOLOv8 instance segmentation framework as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Yellow River ice semantic segmentation method based on multi-attention mechanism double-flow fusion network
CN111160311A
Real-time semantic segmentation method for unmanned aerial vehicle aerial image of Yellow River ice
CN114943835A
Instance segmentation method based on detection enhancement and multi-stage bounding box feature refinement
CN115797629A
River ice distribution intelligent extraction method, device and equipment and storage medium
CN116012738A
Semantic segmentation method, device and equipment based on double-branch feature fusion
CN116229056A