A complex scene detection method combining boundary perception and semantic segmentation

By combining complex scene detection methods with boundary perception and semantic segmentation, a dual-flow detection network is built, which solves the problems of uneven brightness and similarity of object reflection in inland river environments, and realizes high-precision and high-speed water shoreline detection, which is suitable for real-time applications of unmanned ships.

CN115131321BActive Publication Date: 2025-06-06HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210777680.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-04
Publication Date
2025-06-06
Estimated Expiration
2042-07-04

AI Technical Summary

Technical Problem

In the inland environment, existing unmanned ships have difficulty in effectively dealing with the uneven brightness problems caused by trees and grass and building shadows on the river bank. In the inland environment, the reflection similarity between objects and water surfaces is high, resulting in low detection accuracy and difficulty in deploying on unmanned ships to operate in real time.

Method used

Using complex scene detection methods combining boundary perception and semantic segmentation, a dual-flow detection network is built, and features are extracted using the pre-trained ResNet-18 network, and a boundary perception network and a decoder network are introduced. Combined with global average pool and maximum pool operation, a double loss function is designed to optimize boundary learning and achieve high-precision detection of water shorelines.

Benefits of technology

This method can effectively detect water shorelines under different structures and lighting conditions, improve detection accuracy, and have a high detection speed, which is suitable for real-time applications of unmanned ships.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115131321B_ABST
    Figure CN115131321B_ABST
Patent Text Reader

Abstract

The present invention discloses a complex scene detection method combining boundary perception and semantic segmentation, comprising the following steps: S1, creating an image data set; S2, preprocessing the image data in the data set; S3, designing a detection network; S4, training the detection network to obtain an optimal convergence model; S5, loading the optimal convergence model obtained by training, inputting the image to be predicted into the prediction network for prediction to obtain a segmentation map; S6, post-processing the segmentation map to obtain the final water shoreline. This method solves the problem of detection failure caused by the similarity between the object and the water surface reflection. This method can effectively detect water shorelines under different structures and different lighting conditions, and has a high detection speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of visual image processing, and in particular to a complex scene detection method combining boundary perception and semantic segmentation. Background Art

[0002] Unmanned boats are autonomous surface robots that can navigate on the water according to preset tasks without remote control and with the help of their own sensors. Their ability to perform dangerous and time-consuming tasks is becoming increasingly prominent. Due to the huge demand from the business, scientific and environmental communities, researchers have accelerated the development of unmanned boat applications, such as hydrographic surveying and mapping, water quality monitoring, floating waste removal and water search and rescue operations. The water shoreline of inland rivers is equivalent to the sea-sky line detected by the sea surface environment, which is of great significance.

[0003] In recent years, with the development of machine learning theory and computer equipment, various machine learning-based methods have been proposed for the environmental perception of unmanned ships. However, most of the existing unmanned ship vision research is based on the marine environment. Compared with coastal and marine unmanned ships, unmanned ships in inland environments are more closely related to human life and have great potential value.

[0004] However, compared to the ocean, the inland river environment is more complex. The shadows of trees, grass, and buildings on the river bank will cover the water surface, causing the brightness value of the water surface to be significantly reduced. The uneven brightness distribution will inevitably reduce the accuracy of segmentation. Therefore, coastal and ocean unmanned ships cannot be directly applied to inland river shoreline detection. In addition, the shoreline detection of inland rivers is different from the ocean sea-sky line detection task. The similarity between objects on the river bank and the water surface reflection makes it difficult to accurately detect the water boundary. The complexity of the inland river environment makes it difficult to establish a reliable shoreline detection network. In addition, existing methods are difficult to deploy on unmanned ships for real-time operation. In summary, there is an urgent need for a complex scene detection method that combines boundary perception and semantic segmentation. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention proposes a complex scene detection method that combines boundary perception and semantic segmentation. The method solves the problem of detection failure caused by the similarity between objects and water surface reflections. The method can effectively detect water shorelines with different structures and under different lighting conditions, and has a high detection speed.

[0006] In order to solve the above technical problems, the technical solution of the present invention is:

[0007] A complex scene detection method combining boundary perception and semantic segmentation includes the following steps:

[0008] S1, create an image dataset;

[0009] S2, preprocessing the image data in the data set;

[0010] S3. Build a detection network. The detection network adopts a dual-stream structure and uses a deep neural network structure built by the pytorch framework. The detection network includes an encoder network, a decoder network, and a boundary perception network.

[0011] The encoder network is used to extract features. The encoder network uses a pre-trained ResNet-18 network, which can quickly downsample feature information to obtain a large receptive field. Four convolutional blocks are used to downsample the image, and each downsampling obtains feature maps of different scales.

[0012] The boundary-aware network is used to focus on the boundaries of water areas and discard a lot of redundant information. The boundary-aware network promotes the conversion of semantic information into boundary information through boundary convolution.

[0013] The decoder network fuses features of different sizes with boundary information, proposes global average pooling and maximum pooling operations, and refines them into the final segmentation result;

[0014] S4. Train the detection network to obtain the optimal convergence model

[0015] S4-1. Design of dual loss function

[0016] In the dual loss function, the semantic loss uses standard cross entropy, and the boundary loss uses binary cross entropy and dice loss to jointly optimize boundary learning. The expression is as follows:

[0017] L boundary (p d ,g d )=λ 1 L dice (p d ,g d )+λ 2 L bce (p d ,g d )

[0018]

[0019] Among them, L boundary represents the boundary loss, H and W represent the length and width of the image in the dataset, respectively, and L bce represents the binary cross entropy loss, L dice represents the dice loss, p d ∈R H×W represents the predicted boundary map, g d ∈R H×W represents the ground-truth boundary map obtained from the labeled segmentation mask, where i represents the i-th pixel and λ 1 , 2, c is a hyperparameter;

[0020] S4-2, determine the hyperparameters and use the Adam optimizer to train the detection network to minimize the loss of the detection network, that is, the optimal convergence model;

[0021] S5, loading the optimal convergence model obtained through training, and inputting the image to be predicted into the optimal convergence model for prediction to obtain a segmentation map;

[0022] S6. Post-process the segmented image to obtain the final shoreline.

[0023] Preferably, the dataset created in step S1 uses the USV Inland dataset as the shoreline detection dataset.

[0024] Preferably, the method for preprocessing the image data is: increasing the quantity and diversity of images by a data enhancement method, wherein the data enhancement method includes cropping, mirroring, brightness and contrast adjustment.

[0025] Preferably, the network constructed in step S3 includes an encoder network, a boundary perception network and a decoder network, the encoder network includes 4 sequentially connected convolutional layers and 1 pooling layer, the boundary perception network includes 2 boundary convolution blocks, the decoder network includes an attention refinement module, a feature fusion module and upsampling, the first and third convolutional layers in the encoder network are respectively inserted with two boundary convolution blocks to generate a boundary feature map, the attention refinement module is respectively connected to the pooling layer and the fourth convolutional layer to obtain semantic information, the feature fusion module reads the boundary feature map and semantic information for fusion, and outputs the segmentation map through upsampling.

[0026] Preferably, the pooling layer used at the end of the encoder network to increase the receptive field adopts global average pooling. Average pooling can summarize spatial information, and maximum pooling can obtain more refined channel attention.

[0027] Preferably, the post-processing method in step S6 is to extract the shoreline in the segmentation image through an edge detection algorithm.

[0028] The present invention has the following characteristics and beneficial effects:

[0029] The above technical solution is adopted to extract features using a lightweight backbone network to meet the real-time requirements of USV. The present invention designs an attention refinement module activated by global average pooling and maximum pooling and a boundary perception module that focuses on the feature information of the region of interest, thereby improving the detection accuracy of shadow and reflection areas. In addition, a feature fusion module is introduced to better fuse the feature information in the decoder network. In addition, the present invention designs a boundary perception loss function, which allows the network to focus on the scene boundary and strengthen the network to learn detailed information about the region of interest. This method can produce clearer predictions at the boundary of the water area and improve the performance of water shoreline detection. Compared with the traditional sea-sky line detection method, the unmanned ship water shoreline detection method proposes a new detection idea, which obtains the category information of pixel points through image semantic segmentation, and then corrects the boundary pixels in combination with boundary perception. Finally, the segmentation result map is post-processed to obtain the water shoreline. This detection method can effectively avoid the misdetection caused by interference such as shadows and reflections while maintaining good detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0031] Figure 1 Flowchart of complex scene detection method combining boundary perception and semantic segmentation;

[0032] Figure 2 Schematic diagram of the water and shoreline detection network structure;

[0033] Figure 3 Boundary convolution block, attention refinement module and feature fusion module. DETAILED DESCRIPTION

[0034] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.

[0035] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and the like are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, features defined as "first", "second", and the like may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "multiple" means two or more.

[0036] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood by specific circumstances.

[0037] The present invention provides a complex scene detection method combining boundary perception and semantic segmentation, such as Figure 1 As shown, the following steps are included:

[0038] S1. Create an image dataset. The created dataset uses the USVInland dataset as the shoreline detection dataset. This dataset is the first multi-sensor dataset for inland waterway unmanned ships. In the water segmentation sub-dataset, there are 518 relatively low-resolution (640x320) and 182 relatively high-resolution (1280x640) images, all of which are pixel-by-pixel labeled as water and non-water pixels. The USVInland dataset contains most of the complex scenes, such as water surface reflections, object mirrors, fog or mist. Unlike most datasets for seawater segmentation (where the sea-sky line is mainly straight lines), USVInland covers inland waterways with various waterline structures;

[0039] S2. Preprocess the image data in the dataset.

[0040] Specifically, the image data preprocessing method is: increasing the quantity and diversity of images by a data enhancement method, and the data enhancement method includes cropping, mirroring, brightness and contrast adjustment.

[0041] It can be understood that through the above data enhancement method, more labeled data can be obtained, which greatly reduces the time of manual annotation. At the same time, data enhancement is also one of the important methods to enhance the generalization ability of network models.

[0042] S3. Build a detection network. The detection network adopts a dual-stream structure and uses a deep neural network structure built by the pytorch framework. The detection network includes an encoder network, a decoder network, and a boundary perception network.

[0043] The encoder network is used to extract features. The encoder network uses a pre-trained ResNet-18 network, which can quickly downsample feature information to obtain a large receptive field. Four convolutional blocks are used to downsample the image, and each downsampling obtains feature maps of different scales.

[0044] The boundary-aware network is used to focus on the boundaries of water areas and discard a lot of redundant information. The boundary-aware network promotes the conversion of semantic information into boundary information through boundary convolution.

[0045] The decoder network fuses features of different sizes with boundary information, proposes global average pooling and maximum pooling operations, and refines them into the final segmentation result;

[0046] It should be noted that the detection network structure can be understood as Y (output) = x1 + x2 + x3^3 + ... xn, where x is a parameter. If we train these parameters ourselves, we will randomly initialize them, that is, we will give them a random value. This effect will definitely not be good, and then we will continue to train and iterate to get a good parameter value. Pre-training is a network where many people use a large amount of data experiments to conclude that this set of data has a better general effect. You can use it directly in the future. Continuing training on the basis of pre-training can get better results. The tasks of each network are different. The general effect of pre-training is better than the effect of randomly initialized parameters, but it needs to be modified for specific tasks.

[0047] S4. Train the detection network to obtain the optimal convergence model

[0048] S4-1. Design a dual loss function and use it to evaluate the error of the network detection results

[0049] In the dual loss function, the semantic loss uses standard cross entropy, and the boundary loss uses binary cross entropy and dice loss to jointly optimize boundary learning. The expression is as follows:

[0050] L boundary (p d ,g d )=λ 1 L dice (p d ,g d )+λ2 L bce (p d ,g d )

[0051]

[0052] Among them, L boundary represents the boundary loss, H and W represent the length and width of the image in the dataset, respectively, and L bce represents the binary cross entropy loss, L dice represents the dice loss, p d ∈R H×W represents the predicted boundary map, g d ∈R H×W represents the ground-truth boundary map obtained from the labeled segmentation mask, where i represents the i-th pixel and λ 1 , 2 , c is a hyperparameter;

[0053] S4-2, determine the hyperparameters and use the Adam optimizer to train the detection network to minimize the loss of the detection network, that is, the optimal convergence model;

[0054] Specifically, the network, loss function, optimizer and hyperparameters were determined. The network was trained 60 times with an initial learning rate of 5*10 -5 After 30 trainings, the learning rate is set to 2*10 -5 , the training batch size is 8.

[0055] S5, loading the optimal convergence model obtained through training, and inputting the image to be predicted into the optimal convergence model for prediction to obtain a segmentation map;

[0056] Furthermore, the pooling layer adopts global average pooling.

[0057] In the above technical solution, the image is downsampled by using a ResNet-18 network composed of 4 convolutional blocks to obtain feature maps of different scales. The target water area is a large area, and a large receptive field is required to recognize the entire water area. Therefore, a global average pool is added at the end of the network, which can provide the maximum receptive field of global semantic information.

[0058] Specifically, each of these four convolution blocks outputs a feature map of a different size. The input image passes through the first convolution block to the original Figure 1 The half-sized feature map is then passed through the second convolution block to obtain the original feature map. Figure 1 / 4 size feature map... and so on, the output feature maps of the four convolution blocks are 1 / 2, 1 / 4, 1 / 8, and 1 / 16 of the original image.

[0059] It is understandable that compared with the ResNet-101 network, the ResNet-18 network has fewer parameters and is more suitable for deployment on unmanned ships. In order to compensate for the loss of feature information caused by the lightweight network, this embodiment uses an additional boundary perception block to only focus on the water boundary, thereby discarding a large amount of redundant information. The boundary perception network uses a boundary convolution block, such as Figure 3 As shown in (a), it promotes the conversion of semantic information into boundary information.

[0060] It is conceivable that Figure 2 As shown in the figure, the boundary convolution layer is inserted at the first and third convolution layers in the encoder network to generate the boundary feature map. Then the true value of the boundary is used as the guidance of the boundary feature map to guide the network to learn the boundary features of the water area. Finally, the boundary information is fused with the semantic information to obtain a better segmentation effect.

[0061] The decoder network fuses features of different sizes with boundary information and refines them to obtain the final segmentation map. For better feature fusion, the fusion block is extended with additional modules based on the upsampling of the decoder network. Its architecture is based on a progressive fusion method, using an attention refinement module to learn the best fusion scheme between boundaries and different feature channels. In the attention mechanism, average pooling is generally used to summarize spatial information. Max pooling collects another important clue about the target feature and can obtain more refined channel attention. Therefore, the global average pooling and max pooling operations proposed in this invention, such as Figure 3 (b) is shown. More specifically, for unmanned ship shoreline detection, the global average pooling can better understand the water location information; while the maximum pooling operation can improve the local feature recognition of details such as water boundaries and edges. Using these two features can improve the representation ability of the network and obtain higher shoreline detection accuracy. Feature fusion module, such as Figure 3 As shown in (c), the semantic feature map extracted by the encoder network and the boundary detail features extracted by the boundary perception network are better fused. In the detection of unmanned inland waters, the water boundary information in the shallow features can correct the problem of blurred shoreline detection. By extracting the water boundary information of the shallow features as a guide for the deep features, the final segmentation result is more accurate. Reducing the redundant information in the shallow features for fusion reduces the computational overhead and increases the real-time performance of the system.

[0062] S6. Post-process the segmented image to obtain the final shoreline.

[0063] Among them, the post-processing method is to extract the water shoreline in the segmentation image through an edge detection algorithm.

[0064] As can be imagined, in this embodiment, three commonly used indicators in semantic segmentation networks are used to evaluate the performance of the method of the present invention: mean intersection over union (MIoU), F 1The score and accuracy (ACC) are defined as follows:

[0065]

[0066] Among them, k is the number of sample categories, TP represents the number of positive samples predicted as positive, FP represents the number of negative samples predicted as positive, and FN represents the number of positive samples predicted as negative. Since the method proposed in this study aims to predict high-quality water boundaries, we introduce another evaluation indicator, the boundary root mean square error (RMSE).

[0067]

[0068] Where gap[i] represents the pixel difference between the true edge and the predicted edge value in the vertical direction of the image.

[0069] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, various changes, modifications, substitutions and variations of these embodiments including components are made without departing from the principles and spirit of the present invention, and still fall within the scope of protection of the present invention.

Claims

1. A complex scene detection method combining boundary perception and semantic segmentation, It is characterized in that The steps include: S1. Create an image dataset; S2, preprocessing the image data in the data set; S3. Build a detection network. The detection network adopts a dual-stream structure and uses a deep neural network structure built by the pytorch framework. The detection network includes an encoder network, a decoder network, and a boundary perception network. The encoder network is used to extract features. The encoder network uses a pre-trained ResNet-18 network to quickly downsample feature information to obtain a large receptive field. Four convolutional blocks are used to downsample the image, and each downsampling obtains feature maps of different scales. The boundary-aware network is used to focus on the boundaries of water areas and discard a lot of redundant information. The boundary-aware network promotes the conversion of semantic information into boundary information through boundary convolution. The decoder network fuses features of different sizes with boundary information, proposes global average pooling and maximum pooling operations, and refines them into the final segmentation result; S4. Train the detection network to obtain the optimal convergence model S4-1. Design of dual loss function In the dual loss function, the semantic loss uses standard cross entropy, and the boundary loss uses binary cross entropy and dice loss to jointly optimize boundary learning. The expression is as follows: L boundary ( p d ,g d )−λ 1 L dice ( p d ,g d )+λ 2 L bce ( p d ,g d ) Among them, L boundary represents the boundary loss, H and W represent the length and width of the image in the dataset, respectively, and L bce represents the binary cross entropy loss, L dice represents the dice loss, p d ∈R H×W represents the predicted boundary map, g d ∈R H×W represents the ground-truth boundary map obtained from the labeled segmentation mask, where i represents the i-th pixel and λ 1 , 2 , c is a hyperparameter; S4-2, determine the hyperparameters and use the Adam optimizer to train the detection network to minimize the loss of the detection network, that is, the optimal convergence model; S5, loading the optimal convergence model obtained through training, and inputting the image to be predicted into the optimal convergence model for prediction to obtain a segmentation map; S6. Post-process the segmented image to obtain the final shoreline.

2. The complex scene detection method combining boundary perception and semantic segmentation according to claim 1, It is characterized in that The dataset created in step S1 uses the USV Inland dataset as the shoreline detection dataset.

3. The complex scene detection method combining boundary perception and semantic segmentation according to claim 1, It is characterized in that The method for preprocessing the image data is: increasing the quantity and diversity of images by a data enhancement method, wherein the data enhancement method includes cropping, mirroring, brightness and contrast adjustment.

4. The complex scene detection method combining boundary perception and semantic segmentation according to claim 1, It is characterized in that The network constructed in step S3 includes an encoder network, a boundary perception network and a decoder network. The encoder network includes 4 convolutional layers and 1 pooling layer connected in sequence. The boundary perception network includes 2 boundary convolution blocks. The decoder network includes an attention refinement module, a feature fusion module and upsampling. The first and third convolutional layers in the encoder network are respectively inserted with two boundary convolution blocks to generate a boundary feature map. The attention refinement module is respectively connected to the pooling layer and the fourth convolutional layer to obtain semantic information. The feature fusion module reads the boundary feature map and the semantic information for fusion, and outputs the segmentation map through upsampling.

5. The complex scene detection method combining boundary perception and semantic segmentation according to claim 1, It is characterized in that The pooling layer used at the tail of the encoder network to increase the receptive field adopts global average pooling.

6. The complex scene detection method combining boundary perception and semantic segmentation according to claim 1, It is characterized in that The post-processing method in step S6 is to extract the shoreline in the segmentation image through an edge detection algorithm.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method based on attention multi-scale feature fusion

    CN111127493A

  • RGBD image semantic segmentation method based on boundary attention

    CN113822284A