An Adaptive Perception Method and System for Marine Environment Based on Boundary Enhancement

Through the method of strengthening boundary feature learning in marine environment perception, a semantic-boundary multi-level reinforcement image understanding model is built, which solves the problem of insufficient boundary feature learning in marine environment perception, and achieves higher environmental perception accuracy and autonomous cruise capabilities of unmanned surface boats.

CN120071304BActive Publication Date: 2025-07-01SHANDONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510533490.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-01
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

The prior art ignores the learning of boundary features in images in marine environment perception, which makes it difficult for category perception accuracy to reach an ideal level, and it is difficult for unmanned surface boats to accurately plan routes and cruise in different mission scenarios.

Method used

A method of adaptive perception of the marine environment based on boundary reinforcement is proposed. By designing the boundary-semantic information fusion of images and supervision of boundary feature learning, a ship-borne camera image understanding model with semantic-boundary multi-level reinforcement is built, and a module such as boundary feature encoding stream, semantic feature encoding stream, boundary-semantic feature fusion stream and fusion feature decoder can be used to realize accurate perception of the marine environment.

Benefits of technology

It effectively improves the adaptability of unmanned surface boats to the maritime environment, enhances boundary perception, improves the accuracy of environmental semantic category classification, provides more reliable decision-making support, and provides more accurate and efficient support for unmanned ship path planning and autonomous cruise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071304B_ABST
    Figure CN120071304B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of autonomous cruising of unmanned surface vessels and marine environment perception, and specifically discloses a method and system for adaptive perception of marine environment based on boundary enhancement. In the method, an image understanding model of a shipborne camera with multi-level enhancement of semantics and boundaries is proposed. In the training stage, the original image and the boundary ground truth are respectively encoded through a semantic feature encoding stream and a boundary feature encoding stream, and feature interaction is carried out through a boundary-semantic feature fusion stream to enhance the boundary perception ability; the results of the encoding stream and the fusion stream are input into a fusion feature decoder, and after being enhanced by three-dimensional feature attention of direction, position and channel, the environmental category understanding result is output. The predicted output of the model generates a predicted boundary through boundary extraction, and a quantization index is constructed by combining the boundary ground truth in a boundary feature discriminator, and the model is optimized through negative feedback of a multi-task loss function. The present invention covers boundary information in different scenarios and improves the adaptability of the unmanned ship to the marine environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous cruising of unmanned surface vessels (USVs) and marine environment perception, and particularly relates to a marine environment adaptive perception method and system based on boundary enhancement. Background Art

[0002] With the continuous evolution of artificial intelligence and unmanned driving technologies, USVs have been increasingly widely used in practical applications. With their significant advantages of miniaturization and high scalability, USVs play an important role in key fields such as marine environment monitoring and water area inspection. When performing tasks, USVs usually carry multi-source sensors including cameras, collect images of marine scenes through the cameras, and perform environmental perception and understanding to generate category masks including different categories such as water surface, obstacles, sky, etc. These mask information provides key environmental perception capabilities for USVs, helping them accurately identify navigable waters and restricted areas, and achieve autonomous navigation and efficient task execution. In recent years, although the research on perception and understanding of the marine environment has been continuously improved, most studies have the problem of ignoring the learning of boundary features between various categories, resulting in the category perception accuracy being difficult to reach an ideal level.

[0003] In the research on environmental perception assisted by boundary features, in-depth explorations have been carried out in various fields. Bovcon et al. simulated the edge information of the horizontal plane with an inertial measurement unit (IMU) and fused it with semantic information; He et al. generated a pseudo water line using radar data and fused its feature encoding with semantic information. Both of these two studies have improved the category perception accuracy of the water surface edge to a certain extent, but in the process of obtaining boundary information, they both ignored the boundary features originally contained in the image, such as the boundary between the sea surface and marine obstacles, the boundary between the sky and shore obstacles, etc. In addition, in the field of unmanned vehicle technology, rapid progress has been made in the research on autonomous navigation route planning based on sensors such as on-vehicle cameras. However, due to the essential differences between the marine environment and the land environment, the related technologies of unmanned vehicles are difficult to be directly transferred and applied to USVs.

[0004] Currently, the mission scenarios of USVs at sea can be divided into two categories: (1) Inspection tasks in the coastal area near the sea, where the perception effect of shore buildings will directly affect the accuracy and reliability of USV path planning; (2) Environmental data collection tasks in the open sea area, where the understanding effect of obstacles on the sea surface will determine the accuracy of USV autonomous navigation. Thus, it can be seen that at the present stage, only using the edge features of the simulated sea level cannot effectively guide the USV to understand the images of on-board cameras in different mission scenarios, and thus cannot correctly guide the USV to perform route planning and autonomous cruising. Summary of the Invention

[0005] The object of the present invention is to propose an adaptive perception method for the marine environment based on boundary enhancement, which is designed through the fusion of boundary-semantic information in images and the supervision of boundary feature learning to cover boundary information in different scenarios, improve the adaptability of unmanned ships to the marine environment, and thus more effectively guide the unmanned ships to perform operations such as path planning.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] An adaptive perception method for the marine environment based on boundary enhancement, comprising the following steps:

[0008] Step 1. Obtain the original image dataset for marine environment perception, perform data augmentation on the original images, and obtain the boundary ground truth for each original image. The boundary ground truth and the original image correspond one by one to form a training dataset;

[0009] Step 2. Build an on-board camera image understanding model with multi-level semantic-boundary reinforcement, which includes a boundary feature encoding stream, a semantic feature encoding stream, a boundary-semantic feature fusion stream, and a fusion feature decoder;

[0010] The boundary feature encoding stream and the semantic feature encoding stream are respectively used to encode the boundary ground truth and the original image input into the model; each encoding module includes four cascaded feature extraction layers;

[0011] The boundary-semantic feature fusion stream includes four cascaded feature attention modules;

[0012] Each layer of the feature attention module is respectively used to receive the outputs of the corresponding layers of the boundary feature encoding stream and the semantic feature encoding stream, and is used to realize the feature interaction between the boundary features and the semantic features of the corresponding layer to enhance the boundary perception ability;

[0013] The outputs of each layer of the feature attention module are respectively fed back to the corresponding layer of the semantic feature encoding stream, and together with the output of the corresponding layer of the semantic feature encoding stream, are used as the input of the lower layer of the semantic feature encoding stream;

[0014] The output results of the boundary feature encoding stream, the semantic feature encoding stream, and the boundary-semantic feature fusion stream are input into the fusion feature decoder. After being enhanced by three-dimensional feature attention in terms of direction, position, and channel, the classification result of marine environment perception is output;

[0015] Step 3. Use the training dataset to train the on-board camera image understanding model with multi-level semantic-boundary reinforcement; input the image data of the marine environment perception obtained in real time into the trained model to obtain the adaptive perception result of the marine environment.

[0016] In addition, based on the above-mentioned method for adaptively sensing the marine environment based on boundary reinforcement, the present invention also proposes a corresponding adaptively sensing system for the marine environment based on boundary reinforcement, which adopts the following technical solutions:

[0017] A marine environment adaptive perception system based on boundary reinforcement, carried on a USV, includes a camera and a computer device; the camera is used to obtain image data of marine environment perception in real time and upload it to the computer device;

[0018] The computer device includes a memory and a processor; executable code is stored in the memory, and when the processor executes the executable code, it is used to implement the steps of adaptive perception of the marine environment based on boundary reinforcement as described above.

[0019] The present invention has the following advantages:

[0020] As described above, the present invention relates to a method and system for adaptively perceiving a marine environment based on boundary reinforcement. In order to achieve more accurate environmental perception in the multi-scenario application of unmanned ships at sea, the present invention proposes a semantic-boundary multi-level reinforced shipborne camera image understanding model. During training, the model realizes multi-branch and multi-level fusion and reinforcement of boundary features and semantic information at the input level, encoding, decoding level, and output level, effectively solving the problem of inaccurate classification of environmental semantic information due to fuzzy category boundaries in nearshore and offshore environments. In addition, in order to solve the problem of missing boundary information, at the input level, the present invention introduces the boundary truth value of the image as input into the model during the training stage, and encodes it in parallel with the original image. At the encoding level, the present invention designs a boundary-semantic feature fusion flow in combination with the attention mechanism, realizes the interaction and enhancement of the original image features (i.e., semantic features) and the boundary features in the layer-by-layer encoding process, and gives the model a deep semantic understanding ability and boundary sensitivity. In addition, at the decoding level, in order to effectively integrate semantics, boundaries, and semantic-boundary features, the present invention also designs a multi-branch cross-channel attention module, which complements the features at the channel, direction, and position levels, and synergistically improves the accuracy of the environmental perception category results. At the output level, the present invention designs a boundary feature identifier, which dynamically adjusts the weight distribution of boundary features by quantifying the differences in boundary features, thereby improving the model's sensitivity and understanding of boundary information. The model of the present invention can effectively learn the boundary features in shipborne camera images, help unmanned ships adaptively understand the implicit boundary information of images in multiple application scenarios, and provide more reliable decision support for subsequent unmanned ship path planning by accurately and efficiently guiding the classification of environmental semantic categories, thereby effectively improving the accuracy and execution efficiency of path planning. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1Flow chart of adaptive perception of marine environment based on boundary enhancement in the embodiment of the present invention;

[0022] Figure 2 Structural diagram of shipborne camera image understanding model with semantic-boundary multi-level enhancement in the embodiment of the present invention;

[0023] Figure 3 Structural diagram of boundary-semantic feature fusion flow in the embodiment of the present invention;

[0024] Figure 4 Structural diagram of feature attention module in the embodiment of the present invention;

[0025] Figure 5 Structural diagram of multi-level attention fusion flow in the fusion feature decoder in the embodiment of the present invention;

[0026] Figure 6 Distribution diagram of average value of sample similarity scores in the embodiment of the present invention. Detailed implementation manners

[0027] The present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners:

[0028] Embodiment 1

[0029] This embodiment describes a method for adaptive perception of marine environment based on boundary enhancement. In this method, a shipborne camera image understanding model with semantic-boundary multi-level enhancement is proposed, which includes a boundary feature encoding flow, a semantic feature encoding flow, a boundary-semantic feature fusion flow, a fusion feature decoder, a boundary feature discriminator, etc. Among them, in the training stage, the present invention adopts a dual-channel feature interaction architecture, that is, the original image and the boundary ground truth are respectively encoded through the semantic feature encoding flow and the boundary feature encoding flow, and feature interaction is performed through the boundary-semantic feature fusion flow to enhance the boundary perception ability; the encoding results of the boundary feature encoding flow, the semantic feature encoding flow, and the boundary-semantic feature fusion flow are input into the fusion feature decoder, and after being enhanced by three-dimensional feature attention in the direction, position, and channel in the fusion feature decoder, the environmental category understanding result is output. In addition, in the model training stage, the present invention also generates a predicted boundary by boundary extraction of the predicted output of the model, constructs a quantization index in the boundary feature discriminator in combination with the boundary ground truth, and optimizes the model through negative feedback of the multi-task loss function. In the model inference stage, only the semantic feature encoding flow and the fusion feature decoder are retained, and the optimized boundary feature parameters in the training stage are used to improve the environmental category understanding accuracy while reducing the number of parameters, effectively guiding the route planning and autonomous cruising of the unmanned ship.

[0030] As Figure 1 shown, the method for adaptive perception of marine environment based on boundary enhancement in this embodiment includes the following steps:

[0031] Step 1. Obtain the original image dataset for maritime environment perception, perform data augmentation on the original images, and obtain the ground truth boundaries for each original image. The ground truth boundaries and the original images are paired one by one to form a training dataset.

[0032] In this embodiment, the original image dataset used is, for example, the MaSTr1325 dataset commonly used in the field of maritime environment perception, with a quantity of 1325 images, covering image data of various time periods within a day and diverse weather conditions.

[0033] Aiming at the problems of insufficient sample size in the MaSTr1325 dataset and differences in imaging parameters of different models of cameras, the present invention performs data augmentation by adjusting image attributes such as exposure, contrast, saturation, and hue, expanding the sample quantity while maintaining the original semantic information, and effectively improving the generalization ability and environmental adaptability of the model to different imaging conditions.

[0034] Specifically, the exposure (exposure adjustment coefficient k ∈ [0.6, 1.4]) and contrast (contrast factor c ∈ [0.8, 1.2]) are adjusted through the linear transformation clip function in python. The saturation (0.7 - 1.3 times) is controlled by adjusting the S channel and the hue is fine-tuned (±15°) by adjusting the H channel in the HSV space of the picture. For each image, enhanced versions are generated for single-parameter adjustment (2 intensities for each of the 4 types of parameters), two-parameter combinations (2 intensities for each of the 6 combinations), three-parameter combinations (2 intensities for each of the 4 combinations), and four-parameter combinations (2 intensities for 1 combination). Finally, the improvement from 1325 original images to 41075 expanded samples is achieved.

[0035] Through the above data augmentation method, the data deviation is effectively alleviated and the generalization ability of the model is enhanced. The ground truth of the MaSTr1325 dataset is a grayscale image with pixel-by-pixel class annotations. The dataset includes:

[0036] ① The color image [original image] taken by the camera, which has three RGB channels:

[0037] ② The ground truth for classifying each pixel of the original image, which has only one channel, and the value of each pixel is 1, 2, or 3 to distinguish which class this pixel belongs to. The numbers 1, 2, and 3 represent water surface, sky, and obstacle respectively.

[0038] In order to obtain the ground truth boundaries, the present invention first converts the grayscale image into a multi-class RGB image through pseudo-color coding, then applies the Canny edge detection algorithm to extract the boundaries, and finally obtains the ground truth boundaries.

[0039] Step 2. Build an on-board camera image understanding model with multi-level reinforcement of semantics and boundaries.

[0040] AsFigure 2 As shown in Figure 2 , the shipborne camera image understanding model with multi-level semantic-boundary reinforcement established in this embodiment includes a boundary feature encoding stream, a semantic feature encoding stream, a boundary-semantic feature fusion stream, and a fusion feature decoder.

[0041] In this embodiment, the shipborne camera image (i.e., the original image) and the corresponding boundary ground truth are used as the data sources of the shipborne camera image understanding model, where the input data sizes are correspondingly denoted as 3×h×w and 1×h×w.

[0042] The boundary feature encoding stream and the semantic feature encoding stream are respectively used to encode the boundary ground truth and the original image input into the model; each encoding module includes four cascaded feature extraction layers.

[0043] The boundary-semantic feature fusion stream includes four cascaded feature attention modules.

[0044] Each layer of the feature attention module is respectively used to receive the outputs of the corresponding layers of the boundary feature encoding stream and the semantic feature encoding stream, and is used to implement the feature interaction between the boundary features and the semantic features of the corresponding layer to enhance the boundary perception ability.

[0045] The outputs of each layer of the feature attention module are respectively fed back to the corresponding layer of the semantic feature encoding stream, and together with the output of the corresponding layer of the semantic feature encoding stream, are used as the input of the lower layer of the semantic feature encoding stream.

[0046] The output results of the boundary feature encoding stream, the semantic feature encoding stream, and the boundary-semantic feature fusion stream are input into the fusion feature decoder. After being enhanced by three-dimensional feature attention in the direction, position, and channel, an RGB image is output as the classification result of the marine environment perception, so that the same type of targets present consistent RGB color coding in this image.

[0047] Since the boundary ground truth features are relatively simple, the boundary feature encoding stream uses a ResNet18 network with fewer parameters; the first to fourth residual convolutional layers of the ResNet18 network are used as the four-layer feature extraction layers of the boundary feature encoding stream.

[0048] The output of each residual convolutional layer is respectively input into the boundary-semantic feature fusion stream to realize the gradual interaction between the boundary ground truth and the semantic features, promote bidirectional feature learning, and perform spatial downsampling based on the output size of the last residual convolutional layer. Finally, the multi-scale boundary feature maps formed by each residual convolutional layer are spliced and used as the output of the boundary feature encoding stream.

[0049] The boundary feature encoding stream is conducive to enhancing the network's understanding of boundary details through the splicing of multi-scale boundary feature maps.

[0050] In this embodiment, the basic component unit of the boundary feature encoding stream is the residual convolutional layer.

[0051] Each residual convolutional layer is composed of several convolutional layers. In the structure of ResNet-18, the conv1 layer contains a convolutional layer with a kernel size of 7×7 and a stride of 2; conv2_x contains 2 residual blocks, each residual block is composed of two convolutional layers with a kernel size of 3×3, and each layer outputs 64 channels; conv3_x contains 2 residual blocks, each residual block is composed of two convolutional layers with a kernel size of 3×3, and each layer outputs 128 channels; conv4_x contains 2 residual blocks, each residual block is composed of two convolutional layers with a kernel size of 3×3, and each layer outputs 256 channels. For the downsampling layer of conv2_x, it is designed as a convolutional layer with a kernel size of 8 and a stride of 8; for the downsampling layer of conv3_x, it is designed as a convolutional layer with a kernel size of 4 and a stride of 4; for the downsampling layer of conv4_x, it is designed as a convolutional layer with a kernel size of 2 and a stride of 2.

[0052] The semantic feature encoding stream is constructed based on the Conformer-S architecture, and the multi-head Transformer module adopted by this architecture contains a 6-head attention mechanism. Conformer combines the convolutional neural network and the self-attention mechanism of Transformer, giving full play to the advantages of local feature extraction and global context modeling, enabling the model to accurately capture complex dynamic features such as sea surface fluctuations and lighting changes, and at the same time effectively modeling long-distance dependencies, thereby enhancing the understanding and parsing ability of the marine environment.

[0053] The first four Conformer layers of the Conformer-S architecture are used as the four feature extraction layers of the semantic feature encoding stream; the output of each Conformer layer is respectively input into the boundary-semantic feature fusion stream to achieve the hierarchical interaction between the boundary ground truth and the semantic features, and promote two-way feature learning. The output of the third Conformer layer is used as the input of the feature loss function of.

[0054] At the same time, each Conformer layer also receives the feedback information of the corresponding layer of the boundary-semantic feature fusion stream.

[0055] The semantic feature encoding stream performs spatial downsampling based on the output size of the last Conformer layer. Finally, the multi-scale semantic feature maps formed by each Conformer layer are concatenated as the output of the semantic feature encoding stream.

[0056] By fusing multi-scale features through downsampling with a unified size, it effectively integrates shallow details (boundary positions) and deep semantics (category information), enhances the richness and robustness of feature representation, and at the same time avoids complex upsampling operations, improving the computational efficiency.

[0057] The basic building blocks of the semantic feature encoding stream are Conformer layers c1 to c4. C1 contains a convolutional layer with a 7×7 convolutional kernel and a stride of 2, followed by a 3×3 max-pooling layer with a stride of 2. In c2, the CNN Branch contains a module composed of a 1×1 convolution, a 3×3 convolution, and a 1×1 convolution in sequence; the Feature Coupling Unit (FCU) contains a 1×1 convolution; the Transformer Branch first passes through a 4×4 convolution with a stride of 4, and then connects to a module composed of a multi-head attention mechanism (MHSA-6), a 1×1 convolution, and a 1×1 convolution, and this module is repeated 1 time. In c3, the CNN Branch contains a module composed of a 1×1 convolution, a 3×3 convolution, and a 1×1 convolution in sequence; the FCU contains a 1×1 convolution; the Transformer Branch is composed of the module of MHSA-6, 1×1 convolution, and 1×1 convolution repeated 4 times. In c4, the CNN Branch contains a module composed of a 1×1 convolution, a 3×3 convolution, and a 1×1 convolution in sequence; the FCU contains a 1×1 convolution; the Transformer Branch is composed of the module of MHSA-6, 1×1 convolution, and 1×1 convolution repeated 3 times. For the downsampling layer of c2, it is designed as a convolutional layer with a convolutional kernel size of 8 and a stride of 8; for the downsampling layer of c3, it is designed as a convolutional layer with a convolutional kernel size of 4 and a stride of 4; for the downsampling layer of c4, it is designed as a convolutional layer with a convolutional kernel size of 2 and a stride of 2.

[0058] The original Conformer-S architecture and Resnet18 architecture have a total of five layers, and the output result size is c×7×7 at this time.

[0059] The present invention only adopts the first four layers of the two architectures, and the output result size is c×14×14. Reducing the number of layers to improve the resolution of the feature map can retain richer spatial details and alleviate the problem of information loss of small targets in the image.

[0060] As Figure 3 shown, the boundary-semantic feature fusion stream contains 4 cascaded feature attention modules. Each layer of the feature attention module is used to receive the outputs of the corresponding levels of the boundary feature encoding stream and the semantic feature encoding stream, and perform feature interaction. On the one hand, the processing results are fed back to the corresponding levels of the semantic feature encoding stream. On the other hand, each layer of the feature attention module in the boundary-semantic feature fusion stream is downsampled according to the spatial size of the last layer of the feature attention module, and finally the feature maps output by each layer of the feature attention module are concatenated in the channel dimension to form the multi-scale feature output of the boundary-semantic feature fusion stream.

[0061] During subsequent model inference, the overall input and output only pass through the semantic feature encoding stream and do not go through the boundary feature encoding stream and the boundary-semantic feature fusion stream. Therefore, during the model training process, during the fusion feature transfer process through the boundary-semantic feature fusion stream, it only interacts with the semantic feature encoding stream. This architecture enables the semantic feature encoding stream to synchronously learn boundary feature parameters during the encoding stage while reducing semantic feature interference during boundary encoding, thereby effectively improving the training efficiency.

[0062] The feature attention module designs an attention-based feature aggregation mechanism, and its structure is as Figure 4 shown.

[0063] First, the boundary feature and the semantic feature input into the feature attention module are added element-wise to form the combined feature one.

[0064] The combined feature one is divided into two branches. One branch is processed by introducing the ReLU activation function of non-linear transformation to enable the model to enhance the expression ability of the fine structure of the boundary. The other branch is aggregated with global semantic information through global average pooling and then activated through ReLU. Further, the abstract expression of the semantic feature is mined through non-linear operations to enhance the representation ability of the semantic feature for abstract semantic content.

[0065] The processing results of the two branches are added element-wise to obtain the combined feature two.

[0066] The combined feature two is multiplied element-wise with the combined feature one to generate the attention weight based on the post-interaction feature, thereby measuring the key degree of the boundary and semantic features. Then, they are multiplied with the boundary and semantic features respectively to specifically strengthen the effective information of the boundary details and semantic content. Finally, they are added element-wise to integrate the localization details of the boundary feature and the abstract understanding of the semantic feature to obtain the fusion feature, so that the fusion feature simultaneously has refined boundary information and high-level semantic representation.

[0067] The semantic information and boundary information at each level are processed through combined features, global average pooling, and activation operations to strengthen the key feature weights and achieve the efficient cross-fusion of semantic and boundary information. Therefore, more representative fusion features can be output.

[0068] The boundary-semantic feature fusion stream is mainly composed of four consecutive feature attention modules, which respectively receive the output results of the residual convolutional layers conv1, conv2_x, conv3_x, conv4_x from the boundary feature encoding stream, and the output results of the Conformer layers c1 to c4 from the semantic feature encoding stream as inputs.

[0069] For a single feature attention module, after receiving the boundary features and semantic features, first add the two element-wise to generate "joint feature one", and process it through the ReLU activation function; then apply a 1×1 convolution operation to the joint feature 1, and complete feature normalization with BatchNorm. After that, the feature processing is divided into two paths: one path performs global average pooling on the features to extract global semantic information, and then activates it through ReLU; the other path retains local feature details. The features of the two paths are fused through a multiplication operation, and "joint feature two" is generated through the Sigmoid activation function. Finally, through two multiplication interaction operations, the feature information of different levels is gradually fused, and finally the refined features that fuse the boundary and semantic information are output. The downsampling structure output by each stream: A 3×3 convolution with a stride of 2 can be selected to ensure that the spatial dimensions of the four streams are completely aligned after downsampling, and finally the four downsampled features are concatenated in the channel dimension.

[0070] The fused feature decoder includes a multi-level attention fusion stream and an atrous convolution pooling pyramid module, where Figure 5 shows the network structure of the multi-level attention fusion stream, including a multi-branch cross-channel attention module and co-attention.

[0071] After receiving the outputs of the boundary feature encoding stream, semantic feature encoding stream, and boundary-semantic feature fusion stream, the multi-level attention fusion module fuses the three features and enhances the model's understanding of the correlation between them.

[0072] The multi-level attention fusion module includes a multi-branch cross-channel attention module and a co-attention module.

[0073] The multi-branch cross-channel attention module uses a multi-branch cross-channel interaction structure to adaptively learn channel weights, establish a direct association between channels and their weights, and reflect the global contributions of semantic features and boundary features in terms of scale and type dimensions.

[0074] The co-attention module captures spatial features in two directions and generates an attention map to complementarily enhance the representation of the objects of interest in the input feature map, so as to solve the deficiencies of the multi-branch cross-channel attention module in processing direction and position information.

[0075] The above two attention modules complement each other and jointly improve the accuracy of the category perception results in the marine environment.

[0076] The atrous convolution pooling pyramid module is used to receive the outputs from the multi-level attention fusion stream and the semantic feature encoding stream, while expanding the receptive field and ensuring the resolution of the output image.

[0077] Finally, after adjusting the number of channels through a convolution layer with a kernel size of 1×1, the classification results of marine environment perception are output.

[0078] Figure 5 It is a structural diagram of a multi-level attention fusion stream, which includes a multi-branch cross-channel attention module and collaborative attention. As Figure 5 shown in the figure above, the inputs of the multi-branch cross-channel attention module are the results of the boundary feature encoding stream, the semantic feature encoding stream, and the boundary-semantic feature fusion stream , all with dimensions of C×H×W. First, global average pooling is performed on all results respectively, and the results are concatenated into a 3×C feature, which is processed by batch normalization and channel shuffle to generate feature p. Then, a convolution operation is performed on p and another circular feature q using a circular convolution (convolution kernel size k = 3), followed by processing through a Sigmoid activation function and an inverse shuffle operation to output three 1×C attention features. The initial input is multiplied by the obtained attention features channel by channel and concatenated in the channel dimension, and the final result is output to the collaborative attention module.

[0079] As Figure 5 shown in the figure below, the input of the collaborative attention module is , which first passes through two adaptive average pooling layers respectively for pooling in the height and width dimensions, reducing the height and width dimensions of the feature map to 1, and the results are set as the height and width attention feature maps and , with dimensions of C×H×1 and C×1×W respectively. Calculate the intermediate number of channels mip, which takes the larger value between the result of dividing the input number of channels C by the reduction rate r and 8. Swap the second and third dimensions of to obtain a height attention feature map with dimensions of C×1×H . Concatenate and at the channel level and pass through a convolutional layer with a convolution kernel size of 1×1 to reduce the number of channels of the input feature map from C to C / r. The result is input to a batch normalization layer to normalize the output of the convolutional layer, and then the h_swish activation function is used for activation. The feature map after convolution, normalization, and activation is split according to the height and width dimensions. Swap the width and height dimensions of the width attention feature map to restore the original dimension order. Perform 1×1 convolution on the height and width attention feature maps to expand the number of channels of the intermediate feature map from mip to C and convert it to an attention weight through the Sigmoid function. Multiply the input feature map by the height and width attention weights to obtain the final output feature map with dimensions of C×H×W.

[0080] The basic components of the dilated convolutional pooling pyramid include a 1×1 convolutional module, a multi-scale dilated convolutional module, and a pooling module. 1×1 convolutional module: It contains a convolutional layer with a kernel size of 1×1 and a stride of 1, followed by a batch normalization layer and a ReLU activation function. Multi-scale dilated convolutional module: It consists of multiple ASPPConv sub-modules. Each ASPPConv sub-module contains a dilated convolutional layer with a kernel size of 3×3, a stride of 1, and a padding size equal to the dilation rate corresponding to the sub-module, followed by a batch normalization layer and a ReLU activation function. The dilation rates of different sub-modules are specified by the input atrous_rates list, and the atrous_rates applied in this method are 1, 1, 4, 8. Pooling module: First, the input feature map is pooled to a size of 1×1 through adaptive average pooling, followed by a convolutional layer with a kernel size of 1×1 and a stride of 1, then a batch normalization layer and a ReLU activation function, and finally, the feature map is upsampled to the original size of the input feature map using bilinear interpolation. Finally, it is output through a convolutional layer with a kernel size of 1×1 and 3 channels as the final output of the model.

[0081] Step 3. Use the training dataset to train the shipborne camera image understanding model with semantic-boundary multi-level reinforcement; input the image data of real-time marine environment perception into the trained model to obtain the adaptive perception result of the marine environment.

[0082] This model is trained on the enhanced MaSTr1325 dataset. The MaSTr1325 dataset is divided into a training set, a test set, and a validation set using the random sampling method, and the ratio of the three is 8:1:1. The experiment is implemented on torch1.13.0 and CUDA11.6, with a 12th Gen Intel(R) Core(TM) i7-12700K for the CPU and an NVIDIA GeForce RTX4090 for the GPU.

[0083] Early stopping training for 10 training epochs is performed using the RMSProp optimizer, and the initial learning rate is set to , using the standard polynomial decay strategy with a decay coefficient of 0.9. The optimizer decay coefficient is 0.9.

[0084] The training epoch is 100, the batch size is set to 8, and the boundary feature discriminator is activated after the 80th training epoch.

[0085] In the later stage of model training, discontinuous or abrupt perception errors often occur at the boundaries between environmental categories. This kind of error reflects the model's insufficient understanding of the boundary features of categories, resulting in mistakes in the subsequent path planning of unmanned ships.

[0086] Therefore, in order to further optimize the boundary perception accuracy and maintain the accuracy of semantic features, the present invention also designs a boundary feature discriminator applicable to the later stage of training, as Figure 2 shown. The boundary feature discriminator quantifies the difference in boundary features by comparing the similarity between the boundary prediction value of the model and the boundary ground truth, and dynamically adjusts the weight of the boundary features in the feature fusion and decoding stages, thereby enhancing the sensitivity and understanding ability of the model to boundary information.

[0087] The boundary feature discriminator utilizes the shape context method, which provides a relatively robust shape matching strategy based on the mutual relationship (such as distance, gradient) between the sampling points of the boundary ground truth and the boundary prediction result.

[0088] First, boundary extraction and boundary point sampling are performed on the prediction output of the maritime environment perception model.

[0089] The specific process of boundary extraction is as follows: Since the classification result of the maritime environment perception output by the model is a grayscale image, the class information of each pixel point needs to be mapped to the corresponding RGB color value for pseudo-color rendering, which serves as the input for boundary extraction. In this embodiment, the Canny operator is used to detect the boundary of the pseudo-color image, and finally the boundary prediction value is obtained.

[0090] Due to the uneven distribution of the boundary point set, the present invention adopts Jitendra's sampling method. This method randomly arranges the point set, calculates the distance matrix between point pairs, and gradually deletes the closest point pairs to ensure the uniform distribution of sampling points.

[0091] After obtaining the sampling points, the global similarity of each point is calculated first, and the process is as follows:

[0092] For each sampling point, a polar logarithm coordinate system is constructed, the neighboring sampling points are mapped to each region, the number of points in each region is counted and normalized, a shape context histogram matrix is generated, and the chi-square statistic is used to quantify the boundary similarity.

[0093] Suppose a total of points are sampled, each point has angles and distance regions, then the shape context matrices of the boundary ground truth and the boundary prediction value are obtained, denoted as and respectively, and substituted into the chi-square formula to calculate the similarity metric matrix ; ;

[0094] .

[0095] Through the above calculation formula, a similarity metric matrix of size N×N for global boundary features is obtained.

[0096] The local direction is used to evaluate the similarity of local features of the boundary, and the process is as follows:

[0097] First, calculate the gradient components of the boundary ground truth and the boundary prediction result, and extract the gradient values of the contour points to obtain the tangential angle , , and define the tangential angle dissimilarity function as the similarity measure of the local direction ;

[0098]

[0099]

[0100] .

[0101] Perform weighted summation on the similarity measure matrices , to output the total similarity measure matrix , and the formula is as follows:

[0102] .

[0103] Among them is the weight, which is set to 0.3 in this embodiment.

[0104] Abstract the total similarity measure matrix into a distance matrix, and solve it through the Hungarian algorithm introducing pseudo points to obtain the point pair matching relationship that minimizes the total loss; extract the similarity scores of the sampling points from the matrix according to the matching index .

[0105] Among them, the lower the similarity score between sampling points, the higher the matching rate.

[0106] On the marine image dataset LaRS, the present invention screened about 2000 groups of prediction results, each group including two types of samples with perception errors and samples highly close to the ground truth, and performed similarity matching with the boundary ground truth (sampling points 800). The results are as Figure 6 shown. There are significant differences in the average similarity scores of the two types of samples, and the calculated similarity threshold Ts is 9.

[0107] In the last stage of the multi-branch cross-channel attention module, the result after element-wise multiplication of the feature input and the attention vector is adjusted in the number of channels through a 1×1 convolutional layer and then concatenated. When the boundary feature discriminator is not activated, the initial channel number ratio of the preset boundary feature, semantic feature, and fusion feature is set to 1:1:2 in this embodiment.

[0108] After activating the boundary feature discriminator, the channel number ratio of the boundary feature, semantic feature, and fusion feature is dynamically adjusted according to the similarity score and the similarity threshold Ts; if > Ts, indicating that the boundary prediction does not match the ground truth well, then increase the weight of the boundary feature in the overall feature, and adjust the ratio to 3:1:1; otherwise, if S ≤ Ts, keep the initial channel number ratio unchanged.

[0109] Considering that the water surface has complex and dynamically changing features such as waves and reflections, the feature learning is difficult and given the significant differences between the water surface features and obstacle features, a water-obstacle perception loss is introduced as the feature loss of the network .

[0110] When the network learns class features, it is vulnerable to the interference of the class imbalance problem, resulting in a decline in classification performance. Therefore, a weighted focal loss function is introduced, and the result is calculated with the classification ground truth to obtain the result of this loss function.

[0111] The total loss function of the model is obtained by weighted summation of the three:

[0112] .

[0113] Among them, the loss function weights are and , where is 1, is 0.1.

[0114] During the model training process, a water-obstacle perception loss is introduced as the feature loss of the model Among them, the output of the third Conformer layer is used as the input of the feature loss function .

[0115] After training the model, the image data of the maritime environment perception obtained in real time (without inputting the boundary ground truth) is input into the trained model to obtain the classification result of the maritime environment perception. Specifically, the semantic feature of the image data input into the model only passes through the semantic feature encoding stream and the dilated convolutional pooling pyramid, and finally, after adjusting the number of channels through a convolutional layer with a kernel size of 1×1, the classification result of the maritime environment perception is output, without passing through the boundary feature encoding stream, the boundary-semantic feature fusion stream, and the multi-level attention fusion stream in the fusion feature decoder, as Figure 2 shown.

[0116] Based on the images collected by an on - ship camera, the present invention uses an adaptive understanding model for the marine environment with multi - level enhancement of boundary - semantic features, which well solves the problem of fuzzy boundary perception in understanding marine image categories and provides support for the autonomous cruise of USV.

[0117] In addition, in order to verify the effectiveness of the semantic - boundary multi - level enhanced on - ship camera image understanding model proposed in the present invention, the following experimental process is also given. Model evaluation is carried out on the LaRS dataset and the MODD2 dataset respectively. Three existing models, Segformer, YOLOv11m - seg, and WaSRNet, are introduced for comparative analysis in the evaluation. Among them, Segformer realizes a concise and efficient environmental perception model design through the organic combination of Transformer and a lightweight multi - layer perceptron decoder; YOLOv11m - seg continues the end - to - end architecture paradigm, adopts a multi - scale feature pyramid network and a dynamic matching mechanism, and strengthens the feature representation ability while maintaining the lightweight characteristics; WaSRNet, as a dedicated model for marine environment perception, on the basis of adopting a ResNet101 deep encoding architecture, incorporates the horizontal plane angle information simulated by an inertial sensor, improving the perception accuracy of the water surface scene. Through the comparative experiment with the three models, empirical support is provided for the effectiveness of the model designed in the present invention.

[0118] Step 1. Conduct model evaluation on the LaRS dataset.

[0119] The LaRS dataset collected many waterborne navigation scenarios from publicly available videos, covering scene types such as canals, oceans, lakes, crowded scenes, strong reflections, and low visibility conditions (such as rainy days and night navigation).

[0120] Use evaluation metrics and to evaluate the class perception accuracy and quality of the network, and their definitions are as follows:

[0121] .

[0122] .

[0123] TP is the number of samples that the model correctly predicts as positive classes. FP is the number of samples that the model wrongly predicts as positive classes. FN is the number of samples that the model wrongly predicts as negative classes. IoU is the intersection - over - union ratio, which measures the overlap degree between the model prediction result and the true result. C is the total number of classes in the task. is the IoU value of the c - th class, and the intersection - over - union ratio of the prediction result of this class is calculated separately.

[0124] Table 1 shows the performance of different networks on the LaRS dataset.

[0125] Experimental result diagram of the LaRS dataset

[0126]

[0127] It can be easily seen from Table 1 that thanks to the semantic-boundary feature fusion strategy and the supervision mechanism of boundary features proposed in the present invention, although the LaRS dataset contains a large number of complex scenarios with strong reflections and adverse weather conditions, the method of the present invention still performs excellently in various indicators. Specifically, the method of the present invention has achieved significant improvements in the mean pixel accuracy (mPA: 96.91%) and the mean intersection over union (mIoU: 94.56%) compared with WaSRNet (mPA: 94.91%, mIoU: 93.16%) designed specifically for marine environment perception. At the same time, when dealing with the perception tasks of small target obstacles on the sea (the first row) and inland water surface obstacles (the second and third rows), the method of the present invention can provide clearer and more accurate perception results at the junction of obstacles and the water surface, which provides more accurate support for the water surface path planning of unmanned ships.

[0128] Step 2. Conduct model evaluation on the MODD2 dataset.

[0129] The MODD2 dataset contains 80,828 images collected from the ports of Koper and Strunjan, covering data under different weather conditions such as sunny, cloudy, and foggy days, as well as extreme conditions (such as motion distortion, sunlight, and reflection). The MODS evaluation system is used for testing on the MODD2 dataset, which mainly includes the following indicators:

[0130] ① Obstacle-water edge localization accuracy 。

[0131] The true value of the evaluation dataset contains the water edge position of each picture.

[0132] The obstacle-water edge localization accuracy is defined as:

[0133] For the mask of the water surface category in the prediction mask, the square root of the average squared distance between the layer of pixels with the closest vertical distance to the true water edge among all its pixels and the true water edge. Its definition is as follows:

[0134] 。

[0135] Where n is the total number of pixels of the water surface category in the prediction mask, and these pixels are the objects participating in the calculation of the water edge localization accuracy; is the i-th pixel of the water surface category in the prediction mask, and its vertical distance to the layer of pixels closest to the true water edge reflects the position difference between a single pixel and the true water edge. is the average value (arithmetic mean) for all , representing the average level of the distance between the water surface pixels in the predicted mask and the true water edge.

[0136] ② Accuracy .

[0137] The accuracy of the model for environmental perception is defined as follows:

[0138] .

[0139] ③ Recall .

[0140] The recall of the model for environmental perception reflects the ability of the model to successfully identify all positive example samples. It is defined as follows:

[0141] .

[0142] ④ Error rate .

[0143] The average number of false positive detections per 100 images.

[0144] ⑤ F1 score .

[0145] F1 is an indicator that comprehensively considers precision and recall and is used to comprehensively evaluate the performance of the model. It is defined as follows:

[0146] .

[0147] Table 2 shows the verification results of different networks on the MODD2 dataset based on the MODS evaluation system.

[0148] Table 2 Experimental result diagram of the MODD2 dataset

[0149]

[0150] It can be easily seen from Table 2 that although UNet is slightly better than the method of the present invention in terms of Re (95.1%) and TPr (62.1) (Re: 92.2, TPr: 61.4), in terms of FPr, UNet (34.9) is significantly inferior to the method of the present invention (2.4). This advantage not only benefits from the in-depth learning of the boundary features and the full fusion of semantic features in the present invention, but also attributes to the introduction of a loss function specifically for separating the water surface and obstacle features, which improves the adaptability of the network to the complex water environment.

[0151] The present invention conducts empirical tests on two public datasets, LaRS and MODD2, by introducing the MODS evaluation system and the mIoU metric, and conducts a comparative analysis with networks such as WaSRNet. The experimental data shows that the model of the present invention has a result of 94.56% in the mIoU metric, and the metric reaches 16.3 pixels, thus verifying its effectiveness and superiority.

[0152] Of course, the above description is only a preferred embodiment of the present invention. The present invention is not limited to listing the above embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any person skilled in the art under the teaching of this specification fall within the substantial scope of this specification and should be protected by the present invention.

Claims

1. A method for adaptively sensing marine environment based on boundary reinforcement, characterized in that: The steps include: Step 1. Obtain the original image dataset of marine environment perception, perform data enhancement on the original images, and obtain the boundary truth value for each original image. The boundary truth value and the original image correspond one-to-one to form a training dataset; Step 2. Build a semantic-boundary multi-level enhanced shipborne camera image understanding model, which includes a boundary feature encoding stream, a semantic feature encoding stream, a boundary-semantic feature fusion stream, and a fusion feature decoder; The boundary feature encoding stream and the semantic feature encoding stream are used to encode the boundary truth value and the original image of the input model respectively; each encoding stream includes four cascaded feature extraction layers; The boundary-semantic feature fusion flow includes a cascade of four layers of feature attention modules; Each layer of feature attention module is used to receive the output of the corresponding layer of the boundary feature encoding stream and the semantic feature encoding stream, and is used to realize the feature interaction of the boundary feature and the semantic feature of the corresponding layer to enhance the boundary perception ability; The outputs of the feature attention modules at each layer are fed back to the corresponding layers of the semantic feature encoding stream, and together with the outputs of the corresponding layers of the semantic feature encoding stream, serve as the inputs of the lower layers of the semantic feature encoding stream; The output results of the boundary feature encoding stream, semantic feature encoding stream, and boundary-semantic feature fusion stream are input into the fusion feature decoder, and after being enhanced by the three-dimensional feature attention of direction, position, and channel, the marine environment perception classification result is output; Step 3. Use the training data set to train the semantic-boundary multi-level reinforced shipborne camera image understanding model; input the real-time acquired marine environment perception image data into the trained model to obtain the marine environment adaptive perception results.

2. The method for adaptively sensing marine environment based on boundary reinforcement according to claim 1 is characterized in that: In step 1, the training data set is obtained as follows: Firstly, the original image dataset for marine environment perception is obtained, and the original image is preprocessed with data enhancement to expand the amount of original image data; the true value of the dataset is a grayscale image with pixel-by-pixel category annotation; Data augmentation by adjusting image properties, including exposure, contrast, saturation, and / or hue; At the same time, the boundary truth value is obtained for each original image, and the boundary truth value corresponds to the original image one by one; The method for obtaining the boundary truth value is as follows: first, the grayscale image is converted into a multi-category RGB image through pseudo-color coding, and then the Canny edge detection algorithm is applied to extract the boundary, and finally the boundary truth value is obtained.

3. The method for adaptively sensing marine environment based on boundary reinforcement according to claim 1 is characterized in that: The boundary feature encoding stream adopts the ResNet18 network; the 1st to 4th residual convolution layers of the ResNet18 network are used as the four-layer feature extraction layers of the boundary feature encoding stream; the output of each residual convolution layer is respectively input into the boundary-semantic feature fusion stream to realize the step-by-step interaction between the boundary truth value and the semantic feature, and promote bidirectional feature learning; The boundary feature encoding stream is spatially downsampled based on the output size of the last residual convolution layer, and finally the multi-scale boundary feature maps formed by each residual convolution layer are concatenated as the output of the boundary feature encoding stream.

4. The method for adaptively sensing marine environment based on boundary reinforcement according to claim 1 is characterized in that: The semantic feature encoding stream is constructed based on the Conformer-S architecture; the first four Conformer layers of the Conformer-S architecture are used as four feature extraction layers of the semantic feature encoding stream; the output of each Conformer layer is respectively input into the boundary-semantic feature fusion stream to realize the step-by-step interaction between the boundary truth value and the semantic feature, and promote bidirectional feature learning; At the same time, each Conformer layer also receives feedback information from the corresponding layer of the boundary-semantic feature fusion flow; The semantic feature encoding stream is spatially downsampled based on the output size of the last Conformer layer, and finally the multi-scale semantic feature maps formed by each Conformer layer are concatenated as the output of the semantic feature encoding stream.

5. The method for adaptively sensing marine environment based on boundary reinforcement according to claim 1 is characterized in that: The boundary-semantic feature fusion stream includes 4 layers of cascaded feature attention modules, each layer of feature attention modules is used to receive the output of the corresponding layer of the boundary feature encoding stream and the semantic feature encoding stream, and perform feature interaction; The boundary-semantic feature fusion flow is downsampled according to the spatial size of the last layer of feature attention modules, and finally the feature maps output by the feature attention modules of each layer are spliced ​​in the channel dimension to form a multi-scale feature output of the boundary-semantic feature fusion flow.

6. The method for adaptively sensing marine environment based on boundary reinforcement according to claim 1 is characterized in that: The processing flow of the feature attention module is as follows: First, the boundary features and semantic features input into the feature attention module are added element by element to form joint feature 1; The joint feature is divided into two branches. One branch introduces the ReLU activation function with nonlinear transformation to enhance the model's ability to express fine boundary structures. The other branch aggregates global semantic information through global average pooling and then activates through ReLU. The two processing results are added element by element to obtain joint feature 2; The joint feature 2 is multiplied element by element by the joint feature 1 to generate the attention weight based on the interactive feature, which is then multiplied with the boundary and semantic features respectively to specifically enhance the effective information of the boundary details and semantic content; Finally, we add them element by element to integrate the positioning details of the boundary features with the abstract understanding of the semantic features to obtain the fused features.

7. The method for adaptively sensing marine environment based on boundary reinforcement according to claim 1 is characterized in that: The fusion feature decoder includes a multi-level attention fusion module and a dilated convolution pooling pyramid module; After receiving the output of the boundary feature encoding stream, the semantic feature encoding stream, and the boundary-semantic feature fusion stream, the multi-level attention fusion module fuses the three features and enhances the model's understanding of the correlation between them; The multi-level attention fusion module includes a multi-branch cross-channel attention module and a collaborative attention module; The multi-branch cross-channel attention module uses the multi-branch cross-channel interaction structure to adaptively learn channel weights, establish a direct association between channels and their weights, and reflect the global contribution of semantic features and boundary features in terms of scale and type dimensions; The collaborative attention module captures the spatial features in two directions and generates an attention map to complementarily enhance the representation of the object of interest in the input feature map to address the shortcomings of the multi-branch cross-channel attention module in processing direction and position information; The dilated convolutional pooling pyramid module is used to receive the output from the multi-level attention fusion module and the semantic feature encoding stream, while expanding the receptive field and ensuring the resolution of the output image; Finally, after adjusting the number of channels through a convolution layer with a convolution kernel size of 1×1, the marine environment perception classification results are output.

8. The method for adaptively sensing marine environment based on boundary reinforcement according to claim 1 is characterized in that: In the step 3, a boundary feature identifier is introduced in the model training process; wherein the marine environment perception classification result output by the semantic-boundary multi-level enhanced shipborne camera image understanding model is a grayscale image; The category information of each pixel in the grayscale image is mapped to the corresponding RGB color value for pseudo-color rendering. As the input of boundary extraction, the Canny operator is used to perform boundary detection on the pseudo-color image to finally obtain the boundary prediction value. Then, Jitendra's sampling method is used to sample boundary points. By randomly arranging the point set, the distance matrix between point pairs is calculated, and the closest point pairs are gradually deleted to ensure that the sampling points are evenly distributed. The boundary feature discriminator compares the similarity between the boundary prediction value and the boundary true value, quantifies the difference of boundary features, and dynamically adjusts the weight of boundary features in the feature fusion and decoding stages to improve the model's sensitivity and understanding of boundary information.

9. The method for adaptively sensing marine environment based on boundary reinforcement according to claim 8, characterized in that: In step 3, after obtaining the sampling points, the global similarity of each point is first calculated, and the process is as follows: Construct an extreme logarithmic coordinate system for each sampling point, map neighboring sampling points to each region, count the number of points in each region and normalize them, generate a shape context histogram matrix, and use chi-square statistics to quantify boundary similarity; Assuming co-sampling points, each point is Angles and distance region, and obtain the boundary truth value and boundary prediction value The shape context matrix is ​​recorded as and , substitute into the chi-square formula to calculate the similarity measure matrix ; Use local directions to evaluate the similarity of local features of the boundary. The process is as follows: First, the gradient component of the boundary true value and the boundary prediction knot value is calculated, and the gradient value of the contour point is extracted to obtain the tangent angle and , define the tangential angle dissimilarity function as the similarity measure of local directions ; The similarity measure matrix , Perform weighted summation and output the total similarity measurement matrix ; The total similarity measurement matrix Abstracted as a distance matrix, and solved by introducing the pseudo-point Hungarian algorithm to obtain the point pair matching relationship that minimizes the total loss; according to the matching index from the matrix Extract the similarity scores of the sampling points from ; Defining Similarity Threshold ,In the final stage of the multi-branch cross-channel attention module, the result of element-wise multiplication of the feature input and the attention vector is concatenated after adjusting the number of channels through a 1×1 convolutional layer; When the boundary feature discriminator is not activated, the ratio of the initial number of channels of the boundary feature, semantic feature, and fusion feature is preset; When the boundary feature discriminator is activated, the ratio of the number of channels of boundary features, semantic features and fusion features is calculated according to and Make dynamic adjustments; if > , increase the weight of boundary features in the overall features; otherwise, keep the ratio of the initial number of channels unchanged.

10. A marine environment adaptive perception system based on boundary reinforcement, carried on a USV, comprising a camera and a computer device; the camera is used to obtain image data of marine environment perception in real time and upload it to the computer device; in, The computer device includes a memory and a processor; An executable code is stored in the memory, characterized in that when the processor executes the executable code, it is used to implement the steps of adaptive perception of the marine environment based on boundary enhancement as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method based on attention multi-scale feature fusion

    CN111127493A

  • Boundary-driven neural network structure for unmanned ship environment understanding

    CN118396071A