Marine environment adaptive sensing method and system based on boundary enhancement

By adopting boundary reinforcement methods in maritime environment perception, a semantic-boundary multi-level reinforcement image understanding model is built, which solves the problem of insufficient learning of boundary features in marine environment perception, and achieves more accurate environmental perception and unmanned ship path planning, improving the accuracy and efficiency of autonomous cruises.

CN120071304AActive Publication Date: 2025-05-30SHANDONG UNIV OF SCI & TECH

Patent Information

Application Number
CN202510533490.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-05-30
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

The prior art ignores the learning of boundary features in the image in marine environment perception, which makes it difficult for category perception accuracy to reach an ideal level, and unmanned surface boats find it difficult to accurately understand the onboard camera images in different mission scenarios, affecting the accuracy of path planning and autonomous cruise.

Method used

Adaptive perception method of maritime environment based on boundary reinforcement is adopted, and the marine camera image understanding model with semantic and boundary multi-level reinforcement is designed by designing the image boundary-semantic information fusion and supervision of boundary feature learning, and the accurate perception of the maritime environment is achieved using boundary feature encoding streams, semantic feature encoding streams, boundary-semantic feature fusion streams and fusion feature decoders.

Benefits of technology

It effectively improves the adaptability of unmanned ships to the maritime environment, enhances boundary perception, improves the accuracy of environmental semantic category classification, provides more reliable decision-making support, and provides more accurate and efficient support for unmanned ship path planning and autonomous cruise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071304A_ABST
    Figure CN120071304A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of unmanned surface vessel autonomous cruise and marine environment perception, and particularly discloses a marine environment self-adaptive perception method and system based on boundary enhancement. In the method, a semantic-boundary multi-level enhanced shipborne camera image understanding model is provided. In the training stage, an original image and a boundary truth value are coded through a semantic feature coding stream and a boundary feature coding stream respectively, and feature interaction is performed through a boundary-semantic feature fusion stream to enhance the boundary perception ability; the results of the coding stream and the fusion stream are input into a fusion feature decoder, after direction, position and channel three-dimensional feature attention enhancement, an environment category understanding result is output, prediction output of the model generates a prediction boundary through boundary extraction, a boundary truth value is combined in a boundary feature discriminator to construct a quantitative index, and the prediction boundary is subjected to direction, position and channel three-dimensional feature attention enhancement. And performing negative feedback optimization on the model through a multi-task loss function. According to the invention, boundary information in different scenes is covered, and the adaptability of the unmanned ship to the marine environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous cruising of unmanned surface vessels (USVs) and marine environment perception, and particularly relates to a method and system for adaptive marine environment perception based on boundary enhancement. Background Art

[0002] With the continuous evolution of artificial intelligence and unmanned driving technologies, USVs have been increasingly widely used in practical applications. With its significant advantages of miniaturization and high scalability, USVs play an important role in key fields such as marine environment monitoring and water area inspection. When performing tasks, USVs usually carry multi-source sensors including cameras, collect images of marine scenes through the cameras, and perform environmental perception and understanding to generate class masks including different categories such as water surface, obstacles, and sky. These mask information provides key environmental perception capabilities for USVs, helping them accurately identify navigable waters and restricted areas, and achieve autonomous navigation and efficient task execution. In recent years, although the research on perception and understanding of the marine environment has been continuously improved, most of the research ignores the learning of boundary features between various categories, resulting in the class perception accuracy being difficult to reach an ideal level.

[0003] In the research on environmental perception assisted by boundary features, in-depth explorations have been carried out in various fields. Bovcon et al. simulated the edge information of the horizontal plane with an inertial measurement unit (IMU) and fused it with semantic information; He et al. generated a pseudo water line using radar data and fused its feature encoding with semantic information. Both of these studies have improved the class perception accuracy of the water surface edge to a certain extent, but in the process of obtaining boundary information, they both ignore the boundary features originally contained in the image, such as the boundary between the sea surface and marine obstacles, the boundary between the sky and shore obstacles, etc. In addition, in the field of unmanned vehicle technology, rapid progress has been made in the research on autonomous navigation route planning based on sensors such as on-vehicle cameras. However, due to the essential differences between the marine environment and the land environment, the related technologies of unmanned vehicles are difficult to be directly transferred and applied to USVs.

[0004] Currently, the mission scenarios of USVs at sea can be divided into two categories: (1) inspection tasks in the coastal area near the sea, and at this time, the perception effect of shore buildings will directly affect the accuracy and reliability of USV path planning; (2) environmental data collection tasks in the open sea area, and at this time, the understanding effect of obstacles on the sea surface will determine the accuracy of USV autonomous navigation. It can be seen that at the present stage, only using the edge features of the simulated sea level cannot effectively guide USVs to understand the images of on-board cameras in different mission scenarios, and thus cannot correctly guide USVs to perform route planning and autonomous cruising. Summary of the Invention

[0005] The object of the present invention is to propose an adaptive perception method for the marine environment based on boundary enhancement, which is designed by fusing boundary-semantic information in images and supervising the learning of boundary features, so as to cover boundary information in different scenarios, improve the adaptability of unmanned ships to the marine environment, and thus more effectively guide the unmanned ships to perform operations such as path planning.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions: An adaptive perception method for the marine environment based on boundary enhancement, comprising the following steps: Step 1. Obtain the original image dataset for marine environment perception, perform data augmentation on the original images, and obtain the boundary ground truth for each original image. The boundary ground truth and the original image correspond one by one to form a training dataset; Step 2. Build an on-board camera image understanding model with multi-level reinforcement of semantics and boundaries, which includes a boundary feature encoding stream, a semantic feature encoding stream, a boundary-semantic feature fusion stream, and a fusion feature decoder; The boundary feature encoding stream and the semantic feature encoding stream are respectively used to encode the boundary ground truth and the original image input into the model; each encoding module includes four cascaded feature extraction layers; The boundary-semantic feature fusion stream includes four cascaded feature attention modules; Each layer of the feature attention module is respectively used to receive the outputs of the corresponding layers of the boundary feature encoding stream and the semantic feature encoding stream, and is used to realize the feature interaction between the boundary features and the semantic features of the corresponding layers, so as to enhance the boundary perception ability; The outputs of each layer of the feature attention module are respectively fed back to the corresponding layer of the semantic feature encoding stream, and are used as the lower-layer input of the semantic feature encoding stream together with the output of the corresponding layer of the semantic feature encoding stream; The output results of the boundary feature encoding stream, the semantic feature encoding stream, and the boundary-semantic feature fusion stream are input into the fusion feature decoder, and after being strengthened by three-dimensional feature attention in terms of direction, position, and channel, the classification result of marine environment perception is output; Step 3. Use the training dataset to train the on-board camera image understanding model with multi-level reinforcement of semantics and boundaries; input the image data of the marine environment perception obtained in real time into the trained model to obtain the adaptive perception result of the marine environment.

[0007] In addition, on the basis of the above adaptive perception method for the marine environment based on boundary enhancement, the present invention also proposes a corresponding adaptive perception system for the marine environment based on boundary enhancement, which adopts the following technical solutions: An adaptive maritime environment perception system based on boundary enhancement, mounted on a USV, includes a camera and a computer device; the camera is used to obtain real-time image data for maritime environment perception and upload it to the computer device; Among them, the computer device includes a memory and a processor; executable code is stored on the memory, and when the processor executes the executable code, it is used to implement the steps of the above-mentioned adaptive maritime environment perception based on boundary enhancement.

[0008] The present invention has the following advantages: As described above, the present invention relates to a method and system for adaptive maritime environment perception based on boundary enhancement. Among them, in order to achieve more accurate environment perception in multi-scenario applications of unmanned ships at sea, the present invention proposes a semantic-boundary multi-level enhanced onboard camera image understanding model. This model realizes multi-branch multi-level fusion and enhancement of boundary features and semantic information at the input level, encoding level, decoding level, and output level during training, effectively solving the problem of inaccurate classification of environmental semantic information due to fuzzy class boundaries in nearshore and open sea environments. In addition, in order to solve the problem of missing boundary information, at the input level, the present invention introduces the boundary ground truth of the image as an input into the model during the training phase and encodes it in parallel with the original image. At the encoding level, the present invention designs a boundary-semantic feature fusion stream in combination with the attention mechanism to realize the interaction and enhancement of the original image features (i.e., semantic features) and boundary features during the layer-by-layer encoding process, endowing the model with deep semantic understanding ability and boundary sensitivity. In addition, at the decoding level, in order to effectively fuse semantic, boundary, and semantic-boundary features, the present invention also designs a multi-branch cross-channel attention module to fuse features complementarily at the channel, direction, and position levels, synergistically improving the accuracy of the environmental perception category results. At the output level, the present invention designs a boundary feature discriminator to dynamically adjust the weight distribution of boundary features by quantifying the boundary feature differences, enhancing the model's sensitivity and understanding ability to boundary information. The model of the present invention can effectively learn the boundary features in the onboard camera images, help unmanned ships adaptively understand the implicit boundary information in the images in multiple application scenarios, and provide more reliable decision-making support for subsequent unmanned ship path planning by accurately and efficiently guiding environmental semantic category classification, effectively improving the accuracy and execution efficiency of path planning. Description of the Drawings

[0009] Figure 1 It is a flowchart of the adaptive maritime environment perception based on boundary enhancement in the embodiment of the present invention; Figure 2 It is a structural diagram of the semantic-boundary multi-level enhanced onboard camera image understanding model in the embodiment of the present invention; Figure 3It is the boundary-semantic feature fusion flow structure diagram in the embodiment of the present invention; Figure 4 It is the structure diagram of the feature attention module in the embodiment of the present invention; Figure 5 It is the structure diagram of the multi-level attention fusion flow in the fusion feature decoder in the embodiment of the present invention; Figure 6 It is the distribution diagram of the average value of the sample similarity scores in the embodiment of the present invention. Detailed implementation manners

[0010] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners: Embodiment 1 This embodiment describes an adaptive perception method for the marine environment based on boundary enhancement. In this method, a shipborne camera image understanding model with multi-level semantic-boundary enhancement is proposed, which includes a boundary feature encoding flow, a semantic feature encoding flow, a boundary-semantic feature fusion flow, a fusion feature decoder, and a boundary feature discriminator, etc. Among them, in the training stage, the present invention adopts a dual-channel feature interaction architecture, that is, the original image and the boundary ground truth are respectively encoded through the semantic feature encoding flow and the boundary feature encoding flow, and feature interaction is performed through the boundary-semantic feature fusion flow to enhance the boundary perception ability; the encoding results of the boundary feature encoding flow, the semantic feature encoding flow, and the boundary-semantic feature fusion flow are input into the fusion feature decoder, and after being enhanced by three-dimensional feature attention in the fusion feature decoder in terms of direction, position, and channel, the environmental category understanding result is output. In addition, in the model training stage, the present invention also generates a predicted boundary by boundary extraction from the predicted output of the model, constructs a quantization index in the boundary feature discriminator in combination with the boundary ground truth, and optimizes the model through negative feedback of the multi-task loss function. In the model inference stage, only the semantic feature encoding flow and the fusion feature decoder are retained, and the optimized boundary feature parameters in the training stage are used to improve the environmental category understanding accuracy while reducing the number of parameters, effectively guiding the route planning and autonomous cruising of the unmanned ship.

[0011] As Figure 1 shown, the adaptive perception method for the marine environment based on boundary enhancement in this embodiment includes the following steps: Step 1. Obtain the original image dataset for marine environment perception, perform data augmentation on the original images, and obtain the boundary ground truth for each original image. The boundary ground truth and the original image form a training dataset in one-to-one correspondence.

[0012] The original image dataset adopted in this embodiment is, for example, the commonly used MaSTr1325 dataset in the field of marine environment perception, with a quantity of 1325 images, covering image data of various time periods within a day and diverse weather conditions.

[0013] Aiming at the problems of insufficient sample size in the MaSTr1325 dataset and differences in imaging parameters of different models of cameras, the present invention performs data augmentation by adjusting image attributes such as exposure, contrast, saturation, and hue, expands the number of samples while maintaining the original semantic information, and effectively improves the generalization ability and environmental adaptability of the model to different imaging conditions.

[0014] Specifically, the exposure (exposure adjustment coefficient k ∈ [0.6, 1.4]) and contrast (contrast factor c ∈ [0.8, 1.2]) are adjusted through the linear transformation clip function in python. The saturation (0.7 - 1.3 times) is controlled by adjusting the S channel and the hue is finely tuned (±15°) by adjusting the H channel in the HSV space of the picture. Enhanced versions with single-parameter adjustment (2 intensities for each of the 4 types of parameters), two-parameter combinations (2 intensities for each of the 6 combinations), three-parameter combinations (2 intensities for each of the 4 combinations), and four-parameter combinations (2 intensities for 1 combination) are generated for each image. Finally, an improvement from 1325 original images to 41075 expanded samples is achieved.

[0015] Through the above data augmentation method, the data deviation is effectively alleviated and the generalization ability of the model is enhanced. The ground truth of the MaSTr1325 dataset is a grayscale image with pixel-by-pixel class annotations. The dataset includes: ① Color images [original images] taken by the camera, which have three RGB channels: ② The ground truth for classifying each pixel of the original image, which has only one channel, and the value of each pixel is 1, 2, or 3 to distinguish which class this pixel belongs to. The numbers 1, 2, and 3 represent water surface, sky, and obstacle respectively.

[0016] In order to obtain the boundary ground truth, the present invention first converts the grayscale image into a multi-class RGB image through pseudo-color coding, then applies the Canny edge detection algorithm to extract the boundary, and finally obtains the boundary ground truth.

[0017] Step 2. Build a shipborne camera image understanding model with multi-level reinforcement of semantics and boundaries.

[0018] As Figure 2 shown, the shipborne camera image understanding model with multi-level reinforcement of semantics and boundaries built in this embodiment includes a boundary feature encoding stream, a semantic feature encoding stream, a boundary-semantic feature fusion stream, and a fusion feature decoder.

[0019] In this embodiment, the shipborne camera images (i.e., original images) and the corresponding boundary ground truth inputs are used as the data sources of the shipborne camera image understanding model, where the sizes of the input data are correspondingly denoted as 3×h×w and 1×h×w.

[0020] The boundary feature encoding stream and the semantic feature encoding stream are respectively used to encode the boundary ground truth and the original image of the input model; each encoding module includes four cascaded feature extraction layers.

[0021] The boundary-semantic feature fusion stream includes four cascaded feature attention modules.

[0022] Each layer of the feature attention module is respectively used to receive the outputs of the corresponding layers of the boundary feature encoding stream and the semantic feature encoding stream, and is used to implement the feature interaction between the boundary features and the semantic features of the corresponding layer, so as to enhance the boundary perception ability.

[0023] The outputs of each layer of the feature attention module are respectively fed back to the corresponding layer of the semantic feature encoding stream, and together with the output of the corresponding layer of the semantic feature encoding stream, they are used as the input of the lower layer of the semantic feature encoding stream.

[0024] The output results of the boundary feature encoding stream, the semantic feature encoding stream, and the boundary-semantic feature fusion stream are input into the fusion feature decoder. After being enhanced by three-dimensional feature attention in terms of direction, position, and channel, an RGB image is output as the classification result of marine environment perception, so that the same type of targets present consistent RGB color coding in this image.

[0025] Since the boundary ground truth features are relatively simple, the boundary feature encoding stream uses the ResNet18 network with fewer parameters; the first to fourth residual convolutional layers of the ResNet18 network are used as the four-layer feature extraction layers of the boundary feature encoding stream.

[0026] The outputs of each residual convolutional layer are respectively input into the boundary-semantic feature fusion stream to realize the hierarchical interaction between the boundary ground truth and the semantic features, promote bidirectional feature learning, and perform spatial downsampling based on the output size of the last residual convolutional layer. Finally, the multi-scale boundary feature maps formed by each residual convolutional layer are spliced and used as the output of the boundary feature encoding stream.

[0027] The boundary feature encoding stream is conducive to enhancing the network's understanding of boundary details through the splicing of multi-scale boundary feature maps.

[0028] In this embodiment, the basic component unit of the boundary feature encoding stream is the residual convolutional layer.

[0029] Each residual convolutional layer consists of several convolutional layers. In the structure of ResNet-18, the conv1 layer contains a convolutional layer with a kernel size of 7×7 and a stride of 2; conv2_x contains 2 residual blocks, each residual block consists of two convolutional layers with a kernel size of 3×3, and each layer outputs 64 channels; conv3_x contains 2 residual blocks, each residual block consists of two convolutional layers with a kernel size of 3×3, and each layer outputs 128 channels; conv4_x contains 2 residual blocks, each residual block consists of two convolutional layers with a kernel size of 3×3, and each layer outputs 256 channels. For the downsampling layer of conv2_x, it is designed as a convolutional layer with a kernel size of 8 and a stride of 8; for the downsampling layer of conv3_x, it is designed as a convolutional layer with a kernel size of 4 and a stride of 4; for the downsampling layer of conv4_x, it is designed as a convolutional layer with a kernel size of 2 and a stride of 2.

[0030] The semantic feature encoding stream is constructed based on the Conformer-S architecture, and the multi-head Transformer module adopted by this architecture contains a 6-head attention mechanism. Conformer combines the convolutional neural network and the self-attention mechanism of Transformer, giving full play to the advantages of local feature extraction and global context modeling, enabling the model to accurately capture complex dynamic features such as sea surface fluctuations and lighting changes, and effectively modeling long-distance dependencies, thereby enhancing the understanding and parsing ability of the marine environment.

[0031] Take the first four Conformer layers of the Conformer-S architecture as the four feature extraction layers of the semantic feature encoding stream; the output of each Conformer layer is respectively input into the boundary-semantic feature fusion stream to achieve the hierarchical interaction between the boundary ground truth and the semantic features, and promote bidirectional feature learning. The output of the third Conformer layer is used as the input of the feature loss function of.

[0032] At the same time, each Conformer layer also receives the feedback information of the corresponding layer of the boundary-semantic feature fusion stream.

[0033] The semantic feature encoding stream performs spatial downsampling based on the output size of the last Conformer layer, and finally splices the multi-scale semantic feature maps formed by each Conformer layer as the output of the semantic feature encoding stream.

[0034] By fusing multi-scale features through downsampling of a unified size, it effectively integrates shallow details (boundary positions) and deep semantics (category information), enhances the richness and robustness of feature representation, and at the same time avoids complex upsampling operations, improving computational efficiency.

[0035] The basic building blocks of the semantic feature encoding stream are Conformer layers c1 to c4. C1 contains a convolutional layer with a 7×7 convolutional kernel and a stride of 2, followed by a 3×3 max pooling layer with a stride of 2. In c2, the CNN Branch contains a module composed of a 1×1 convolution, a 3×3 convolution, and a 1×1 convolution in sequence; the Feature Coupling Unit (FCU) contains a 1×1 convolution; the Transformer Branch first passes through a 4×4 convolution with a stride of 4, and then connects to a module composed of a multi-head attention mechanism (MHSA-6), a 1×1 convolution, and a 1×1 convolution, and this module is repeated 1 time. In c3, the CNN Branch contains a module composed of a 1×1 convolution, a 3×3 convolution, and a 1×1 convolution in sequence; the FCU contains a 1×1 convolution; the Transformer Branch is composed of the module of MHSA-6, 1×1 convolution, and 1×1 convolution repeated 4 times. In c4, the CNN Branch contains a module composed of a 1×1 convolution, a 3×3 convolution, and a 1×1 convolution in sequence; the FCU contains a 1×1 convolution; the Transformer Branch is composed of the module of MHSA-6, 1×1 convolution, and 1×1 convolution repeated 3 times. For the downsampling layer of c2, it is designed as a convolutional layer with a convolutional kernel size of 8 and a stride of 8; for the downsampling layer of c3, it is designed as a convolutional layer with a convolutional kernel size of 4 and a stride of 4; for the downsampling layer of c4, it is designed as a convolutional layer with a convolutional kernel size of 2 and a stride of 2.

[0036] The original Conformer-S architecture and the Resnet18 architecture have a total of five layers, and the output result size is c×7×7 at this time.

[0037] The present invention only adopts the first four layers of the two architectures, and the output result size is c×14×14. Reducing the number of layers to improve the resolution of the feature map can retain richer spatial details and alleviate the problem of information loss of small targets in the image.

[0038] As Figure 3 shown, the boundary-semantic feature fusion stream contains 4 cascaded feature attention modules. Each layer of the feature attention module is used to receive the outputs of the corresponding levels of the boundary feature encoding stream and the semantic feature encoding stream, and perform feature interaction. On the one hand, the processing result is fed back to the corresponding level of the semantic feature encoding stream. On the other hand, each layer of the feature attention module of the boundary-semantic feature fusion stream is downsampled with the spatial size of the last layer of the feature attention module, and finally the feature maps output by each layer of the feature attention module are concatenated in the channel dimension to form the multi-scale feature output of the boundary-semantic feature fusion stream.

[0039] During subsequent model inference, the overall input and output only pass through the semantic feature encoding stream and do not go through the boundary feature encoding stream and the boundary-semantic feature fusion stream. Therefore, during the model training process, during the fusion feature transfer process through the boundary-semantic feature fusion stream, it only interacts with the semantic feature encoding stream. This architecture enables the semantic feature encoding stream to synchronously learn boundary feature parameters during the encoding stage, reducing semantic feature interference during boundary encoding, thereby effectively improving training efficiency.

[0040] The feature attention module designs an attention-based feature aggregation mechanism, and its structure is as Figure 4 shown.

[0041] First, the boundary feature and the semantic feature input into the feature attention module are added element-wise to form joint feature one.

[0042] Joint feature one is divided into two branches. One branch is processed by introducing the ReLU activation function of non-linear transformation to enable the model to enhance the expression ability of the fine structure of the boundary. The other branch is aggregated with global semantic information through global average pooling and then activated by ReLU. Further, the abstract expression of semantic features is mined through non-linear operations to enhance the representation ability of semantic features for abstract semantic content.

[0043] The processing results of the two branches are added element-wise to obtain joint feature two.

[0044] Joint feature two is multiplied element-wise with joint feature one to generate attention weights based on the post-interaction features, thereby measuring the key degree of boundary and semantic features. Then, they are multiplied with the boundary and semantic features respectively to specifically strengthen the effective information of boundary details and semantic content. Finally, they are added element-wise to integrate the localization details of the boundary feature and the abstract understanding of the semantic feature, obtaining the fusion feature, so that the fusion feature simultaneously has refined boundary information and high-level semantic representation.

[0045] The semantic information and boundary information at each level are processed through joint features, global average pooling, and activation operations to strengthen the key feature weights and achieve the efficient cross-fusion of semantic and boundary information. Therefore, more representative fusion features can be output.

[0046] The boundary-semantic feature fusion stream is mainly composed of four consecutive feature attention modules, which respectively accept the output results of the residual convolutional layers conv1, conv2_x, conv3_x, conv4_x from the boundary feature encoding stream, and the output results of the Conformer layers c1 to c4 from the semantic feature encoding stream as inputs.

[0047] For a single feature attention module, after receiving the boundary feature and the semantic feature, first add the two element-wise to generate "combined feature one", and process it through the ReLU activation function; then apply a 1×1 convolution operation to the combined feature 1, and complete feature normalization with BatchNorm. After that, the feature processing is divided into two paths: one path performs global average pooling on the feature to extract global semantic information, and then activates it through ReLU; the other path retains local feature details. The features of the two paths are fused through a multiplication operation, and "combined feature two" is generated through the Sigmoid activation function. Finally, through two multiplication interaction operations, the feature information of different levels is gradually fused, and finally the refined feature that combines the boundary and semantic information is output. The downsampling structure output by each stream: A 3×3 convolution with a stride of 2 can be selected to ensure that the spatial dimensions of the four streams are completely aligned after downsampling, and finally the four downsampled features are concatenated in the channel dimension.

[0048] The fused feature decoder includes a multi-level attention fusion stream and an atrous convolution pooling pyramid module, where Figure 5 shows the network structure of the multi-level attention fusion stream, including a multi-branch cross-channel attention module and co-attention.

[0049] After receiving the outputs of the boundary feature encoding stream, the semantic feature encoding stream, and the boundary-semantic feature fusion stream, the multi-level attention fusion module fuses the three features and enhances the model's understanding of the correlation between them.

[0050] The multi-level attention fusion module includes a multi-branch cross-channel attention module and a co-attention module.

[0051] The multi-branch cross-channel attention module uses a multi-branch cross-channel interaction structure to adaptively learn channel weights, establish a direct association between channels and their weights, and reflect the global contributions of semantic features and boundary features in terms of scale and type dimensions.

[0052] The co-attention module captures spatial features in two directions and generates an attention map to complementarily enhance the representation of the objects of interest in the input feature map, so as to address the deficiencies of the multi-branch cross-channel attention module in processing direction and position information.

[0053] The above two attention modules complement each other and jointly improve the accuracy of the category perception results in the marine environment.

[0054] The atrous convolution pooling pyramid module is used to receive the outputs from the multi-level attention fusion stream and the semantic feature encoding stream, while expanding the receptive field and ensuring the resolution of the output image.

[0055] Finally, after adjusting the number of channels through a convolution layer with a kernel size of 1×1, the classification result of the marine environment perception is output.

[0056] Figure 5 It is a structural diagram of a multi-level attention fusion stream, which includes a multi-branch cross-channel attention module and collaborative attention. As shown in Figure 5 the above figure, the inputs of the multi-branch cross-channel attention module are the results of the boundary feature encoding stream, the semantic feature encoding stream, and the boundary-semantic feature fusion stream , all with the size of C×H×W. First, global average pooling is performed on all results respectively, and the results are concatenated into a 3×C feature, which is processed by batch normalization and channel shuffling to generate feature p. Then, circular convolution (convolution kernel size k = 3) is used to perform convolution operations on p and another circular feature q, and then processed by the Sigmoid activation function and through the inverse shuffling operation to output 3 1×C attention features. The initial input is multiplied by the obtained attention features channel by channel and concatenated in the channel dimension, and the final result is output to the collaborative attention module.

[0057] As shown in Figure 5 the following figure, the input of the collaborative attention module is , which first passes through two adaptive average pooling layers respectively, used for pooling in the height and width dimensions, reducing the height and width dimensions of the feature map to 1, and the results are set as the height and width attention feature maps and , with the sizes of C×H×1 and C×1×W respectively. Calculate the intermediate number of channels mip, taking the larger value between the result of dividing the number of input channels C by the reduction rate r and 8. Swap the second and third dimensions of to obtain the height attention feature map with the size of C×1×H . Concatenate with at the channel level and pass through a convolutional layer with a convolution kernel size of 1×1, used to reduce the number of channels of the input feature map from C to C / r. The result is input to the batch normalization layer, used to normalize the output of the convolutional layer, and then activated using the h_swish activation function. Split the feature map after convolution, normalization, and activation in the height and width dimensions. Swap the width and height dimensions of the width attention feature map to restore the original dimension order. Perform 1×1 convolution on the height and width attention feature maps to expand the number of channels of the intermediate feature map from mip to C and convert it into attention weights through the Sigmoid function. Multiply the input feature map by the height and width attention weights to obtain the final output feature map with the size of C×H×W.

[0058] The basic components of the Atrous Convolution Pooling Pyramid include a 1×1 convolution module, a multi-scale atrous convolution module, and a pooling module. 1×1 convolution module: It contains a convolution layer with a 1×1 convolution kernel and a stride of 1, followed by a batch normalization layer and a ReLU activation function. Multi-scale atrous convolution module: It consists of multiple ASPPConv sub-modules. Each ASPPConv sub-module contains an atrous convolution layer with a 3×3 convolution kernel, a stride of 1, and a padding size equal to the dilation rate corresponding to this sub-module, followed by a batch normalization layer and a ReLU activation function. The dilation rates of different sub-modules are specified by the input atrous_rates list, and the atrous_rates applied in this method are 1, 1, 4, 8. Pooling module: First, the input feature map is pooled into a size of 1×1 through adaptive average pooling, followed by a convolution layer with a 1×1 convolution kernel and a stride of 1, then a batch normalization layer and a ReLU activation function, and finally, the feature map is upsampled to the original size of the input feature map using bilinear interpolation. Finally, it is output through a convolution layer with a convolution kernel size of 1×1 and 3 channels as the final output of the model.

[0059] Step 3. Use the training dataset to train the shipborne camera image understanding model with semantic-boundary multi-level reinforcement; input the image data of real-time marine environment perception into the trained model to obtain the marine environment adaptive perception result.

[0060] This model is trained on the enhanced MaSTr1325 dataset. The MaSTr1325 dataset is divided into a training set, a test set, and a validation set using the random sampling method, and the ratio of the three is 8:1:1. The experiment is implemented on torch1.13.0 and CUDA11.6, with a 12th Gen Intel(R) Core(TM) i7-12700K for the CPU and an NVIDIA GeForce RTX4090 for the GPU.

[0061] Early stopping training for 10 training epochs is performed using the RMSProp optimizer, and the initial learning rate is set to , using the standard polynomial decay strategy with a decay coefficient of 0.9. The optimizer decay coefficient is 0.9.

[0062] The training epoch is 100, the batch size is set to 8, and the boundary feature discriminator is activated after the 80th training epoch.

[0063] In the later stage of model training, discontinuous or abrupt perception errors often occur at the boundaries between environmental categories. This kind of error reflects the model's insufficient understanding of the boundary features of categories, resulting in mistakes in the subsequent path planning of unmanned ships.

[0064] Therefore, in order to further optimize the boundary perception accuracy and maintain the accuracy of semantic features, the present invention also designs a boundary feature discriminator applicable to the later stage of training, as Figure 2 shown. The boundary feature discriminator quantifies the difference in boundary features by comparing the similarity between the boundary prediction value of the model and the boundary ground truth, and dynamically adjusts the weight of the boundary features in the feature fusion and decoding stages, thereby enhancing the sensitivity and understanding ability of the model to boundary information.

[0065] The boundary feature discriminator utilizes the shape context method, which provides a relatively robust shape matching strategy based on the mutual relationship (such as distance, gradient) between the sampling points of the boundary ground truth and the boundary prediction result.

[0066] First, boundary extraction and boundary point sampling are performed on the prediction output of the maritime environment perception model.

[0067] The specific process of boundary extraction is as follows: Since the classification result of the maritime environment perception output by the model is a grayscale image, the class information of each pixel point needs to be mapped to the corresponding RGB color value for pseudo-color rendering, which is used as the input for boundary extraction. In this embodiment, the Canny operator is used to detect the boundary of the pseudo-color image, and finally the boundary prediction value is obtained.

[0068] Due to the uneven distribution of the boundary point set, the present invention adopts Jitendra's sampling method. This method randomly arranges the point set, calculates the distance matrix between point pairs, and gradually deletes the closest point pairs to ensure the uniform distribution of the sampling points.

[0069] After obtaining the sampling points, the global similarity of each point is calculated first, and the process is as follows: For each sampling point, a polar logarithmic coordinate system is constructed, the neighboring sampling points are mapped to each region, the number of points in each region is counted and normalized, a shape context histogram matrix is generated, and the chi-square statistic is used to quantify the boundary similarity.

[0070] Assume that a total of points are sampled, and each point has angles and distance regions, then the shape context matrices of the boundary ground truth and the boundary prediction value are obtained, denoted as and respectively, and substituted into the chi-square formula to calculate the similarity metric matrix ; .

[0071] Through the above calculation formula, a similarity metric matrix of size N×N for global boundary features is obtained.

[0072] The local direction is used to evaluate the similarity of local features of the boundary, and the process is as follows: First, calculate the gradient components of the boundary ground truth and the boundary prediction result, and extract the gradient values of the contour points to obtain the tangential angle , , and define the tangential angle dissimilarity function as the similarity measure of the local direction ;

[0073]

[0074] .

[0075] Perform weighted summation on the similarity measure matrices , to output the total similarity measure matrix , and the formula is as follows: .

[0076] Where is the weight, which is set to 0.3 in this embodiment.

[0077] Abstract the total similarity measure matrix into a distance matrix, and solve it through the Hungarian algorithm introducing pseudo points to obtain the point pair matching relationship that minimizes the total loss; extract the similarity scores of the sampling points from the matrix according to the matching index .

[0078] Among them, the lower the similarity score between sampling points, the higher the matching rate.

[0079] On the maritime image dataset LaRS, the present invention screened about 2000 groups of prediction results, each group including two types of samples with perception errors and being highly close to the ground truth, and performed similarity matching with the boundary ground truth (sampling points 800). The results are as Figure 6 shown. There are significant differences in the mean similarity scores of the two types of samples, and the calculated similarity threshold Ts is 9.

[0080] In the last stage of the multi-branch cross-channel attention module, the result after element-wise multiplication of the feature input and the attention vector is adjusted in the number of channels through a 1×1 convolutional layer and then concatenated. When the boundary feature discriminator is not activated, the initial channel number ratio of the preset boundary feature, semantic feature, and fusion feature is set to 1:1:2 in this embodiment.

[0081] When the boundary feature discriminator is activated, the channel number ratio of the boundary feature, semantic feature, and fusion feature is dynamically adjusted according to the similarity score and the similarity threshold Ts; if > Ts indicates that the boundary prediction has a poor match with the ground truth. Then, the weight of the boundary feature in the overall feature is increased, and the adjusted ratio is 3:1:1. Otherwise, if S ≤ Ts, the initial channel number ratio remains unchanged.

[0082] Considering that the water surface has complex and dynamically changing characteristics such as waves and reflections, the feature learning is difficult. And in view of the significant differences between the water surface features and the obstacle features, a water-obstacle perception loss is introduced as the feature loss of the network. 。

[0083] When the network learns category features, it is easily interfered by the problem of class imbalance, resulting in a decline in classification performance. Therefore, a weighted focal loss function is introduced. , and the result is calculated with the classification ground truth to obtain the result of this loss function.

[0084] The total loss function of the model is obtained by weighted summation of the three: 。

[0085] Among them, the weights of the loss functions are and , where is 1, is 0.1.

[0086] During the model training process, a water-obstacle perception loss is introduced as the feature loss of the model. Among them, the output of the third Conformer layer is used as the input of the feature loss function. of.

[0087] After the model is trained, the image data of the marine environment perception obtained in real time (without inputting the boundary ground truth) is input into the trained model to obtain the classification result of the marine environment perception. Specifically, the semantic features of the image data input into the model only pass through the semantic feature encoding stream and the dilated convolution pooling pyramid, and finally, after adjusting the number of channels through the convolution layer with a convolution kernel size of 1×1, the classification result of the marine environment perception is output, without going through the boundary feature encoding stream, the boundary-semantic feature fusion stream, and the multi-level attention fusion stream in the fusion feature decoder, as Figure 2 shown.

[0088] Based on the images collected by the on-board camera, the present invention uses an adaptive understanding model of the marine environment with multi-level reinforcement of boundary-semantic features, which well solves the problem of fuzzy boundary perception in the understanding of marine image categories and provides support for the autonomous cruise of USV.

[0089] In addition, to verify the effectiveness of the proposed semantic-boundary multi-level enhanced shipborne camera image understanding model of the present invention, the following experimental procedures are also given, and the model is evaluated on the LaRS dataset and the MODD2 dataset respectively. In the evaluation, three existing models, namely Segformer, YOLOv11m-seg, and WaSRNet, are introduced for comparative analysis. Among them, Segformer realizes a simple and efficient environmental perception model design through the organic combination of Transformer and a lightweight multi-layer perceptron decoder; YOLOv11m-seg continues the end-to-end architecture paradigm, adopts a multi-scale feature pyramid network and a dynamic matching mechanism, and strengthens the feature representation ability while maintaining the lightweight characteristics; WaSRNet, as a dedicated model for maritime environmental perception, incorporates the horizontal plane angle information simulated by inertial sensors on the basis of adopting the ResNet101 deep encoding architecture, improving the perception accuracy of the water surface scene. Through the comparative experiment with the three models, empirical support is provided for the effectiveness of the model designed by the present invention.

[0090] Step 1. Evaluate the model on the LaRS dataset.

[0091] The LaRS dataset collects many waterborne navigation scenarios from publicly available videos, covering scene types such as canals, oceans, lakes, crowded scenes, strong reflections, and low visibility conditions (such as rainy days and night navigation).

[0092] Use evaluation metrics and to evaluate the class perception accuracy and quality of the network, and their definitions are as follows: .

[0093] .

[0094] TP is the number of samples correctly predicted as positive by the model. FP is the number of samples wrongly predicted as positive by the model. FN is the number of samples wrongly predicted as negative by the model. IoU is the intersection over union, which measures the overlap degree between the model prediction result and the true result. C is the total number of classes in the task. is the IoU value of the c-th class, and the intersection over union of the prediction result of this class is calculated separately.

[0095] Table 1 shows the performance of different networks on the LaRS dataset.

[0096] Table 1 Experimental result diagram of the LaRS dataset

[0097] As can be easily seen from Table 1, thanks to the semantic-boundary feature fusion strategy and the supervision mechanism of boundary features proposed in the present invention, although the LaRS dataset contains a large number of complex scenarios with strong reflections and adverse weather conditions, the method of the present invention still performs excellently in various indicators. Specifically, the method of the present invention has achieved significant improvements in terms of mean pixel accuracy (mPA: 96.91%) and mean intersection over union (mIoU: 94.56%) compared to WaSRNet (mPA: 94.91%, mIoU: 93.16%) designed specifically for marine environment perception. At the same time, when dealing with the perception tasks of small target obstacles on the sea (the first row) and inland water surface obstacles (the second and third rows), the method of the present invention can provide clearer and more accurate perception results at the junction of obstacles and the water surface, which provides more accurate support for the water surface path planning of unmanned ships.

[0098] Step 2. Conduct model evaluation on the MODD2 dataset.

[0099] The MODD2 dataset contains 80,828 images collected from the ports of Koper and Strunjan, covering different weather conditions such as sunny, cloudy, foggy days, and data under extreme conditions (such as motion distortion, sunlight, and reflection). The MODS evaluation system is used for testing on the MODD2 dataset, which mainly includes the following indicators: ① Obstacle-water edge localization accuracy 。

[0100] The ground truth of the evaluation dataset contains the water edge position of each picture.

[0101] The obstacle-water edge localization accuracy is defined as: For the mask of the water surface category in the prediction mask, the square root of the average squared distance between the layer of pixels with the closest vertical distance to the water edge ground truth among all its pixels and the water edge ground truth. Its definition is as follows: 。

[0102] Where n is the total number of pixels of the water surface category in the prediction mask, and these pixels are the objects participating in the calculation of the water edge localization accuracy; is the i-th pixel of the water surface category in the prediction mask, and its vertical distance to the layer of pixels closest to the water edge ground truth reflects the position difference between a single pixel and the real water edge. is the average value (arithmetic mean) of all , representing the average level of the distance between the water surface pixels in the prediction mask and the water edge ground truth.

[0103] ② Accuracy 。

[0104] The accuracy of the model for environmental perception is defined as follows: 。

[0105] ③ Recall rate 。

[0106] The recall rate of the model for environmental perception reflects the ability of the model to successfully identify all positive example samples. It is defined as follows: 。

[0107] ④ Error rate 。

[0108] The average number of false positive detections per 100 images.

[0109] ⑤ F1 score 。

[0110] F1 is an index that comprehensively considers precision and recall rate and is used to comprehensively evaluate the performance of the model. It is defined as follows: 。

[0111] Table 2 shows the verification results of different networks on the MODD2 dataset based on the MODS evaluation system.

[0112] Table 2 Experimental result diagram of the MODD2 dataset

[0113] It can be easily seen from Table 2 that although UNet is slightly better than the method of the present invention in terms of Re(95.1%) and TPr(62.1) (Re: 92.2, TPr: 61.4), in terms of FPr, UNet (34.9) is significantly inferior to the method of the present invention (2.4). This advantage not only benefits from the in-depth learning of boundary features and the full fusion of semantic features in the present invention, but also attributes to the introduction of a loss function specifically for separating water surface and obstacle features, which improves the adaptability of the network to complex water environments.

[0114] The present invention introduces the MODS evaluation system and the mIoU index, conducts empirical tests on two public datasets, LaRS and MODD2, and conducts comparative analysis with networks such as WaSRNet. Experimental data shows that the model of the present invention has a result of 94.56% in terms of the mIoU index, The index reaches 16.3 pixels, thus verifying its effectiveness and superiority.

[0115] Of course, the above description is only a preferred embodiment of the present invention. The present invention is not limited to listing the above embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any person skilled in the art under the teaching of this specification fall within the substantial scope of this specification and should be protected by the present invention.

Claims

1. A method for adaptively sensing marine environment based on boundary reinforcement, characterized in that: The steps include: Step 1. Obtain the original image dataset of marine environment perception, perform data enhancement on the original images, and obtain the boundary truth value for each original image. The boundary truth value and the original image correspond one-to-one to form a training dataset; Step 2. Build a semantic-boundary multi-level enhanced shipborne camera image understanding model, which includes a boundary feature encoding stream, a semantic feature encoding stream, a boundary-semantic feature fusion stream, and a fusion feature decoder; The boundary feature encoding stream and the semantic feature encoding stream are used to encode the boundary truth value and the original image of the input model respectively; each encoding stream includes four cascaded feature extraction layers; The boundary-semantic feature fusion flow includes a cascade of four layers of feature attention modules; Each layer of feature attention module is used to receive the output of the corresponding layer of the boundary feature encoding stream and the semantic feature encoding stream, and is used to realize the feature interaction of the boundary feature and the semantic feature of the corresponding layer to enhance the boundary perception ability; The outputs of the feature attention modules at each layer are fed back to the corresponding layers of the semantic feature encoding stream, and together with the outputs of the corresponding layers of the semantic feature encoding stream, serve as the inputs of the lower layers of the semantic feature encoding stream; The output results of the boundary feature encoding stream, semantic feature encoding stream, and boundary-semantic feature fusion stream are input into the fusion feature decoder, and after being enhanced by the three-dimensional feature attention of direction, position, and channel, the marine environment perception classification result is output; Step 3. Use the training data set to train the semantic-boundary multi-level reinforced shipborne camera image understanding model; input the real-time acquired marine environment perception image data into the trained model to obtain the marine environment adaptive perception results.

2. The method for adaptively sensing marine environment based on boundary reinforcement according to claim 1 is characterized in that: In step 1, the training data set is obtained as follows: Firstly, the original image dataset for marine environment perception is obtained, and the original image is preprocessed with data enhancement to expand the amount of original image data; the true value of the dataset is a grayscale image with pixel-by-pixel category annotation; Data augmentation by adjusting image properties, including exposure, contrast, saturation, and / or hue; At the same time, the boundary truth value is obtained for each original image, and the boundary truth value corresponds to the original image one by one; The method for obtaining the boundary truth value is as follows: first, the grayscale image is converted into a multi-category RGB image through pseudo-color coding, and then the Canny edge detection algorithm is applied to extract the boundary, and finally the boundary truth value is obtained.

3. The method for adaptively sensing marine environment based on boundary reinforcement according to claim 1 is characterized in that: The boundary feature encoding stream adopts the ResNet18 network; the 1st to 4th residual convolution layers of the ResNet18 network are used as the four-layer feature extraction layers of the boundary feature encoding stream; the output of each residual convolution layer is respectively input into the boundary-semantic feature fusion stream to realize the step-by-step interaction between the boundary truth value and the semantic feature, and promote bidirectional feature learning; The boundary feature encoding stream is spatially downsampled based on the output size of the last residual convolution layer, and finally the multi-scale boundary feature maps formed by each residual convolution layer are concatenated as the output of the boundary feature encoding stream.

4. The method for adaptively sensing marine environment based on boundary reinforcement according to claim 1 is characterized in that: The semantic feature encoding stream is constructed based on the Conformer-S architecture; the first four Conformer layers of the Conformer-S architecture are used as four feature extraction layers of the semantic feature encoding stream; the output of each Conformer layer is respectively input into the boundary-semantic feature fusion stream to realize the step-by-step interaction between the boundary truth value and the semantic feature, and promote bidirectional feature learning; At the same time, each Conformer layer also receives feedback information from the corresponding layer of the boundary-semantic feature fusion flow; The semantic feature encoding stream is spatially downsampled based on the output size of the last Conformer layer, and finally the multi-scale semantic feature maps formed by each Conformer layer are concatenated as the output of the semantic feature encoding stream.

5. The method for adaptively sensing marine environment based on boundary reinforcement according to claim 1 is characterized in that: The boundary-semantic feature fusion stream includes 4 layers of cascaded feature attention modules, each layer of feature attention modules is used to receive the output of the corresponding layer of the boundary feature encoding stream and the semantic feature encoding stream, and perform feature interaction; The boundary-semantic feature fusion flow is downsampled according to the spatial size of the last layer of feature attention modules, and finally the feature maps output by the feature attention modules of each layer are spliced ​​in the channel dimension to form a multi-scale feature output of the boundary-semantic feature fusion flow.

6. The method for adaptively sensing marine environment based on boundary reinforcement according to claim 1, characterized in that: The processing flow of the feature attention module is as follows: First, the boundary features and semantic features input into the feature attention module are added element by element to form joint feature 1; The joint feature is divided into two branches. One branch introduces the ReLU activation function with nonlinear transformation to enhance the model's ability to express fine boundary structures. The other branch aggregates global semantic information through global average pooling and then activates through ReLU. The two processing results are added element by element to obtain joint feature 2; The joint feature 2 is multiplied element by element by the joint feature 1 to generate the attention weight based on the interactive feature, which is then multiplied with the boundary and semantic features respectively to specifically enhance the effective information of the boundary details and semantic content; Finally, we add them element by element to integrate the positioning details of the boundary features with the abstract understanding of the semantic features to obtain the fused features.

7. The method for adaptively sensing marine environment based on boundary reinforcement according to claim 1 is characterized in that: The fusion feature decoder includes a multi-level attention fusion module and a dilated convolution pooling pyramid module; After receiving the output of the boundary feature encoding stream, the semantic feature encoding stream, and the boundary-semantic feature fusion stream, the multi-level attention fusion module fuses the three features and enhances the model's understanding of the correlation between them; The multi-level attention fusion module includes a multi-branch cross-channel attention module and a collaborative attention module; The multi-branch cross-channel attention module uses the multi-branch cross-channel interaction structure to adaptively learn channel weights, establish a direct association between channels and their weights, and reflect the global contribution of semantic features and boundary features in terms of scale and type dimensions; The collaborative attention module captures the spatial features in two directions and generates an attention map to complementarily enhance the representation of the object of interest in the input feature map to address the shortcomings of the multi-branch cross-channel attention module in processing direction and position information; The dilated convolutional pooling pyramid module is used to receive the output from the multi-level attention fusion module and the semantic feature encoding stream, while expanding the receptive field and ensuring the resolution of the output image; Finally, after adjusting the number of channels through a convolution layer with a convolution kernel size of 1×1, the marine environment perception classification results are output.

8. The method for adaptively sensing marine environment based on boundary reinforcement according to claim 1 is characterized in that: In the step 3, a boundary feature identifier is introduced in the model training process; wherein the marine environment perception classification result output by the semantic-boundary multi-level enhanced shipborne camera image understanding model is a grayscale image; The category information of each pixel in the grayscale image is mapped to the corresponding RGB color value for pseudo-color rendering. As the input of boundary extraction, the Canny operator is used to perform boundary detection on the pseudo-color image to finally obtain the boundary prediction value. Then, Jitendra's sampling method is used to sample boundary points. By randomly arranging the point set, the distance matrix between point pairs is calculated, and the closest point pairs are gradually deleted to ensure that the sampling points are evenly distributed. The boundary feature discriminator compares the similarity between the boundary prediction value and the boundary true value, quantifies the difference of boundary features, and dynamically adjusts the weight of boundary features in the feature fusion and decoding stages to improve the model's sensitivity and understanding of boundary information.

9. The method for adaptively sensing marine environment based on boundary reinforcement according to claim 8, characterized in that: In step 3, after obtaining the sampling points, the global similarity of each point is first calculated, and the process is as follows: Construct an extreme logarithmic coordinate system for each sampling point, map neighboring sampling points to each region, count the number of points in each region and normalize them, generate a shape context histogram matrix, and use chi-square statistics to quantify boundary similarity; Assuming co-sampling points, each point is Angles and distance region, and obtain the boundary truth value and boundary prediction value The shape context matrix is ​​recorded as and , substitute into the chi-square formula to calculate the similarity measure matrix ; Use local directions to evaluate the similarity of local features of the boundary. The process is as follows: First, the gradient component of the boundary true value and the boundary prediction knot value is calculated, and the gradient value of the contour point is extracted to obtain the tangent angle and , define the tangential angle dissimilarity function as the similarity measure of local directions ; The similarity measure matrix , Perform weighted summation and output the total similarity measurement matrix ; The total similarity measurement matrix Abstracted as a distance matrix, and solved by introducing the pseudo-point Hungarian algorithm to obtain the point pair matching relationship that minimizes the total loss; according to the matching index from the matrix Extract the similarity scores of the sampling points from ; Defining Similarity Threshold ,In the final stage of the multi-branch cross-channel attention module, the result of element-wise multiplication of the feature input and the attention vector is concatenated after adjusting the number of channels through a 1×1 convolutional layer; When the boundary feature discriminator is not activated, the ratio of the initial number of channels of the boundary feature, semantic feature, and fusion feature is preset; When the boundary feature discriminator is activated, the ratio of the number of channels of boundary features, semantic features and fusion features is calculated according to and Make dynamic adjustments; if > , increase the weight of boundary features in the overall features; otherwise, keep the ratio of the initial number of channels unchanged.

10. A marine environment adaptive perception system based on boundary reinforcement, carried on a USV, comprising a camera and a computer device; the camera is used to obtain image data of marine environment perception in real time and upload it to the computer device; in, The computer device includes a memory and a processor; An executable code is stored in the memory, characterized in that when the processor executes the executable code, it is used to implement the steps of adaptive perception of the marine environment based on boundary enhancement as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method based on attention multi-scale feature fusion

    CN111127493A

  • Boundary-driven neural network structure for unmanned ship environment understanding

    CN118396071A

  • Three-dimensional point-cloud semantic segmentation method based on multi-level boundary enhancement for unstructured environment

    WO2024230038A1

Cited By

  • Image semantic segmentation optimization method and system based on deep learning

    CN120510391A

  • Molybdenum target image classification system and method

    CN121280811A

  • A molybdenum target image classification system and method

    CN121280811B