An Underwater Optical Image Target Detection Method Based on Metric Learning Hybrid Network

By building a hybrid network based on metric learning, the residual global module and the front-background encoding module extract the underwater optical image features are solved, and the problem of foreground confusion and occlusion target detection in the underwater environment is improved, and the reliability and robustness of the detection is improved, which is suitable for underwater robot systems.

CN116630788BActive Publication Date: 2025-07-11NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310424311.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-20
Publication Date
2025-07-11
Estimated Expiration
2043-04-20

AI Technical Summary

Technical Problem

The existing underwater optical image object detection methods have problems of confusion between prospects and background in underwater environments, and it is difficult to effectively extract the contextual relationship of the target, resulting in poor detection results, especially for unsatisfactory detection effects for occluding targets.

Method used

A hybrid network based on metric learning is constructed, including residual global module and pre-background coding module, extract the features of underwater optical images through the global self-attention mechanism and sparse converter, use metric learning ideas to distinguish the foreground and background, and obtain target position and category information through position coding.

Benefits of technology

It effectively solves the problem of confusion between the foreground and background in underwater optical images, improves the reliability and robustness of the model, can better detect occlusion targets, and improves the intelligence and operation reliability of the underwater robot system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116630788B_ABST
    Figure CN116630788B_ABST
Patent Text Reader

Abstract

The present invention discloses an underwater optical image target detection method based on metric learning. First, a residual global module composed of two hybrid modules and two global self-attention modules is constructed for feature extraction of underwater optical images. Then, a foreground-background encoding module and a position encoding module are constructed for classification. Furthermore, a hybrid network based on metric learning is constructed using the residual global module, the foreground-background encoding module, and the position encoding module. The underwater optical image to be detected is input into the trained hybrid network, and finally the detection result is obtained to achieve target detection in underwater optical images. The present invention effectively solves the problem of foreground-background confusion often occurring in underwater optical images in the model, and improves the reliability and robustness of the model in underwater optical image target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of pattern recognition, and particularly relates to an underwater optical image target detection method. Background Art

[0002] With the development of marine technology, target detection has gradually become an important technology in the marine economy. In fishery and aquaculture, underwater unmanned robots equipped with optical cameras can replace humans to complete many high-risk tasks. Among the many functions of underwater unmanned robots, target detection is one of its core functions. Compared with general target detection in the atmospheric environment, due to the particularity of the underwater environment, underwater optical target detection still faces many difficulties. Due to the influence of forward scattering, backscattering, water absorption, and turbidity interference, underwater optical images often show serious degradation. Scattering causes the image to be blurred and the target features to degenerate. The window effect causes underwater optical images to show a blue-green bias, and the target is similar to the background, further increasing the recognition difficulty, resulting in easy confusion between the background and the foreground target, causing false alarms and misjudgments.

[0003] Although the existing optical detection methods based on convolutional neural networks have achieved quite satisfactory results, the model does not specifically distinguish between the foreground and the background. The detection effect of targets similar to the background and aggregated and occluded targets is usually not good. The important source of the error rate is the confusion error between the foreground and the background. At the same time, the existing models for underwater optical image detection are all based on convolutional neural networks and are usually not good at extracting long-range relationships such as the context of the target, resulting in poor detection effects for occluded targets. Summary of the Invention

[0004] In order to overcome the deficiencies of the prior art, the present invention provides an underwater optical image target detection method based on metric learning. First, a residual global module composed of two hybrid modules and two global self-attention modules is constructed for feature extraction of underwater optical images; then a foreground-background encoding module and a position encoding module are constructed for classification; then a hybrid network based on metric learning is constructed using the residual global module, the foreground-background encoding module, and the position encoding module; the underwater optical image to be detected is input into the trained hybrid network, and finally the detection result is obtained to achieve target detection in underwater optical images. The present invention effectively solves the problem of foreground-background confusion often occurring in underwater optical images and improves the reliability and robustness of the model in underwater optical image target detection.

[0005] The technical solutions adopted by the present invention to solve its technical problems include the following steps:

[0006] Step 1: Construct a residual global module composed of two hybrid modules and two global self-attention modules for feature extraction of underwater optical images;

[0007] The input image is simultaneously input into the query end of the first global self-attention module and the first hybrid module; the output of the first hybrid module is simultaneously input into the index and value ends of the first global self-attention module; the output of the first global self-attention module and the output of the first hybrid module are fused additively and input into the second hybrid module; the output of the first global self-attention module serves as the index and value ends of the second global self-attention module; the output of the second hybrid module serves as the query end of the second global self-attention module; the output of the second hybrid module, the second global self-attention module, and the input image are fused additively and used as the final output of the residual global module.

[0008] The hybrid module is divided into two stages. In the first stage, two consecutive low-rank convolutions of 1x3 and 3x1 and a non-linear layer are used to extract the initial rectangular feature map of the image. In the second stage, the rectangular feature map is first converted into a feature column vector through a deformation layer, then input into the global attention mechanism layer for global feature extraction. After passing through the non-linear layer GLU, the global feature is deformed into a rectangular feature map through another deformation layer. A 1x1 point convolution is used to connect the two stages, and the output of each stage is directly added to the input of this stage.

[0009] Step 2: Construct a foreground-background encoding module and a position encoding module.

[0010] The feature map is simultaneously input into the parallel foreground-background sparse transformer and classification sparse transformer. The foreground-background sparse transformer and the classification sparse transformer do not share weights but have the same structure, both of which are composed of 4 basic modules in series, which are sequentially composed of 3x3 depthwise separable convolution, grouped point convolution, and non-linear layer. After the feature map X passes through the two transformers respectively, a foreground-background discrimination feature map and an intermediate feature map are generated. Then, the foreground-background discrimination feature map and the intermediate feature map are multiplicatively fused, and after passing through a 3x3 convolution, a classification feature map is obtained for classification.

[0011] The position encoding module has the same structure as the foreground-background encoding module.

[0012] Step 3: Construct a metric learning-based hybrid network.

[0013] The image is first input into a backbone network composed of 4 residual global modules in sequence to extract high-level features. Then, the generated high-level features are respectively sent into the foreground-background encoding module and the position encoding module for semantic transformation and position encoding, and finally the target position and category information are obtained.

[0014] Step 4: Train the hybrid network based on metric learning. By adjusting the hyperparameters of the hybrid network, improve the stability of network learning; obtain the trained hybrid network;

[0015] Step 5: Input the underwater optical image to be detected into the trained hybrid network, and finally obtain the detection result to achieve target detection in the underwater optical image.

[0016] Preferably, the hyperparameters include learning rate, number of channels, number of network training rounds, training batch size, input image size, and optimizer.

[0017] The beneficial effects of the present invention are as follows:

[0018] 1. The present invention provides a method for underwater optical image target detection based on a hybrid network of metric learning. This method extracts the global context information of the target through the hybrid network, and can obtain robust features of the target, especially occluded targets. The model can better handle the situation of a large number of target occlusions in actual underwater optical images. In addition, the present invention uses the idea of metric learning to propose foreground and background encoders, effectively solving the problem of foreground and background confusion that often occurs in underwater optical images, and improving the reliability and robustness of the model in underwater optical image target detection.

[0019] 2. The present invention can be used in an underwater robot system. By using the provided underwater optical image target detection model sensitive to foreground and background, continuously perceive the object categories and spatial positions to be grasped and operated in the environment, making the underwater robot more intelligent, improving the reliability and accuracy of the underwater robot during operation, avoiding potential dangers caused by foreground and background confusion, and reducing possible economic losses. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a schematic diagram of the foreground-background encoding module of the present invention.

[0021] Figure 2 It is a schematic diagram of the foreground-background sparse converter of the present invention.

[0022] Figure 3 It is a schematic diagram of the hybrid module of the present invention.

[0023] Figure 4 It is a schematic diagram of the residual global module of the present invention.

[0024] Figure 5 It is a schematic diagram of the detection process of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0025] The present invention will be further described below with reference to the drawings and embodiments.

[0026] The present invention designs an underwater optical target detection method that can meet the actual application requirements to achieve robust detection of target objects in underwater optical images. Aiming at the problem of insufficient robustness of foreground and background features in current underwater optical target detection algorithms, a method for underwater optical image target detection based on metric learning and a hybrid network is proposed. To solve the problem of similar foreground and background features, the present invention proposes a foreground-background encoding module; to solve the problem of non-robust features of underwater occluded targets, the present invention proposes a residual global module to extract the global context relationship of the targets, improving the robustness and generalization ability of the model for detecting occluded targets.

[0027] The present invention adopts the idea of high-dimensional embedding and uses contrast mapping to map foreground and background objects to different feature subspaces respectively to achieve highly reliable separation of foreground and background. Through the organic integration of convolution and self-attention mechanism, the complete global context long-range features of occluded targets are obtained, realizing a more robust feature representation and more accurate positioning for the case where the target is occluded. The main steps of the present invention are as follows:

[0028] Step 1: Construct a residual global module composed of two hybrid modules and two global self-attention modules for feature extraction of underwater optical images;

[0029] The input image is simultaneously input into the query end of the first global self-attention module and the first hybrid module; the output of the first hybrid module is simultaneously input into the index and value ends of the first global self-attention module; the output of the first global self-attention module and the output of the first hybrid module are input into the second hybrid module by additive fusion; the output of the first global self-attention module serves as the index and value ends of the second global self-attention module; the output of the second hybrid module serves as the query end of the second global self-attention module; the output of the second hybrid module, the second global self-attention module, and the input image are additively fused and then used as the final output of the residual global module;

[0030] The hybrid module is divided into two stages. In the first stage, the initial rectangular feature map of the image is extracted through two consecutive low-rank convolutions of 1x3 and 3x1 and a non-linear layer; in the second stage, the rectangular feature map is first converted into a feature column vector through a deformation layer, then input into the global attention mechanism layer for global feature extraction, and after passing through the non-linear layer GLU, the global features are deformed into a rectangular feature map through another deformation layer; a 1x1 point convolution is used to connect the two stages, and the output of each stage is directly added to the input of this stage;

[0031] Step 2; Construct a foreground-background encoding module and a position encoding module;

[0032] Input the feature map into the parallel foreground-background sparse transformer and classification sparse transformer simultaneously; the foreground-background sparse transformer and the classification sparse transformer do not share weights, but have the same structure, both of which are composed of 4 basic modules connected in series, which are composed of 3x3 depthwise separable convolution, grouped point convolution and non-linear layer in sequence; after the feature map X passes through the two transformers respectively, a foreground-background discrimination feature map and an intermediate feature map are generated, and then the foreground-background discrimination feature map and the intermediate feature map are multiplicatively fused, and after a 3x3 convolution, a classification feature map is obtained for classification;

[0033] The position encoding module has the same structure as the foreground-background encoding module;

[0034] Step 3: Construct a metric learning-based hybrid network;

[0035] Input the image into the backbone network composed of 4 residual global modules in sequence first to extract high-level features; then send the generated high-level features into the foreground-background encoding module and the position encoding module respectively for semantic transformation and position encoding, and then obtain the final target position and category information;

[0036] Step 4: Train the metric learning-based hybrid network, and improve the stability of network learning by adjusting the hyperparameters of the hybrid network, including the number of channels, the input image size, and the optimizer; obtain the trained hybrid network;

[0037] Step 5: Input the underwater optical image to be detected into the trained hybrid network, and finally obtain the detection result to realize the target detection in the underwater optical image. Specific embodiment:

[0039] The present invention provides an underwater optical target detection method based on a metric learning-based hybrid network. This method uses a foreground-background encoder to realize the display encoding of the foreground and background, and then effectively distinguishes the foreground and background targets, avoiding the confusion between the foreground target and the background. In addition, a basic network composed of a residual global module and a hybrid module is introduced, which can effectively improve the model's ability to extract target information.

[0040] Specifically, the foreground-background encoding module is as Figure 1 shown. In Figure 1 , for a given feature map input X, it is sent into the foreground-background sparse transformer and the classification sparse transformer at the same time. The foreground-background sparse transformer and the classification sparse transformer do not share weights, but have the same structure, both of which are composed of 4 basic modules connected in series, which are composed of 3x3 depthwise separable convolution and grouped point convolution, as Figure 2As shown, after the feature map X passes through two transformers, under the influence of the corresponding loss function, two intermediate products, namely the foreground / background discrimination feature map and the intermediate feature map, are generated. Further, the two are multiplicatively fused, and then fused through a 3x3 ordinary convolution to obtain a classification feature map with foreground / background sensitivity characteristics for classification.

[0041] To further improve the detection performance of the model in complex environments, a residual global module composed of two parallel structures, namely a hybrid module and a global self-attention module, is introduced. The hybrid module is as Figure 3 shown. In Figure 3 , for an image input I, it first passes through two consecutive low-rank convolutions of 1x3 and 3x1 and a non-linear layer to extract preliminary local features. Then, in the second stage, the deformable layer is used to convert the rectangular feature map into a feature column vector, which is then input into the global attention mechanism layer for global feature extraction. After passing through the non-linear layer GLU, the column vector is deformed into a rectangular feature map through another deformable layer. Note that a 1x1 point convolution is used to connect these two stages, and each stage is added to its own input.

[0042] In the residual global module, in addition to the hybrid module, the global self-attention module makes an important contribution to underwater target detection in complex environments. As can be seen from Figure 4 , for an image input I, in addition to entering the hybrid module, it also enters the query input of the global self-attention module that is parallel to it. In the residual global module, the hybrid module and the global attention module appear alternately, and in the global attention module, the two inputs of query and value, the output of the hybrid module and the output of the global self-attention module alternately enter the next global self-attention module. In addition, the output of the global self-attention module also uses additive fusion to further fuse with the hybrid module, thereby obtaining the global feature signal in the underwater optical image. Each residual global branch consists of 2 global self-attention modules and 2 hybrid modules, and the two structures are parallel, and the input and output of the residual global branch are also short-connected.

[0043] The overall process is as shown in Figure 5. For an input image, it is first input into the backbone network composed of 4 residual global modules to extract high-level features. Then, the generated intermediate high-level features are respectively sent into the foreground / background encoding module and the position encoding module for semantic conversion and position encoding, thereby obtaining the final target position and category information. Note that the position encoding module has the same structure as the foreground / background sparse transformer.

[0044] By setting the hyperparameters of the hybrid network, the performance of the model can be changed. Specifically, set the learning rate, the number of network training epochs, the training batch size, the optimization algorithm, etc.

[0045] In the training of the hybrid network based on metric learning, the learning rate is set to 0.01, the momentum parameter is set to 0.09, the training batch size is 32, and the total number of training epochs is 237. The mini-batch stochastic gradient descent optimization method is used to gradually optimize the model, and the learning rate is reduced by 0.1 times at the 200th and 230th epochs until convergence. The network after training can perform object detection on the input underwater optical image.

Claims

1. A method for underwater optical image target detection based on metric learning, characterized in that It includes the following steps: Step 1: Construct a residual global module composed of two hybrid modules and two global self-attention modules for feature extraction of underwater optical images to obtain feature maps; The input image is simultaneously input into the query end of the first global self-attention module and the first hybrid module; the output of the first hybrid module is simultaneously input into the index and value ends of the first global self-attention module; the output of the first global self-attention module and the output of the first hybrid module are input into the second hybrid module through additive fusion; the output of the first global self-attention module serves as the index and value ends of the second global self-attention module; the output of the second hybrid module serves as the query end of the second global self-attention module; the output of the second hybrid module, the second global self-attention module, and the input image are additively fused and used as the final output feature map of the residual global module; The hybrid module is divided into two stages. In the first stage, two consecutive 1x3 and 3x1 low-rank convolutions and a non-linear layer are used to extract the initial rectangular feature map of the image. In the second stage, the rectangular feature map is first converted into a feature column vector through a deformation layer, then input into the global attention mechanism layer for global feature extraction. After passing through the non-linear layer GLU, the global feature is deformed into a rectangular feature map through another deformation layer. A 1x1 point convolution is used to connect the two stages, and the output of each stage is directly added to the input of this stage; Step 2; Construct a foreground-background encoding module and a position encoding module; The feature map is simultaneously input into the parallel foreground-background sparse transformer and classification sparse transformer; the foreground-background sparse transformer and the classification sparse transformer do not share weights but have the same structure, which is composed of 4 basic modules sequentially formed by 3x3 depthwise separable convolution, grouped point convolution, and non-linear layer in series. After the feature map passes through the two transformers respectively, a foreground-background discrimination feature map and an intermediate feature map are generated. Then, the foreground-background discrimination feature map and the intermediate feature map are multiplicatively fused and, after passing through a 3x3 convolution, a classification feature map is obtained for classification; The structure of the position encoding module is the same as that of the foreground-background encoding module; Step 3: Construct a metric learning-based hybrid network; The image is first input into a backbone network composed of 4 residual global modules in series to extract high-level features; then the generated high-level features are respectively sent into the foreground-background encoding module and the position encoding module for semantic transformation and position encoding, and finally the target position and category information are obtained; Step 4: Train the metric learning-based hybrid network, and improve the stability of network learning by adjusting the hyperparameters of the hybrid network; Obtain the trained hybrid network; Step 5: Input the underwater optical image to be detected into the trained hybrid network, and finally obtain the detection result to realize the target detection in the underwater optical image.

2. The underwater optical image target detection method based on metric learning hybrid network according to claim 1, characterized in that The hyperparameters include learning rate, number of channels, number of network training rounds, training batch size, input image size, and optimizer.

Citation Information

Patent Citations

  • A target detection method base on alternately updated densely connected slave zero training network

    CN109376576A

  • Two-stage interactive multi-modal hybrid encoder and encoding method for multi-modal neural machine translation

    CN115034235A