Real-time underwater target detection method based on efficient codec
By using an efficient codec in underwater object detection, a full-dimensional dynamic convolution module and an efficient hybrid encoder are built, which solves the problems of category imbalance and low image quality in underwater object detection, and achieves higher detection accuracy and speed.
Patent Information
- Application Number
- CN202510132496.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-30
AI Technical Summary
Underwater target detection faces problems such as insufficient data, low image quality, complexity of target recognition and category imbalance, resulting in low detection accuracy and efficiency.
Real-time underwater object detection method based on high-efficiency codecs is adopted, and the detection accuracy and processing capability of the model are improved by building a full-dimensional dynamic convolution module, a full-dimensional dynamic residual network, an in-scale interaction module and an efficient hybrid encoder, combining a multi-head attention mechanism and a gated dynamic convolution feedforward network.
It effectively alleviates the problem of imbalance in the category of underwater targets, improves the detection model's ability to identify overlapping targets, and significantly improves the overall accuracy and speed of underwater target detection.
Smart Images

Figure CN120071109A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the technical field of underwater target detection, and particularly to a real-time underwater target detection method based on an efficient codec. Background Art
[0002] Traditional marine environment exploration mainly relies on human diving operations, and all tasks related to underwater target detection and recognition must be manually completed by oceanographers. This highly labor-intensive underwater target detection method is not only inefficient and difficult to obtain richer information, but also places extremely high requirements on the professional skills and physical fitness of divers. In the deep-sea environment, it is obviously unrealistic to rely on human operations for long-term exploration. The prosperity of machine vision technology enables humans to remotely operate underwater detection equipment such as submersibles and autonomous underwater vehicles to develop marine ecological resources and digitally simulate the marine environment, making marine biological monitoring more accurate and rapid. With the vigorous development of technology, using optical technology to achieve accurate identification and rapid tracking of underwater targets has become the mainstream technology in the current underwater field. Autonomous underwater robots are even more prominent, capable of better simulating the underwater ecosystem, achieving precise capture of underwater organisms, maintenance of underwater engineering facilities, real-time observation of the underwater environment, etc., thus better supporting people's underwater activities. These technologies are now widely used in various operations in the marine environment, such as aquaculture, offshore fishing, marine species monitoring, underwater archaeology, etc.
[0003] Underwater target detection not only requires identifying the target category but also predicting its spatial position. Due to the complex and variable underwater environment, underwater target detection remains one of the most challenging tasks in machine vision, and the captured images often suffer from severe blurring, color distortion, and reduced visibility. In recent years, driven by the progress of computing power and data availability, artificial intelligence-based underwater target detection algorithms have grown exponentially. Artificial intelligence is a technology that enables computers to simulate human intelligence and has attracted wide attention since its inception in the 1950s. By 1995, with the revival of underwater target detection, artificial intelligence technology began to be integrated into the field of underwater target detection technology. Nowadays, artificial intelligence, especially machine learning, has been widely applied in the field of underwater target detection technology. Among various artificial intelligence technologies, deep learning has attracted great attention and extensive research in the field of underwater target detection technology. Deep learning models require a large amount of training data but have obvious advantages over traditional methods in terms of detection accuracy. There are mainly two research directions for underwater target detection. The first research line is general deep detection networks, such as Faster RCNN, SSD, YOLO and their variants, which have been used for underwater target detection, significantly improving the detection performance. The second research line focuses on developing special network backbones, loss functions or learning strategies for underwater target detection.
[0004] Although the deep learning-based underwater target detection technology has great potential in academic and commercial applications, it still faces many difficulties at present. The first point is the lack of underwater image data. The complex underwater environment has high requirements for underwater imaging devices. Compared with atmospheric optical images, this makes it very difficult to obtain underwater images. Therefore, it is difficult for data-driven deep learning models to achieve satisfactory results. The second point is the low quality of underwater images. Underwater images are usually affected by uneven lighting conditions, blur, low contrast, and color deviation, resulting in much worse performance in terms of color and texture information than atmospheric optical images. The third point is the problem of small and clustered target recognition. Some underwater targets are usually clustered and small in size, making it difficult to extract rich details. The fourth point is the imbalance in the number of different underwater targets. The class imbalance of underwater targets makes it difficult for underwater target detectors to learn the features of classes with a small number of samples, which has a negative impact on their performance. These difficulties pose great challenges to the underwater target detection task. Summary of the Invention
[0005] To solve the above technical problems, an embodiment of the present application proposes a real-time underwater target detection method based on an efficient encoder-decoder, which can alleviate the problem of underwater target class imbalance and improve the ability of the detection model to detect overlapping targets, thereby effectively improving the overall detection accuracy of the detection model.
[0006] To achieve the above object, an embodiment of the present application proposes a real-time underwater target detection method based on an efficient encoder-decoder, the method comprising the following steps: using a multi-dimensional attention mechanism and a parallel strategy to learn the attention of the convolutional kernel along four dimensions of the kernel space in any convolutional layer to construct a full-dimensional dynamic convolution module; wherein the four dimensions are respectively the convolutional kernel number dimension, the spatial dimension of the convolutional kernel, the input channel dimension, and the output channel dimension; replacing the ordinary convolution module in the residual network with the full-dimensional dynamic convolution module to obtain a full-dimensional dynamic residual network; using 1×1 convolution, 3×3 convolution, and Gaussian error linear unit activation function to design a gated dynamic convolution feedforward network, and using a multi-head attention mechanism and the gated dynamic convolution feedforward network to construct an intra-scale interaction module; stacking the intra-scale interaction modules learned across spaces in the form of a feature pyramid to construct an efficient hybrid encoder; combining the full-dimensional dynamic residual network, the efficient hybrid encoder, and a decoder based on target queries to obtain an underwater target detection model, and iteratively training the underwater target detection model based on a pre-acquired training sample set until convergence to obtain a trained underwater target detection model; inputting the data to be recognized into the trained underwater target detection model to obtain the detection result of the data to be recognized output by the trained underwater target detection model.
[0007] To achieve the above object, an embodiment of the present application further provides a real-time underwater target detection system based on an efficient codec. The system includes: a full-dimensional dynamic convolution module construction module, configured to use a multi-dimensional attention mechanism and a parallel strategy to learn the attention of the convolution kernel along four dimensions of the kernel space in any convolutional layer, and construct a full-dimensional dynamic convolution module, where the four dimensions are respectively the convolution kernel number dimension, the spatial dimension of the convolution kernel, the input channel dimension, and the output channel dimension; a full-dimensional dynamic residual network construction module, configured to replace the ordinary convolution module in the residual network with the full-dimensional dynamic convolution module to obtain a full-dimensional dynamic residual network; a scale-in interaction module construction module, configured to design a gated dynamic convolution feedforward network using 1×1 convolution, 3×3 convolution, and Gaussian error linear unit activation function, and construct a scale-in interaction module using a multi-head attention mechanism and the gated dynamic convolution feedforward network; an efficient hybrid encoder construction module, configured to stack the scale-in interaction modules learned across spaces in the form of a feature pyramid to construct an efficient hybrid encoder; an underwater target detection model construction module, configured to combine the full-dimensional dynamic residual network, the efficient hybrid encoder, and a decoder based on target queries to obtain an underwater target detection model; a model training module, configured to iteratively train the underwater target detection model based on a pre-acquired training sample set until convergence to obtain a trained underwater target detection model; a model deployment and usage module, configured to input the data to be recognized into the trained underwater target detection model to obtain the detection result of the data to be recognized output by the trained underwater target detection model.
[0008] To achieve the above object, an embodiment of the present application further provides an electronic device. The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the above-mentioned real-time underwater target detection method based on an efficient codec.
[0009] To achieve the above object, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it can implement the above-mentioned real-time underwater target detection method based on an efficient codec.
[0010] An underwater target detection method based on an efficient codec proposed in an embodiment of the present application constructs a full-dimensional dynamic convolution module, a full-dimensional dynamic residual network, an intra-scale interaction module, and an efficient hybrid encoder in sequence. By combining the full-dimensional dynamic residual network, the efficient hybrid encoder, and a decoder based on target queries, an underwater target detection model can be obtained. After iteratively training the underwater target detection model until convergence, a trained underwater target detection model is obtained and deployed to scenarios in need to perform underwater target detection tasks. The full-dimensional dynamic residual network is responsible for feature extraction. By dynamically adjusting the convolution kernel, it can enhance the expressiveness of the traditional convolutional neural network during the feature extraction process, extract features of different image regions more accurately, thereby alleviating the problem of unbalanced target categories and improving the model's ability to detect overlapping targets. The efficient hybrid encoder combines a Transformer encoder and multi-scale feature extraction, and can effectively process spatial and channel information in images. By introducing a target query mechanism and position encoding, the model can more effectively capture global context information and long-range dependencies in images, further enhancing the richness of feature representation. Constructing, training, deploying, and using such an underwater target detection model for underwater target detection effectively improves the detection speed and detection accuracy.
[0011] Optionally, using a multi-dimensional attention mechanism and a parallel strategy, learn the attention of the convolution kernel along the four dimensions of the kernel space in any convolutional layer, and construct a full-dimensional dynamic convolution module, including:
[0012] Construct a channel attention unit, use a 1×1 convolution on the input to generate a first feature tensor with a shape of [batch_size, in_planes, 1, 1], and then use a sigmoid activation function to normalize the first feature tensor to obtain channel attention weights;
[0013] Construct a filter attention unit, perform adaptive max pooling and adaptive average pooling on the input respectively, extract a second feature tensor and a third feature tensor with a shape of [batch_size, in_planes, 1, 1], use a 1×1 convolution on the second feature tensor and the third feature tensor to generate a max pooling feature and an average pooling feature with a shape of [batch_size, out_planes, 1, 1], add the max pooling feature and the average pooling feature to obtain a shape fusion tensor, and finally use a sigmoid activation function to normalize the fusion tensor to obtain filter attention weights;
[0014] Construct a spatial attention unit, use 1×1 convolution to extract spatial-related features to construct a spatial attention layer. After the input is processed by the spatial attention layer, a first weight vector with the shape of [batch_size, kernel_size*kernel_size, 1, 1] is generated. Reshape the shape of the first weight vector to [batch_size, 1, 1, 1, kernel_size, kernel_size], and finally use the sigmoid activation function for normalization to obtain the spatial attention weight;
[0015] Construct a convolutional kernel attention unit, use 1×1 convolution to extract convolutional kernel-related features to construct a convolutional kernel attention layer. After the input is processed by the convolutional kernel attention layer, a second weight vector with the shape of [batch_size, kernel_num, 1, 1] is generated. Reshape the shape of the second weight vector to [batch_size, kernel_num, 1, 1, 1, 1], and finally use the softmax function for activation to obtain the convolutional kernel attention weight;
[0016] Implement dynamic convolution based on the channel attention weight, filter attention weight, spatial attention weight, and convolutional kernel attention weight, and construct a full-dimensional dynamic convolution module.
[0017] Optionally, implement dynamic convolution based on the channel attention weight, filter attention weight, spatial attention weight, and convolutional kernel attention weight, and construct a full-dimensional dynamic convolution module, including:
[0018] Adjust the feature intensity of each channel based on the channel attention weight, and generate a spatially specific weight matrix for each convolutional kernel based on the spatial attention weight;
[0019] Adjust the shape of the input to [1, batch_size*in_planes, height, width], use the convolutional kernel attention weight to convolve the input, and obtain an intermediate output with the shape of [batch_size*groups, out_planes / / groups, out_height, out_width];
[0020] Adjust the intermediate output back to the standard shape [batch_size, out_planes, out_height, out_width], multiply it element-wise by the filter attention weight to form the final output with the shape of [batch_size, out_planes, out_height, out_width], that is, complete the construction of the full-dimensional dynamic convolution module.
[0021] Optionally, the constructed full-dimensional dynamic convolution module includes two 1×1 full-dimensional dynamic convolution modules and one 3×3 full-dimensional dynamic convolution module. The ordinary convolution module in the residual network is replaced with the full-dimensional dynamic convolution module to obtain the full-dimensional dynamic residual network, including:
[0022] Use two 1×1 full-dimensional dynamic convolution modules and one 3×3 full-dimensional dynamic convolution module to construct a residual convolution module. Among them, assume the shape of the input feature of the residual convolution module is [N, c1, H, W], where N is the batch size, c1 is the number of input channels, and H and W are the height and width of the input feature respectively. The input feature of the residual convolution module first passes through a 1×1 full-dimensional dynamic convolution module for processing, and the output feature map has a shape of [N, c2, H, W]. Then it passes through a 3×3 full-dimensional dynamic convolution module with a stride of s for processing, and the output feature map has a shape of [N, c2, H / s, W / s]. Finally, it passes through a 1×1 full-dimensional dynamic convolution module with a stride of 1 for processing, and the output feature map has a shape of [N, c3, H, W].
[0023] Based on the residual convolution module, construct a feature extraction backbone network, and finally obtain the full-dimensional dynamic residual network. Among them, assume the shape of the input feature of the feature extraction backbone network is [N, c1, H, W]. The input feature of the feature extraction backbone network first passes through a convolution operation with a kernel size of 7×7, a stride of 2, and a padding of 3, and the output feature map has a shape of [N, c2, H / 2, W / 2]. Subsequently, it passes through a max pooling layer with a kernel size of 3×3, a stride of 2, and a padding of 1, and the output feature map has a shape of [N, c2, H / 4, W / 4].
[0024] The full-dimensional dynamic residual network undergoes four stages of processing. The first stage consists of 3 residual convolution modules with a stride of 1, the second stage consists of 4 residual convolution modules with a stride of 2, the third stage consists of 6 residual convolution modules with a stride of 2, and the fourth stage consists of 3 residual convolution modules with a stride of 2.
[0025] Optionally, use 1×1 convolution, 3×3 convolution, and Gaussian error linear unit activation function to design a gated dynamic convolution feedforward network, and use the multi-head attention mechanism and the gated dynamic convolution feedforward network to construct a scale-internal interaction module, including:
[0026] Assume the shape of the input feature is [N, C, H, W], where N is the batch size, C is the number of input channels, and H and W are the height and width of the input feature respectively.
[0027] First, pass the input features through a 2D sine-cosine positional encoding module to generate positional embeddings of shape [1, H×W, C], and flatten the input features into a tensor of shape [N, H×W, C]. Then, pass through a Transformer encoder layer based on the multi-head self-attention mechanism to obtain the globally modeled features;
[0028] Restore the globally modeled features to the original spatial shape [N, C, H, W]. Then, pass the input features through a 1×1 convolution to expand the number of channels by 4 times, and the output is a feature of shape [N, 4×C, H, W]. Then, pass through a 3×3 depth convolution, and the output shape is [N, 4×C, H, W];
[0029] The features output by the depth convolution are divided into two parts in the channel dimension. One part is element-wise multiplied with the other part after passing through the GELU activation to obtain a feature of shape [N, 2×C, H, W]. Then, compress the number of channels back to C through a 1×1 convolution to obtain a restored feature of shape [N, C, H, W]. Finally, add the input features and the restored features through a residual connection and a normalization module to obtain the final feature of shape [N, C, H, W]. Thus, the intra-scale interaction module is constructed.
[0030] Optionally, the entire network of the efficient hybrid encoder is divided into four stages;
[0031] First, perform channel mapping on each input feature through a 1×1 convolutional layer, so that the corresponding output shape becomes [N, hidden_dim, H, W], where hidden_dim is the intermediate number of channels of the network;
[0032] For a specified encoder layer index, extract the features of a specific layer, and the output shape is [N, H, W, hidden_dim];
[0033] Generate sine-cosine positional embeddings for each pixel, with a shape of [N, hidden_dim, H, W]. The input features and the positional encoding are input into the Transformer encoder together. After encoding, the output shape is still [N, H, W, hidden_dim]. Finally, readjust the shape of the encoded features to [N, H, W, hidden_dim];
[0034] Perform top - down fusion on the multi - layer feature maps. The deepest layer of feature maps is directly used as the initial output and fused layer by layer from bottom to top. The high - resolution features of each layer are adjusted in channels through a 1×1 convolution, and then concatenated with the upsampled deeper - layer features, with the output shape remaining unchanged. The shallower features of each layer are downsampled through a 3×3 convolution and then concatenated and fused with the deeper - layer features, with the output shape being the same. The final output is a multi - layer feature map with the shape of [N, hidden_dim, H, W]. The resolution of each layer is the same as that of the corresponding input layer, and the number of feature channels is hidden_dim for all layers.
[0035] Optionally, input the data to be recognized into the trained underwater target detection model to obtain the detection result of the data to be recognized output by the trained underwater target detection model, including:
[0036] Input the data to be recognized into the trained underwater target detection model. The trained underwater target detection model uses a full - dimensional dynamic residual network to extract the features of the data to be recognized, enhances the expression ability of the features through the way of residual learning, then uses an efficient hybrid encoder to capture multi - scale features, and performs sequence modeling through Transformer to further enhance the spatial and channel expression abilities. Finally, it performs target query through a decoder based on target queries to focus on the target regions in the data to be recognized, locates and classifies the targets, and outputs the detection result of the data to be recognized. Brief Description of the Drawings
[0037] To more clearly illustrate the embodiments of the present application or the technical solutions in the related art, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or the related art. Obviously, the following - described drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0038] Figure 1 is a flowchart of a real - time underwater target detection method based on an efficient encoder - decoder provided in an embodiment of the present application;
[0039] Figure 2 is a schematic structural diagram of a full - dimensional dynamic convolution module provided in an embodiment of the present application;
[0040] Figure 3 is a schematic structural diagram of a scale - internal interaction module provided in an embodiment of the present application;
[0041] Figure 4 is a schematic structural diagram of an underwater target detection model provided in an embodiment of the present application;
[0042] Figure 5 It is a schematic diagram of the training effect of the underwater target detection model provided in an embodiment of the present application;
[0043] Figure 6 It is a schematic diagram of the detection result output by the underwater target detection model provided in an embodiment of the present application;
[0044] Figure 7 It is a schematic structural diagram of a real-time underwater target detection system based on an efficient codec provided in another embodiment of the present application;
[0045] Figure 8 It is a schematic structural diagram of an electronic device provided in another embodiment of the present application. Detailed implementation manners
[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. In various embodiments of the present application, many technical details are proposed for readers to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions required to be protected by the present application can be implemented. The division of the following embodiments is only for convenient description and should not constitute any limitation on the specific implementation manners of the present application. The various embodiments can be combined and cross-referenced with each other on the premise of not contradicting each other.
[0047] An embodiment of the present application proposes a real-time underwater target detection method based on an efficient codec, which is applied to an electronic device. Among them, the electronic device can be a terminal or a server. In this embodiment and the following embodiments, the electronic device is described by taking the server as an example. The implementation details of the real-time underwater target detection method based on an efficient codec proposed in this embodiment will be specifically described below. The following content is only the implementation details provided for convenient understanding and is not necessary for implementing this solution.
[0048] The specific process of the real-time underwater target detection method based on an efficient codec proposed in this embodiment can be as Figure 1 shown and includes:
[0049] Step 101, using a multi-dimensional attention mechanism and a parallel strategy, learn the attention of the convolution kernel along the four dimensions of the kernel space in any convolutional layer, and construct a full-dimensional dynamic convolution module.
[0050] In a specific implementation, the server first needs to utilize the multi-dimensional attention mechanism and parallel strategy to learn the attention of the convolutional kernel along the four dimensions of the kernel space in any convolutional layer, and construct a full-dimensional dynamic convolution module. Among them, the four dimensions are the number of convolutional kernels dimension, the spatial dimension of the convolutional kernel, the input channel dimension, and the output channel dimension. That is to say, the server needs to calculate the channel attention, filter attention, spatial attention, and convolutional kernel attention respectively to construct the full-dimensional dynamic convolution module.
[0051] In one example, as Figure 2 shown, the full-dimensional dynamic convolution module is actually composed of a channel attention unit, a filter attention unit, a spatial attention unit, and a convolutional kernel attention unit. The server needs to construct the channel attention unit, the filter attention unit, the spatial attention unit, and the convolutional kernel attention unit in sequence.
[0052] Construct the channel attention unit, use 1×1 convolution on the input to generate the first feature tensor with the shape of [batch_size, in_planes, 1, 1], and then use the sigmoid activation function to normalize the first feature tensor to obtain the channel attention weight.
[0053] Construct the filter attention unit, perform adaptive max pooling and adaptive average pooling on the input respectively to extract the second feature tensor and the third feature tensor with the shape of [batch_size, in_planes, 1, 1], use 1×1 convolution on the second feature tensor and the third feature tensor to generate the max pooling feature and the average pooling feature with the shape of [batch_size, out_planes, 1, 1], add the max pooling feature and the average pooling feature to obtain the shape fusion tensor, and finally use the sigmoid activation function to normalize the fusion tensor to obtain the filter attention weight.
[0054] Construct the spatial attention unit, use 1×1 convolution to extract spatial-related features to construct the spatial attention layer. After the input is processed by the spatial attention layer, generate the first weight vector with the shape of [batch_size, kernel_size * kernel_size, 1, 1], change the shape of the first weight vector to [batch_size, 1, 1, 1, kernel_size, kernel_size], and finally use the sigmoid activation function for normalization to obtain the spatial attention weight.
[0055] Construct a convolutional kernel attention unit. Use 1×1 convolution to extract features related to the convolutional kernel to construct a convolutional kernel attention layer. After the input is processed by the convolutional kernel attention layer, a second weight vector with the shape of [batch_size, kernel_num, 1, 1] is generated. Reshape the shape of the second weight vector to [batch_size, kernel_num, 1, 1, 1, 1], and finally activate it using the softmax function to obtain the convolutional kernel attention weights.
[0056] So far, the channel attention unit, filter attention unit, spatial attention unit, and convolutional kernel attention unit have all been constructed. The server realizes dynamic convolution based on the channel attention weights, filter attention weights, spatial attention weights, and convolutional kernel attention weights, and then the full-dimensional dynamic convolution module can be constructed.
[0057] In an example, the implementation of dynamic convolution needs to be based on the channel attention weights, filter attention weights, spatial attention weights, and convolutional kernel attention weights. The use of the four attention weights is as follows:
[0058] Adjust the feature intensity of each channel based on the channel attention weights.
[0059] Generate a spatially specific weight matrix for each convolutional kernel based on the spatial attention weights.
[0060] Adjust the shape of the input to [1, batch_size * in_planes, height, width], and use the convolutional kernel attention weights to convolve the input to obtain an intermediate output with the shape of [batch_size * groups, out_planes / / groups, out_height, out_width]. The convolutional kernel attention weights dynamically select a combination of multiple convolutional kernels.
[0061] Adjust the intermediate output back to the standard shape [batch_size, out_planes, out_height, out_width], and multiply it element-wise by the filter attention weights to form the final output with the shape of [batch_size, out_planes, out_height, out_width], thus completing the construction of the full-dimensional dynamic convolution module.
[0062] Step 102, replace the ordinary convolution module in the residual network with the full-dimensional dynamic convolution module to obtain a full-dimensional dynamic residual network.
[0063] In a specific implementation, after the server constructs the full-dimensional dynamic convolution module, it needs to replace the ordinary convolution module in the residual network with the full-dimensional dynamic convolution module to obtain a full-dimensional dynamic residual network.
[0064] In one example, the full-dimensional dynamic convolution module constructed by the server includes two 1×1 full-dimensional dynamic convolution modules and one 3×3 full-dimensional dynamic convolution module.
[0065] The server first uses two 1×1 full-dimensional dynamic convolution modules and one 3×3 full-dimensional dynamic convolution module to construct a residual convolution module. Assume the shape of the input feature of the residual convolution module is [N, c1, H, W], where N is the batch size, c1 is the number of input channels, and H and W are the height and width of the input feature respectively. The input feature of the residual convolution module first passes through a 1×1 full-dimensional dynamic convolution module for processing, and the output feature map has a shape of [N, c2, H, W]. Then it passes through a 3×3 full-dimensional dynamic convolution module with a stride of s for processing, and the output feature map has a shape of [N, c2, H / s, W / s]. Finally, it passes through a 1×1 full-dimensional dynamic convolution module with a stride of 1 for processing, and the output feature map has a shape of [N, c3, H, W].
[0066] Next, the server needs to construct a feature extraction backbone network based on the residual convolution module, and finally obtain a full-dimensional dynamic residual network. Assume the shape of the input feature of the feature extraction backbone network is [N, c1, H, W]. The input feature of the feature extraction backbone network first undergoes a convolution operation with a kernel size of 7×7, a stride of 2, and a padding of 3, and the output feature map has a shape of [N, c2, H / 2, W / 2]. Subsequently, it passes through a max-pooling layer with a kernel size of 3×3, a stride of 2, and a padding of 1, and the output feature map has a shape of [N, c2, H / 4, W / 4].
[0067] It should be noted that the full-dimensional dynamic residual network needs to go through four stages of processing. The first stage consists of 3 residual convolution modules with a stride of 1, the second stage consists of 4 residual convolution modules with a stride of 2, the third stage consists of 6 residual convolution modules with a stride of 2, and the fourth stage consists of 3 residual convolution modules with a stride of 2.
[0068] Step 103: Design a gated dynamic convolution feed-forward network using 1×1 convolution, 3×3 convolution, and Gaussian error linear unit activation function, and construct an intra-scale interaction module using the multi-head attention mechanism and the gated dynamic convolution feed-forward network.
[0069] In a specific implementation, in addition to constructing the server's full-dimensional dynamic residual network, the server also needs to design a gated dynamic convolution feed-forward network using 1×1 convolution, 3×3 convolution, and Gaussian error linear unit activation function, and construct an intra-scale interaction module using the multi-head attention mechanism and the gated dynamic convolution feed-forward network.
[0070] In one example, when the server constructs the within-scale interaction module, assume the shape of the input feature is [N, C, H, W], where N is the batch size, C is the number of input channels, and H and W are the height and width of the input feature respectively.
[0071] The server first passes the input feature through a 2D sine-cosine position encoding module to generate a position embedding with the shape of [1, H×W, C], and flattens the input feature into a tensor with the shape of [N, H×W, C]. Then, it passes through a Transformer encoder layer based on the multi-head self-attention mechanism to obtain the globally modeled feature, whose shape remains [N, H×W, C].
[0072] Next, the server restores the globally modeled feature to the original spatial shape [N, C, H, W], then passes the input feature through a 1×1 convolution to expand the number of channels by 4 times, and the output is a feature with the shape of [N, 4×C, H, W]. Then, it passes through a 3×3 depth convolution, and the output shape is [N, 4×C, H, W]. The feature output by the depth convolution is divided into two parts in the channel dimension. One part is multiplied element-wise with the other part after passing through the GELU activation to obtain a feature with the shape of [N, 2×C, H, W]. Then, it compresses the number of channels back to C through a 1×1 convolution to obtain a restored feature with the shape of [N, C, H, W]. Finally, the input feature and the restored feature are added through a residual connection and a normalization module to obtain the final feature with the shape of [N, C, H, W], and thus the within-scale interaction module is constructed.
[0073] In one example, the structure of the within-scale interaction module can be as Figure 3 shown.
[0074] Step 104: Stack the within-scale interaction modules for cross-space learning in the form of a feature pyramid to construct an efficient hybrid encoder.
[0075] In a specific implementation, after the server constructs the within-scale interaction module, it needs to stack the within-scale interaction modules for cross-space learning in the form of a feature pyramid to construct an efficient hybrid encoder.
[0076] The entire network of the efficient hybrid encoder is divided into four stages.
[0077] First, each input feature is subjected to channel mapping through a 1×1 convolutional layer, so that the corresponding output shape becomes [N, hidden_dim, H, W], where hidden_dim is the intermediate number of channels in the network.
[0078] For a specified encoder layer index, extract the features of a specific layer, and the output shape is [N, H, W, hidden_dim];
[0079] Generate sine-cosine positional embeddings for each pixel, with the shape of [N, hidden_dim, H, W]. The input features and the positional encoding are input into the Transformer encoder together. After encoding, the output shape remains [N, H, W, hidden_dim]. Finally, the shape of the encoded features is readjusted to [N, H, W, hidden_dim].
[0080] Perform top-down fusion on multi-layer feature maps. The deepest layer of the feature map is directly used as the initial output and fused layer by layer from bottom to top. The high-resolution features of each layer are adjusted in channels through a 1×1 convolution, and then concatenated with the upsampled deeper-layer features. The output shape remains unchanged. The shallower features of each layer are downsampled through a 3×3 convolution and then concatenated and fused with the deeper-layer features. The output shape is the same. The final output is a multi-layer feature map with the shape of [N, hidden_dim, H, W]. The resolution of each layer is the same as that of the corresponding input layer, and the number of feature channels is hidden_dim.
[0081] Step 105: Combine the full-dimensional dynamic residual network, the efficient hybrid encoder, and the decoder based on object queries to obtain an underwater object detection model, and iteratively train the underwater object detection model based on a pre-obtained training sample set until convergence to obtain a trained underwater object detection model.
[0082] In a specific implementation, after the server constructs the full-dimensional dynamic residual network and the efficient hybrid encoder, it can combine the full-dimensional dynamic residual network, the efficient hybrid encoder, and the decoder based on object queries to obtain an underwater object detection model, and iteratively train the underwater object detection model based on a pre-obtained training sample set until convergence to obtain a trained underwater object detection model.
[0083] Step 106: Input the data to be recognized into the trained underwater object detection model to obtain the detection result of the data to be recognized output by the trained underwater object detection model.
[0084] In an example, the structure of the underwater object detection model is as Figure 4 shown. The server inputs the data to be recognized into the trained underwater object detection model. The trained underwater object detection model uses the full-dimensional dynamic residual network to extract the features of the data to be recognized, enhances the expression ability of the features through the way of residual learning, then uses the efficient hybrid encoder to capture multi-scale features, and performs sequence modeling through the Transformer to further enhance the expression ability of space and channels. Finally, it performs object queries through the decoder based on object queries to focus on the target area in the data to be recognized, performs target localization and classification, and outputs the detection result of the data to be recognized.
[0085] A real-time underwater target detection method based on an efficient codec proposed in this embodiment constructs a full-dimensional dynamic convolution module, a full-dimensional dynamic residual network, an intra-scale interaction module, and an efficient hybrid encoder in sequence. By combining the full-dimensional dynamic residual network, the efficient hybrid encoder, and a decoder based on target queries, an underwater target detection model can be obtained. After iteratively training the underwater target detection model until convergence, a trained underwater target detection model is obtained and deployed to scenarios in need to perform underwater target detection tasks. The full-dimensional dynamic residual network is responsible for feature extraction. By means of dynamic full-dimensional convolution, it enhances the expressiveness of the traditional convolutional neural network in the feature extraction process. By dynamically adjusting the convolution kernel, it can more accurately extract the features of different image regions, thereby alleviating the problem of target category imbalance and improving the model's ability to detect overlapping targets at the same time. The efficient hybrid encoder combines the Transformer encoder and multi-scale feature extraction, and can effectively process the spatial and channel information in the image. By introducing a target query mechanism and position encoding, the efficient hybrid encoder enables the model to more effectively capture the global context information and long-range dependencies in the image, further enhancing the richness of feature representation. Constructing, training, deploying, and using such an underwater target detection model for underwater target detection effectively improves the detection speed and detection accuracy.
[0086] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, it is within the protection scope of this application; adding insignificant modifications or introducing insignificant designs to the algorithm or process, but not changing the core design of its algorithm and process are within the protection scope of this application.
[0087] In one embodiment, to prove the superiority of a real-time underwater target detection method based on an efficient codec proposed in this application, we use RUOD as the experimental dataset and design a comparative experiment to verify the effectiveness of the full-dimensional dynamic convolution network, the efficient hybrid encoder, and the underwater target detection model obtained by combining the full-dimensional dynamic residual network, the efficient hybrid encoder, and a decoder based on target queries.
[0088] According to Figure 5 The training results shown indicate that the underwater target detection model proposed in this application has a fast convergence speed and there is no large loss fluctuation. Figure 6 It is the detection result graph output by the trained underwater target detection model.
[0089] In addition, we also conducted model ablation experiments. As can be seen from Table 1, the full-dimensional dynamic convolutional network is very effective in enhancing the detection accuracy of the model for few-shot classes, thereby optimizing the overall performance of the model. At the same time, the ability of the efficient hybrid encoder to improve the object detection accuracy, especially the improvement in mAP@0.5:0.95, indicates that the underwater object detection model can improve the mAP under different localization accuracy requirements (corresponding to different IoU thresholds).
[0090] Table 1: Results of model ablation experiments
[0091]
[0092] Another embodiment of the present application proposes a real-time underwater object detection system based on an efficient encoder-decoder. The implementation details of the real-time underwater object detection system based on an efficient encoder-decoder proposed in this embodiment will be specifically described below. The following content is only the implementation details provided for convenient understanding and is not necessary for implementing this example. Figure 7 is a schematic structural diagram of a real-time underwater object detection system based on an efficient encoder-decoder proposed in this embodiment. The system includes: a full-dimensional dynamic convolution module construction module 201, a full-dimensional dynamic residual network construction module 202, a scale-in interaction module construction module 203, an efficient hybrid encoder construction module 204, an underwater object detection model construction module 205, a model training module 206, and a model deployment and usage module 207.
[0093] The full-dimensional dynamic convolution module construction module 201 is used to utilize the multi-dimensional attention mechanism and parallel strategy to learn the attention of the convolution kernel along the four dimensions of the kernel space in any convolution layer, and construct a full-dimensional dynamic convolution module, where the four dimensions are the convolution kernel number dimension, the spatial dimension of the convolution kernel, the input channel dimension, and the output channel dimension.
[0094] The full-dimensional dynamic residual network construction module 202 is used to replace the ordinary convolution module in the residual network with a full-dimensional dynamic convolution module to obtain a full-dimensional dynamic residual network.
[0095] The scale-in interaction module construction module 203 is used to design a gated dynamic convolution feedforward network by using 1×1 convolution, 3×3 convolution, and Gaussian error linear unit activation function, and construct a scale-in interaction module by using the multi-head attention mechanism and the gated dynamic convolution feedforward network.
[0096] The efficient hybrid encoder construction module 204 is used to stack the scale-in interaction modules for cross-space learning in the form of a feature pyramid to construct an efficient hybrid encoder.
[0097] The underwater target detection model construction module 205 is used to combine a full-dimensional dynamic residual network, an efficient hybrid encoder, and a decoder based on target queries to obtain an underwater target detection model.
[0098] The model training module 206 is used to iteratively train the underwater target detection model based on a pre-obtained training sample set until convergence, obtaining a trained underwater target detection model.
[0099] The model deployment and usage module 207 is used to input the data to be recognized into the trained underwater target detection model to obtain the detection result of the data to be recognized output by the trained underwater target detection model.
[0100] It is worth mentioning that each module involved in this embodiment is a logical module. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or implemented as a combination of multiple physical units. In addition, to highlight the innovative part of this application, units that are not closely related to solving the technical problems proposed in this application are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.
[0101] It is not difficult to find that this embodiment is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above method embodiment. The relevant technical details and technical effects mentioned in the above embodiments are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiments.
[0102] Another embodiment of this application proposes an electronic device, and its specific structure is as Figure 8 shown, including: at least one processor 301; and a memory 302 communicatively connected to the at least one processor 301; wherein, the memory 302 stores instructions executable by the at least one processor 301, and the instructions are executed by the at least one processor 301 to enable the at least one processor 301 to execute a real-time underwater target detection method based on an efficient codec as described in the above embodiments.
[0103] Among them, the memory and the processor can be connected in a bus manner. The bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and the memory together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, etc., which are well known in the art and will not be further described herein. The bus interface is responsible for providing an interface between the bus and the transceiver. The transceiver can be a single component or multiple components, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices over a transmission medium. The data processed by the processor is transmitted over a wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.
[0104] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory can be used to store the data used by the processor when executing operations.
[0105] Another embodiment of the present application proposes a computer-readable storage medium storing a computer program, which when executed by a processor, can implement a real-time underwater target detection based on an efficient codec as described in the above method embodiments.
[0106] That is, those skilled in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by instructing relevant hardware through a program. This program is stored in a storage medium and includes several instructions to enable a device (such as a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, ROM (Read-Only Memory), RAM (Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.
[0107] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application, and in practical applications, various changes can be made in form and details without departing from the spirit and scope of the present application.
Claims
1. A real-time underwater target detection method based on an efficient codec, characterized in that: include: Using multi-dimensional attention mechanism and parallel strategy, the attention of convolution kernel is learned along four dimensions of kernel space in any convolution layer to build a full-dimensional dynamic convolution module; the four dimensions are the number of convolution kernels, the spatial dimension of convolution kernels, the input channel dimension, and the output channel dimension; Use the full-dimensional dynamic convolution module to replace the ordinary convolution module in the residual network to obtain a full-dimensional dynamic residual network; Using 1×1 convolution, 3×3 convolution and Gaussian error linear unit activation function, a gated dynamic convolution feedforward network is designed, and a multi-head attention mechanism and a gated dynamic convolution feedforward network are used to construct an intra-scale interaction module. The scale interaction modules learned across spaces are stacked in a feature pyramid to construct an efficient hybrid encoder. The full-dimensional dynamic residual network, the efficient hybrid encoder, and the target query-based decoder are combined to obtain an underwater target detection model, and the underwater target detection model is iteratively trained based on a pre-acquired training sample set until convergence to obtain a trained underwater target detection model; The data to be identified is input into the trained underwater target detection model to obtain the detection result of the data to be identified output by the trained underwater target detection model.
2. A real-time underwater target detection method based on an efficient codec according to claim 1, characterized in that: Using multi-dimensional attention mechanism and parallel strategy, we learn the attention of convolution kernels along four dimensions of kernel space in any convolution layer and build a full-dimensional dynamic convolution module, including: Construct a channel attention unit, use 1×1 convolution on the input, generate the first feature tensor with shape [batch_size, in_planes, 1, 1], and then use the sigmoid activation function to normalize the first feature tensor to get the channel attention weight; Construct a filter attention unit, perform adaptive maximum pooling and adaptive average pooling on the input, extract the second and third feature tensors with shapes of [batch_size, in_planes, 1, 1], use 1×1 convolution on the second and third feature tensors to generate maximum pooling features and average pooling features with shapes of [batch_size, out_planes, 1, 1], add the maximum pooling features and average pooling features to obtain a shape fusion tensor, and finally use the sigmoid activation function to normalize the fusion tensor to obtain the filter attention weight; Construct a spatial attention unit, use 1×1 convolution to extract spatially related features to construct a spatial attention layer, and after the input is processed by the spatial attention layer, generate a first weight vector with a shape of [batch_size, kernel_size*kernel_size, 1,1], change the shape of the first weight vector to [batch_size, 1,1,1, kernel_size, kernel_size], and finally use the sigmoid activation function for normalization to obtain the spatial attention weight; Construct a convolution kernel attention unit, use 1×1 convolution to extract convolution kernel related features to construct a convolution kernel attention layer, and generate a second weight vector of shape [batch_size, kernel_num, 1, 1] after the input is processed by the convolution kernel attention layer. Reshape the second weight vector to [batch_size, kernel_num, 1, 1, 1, 1], and finally activate it with the softmax function to obtain the convolution kernel attention weight. Dynamic convolution is implemented based on channel attention weights, filter attention weights, spatial attention weights, and convolution kernel attention weights to construct a full-dimensional dynamic convolution module.
3. A real-time underwater target detection method based on an efficient codec according to claim 2, characterized in that: Dynamic convolution is implemented based on channel attention weights, filter attention weights, spatial attention weights, and convolution kernel attention weights to construct a full-dimensional dynamic convolution module, including: Adjust the feature strength of each channel based on the channel attention weight, and generate a space-specific weight matrix for each convolution kernel based on the spatial attention weight; Reshape the input to [1, batch_size*in_planes, height, width], convolve the input using the convolution kernel attention weight, and get an intermediate output of shape [batch_size*groups,out_planes / / groups,out_height,out_width]; Adjust the intermediate output back to the standard shape [batch_size, out_planes, out_height, out_width], and multiply the filter attention weight element by element to form the final output of shape [batch_size, out_planes, out_height, out_width], thus completing the construction of the full-dimensional dynamic convolution module.
4. The real-time underwater target detection method based on an efficient codec according to claim 1, characterized in that: The constructed full-dimensional dynamic convolution module includes two 1×1 full-dimensional dynamic convolution modules and one 3×3 full-dimensional dynamic convolution module. The full-dimensional dynamic convolution module is used to replace the ordinary convolution module in the residual network to obtain a full-dimensional dynamic residual network, including: Use two 1×1 full-dimensional dynamic convolution modules and one 3×3 full-dimensional dynamic convolution module to construct a residual convolution module; where the shape of the input feature of the residual convolution module is assumed to be [N, c1, H, W], N is the batch size, c1 is the number of input channels, H and W are the height and width of the input feature respectively, the input feature of the residual convolution module is first processed by a 1×1 full-dimensional dynamic convolution module, and the output shape is a [N, c2, H, W] feature map, and then processed by a 3×3 full-dimensional dynamic convolution module with a step size of s, and the output shape is a [N, c2, H / s, W / s] feature map, and finally processed by a 1×1 full-dimensional dynamic convolution module with a step size of 1, and the output shape is a [N, c3, H, W] feature map; A feature extraction backbone network is constructed based on the residual convolution module, and finally a full-dimensional dynamic residual network is obtained; wherein, the shape of the input feature of the feature extraction backbone network is assumed to be [N, c1, H, W], and the input feature of the feature extraction backbone network first undergoes a convolution operation with a kernel size of 7×7, a step size of 2, and a padding of 3, and the output shape is a feature map of [N, c2, H / 2, W / 2], and then passes through a maximum pooling layer with a kernel size of 3×3, a step size of 2, and a padding of 1, and the output shape is a feature map of [N, c2, H / 4, W / 4]; The full-dimensional dynamic residual network performs four stages of processing. The first stage consists of 3 residual convolution modules with a step size of 1, the second stage consists of 4 residual convolution modules with a step size of 2, the third stage consists of 6 residual convolution modules with a step size of 2, and the fourth stage consists of 3 residual convolution modules with a step size of 2.
5. The real-time underwater target detection method based on an efficient codec according to claim 1, characterized in that: Using 1×1 convolution, 3×3 convolution and Gaussian error linear unit activation function, a gated dynamic convolution feedforward network is designed, and a multi-head attention mechanism and a gated dynamic convolution feedforward network are used to construct an intra-scale interaction module, including: Assume the shape of the input feature is [N, C, H, W], where N is the batch size, C is the number of input channels, and H and W are the height and width of the input feature respectively; First, the input features are passed through a 2D sine-cosine positional encoding module to generate a positional embedding with a shape of [1, H×W, C], and the input features are flattened into a tensor with a shape of [N, H×W, C]. Then, they are passed through a Transformer encoder layer based on a multi-head self-attention mechanism to obtain the features after global modeling. The features after global modeling are restored to the original spatial shape [N, C, H, W]. Then the input features are subjected to a 1×1 convolution to expand the number of channels by 4 times and output features with a shape of [N, 4×C, H, W]. Then, a 3×3 depth convolution is performed to output features with a shape of [N, 4×C, H, W]. The features output by the deep convolution are divided into two parts in the channel dimension. One part is activated by GELU and then multiplied element-by-element with the other part to obtain a feature with a shape of [N, 2×C, H, W]. The number of channels is then compressed back to C through a 1×1 convolution to obtain a restored feature with a shape of [N, C, H, W]. Finally, the input feature is added to the restored feature through the residual connection and normalization module to obtain the final feature with a shape of [N, C, H, W]. The in-scale interaction module is thus constructed.
6. A real-time underwater target detection method based on an efficient codec according to claim 5, characterized in that: The whole network of the efficient hybrid encoder is divided into four stages; First, each input feature is channel mapped through a 1×1 convolution layer, so that the corresponding output shape becomes [N, hidden_dim, H, W], where hidden_dim is the number of intermediate channels of the network; For the specified encoder layer index, extract the features of a specific layer, and the output shape is [N, H, W, hidden_dim]; Generate sine-cosine position embedding for each pixel, with a shape of [N, hidden_dim, H, W]. The input feature and position encoding are input into the Transformer encoder together. The output shape after encoding is still [N, H, W, hidden_dim]. Finally, the shape of the encoded feature is resized to [N, H, W, hidden_dim]. The multi-layer feature maps are fused from top to bottom. The deepest feature map is directly used as the initial output and fused layer by layer from bottom to top. The high-resolution features of each layer are channel-adjusted through a 1×1 convolution, and then concatenated with the upsampled deeper features. The output shape remains unchanged. The shallower features of each layer are downsampled through a 3×3 convolution, and then concatenated and fused with the deeper features. The output shape is the same. The final output is a multi-layer feature map with a shape of [N, hidden_dim, H, W]. The resolution of each layer is the same as the corresponding layer of the input, and the number of feature channels is hidden_dim.
7. A real-time underwater target detection method based on an efficient codec according to any one of claims 1 to 6, characterized in that: The data to be identified is input into the trained underwater target detection model, and the detection result of the data to be identified output by the trained underwater target detection model is obtained, including: The data to be identified is input into the trained underwater target detection model. The trained underwater target detection model uses a full-dimensional dynamic residual network to extract the features of the data to be identified, and enhances the feature expression capability through residual learning. Then, an efficient hybrid encoder is used to capture multi-scale features, and sequence modeling is performed through Transformer to further enhance the spatial and channel expression capabilities. Finally, a target query-based decoder is used to perform target query to focus on the target area in the data to be identified, locate and classify the target, and output the detection results of the data to be identified.
8. A real-time underwater target detection system based on an efficient codec, characterized in that: include: The full-dimensional dynamic convolution module building module is used to learn the attention of the convolution kernel along the four dimensions of the kernel space at any convolution layer using the multi-dimensional attention mechanism and parallel strategy to build a full-dimensional dynamic convolution module, where the four dimensions are the number of convolution kernels, the spatial dimension of the convolution kernel, the input channel dimension, and the output channel dimension; A full-dimensional dynamic residual network construction module is used to replace the ordinary convolution module in the residual network with the full-dimensional dynamic convolution module to obtain a full-dimensional dynamic residual network; The intra-scale interaction module construction module is used to design a gated dynamic convolutional feedforward network using 1×1 convolution, 3×3 convolution and Gaussian error linear unit activation function, and use the multi-head attention mechanism and the gated dynamic convolutional feedforward network to build the intra-scale interaction module; An efficient hybrid encoder construction module is used to stack the scale interaction modules of cross-space learning in the form of feature pyramids to construct an efficient hybrid encoder; An underwater target detection model building module is used to combine a full-dimensional dynamic residual network, an efficient hybrid encoder, and a target query-based decoder to obtain an underwater target detection model; A model training module is used to iteratively train the underwater target detection model based on a pre-acquired training sample set until convergence, thereby obtaining a trained underwater target detection model; The model deployment module is used to input the data to be identified into the trained underwater target detection model, and obtain the detection result of the data to be identified output by the trained underwater target detection model.
9. An electronic device, characterized in that: include: at least one processor; and, a memory communicatively coupled to the at least one processor; Wherein, the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a real-time underwater target detection method based on an efficient codec as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it can implement a real-time underwater target detection method based on a high-efficiency codec as described in any one of claims 1 to 7.