Underwater fish target detection method and device
Through the target recognition model of the backbone network and the fusion module, the VSS block and CPB block are used to extract features and restore low-frequency information on the underwater image, which solves the problem of insufficient adaptability of underwater fish target detection and realizes high-precision underwater fish target detection.
Patent Information
- Application Number
- CN202510837682.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-23
AI Technical Summary
The existing underwater fish target detection technology is insufficiently adaptable in complex and changing underwater environments and cannot effectively deal with unknown degradation of underwater background source information interference, resulting in loss of target feature information and degradation of detection performance.
The target recognition model of the backbone network and fusion module is adopted to extract and restore the underwater image through feature extraction and recovery network, and feature map fusion is combined with VSS blocks and CPB blocks to dynamically restore feature information layer by layer to enhance detection capabilities.
It significantly reduces feature loss, improves the reliability and accuracy of underwater fish target detection, and can achieve high-precision detection under multiple degradation conditions.
Smart Images

Figure CN120356083A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and in particular, to an underwater fish target detection method and device. Background Art
[0002] Underwater fish target detection is an important and challenging problem in marine ecological monitoring and underwater robot research. However, the complex and changeable underwater environment (such as the complexity of the seabed background, the diversity of fish postures, the large variation in target size, and the limited image resolution) and the interference of background source domain information will lead to the loss of details of target low-frequency features. Moreover, as the network depth increases, the target feature information becomes chaotic, thus affecting the detection performance.
[0003] Although the existing underwater fish target detection has improved the ability of the model to extract target feature information, the underwater fish target detection image is essentially an unknown degraded image, showing insufficient adaptability when facing unknown multiple degraded underwater images and unable to meet the requirements of the real-world environment.
[0004] On the other hand, the unknown degraded underwater background source domain information will seriously interfere with the learning of target feature information, and the interference becomes more serious as the depth of the model network increases, and further shows more chaos in the visualization of high-dimensional space features, resulting in serious loss of details of target feature information. Summary of the Invention
[0005] The present invention provides an underwater fish target detection method and device to solve the defect in the prior art that only focuses on the extraction of fish target feature information, resulting in insufficient adaptability of underwater fish target detection, and to achieve reliable and accurate detection of underwater fish targets.
[0006] In a first aspect, the present invention provides an underwater fish target detection method, including: Inputting an underwater image into a target recognition model to obtain the fish species recognition result and the position information of the fish in the underwater image output by the target recognition model; wherein, the target recognition model includes a backbone network and a fusion module; The backbone network includes a plurality of serially connected feature extraction and restoration networks, and each feature extraction and restoration network is used for: extracting features of the underwater image to obtain a fish feature map, and restoring low-frequency information of the underwater image to obtain a low-frequency texture feature map; fusing the fish feature map and the low-frequency texture feature map to obtain a restored feature map; The fusion module is used for fusing the restored feature maps of each feature extraction and restoration network to obtain the fish species recognition result and the position information of the fish in the underwater image.
[0007] In one embodiment, in the multiple cascaded feature extraction and restoration networks: The input of the current feature extraction and restoration network is the output of the previous feature extraction and restoration network, and the input of the first feature extraction and restoration network is the underwater image; The sizes of the restored feature maps of each feature extraction and restoration network are different.
[0008] In one embodiment, the feature extraction and restoration network includes a Visual Spatial State (VSS) block; The VSS block is used for: Decomposing the input image into multiple image blocks and expanding the multiple image blocks into sequences along different traversal paths; Parallelly processing each sequence based on the S6 block, and reshaping and merging the processing results of each sequence to obtain the fish feature map.
[0009] In one embodiment, the feature extraction and restoration network includes a Content-driven Prompting Block (CPB); The CPB is used for: Enhancing the spatial feature information of the input image to obtain a first image; Fusing the input image, the first image, and the Chain of Thought (COT) prompting feature to obtain a second image; the COT prompting feature is obtained by the COT module learning the background source domain information of the input image; Performing spatial channel enhancement on the second image to obtain the low-frequency texture feature map.
[0010] In one embodiment, the backbone network further includes an enhancement block, and the enhancement block is the VSS block or the CPB; The enhancement block is located between the last feature extraction and restoration network and the fusion module; The fusion module is further configured to fuse the restored feature maps of each feature extraction and restoration network and the output image of the enhancement block to obtain the fish species recognition result and the position information of the fish in the underwater image.
[0011] In one embodiment, the method further includes: Using underwater images in different scenarios as samples, and the corresponding fish species and the position of the fish in the underwater image as labels to train the target recognition model.
[0012] In a second aspect, the present invention provides an underwater fish target detection device, including: A detection module, configured to input an underwater image into the target recognition model to obtain the fish species recognition result output by the target recognition model and the position information of the fish in the underwater image; Among them, the target recognition model includes a backbone network and a fusion module; The backbone network includes a plurality of serially connected feature extraction and restoration networks, and each feature extraction and restoration network is used for: extracting features of the underwater image to obtain a fish feature map, and restoring low-frequency information of the underwater image to obtain a low-frequency texture feature map; fusing the fish feature map and the low-frequency texture feature map to obtain a restored feature map; The fusion module is used for fusing the restored feature maps of each feature extraction and restoration network to obtain the fish species recognition result and the position information of the fish in the underwater image.
[0013] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the underwater fish target detection method described in the first aspect above is implemented.
[0014] In a fourth aspect, the present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the underwater fish target detection method described in the first aspect above is implemented.
[0015] In a fifth aspect, the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the underwater fish target detection method described in the first aspect above is implemented.
[0016] The underwater fish target detection method and device provided by the present invention can process the loss of feature information caused by degradation at the feature map level by restoring low-frequency information of the underwater image to obtain a low-frequency texture feature map and fusing the low-frequency texture feature map with the fish feature map, thereby enhancing the details of high-dimensional features. In addition, by fusing the restored feature maps of each feature extraction and restoration network, the feature information can be dynamically restored layer by layer, significantly reducing feature loss and avoiding further degradation of features at deeper levels. Therefore, compared with the prior art that only focuses on the extraction of fish target feature information, the underwater fish target detection method provided by the present invention can achieve reliable and accurate detection of underwater fish targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 It is a schematic flow chart of the underwater fish target detection method provided by the present invention.
[0019] Figure 2 It is one of the schematic structural diagrams of the target recognition model provided by the present invention.
[0020] Figure 3 It is a schematic structural diagram of the feature extraction and restoration network provided by the present invention.
[0021] Figure 4 It is a schematic structural diagram of the COT module provided by the present invention.
[0022] Figure 5 It is the second schematic structural diagram of the target recognition model provided by the present invention.
[0023] Figure 6 It is a schematic structural diagram of the underwater fish target detection system provided by the present invention.
[0024] Figure 7 It is a schematic structural diagram of the illuminance transmitter provided by the present invention.
[0025] Figure 8 It is a schematic structural diagram of the underwater fish target detection device provided by the present invention.
[0026] Figure 9 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0027] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0028] Figure 1 It is a schematic flow chart of the underwater fish target detection method provided by the present invention. As Figure 1 shown, the method may include the following steps: Step 110: Input the underwater image into the target recognition model to obtain the fish species recognition result output by the target recognition model and the position information of the fish in the underwater image; Among them, the target recognition model includes a backbone network and a fusion module; The backbone network includes a plurality of serially connected feature extraction and restoration networks, and each feature extraction and restoration network is used for: Extract features from the underwater image to obtain a fish feature map, and recover the low-frequency information of the underwater image to obtain a low-frequency texture feature map; Fuse the fish feature map and the low-frequency texture feature map to obtain a restored feature map; A fusion module, which is used to fuse the restored feature maps of each feature extraction and restoration network to obtain the fish species recognition result and the position information of the fish in the underwater image.
[0029] It should be noted that the execution subject of the above underwater fish target detection method can be a computer device, such as a mobile phone, a tablet computer, a notebook computer, a palm computer, etc.
[0030] In step 110, the underwater image can be an image of any underwater scene containing fish. Specifically, the underwater image can be an underwater image of the same location or different locations with different time periods, angles, and light attenuation, or an underwater image of different underwater environments such as the deep sea, the shallow sea, and the aquaculture farm. The present invention does not limit the specific type of the underwater image.
[0031] After obtaining the underwater image, certain preprocessing operations can be performed on the underwater image first. For example, the Stem layer can be used to perform preprocessing operations such as bottom layer feature extraction, spatial dimensionality reduction, multi-scale fusion, and modular adaptation on the underwater image. According to the type of the target recognition model, the structure and connection method of the Stem layer can be adjusted to balance the calculation efficiency and the feature expression ability, so as to provide optimized input features for the subsequent deep network.
[0032] The target recognition model can be various models with image recognition capabilities, such as Transformer, Mamba, RNN, CNN, etc.
[0033] Taking the target recognition model as the Mamba network as an example, the technical solution of the present invention will be described below. In the following text, the target recognition model can be called "Fish Mamba".
[0034] After preprocessing, the underwater image can be input into Fish Mamba, so as to obtain the fish species recognition result output by Fish Mamba and the position information of the fish in the underwater image.
[0035] Specifically, as Figure 2 shown, Fish Mamba can include a backbone network 210 and a fusion module 220.
[0036] The backbone network 210 can include a number of serially connected feature extraction and restoration networks 211 ( Figure 2The three feature extraction and restoration networks 211 shown are only for illustrative purposes and do not limit the specific number of the feature extraction and restoration networks 211. Each feature extraction and restoration network 211 is used for: Extract features from the underwater image to obtain a fish feature map, and restore the low-frequency information of the underwater image to obtain a low-frequency texture feature map; Specifically, the feature extraction and restoration network 211 can use various networks (functional modules) such as VSS (Visual Spatial State), CNN, Transformer, ResNet, YOLO, etc. and their combinations to extract features from the underwater image to obtain a fish feature map.
[0037] The feature extraction and restoration network 211 can use various networks (functional modules) such as CPB (Content-driven Prompt Block), WDMamba, FocalNet, HLNet, MSB, SwinIR, etc. and their combinations to restore the low-frequency information of the underwater image to obtain a low-frequency texture feature map.
[0038] After obtaining the fish feature map and the low-frequency texture feature map, the feature extraction and restoration network 211 can fuse the two to obtain a restored feature map.
[0039] The fusion module 220 can use various networks (functional modules) such as DINO network, TSJNet, EfficientViT, CoCoNet, etc. and their combinations to fuse the restored feature maps of each feature extraction and restoration network 211 to obtain the fish species recognition result and the position information of the fish in the underwater image.
[0040] The underwater fish target detection method provided by the present invention can process the loss of feature information caused by degradation at the feature map level by restoring the low-frequency information of the underwater image to obtain a low-frequency texture feature map and fusing the low-frequency texture feature map with the fish feature map, thereby enhancing the details of the high-dimensional features. In addition, by fusing the restored feature maps of each feature extraction and restoration network, the feature information can be dynamically restored layer by layer, significantly reducing feature loss and avoiding further degradation of features at deeper levels. Therefore, compared with the prior art that only focuses on the extraction of fish target feature information, the underwater fish target detection method provided by the present invention can achieve reliable and accurate detection of underwater fish targets.
[0041] In one embodiment, among multiple serially connected feature extraction and restoration networks 211: The input of the current feature extraction and restoration network 211 is the output of the previous feature extraction and restoration network 211, and the input of the first feature extraction and restoration network 211 is the underwater image; The sizes of the restored feature maps of each feature extraction and restoration network 211 are different.
[0042] As Figure 2 shown, for the first feature extraction and restoration network 211, its input is the underwater image, for the second feature extraction and restoration network 211, its input is the output of the first feature extraction and restoration network 211, and for the third feature extraction and restoration network 211, its input is the output of the second feature extraction and restoration network 211.
[0043] Assume that the size of the underwater image is W H C, where W is the image width, H is the image height, and C is the number of image channels (for example, a visible light image is usually composed of three colors: blue, red, and yellow, then the value of C is 3). Then, the size of the restored feature map output by the first feature extraction and restoration network 211 can be W / 2 H / 2 2C, the size of the restored feature map output by the second feature extraction and restoration network 211 can be W / 4 H / 4 4C, and the size of the restored feature map output by the third feature extraction and restoration network 211 can be W / 8 H / 8 8C.
[0044] It can be understood that by connecting the inputs and outputs of each feature extraction and restoration network 211 in series and making the sizes of the restored feature maps of each feature extraction and restoration network different, the orderly progress of feature extraction and restoration can be achieved, which helps to improve the learning ability of Fish Mamba.
[0045] In one embodiment, the feature extraction and restoration network 211 may include a VSS block; The VSS block is used for: decomposing the input image into multiple image blocks and expanding the multiple image blocks into sequences along different traversal paths; processing each sequence in parallel based on the S6 block and reshaping and merging the processing results of each sequence to obtain the fish feature map.
[0046] Figure 3 is a schematic structural diagram of the feature extraction and restoration network provided by the present invention. Among them, Figure 3 can correspond to the first feature extraction and restoration network 211.
[0047] AsFigure 3 As shown, the VSS block first decomposes the input image into multiple image patches (patches) ( Figure 2 shown as 9 for example only), and unfolds them into sequences along 4 different traversal paths (selective scanning, i.e., Cross-Scan).
[0048] Next, the VSS block uses a separate S6 block to process each image patch sequence in parallel, and reshapes and merges the processing results of each sequence to obtain a fish feature map (i.e., Cross-Merge).
[0049] Furthermore, by using complementary one-dimensional traversal paths in the S6 block, the VSS block can use SS2D to enable each pixel in the image to effectively integrate the information of all other pixels from different directions, thereby promoting the fish feature information for establishing a global receptive field in the two-dimensional space (i.e., the fish feature map ).
[0050] Among them, the S6 block is a spatial state model for describing a linear time-invariant (LTI) system. Specifically in the present invention: (1) Among them, is an image patch with a temporal order, is the feature information after state update, represents the state vector or hidden state of the system, y represents the output variable, L represents the number of output variables, and the matrix A represents how the current state evolves over time, B , C and D matrices respectively record each batch and sequence position in different cases, represents the step size.
[0051] The underwater fish target detection method provided by the present invention can effectively integrate the fish feature information from different directions by using the VSS block to extract features from the underwater image to obtain a fish feature map, making the acquisition of fish features accurate and comprehensive, and thus improving the performance of Fish Mamba.
[0052] In one embodiment, the feature extraction and recovery network 211 includes a CPB; The CPB is used for: Strengthening the spatial feature information of the input image to obtain a first image; Fuse the input image, the first image, and the COT (Chain of Thought) prompt features to obtain the second image; the COT prompt features are obtained by the COT module learning the background source domain information of the input image. Perform spatial-channel enhancement on the second image to obtain a low-frequency texture feature map.
[0053] Specifically, as Figure 3 shown, CPB first enhances the input image at the spatial feature level through the Enhancement of spatial feature information module to obtain the first image . At the same time, CPB fuses the first image , the COT prompt features , and the input image to obtain the second image .
[0054] Next, CPB inputs the second image into the spatial-channel enhancement module for spatial-channel enhancement. Specifically, the spatial-channel enhancement module divides the second image into n groups, and each group ( j = 1, 2, ..., n) respectively uses an independent Transformer block for spatial-channel enhancement, and finally fuses the outputs ( ) of each Transformer block to obtain the low-frequency texture feature map . Among them, the fish feature map and the low-frequency texture feature map can be fused to obtain the restored feature map .
[0055] It should be noted that different from conventional underwater image enhancement that requires preprocessing in advance or information restoration at the image visual level, the above image low-frequency information restoration is only at the feature level, thereby enhancing the target feature information and is end-to-end.
[0056] The process of the above image low-frequency information restoration can be expressed as follows: (2) Among them, represents the feature information after attention learning, (the first image) represents the feature information after adding prompt learning, represents the low-frequency information containing the image texture features, A function of the attention mechanism module representing the feature level (i.e., the image width W and height H). A function of the spatial (i.e., channel C) attention mechanism module.
[0057] For the CPB provided by the present invention, it is designed to facilitate the input image and the COT prompt features to interact, enabling Fish Mamba to adjust its enhancement strategy according to the degradation type.
[0058] To better utilize the input content, CPB calculates the spatial importance map of each channel and fully mixes the attention weights in the channel dimension ( Figure 3 in ) and the attention weights in the spatial dimension ( Figure 3 in ) to comprehensively capture features.
[0059] Next, denotes being split into n equal parts along the channel dimension direction, and each part is fed into an independent Transformer block, and each Transformer block uses the degradation information encoded in the COT prompt for learning.
[0060] Finally, all the output results are concatenated along the channel dimension.
[0061] This design of the CPB provided by the present invention brings at least the following benefits: 1), Each part can focus on the correlations of different channels and features, thus increasing the expressive power of the model; 2), Reducing the number of parameters and computational complexity; 3), Each part is calculated independently, significantly reducing the training time.
[0062] Furthermore, in the CPB of the present invention, the COT module is used to encode the context information of specific degradation.
[0063] Specifically, COT can sample and learn the background source domain information of the input image by itself. Since it is unknown, it is initially a random image feature vector.
[0064] As Figure 4 shown, the COT module generates multi-scale prompts through a series of transposed convolutional layers and applies the Hardswish activation function after all transposed convolutions (TConv) to control the COT prompt features input to the next stage.
[0065] The size of the COT prompt features generated in each stage is the same as that of the corresponding first image They have the same size, and the sizes of the COT prompt features at each stage are different. This not only helps Fish Mamba better understand the degradation types in a coherent and step-by-step manner but also assists Fish Mamba in learning hierarchical representations.
[0066] The CPB provided by the present invention can adaptively learn according to different underwater background source domain information by integrating the COT module and generate corresponding prompts, which can drive the CPB to recover more abundant low-frequency feature information of fish targets, especially under the degradation phenomena such as insufficient light, color attenuation, and blurring commonly found in underwater environments. Combining the prompt learning of different background source domain information, the learning and recovery of low-frequency feature information make up for the loss of low-frequency features caused by light and environmental degradation, achieving high-precision underwater fish target detection by Fish Mamba under multiple degradation conditions.
[0067] In summary, the feature extraction and recovery network 211 provided by the present invention, combined with the CPB-VSS architecture proposed by the VSS module and the COT module, can prompt and drive the CPB to generate lost low-frequency image feature information according to underwater source domain information. At the same time, it also uses the correlation of elements in the sequence to selectively retain or filter information, strengthening the capture and recovery of target features.
[0068] In one embodiment, the backbone network 210 further includes an enhancement block 212, and the enhancement block 212 can be a VSS block or a CPB; The enhancement block 212 is located between the last feature extraction and recovery network 211 and the fusion module 220; The fusion module 220 is further configured to fuse the recovered feature maps of each feature extraction and recovery network 211 and the output image of the enhancement block 212 to obtain the fish species recognition result and the position information of the fish in the underwater image.
[0069] Specifically, as Figure 5 shown, after connecting multiple feature extraction and recovery networks 211 in series, an independent enhancement block 212 can also be set between the last feature extraction and recovery network 211 and the fusion module 220.
[0070] When the enhancement block 212 is a VSS block, the enhancement block 212 can perform further fish feature extraction on the recovered feature map of the last feature extraction and recovery network 211, thereby strengthening Fish Mamba's recognition of fish features.
[0071] When the enhancement block 212 is a CPB block, the enhancement block 212 can perform further recovery of low-frequency texture feature maps on the recovered feature map of the last feature extraction and recovery network 211, thereby strengthening Fish Mamba's recovery of low-frequency texture features.
[0072] It is understandable that assuming the size of the underwater image is W H C, then in Figure 5 the size of the restored feature map output by the first feature extraction and restoration network 211 can be W / 2 H / 2 2C, the size of the restored feature map output by the second feature extraction and restoration network 211 can be W / 4 H / 4 4C, the size of the restored feature map output by the third feature extraction and restoration network 211 can be W / 8 H / 8 8C, the size of the output image of the enhancement block 212 can be W / 16 H / 16 16C.
[0073] For the underwater fish target detection method provided by the present invention, by setting the enhancement block 212 between the last feature extraction and restoration network 211 and the fusion module 220 to enhance the recognition ability of Fish Mamba for fish features or improve the restoration ability of low-frequency texture features, Fish Mamba can meet various requirements, thereby improving the practicability.
[0074] In one embodiment, the underwater fish target detection method provided by the present invention may further include: Using underwater images in different scenarios as samples, and the corresponding fish species and the positions of the fish in the underwater images as labels to train the target recognition model.
[0075] For example, a camera device can be used to capture underwater images containing fish targets in various underwater scenarios, and the corresponding fish species and the positions of the fish in the underwater images can be used as labels to construct an underwater image training sample set.
[0076] Then, the underwater image training sample set can be used as the input of Fish Mamba, set the initial network parameters and use the SGD optimizer to train Fish Mamba.
[0077] Specifically, the captured underwater images can be subjected to picture conversion, and the converted pictures can be labeled with fish species and the position information of the fish in the underwater images.
[0078] Subsequently, according to a certain ratio, such as 0.85:0.15, etc., the underwater image training sample set is redistributed into a training set and a test set.
[0079] Finally, adjust the short side of the images in the training set to W, limit the long side to H, and perform image flipping augmentation. After the training of Fish Mamba is completed, use the test set to verify the recognition effect of Fish Mamba.
[0080] The underwater fish target detection method provided by the present invention can effectively improve the generalization ability of the target recognition model and the recognition accuracy of the target recognition model by training the target recognition model with underwater images in various scenarios.
[0081] Figure 6 It is a schematic structural diagram of the underwater fish target detection system provided by the present invention. As Figure 6 shown, to implement the above underwater fish target detection method, the underwater fish target detection system provided by the present invention may include: a light source 1, an underwater shooting device 2, a light intensity transmitter 3, a control processor 4, and an oxygen generation device 5.
[0082] The control processor 4 is respectively communicatively connected to the underwater shooting device 2, the light source 1, and the light intensity transmitter 3.
[0083] The underwater shooting device 2 can collect the motion video of the fish target in real time under the control of the control processor 4 and extract the fish target video into a video stream.
[0084] The light source 1 is used to supplement light for the underwater shooting device 2. The light intensity transmitter 3 can sense the light intensity of the environment and transmit the light intensity information to the control processor 4. The control processor 4 controls the switch of the light source 1 and the light intensity according to the light intensity information. The control processor 4 can perform fish target detection and recognition according to the trained Fish Mmaba.
[0085] Further, as Figure 7 shown, the light intensity transmitter 3 includes a light intensity sensor 31, a microcontroller 32, and a communication interface 33 that are communicatively connected in sequence. The microcontroller 32 is respectively connected to the light intensity sensor 31 and the communication interface 33. The microcontroller 32 can control the light intensity sensor 31 to collect data and transmit the data collected by the light intensity sensor to the control processor 4 through the communication interface 33.
[0086] The underwater fish target detection system provided by the present invention has a refined structure and strong practicability, and can be effectively applied to the aquaculture environment, providing a reliable and accurate technical means for fish target detection in an underwater unknown degradation domain.
[0087] Next, the underwater fish target detection device provided by the present invention will be described. The underwater fish target detection device described below can be mutually corresponding and referenced with the underwater fish target detection method described above, and can achieve the same technical effects, which will not be elaborated here.
[0088] Figure 8 This is a schematic structural diagram of the underwater fish target detection device provided by the present invention. As Figure 8 shown, the device may include: A detection module 810, configured to input an underwater image into a target recognition model, and obtain a fish species recognition result and position information of the fish in the underwater image output by the target recognition model; Wherein, the target recognition model includes a backbone network and a fusion module; The backbone network includes a plurality of cascaded feature extraction and restoration networks, and each feature extraction and restoration network is used for: Performing feature extraction on the underwater image to obtain a fish feature map, and performing low-frequency information restoration on the underwater image to obtain a low-frequency texture feature map; Fusing the fish feature map and the low-frequency texture feature map to obtain a restored feature map; The fusion module is configured to fuse the restored feature maps of each feature extraction and restoration network to obtain the fish species recognition result and the position information of the fish in the underwater image.
[0089] In one embodiment, in the plurality of cascaded feature extraction and restoration networks: The input of the current feature extraction and restoration network is the output of the previous feature extraction and restoration network, and the input of the first feature extraction and restoration network is the underwater image; The sizes of the restored feature maps of each feature extraction and restoration network are different.
[0090] In one embodiment, the feature extraction and restoration network includes a Visual Spatial State (VSS) block; The VSS block is used for: Decomposing the input image into a plurality of image blocks, and expanding the plurality of image blocks into sequences along different traversal paths; Parallelly processing each sequence based on an S6 block, and reshaping and merging the processing results of each sequence to obtain the fish feature map.
[0091] In one embodiment, the feature extraction and restoration network includes a Content-driven Prompting Block (CPB); The CPB is used for: Strengthening the spatial feature information of the input image to obtain a first image; Fusing the input image, the first image, and a Chain of Thought (COT) prompting feature to obtain a second image; the COT prompting feature is obtained by the COT module learning background source domain information of the input image; Performing spatial channel strengthening on the second image to obtain the low-frequency texture feature map.
[0092] In one embodiment, the backbone network further includes an enhancement block, and the enhancement block is the VSS block or the CPB; The enhancement block is located between the last feature extraction and restoration network and the fusion module; The fusion module is further configured to fuse the restored feature maps of each feature extraction and restoration network and the output image of the enhancement block to obtain the fish species recognition result and the position information of the fish in the underwater image.
[0093] In one embodiment, the method further includes a training module (not shown in the figure) for: Using underwater images in different scenarios as samples, and the corresponding fish species and the position of the fish in the underwater image as labels, to train the target recognition model.
[0094] Figure 9 An example of the physical structure diagram of an electronic device is shown as Figure 9 shown. The electronic device may include: a processor 910, a communication interface 920, a memory 930, and a communication bus 940. Among them, the processor 910, the communication interface 920, and the memory 930 communicate with each other through the communication bus 940. The processor 910 can call the logical instructions in the memory 930 to execute the underwater fish target detection method described in any of the above embodiments, for example, including: Inputting the underwater image into the target recognition model to obtain the fish species recognition result and the position information of the fish in the underwater image output by the target recognition model; Among them, the target recognition model includes a backbone network and a fusion module; The backbone network includes a plurality of cascaded feature extraction and restoration networks, and each feature extraction and restoration network is used for: Performing feature extraction on the underwater image to obtain a fish feature map, and performing low-frequency information restoration on the underwater image to obtain a low-frequency texture feature map; Fusing the fish feature map and the low-frequency texture feature map to obtain a restored feature map; The fusion module is used to fuse the restored feature maps of each feature extraction and restoration network to obtain the fish species recognition result and the position information of the fish in the underwater image.
[0095] In addition, when the logical instructions in the above-mentioned memory 930 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0096] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the underwater fish target detection method described in any of the above embodiments, for example, including: Input the underwater image into the target recognition model to obtain the fish species recognition result output by the target recognition model and the position information of the fish in the underwater image; Among them, the target recognition model includes a backbone network and a fusion module; The backbone network includes multiple cascaded feature extraction and restoration networks, and each feature extraction and restoration network is used for: Extract features from the underwater image to obtain a fish feature map, and restore the low-frequency information of the underwater image to obtain a low-frequency texture feature map; Fuse the fish feature map and the low-frequency texture feature map to obtain a restored feature map; The fusion module is used to fuse the restored feature maps of each feature extraction and restoration network to obtain the fish species recognition result and the position information of the fish in the underwater image.
[0097] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the underwater fish target detection method described in any of the above embodiments, for example, including: Input the underwater image into the target recognition model to obtain the fish species recognition result output by the target recognition model and the position information of the fish in the underwater image; Among them, the target recognition model includes a backbone network and a fusion module; The backbone network includes a plurality of serially connected feature extraction and restoration networks, and each feature extraction and restoration network is used for: extracting features from the underwater image to obtain a fish feature map, and restoring low-frequency information of the underwater image to obtain a low-frequency texture feature map; fusing the fish feature map and the low-frequency texture feature map to obtain a restored feature map; The fusion module is used for fusing the restored feature maps of each feature extraction and restoration network to obtain the fish species recognition result and the position information of the fish in the underwater image.
[0098] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.
[0099] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An underwater fish target detection method, characterized in that, Including: Input the underwater image into the target recognition model to obtain the fish species recognition result output by the target recognition model and the position information of the fish in the underwater image; Wherein, the target recognition model includes a backbone network and a fusion module; The backbone network includes a plurality of cascaded feature extraction and restoration networks, and each feature extraction and restoration network is used for: Extract features from the underwater image to obtain a fish feature map, and restore low-frequency information of the underwater image to obtain a low-frequency texture feature map; Fuse the fish feature map and the low-frequency texture feature map to obtain a restored feature map; The fusion module is used to fuse the restored feature maps of each feature extraction and restoration network to obtain the fish species recognition result and the position information of the fish in the underwater image.
2. The underwater fish target detection method according to claim 1, characterized in that, Among the plurality of cascaded feature extraction and restoration networks: The input of the current feature extraction and restoration network is the output of the previous feature extraction and restoration network, and the input of the first feature extraction and restoration network is the underwater image; The sizes of the restored feature maps of each feature extraction and restoration network are different.
3. The underwater fish target detection method according to claim 2, characterized in that The feature extraction and restoration network includes a Visual Spatial State (VSS) block; The VSS block is used for: Decompose the input image into multiple image blocks, and expand the multiple image blocks into sequences along different traversal paths; Parallel process each sequence based on the S6 block, and reshape and merge the processing results of each sequence to obtain the fish feature map.
4. The underwater fish target detection method according to claim 3, wherein The feature extraction and restoration network includes a Content-driven Prompting Block (CPB); The CPB is used for: Enhance the spatial feature information of the input image to obtain a first image; Fuse the input image, the first image, and the Chain of Thought (COT) prompting feature to obtain a second image; the COT prompting feature is obtained by the COT module learning the background source domain information of the input image; Perform spatial channel enhancement on the second image to obtain the low-frequency texture feature map.
5. The underwater fish target detection method according to claim 4, wherein The backbone network further includes an enhancement block, and the enhancement block is the VSS block or the CPB; The enhancement block is located between the last feature extraction and restoration network and the fusion module; The fusion module is further used to fuse the restored feature maps of each feature extraction and restoration network and the output image of the enhancement block to obtain the fish species recognition result and the position information of the fish in the underwater image.
6. The underwater fish target detection method according to any one of claims 1 to 5, characterized in that The method further includes: Use underwater images in different scenarios as samples, and the corresponding fish species and the position of the fish in the underwater image as labels to train the target recognition model.
7. An underwater fish target detection device, characterized in that, Including: A detection module for inputting the underwater image into the target recognition model to obtain the fish species recognition result output by the target recognition model and the position information of the fish in the underwater image; Wherein, the target recognition model includes a backbone network and a fusion module; The backbone network includes a plurality of cascaded feature extraction and restoration networks, and each feature extraction and restoration network is used for: Extract features from the underwater image to obtain a fish feature map, and restore low-frequency information of the underwater image to obtain a low-frequency texture feature map; Fuse the fish feature map and the low-frequency texture feature map to obtain a restored feature map; The fusion module is used to fuse the restored feature maps of each feature extraction and restoration network to obtain the fish species recognition result and the position information of the fish in the underwater image.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the underwater fish target detection method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the underwater fish target detection method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the underwater fish target detection method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Underwater fish target detection method and device and storage medium
CN113869330A
Fish individual identity recognition method based on improved YOLOv4 and FIRN
CN115631404A
Lightweight underwater target detection method and system based on degraded image enhancement
CN116543295A
Night storage robot target detection method and system based on integral network
CN118097089A
Image enhancement processing method for fishery resource statistics
CN119107270A