An underwater dynamic target detection and identification method
By constructing a spatiotemporal optical flow divergence-scattering coupling entropy and attention weight matrix, the coupling interference problem between optical scattering and motion distortion in underwater dynamic target detection is solved, achieving high-precision and high-robust target detection in complex underwater environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU NONGKEN FISHERY TECHNOLOGY CO LTD
- Filing Date
- 2026-04-27
- Publication Date
- 2026-05-29
Smart Images

Figure CN122116108A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image recognition technology, and in particular to a method for detecting and recognizing underwater dynamic targets. Background Technology
[0002] In underwater inspection scenarios, remotely operated vehicles (ROVs) need to achieve real-time detection and identification of dynamic targets. Mainstream solutions employ underwater optical imaging combined with deep learning and optical flow methods for target extraction and motion analysis. Underwater imaging is affected by backscattering from suspended particles in the water, resulting in large-area light spot interference. Simultaneously, water turbulence generates non-rigid pseudo-motion distortion, and the combination of these two factors leads to severe degradation of visual features. Existing detection methods mostly rely on pure deep learning or single physical compensation. While these can achieve basic target identification, they struggle to cope with the dual interference of optical scattering and motion distortion in complex underwater environments with high turbidity and strong disturbances. Dynamic target detection requires eliminating the influence of the vehicle's own motion and distinguishing between real targets and water disturbances. However, current technologies lack a coupled correlation model for underwater interference, becoming a core bottleneck restricting the accuracy and robustness of underwater dynamic target detection.
[0003] Existing underwater dynamic target detection technologies suffer from four core flaws: First, they only address backscattered optical interference or optical flow distortion in isolation, failing to recognize the strong coupling between the two and thus failing to uniformly suppress pseudo-dynamic features resulting from this coupling, leading to incomplete interference removal. Second, vehicle motion compensation only considers translational components, ignoring global pixel shifts caused by rotational motion, leaving significant rigid motion errors in the optical flow field and distorting apparent current divergence calculations. Third, deep learning networks directly input original image features, causing scattered light spots and turbulent pseudo-motion features to be continuously amplified during multi-scale fusion, resulting in numerous false target detections and a significant decrease in localization and classification accuracy. Fourth, dynamic target matching relies solely on the detection box position without physical prior constraints, leading to high mismatch rates between adjacent frames and significant errors in motion vector calculation. Furthermore, existing methods lack quantitative characterization of interference coupling, hindering adaptive feature suppression and exhibiting poor versatility and stability in complex underwater environments, making it difficult to meet the practical needs of high-confidence underwater dynamic target detection. Summary of the Invention
[0004] The main objective of this invention is to provide an underwater dynamic target detection and recognition method. It innovatively constructs a spatiotemporal optical flow divergence-scattering coupling entropy to quantitatively characterize the coupling degree between underwater scattering and turbulent pseudo-motion, accurately locating high-interference regions and fundamentally distinguishing real targets from water disturbances. It fully compensates for the vehicle's translational and rotational global motion, eliminating self-motion interference pixel by pixel to obtain a pure relative optical flow field, thus improving the accuracy of divergence calculation. A coupling entropy-driven physical attention weight matrix directly suppresses the activation of high-entropy interference regions in the shallow feature stage, avoiding the amplification of pseudo-features in the network and reducing false detection and false negative rates. Based on accurate IOU matching and pure feature output, it ensures one-to-one matching of dynamic targets in adjacent frames, making motion vector calculation accurate and reliable. Furthermore, this solution integrates an underwater physical imaging model and a deep learning architecture, retaining the rigor of physical priors while possessing the feature extraction capabilities of neural networks. It maintains high detection accuracy and robustness even in high-turbidity and highly disturbed underwater environments, adapting to various underwater dynamic target detection scenarios, and its practicality and versatility surpass existing technologies.
[0005] The technical solution of the present invention is as follows:
[0006] A method for underwater dynamic target detection and recognition is proposed, which includes the following steps:
[0007] S1. An underwater vehicle equipped with an imaging acquisition unit, an active illumination unit, a depth sensing unit, and a motion sensing unit acquires underwater images and obtains underwater multi-source physical signals. Based on the underwater multi-source physical signals, the backscattered light intensity distribution field and the surface divergence field are extracted.
[0008] S2. Based on the backscattered light intensity distribution field and the apparent light divergence field, the spatiotemporal light divergence-scattering coupling entropy is calculated.
[0009] S3. Generate an attention weight matrix based on the spatiotemporal optical flow divergence-scattering coupling entropy, and modulate the attention weight matrix element-wise with the shallow feature map extracted by the target detection network to obtain a pure target feature map.
[0010] S4. Based on the clean target feature map, perform multi-scale feature extraction and classification regression processing, and output the dynamic target category probability and the corresponding target bounding box coordinates and motion vector.
[0011] A further improvement of the present invention is that step S1 includes the following specific steps:
[0012] S11. An underwater vehicle equipped with an imaging acquisition unit, an active illumination unit, a depth sensing unit, and a motion sensing unit acquires underwater images and obtains underwater multi-source physical signals, wherein the underwater multi-source physical signals include the original observed light intensity sequence. Active light source radiation power Target space depth The vehicle's velocity vector and the vehicle's rotational angular velocity vector ;in, This represents the horizontal coordinate index within the pixel plane. This represents the vertical coordinate index within the pixel plane. This refers to the time corresponding to the acquisition time sequence;
[0013] S12. Based on the Jaffe-McGlamery underwater imaging model, select the dark background area in the underwater image whose spatial distance is greater than 10 times the imaging sharpness distance, and calculate the water attenuation coefficient corresponding to the emission wavelength of the active light source center. Bilateral filtering was used to process the original observed light intensity sequence. Spatial domain filtering and separation were performed to extract the backscattered light intensity distribution field. The formula for calculating the water body attenuation coefficient is as follows:
[0014] ;
[0015] in, The center emission wavelength of the active lighting unit The corresponding water body attenuation coefficient, The average illuminance of the dark background area. The average spatial distance of the dark background area. For the optical system efficiency of the imaging acquisition unit, The angle between the optical axis of the active illumination unit and the optical axis of the imaging acquisition unit. The solid angle corresponding to a single pixel in the imaging acquisition unit;
[0016] S13. Extract the original observed light intensity sequence corresponding to two adjacent underwater images. , The initial surface flow field was calculated using the Lucas-Kanade algorithm. The motion parameters in the camera coordinate system are obtained by transforming a pre-calibrated extrinsic parameter matrix. Then, the three-dimensional coordinates of the corresponding spatial points in the camera coordinate system are calculated pixel by pixel using the pinhole camera back projection formula. Based on the motion parameters and three-dimensional coordinates in the camera coordinate system, the pixel-by-pixel global motion offset vector is calculated. For the initial surface sightseeing flow field Compensation is performed to obtain the relative optical flow field The apparent optical divergence field is obtained by calculating the divergence of the relative optical flow field. ,in, This represents the time interval between the acquisition of two adjacent underwater images.
[0017] A further improvement of the present invention is that step S2 includes the following specific steps:
[0018] S21, Set in pixels Local sliding window in the spatiotemporal domain centered on The spatiotemporal local sliding window The time dimension includes the current time t and the previous time. For local sliding windows in the spatiotemporal domain Backscattered light intensity distribution field within The normalization process is performed using the following formula: ;in, This is the normalized value for backscattering. , Let [0,1] be the minimum and maximum values of backscattered light intensity within a local sliding window in the spatiotemporal domain, respectively. Divide [0,1] into N equal intervals and count the proportion of pixels in each interval to the total number of pixels within the local sliding window in the spatiotemporal domain to obtain the interval probability distribution of the backscattered normalized value. Where i is the interval index, and its value ranges from 1 to N;
[0019] S22. Local sliding window in the spatiotemporal domain The surface of the tourist flow field The normalization process is performed using the following formula: ; , Local sliding windows in the spatiotemporal domain Minimum and maximum absolute values of internal optical divergence. The normalized absolute value of optical flow divergence is obtained by statistically analyzing the proportion of pixels within N intervals to the total number of pixels within a local sliding window in the spatiotemporal domain. Where j is the interval index, with a value range of 1-N;
[0020] S23. Interval probability distribution of normalized backscatter values Interval probability distribution of the normalized absolute value of optical divergence Substituting into the joint information entropy formula, the spatiotemporal optical flow divergence-scattering coupling entropy is calculated. The formula is: .
[0021] A further improvement of the present invention is that step S3 includes the following specific steps:
[0022] S31, the original observed light intensity sequence The shallow feature map is extracted after inputting the backbone layer of the YOLOv8 object detection network and performing three downsampling operations. ;
[0023] S32, Spatiotemporal optical divergence-scattering coupling entropy Perform inverse exponential mapping to generate the attention weight matrix. The formula for calculating the attention weight matrix is as follows: ;in, The attention weight matrix is adjusted using bilinear interpolation to match the shallow features. Figure One The spatial dimensions are adjusted, and the attention weight matrix is output after resizing. ;
[0024] S33. Adjust the size of the attention weight matrix. With shallow feature map Element-wise multiplication along the spatial dimension yields a clean target feature map. , ;in, This is the element-wise multiplication operator.
[0025] A further improvement of the present invention is that the specific content of S4 is as follows:
[0026] S41. Extract the clean target feature map. The YOLOv8-based PAFPN multi-scale feature pyramid network is input, and bidirectional feature fusion is performed through a top-down upsampling path and a bottom-up downsampling path, outputting three sets of multi-scale feature maps at different scales. These multi-scale feature maps are then simultaneously input into a decoupled classification head and regression head. The classification head, after passing through three consecutive convolutional layers and a sigmoid activation layer, outputs the dynamic target class probability corresponding to each detection box. The regression head outputs the target bounding box coordinates corresponding to each detection box through three consecutive convolutional layers and linear activation layers;
[0027] S42. Match the bounding boxes of targets in adjacent frames using a preset matching threshold. Analyze the changes in the center coordinates of the adjacent frames of successfully matched targets against the time interval of the acquisition. Calculate the target motion vector The final output is the dynamic target category probability. The corresponding target bounding box coordinates and motion vectors The formula for calculating the target motion vector is: ;in, Let be the center pixel coordinates of the target bounding box at time t. for The center pixel coordinates of the bounding box of the same target that is successfully matched at any given time.
[0028] The technical effects of this invention are as follows:
[0029] A novel underwater dynamic target detection and recognition method is constructed, innovatively building a spatiotemporal optical flow divergence-scattering coupling entropy to quantitatively characterize the coupling degree between underwater scattering and turbulent pseudo-motion, accurately locating high-interference regions and fundamentally distinguishing real targets from water disturbances. The method fully compensates for the vehicle's translational and rotational global motions, removing self-motion interference pixel by pixel to obtain a pure relative optical flow field, thus improving divergence calculation accuracy. A coupling entropy-driven physical attention weight matrix directly suppresses the activation of high-entropy interference regions in the shallow feature stage, avoiding the amplification of pseudo-features in the network and reducing false detection and false negative rates. Based on accurate IOU matching and pure feature output, it ensures one-to-one matching of dynamic targets in adjacent frames, ensuring reliable motion vector calculation. Furthermore, this scheme integrates an underwater physical imaging model and a deep learning architecture, retaining the rigor of physical priors while possessing the feature extraction capabilities of neural networks. It maintains high detection accuracy and robustness in high-turbidity and highly disturbed underwater environments, adapting to various underwater dynamic target detection scenarios, and its practicality and versatility surpass existing technologies. Attached Figure Description
[0030] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0031] Figure 1 This is a flowchart illustrating an underwater dynamic target detection and recognition method according to Embodiment 1 of the present invention. Detailed Implementation
[0032] Example 1: This example proposes an underwater dynamic target detection and recognition method. It innovatively constructs a spatiotemporal optical flow divergence-scattering coupling entropy to quantitatively characterize the coupling degree between underwater scattering and turbulent pseudo-motion, accurately locating high-interference regions and fundamentally distinguishing real targets from water disturbances. It fully compensates for the vehicle's translational and rotational global motion, removing self-motion interference pixel by pixel to obtain a pure relative optical flow field, improving divergence calculation accuracy. A coupling entropy-driven physical attention weight matrix directly suppresses the activation of high-entropy interference regions in the shallow feature stage, avoiding the amplification of pseudo-features in the network and reducing false detection and false negative rates. Based on accurate IOU matching and pure feature output, it ensures one-to-one matching of dynamic targets in adjacent frames, making motion vector calculation accurate and reliable. Simultaneously, this scheme integrates an underwater physical imaging model and a deep learning architecture, retaining the rigor of physical priors while possessing the feature extraction capabilities of neural networks. It maintains high detection accuracy and robustness in high-turbidity, highly disturbed underwater environments, adapting to various underwater dynamic target detection scenarios, and its practicality and versatility surpass existing technologies. Specifically, such as... Figure 1 As shown, the underwater dynamic target detection and recognition method proposed in this embodiment includes the following specific steps:
[0033] S1. An underwater vehicle equipped with an imaging acquisition unit, an active illumination unit, a depth sensing unit, and a motion sensing unit acquires underwater images and obtains underwater multi-source physical signals. Based on the underwater multi-source physical signals, the backscattered light intensity distribution field and the surface divergence field are extracted.
[0034] S2. Based on the backscattered light intensity distribution field and the apparent light divergence field, the spatiotemporal light divergence-scattering coupling entropy is calculated.
[0035] S3. Generate an attention weight matrix based on the spatiotemporal optical flow divergence-scattering coupling entropy, and modulate the attention weight matrix element-wise with the shallow feature map extracted by the target detection network to obtain a pure target feature map.
[0036] S4. Based on the clean target feature map, perform multi-scale feature extraction and classification regression processing, and output the dynamic target category probability and the corresponding target bounding box coordinates and motion vector.
[0037] In this embodiment, step S1 includes the following specific steps:
[0038] S11. An underwater vehicle equipped with an imaging acquisition unit, an active illumination unit, a depth sensing unit, and a motion sensing unit acquires underwater images and obtains underwater multi-source physical signals, wherein the underwater multi-source physical signals include the original observed light intensity sequence. Active light source radiation power Target space depth The vehicle's velocity vector and the vehicle's rotational angular velocity vector ;in, This represents the horizontal coordinate index within the pixel plane. This represents the vertical coordinate index within the pixel plane. This refers to the time corresponding to the data collection time sequence.
[0039] In this embodiment, the original observed light intensity sequence Let be the total illuminance received at pixel coordinates (x, y) in the image captured by the imaging unit at time t, expressed in units of . Active light source radiant power The actual luminous power of the active illumination unit at time t, in W; target space depth. The spatial distance from the optical center of the imaging acquisition unit to the corresponding underwater entity at the pixel coordinates (x, y) at time t, in meters; the vehicle's velocity vector. Let be the translational velocity vector of the underwater vehicle in the three-dimensional world coordinate system at time t, a 3×1 column vector, with units of m / s, and components in the form of... ; This vehicle's rotational angular velocity vector Let be the angular velocity vector of the underwater vehicle at time t in the three-dimensional world coordinate system, a 3×1 column vector, with units of rad / s, and components in the form of... .
[0040] S12. Based on the Jaffe-McGlamery underwater imaging model, select the dark background area in the underwater image whose spatial distance is greater than 10 times the imaging sharpness distance, and calculate the water attenuation coefficient corresponding to the emission wavelength of the active light source center. Bilateral filtering was used to process the original observed light intensity sequence. Spatial domain filtering and separation were performed to extract the backscattered light intensity distribution field. The formula for calculating the water body attenuation coefficient is as follows:
[0041] ;
[0042] in, The center emission wavelength of the active lighting unit The corresponding water body attenuation coefficient, The average illuminance of the dark background area. The average spatial distance of the dark background area. The optical system efficiency of the imaging acquisition unit is preferably in the range of 0.6-0.85; in this embodiment, it is fixed at 0.75. The angle between the optical axis of the active illumination unit and the optical axis of the imaging acquisition unit is preferably in the range of 5°-15°. In this embodiment, it is fixed at 10°. The solid angle corresponding to a single pixel of the imaging acquisition unit is expressed in sr (steradian degrees).
[0043] In this embodiment, based on the Jaffe-McGlamery underwater imaging model, the total light intensity of underwater imaging is decomposed into direct propagation components, forward scattering components, and backscattering components. The direct propagation component of distant dark background areas approaches zero, and the total light intensity is approximately equal to the backscattering component. Therefore, areas in the underwater image with a spatial distance greater than 10 times the imaging sharpness distance are selected as dark background areas, and the average illuminance of these dark background areas is statistically analyzed. Average spatial distance Substituting into the formula for calculating the water body attenuation coefficient, we get After obtaining the water body attenuation coefficient, a bilateral filter was used to process the original observed light intensity sequence. Spatial domain filtering separation is performed. The preferred range for the bilateral filter kernel size is 5×5-15×15, but in this embodiment, it is fixed at 9×9. The preferred spatial domain standard deviation of the filter is 3, and the preferred range standard deviation is 0.1. After filtering, the backscattered light intensity distribution field is obtained. , Let be the backscattered light intensity distribution field at pixel coordinates (x, y) at time t. The light emitted by the active illumination unit enters the imaging acquisition unit after being reflected by suspended particles in the water.
[0044] S13. Extract the original observed light intensity sequence corresponding to two adjacent underwater images. , The initial surface flow field was calculated using the Lucas-Kanade algorithm. The motion parameters in the camera coordinate system are obtained by transforming a pre-calibrated extrinsic parameter matrix. Then, the three-dimensional coordinates of the corresponding spatial points in the camera coordinate system are calculated pixel by pixel using the pinhole camera back projection formula. Based on the motion parameters and three-dimensional coordinates in the camera coordinate system, the pixel-by-pixel global motion offset vector is calculated. For the initial surface sightseeing flow field Compensation is performed to obtain the relative optical flow field The apparent optical divergence field is obtained by calculating the divergence of the relative optical flow field. ,in, The time interval between two adjacent underwater images is expressed in seconds (s). In this embodiment, the frame rate is 30 fps. The value is fixed at 1 / 30s.
[0045] In this embodiment, the initial apparent flow field is calculated using the Lucas-Kanade algorithm. The preferred value for the number of pyramid layers in the algorithm is 3-4, and in this embodiment, it is 3. The preferred value for the calculation window size is 11×11-21×21, and in this embodiment, it is 15×15. Then, using a pre-calibrated extrinsic parameter matrix, the vehicle motion parameters in the world coordinate system are converted to motion parameters in the camera coordinate system. The conversion formula is: ; Where R is the 3×3 rotation extrinsic parameter matrix from the world coordinate system to the camera coordinate system, and T is the 3×1 translation extrinsic parameter vector from the world coordinate system to the camera coordinate system, both of which are pre-calibrated by the imaging acquisition unit; Let be the translational velocity vector of the imaging acquisition unit in the camera coordinate system at time t, a 3×1 column vector, with components in the form of... ; Let be the rotational angular velocity vector of the imaging acquisition unit in the camera coordinate system at time t, a 3×1 column vector, with components in the form of... Then, using the pinhole camera back projection formula, the three-dimensional coordinates of the spatial point corresponding to pixel (x,y) in the camera coordinate system are calculated. , The back projection formula is: ; ; , These are the equivalent focal lengths of the imaging acquisition unit in the x-axis and y-axis directions of the pixel plane, respectively, in pixels, obtained from the camera calibration results; , These are the coordinates of the principal point of the imaging acquisition unit along the x-axis and y-axis of the pixel plane, respectively, in pixels, obtained from the camera calibration results; based on the motion parameters and 3D coordinates in the camera coordinate system, a pixel-by-pixel global motion offset vector is calculated. , ; ; This formula incorporates pixel offsets caused by the translational and rotational motions of the underwater vehicle, ensuring that all global rigid motions are covered. Subsequently, the initial apparent optical flow field is compensated to obtain the relative optical flow field after removing its own motion. , ; Finally, the divergence of the relative optical flow field is calculated to generate the apparent optical flow divergence field. , The partial derivatives are calculated using the central difference method, with a difference step size of 1 pixel unit.
[0046] In this embodiment, step S2 includes the following specific steps:
[0047] S21, Set in pixels Local sliding window in the spatiotemporal domain centered on The spatiotemporal local sliding window The time dimension includes the current time t and the previous time. For local sliding windows in the spatiotemporal domain Backscattered light intensity distribution field within The normalization process is performed using the following formula: ;in, This is the normalized value for backscattering. , Let [0,1] be the minimum and maximum values of backscattered light intensity within a local sliding window in the spatiotemporal domain, respectively. Divide [0,1] into N equal intervals and count the proportion of pixels in each interval to the total number of pixels within the local sliding window in the spatiotemporal domain to obtain the interval probability distribution of the backscattered normalized value. , where i is the interval index, with a value range of 1-N.
[0048] In this embodiment, a local sliding window in the spatiotemporal domain The preferred spatial dimensions are between 7×7 and 13×13, but in this embodiment, they are fixed at 9×9. To prevent the constant from being divided by zero in the denominator, the value is taken as... The value of N is 16.
[0049] S22. Local sliding window in the spatiotemporal domain The surface of the tourist flow field The normalization process is performed using the following formula: ; , Local sliding windows in the spatiotemporal domain Minimum and maximum absolute values of internal optical divergence. The normalized absolute value of optical flow divergence is obtained by statistically analyzing the proportion of pixels within N intervals to the total number of pixels within a local sliding window in the spatiotemporal domain. Where j is the interval index, with a value range of 1-N;
[0050] S23. Interval probability distribution of normalized backscatter values Interval probability distribution of the normalized absolute value of optical divergence Substituting into the joint information entropy formula, the spatiotemporal optical flow divergence-scattering coupling entropy is calculated. The formula is: Spatiotemporal optical flow divergence-scattering coupling entropy is a scalar field that quantifies the degree of interference coupling. In this formula, the closer the joint probability distribution is to a uniform distribution, the higher the calculated entropy value, and the corresponding region is the interference region of strong scattering and high divergence turbulence coupling; the more concentrated the joint probability distribution is, the lower the entropy value, and the corresponding region is the effective region of rigid targets.
[0051] In this embodiment, step S3 includes the following specific steps:
[0052] S31, the original observed light intensity sequence The shallow feature map is extracted after inputting the backbone layer of the YOLOv8 object detection network and performing three downsampling operations. ;
[0053] S32, Spatiotemporal optical divergence-scattering coupling entropy Perform inverse exponential mapping to generate the attention weight matrix. The formula for calculating the attention weight matrix is as follows: ;in, The attention weight matrix is adjusted using bilinear interpolation to match the shallow features. Figure One The spatial dimensions are adjusted, and the attention weight matrix is output after resizing. ;
[0054] S33. Adjust the size of the attention weight matrix. With shallow feature map Element-wise multiplication along the spatial dimension yields a clean target feature map. , ;in, This is the element-wise multiplication operator.
[0055] In this embodiment, the original observed light intensity sequence The corresponding underwater image is scaled to 640×640 pixels and input into the backbone layer of the YOLOv8 object detection network. The backbone layer contains Conv convolutional layers, C2f modules, and SPPF modules. After the input image is downsampled three times by the backbone layer, a shallow feature map with a size of 80×80 and 256 channels is output. ; Spatiotemporal optical flow divergence-scattering coupling entropy Generate attention weight matrix by performing inverse exponential mapping The formula for calculating the attention weight matrix is: ; The adjustment coefficient for the attention weight calculation is preferably in the range of 0.5-2.0. In this embodiment, it is fixed at 1. This mapping makes the weight of the high-entropy interference region approach 0 and the weight of the low-entropy target region approach 1, thereby achieving interference suppression and target enhancement. Subsequently, the initial attention weight matrix is adjusted to match the shallow features through bilinear interpolation. Figure One Given an 80×80 spatial size, output the resized attention weight matrix. The resized attention weight matrix is then multiplied element-wise with the shallow feature map along the spatial dimension. During the operation, the weight matrix is broadcast along the channel dimension to ensure that each channel of the multi-channel feature map completes its corresponding modulation, resulting in a clean target feature map. .
[0056] In this embodiment, the specific content of S4 is as follows:
[0057] S41. Extract the clean target feature map. The YOLOv8-based PAFPN multi-scale feature pyramid network is input, and bidirectional feature fusion is performed through a top-down upsampling path and a bottom-up downsampling path, outputting three sets of multi-scale feature maps at different scales. These multi-scale feature maps are then simultaneously input into a decoupled classification head and regression head. The classification head, after passing through three consecutive convolutional layers and a sigmoid activation layer, outputs the dynamic target class probability corresponding to each detection box. The regression head outputs the target bounding box coordinates corresponding to each detection box through three consecutive convolutional layers and linear activation layers;
[0058] S42. Match the bounding boxes of targets in adjacent frames using a preset matching threshold. Analyze the changes in the center coordinates of the adjacent frames of successfully matched targets against the time interval of the acquisition. Calculate the target motion vector The final output is the dynamic target category probability. The corresponding target bounding box coordinates and motion vectors The formula for calculating the target motion vector is: ;in, Let be the center pixel coordinates of the target bounding box at time t. for The center pixel coordinates of the bounding box of the same target that is successfully matched at any given time.
[0059] In this implementation, the clean target feature map The YOLOv8-based PAFPN multi-scale feature pyramid network is input. This network includes a top-down upsampling path and a bottom-up downsampling path. The top-down path transfers strong semantic features from high-level layers to low-level layers, enhancing the class recognition of small targets. The bottom-up path transfers strong detail features from low-level layers to high-level layers, enhancing the contour localization accuracy of large targets. After bidirectional feature fusion, three sets of multi-scale feature maps at different scales are output, with sizes of 80×80, 40×40, and 20×20, respectively, corresponding to the detection requirements of small, medium, and large targets. The three sets of multi-scale feature maps are simultaneously input into the decoupled classification head and regression head, which share the multi-scale feature input. The classification head and regression head respectively complete the class probability calculation and bounding box regression. The classification head consists of three consecutive 3×3 convolutional layers (each followed by a batch normalization layer and a SiLU activation function) and one 1×1 convolutional layer + Sigmoid activation layer, outputting the dynamic target class probability corresponding to each detection box. The regression head consists of three consecutive 3×3 convolutional layers (each followed by a batch normalization (BN) layer and a SiLU activation function) and one 1×1 convolutional layer plus a linear activation layer. It outputs the target bounding box coordinates for each detection box, including the bounding box center coordinates, width, and height. A matching threshold of 0.5 is set. The Interchange of Units (IOU) is calculated pairwise for target bounding boxes in adjacent frames. Only bounding boxes with an IOU greater than the matching threshold are considered to be the same target, completing the one-to-one matching of targets in adjacent frames. The changes in the center coordinates of successfully matched targets in adjacent frames are correlated with the acquisition time interval. Calculate the target motion vector The final output is the dynamic target category probability. The corresponding target bounding box coordinates and motion vectors .
[0060] It should be noted that in existing underwater dynamic target detection technologies, the backscattered light spots generated by suspended particles in the water and the non-rigid pseudo-motion caused by water turbulence exhibit strong coupling characteristics. The superposition of these two types of interference leads to high false detection and high false negative rates in the detection network. Existing technologies typically handle optical degradation and motion interference separately, failing to effectively suppress coupled interference. This method, based on the coupling characteristics of underwater interference, constructs a spatiotemporal optical flow divergence-scattering coupling entropy to accurately locate the coupled interference region. Then, it generates an attention weight matrix through inverse exponential mapping, directly suppressing the ineffective activation of the interference region at the feature level. This prevents interference information from being amplified in subsequent feature fusion and calculation in the network, fundamentally solving the false detection and false negative problems in underwater dynamic target detection and improving the robustness and detection accuracy of the method in complex underwater environments.
[0061] The threshold and weight settings can be based on the default settings of this invention, or they can be set by the operator.
[0062] Example 2: This example provides an electronic device, including a processor and a memory, wherein the memory stores a computer program that can be called by the processor; the processor executes the above-described underwater dynamic target detection and identification method by calling the computer program stored in the memory.
[0063] The electronic device can vary considerably depending on its configuration or performance. It may include one or more Central Processing Units (CPUs) and one or more memories, wherein the memory stores at least one computer program, which is loaded and executed by the processor to implement the underwater dynamic target detection and identification method provided in the above-described embodiment. The electronic device may also include other components for implementing its functions; for example, it may have wired or wireless network interfaces and input / output interfaces for data input and output. Further details are omitted in this embodiment.
[0064] Those skilled in the art will recognize that this invention can be implemented as a system, method, or computer program product. Therefore, this invention can be implemented in the following forms: it can be entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, this invention can also be implemented as a computer program product contained in one or more computer-readable media, which includes computer-readable program code.
[0065] Any combination of one or more computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0066] This invention is described with reference to flowchart illustrations and block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and block diagrams, as well as combinations of blocks in the flowchart illustrations and block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart. Figure One One or more processes and boxes Figure One A device that provides the functions specified in one or more boxes.
[0067] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure One One or more processes and boxes Figure One The steps of the function specified in one or more boxes.
[0068] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A method for underwater dynamic target detection and recognition, characterized in that: The specific steps include the following: S1. An underwater vehicle equipped with an imaging acquisition unit, an active illumination unit, a depth sensing unit, and a motion sensing unit acquires underwater images and obtains underwater multi-source physical signals. Based on the underwater multi-source physical signals, the backscattered light intensity distribution field and the surface divergence field are extracted. S2. Based on the backscattered light intensity distribution field and the apparent light divergence field, the spatiotemporal light divergence-scattering coupling entropy is calculated. S3. Generate an attention weight matrix based on the spatiotemporal optical flow divergence-scattering coupling entropy, and modulate the attention weight matrix element-wise with the shallow feature map extracted by the target detection network to obtain a pure target feature map. S4. Based on the clean target feature map, perform multi-scale feature extraction and classification regression processing, and output the dynamic target category probability and the corresponding target bounding box coordinates and motion vector.
2. The underwater dynamic target detection and recognition method according to claim 1, characterized in that: S1 includes the following specific steps: S11. An underwater vehicle equipped with an imaging acquisition unit, an active illumination unit, a depth sensing unit, and a motion sensing unit acquires underwater images and obtains underwater multi-source physical signals, wherein the underwater multi-source physical signals include the original observed light intensity sequence. Active light source radiation power Target space depth The vehicle's velocity vector and the vehicle's rotational angular velocity vector ;in, This represents the horizontal coordinate index within the pixel plane. This represents the vertical coordinate index within the pixel plane. This refers to the time corresponding to the acquisition time sequence; S12. Based on the Jaffe-McGlamery underwater imaging model, select the dark background area in the underwater image whose spatial distance is greater than 10 times the imaging sharpness distance, and calculate the water attenuation coefficient corresponding to the emission wavelength of the active light source center. Bilateral filtering was used to process the original observed light intensity sequence. Spatial domain filtering and separation were performed to extract the backscattered light intensity distribution field. The formula for calculating the water body attenuation coefficient is as follows: ; in, The center emission wavelength of the active lighting unit The corresponding water body attenuation coefficient, The average illuminance of the dark background area. The average spatial distance of the dark background area. For the optical system efficiency of the imaging acquisition unit, The angle between the optical axis of the active illumination unit and the optical axis of the imaging acquisition unit. The solid angle corresponding to a single pixel in the imaging acquisition unit; S13. Extract the original observed light intensity sequence corresponding to two adjacent underwater images. , The initial surface flow field was calculated using the Lucas-Kanade algorithm. The motion parameters in the camera coordinate system are obtained by transforming a pre-calibrated extrinsic parameter matrix. Then, the three-dimensional coordinates of the corresponding spatial points in the camera coordinate system are calculated pixel by pixel using the pinhole camera back projection formula. Based on the motion parameters and three-dimensional coordinates in the camera coordinate system, the pixel-by-pixel global motion offset vector is calculated. For the initial surface sightseeing flow field Compensation is performed to obtain the relative optical flow field The apparent optical divergence field is obtained by calculating the divergence of the relative optical flow field. ,in, This represents the time interval between the acquisition of two adjacent underwater images.
3. The underwater dynamic target detection and recognition method according to claim 2, characterized in that: S2 includes the following specific steps: S21, Set in pixels Local sliding window in the spatiotemporal domain centered on The spatiotemporal local sliding window The time dimension includes the current time t and the previous time. For local sliding windows in the spatiotemporal domain Backscattered light intensity distribution field within The normalization process is performed using the following formula: ;in, This is the normalized value for backscattering. , Let [0,1] be the minimum and maximum values of backscattered light intensity within a local sliding window in the spatiotemporal domain, respectively. Divide [0,1] into N equal intervals and count the proportion of pixels in each interval to the total number of pixels within the local sliding window in the spatiotemporal domain to obtain the interval probability distribution of the backscattered normalized value. Where i is the interval index, and its value ranges from 1 to N; S22. Local sliding window in the spatiotemporal domain The surface of the tourist flow field The normalization process is performed using the following formula: ; , Local sliding windows in the spatiotemporal domain Minimum and maximum absolute values of internal optical flow. The normalized absolute value of optical flow divergence is obtained by statistically analyzing the proportion of pixels within N intervals to the total number of pixels within a local sliding window in the spatiotemporal domain. Where j is the interval index, with a value range of 1-N; S23. Interval probability distribution of normalized backscatter values Interval probability distribution of the normalized absolute value of optical divergence Substituting into the joint information entropy formula, the spatiotemporal optical flow divergence-scattering coupling entropy is calculated. The formula is: .
4. The underwater dynamic target detection and recognition method according to claim 3, characterized in that: S3 includes the following specific steps: S31, the original observed light intensity sequence The shallow feature map is extracted from the backbone layer of the YOLOv8 object detection network after three downsampling operations. ; S32, Spatiotemporal optical divergence-scattering coupling entropy Perform inverse exponential mapping to generate the attention weight matrix. The formula for calculating the attention weight matrix is as follows: ;in, To adjust the coefficients, the attention weight matrix is adjusted to a spatial size consistent with the shallow feature map using bilinear interpolation, and the adjusted attention weight matrix is output. ; S33. Adjust the size of the attention weight matrix. With shallow feature map Element-wise multiplication along the spatial dimension yields a clean target feature map. , ;in, This is the element-wise multiplication operator.
5. The underwater dynamic target detection and recognition method according to claim 4, characterized in that: The specific content of S4 is as follows: S41. Extract the clean target feature map. The YOLOv8-based PAFPN multi-scale feature pyramid network is input, and bidirectional feature fusion is performed through a top-down upsampling path and a bottom-up downsampling path, outputting three sets of multi-scale feature maps at different scales. These multi-scale feature maps are then simultaneously input into a decoupled classification head and regression head. The classification head, after passing through three consecutive convolutional layers and a sigmoid activation layer, outputs the dynamic target class probability corresponding to each detection box. The regression head outputs the target bounding box coordinates corresponding to each detection box through three consecutive convolutional layers and linear activation layers; S42. Match the bounding boxes of targets in adjacent frames using a preset matching threshold. Analyze the changes in the center coordinates of the adjacent frames of successfully matched targets against the time interval of the acquisition. Calculate the target motion vector The final output is the dynamic target category probability. The corresponding target bounding box coordinates and motion vectors The formula for calculating the target motion vector is: ;in, Let be the center pixel coordinates of the target bounding box at time t. for The center pixel coordinates of the bounding box of the same target that is successfully matched at any given time.