A method for underwater target detection and positioning using a monocular camera based on motion inversion

Through the motion inversion-based monocular camera underwater target detection method and the detection missed frame compensation algorithm of kinematic inversion, the problem of discontinuous target detection in underwater environment is solved, and the accurate and stable positioning of the target is achieved.

CN119360191BActive Publication Date: 2025-09-12NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411410847.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-10
Publication Date
2025-09-12
Estimated Expiration
2044-10-10

AI Technical Summary

Technical Problem

Existing target detection models cannot perform continuous and stable target detection and positioning in underwater environments due to image distortion caused by water waves, turbulence, and refraction.

Method used

A monocular camera underwater target detection method based on motion inversion is adopted. Through continuous multi-frame image detection, the target center coordinates are calculated using the detection missed frame compensation algorithm of kinematic inversion. The actual physical size and position of the target are obtained by combining the pixel width of the target in the image and the depth change of the vehicle. The ground, carrier and image coordinate systems are established to locate the target.

Benefits of technology

It realizes continuous and stable detection and positioning of targets in underwater environments, and improves the accuracy and continuity of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360191B_ABST
    Figure CN119360191B_ABST
Patent Text Reader

Abstract

The present invention relates to the fields of underwater robot navigation and computer vision technology, and more specifically to a method for underwater target detection and positioning using a monocular camera based on motion inversion. The method comprises the following steps: capturing underwater images; detecting and compensating for targets in the underwater images, and obtaining the target's actual physical width and height; obtaining the lateral deviation angle of the target's center relative to the camera's center; obtaining the straight-line distance from the target's center to the camera's center; obtaining the longitudinal height between the target's center and the camera's center; and locating the target. The present invention can simultaneously detect targets and robustly estimate their size and position, ensuring the continuity and accuracy of target detection and positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of underwater robot navigation and computer vision technology, and in particular to a monocular camera underwater target detection and positioning method based on motion inversion. Background Art

[0002] Underwater vehicles are indispensable tools for marine science, submarine pipeline inspection, resource exploration, underwater archaeology, and even reconnaissance. They can perform tasks in deep-sea environments beyond human reach, collecting data, conducting observations, and conducting interventions. When performing underwater missions, underwater vehicles typically need to be able to accurately identify and locate target objects. Cameras, as visual perception sensors, are widely equipped in these vehicles. They use visual information to perceive and understand the underwater environment, enabling them to perform complex tasks such as target detection, obstacle avoidance, charging and docking, and intervention.

[0003] Although binocular vision-based positioning methods can provide depth information and achieve positioning in three-dimensional space, they are relatively costly and bulky, making them difficult to integrate into small or micro underwater vehicles, and they perform poorly in low-light or turbid underwater environments. LiDAR can provide high-precision distance and direction information and perform stably under various lighting conditions, but it is expensive and bulky, making it unsuitable for small underwater vehicles. It also suffers from severe attenuation in underwater environments, and its effective working range is limited. Traditional positioning methods based on monocular cameras (such as the similar triangle method) are low-cost and easy to install and maintain, but they require prior knowledge of camera intrinsic parameters such as focal length, which is usually not feasible in autofocus digital cameras, and information such as relative depth and lateral deviation cannot be directly obtained. Mainstream computer vision task datasets, such as COCO and VOC, although relatively rich and mature, are all generated from objects in the air. In summary, when object detection models trained with these datasets are used to detect underwater targets, the images are distorted due to factors such as water waves, turbulence, and refraction, resulting in the inability to perform continuous and stable detection, thereby affecting the continuity and accuracy of target detection and positioning.

[0004] Therefore, it is necessary to provide a monocular camera underwater target detection and positioning method based on motion inversion to solve the above problems. Summary of the Invention

[0005] The present invention provides a method for underwater target detection and positioning using a monocular camera based on motion inversion, so as to solve the problem that when detecting underwater targets using existing target detection models trained with data sets, images are distorted due to factors such as water waves, turbulence, and refraction, resulting in the inability to perform continuous and stable detection, thereby affecting the continuity and accuracy of target detection and positioning.

[0006] The present invention provides a method for underwater target detection and positioning using a monocular camera based on motion inversion, which adopts the following technical solutions, including:

[0007] Collect multiple frames of continuous underwater images;

[0008] Target detection is performed on underwater images using a target detection algorithm. If a target is detected in a current underwater image, but not in the next underwater image of the current underwater image, and the interval between the previous underwater image and the next underwater image is less than a preset time, a kinematic inversion detection omission frame compensation algorithm is used to infer the target center coordinates of the next underwater image. If the next two underwater images also do not detect a target, and the interval between the current underwater image and the next two underwater images is less than a preset time, a kinematic inversion detection omission frame compensation algorithm is used to infer the target center coordinates of the next two underwater images until the target is detected in the underwater image. The target detection result includes a target label, a target confidence, a central horizontal coordinate and a central vertical coordinate of the target, and a pixel width and pixel height of the target in the image. The actual physical width and actual physical height of the target are obtained based on the pixel width and pixel height of the target in the image, the change in the vertical coordinate of the target center in two adjacent frames, and the change in the depth of the vehicle.

[0009] Obtain a lateral deviation angle of the target center relative to the camera center according to the central abscissa coordinate of the target in each frame of underwater image, the central abscissa coordinate and lateral size of each frame of underwater image, and the lateral field of view angle of the camera;

[0010] Obtain the straight-line distance from the center of the target to the center of the camera according to the lateral size of the underwater image of the target frame, the width of the target in the image, the actual physical width of the target, and the lateral field of view of the camera;

[0011] Obtain the vertical height between the center of the target and the center of the camera according to the vertical size of the underwater image of the target frame, the vertical field of view of the camera, the vertical coordinate of the center of the target, the vertical coordinate of the center of the underwater image of the target frame, and the straight-line distance from the center of the target to the center of the camera;

[0012] The target is located based on the longitudinal height between the target center and the camera center, the straight-line distance from the target center to the camera center, and the lateral deviation angle of the target center relative to the camera center.

[0013] Preferably, the steps of calculating the target center coordinates of the next underwater image frame using the detection missed frame compensation algorithm based on kinematic inversion are:

[0014] Establish a ground coordinate system and a carrier coordinate system: the ground coordinate system takes the initial buoyancy center of the vehicle as its origin, the x-axis of the ground coordinate system points to due north, the y-axis of the ground coordinate system is perpendicular to the x-axis and points to due east, and the z-axis of the ground coordinate system is perpendicular to the x-axis and y-axis to form a right-handed system; the carrier coordinate system takes the buoyancy center of the vehicle as its origin, the x-axis of the carrier coordinate system is along the length direction of the underwater vehicle and points forward, the y-axis of the carrier coordinate system is perpendicular to the x-axis and points to starboard, and the z-axis of the carrier coordinate system is perpendicular to the x-axis and y-axis to form a right-handed system;

[0015] The target center coordinates of the next underwater image frame are calculated based on the pixels corresponding to the horizontal unit angle and vertical unit angle of the current underwater image frame, the center horizontal coordinate and center vertical coordinate of the target in the current underwater image frame, the roll angle, pitch angle, navigation angle, depth of the vehicle when collecting two adjacent underwater image frames, and the distance from the center of the target in the current frame to the vehicle.

[0016] Preferably, the expression of the target center coordinates of the next frame of underwater image is:

[0017]

[0018] Where, pD x Indicates the corresponding pixel per horizontal unit angle of the underwater image; pD y Represents the corresponding pixel of the vertical unit angle of the underwater image; Indicates the The central vertical coordinate of the target in the frame underwater image; Indicates the The central vertical coordinate of the target in the frame underwater image; Indicates the The central vertical coordinate of the target in the frame underwater image; Indicates the The central horizontal coordinate of the target in the frame underwater image; Indicates collection of Frame underwater image to capture the first The change in the vehicle's roll angle when framing underwater images; Indicates collection of Frame underwater image to capture the first The pitch angle change of the vehicle when framing underwater images; Ψ Indicates collection of Frame underwater image to capture the first The change in the yaw angle of the vehicle when framing underwater images; Indicates collection of Frame underwater image to capture the first The depth change of the vehicle when framing underwater images; Indicates the The distance from the vehicle to the center of the target in the underwater image.

[0019] Preferably, obtaining the actual physical width and actual physical height of the target includes:

[0020]

[0021]

[0022] Where, Indicates the actual physical width of the target; Indicates the actual physical height of the target; Indicates the pixel width of the target in the underwater image; Indicates the pixel height of the target in the underwater image; Indicates collection of Frame underwater image to capture the first The depth change of the vehicle when framing underwater images; Indicates collection of Frame underwater image to capture the first The change in the vertical coordinate of the target center when framing the underwater image; Indicates the physical distance corresponding to the unit pixel.

[0023] Preferably, obtaining the lateral deviation angle of the target center relative to the camera head center includes:

[0024]

[0025] Where, Indicates the lateral deviation angle of the target center relative to the camera center; Represents the central horizontal coordinate of the underwater image in the image coordinate system; Indicates the horizontal coordinate of the target center in the image coordinate system; is the horizontal field of view of the camera; is the horizontal pixel size of the underwater image.

[0026] Preferably, obtaining the straight-line distance from the center of the target to the center of the camera includes:

[0027]

[0028] Where, Indicates the straight-line distance from the center of the target to the center of the camera; is the horizontal field of view of the camera; is the horizontal pixel size of the underwater image; Indicates the actual physical width of the target; Indicates the pixel width of the target in the underwater image.

[0029] Preferably, obtaining the longitudinal height between the center of the target and the center of the camera head includes:

[0030]

[0031] Where, Indicates the vertical height between the center of the target and the center of the camera lens; Indicates the straight-line distance from the center of the target to the center of the camera; is the vertical field of view of the camera; is the vertical pixel size of the underwater image; y c Represents the central vertical coordinate of the underwater image in the image coordinate system; y t Indicates the vertical coordinate of the target center in the image coordinate system.

[0032] Preferably, the target detection algorithm adopts the YOLOv4-tiny algorithm.

[0033] The beneficial effects of the present invention are:

[0034] The target is acquired based on the target detection algorithm. In the continuous frames of underwater images, the target is detected in the previous frame of underwater image, but not in the next frame of underwater image. Then, the target center coordinates of the next frame of underwater image are obtained and estimated based on the detection missed frame compensation algorithm based on kinematic inversion. The actual physical width and actual physical height of the target are obtained according to the pixel width and pixel height of the target in the image, the vertical coordinate change of the target center in two adjacent frames and the depth change of the vehicle. Finally, the longitudinal height between the target center and the camera center, the straight-line distance from the target center to the camera center, and the lateral deviation angle of the target center relative to the camera center are obtained based on the positioning algorithm. The target is positioned according to the longitudinal height between the target center and the camera center, the straight-line distance from the target center to the camera center, and the lateral deviation angle of the target center relative to the camera center. In summary, the present invention can simultaneously detect the target and robustly estimate the size and position, thereby ensuring the continuity and accuracy of target detection and positioning. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0036] Figure 1 This is a flow chart of a method for underwater target detection and positioning using a monocular camera based on motion inversion according to the present invention;

[0037] Figure 2 An underwater image of a starfish as a target in an embodiment of the present invention;

[0038] Figure 3 Schematic diagram of the ground coordinate system, carrier coordinate system and image coordinate system in an embodiment of the present invention;

[0039] Figure 4 Schematic diagram of the lateral deviation angle of the target center relative to the camera head center in an embodiment of the present invention;

[0040] Figure 5 Schematic diagram of the straight-line distance from the center of the target to the center of the camera head in an embodiment of the present invention;

[0041] Figure 6 Schematic diagram of the longitudinal height between the center of the target and the center of the camera head in an embodiment of the present invention;

[0042] Figure 7 A comparison chart of the detection results of the YOLOv4-tiny algorithm for underwater images in an embodiment of the present invention and the results of the detection and missing frame compensation algorithm for underwater images using kinematic inversion;

[0043] Figure 8 Result diagram of lateral position estimation error corresponding to the detection missing frame compensation algorithm based on the inversion of the static model and the detection missing frame compensation algorithm based on the inversion of the motion model during the experiment of the embodiment of the present invention;

[0044] Figure 9 Result diagram of longitudinal position estimation error corresponding to the detection missing frame compensation algorithm based on the inversion of the static model and the detection missing frame compensation algorithm based on the inversion of the motion model during the experiment of the embodiment of the present invention;

[0045] Figure 10 Schematic diagram of the algorithm framework of a monocular camera underwater target detection and positioning method based on motion inversion of the present invention. DETAILED DESCRIPTION

[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0047] An embodiment of a method for underwater target detection and positioning using a monocular camera based on motion inversion of the present invention is as follows: Figure 1 and Figure 10 Shown, including:

[0048] S1, collecting underwater images;

[0049] Specifically, in this embodiment, a monocular camera of the vehicle is used to capture multiple frames of continuous underwater images.

[0050] S2. Detect and compensate the target in the underwater image, and obtain the actual physical width and actual physical height of the target;

[0051] Step 21: The steps of detecting and compensating the target in the underwater image are as follows:

[0052] like Figure 10 As shown, the target detection algorithm is used to detect the target in the underwater image. If the current frame underwater image detects the target, the next frame underwater image of the current frame underwater image does not detect the target, and the interval time between the previous frame underwater image and the next frame underwater image is less than the preset time length, the kinematic inversion detection missing frame compensation algorithm is used to calculate the target center coordinates of the next frame underwater image; if the next two frames underwater image also do not detect the target, and the interval time between the current frame underwater image and the next two frames underwater image is less than the preset time length, the kinematic inversion detection missing frame compensation algorithm is used to calculate the target center coordinates of the next two frames underwater image, until The underwater image detects the target. If the next three frames of underwater images also do not detect the target, and the interval time between the current frame underwater image and the next three frames underwater image is less than the preset duration, the kinematic inversion detection missed frame compensation algorithm is used to infer the target center coordinates of the next three frames underwater images, and so on, until the underwater image detects the target. If the target is not detected subsequently, the interval time is calculated with the most recent frame underwater image in which the target is detected. The preset duration in this embodiment is 2s. It should be noted that, in this embodiment, the target detection algorithm is yolov4-tiny. The detection frequency of yolov4-tiny when running on jetson nano software is about 15hz, that is, 15 frames of underwater images can be detected in one second, and 30 frames of underwater images are detected in 2s. If no target is detected in 30 consecutive underwater image frames, it means that the target is out of the camera's shooting range and no further inference is performed. The target detection results include the target label, target confidence, the target's center horizontal and vertical coordinates, and the target's pixel width and pixel height in the image. The actual physical width and actual physical height of the target are obtained based on the target's pixel width and pixel height in the image.

[0053] Step 211. In this embodiment, the YOLOv4-tiny algorithm is used to detect targets in underwater images. The YOLO (You Only Look Once) series of algorithms is well-known for its efficient end-to-end detection process. It treats the target detection problem as a regression problem and directly predicts the coordinates of the bounding box and the category probability in a single neural network. YOLOv4-tiny is a lightweight version of YOLOv4. It uses a network structure similar to YOLOv4, but adopts a series of strategies to reduce model complexity and computational complexity, such as reducing the number of network layers and the number of channels. Because YOLOv4-tiny can achieve efficient target detection in an environment with limited computing resources, it can improve the detection rate while maintaining high detection accuracy. It is suitable for deployment on single-board computers with limited computing power. In addition, the technology is relatively mature and the software development is relatively convenient. Therefore, this embodiment will not be described in detail. Therefore, this embodiment selects the YOLOv4-tiny algorithm.

[0054] Step 212, the steps of calculating the target center coordinates of the next underwater image frame (i.e., missing frame compensation) based on the detection missing frame compensation algorithm of kinematic inversion are as follows:

[0055] like Figure 3 As shown, in order to describe the motion of the aircraft, a ground coordinate system { n} and the carrier coordinate system { b}, the units of the coordinate axes are meters. The ground coordinate system is fixed to the earth, where the ground coordinate system { n}Origin o n Select the initial buoyancy center of the spacecraft at the initial moment, axis o n x n Pointing due north, the axis o n y n Perpendicular to the axis o n x n and points due east, the axis o n z n Perpendicular to the axis o n x n and o n y n , direction makes o n x n y nz n Become a right-handed system, pointing directly downward; carrier coordinate system { b}Fixed to the underwater vehicle, coordinate origin o b Select the center of buoyancy of the vehicle, axis o b x b Along the longitudinal axis of the underwater vehicle and pointing forward, the axis o b y b Perpendicular to the axis o b x b and pointed to starboard, axis o b z b Perpendicular to the axis o b x b and o b y b , direction makes o b x b y b z b The right-hand system points downwards. The aircraft has six degrees of freedom, including forward and backward motion, sideways motion, heave motion, roll motion, pitch motion, and yaw motion. The position of the aircraft relative to the ground coordinate system is given by the coordinates ( x n ,y n ,z n ), the attitude relative to the ground coordinate system is described by the three-axis Euler angle ( , , ) description; To describe the position of the target in the underwater image, this embodiment establishes an image coordinate system { p}, the unit of the coordinate axis is pixel. The origin of the image coordinate system o Select the upper left corner of the underwater image, axis ox Pointing due right, axis oy Perpendicular to the axis ox Pointing directly downward. In a frame of underwater image, the position of the target center is represented by its coordinates ( x t ,y t ) is determined, and the image center coordinates are set to ( x c ,y c ), assuming that the size and position of the target remain unchanged, the spacecraft makes a small movement, that is, only changes its three-axis Euler angle and depth in a small range, and the distance from the spacecraft to the target center D The target center position is calculated by the following formula:

[0056]

[0057] in,

[0058]

[0059] Where, pD x Indicates the corresponding pixel per horizontal unit angle of the underwater image; pD y Represents the corresponding pixel of the vertical unit angle of the underwater image; Indicates the The central vertical coordinate of the target in the frame underwater image; Indicates the The central vertical coordinate of the target in the frame underwater image; Indicates the The central vertical coordinate of the target in the frame underwater image; Indicates the The central horizontal coordinate of the target in the frame underwater image; Indicates collection of Frame underwater image to capture the first The change in the vehicle's roll angle when framing underwater images; Indicates collection of Frame underwater image to capture the first The pitch angle change of the vehicle when framing underwater images; Ψ Indicates collection of Frame underwater image to capture the first The change in the yaw angle of the vehicle when framing underwater images; Indicates collection of Frame underwater image to capture the first The depth change of the vehicle when framing underwater images; Indicates that the aircraft has reached The distance between the target center in the underwater image frame is calculated. When the "breathing frame" frequency is high, the time interval between detected target frames is short, and the vehicle only changes its three-axis Euler angles and depth within a small range, which can be considered to meet the assumption of small vehicle motion. As the "breathing frame" frequency increases, the time interval between detected target frames increases, the actual situation deviates from the assumption, and the estimated error gradually increases. When the YOLOv4-tiny algorithm is running on the vehicle, the frame rate is approximately 15fps. After multiple tests, it was found that setting the maximum estimated frame number to 30 is appropriate. That is, the algorithm is used only when the time interval between detected target frames is less than 2s; otherwise, the target is considered lost. It should be noted that the "breathing frame" is the box used by the vehicle to select the target in the image during target detection. When detection is discontinuous, the detection frame in the image appears and disappears, a phenomenon vividly referred to as "breathing frame." Experiments have shown that the breathing frequency is higher when the vehicle's motion is small and lower when the motion is large. The compensation algorithm based on motion inversion is designed to solve the problem of unstable and discontinuous detection. When the target detection algorithm fails to detect the target, the motion information is used to estimate the target position.

[0060] Step 23: Obtain the actual physical width and actual physical height of the target as follows:

[0061] For a single underwater target, its actual physical width is L , the actual physical height is H The target detection result output by the YOLOv4-tiny algorithm is its label, confidence, and the horizontal coordinate of the target center. x t and the vertical axis y t , the pixel width of the target in the underwater image w and pixel height h Combined with the vehicle's own motion information, the actual physical width of the target is L and actual physical height H Calculated by the following formula:

[0062]

[0063] in,

[0064]

[0065] Where, Indicates the actual physical width of the target; Indicates the actual physical height of the target; Indicates the pixel width of the target in the underwater image; Indicates the pixel height of the target in the underwater image; Indicates collection of Frame underwater image to capture the first The depth change of the vehicle when framing underwater images; Indicates collection of Frame underwater image to capture the first The change in the vertical coordinate of the target center when framing the underwater image; Indicates the physical distance corresponding to the unit pixel.

[0066] S3, obtaining the lateral deviation angle of the target center relative to the camera center;

[0067] Specifically, the lateral deviation angle of the target center relative to the camera center is obtained according to the central lateral coordinate of the target, the central lateral coordinate and lateral size of the underwater image of the frame where the target is located, and the lateral field of view of the camera.

[0068] The schematic diagram of the lateral deviation angle of the target center relative to the camera center is as follows: Figure 4 As shown, the expression of the lateral deviation angle of the target center relative to the camera center is:

[0069]

[0070] Where, Indicates the lateral deviation angle of the target center relative to the camera center; Represents the central horizontal coordinate of the underwater image in the image coordinate system; Indicates the horizontal coordinate of the target center in the image coordinate system; is the horizontal field of view of the camera; is the horizontal pixel size of the underwater image.

[0071] S4, obtaining the straight-line distance from the target center to the camera center;

[0072] Specifically, the straight-line distance from the center of the target to the center of the camera is obtained according to the lateral size of the underwater image of the frame where the target is located, the width of the target in the image, the actual physical width of the target, and the lateral field of view of the camera.

[0073] Among them, the schematic diagram of the straight-line distance from the target center to the camera center is as follows Figure 5 As shown, the expression of the straight-line distance from the target center to the camera center is:

[0074]

[0075] Where, Indicates the straight-line distance from the center of the target to the center of the camera; is the horizontal field of view of the camera; is the horizontal pixel size of the underwater image; Indicates the actual physical width of the target; Represents the pixel width of the target in the underwater image. It should be noted that this calculation method approximates the width of the target as an arc, which simplifies the calculation.

[0076] S5, obtaining the longitudinal height between the target center and the camera head center;

[0077] Specifically, the longitudinal height between the center of the target and the center of the camera is obtained according to the longitudinal size of the underwater image of the frame where the target is located, the longitudinal field of view of the camera, the central longitudinal coordinate of the target, the central longitudinal coordinate of the underwater image of the frame where the target is located, and the straight-line distance from the center of the target to the center of the camera.

[0078] The schematic diagram of the vertical height between the target center and the camera center is as follows: Figure 6 As shown, the expression of the vertical height between the target center and the camera center is:

[0079]

[0080] Where, Indicates the vertical height between the center of the target and the center of the camera lens; Indicates the straight-line distance from the center of the target to the center of the camera; is the vertical field of view of the camera; is the vertical pixel size of the underwater image; y c Represents the central vertical coordinate of the underwater image in the image coordinate system; y t Indicates the vertical coordinate of the target center in the image coordinate system

[0081] S6. Target positioning;

[0082] The target can be positioned based on the longitudinal height between the target center and the camera center, the straight-line distance from the target center to the camera center, and the lateral deviation angle of the target center relative to the camera center.

[0083] The following is described with reference to specific embodiments:

[0084] The positioning method of the present invention is encapsulated as a ROS algorithm node and applied to an underwater vehicle. The vehicle can detect and locate starfish using the positioning method of the present invention.

[0085] Conduct experiments to verify the detection missed frame compensation algorithm based on kinematic inversion:

[0086] The algorithm was written into the computer vision module, and a ROS topic named "CvDr" was created. Information about the number of YOLOv4-tiny algorithm detection frames, the number of algorithm-derived frames, and the deviation between the static and motion models was published for rosbag subscription and recording. A starfish model, sealed in a sealed bag, was placed in a pool as a detection target. A remotely controlled vehicle was positioned 40 centimeters in front of the target and remotely controlled. The remotely controlled vehicle dived to a depth of 0.1 meters. Only when the target was about to stray out of the camera's field of view was the remotely controlled yaw angle adjusted to bring the target back to the center. The detection frame was defined as light blue if the target in this frame was detected by the YOLOv4-tiny algorithm; black if the target was derived by the algorithm.

[0087] If a target is detected in one frame of underwater image, but no target is detected in the next frame, the detection missed frame compensation algorithm of kinematic inversion is triggered, and the target position is estimated based on the static model and the motion model for the frames without target within the maximum estimated frame number, until the target is detected again in a certain frame of underwater image. The target position output by the underwater image is taken as the true value, and the difference is made with the estimated results of the static model and the motion model, and the absolute value is taken as the estimated deviation of the two models. If the estimated frame number exceeds the maximum estimated frame number, it is determined that the target is lost, and the model estimation deviation is not calculated. Figure 7 As shown, Figure 7 The detection results of the YOLOv4-tiny algorithm for a frame of underwater image and the inference results of the detection missing frame compensation algorithm for a frame of underwater image through kinematic inversion are shown. Figure 7 It can be seen that there is basically no difference between the two detection results; use MATLAB to draw the lateral position estimation error corresponding to the detection missing frame compensation algorithm of the inversion of the static model and the detection missing frame compensation algorithm of the inversion of the motion model during the experiment. Figure 8 As shown, MATLAB is used to draw the longitudinal position estimation errors corresponding to the detection missing frame compensation algorithm of the inversion of the static model and the detection missing frame compensation algorithm of the inversion of the motion model during the experiment. Figure 9 As shown in the figure, a motion model of an underwater vehicle is established using the detection missing frame compensation algorithm of kinematic inversion to estimate the target center position. For comparison, a static model (assuming that the vehicle is not moving, which is the simplest motion model) is established and estimated. Figure 8 and Figure 9 As can be seen from the figure, the error curve of the static model almost covers the error curve of the motion model, indicating that the motion model established by the present invention can better describe the motion of the underwater vehicle and improve the accuracy of missed frame compensation.

[0088] A performance comparison of the static and motion models is shown in Table 1. Compared to the static model, the motion model-based extrapolation algorithm reduces the average horizontal error by approximately 32% and the average horizontal error by approximately 6%. Approximately 40% of the output detection results come from the extrapolation algorithm, significantly filling the detection gaps between "breathing frames."

[0089] Table 1

[0090]

[0091] An experiment was conducted to test the monocular camera target localization algorithm based on target detection. A tape measure and a protractor were used to measure the actual orientation of the target. The orientation output by the algorithm was recorded five times and the average value was calculated as the algorithm output result. The results are shown in Table 2.

[0092] Table 2

[0093]

[0094] Experimental results show that the algorithm's calculation error for lateral deflection angle and longitudinal height difference is less than 10%, and the calculation error for straight-line distance is within 15%.

[0095] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for underwater target detection and positioning using a monocular camera based on motion inversion, characterized in that: include: Collect multiple frames of continuous underwater images; Target detection is performed on underwater images using a target detection algorithm. If a target is detected in a current underwater image, but not in the next underwater image of the current underwater image, and the interval between the previous underwater image and the next underwater image is less than a preset time, a kinematic inversion detection omission frame compensation algorithm is used to infer the target center coordinates of the next underwater image. If the next two underwater images also do not detect a target, and the interval between the current underwater image and the next two underwater images is less than a preset time, a kinematic inversion detection omission frame compensation algorithm is used to infer the target center coordinates of the next two underwater images until the target is detected in the underwater image. The target detection result includes a target label, a target confidence, a central horizontal coordinate and a central vertical coordinate of the target, and a pixel width and pixel height of the target in the image. The actual physical width and actual physical height of the target are obtained based on the pixel width and pixel height of the target in the image, the change in the vertical coordinate of the target center in two adjacent frames, and the change in the depth of the vehicle. Obtain a lateral deviation angle of the target center relative to the camera center according to the central abscissa coordinate of the target in each frame of underwater image, the central abscissa coordinate and lateral size of each frame of underwater image, and the lateral field of view angle of the camera; Obtain the straight-line distance from the center of the target to the center of the camera according to the lateral size of the underwater image of the target frame, the width of the target in the image, the actual physical width of the target, and the lateral field of view of the camera; Obtain the vertical height between the center of the target and the center of the camera according to the vertical size of the underwater image of the target frame, the vertical field of view of the camera, the vertical coordinate of the center of the target, the vertical coordinate of the center of the underwater image of the target frame, and the straight-line distance from the center of the target to the center of the camera; The target is located based on the longitudinal height between the target center and the camera center, the straight-line distance from the target center to the camera center, and the lateral deviation angle of the target center relative to the camera center.

2. The method for underwater target detection and positioning using a monocular camera based on motion inversion according to claim 1, characterized in that: The steps for calculating the target center coordinates of the next underwater image frame using the detection missed frame compensation algorithm based on kinematic inversion are as follows: Establish a ground coordinate system and a carrier coordinate system: the ground coordinate system takes the initial buoyancy center of the vehicle as its origin, the x-axis of the ground coordinate system points to due north, the y-axis of the ground coordinate system is perpendicular to the x-axis and points to due east, and the z-axis of the ground coordinate system is perpendicular to the x-axis and y-axis to form a right-handed system; the carrier coordinate system takes the buoyancy center of the vehicle as its origin, the x-axis of the carrier coordinate system is along the length direction of the underwater vehicle and points forward, the y-axis of the carrier coordinate system is perpendicular to the x-axis and points to starboard, and the z-axis of the carrier coordinate system is perpendicular to the x-axis and y-axis to form a right-handed system; The target center coordinates of the next underwater image frame are calculated based on the pixels corresponding to the horizontal unit angle and vertical unit angle of the current underwater image frame, the center horizontal coordinate and center vertical coordinate of the target in the current underwater image frame, the roll angle, pitch angle, navigation angle, depth of the vehicle when collecting two adjacent underwater image frames, and the distance from the center of the target in the current frame to the vehicle.

3. The method for underwater target detection and positioning using a monocular camera based on motion inversion according to claim 2, wherein: The expression of the target center coordinates of the next frame of underwater image is: Where, pD x Indicates the corresponding pixel per horizontal unit angle of the underwater image; pD y Represents the corresponding pixel of the vertical unit angle of the underwater image; Indicates the The central vertical coordinate of the target in the frame underwater image; Indicates the The central vertical coordinate of the target in the frame underwater image; Indicates the The central vertical coordinate of the target in the frame underwater image; Indicates the The central horizontal coordinate of the target in the frame underwater image; Indicates collection of Frame underwater image to capture the first The change in the vehicle's roll angle when framing underwater images; Indicates collection of Frame underwater image to capture the first The pitch angle change of the vehicle when framing underwater images; Ψ Indicates collection of Frame underwater image to capture the first The change in the yaw angle of the vehicle when framing underwater images; Indicates collection of Frame underwater image to capture the first The depth change of the vehicle when framing underwater images; Indicates that the aircraft has reached The distance to the center of the target in the underwater image frame.

4. The method for underwater target detection and positioning using a monocular camera based on motion inversion according to claim 1, wherein: Get the actual physical width and actual physical height of the target, including: Where, Indicates the actual physical width of the target; Indicates the actual physical height of the target; Indicates the pixel width of the target in the underwater image; Indicates the pixel height of the target in the underwater image; Indicates collection of Frame underwater image to capture the first The depth change of the vehicle when framing underwater images; Indicates collection of Frame underwater image to capture the first The change in the vertical coordinate of the target center when framing the underwater image; Indicates the physical distance corresponding to the unit pixel.

5. The method for underwater target detection and positioning using a monocular camera based on motion inversion according to claim 1, wherein: Obtaining the lateral deviation angle of the target center relative to the camera center includes: Where, Indicates the lateral deviation angle of the target center relative to the camera center; Represents the central horizontal coordinate of the underwater image in the image coordinate system; Indicates the horizontal coordinate of the target center in the image coordinate system; is the horizontal field of view of the camera; is the horizontal pixel size of the underwater image.

6. The method for underwater target detection and positioning using a monocular camera based on motion inversion according to claim 1, wherein: Obtaining the straight-line distance from the target center to the camera center includes: Where, Indicates the straight-line distance from the center of the target to the center of the camera; is the horizontal field of view of the camera; is the horizontal pixel size of the underwater image; Indicates the actual physical width of the target; Indicates the pixel width of the target in the underwater image.

7. The method for underwater target detection and positioning using a monocular camera based on motion inversion according to claim 1, wherein: Getting the vertical height between the target center and the camera center includes: Where, Indicates the vertical height between the center of the target and the center of the camera lens; Indicates the straight-line distance from the center of the target to the center of the camera; is the vertical field of view of the camera; is the vertical pixel size of the underwater image; y c Represents the central vertical coordinate of the underwater image in the image coordinate system; y t Indicates the vertical coordinate of the target center in the image coordinate system.

8. The method for underwater target detection and positioning using a monocular camera based on motion inversion according to claim 1, wherein: The target detection algorithm uses the YOLOv4-tiny algorithm.

Citation Information

Patent Citations

  • Method for correcting navigational positioning of underwater vehicle based on single-eye vision

    CN109211240A

  • Target ranging method based on monocular vision

    CN111982072A