Zoom binocular vision ship target tracking method
By using YOLOv5 and DeepSort algorithms combined with binocular vision technology in the ship target tracking system, the problem of the acquisition of ship information is susceptible to noise and lack of continuous monitoring in traditional methods, and high-precision suspicious ship identification and tracking is achieved, and detailed ship information and risk analysis is provided.
Patent Information
- Application Number
- CN202510222852.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-20
AI Technical Summary
Traditional ship information acquisition methods have signal interference and noise influences in complex water environments, making it difficult to effectively identify and track suspicious ships, and lack the ability to continuously monitor suspicious ships.
The target detection model based on the YOLOv5 algorithm is used to combine the DeepSort target tracking algorithm, and the identification and tracking of suspicious ships is used to use binocular vision technology. The target is clear in the field of view through gimbal rotation and focal length adjustment, and the ship's size, speed and tonnage are calculated.
It improves the accuracy and robustness of ship target detection, realizes effective identification and continuous monitoring of suspicious ships, and provides three-dimensional information and risk analysis data of ships.
Smart Images

Figure CN120182925A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of binocular stereo vision, and particularly relates to a ship target tracking method for zoom binocular vision. Background Art
[0002] With the rapid development of the shipping industry, the volume of cargo transported on the sea and inland waterways is increasing day by day. The complex water environment and huge traffic volume have led to frequent occurrence of water safety accidents, especially the safety accidents between ships and bridges have the greatest impact. In order to avoid such accidents, a ship-bridge collision warning system has been developed and improved. Among them, the system mainly obtains ship information by means of radar, Automatic Identification System (AIS) of ships, Global Positioning System (GPS), etc. However, these methods have certain limitations. For example, AIS is vulnerable to weather factors, island occlusion, etc., resulting in signal weakening or loss, and the system can only detect the position of ships with correctly equipped and turned-on AIS devices. The ship information detected by radar has poor intuitiveness, and the radar signal may be attenuated by weather. In addition, the cost of radar is relatively high, and it is less economical if it wants to be deployed on a large scale. Moreover, these traditional ship information acquisition methods cannot actively distinguish which are suspicious ships and lack the ability to continuously monitor suspicious ships. Therefore, finding an intuitive and efficient way to discover and track suspicious ships is of great significance for the development and improvement of the ship-bridge collision warning system.
[0003] In recent years, the continuous development of computer technology has made the object detection technology based on deep learning gradually come into the public eye. Object detection is one of the core problems in the field of computer vision. Its main task is to classify and locate objects in images, and it has been widely used in artificial intelligence, autonomous driving, etc. However, although object detection can locate objects in images, it cannot detect the specific position of the object in space. Therefore, it is usually necessary to use monocular vision and binocular vision to identify the object in the image and determine the three-dimensional coordinates of the object in space. Among them, binocular vision has been widely studied with relatively high accuracy. Binocular vision uses two fixed-focus lenses with relatively fixed positions to detect objects and restores the true position of the object in space according to similar triangles. However, due to the limited focal length, the farther the distance, the lower the efficiency of object detection, and the worse the position restoration accuracy. Therefore, in order to more accurately restore the spatial position of objects at different distances, the variable focal length binocular vision technology has emerged and received more and more attention. Summary of the Invention
[0004] Therefore, the present invention provides a ship target tracking method for zoom binocular vision to solve the problems raised in the background art.
[0005] To achieve the above object, the present invention provides the following technical solution: A ship target tracking method for zoom binocular vision, comprising the following steps:
[0006] (1) Use an object detection model based on the YOLOv5 algorithm for ship target detection, use Mosaic data augmentation to expand the dataset, and improve the generalization ability of the model in different scenarios; preprocess the input image during the detection stage to remove noise interference;
[0007] (2) Use the calibration objects arranged on-site to achieve camera calibration and update of calibration data;
[0008] (3) Select the most interesting one as the target from the detected multiple ships, and use the DeepSort object tracking algorithm to track the target;
[0009] (4) Calculate the pan-tilt rotation angle based on the position change amount of the target tracked by the DeepSort object tracking algorithm in the image, and change the focal length based on the proportion of the target in the image, so that the target is always within the field of view and remains clear;
[0010] (5) Calculate the size, speed and tonnage of the target using the binocular ranging principle.
[0011] Preferably, step (1) specifically includes the following content:
[0012] (1.1) In the construction of the dataset, provide pictures of the same type of ship at different angles and under different weather conditions, and use Mosaic augmentation data during the model training stage to improve the detection accuracy and robustness.
[0013] (1.2) In the detection stage, preprocess the images in the object detection model, including grayscale processing and filtering processing, to reduce noise interference in the images.
[0014] Preferably, step (2) specifically includes the following content:
[0015] (2.1) Keep the relative positions of the left and right cameras unchanged, fix the two cameras on an electric pan-tilt that supports secondary development of the SDK, and record the initial position of the pan-tilt;
[0016] (2.2) When the pan-tilt is in the initial position, place calibration objects in the overlapping area of the left and right cameras, synchronously zoom the left and right cameras to each preset focal length, observe the change of the overlapping area, and move the calibration objects so that they are always in the overlapping area of the left and right cameras;
[0017] (2.3) Control the pan-tilt to fine-tune, and take 20 groups of photos of the calibration objects from different angles at the same time;
[0018] (2.4) Use the calibration function in OpenCv to input the image and obtain the internal parameter coefficients of the left and right cameras, as well as the rotation matrix and translation vector describing the pose position.
[0019] Preferably, set a time interval to make the pan-tilt return to the initial position for camera calibration to avoid the position offset between the two cameras during long-term operation.
[0020] Preferably, step (3) specifically includes the following content:
[0021] (3.1) Based on the object detection model described in claim 2, call the binocular camera, and after preprocessing the acquired video stream image, segment it into left and right images;
[0022] (3.2) Use the left camera image for object detection and return the type and position of the target ship;
[0023] (3.2) When the camera perspective is fixed, objects closer in distance are located below in the image coordinates, that is, the closer the target distance, the larger the y coordinate value; therefore, at the initial position, select the target with the largest y coordinate of the target center from the detected multiple ship targets as the target of interest and track it;
[0024] (3.3) After determining the target to be tracked, input the ship image detected by the object detection model into the feature extraction network of the DeepSort object tracking algorithm, extract and save the ship features;
[0025] (3.4) Use the Kalman filter to predict the position of the ship of interest in the next frame of the image;
[0026] (3.5) Match the position of the ship of interest predicted by the Kalman filter in the previous frame with the image features of all detected targets in the current frame, and mark the successfully matched targets.
[0027] Preferably, step (4) specifically includes the following content:
[0028] (4.1) Define a range for the ratio between the target and the image. When the ratio is less than the lower limit of the range, control the left and right cameras to synchronously zoom to the next focal length segment. When the ratio is greater than the upper limit of the range, control the left and right cameras to synchronously zoom to the previous focal length segment;
[0029] (4.2) After the ratio of the target in the image field of view meets the requirements, use the depth map to return the depth of the target center, and calculate the three-dimensional coordinates of the center point in combination with the coordinates (x, y) of the center point in the image;
[0030] (4.3) To achieve continuous tracking of the target, restrictions are imposed on the offset of the target center point coordinates (x, y) relative to the origin of the image coordinate system. Since the origin of the image coordinate system is located at the center of the image, the absolute values of x and y in the current frame are obtained.
[0031] (4.4) When the offsets in the x and y directions exceed the threshold, calculate the flipping angle and pitching angle of the pan-tilt head in the horizontal direction so that the target always appears in the field of view.
[0032] Preferably, step (5) specifically includes the following content:
[0033] (5.1) Since in any frame, the depth value of each point of the target is obtained from the depth map, and further the three-dimensional coordinates of this point are obtained according to similar triangles; therefore, the true size of the target is obtained by the coordinate values on the boundary of the detection box.
[0034] (5.2) According to the frame rate of the video stream, obtain the time interval between adjacent frames. By comparing the change amount of the three-dimensional coordinates of the target center point in adjacent frames, approximately obtain the speed of the target in the three coordinate directions.
[0035] (5.3) Construct a database for the types of ships involved in the data set, store the length, width, and height data of each type of ship in the database. Based on the type obtained from the target detection model and the calculated ship size, match with the sample information in the database. The difference between the true height of the ship obtained by matching and the height of the ship above the water surface detected is the draft. According to the draft, as well as the length and width of the ship, approximately calculate the tonnage of the ship.
[0036] Compared with the existing methods, the present invention has the following advantages:
[0037] The present invention takes into account the defects that the ship-bridge collision warning system is vulnerable to condition noises such as rain, snow, fog, and light during the process of obtaining suspicious ship information, and lacks continuous monitoring of suspicious ships. The YOLOv5 target detection algorithm is used to identify suspicious ships, combined with the binocular stereo vision technology, which improves the detection accuracy and makes up for the defect of the lack of traceability of traditional methods. The beneficial effects are as follows:
[0038] 1. Use the YOLOv5 algorithm for ship target detection. During the training stage of the detection model, use Mosaic to expand the data set, which improves the detection ability of the model for ships under different weather and angles. During the detection stage, perform preprocessing operations such as grayscale and filtering on the input image, which can improve the detection speed to a certain extent and reduce the noise interference in the image.
[0039] 2. Implement the tracking of suspicious ships based on the DeepSort object tracking algorithm. Predict the target features extracted by YOLOv5 in the current frame using the Kalman filter, and achieve the unique coding of the ship by matching with multiple targets detected in the next frame. Combine the pan-tilt rotation to achieve the tracking of the ship target.
[0040] 3. Adopt variable-focus binocular vision, select an appropriate focal length according to the distance of the target, and achieve detection and depth calculation at different distances.
[0041] Using binocular vision, the three-dimensional information of any point of the target can be obtained, and further the speed and tonnage of the suspicious ship can be obtained, providing a basis for estimating the risk of ship-bridge collision. Description of the Drawings
[0042] Figure 1 It is the overall flowchart of a ship target tracking method based on zoom binocular vision provided by the present invention;
[0043] Figure 2 It is the flowchart of using the DeepSort object tracking algorithm to achieve the tracking of ships of interest provided by the present invention. Detailed Embodiments
[0044] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0045] As Figure 1 shown, this embodiment proposes a ship target tracking method based on zoom binocular vision, providing an information acquisition approach for constructing a ship-bridge anti-collision warning system. To evaluate the risk of a ship colliding with a bridge, first, identify the suspicious ship, and then further conduct continuous monitoring, using information such as the ship's speed, size, and tonnage as data support for the research on the risk of ship-bridge collision. Considering problems such as traditional object detection models being easily interfered by noise such as weather, and the short working distance of fixed-focus binocular vision. Therefore, this embodiment proposes a ship target tracking method based on zoom binocular vision. First, train the model on the YOLOv5 network and call the binocular camera to achieve ship target detection; then extract the features of the detected ships and match them with the ships predicted by the Kalman filter to achieve target tracking; subsequently, combine the zoom strategy and the pan-tilt rotation to ensure that the tracked target is always clear, and the obtained ship size, speed, and tonnage information can be used as data for the risk analysis of the ship-bridge anti-collision warning system.
[0046] I. Ship Detection Model Based on YOLOv5 Algorithm
[0047] Use YOLOv5 for ship target detection. During the training stage of the detection model, perform Mosaic augmentation on the data in the dataset to improve the generalization of the model. During the detection stage, preprocess the input image to remove noise interference. The specific contents are as follows:
[0048] (1.1) In dataset construction, provide pictures of the same type of ship under different angles and weather conditions. During the model training stage, use Mosaic to augment the data to improve the accuracy and robustness of detection.
[0049] (1.2) In the detection stage, preprocess the image input into the model, including grayscale processing and filtering, to reduce noise interference in the image.
[0050] Mosaic is a data augmentation technique that performs splicing through methods such as random scaling, cropping, and flipping. This method can greatly enrich the dataset. Especially random scaling adds many small-sized targets, making the model more robust. In dataset construction, collect pictures of common ships under different angles, directions, and weather conditions, and use the Image program to randomly perform transformation operations on the original images, such as flipping, scaling, simulating rain, simulating snow, changing brightness and contrast, etc., to increase image diversity.
[0051] Image preprocessing is a commonly used method in the field of target detection, usually including image binarization, image denoising, etc. The purpose is to reduce interference and enhance contrast before sending the image into the detection model, so as to speed up the detection speed and improve the detection efficiency. Grayscale processing is a method of image preprocessing. Specifically, it is a method of converting a color image into a grayscale image. This method can greatly reduce the amount of data and can speed up the model detection speed. The filtering algorithm is mainly used to smooth the image and eliminate noise interference. Currently, the filtering processing algorithms are divided into two types: linear and non-linear. The median filtering algorithm is a non-linear filtering technology, which is commonly used in the image preprocessing stage. Its purpose is to smooth the image and is very effective in eliminating salt-and-pepper noise (similar to the bright or dark spots in image wave noise). It replaces the pixel value by taking the median of the surrounding neighborhood of the pixel. Compared with other filtering algorithms, it can reduce the degree of blurring of the image edge. Processing the image input into the detection model with the median filtering algorithm can reduce the influence of wave noise in the image and improve the detection accuracy. Therefore, performing grayscale processing and median filtering algorithm processing on the image input into the detection model can improve the detection speed and detection accuracy.
[0052] II. On-site Calibration and Data Update of Binocular Cameras
[0053] (2.1) Keep the relative positions of the left and right cameras unchanged, fix the two cameras on an electric pan-tilt that supports secondary development of the SDK, and record the initial position of the pan-tilt.
[0054] (2.2) When the pan-tilt is in the initial position, place a calibration object within the overlapping area of the left and right cameras, synchronously zoom the left and right cameras to each preset focal length, observe the change in the overlapping area, and move the calibration object so that it is always within the overlapping area of the left and right cameras.
[0055] (2.3) Control the pan-tilt to make fine adjustments and take 20 groups of photos of the calibration object from different angles simultaneously.
[0056] (2.4) Use the calibration function in OpenCv to input the images to obtain the internal parameter coefficients of the left and right cameras, as well as the rotation matrix and translation vector describing the pose position.
[0057] (2.5) To avoid the position offset between the two cameras caused by long-term operation, a certain time interval can be set to make the pan-tilt return to the initial position for camera calibration.
[0058] Camera calibration is the basis for binocular ranging. The calibration of a binocular camera is divided into two parts. First, single-object calibration is performed on each of the left and right cameras. The purpose is to obtain the internal parameters of the left and right cameras, such as the optical center, focal length, distortion coefficient, etc., which can be used for subsequent distortion correction of the cameras. Second, binocular calibration is performed on the two cameras. The purpose is to obtain the rotation matrix and translation vector between the left and right cameras, which can help describe the positional relationship between the cameras. The principle of binocular camera calibration is to solve the rotation matrix and translation vector based on the position difference of an object with known size information in the left and right cameras, which can be used for subsequent row alignment of the pixel points of the same point in the left and right cameras, that is, stereo matching. When the relative positional relationship between the two cameras changes, calibration generally needs to be performed again.
[0059] This embodiment proposes a calibration method for a zoom binocular camera. First, keep the relative positions of the left and right cameras unchanged and fix them on an electric pan-tilt that supports secondary development of the SDK. Specify an initial position, and debug the pan-tilt so that the roll angle and pitch angle are zero in the initial position. Subsequently, simultaneously adjust the preset fixed focal lengths of the binocular cameras, find the common field-of-view overlapping area of the left and right cameras at several focal lengths, and place the calibration object in this area. In the initial position, control the pan-tilt to make fine adjustments so that the left and right cameras synchronously take 20 groups of photos of the calibration object at different focal lengths, import the taken photos into the calibration tool of OpenCV to obtain the internal and external parameters of the cameras at different focal lengths, and then perform distortion correction and stereo matching on the cameras based on the calibration data.
[0060] Considering that the long-term operation of the device will cause the relative position between the left and right cameras to move, a certain time interval is set to return to the initial position for updating the camera calibration data.
[0061] III. Identification and Tracking of Suspicious Vessels
[0062] (3.1) Based on the detection model, call the binocular camera, and preprocess the acquired video stream image and segment it into left and right images;
[0063] (3.2) Use the left camera image for object detection and return the type and position of the target vessel;
[0064] (3.2) When the camera view is fixed, objects closer in distance are located below in the image coordinates, that is, in the area with a larger y coordinate. Therefore, at the initial position, select the target with the largest y coordinate of the target center from the multiple detected vessel targets as the target of interest and track it;
[0065] (3.3) After determining the target to be tracked, input the image of this vessel detected by the YOLOv5 algorithm into the feature extraction network of the DeepSort algorithm, extract the vessel features and save them;
[0066] (3.4) Use the Kalman filter to predict the position of the vessel of interest in the next frame image;
[0067] (3.5) Match the position of the vessel of interest predicted by the Kalman filter in the previous frame with the image features of all detected targets in the current frame, and mark the targets with successful matches.
[0068] Call the binocular camera using the trained model. Preprocess the video stream image of the binocular camera and segment it into left and right camera images. Feeding the left camera image into the detection model can detect multiple vessel targets. When the camera position remains unchanged, the closer the target is, the lower its position in the image. In the case where information such as the speed and tonnage is unknown, rely on the distance of the target from the camera to determine the suspicious vessel, that is, take the target with the largest ordinate of the center point of the detection frame as the suspicious vessel for tracking.
[0069] The Kalman filter is the core of the DeepSort algorithm. It uses the minimum variance for optimal value estimation, predicts the state of the object at the next moment based on its state at the previous moment, and tracks the target through the motion equation and probability estimation. Predict the position of the suspicious vessel at the next moment through the Kalman filter and match this predicted position with the extracted features of the suspicious vessel to achieve the tracking of the suspicious vessel. In addition, the noise processing algorithm in DeepSort can effectively reduce external interference, handle the occlusion problem well, and meet the task requirements of real-time tracking of vessels.
[0070] IV. Ensuring the Clarity of the Tracking Target Based on the PTZ Rotation and Zoom Strategy
[0071] To keep the suspicious ship always clear in the camera's field of view, the method of changing the focal length and rotating the pan-tilt head is adopted to achieve this.
[0072] (4.1) To ensure the measurement accuracy, the proportion of the target in the image is controlled at about 50%. Therefore, a range is delimited for the proportion of the two. When the proportion is less than the lower limit of the range, the left and right cameras are controlled to synchronously zoom to the next focal length segment. When it is greater than the upper limit of the range, the left and right cameras are controlled to synchronously zoom to the previous focal length segment.
[0073] (4.2) After the proportion of the target in the image field of view meets the requirements, the depth of the target center is returned using the depth map, and the three-dimensional coordinates of the center point are calculated by combining the coordinates (x, y) of the center point in the image.
[0074] (4.3) To achieve continuous tracking of the target, restrictions are placed on the offset of the coordinates (x, y) of the target center point relative to the origin of the image coordinate system. Since the origin of the image coordinate system is located at the center of the image, that is, the absolute values of x and y in the current frame are obtained.
[0075] (4.4) When the offsets in the x and y directions exceed the threshold, the pan angle and pitch angle of the pan-tilt head in the horizontal direction are calculated to keep the target always in the field of view.
[0076] The accuracy of binocular ranging is related to the proportion of the target in the image. Therefore, to ensure the subsequent ranging accuracy, the proportion of the target in the image is maintained within a certain range by controlling the camera focal length. Since the larger the focal length, the larger the proportion of the target, and the smaller the focal length, the smaller the proportion of the target. Therefore, when the proportion of the detection frame in the image is less than the lower limit of the proportion range, the focal length is increased, and when it is greater than the upper limit of the proportion range, the focal length is decreased. Since binocular ranging will be carried out in the subsequent steps, the changes of the left and right cameras should be consistent.
[0077] The characteristic of binocular vision is that it can accurately measure the depth information of the target. According to the camera imaging principle, the intersection of the line connecting a point p on the target and the optical center O of the left camera with the left imaging plane is (x1, y1), and the intersection of the line connecting it and the optical center O of the right camera with the right imaging plane is (x l , y r ). Since the left and right imaging planes are coplanar after the left and right cameras are calibrated and stereoscopically corrected, the difference between the intersections of point p with the left and right imaging planes is d = x1 - x r . The distance b between the optical centers of the left and right cameras can be obtained by measurement during equipment installation, and the focal lengths of the left and right cameras are the same and can be obtained from the calibrated internal parameter matrix. According to the principle of similar triangles, it can be known that the distance z r from the target to the camera is z r = f * b / d. Based on this, x c = x1 * z / f, y c = y1 * z / f. c= y1 * z / f. Therefore, the coordinate value of any point of the target can be obtained according to the binocular information.
[0078] To keep the object always within the camera's field of view is mainly achieved by rotating the electric pan-tilt head. After the binocular camera detects a suspicious ship, the three-dimensional coordinates of the center point of the ship detection frame can be calculated. Since the coordinates (x, y) of the ship in the image are known, the offset of the ship relative to the center point can be calculated by combining with the coordinates of the center point of the image. When the offset exceeds the threshold limit, the rotation angle of the pan-tilt head is calculated to align with the target.
[0079] V. Calculation of Ship Dimensions, Speed, and Tonnage
[0080] (5.1) Since in any frame, the depth value of each point of the target can be obtained from the depth map, and the three-dimensional coordinates of this point can be further calculated according to similar triangles. Therefore, the true size of the target can be obtained by the coordinate values on the boundary of the detection frame;
[0081] (5.2) According to the frame rate of the video stream, the time interval between adjacent frames can be known. By comparing the change amount of the three-dimensional coordinates of the center point of the target in adjacent frames, the speed of the target in the three coordinate directions can be approximately obtained;
[0082] (5.3) Build a database for the ship types involved in the data set, store the length, width, and height of each type of ship in the database. Based on the type obtained from the target detection and the calculated ship dimensions, match with the sample information in the database. The difference between the true height of the ship obtained by matching and the height of the ship above the water surface detected is the draft. According to the draft, as well as the length and width of the ship, the tonnage of the ship can be approximately calculated.
[0083] The calculation of ship tonnage depends on the acquisition of the draft. Generally, the method used to detect the draft is to utilize the ship's own sensors. The existing specific method for detecting the draft using images is to further perform image segmentation on the detected ship target, use the edge algorithm to mark the waterline position, and perform depth detection on the hull contour below the waterline, or directly read the scale on the ship's hull using computer vision. However, these methods have high requirements for the image resolution, and it is difficult for the detection speed to meet the real-time requirement.
[0084] This embodiment proposes a new method for detecting the draft of a ship. First, a sample library is built for various types of ships in the data set, and the length, width, and height of each type of ship are stored in the sample library. After detecting a suspicious ship and tracking it, based on binocular vision, the type, length, width, and the height of the ship above the water of the ship can be calculated. Matching the obtained ship type, length, and width information with the sample library can obtain the actual height of this type of ship. Subtracting the height of the ship above the water from the actual height of the ship can obtain the draft of the ship.
[0085] Furthermore, the tonnage of the ship can be calculated according to formula (1):
[0086]
[0087] In formula (1), M represents the displacement of the ship; L represents the length of the ship; W represents the width of the ship; H C represents the draft of the ship; C represents the block coefficient of the ship, which is related to the ship type; ρ is the reciprocal of the liquid density, where 0.9756 is taken for seawater and 1 is taken for fresh water.
[0088] Although the present invention has been described in detail above with general descriptions and specific embodiments, on the basis of the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.
Claims
1. A ship target tracking method using zoom binocular vision, characterized in that: The following steps are involved: (1) A target detection model based on the YOLOv5 algorithm is used to detect ship targets. Mosaic data enhancement is used to expand the dataset to improve the generalization ability of the model in different scenarios. The input image is preprocessed in the detection stage to remove noise interference. (2) Using the calibration objects placed on site to calibrate the camera and update the calibration data; (3) Select the most interesting target from the multiple detected ships and track it using the DeepSort target tracking algorithm; (4) Based on the change in the position of the target tracked by the DeepSort target tracking algorithm in the image, the pan / tilt rotation angle is calculated, and the focal length is changed based on the proportion of the target in the image, so that the target is always within the field of view and remains clear; (5) Use binocular ranging principle to calculate the size, speed and tonnage of the target.
2. The ship target tracking method using zoom binocular vision according to claim 1, characterized in that: Step (1) specifically includes the following contents: (1.1) In the dataset construction, pictures of the same type of ships from different angles and under different weather conditions are provided, and Mosaic is used to enhance the data during the model training stage; (1.2) In the detection stage, the image in the target detection model is preprocessed, including grayscale processing and filtering processing.
3. The ship target tracking method using zoom binocular vision according to claim 1, characterized in that: Step (2) specifically includes the following contents: (2.1) Keep the relative positions of the left and right cameras unchanged, fix the two cameras on an electric gimbal that supports SDK secondary development, and record the initial position of the gimbal; (2.2) When the gimbal is in the initial position, place a calibration object in the overlapping area of the left and right cameras, so that the left and right cameras can be zoomed to each preset focal length synchronously, observe the changes in the overlapping area, and move the calibration object so that it is always in the overlapping area of the left and right cameras; (2.3) Control the gimbal to fine-tune and take photos of the calibration object from different angles at the same time; (2.4) Using the calibration function in OpenCv, the image input is used to obtain the intrinsic coefficients of the left and right cameras as well as the rotation matrix and translation vector that describe the posture position.
4. The ship target tracking method using zoom binocular vision according to claim 3 is characterized in that: Set a time interval to return the gimbal to its initial position for camera calibration to avoid positional offset between the two cameras caused by long-term operation.
5. The ship target tracking method with zoom binocular vision according to claim 3 is characterized by: Step (3) specifically includes the following contents: (3.1) Based on the target detection model described in claim 2, call the binocular camera and pre-process the acquired video stream image and split it into left and right images; (3.2) Use the left camera image to perform target detection and return the target ship type and position; (3.2) When the camera angle of view is fixed, the closer the object is, the larger the y coordinate value in the image coordinate system. Therefore, at the initial position, the target with the largest y coordinate of the center is selected from the multiple detected ship targets as the target of interest and tracked; (3.3) After determining the target to be tracked, the image of the ship detected by the target detection model is input into the feature extraction network of the DeepSort target tracking algorithm to extract and save the ship features; (3.4) Use the Kalman filter to predict the position of the ship of interest in the next frame of image; (3.5) Match the position of the ship of interest predicted by the Kalman filter in the previous frame with the image features of all targets detected in the current frame, and mark the successfully matched targets.
6. The ship target tracking method with zoom binocular vision according to claim 3 is characterized by: Step (4) specifically includes the following contents: (4.1) A range is defined for the ratio of the target to the image. When the ratio is less than the lower limit of the range, the left and right cameras are controlled to zoom to the next focal length synchronously. When the ratio is greater than the upper limit of the range, the left and right cameras are controlled to zoom to the previous focal length synchronously. (4.2) When the proportion of the target in the image field meets the requirements, the depth map is used to return the center depth of the target, and the three-dimensional coordinates of the center point are calculated by combining the coordinates (x, y) of the center point in the image; (4.3) In order to achieve continuous tracking of the target, the offset of the target center point coordinates (x, y) relative to the origin of the image coordinate system is restricted. Since the origin of the image coordinate system is located at the center of the image, the absolute values of x and y in the current frame are calculated; (4.4) When the x- and y-direction offsets exceed the threshold, the horizontal flip angle and pitch angle of the gimbal are calculated so that the target always appears in the field of view.
7. The ship target tracking method using zoom binocular vision according to claim 6 is characterized by: Step (5) specifically includes the following contents: (5.1) In any frame, the depth value of each point of the target is obtained according to the depth map, and the three-dimensional coordinates of the point are further calculated according to similar triangles; therefore, the real size of the target is calculated by the coordinate value on the boundary of the detection box; (5.2) According to the frame rate of the video stream, the time interval between adjacent frames is obtained. By comparing the changes in the three-dimensional coordinates of the target center point in adjacent frames, the speed of the target in the three coordinate directions can be approximately calculated; (5.3) A database is constructed for the types of ships involved in the data set. The length, width, and height data of each type of ship are stored in the database. The type obtained based on the target detection model and the calculated ship size are matched with the sample information in the database. The difference between the true height of the matched ship and the detected height of the ship above the water surface is the draft. The tonnage of the ship is approximately calculated based on the draft and the length and width of the ship.
Citation Information
Patent Citations
Safety helmet wearing detection method and system in construction scene
CN115171022A
Binocular camera parallax correction method suitable for mowing robot
CN116996657A