Water surface multi-target detection and automatic labeling method based on threshold segmentation

Through the multi-objective detection and automatic labeling method based on threshold segmentation and masking technology, the problems of traditional manual labeling are solved, and the surface target monitoring with high accuracy and robustness are achieved, and high-quality data sets are generated.

CN120163969AActive Publication Date: 2025-06-17SUZHOU PERI LINGZHEN TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510257957.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-17
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

Traditional manual data annotation methods are inefficient and unstable in accuracy, making it difficult to adapt to real-time changes in dynamic water surface environments.

Method used

The water surface multi-object detection and automatic labeling method based on threshold segmentation is adopted. By acquiring and pre-treating the visible light monitoring video of the water surface waterway, threshold segmentation and connection domain analysis are performed, background suppression and target extraction are achieved in combination with mask diagram technology, and pre-trained SVM model is used for automatic labeling.

Benefits of technology

It significantly improves the accuracy and robustness of surface target monitoring, reduces the workload and time of manual labeling, can monitor and evaluate water traffic conditions in real time, and generate high-quality surface target sample data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163969A_ABST
    Figure CN120163969A_ABST
Patent Text Reader

Abstract

The invention relates to a water surface multi-target detection and automatic labeling method based on threshold segmentation, belongs to the field of water surface target detection and deep learning, and solves the problems of low efficiency and unstable precision of a traditional manual data labeling method. Comprising the steps that a water surface channel visible light monitoring video is acquired and preprocessed, and a multi-frame binary image with a water surface target and a background separated is obtained; performing connected domain analysis on each frame of binarized image to obtain a plurality of target mark box sets; cyclically combining and screening the plurality of target mark frames to obtain a mark frame of each real water surface target; performing target tracking prediction on the real water surface target marking frame in each frame to obtain predicted water surface target position coordinates and size information; and performing water surface target image cutting, water surface target position coordinate normalization conversion and automatic labeling on the basis of the water surface target position coordinate and size information to obtain a labeled data set. Water surface target sample data can be automatically obtained, and labels can be automatically labeled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of surface target detection and deep learning, and particularly to a method for multi-target detection and automatic annotation of water surfaces based on threshold segmentation. Background Art

[0002] With the continuous increase in the number of navigable vessels, the complexity and potential risks of water transportation are gradually rising. The problem of ship-bridge collisions has gradually become a serious safety hazard, especially in densely navigable areas such as urban rivers and lakes. To effectively prevent the occurrence of ship-bridge collision accidents, there is an urgent need for more intelligent monitoring means to real-time monitor and evaluate the water transportation situation. This need has promoted the research on new monitoring systems and intelligent algorithms, aiming to improve the accuracy and efficiency of monitoring.

[0003] In the field of deep learning, the production of data sets is crucial for model training. Traditionally, many application scenarios rely on the method of manually generating data sets. However, this method has various drawbacks. Firstly, manual annotation usually requires a large amount of time and human resources. Especially in complex scenarios, the accuracy and consistency of annotation are often difficult to guarantee. Secondly, manual annotation is easily affected by subjective factors, resulting in uneven quality of data sets, thus affecting the performance of the model. In addition, in a dynamic environment, such as the changes and unpredictable movements of surface targets, manual annotation is difficult to adapt to real-time changes, reducing the timeliness and applicability of data sets.

[0004] In this context, computer-intelligent data generation methods have emerged as an important means to solve the above problems. Existing computer-aided data generation methods mainly include synthetic data generation, automatic annotation, and semi-supervised learning. These methods utilize advanced computer vision and deep learning technologies, which can greatly improve the efficiency and accuracy of data generation. For example, synthetic data generation can simulate the environment and objects to achieve a large number of high-quality training samples. Automatic annotation technology uses existing models for preliminary annotation, reducing manual intervention, thus accelerating the process of data set production.

[0005] These intelligent data generation methods still have certain limitations in practical applications. In the field of image object detection, common automatic labeling methods have their own advantages and disadvantages. Automatic labeling based on weak supervision uses image-level labels to train a model to generate pseudo bounding box annotations, which can reduce the workload of manual labeling. However, the positioning accuracy may be limited, and the accuracy of the generated pseudo annotations needs to be improved. Semi-supervised learning automatic labeling combines a small amount of manual labeling with a large amount of unlabeled data, which can effectively utilize data resources to improve the model performance. However, the quality of pseudo labels depends on the initial model. If the initial model performance is poor, it may lead to error propagation. Automatic labeling based on active learning helps to focus on key samples to reduce the number of labels and improve the labeling efficiency. However, it relies on manual labeling of the samples selected by the algorithm and is not completely automated. Automatic labeling based on the generative adversarial network GAN (Generative Adversarial Network) can generate pseudo samples and annotations for auxiliary training. However, its training process is complex, and the generated pseudo annotations may deviate from the real data. The method of combining crowdsourcing and automatic screening can quickly obtain labeled data with the help of a large number of annotators and ensure a certain quality through automatic screening. However, the quality of crowdsourcing labels varies, and the automatic screening algorithm may also make misjudgments, and manual review and correction are still required to ensure the reliability of the labels. Summary of the Invention

[0006] In view of the above analysis, embodiments of the present invention aim to provide a method for multi-object detection and automatic annotation of water surfaces based on threshold segmentation to solve the technical problems of low efficiency and unstable accuracy existing in traditional manual data annotation methods.

[0007] The object of the present invention is mainly achieved through the following technical solutions:

[0008] The present invention provides a method for multi-object detection and automatic annotation of water surfaces based on threshold segmentation, including:

[0009] Obtain the visible light monitoring video of the water surface channel and perform preprocessing to obtain multiple binary images with the water surface targets separated from the background;

[0010] Perform connected component analysis on each binary image to obtain multiple sets of target bounding boxes; circularly merge and screen the multiple target bounding boxes to obtain the bounding boxes of each real water surface target; perform target tracking and prediction on the real water surface target bounding boxes in each frame to obtain the predicted position coordinates and size information of the water surface targets;

[0011] Based on the position coordinates and size information of the water surface targets, perform water surface target image cropping, normalization conversion of the water surface target position coordinates, and automatic annotation labeling to obtain a labeled water surface target data set.

[0012] Further, the preprocessing of the visible light monitoring video of the water surface channel includes:

[0013] Calculate the total number of video frames based on the total duration and frame rate of the visible light monitoring video of the water surface channel;

[0014] Based on the set frame extraction interval, extract frames from the visible light monitoring video of the water surface channel to obtain corresponding multiple frames of color RGB images;

[0015] Convert each frame of the color RGB image into a grayscale image;

[0016] Perform color inversion processing on each frame of the grayscale image to obtain the grayscale image after color inversion processing;

[0017] Based on the binary mask map designed for perspective optimization, perform threshold segmentation on each frame of the grayscale image after color inversion processing to obtain multiple binary images with the water surface target separated from the background.

[0018] Furthermore, designing a binary mask map based on perspective optimization includes:

[0019] Based on the installation position and perspective of the bridge camera that captured the visible light monitoring video, calibrate the water surface channel range;

[0020] Create a binary matrix mask map with the same size as the color RGB image, assign the pixel points within the water surface channel range the value of 1, and assign the pixel points outside the water surface channel range the value of 0.

[0021] Furthermore, automatically annotating labels for the water surface target images includes:

[0022] Extract HOG features from a small number of water surface target images to obtain corresponding feature vectors, and corresponding manually added water surface target labels to form a pre-training sample set;

[0023] Use the pre-training sample set to train the SVM model;

[0024] When the detection accuracy and recall rate of the SVM model meet the requirements, obtain the pre-trained SVM model;

[0025] Input the feature vectors obtained by extracting HOG features from the unlabeled water surface target images into the pre-trained SVM model to obtain the detection results of the corresponding water surface target images as the class labels of the water surface target images;

[0026] Generate YOLO-format annotation labels from the class labels of the water surface target images, the coordinates and size information of the water surface target images, and the frame numbers of the water surface target images.

[0027] Furthermore, performing threshold segmentation on the grayscale image after color inversion processing includes:

[0028] Set the target background segmentation threshold T;

[0029] When the gray value in the inverted gray image is higher than T and the element in the corresponding binary mask image is 0, it is marked as 1, indicating the detected water surface target; otherwise, it is marked as 0, indicating the detected background. The segmentation process of the water surface target and the background is as follows:

[0030]

[0031] Among them, I binary (x, y) is the binary image obtained by separating the water surface target and the background after threshold segmentation, and M(x, y) is the binary mask image.

[0032] Furthermore, perform connected component analysis on the binary image to obtain a set of multiple water surface target bounding boxes, including:

[0033] Perform connected component analysis on the binary image, extract all target connected regions, and obtain the minimum bounding rectangle BBox of the target region. For each connected region C i , it is expressed as follows:

[0034]

[0035] Among them, N is the number of detected connected components, (x min , y min ), (x max , y max ) are the upper left and lower right coordinates of BBox respectively;

[0036] Based on C i , calculate the area S i of BBox; if S i is less than the preset area threshold S th , it is determined as noise or non-water surface target and excluded;

[0037] After the exclusion process, a set of multiple water surface target bounding boxes is obtained.

[0038] Furthermore, perform cyclic merging and screening on multiple target bounding boxes to obtain the bounding box of each real water surface target, including:

[0039] Calculate the intersection and union areas of any two target bounding boxes B a and B b in the set of multiple water surface target bounding boxes, and obtain the intersection over union IoU(B a , B b ) of any two bounding boxes;

[0040] If the intersection over union of two target bounding boxes is greater than the intersection over union threshold IoUthreshold , it is determined that the two target bounding boxes belong to the same target, and the smaller one of the two target bounding boxes is merged into the larger bounding box; multiple merged BBoxes are obtained;

[0041] After the merging is completed, if the area of the merged BBox is less than the preset area threshold S m , it is determined as a noise or interference area and removed, and the bounding box BBox of each real water surface target is obtained.

[0042] Furthermore, target tracking prediction is performed on the bounding box of the real water surface target to obtain the predicted position coordinates and size information of the water surface target, including:

[0043] The set of real water surface target boxes in the current frame t is B t , The set of target boxes in the previous frame t - 1 is B t-1,

[0044] Calculate the Euclidean distance d(B t i , B j t-1 ) between each box in the current frame and the previous frame;

[0045] If meets the distance threshold d th , select the smallest B t i as the matching object;

[0046] Based on the Euclidean distance and the time interval between two frames, predict the next frame speed and acceleration of the water surface target corresponding to the matching object; calculate the speed and acceleration in the horizontal and vertical directions based on the speed and acceleration;

[0047] Based on the speed and acceleration of the water surface target in the horizontal and vertical directions, predict the position of the water surface target in the next frame t + 1 and the estimated results of the width and height dimensions of the target box BBox.

[0048] Furthermore, based on the position coordinates and size information of the water surface target, cut out the water surface target image from the water surface image with an area larger than the preset area S;

[0049] The label annotation of each water surface target image meets the YOLO format, and the label of the water surface target image is <frame number> <class><x center ><y center > <width> <height>;

[0050] Among them, the frame number is the frame sequence number corresponding to the water surface target image; class is the category label of the water surface icon image; (x center , y center ) is the center point coordinate of the bounding box BBox of the water surface target; width and height are the normalized width and height of the BBox respectively.

[0051] Furthermore, the water surface target dataset includes multiple frames of color RGB images obtained by frame extraction and the water surface target images, as well as corresponding category labels.

[0052] Compared with the prior art, the present invention can at least achieve one of the following beneficial effects:

[0053] 1. The water surface multi-target detection and annotation method disclosed by the present invention realizes effective background suppression and target extraction by combining threshold segmentation and mask map technology, significantly improving the accuracy and robustness of water surface target monitoring, and is particularly suitable for dynamically changing water surface environments;

[0054] 2. The present invention automatically annotates the labels of water surface targets through a pre-trained SVM model, greatly reducing the workload and time of manual annotation and improving the annotation efficiency;

[0055] 3. The present invention can monitor and evaluate the water traffic conditions in real time, adapt to different environmental and target changes, and has stronger adaptability and practicability;

[0056] 4. The present invention can generate a high-quality water surface target sample dataset, avoiding errors and omissions that may occur in manual annotation, improving the integrity and accuracy of the water surface sample dataset, providing a more reliable data support basis for subsequent model training and analysis; and providing effective technical support for the intelligent development of the bridge anti-collision warning system.

[0057] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combination schemes. Other features and advantages of the present invention will be described in the subsequent specification, and some advantages can be made obvious from the specification or understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the content specifically pointed out in the specification and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] The drawings are only for the purpose of showing specific embodiments and are not considered to be a limitation of the present invention. Throughout the drawings, the same reference signs denote the same components.

[0059] Figure 1 It is a flowchart of a water surface multi-target detection and automatic annotation method based on threshold segmentation in an embodiment of the present invention;

[0060] Figure 2 It is a schematic diagram of the visible light monitoring video of the waterway on the water surface captured by the monitoring equipment deployed on a certain bridge in the embodiment of the present invention;

[0061] Figure 3 It is a schematic diagram of the frame extraction result when the video interval k takes a value of 4 seconds in the embodiment of the present invention;

[0062] Figure 4 It is a schematic diagram of converting a color RGB image of a certain bridge into a grayscale image in the embodiment of the present invention;

[0063] Figure 5 It is a schematic diagram of the mask map design of the monitoring screen of a certain bridge in the embodiment of the present invention;

[0064] Figure 6 It is a schematic diagram of the result of the image after inverse color processing in the embodiment of the present invention;

[0065] Figure 7 It is a schematic diagram of the binary image after threshold segmentation in the embodiment of the present invention;

[0066] Figure 8 It is a schematic diagram of the result of obtaining the minimum circumscribed rectangle of the target from the binary image of the monitoring screen of a certain bridge in the embodiment of the present invention;

[0067] Figure 9 It is a schematic diagram of the result of circularly merging and deleting BBox in the embodiment of the present invention;

[0068] Figure 10 It is a schematic diagram of the representation of BBox in the embodiment of the present invention. Detailed implementation manners

[0069] Next, the preferred embodiments of the present invention will be specifically described in conjunction with the accompanying drawings. Among them, the accompanying drawings form a part of this application and are used together with the embodiments of the present invention to explain the principle of the present invention, rather than to limit the scope of the present invention.

[0070] To solve the above problems, the present invention proposes a method for multi-target detection and tracking auxiliary annotation on the water surface based on threshold segmentation. This method takes the monitoring video captured by the monitoring equipment deployed on the bridge for the waterway on the water surface as input, and uses the method in the present invention to realize the automatic annotation of water surface target samples, no longer requiring manual frame-by-frame target calibration and label generation, and generating a high-quality water surface target data set.

[0071] With the increasing demand for water traffic safety, adopting an intelligent data generation method is the key to improving the effectiveness and safety of the monitoring system. The method of multi-target detection and automatic annotation of the water surface based on threshold segmentation not only provides a new idea for solving the ship-bridge conflict problem, but also provides reference for the production of water surface target datasets in related fields, promoting the progress of intelligent monitoring technology.

[0072] Exemplarily, this application uses the visible light video monitoring of multiple navigable bridges in Guangdong Province. By combining image processing technology and target detection and tracking algorithms, it makes an intelligent water surface target dataset, which can not only improve the sample annotation efficiency, but also ensure the accuracy of sample annotation, meeting the high-quality requirements of the deep learning model for the water surface target dataset.

[0073] The method for multi-target detection and tracking-assisted automatic annotation of the water surface disclosed by the present invention first slices the monitoring video, uses threshold segmentation technology to preliminarily process the water surface, and extracts the candidate areas of the target ships. Then, it uses the target detection algorithm to accurately identify the candidate areas, obtains the position and motion trajectory information of the water surface target ships, and uses the water surface target tracking algorithm to realize the tracking of multiple targets. Through coordinate conversion and automatic label generation, the detected water surface target information is converted into annotation data in the format conforming to the YOLO model. Finally, the constructed annotation data is combined with the original monitoring video to form a high-quality water surface target dataset, which can not only effectively support the training of the subsequent deep learning model, but also significantly shorten the time for dataset production.

[0074] A specific embodiment of the present invention discloses a method for multi-target detection and automatic annotation of the water surface based on threshold segmentation, as Figure 1 shown, including the following steps:

[0075] Step S1: Obtain the visible light monitoring video of the water surface channel and perform preprocessing to obtain multiple frames of binary images with the water surface targets separated from the background;

[0076] Step S2: Perform connected component analysis on each frame of binary image to obtain multiple sets of target bounding boxes; circularly merge and screen the multiple target bounding boxes to obtain the bounding boxes of each real water surface target; perform target tracking prediction on the real water surface target bounding boxes in each frame to obtain the predicted position coordinates and size information of the water surface targets;

[0077] Step S3: Based on the water surface target position coordinates and size information, perform water surface target image cropping, water surface target position coordinate normalization conversion, and perform automatic annotation labels to obtain the annotated water surface target dataset.

[0078] Step S1 is divided into steps S11 - S12.

[0079] Step S11: Obtain the visible light monitoring video of the water surface channel.

[0080] Based on the monitoring devices deployed on the bridge, obtain the visible light monitoring video of the water surface channel. Exemplarily, the present invention selects a real bridge monitoring video as the data source, which is the water surface monitoring video of a certain bridge from July to September 2024.

[0081] As Figure 2 shown, the visible light video of the water surface channel monitored on June 4, 2024 of a certain bridge has rich water surface ship targets and a complex environmental background, which can provide strong support for the testing, verification and dataset generation of the method of the present invention.

[0082] The selection of video data fully considers the diversity and representativeness in the actual application scenario, especially the complex factors common in water surface target detection, such as light changes, water surface fluctuations, debris interference, etc.

[0083] Step S2: Preprocess the obtained visible light monitoring video of the water surface channel to obtain multiple frames of binary images with the water surface targets separated from the background.

[0084] Preprocess the visible light monitoring video of the water surface channel, including:

[0085] Calculate the total number of frames of the video based on the total duration and frequency of the visible light monitoring video of the water surface channel;

[0086] Based on the set frame extraction interval, extract frames from the visible light monitoring video of the water surface channel to obtain corresponding multiple frames of color RGB images;

[0087] Convert each frame of the color RGB image into a grayscale image;

[0088] Perform color inversion processing on each frame of the grayscale image to obtain the grayscale image after color inversion processing;

[0089] Based on the binary mask map designed for perspective optimization, perform threshold segmentation on each frame of the grayscale image after color inversion processing to obtain multiple frames of binary images with the water surface targets separated from the background.

[0090] Preprocess the visible light monitoring video of the water surface channel to ensure the efficiency and accuracy of the subsequent water surface target detection process. Specifically as follows:

[0091] (1) Calculate the total number of frames of each visible light monitoring video of the water surface channel.

[0092] In the tasks of image target detection and dataset generation for visible light monitoring videos of water surface channels, the original monitoring video data is converted into single-frame image data. To ensure the effectiveness of subsequent image processing, the number of frames in the video is first accurately calculated. For example, if the total duration of the monitoring video is T (in seconds), then the total number of frames N of the visible light monitoring video of the water surface channel is calculated as follows:

[0093] N = T × f Equation (1)

[0094] where f is the frame rate of the video, in frames per second.

[0095] (2) Based on the total number of frames of the video, set the frame extraction interval k, and extract frames from the visible light monitoring video of the water surface channel to obtain corresponding multiple frames of color RGB images.

[0096] When processing the visible light monitoring video of the water surface channel, to ensure the continuity of the image sequence and avoid overfitting, set an appropriate frame extraction interval k to improve the processing efficiency while ensuring information integrity.

[0097] The purpose of frame extraction is to reduce redundant information between adjacent frames. Especially when the position of the water surface target object (such as a ship) changes little and the difference is not significant between adjacent frames, frame-by-frame extraction will lead to redundant calculations and overfitting of the model.

[0098] Frame extraction from the monitoring video results in an image format that is conducive to processing. Exemplarily, such as png and jpg image formats, with a size of 2048*2048.

[0099] To ensure the continuity of the image and avoid overfitting when processing the monitoring video, based on specific requirements, set an appropriate frame extraction interval k. The selection of the frame extraction interval k is flexibly adjusted according to training requirements and the characteristics of target motion, especially considering the motion speed of the water surface target object. Through this interval, the required number of frames M can be effectively extracted from the total number of frames, as follows:

[0100]

[0101] In practical applications, the selection of k is adjusted according to training requirements and the target motion speed to balance information integrity and processing efficiency. Exemplarily, assume that the frame rate of the video is 25 frames per second, and the ship moves relatively slowly in the video. It may take about 5 minutes for the ship to enter the field of view and completely disappear. If images are extracted frame by frame, it may result in a large amount of redundant information and increase the computational burden. Exemplarily, k is taken as 4 seconds; setting an appropriate frame extraction interval can effectively reduce the number of frames to be processed while ensuring that each frame of the image has sufficient dynamic change information.

[0102] After frame extraction, the obtained multi-frame color RGB image sequence is subjected to subsequent image analysis and target detection. As Figure 3 shown, the image is a surveillance footage of a video with an interval k of 4 seconds in the same period. It can be clearly seen that the movement of the ship in the lower right corner has changed. This not only ensures the coherence of the image sequence but also improves the algorithm processing efficiency by reducing the extraction of redundant frames, avoiding overfitting during the training process.

[0103] After frame extraction, multiple frames of color RGB images corresponding to each visible light surveillance video of the water surface channel are obtained.

[0104] (3) Convert each frame of the color RGB image to a grayscale image.

[0105] The purpose of the grayscale process is to convert the color information in the color RGB image into a single brightness value for subsequent image analysis and target detection. The grayscale process is completed through the following weighted formula:

[0106] I gray (x,y) = 0.2989·R(x,y) + 0.5870·G(x,y) + 0.1140·B(x,y) Formula (3)

[0107] where (x,y) is the point coordinate of the pixel in the image; I gray (x,y) is the grayscale value at point (x,y) in the 1-channel grayscale image; R(x,y) is the grayscale value of point (x,y) in the R channel, G(x,y) is the grayscale value of point (x,y) in the G channel, and B(x,y) is the grayscale value of point (x,y) in the B channel.

[0108] By converting the color RGB image obtained by frame extraction into a grayscale image through this grayscale process, the redundant color information in the original color RGB image is compressed into brightness information, significantly reducing the computational complexity. At the same time, it helps to improve the efficiency and robustness of the target detection algorithm. The grayscale image is not only visually more concise but also can effectively retain the shape and contour features of the target, facilitating subsequent connected component analysis and water surface target detection. As Figure 4 shown, it is the grayscale image obtained after image grayscale conversion of a certain bridge surveillance image.

[0109] (4) Design a binary mask image based on perspective optimization.

[0110] To achieve accurate water surface target detection and avoid interference from background objects such as docks and shores on the water surface target detection results, this application proposes an image region screening method based on the mask technology. Design a binary mask image (mask) to highlight the area of interest in the image while suppressing the influence of other irrelevant areas.

[0111] Design principle of the binary mask image: The area of interest is represented by black (gray value is 0) in the mask image, while the non - area of interest is represented by white (gray value is 1).

[0112] Optimize the design of the binary mask image based on the perspective, including:

[0113] Calibrate the water surface channel range based on the installation position and perspective of the bridge camera that shoots the visible - light monitoring video;

[0114] Create a binary matrix mask image with the same size as the color RGB image, assign the pixel points within the water surface channel range to 1, and assign the pixel points outside the water surface channel range to 0.

[0115] Identify and extract the channel area in the grayscale image, and mark it as 1 in the mask image to ensure that this area is fully concerned in the subsequent processing. The dock, shore, and other non - target areas are marked as 0, thus effectively suppressing these interference areas from the image.

[0116] Through this mask design, the non - target areas of the image are effectively excluded, thereby improving the accuracy and robustness of the subsequent water surface target detection algorithm. This method can significantly reduce background interference and improve the accuracy of water surface target extraction.

[0117] Based on the monitoring equipment deployed on the bridge, such as the installation position and perspective of the monitoring camera, manually calibrate the water surface channel range (for example, the channel of a certain bridge is a curved river area).

[0118] In the calibration software (such as LabelMe), use the polygon tool to accurately draw the channel boundary and generate a vector coordinate file.

[0119] The binary mask image M is represented as follows:

[0120]

[0121] Create a binary matrix with the same size as the grayscale image according to the channel vector coordinates:

[0122] Pixel points within the channel: Assign 1; Pixel points outside the channel: Assign 0. As Figure 5 Shown is the monitoring screen and mask image design of a certain bridge, where the positions of the dock, riverbank, and non - ship channel are set as non - areas of interest to prevent interference when detecting water surface targets.

[0123] The design of the custom - made binary mask image enables subsequent processing to only focus on the area identified by the mask, optimizing the operation efficiency of the algorithm, and can be adjusted according to the actual situations of different bridges or monitoring perspectives, ensuring the adaptability and accuracy of the mask image.

[0124] (5) Perform inverse color processing on each frame of the grayscale image to obtain the grayscale image after inverse color processing.

[0125] During the target detection process, since the grayscale value of a water surface ship is usually lower than that of the background (such as the river surface), in order to facilitate subsequent binarization processing, an inverse color processing method is adopted. The main purpose of inverse color processing is to invert the grayscale values in the image, so that the grayscale value of the water surface ship target changes from low to high, and the grayscale value of the background area changes from high to low, thereby enhancing the contrast between the water surface target and the background and laying a foundation for subsequent threshold segmentation operations. For each pixel point I gray (x, y) of the grayscale value, the inverse color processing calculation is as follows:

[0126] I invert (x, y) = 255 - I gray (x, y) Formula (5)

[0127] where, I gray (x, y) is the grayscale value at the position (x, y) in the original grayscale image, and I invert (x, y) is the grayscale value after inverse color processing, and 255 is the maximum value of the grayscale image (i.e., the grayscale value of white).

[0128] As Figure 6 shown, after inverse color processing, the grayscale value of the water surface ship target will become relatively high, and the grayscale value of the background area will be relatively low.

[0129] (6) Perform threshold segmentation on each frame of the grayscale image after inverse color processing to obtain multiple frames of binarized images in which the water surface target and the background are separated.

[0130] Performing threshold segmentation on the grayscale image after inverse color processing includes:

[0131] Set the target-background segmentation threshold T;

[0132] When the grayscale value in the grayscale image after inverse color processing is higher than T and the element in the corresponding binary mask image is 0, it is marked as 1, indicating the detected water surface target; otherwise, it is marked as 0, indicating the detected background. The segmentation process of the water surface target and the background is as follows:

[0133]

[0134] where, I binary (x, y) is the binarized image in which the water surface target and the background are separated after threshold segmentation, and M(x, y) is the binary mask image.

[0135] The image is divided into a target region and a background region using threshold segmentation technology. A suitable target-background segmentation threshold T is set. The regions in the image with gray values higher than T are marked as ship surface targets, while the regions with gray values lower than T are marked as the background. Exemplarily, for an 8-bit image (gray value range 0 - 255), the target-background threshold T is set to 120.

[0136] A value of 0 indicates that the background has been effectively suppressed. Threshold segmentation is performed on the grayscale image to separate the water surface target from the background, suppressing the interfering background while retaining the water surface target information, thereby improving the detection accuracy.

[0137] As Figure 7 shown, after background suppression and threshold segmentation, an image with the background suppressed is obtained. At this time, the entire background is black (gray value 0), while the target is a white region (gray value 255).

[0138] The function of step S1 is to separate the water surface target from the background through preprocessing of the visual monitoring video of the water surface channel, generating multiple frames of binary images, providing a basis for subsequent target detection and tracking.

[0139] Step S2 is divided into steps S21 - S23.

[0140] Based on the information of the multiple frames of binary images obtained in step S1, the water surface target is detected and tracked.

[0141] In step S21, connected component analysis is performed on each frame of the binary image to obtain multiple sets of target bounding boxes.

[0142] Performing connected component analysis on the binary image to obtain multiple sets of water surface target bounding boxes, including:

[0143] Performing connected component analysis on the binary image, extracting all target connected regions, obtaining the minimum bounding rectangle BBox of the target region. For each connected region C i , it is expressed as follows:

[0144]

[0145] where N is the number of detected connected components, (x min , y min ), (x max , y max ) are the upper left and lower right coordinates of the BBox respectively;

[0146] Based on C i , the area S i of the BBox is calculated; if S i is less than the preset area threshold S th , it is determined as noise or non-water surface target and excluded;

[0147] After the exclusion process, multiple sets of water surface target bounding boxes are obtained.

[0148] Binary image I binary (x, y), the target area is composed of connected white pixels. For I binary (x, y), perform connected component analysis (Connected Components Labeling, CCL), extract all target areas, and calculate the minimum bounding rectangle BBox.

[0149] For each connected region C i , its BBox is determined by the upper left corner coordinates (x min , y min ) and the lower right corner coordinates (x max , y max ).

[0150] Calculate the area of the BBox:

[0151] S i =(x max -x min )×(y max -y min ) Formula (8)

[0152] Exemplarily, the preset area threshold S th is set to 3.

[0153] After the above processing, multiple geometries of water surface target bounding boxes are obtained. As Figure 8 shown is the result of the minimum bounding rectangle of the target obtained from the monitoring screen of a certain bridge after the above steps. It can be seen that there are mainly three targets in the image, and the existing ship structures are marked with BBox. The whole ship needs to be detected, so it is necessary to merge unnecessary BBoxes.

[0154] Step S22: Circularly merge and screen multiple target bounding boxes to obtain the bounding box of each real water surface target.

[0155] To optimize the generation of target boxes and reduce the number of redundant boxes, the intersection over union (IoU) is used to achieve this. IoU can effectively measure the overlap degree between two target boxes, thus providing a basis for the merging and elimination of targets.

[0156] Circularly merge and screen multiple target bounding boxes to obtain the bounding box of each real water surface target, including:

[0157] Calculate any two target bounding boxes B in the set of multiple water surface target bounding boxes a and B b The intersection and union areas of, to obtain the intersection over union IoU(B a , B b );

[0158] If the intersection over union of two target bounding boxes is greater than the intersection over union threshold IoU threshold , it is determined that the two target bounding boxes belong to the same target, and the smaller one of the two target bounding boxes is merged into the larger one; multiple merged BBoxes are obtained;

[0159] After the merging is completed, if the area of the merged BBox is less than the preset area threshold S m , it is determined as a noise or interference area and eliminated, and the bounding box BBox of each real water surface target is obtained.

[0160] For the intersection over union IoU(B a and B b ), it is as follows: a , B b );

[0161]

[0162] Among them, B a ∩B b , B a ∪B b are the intersection and union of the bounding boxes B a and B b respectively.

[0163] The intersection area |Ba∩Bb| and the union area |Ba∪Bb| are calculated as follows:

[0164]

[0165] Among them, are the maximum and minimum values of the bounding box B a on the horizontal and vertical coordinates respectively; are the maximum and minimum values of the bounding box B b on the horizontal and vertical coordinates respectively.

[0166] |B a ∪B b | = |B a | + |B b | - |B a ∩B b | Formula (10)

[0167] This step obtains the intersection over union of any two bounding boxes.

[0168] Exemplarily, the intersection over union threshold IoU threshold is set to 0.5; if the intersection over union IoU(B a , B b ) of any two bounding boxes is greater than IoU threshold , it is determined that these two bounding boxes represent the same object, and they are merged into a larger bounding box B merged :

[0169] B merged = B a ∪B b Formula (11)

[0170] If the intersection over union of two BBoxes is large, it is determined that the spatial overlap degree of these two BBoxes is high, and if IoU(B large , B small ) > IoU threshold , it is more likely to belong to the same object, then B small is incorporated into B large , and the bounding box B small with a smaller area is directly eliminated and incorporated into the bounding box B large .

[0171] After the merging is completed, all BBoxes are screened to eliminate false detection regions. The screening rule is: if the area of a certain BBox is less than the preset area threshold S m and it cannot be merged with other BBoxes (that is, it does not meet the merging condition that IoU(B a , B b ) is greater than IoU threshold ), it is determined that it belongs to the noise or water surface interference region and is eliminated. Exemplarily, S m is set to 100 pixel areas.

[0172] The merging step ensures that the water surface ship objects will not be split into multiple BBoxes due to water surface fluctuations or noise. The screening step eliminates isolated small-area connected regions to avoid false alarms. The finally output hull boundary BBox is more stable, can accurately describe the hull boundary, and provides high-quality detection results for subsequent object tracking and behavior analysis.

[0173] Step S22 effectively reduces redundant boxes, improves the accuracy of object detection, and ensures that each object is represented by only one bounding box, as Figure 9 shown, avoiding the problem of duplicate counting of water surface objects. This process is iterated among all BBoxes until there are no more BBoxes that meet the conditions and need to be merged. As Figure 9 Schematic diagram of the result of screening BBoxes through cyclic merging. Non-target components on the water surface are deleted, and BBoxes of the hull itself and surrounding structures are retained and merged. A total of 3 watercraft targets are obtained in the image.

[0174] As Figure 10 shown in the schematic diagram of the BBox representation, the number of rows represents the number of water surface targets, and the number of columns is the information of the BBox, which are the upper left coordinates (abscissa, ordinate) and the height and width of the BBox, in pixels.

[0175] In step S22, through the merging and elimination of bounding boxes, multiple BBoxes containing the complete position information of watercraft targets are obtained. First, preliminary BBoxes are generated through connected component extraction. However, affected by water surface noise, the targets may be segmented into multiple fragmented regions or false detections may occur. Therefore, a multi-cycle merging strategy is adopted to screen and fuse according to the size, shape, and intersection over union (IoU) of the BBoxes, ensuring that each real target corresponds to a complete BBox and eliminating false detection regions.

[0176] The input of this process is a binary image, and the output is an optimized set of BBoxes. The physical meaning is to reconstruct the shape of water targets, improve detection integrity, reduce false detections, and provide accurate data for subsequent target tracking and behavior analysis.

[0177] In step S23, target tracking prediction is performed on the bounding boxes of real water surface targets in each frame to obtain the predicted position coordinates and size information of the water surface targets.

[0178] After the detection of real water targets is completed, continuous tracking of water targets is crucial. Since the movement of ship targets in the water surface environment is usually relatively stable, a matching algorithm between multiple frames is adopted to achieve real-time tracking of the targets.

[0179] Performing target tracking prediction on the bounding boxes of the real water surface targets to obtain the predicted position coordinates and size information of the water surface targets, including:

[0180] The set of bounding boxes of real water surface targets in the current frame t is B t , The set of bounding boxes of the previous frame t - 1 is B t-1,

[0181] Calculating the Euclidean distance between each box in the current frame and the previous frame

[0182] If meets the distance threshold d th , select the smallest B t i as Matching object;

[0183] Based on the Euclidean distance and the time interval between two frames, predict the speed and acceleration of the next frame of the water surface target corresponding to the matching object; calculate the speed and acceleration in the horizontal and vertical directions based on the speed and acceleration;

[0184] Based on the speed and acceleration of the water surface target in the horizontal and vertical directions, predict the position of the water surface target in the next frame t+1 and the estimated result of the size data of the width and height of the target box BBox.

[0185] Calculate the Euclidean distance between each box in the current frame and the previous frame as follows:

[0186]

[0187] where i and j are the serial numbers i of the target in the current frame image and the serial number j in the previous frame image respectively, and the maximum value is the number of water surface targets in the current image (i.e., the number of BBoxes after the above steps), and are the center point coordinates of the labeled boxes B i and B j respectively, and the calculation is as follows:

[0188]

[0189] The distance threshold d th usually takes twice the average moving distance of the water surface target. Exemplarily, d th is taken as 8 pixels.

[0190] After the matching is completed, analyze the motion state of the water surface target. By comparing the changes in the box positions of the current frame and the previous frame, calculate the speed v and acceleration a of the target as follows:

[0191]

[0192] where Δt is the time interval between the current frame and the previous frame, v t and v t-1 are the speeds of the current frame and the previous frame respectively; v x , v y are the moving speeds of the target in the horizontal direction and the vertical direction on the image respectively, and v x , v y are the accelerations of the target in the horizontal direction and the vertical direction on the image respectively, and θ is the angle between the direction of the speed and acceleration and the horizontal direction.

[0193] Based on the calculation results of the speed and acceleration, predict the position of the target in the next frame t+1 (x pred , y pred ), and the size of the bounding box:

[0194]

[0195] Among them, is the center coordinate of the current frame, (x pred , y pred ) is the center coordinate of the predicted bounding box of the target in the next frame.

[0196] In addition to the position change, the size of the target bounding box may also change due to the change in the distance between the ship and the camera of the monitoring device. Therefore, it is necessary to predict the width and height of the target bounding box at frame t + 1. By using the method of velocity extrapolation, the calculation is as follows:

[0197]

[0198] Among them, w pred and h pred are the predicted width and height of the next frame t + 1 respectively, w i , h i are the width and height of the target in the current frame respectively, w t - w t-1 and h t - h t-1 represent the change rates of the width and height of the target bounding box between frame t - 1 and t respectively. This method can adapt to the size scaling caused by the change of the ship's perspective. For example, when the ship gradually moves away from the camera, its width and height will shrink accordingly, so as to predict the width w pred and height h pred of the next frame t + 1.

[0199] Through this prediction mechanism, the system can adjust the position and size of the target bounding box in real time, so as to maintain continuous tracking of the ship target.

[0200] Based on the kinematic principle, by calculating the historical motion trend of the target, the position and size of the target in the next frame are predicted. The position prediction infers the center position of the next frame by the motion of the target in the previous two frames, ensuring that the target will not be lost due to short-term occlusion or detection error during the tracking process. The size prediction is based on the change trend of the target bounding box, and adjusts the size of the target bounding box to adapt to the change of the target's distance, preventing tracking failure caused by scale change.

[0201] The function of step S2 is to detect and track the position and size changes of the water surface target based on the binary image through connected component analysis, BBox merging and screening, and target tracking prediction.

[0202] Step S3, specifically.

[0203] After predictive calculation based on the trends of speed and size, the estimated results of the target position and target size of frame t+1 are obtained. Then it is transformed into the water surface target dataset required by the YOLO (You Only Look Once) intelligent network, and the dataset includes multiple water surface target image data and corresponding labels.

[0204] Automatically annotate labels for water surface target images, including:

[0205] Extract HOG features from a small number of water surface target images to obtain corresponding feature vectors and corresponding manually added water surface target labels, forming a pre-training sample set;

[0206] Use the pre-training sample set to train the SVM model;

[0207] When the detection accuracy and recall rate of the SVM model meet the requirements, a pre-trained SVM model is obtained;

[0208] Input the feature vectors obtained by extracting HOG features from unlabeled water surface target images into the pre-trained SVM model to obtain the detection results of the corresponding water surface target images, which are used as the class labels of the water surface target images;

[0209] Generate YOLO-format annotation labels from the class labels of the water surface target images, the coordinates and size information of the water surface target images, and the frame numbers of the water surface target images.

[0210] Exemplarily, use the scikit-learn library in Python to implement the SVM model.

[0211] When pre-training the SVM model, select a small number of water surface target images. Exemplarily, select 100-200 water surface target images.

[0212] Exemplarily, the labels of water surface target images are cargo ships, passenger ships, fishing boats, yachts, etc. They are defined according to specific requirements.

[0213] Based on the position coordinates and size information of the water surface target, cut out water surface target images from water surface images with an area of the water surface target image larger than the preset area S;

[0214] The label annotation of each water surface target image meets the YOLO format, and the label of the water surface target image is <frame number> <class><x center ><y center > <width> <height>;

[0215] Among them, the frame number is the frame sequence number corresponding to the water surface target image; class is the category label of the water surface icon image; (x center , y center ) is the center point coordinate of the water surface target bounding box BBox; width and height are the normalized width and height of the BBox respectively.

[0216] Exemplarily, the preset area S is set to 200 pixel areas, and the shape of the water surface target can be clearly seen.

[0217] In order to avoid excessive storage space for the water surface target dataset, a storage limit needs to be set to limit the number of images of each water surface target ship, and at most N max images are set to be saved. Exemplarily, N max is set to 5.

[0218] In the YOLO object detection algorithm, the training sample data adopts a standardized annotation format. In the folder of each water surface target ship ID, an `annotations.txt` file is generated to store the position information and frame number of each image in which the water surface target ship appears in the surveillance video.

[0219] width and height are normalized to between [0, 1], rather than using pixel units. For each water surface target B i , calculate the normalized coordinates of the center point and the normalized sizes of width and height as follows:

[0220]

[0221] Among them, W and H are the width and height of the water surface target image respectively, is the normalized width; is the normalized height, are the coordinates of the points at the upper left corner and the lower right corner respectively.

[0222] The water surface target dataset includes multiple frames of color RGB images and the water surface target images obtained by frame extraction, and the corresponding category labels.

[0223] For the convenience of management and classification, two main folders will be generated:

[0224] ① Frame extraction image folder: This folder contains all the image frames extracted from the surveillance video.

[0225] ② Ship information folder: This folder is used to store the detailed information of each ship.

[0226] Label storage: The labels of each water surface target image are recorded in the annotations.txt file, meeting the data annotation requirements of the YOLO format.

[0227] In the ship information folder, create a subfolder for each detected ship, named according to the custom ship ID. For example, if the first ship in the surveillance video is 0001, the folder name is "0001". The structure of this folder is as follows:

[0228] / Ship information folder

[0229] ├──0001 /

[0230] │ ├──images / #Subfolder for storing water surface target images

[0231] │ ├──annotations.txt#Text file for storing location information in YOLO format

[0232] ├──0002 /

[0233] │ ├──images /

[0234] │ ├──annotations.txt

[0235] └──...

[0237] This step constructs a high-quality water surface target image dataset. That is, all detected targets are initially assigned the same class label to form a unified dataset. At the same time, to further improve the usability and accuracy of the data, during the target tracking process, the system records different time frames of the same target and stores its large-size morphological information in the key frames. For each target that has been stably tracked, select n representative frames from its complete motion trajectory, and store the cropped tracked target in an independent folder.

[0238] In the data collation stage, each folder corresponds to an independent water surface target instance, containing image samples of the target at different time points. These samples can not only reflect the morphological changes of the target but also provide visual features under different environmental conditions (such as lighting, angle, etc.), thereby enhancing the robustness of the dataset and providing higher-quality input data for the training of deep learning models.

[0239] The function of step S3 is to generate an annotation dataset that conforms to the YOLO format based on the predicted water surface target position and size information, providing high-quality input data for the training of deep learning models.

[0240] In summary, a method for multi-object detection and automatic annotation of water surface based on threshold segmentation according to the embodiments of the present invention has the following beneficial effects:

[0241] 1. The method for multi-object detection and annotation of water surface disclosed by the present invention realizes effective background suppression and target extraction by combining threshold segmentation and mask map technology, significantly improving the accuracy and robustness of water surface target monitoring, and is particularly suitable for dynamically changing water surface environments;

[0242] 2. The present invention automatically annotates the labels of water surface targets through a pre-trained SVM model, greatly reducing the workload and time of manual annotation and improving the annotation efficiency;

[0243] 3. The present invention can monitor and evaluate the water traffic conditions in real time, adapt to different environmental and target changes, and has stronger adaptability and practicality;

[0244] 4. The present invention can generate a high-quality water surface target sample data set, avoiding errors and omissions that may occur in manual annotation, improving the integrity and accuracy of the water surface sample data set, providing a more reliable data support basis for subsequent model training and analysis; and providing effective technical support for the intelligent development of the bridge anti-collision warning system.

[0245] Those skilled in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory or a random access memory, etc.

[0246] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.< / height> < / width> < / class> < / height> < / width> < / class>

Claims

1. A method for detecting and automatically labeling multiple targets on a water surface based on threshold segmentation, characterized in that: include: Obtain visible light monitoring video of the waterway and perform preprocessing to obtain multi-frame binary images with surface targets and backgrounds separated; Perform connected domain analysis on each frame of binary image to obtain multiple target marker box sets; Multiple target marking frames are cyclically merged and screened to obtain the marking frame of each real surface target; Perform target tracking prediction on the real surface target marking frame in each frame to obtain the predicted surface target position coordinates and size information; Based on the surface target position coordinates and size information, the surface target image is cropped, the surface target position coordinates are normalized and converted, and automatic labeling is performed to obtain a labeled surface target data set.

2. The method according to claim 1, characterized in that: Preprocessing of visible light surveillance video of waterway, including: The total number of video frames is calculated based on the total duration and frame rate of the visible light monitoring video of the surface waterway; Based on the set frame extraction interval, the visible light monitoring video of the surface channel is frame extracted to obtain corresponding multi-frame color RGB images; Convert the color RGB image of each frame into a grayscale image; Performing color inversion processing on the grayscale image of each frame to obtain a grayscale image after color inversion processing; Based on the binary mask image designed with optimized viewing angle, threshold segmentation is performed on the grayscale image after the inversion processing of each frame to obtain a multi-frame binary image with the surface target and the background separated.

3. The method according to claim 2, characterized in that: Design a binary mask image based on viewing angle optimization, including: Based on the installation position and viewing angle of the bridge camera that shoots the visible light monitoring video, calibrate the range of the water surface channel; A binary matrix mask image with the same size as the color RGB image is created, and the pixel points within the water surface channel range are assigned a value of 1, and the pixel points outside the water surface channel range are assigned a value of 0.

4. The method according to claim 1, characterized in that: Automatically label surface target images, including: Perform HOG feature extraction on a small number of water surface target images to obtain the corresponding feature vectors and the corresponding manually annotated water surface target labels to form a pre-training sample set; Using the pre-training sample set to train the SVM model; When the detection accuracy and recall rate of the SVM model meet the requirements, a pre-trained SVM model is obtained; Extracting HOG features from the unlabeled water surface target image to obtain a feature vector, inputting the feature vector into the pre-trained SVM model, and obtaining a detection result of the corresponding water surface target image as a category label of the water surface target image; The category label of the surface target image, the coordinates and size information of the surface target image, and the frame number of the surface target image are used to generate a label label in YOLO format.

5. The method according to claim 4, characterized in that: Performing threshold segmentation on the grayscale image after the inversion processing, including: Set the target background segmentation threshold T; When the grayscale value in the grayscale image after the inversion processing is higher than T and the corresponding element in the binary mask image is 0, it is marked as 1, indicating that it is a detected water surface target; otherwise, it is marked as 0, indicating that it is a detected background. The segmentation process of the water surface target and the background is as follows: Among them, I binary (x, y) is the binary image after threshold segmentation, in which the water surface target is separated from the background, and M(x, y) is the binary mask image.

6. The method according to claim 1, characterized in that: Connected domain analysis is performed on the binary image to obtain multiple water surface target marker frame sets, including: Perform connected domain analysis on the binary image, extract all target connected regions, and obtain the minimum bounding rectangle BBox of the target region. i , which is expressed as follows: Where N is the number of connected domains detected, (x min ,y min )、(x max ,y max ) are the coordinates of the upper left corner and lower right corner of BBox respectively; C-based i , calculate the area S of BBox i If S i Smaller than the preset area threshold S th , it is judged as noise or non-surface target and is removed; After elimination, multiple water surface target marking frame sets are obtained.

7. The method according to claim 6, characterized in that: Multiple target marking frames are cyclically merged and screened to obtain the marking frame of each real surface target, including: Calculate any two target marking frames B of the plurality of surface target marking frames set a and B b The intersection and union area of ​​any two marked boxes is obtained by a ,B b ); If the intersection of two target marker boxes is greater than the intersection of two target boxes threshold IoU threshold , then the two target marking frames are determined to belong to the same target, and the smaller area of ​​the two target marking frames is merged into the larger area marking frame; and multiple merged BBoxes are obtained; After the merging is completed, if the area of ​​the merged BBox is less than the preset area threshold S m , it is determined to be a noise or interference area and removed to obtain the marking box BBox of each real surface target.

8. The method according to claim 7, characterized in that: Performing target tracking prediction on the marking frame of the real surface target to obtain the predicted surface target position coordinates and size information, including: The set of real surface target frames in the current frame t is B t , The target box set of the previous frame t-1 is B t-1, Calculate the Euclidean distance of each box between the current frame and the previous frame like Meet the distance threshold d th ,choose The smallest B t i As B j t-1 Matching objects; Based on Euclidean distance and the time interval between two frames, predicting the next frame speed and acceleration of the surface target corresponding to the matching object; and calculating the speed and acceleration in the horizontal direction and the vertical direction based on the speed and acceleration; Based on the horizontal and vertical speeds and accelerations of the surface target, the position of the surface target in the next frame t+1 and the width and height size data estimation results of the target box BBox are predicted.

9. The method according to claim 8, characterized in that: Based on the position coordinates and size information of the water surface target, cutting out the water surface target image whose area is larger than the preset area S; The label annotation of each surface target image satisfies the YOLO format, and the label of the surface target image is <frame number> <class><x center ><y center > <width> <height> ;< / height> < / width> < / class> Among them, frame number is the frame number corresponding to the surface target image; class is the category label of the surface icon image; (x center ,y center ) is the coordinate of the center point of the surface target bounding box BBox; width and height are the normalized width and height of BBox respectively.

10. The method according to claims 1-9, characterized in that: The surface target data set includes multiple frames of color RGB images obtained by frame extraction and the surface target images, and corresponding category labels.

Citation Information

Patent Citations

  • Pedestrian detection and tracking method and device

    CN108985204A

  • Target detection automatic labeling method and device based on moving object detection

    CN110288629A

  • Automatic labeling method for moving target in monitoring video

    CN114693742A

  • SAR image ship detection method based on human visual attention mechanism

    CN115187856A

  • Zone based object tracking and counting

    US20220051026A1

Cited By

  • Water surface target height calculation method, system and device and computer readable storage medium

    CN120672824A