Ship target detection and tracking method and system based on video image processing

Through a high-definition camera combined with a hybrid Gaussian model and Mosaic enhancement method, the Yolov7 model is optimized, combined with a nuclear-related filtering algorithm, the accuracy and real-time problems of ship object detection and tracking in the existing technology are solved, and efficient ship monitoring in complex waterway environments is achieved.

CN120339745APending Publication Date: 2025-07-18SUZHOU UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510383123.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18

Smart Images

  • Figure CN120339745A_ABST
    Figure CN120339745A_ABST
Patent Text Reader

Abstract

The invention discloses a ship target detection and tracking method and system based on video image processing, and the method comprises the steps: 1), collecting a channel panoramic video stream through a high-definition camera, and extracting a continuous image sequence; 2) constructing a dynamic channel background model by using a Gaussian mixture model (GMM), and updating the background in real time to eliminate interference of water surface fluctuation and illumination variation; (3) enhancing the data obtained in the step (1) by adopting a Mosaic enhancement method, and establishing a data set; 4) optimizing a Yov7 model, and introducing a lightweight CBAM attention mechanism to carry out model training; (5) the trained weight model is experimented on the verification set, and target detection of the ship is achieved; and 6) adopting a kernel correlation filtering algorithm to realize real-time tracking of the moving ship. According to the method, the technology in the field of computer vision is utilized, the real-time video image number of the large-flow channel is effectively utilized, multi-target recognition and real-time tracking are achieved, and the method has high accuracy and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of intelligent shipping and computer vision technology, and particularly to a method for ship target detection and tracking based on video image processing. Background Art

[0002] In the modern water transportation system, ship transportation, as an important transportation mode, its safe and efficient operation is crucial. Accurately and real-time detecting and tracking ships plays an irreplaceable role in ensuring waterway safety, optimizing port operation efficiency, strengthening marine resource management, etc. For example, in busy inland waterways, timely detecting ships and accurately tracking their trajectories can effectively prevent ship collision accidents and ensure smooth passage of the waterway; in port areas, accurately grasping the dynamic information of ships helps to reasonably arrange berths and loading / unloading operations, improving the throughput capacity of the port; in the marine environment, continuous monitoring of ships can provide key data support for fishery management, maritime law enforcement and other activities.

[0003] However, traditional waterway monitoring means, such as manual recording, lidar technology, infrared sensing and the Automatic Identification System (AIS), have problems such as high cost, low efficiency, susceptibility to weather and water surface noise, lack of distance information and large statistical errors. Computer vision technology has the capabilities of high stability, high precision and all-weather monitoring, and can directly obtain the three-dimensional information of the target object, and has long been a popular choice for waterway monitoring. In addition, in complex and busy waterway monitoring scenarios, challenges such as inaccurate multi-ship monitoring and huge computational workload are often faced. Therefore, improving and developing computer vision technology for video image processing has become a key and popular field of current research. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies of the prior art, and provide a method for ship target detection and tracking based on video image processing. By comprehensively applying a variety of advanced image processing technologies and algorithm optimization means, accurate detection and stable tracking of ship targets are realized, providing strong technical support for the intelligent management of the water transportation industry.

[0005] To achieve the above purpose, an embodiment of the present invention provides a ship target detection and tracking technology based on video image processing. Specifically, it includes the following steps:

[0006] S1: Use a high-definition camera to collect the panoramic video stream of the waterway and extract a continuous image sequence.

[0007] S2: Use the Gaussian Mixture Model (GMM) to construct a dynamic waterway background model, and update the background in real time to eliminate the interference of water surface fluctuations and illumination changes;

[0008] S3: Use the Mosaic enhancement method to enhance the data obtained in step (1) and build a dataset;

[0009] S4: Optimize the Yolov7 model, introduce the attention mechanism CBAM, and perform model training;

[0010] S5: Experiment with the trained weight model on the validation set to achieve object detection of ships;

[0011] S6: Use the kernel correlation filtering algorithm to achieve real-time tracking of moving ships.

[0012] Furthermore, the ship target detection and tracking method based on video image processing in step S1. The high-definition camera collects the panoramic video stream of the waterway at a certain frame rate (such as 25 frames per second or 30 frames per second). For the accuracy and real-time performance of subsequent processing, continuous image sequences are extracted from the collected video stream. The extraction time interval can be flexibly set according to actual needs, generally between 0.1 - 1 second. If high real-time performance is required and sufficient computing resources are available, a shorter time interval (such as 0.1 second) can be selected to obtain a denser image sequence and capture the movement details of the ship more precisely; if the computing resources are limited or the real-time performance requirement is relatively low, the time interval can be appropriately increased (such as 0.5 - 1 second). The extracted image sequences are stored in a local large-capacity storage device (such as a solid-state drive) for subsequent processing.

[0013] Furthermore, in step S2, the Gaussian mixture model constructs the waterway background model. For the extracted image sequences, the Gaussian mixture model (GMM) is used to construct a dynamic waterway background model. The Gaussian mixture model describes the probability distribution of each pixel point in the image through the weighted combination of multiple Gaussian distributions. In the model initialization stage, K Gaussian distributions are set for each pixel point, and the parameters of these Gaussian distributions, including the mean μ i,t , and the weight ω i,t etc. The mean represents the average gray value or color value of the pixel point, and the weight represents the contribution of each Gaussian distribution to the pixel point.

[0014] As the image sequence is continuously input, the model updates the parameters of the Gaussian distribution in real time. When a new image frame arrives, calculate the matching degree of each pixel point with each Gaussian distribution. Specifically, by calculating the distance between the pixel point and the mean of the Gaussian distribution and performing normalization processing, the Mahalanobis distance between the pixel point x i,t and each Gaussian distribution can be expressed as

[0015]

[0016] If the Mahalanobis distance between a pixel and a certain Gaussian distribution is less than the set threshold, it is determined that the pixel belongs to the background part represented by the Gaussian distribution, and at the same time, the parameters of the Gaussian distribution are updated to make it more conform to the characteristics of the current pixel; if the Mahalanobis distance between the pixel and all Gaussian distributions is greater than the threshold, that is, d i,t ≥T holds for all i, then it is determined that the pixel is a foreground (i.e., it may be a ship target). By continuously updating the model parameters, the Gaussian mixture model can adapt to the dynamic changes of the channel environment, such as the fluctuations of the water surface, the slow changes of illumination, and the small movements of background objects, etc., accurately distinguish the background and ship targets, and provide a clear background reference for subsequent target detection.

[0017] Furthermore, in step S3, the Mosaic enhancement method is used for image data enhancement and a data set is constructed. The principle of this method is to randomly crop and splice four different images into a new image. During the cropping process, the cropping areas of each image are randomly determined, and the size and position of the cropping areas need to ensure that they can contain the complete or partial ship targets, and at the same time cover different background features as much as possible. The four cropped images are spliced according to certain rules, for example, they can be spliced up and down, left and right, or diagonally, etc., and the splicing order and method are also randomly selected. The enhanced data is merged with the original data to construct a data set for model training. When constructing the data set, the images are carefully annotated. The annotation content includes the position of the ship target (represented by the bounding box coordinates, accurate to the pixel level), the category information of the ship (such as cargo ship, passenger ship, fishing boat, etc., if there are multiple types of ships to be distinguished). The annotation process must strictly follow the unified annotation specifications and standards to ensure the accuracy and consistency of the annotation, and provide high-quality data support for subsequent model training.

[0018] Furthermore, in step S4, the Yolov7 model is optimized, and the attention mechanism CBAM is introduced for model training. CBAM is a lightweight attention module that combines channel attention and spatial attention. The channel attention mechanism focuses on the channel dimension of the feature map, processes the feature map through global average pooling and global maximum pooling respectively to obtain two different global feature descriptions, and then passes through a multi-layer perceptron (MLP) and an activation function to obtain the channel attention map for weighting each channel. The spatial attention mechanism focuses on the spatial dimension of the feature map, performs average pooling and maximum pooling operations on the feature map processed by the channel attention, compresses the information along the channel dimension, and then generates a spatial attention map through a convolutional layer and an activation function to weight the feature map spatially. The parameters of the model are continuously adjusted through the backpropagation algorithm, so that the model learns an effective feature representation of the ship target. During the training process, the model gradually adapts to different ship shapes, illumination conditions, and background environments, continuously optimizes the detection ability of the ship target, and improves the detection accuracy.

[0019] Furthermore, in step S5, the trained model is tested using the validation set. The trained weight model is experimented on the validation set. The validation set is a part of the data divided from the constructed dataset. Its data distribution and features are similar to those of the training set but not exactly the same, and it is used to evaluate the performance and generalization ability of the model. During the validation process, the images in the validation set are input into the trained model one by one, and the model outputs information such as the positions and categories of the detected ship targets. The detection results of the model are compared with the true annotation information of the images in the validation set, and the detection accuracy metrics of the model are calculated, such as the mean average precision (mAP), recall rate, accuracy rate, etc. The mean average precision (mAP) is an index comprehensively measuring the detection accuracy of the model on different category targets. It is obtained by calculating the average precision (AP) of each category and then taking the average. The calculation of the average precision (AP) is based on the change curves of the recall rate and the accuracy rate, and is obtained by integrating the accuracy rate at different recall rate thresholds. The recall rate represents the ratio of the number of targets correctly detected by the model to the actual number of targets, reflecting the detection integrity of the model for the targets; the accuracy rate represents the ratio of the number of targets correctly detected by the model to all the targets detected by the model, reflecting the accuracy of the model's detection results.

[0020] Furthermore, according to the validation results, the model is further optimized and adjusted. If the detection accuracy of the model does not meet the expectations, for example, the mAP is lower than the set threshold (such as 0.8), various optimization measures can be taken. One is to appropriately adjust the structure of the model, such as increasing or decreasing the number of attention modules, adjusting the number of layers of the backbone network, etc.; the second is to increase the amount of training data by further data augmentation or collecting more actual image data to expand the dataset; the third is to optimize the training parameters, such as adjusting the learning rate, the number of iterations, the batch size, etc. Then retrain and validate until the model achieves satisfactory detection results on the validation set, realizing accurate target detection of ships.

[0021] Furthermore, in S6, the kernel correlation filtering algorithm is used for ship target tracking. In the initial frame, according to the results of target detection, the ship target area is selected. The KCF algorithm is based on the principle of correlation filtering. By calculating the correlation between the target area and the surrounding areas, a correlation filter is constructed. In subsequent frames, the correlation filter is used to predict and update the position of the ship target. The KCF algorithm uses a kernel function to map the features in the low-dimensional space to the high-dimensional space, thereby improving the expression ability of complex target features. During the tracking process, the parameters of the correlation filter are continuously updated to adapt to the appearance changes of the ship target, such as the attitude changes (such as turning, tilting) of the ship during navigation, partial occlusion, etc. When the ship target is partially occluded, the KCF algorithm can, based on the previously learned target features and combined with the image information of the current frame, accurately predict the position of the target as much as possible, quickly relock the target after the occlusion is removed, and ensure the continuity of tracking.

[0022] The ship target detection and tracking method and system based on video image processing provided by the present invention have the following advantages:

[0023] 1. A high-definition camera is used to collect video streams and extract image sequences. By combining with the mixture Gaussian model to construct a dynamic background model, it can effectively remove the interference of complex channel backgrounds, accurately separate ship targets, and lay a good foundation for subsequent detection and tracking.

[0024] 2. Through data augmentation techniques such as the Mosaic enhancement method, a rich dataset is built, the Yolov7 model is optimized and the attention mechanism is introduced for training, significantly improving the accuracy of ship target detection and being able to accurately identify ship targets under various complex conditions. Using the kernel correlation filtering algorithm combined with the multi-object tracking algorithm for ship target tracking can track moving ships in real time and stably, effectively handle the attitude changes and occlusion problems of ships, and ensure the continuity and accuracy of tracking.

[0025] 3. The whole method comprehensively uses a variety of advanced video image processing technologies, has high adaptability and scalability, can be widely applied to different ship monitoring scenarios, and provides a reliable technical means for the intelligent management of the water transportation industry. Description of the Drawings

[0026] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0027] Figure 1 It is a flowchart of the ship target detection and tracking method based on video image processing;

[0028] Figure 2 It is a schematic diagram of the installation of the high-definition camera channel;

[0029] Figure 3 It is a schematic diagram of the comparison of Gaussian mixture background modeling;

[0030] Figure 4 It is a schematic diagram of the effect of Mosaic image data enhancement;

[0031] Figure 5 It is a schematic diagram of the improved Yolov7 model;

[0032] Figure 6 It is a schematic diagram of the comparison of the experimental effects of the model before and after improvement;

[0033] Figure 7 It is a schematic diagram of the ship tracking process based on kernel correlation filter;

[0034] Figure 8 It is a schematic diagram of the actual effect of ship target detection and tracking.

[0035] Figure 9 It shows the composition diagram of the ship target detection and tracking system based on video image processing according to the embodiment of the present application.

[0036] Figure 10 It shows the schematic structural diagram of an electronic device provided by an embodiment of the present application.

[0037] Figure 11 It shows the schematic diagram of a storage medium provided by an embodiment of the present application. Detailed implementation manners

[0038] Hereinafter, the exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.

[0039] Hereinafter, the present application will be further described in detail with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention and are not a limitation of the invention. In addition, it should be noted that, for the sake of convenience of description, only the parts related to the relevant invention are shown in the drawings.

[0040] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. Hereinafter, the present application will be described in detail with reference to the drawings and embodiments.

[0041] Currently, in the field of waterway monitoring, traditional methods such as manual statistics, infrared sensing, AIS, etc. have prominent problems such as high cost, low efficiency, large impact of weather conditions, lack of distance information, and large statistical errors. Video image processing technology has emerged in the field of inland waterway monitoring due to its advantages such as intuitive image information acquisition, high-resolution visual presentation, and effective dynamic target recognition. Therefore, this invention focuses on the problems of ship flow observation and statistics in inland waterways, and constructs a method and system for waterway ship flow statistics based on video image processing technology. With the help of video image processing technology, high-definition cameras are deployed on inland waterway bridges to achieve 24-hour all-weather automatic monitoring of waterways and collection of ship image data. Through information technology means such as deep learning, real-time ship target detection and accurate perception of motion information are completed. Refer to Figure 1 As shown, a ship target detection and tracking technology based on video image processing provided by an embodiment of this invention has the following specific implementation methods:

[0042] Step S1: Use a high-definition camera to collect the panoramic video stream of the waterway and extract a continuous image sequence. The specific steps are as follows in S101:

[0043] S101: Camera installation.

[0044] Refer to Figure 1 , the high-definition camera collects the panoramic video stream of the waterway at a certain frame rate. For the accuracy and real-time performance of subsequent processing, continuous image sequences are extracted from the collected video stream. The extraction time interval can be flexibly set according to actual needs, generally between 0.1 - 1 second. If high real-time performance is required and sufficient computing resources are available, a shorter time interval (such as 0.1 second) can be selected to obtain a denser image sequence and more precisely capture the motion details of the ship; if computing resources are limited or the real-time performance requirement is relatively low, the time interval can be appropriately increased (such as 0.5 - 1 second). During the installation process, the installation height and angle of the camera need to be precisely adjusted according to the height, tilt angle, ship navigation rules, and monitoring range requirements of the actual installation scenario. For example, when installing on the bridge of an inland waterway, a suitable pier position should be selected, and the camera should be installed at a position 15m above the water surface to ensure that its field of view can cover the waterway range of 500 - 1000 meters upstream and downstream, and minimize the occlusion of shore buildings and vegetation. At the same time, adjust the pitch angle and horizontal rotation angle of the camera so that the captured image can completely and clearly present the navigation state of the ship. The extracted image sequence is stored in a local large-capacity storage device (such as a solid-state drive) for subsequent processing.

[0045] Step S2: Use the Gaussian Mixture Model (GMM) to construct a dynamic waterway background model, refer to Figure 3 , and update the waterway background in real time to eliminate the interference of water surface fluctuations and light changes. The specific steps are as follows in S201 - S202:

[0046] S201: Model initialization.

[0047] For the initial image frame extracted from the video stream, for each pixel point x i,t , initialize the Gaussian mixture model. It is set that each pixel point is described by K Gaussian distributions, that is

[0048]

[0049] where x t represents the feature vector of the pixel point at time t (for a grayscale image, x t is the grayscale value; for a color image, x t is the RGB value vector), ω i,t is the weight of the i-th Gaussian distribution at time t, and η(x t , μ i,t , ∑ i,t ω i,t ) is the Gaussian probability density function with mean μ i,t , covariance ∑ i,t ω i,t . Through the statistical analysis of the pixel points in the initial image frame, the initial parameters of each Gaussian distribution are estimated. For example, for the mean, the average grayscale value or color value within the neighborhood of the pixel point can be used for initialization; the covariance can be estimated according to the variance of the pixel points within the neighborhood; the weight ω i,t can be set to be equal initially, such as

[0050] S202: Model update and background modeling.

[0051] As subsequent image frames are input, the parameters of the Gaussian mixture model are updated dynamically. For each pixel point in each frame of the image, calculate its Mahalanobis distance from each Gaussian distribution

[0052]

[0053] If there exists an i such that d i,t < T, where T is a pre-set threshold, then it is considered that this pixel point belongs to the background part represented by the i-th Gaussian distribution. At this time, the parameters of this Gaussian distribution are updated respectively, including weight update:

[0054] ω i,t =(1 - α)ω i,t-1 +α

[0055] where α is the learning rate, and its value range is between 0.01 and 0.025, which is used to control the speed of model update. Background update also includes mean update:

[0056] μ i,t = (1 - ρ)μ i,t-1 + ρx i,t

[0057] where ρ is the ratio of the learning rate to the weight at the i-th moment. If d i,t ≥ T holds for all i, then the pixel is determined to be the foreground (possibly a ship target), and a new Gaussian distribution is created to describe the pixel. Also, the Gaussian distribution with the smallest weight and relatively small variance is deleted to maintain the complexity and effectiveness of the model. By continuously performing the above processing on subsequent image frames, the Gaussian mixture model can adapt to the dynamic changes in the channel environment, such as water surface fluctuations, slow changes in illumination, and small movements of background objects. Finally, based on information such as the weights and variances of each Gaussian distribution, it is determined which Gaussian distributions represent the background.

[0058] Step S3: Use the Mosaic enhancement method to enhance the obtained ship image data and build a dataset. The specific steps are as follows in S301 - S302:

[0059] S301: Mosaic data enhancement.

[0060] Use the Mosaic enhancement method to process the image sequence extracted in Step S1. Randomly select four different images from the image sequence. For each image, randomly determine the cropping area. The size and position of the cropping area need to comprehensively consider the integrity of the ship target and the diversity of the channel background. For example, the side length of the cropping area can be randomly selected between a certain multiple of the side length of the original image, and the upper left coordinate of the cropping area is randomly generated within the range of the original image, but it is ensured that the cropping area contains at least part of the ship target or background features related to the ship target. Combine the four cropped images into a new image in a random splicing manner. The splicing method can be up-down, left-right splicing, such as placing in the upper left corner, in the upper right corner, in the lower left corner, in the lower right corner; or other methods such as diagonal splicing. The cropping effect is as Figure 4 shown.

[0061] S302: Dataset construction and annotation.

[0062] Merge the data after Mosaic enhancement processing with the original image sequence data to construct a dataset for model training. When annotating the dataset, use the professional image annotation tool LabelImg. For each image in the dataset, if there is a ship target, its position and category information need to be accurately annotated. When annotating the position of the ship target, mark it in the form of a bounding box. The coordinates of the upper left corner and the lower right corner of the bounding box need to be accurate to the pixel level to ensure that the main part of the ship target is accurately framed. For the category of the ship, if there are multiple types (such as cargo ships, passenger ships, fishing boats, etc.), classify and annotate according to the appearance characteristics of the ship (such as ship type, color, on-board equipment, etc.). During the annotation process, formulate detailed annotation specifications and standards, require annotators to strictly follow the specifications, and conduct random inspections and reviews of the annotation results regularly to ensure the accuracy and consistency of the annotations. For example, clearly define the judgment criteria for different types of ships, unify the annotation format, and ensure that each annotated ship target has accurate and consistent annotation information to provide high-quality data support for the subsequent training of the Yolov7 model.

[0063] Step S4: Optimize the Yolov7 model, introduce an attention mechanism for model training, and enhance the ability to capture ship features and the ability to detect small targets. The specific steps include S401 - S402:

[0064] S401: Model structure optimization.

[0065] As Figure 5 shown, CBAM is a lightweight attention module that combines channel attention and spatial attention. The channel attention mechanism focuses on the channel dimension of the feature map, processes the feature map through global average pooling and global max pooling respectively to obtain two different global feature descriptions, and then obtains the channel attention map through a multi-layer perceptron (MLP) and an activation function for weighting each channel. The spatial attention mechanism focuses on the spatial dimension of the feature map, performs average pooling and max pooling operations on the feature map after channel attention processing, compresses the information along the channel dimension, and then generates a spatial attention map through a convolutional layer and an activation function to weight the feature map spatially. Application in Yolov7: The CBAM module can be inserted into the appropriate positions of the backbone network and the neck network of Yolov7, after the residual block or the feature fusion layer. Through the CBAM module, the model can adaptively adjust the feature responses of different channels and spatial positions, enhance the attention to the key features of the ship target, and suppress the interference of background noise, thereby improving the accuracy of target detection. The structure of the CBAM module contains a dual-branch of channel attention and spatial attention, introduces an attention mechanism for model training, and enhances the ability to capture ship features and the ability to detect small targets. The CBAM module enables the model to dynamically focus on key areas such as the ship's contour and draft line, and suppresses water surface ripples and reflection noise.

[0066] S402: Training parameter setting and model training. Use the built dataset to train the optimized Yolov7 model.

[0067] Before training, set appropriate training parameters. Adopt the cosine annealing learning rate decay strategy. As the training progresses, the learning rate gradually decreases. Specifically, the formula for the change of the learning rate with the number of training rounds is:

[0068]

[0069] where is the initial learning rate, and T max is the total number of training rounds. In this embodiment, the total number of training rounds is set to 300 rounds. The batch size is set to 32, which is determined after considering the computing resources and the training effect of the ship model. A smaller batch size may lead to unstable model training, while a larger batch size can improve the training stability but will increase the memory requirement and training time.

[0070] During the training process, use the backpropagation algorithm to adjust the model parameters. The loss function of the model consists of classification loss, regression loss, and confidence loss. The classification loss uses the cross-entropy loss function to measure the difference between the predicted class of the model and the true class. Through continuous iterative training, the model gradually learns the feature representation of ship targets in different scenarios, optimizes the detection ability of ship targets, and improves the detection accuracy. During the training process, you can choose to save the weights of the model regularly for subsequent evaluation and use.

[0071] Step S5: Experiment with the trained weight model on the validation set to achieve object detection of ships. The specific steps include S501 - S502:

[0072] S501: Selection of the validation set and model validation.

[0073] Divide the validation set from the built dataset according to a certain proportion. For example, select 20% of the data as the validation set. The validation set should contain images with different lighting conditions, different ship types, and different background situations to ensure that the performance and generalization ability of the model can be comprehensively evaluated.

[0074] Load the model weights of different rounds saved during the training process into the optimized Yolov7 model to perform detection and validation on the validation set. Input the images in the validation set into the model one by one, and the model outputs information such as the position (represented by the bounding box coordinates), category, and confidence of the detected ship targets.

[0075] S502: Calculation of performance metrics and model optimization.

[0076] Compare the detection results of the model with the true annotation information of the images in the validation set, and calculate the detection accuracy metrics of the model. The mean average precision (mAP) is one of the key metrics for evaluating the model's performance. It is obtained by calculating the average precision (AP) for each class and then taking the average. When calculating the AP, it is necessary to plot the recall-precision curve. The formulas for recall and precision are as follows:

[0077]

[0078] where TP represents the number of samples correctly classified as positive, FN represents the number of samples misclassified as negative, and FP represents the number of samples correctly classified as negative. Calculate the recall and precision at different confidence thresholds, plot the curve, and calculate the area under the curve to obtain the AP value. For example, traverse the confidence threshold from 0 to 1 with a step of 0.01, calculate the recall and precision at each threshold, and then obtain the AP value. According to the validation results, if the detection accuracy of the model does not meet the expectations, such as the mAP is lower than 0.8, then optimize the model. You can try to adjust the model structure, increase or decrease the number of attention modules, or adjust the parameters of the backbone network and the neck network; you can also further increase the amount of training data by performing data augmentation again or collecting more actual image data to expand the dataset; you can also optimize the training parameters, such as adjusting the learning rate decay strategy, changing the batch size, etc. Then retrain and validate until the model achieves satisfactory detection results on the validation set.

[0079] Step S6: Implement real-time tracking of the moving ship using the kernel correlation filtering algorithm. The specific steps include S601 - S602:

[0080] S601: Initialization of the kernel correlation filtering algorithm: After determining the initial position of the ship target in the object detection stage, use the kernel correlation filtering algorithm for object tracking. Based on the ship target area detected in the initial frame, construct a kernel correlation filter.

[0081] The KCF algorithm determines the position of the target by calculating the correlation between the target area and the surrounding areas. When constructing the filter, first extract the features of the target area. Use the HOG feature, which can effectively describe the shape and texture information of the target. Perform HOG feature extraction on the target area, divide the target area into multiple small cells, calculate the histogram of gradient directions in each cell, and then combine these histograms to form a HOG feature vector. According to the HOG features of the target area, construct a kernel correlation filter. The KCF algorithm uses the kernel function to map the features in the low-dimensional space to the high-dimensional space, improving the expression ability of complex target features. In this embodiment, the Gaussian kernel function is used:

[0082]

[0083] where x i and x j are two feature vectors, σ is the bandwidth of the Gaussian kernel, and the parameters of the kernel correlation filter are solved by minimizing the objective function, and the objective function is:

[0084]

[0085] where y i is the expected response value, and for the samples at the target center position, y i is set to 1, and for other positions it is set to 0. f(x i ) is the response of the filter to the sample, λ is the regularization parameter used to prevent overfitting, usually taking values between 0.001 and 0.1, and in this embodiment it takes 0.02. n is the total number of pixels, and ω is the bandwidth of the Gaussian kernel.

[0086] S602: Target tracking and update: In subsequent video frames, the trained kernel correlation filter is used to predict the position of the ship target. According to the response value of the filter, the position of the target in the current frame is determined. The position with the maximum response value is the predicted target center position. As the ship moves and its appearance changes, the parameters of the kernel correlation filter are continuously updated. In each frame, according to the position and features of the target in the current frame, combined with the filter parameters of the previous frame, the filter is updated. Specifically, an incremental learning method is adopted to update the filter according to the new samples. For the change of the appearance model, the bilinear interpolation method is added to the update of the target model. At this time, the filter coefficients and the target observation model are:

[0087] α τ+1 =(1 - λ)α τ-1 +λα τ

[0088] where α is the learning rate used to control the speed of parameter update, λ is the regularization parameter used to prevent overfitting, and τ is the video frame sequence. When the ship target is partially occluded, the KCF algorithm can, based on the target features learned previously and combined with the image information of the current frame, predict the position of the target as accurately as possible. By comparing the similarity between the features of the predicted target region in the current frame and the previously saved target features, the occlusion situation is judged. If the similarity is lower than a certain threshold, it is considered that partial occlusion has occurred. During the occlusion period, the target is continuously tracked according to the prediction result of the filter, and the duration of the occlusion is recorded. When the occlusion is removed, the previously saved target features and the information of the current frame are used to quickly relock the target to ensure the continuity of tracking.

[0089] The application embodiment provides a ship target detection and tracking system based on video image processing, which is used to execute the ship target detection and tracking method based on video image processing described in the above embodiment, such as Figure 9 shown. The system includes:

[0090] An image acquisition module 901, which is used to collect channel image data in real time using a high-definition camera;

[0091] A channel background model construction module 902, which is used to construct a dynamic channel background model using a mixture Gaussian model, and update the channel background in real time to eliminate the interference of water surface fluctuations and illumination changes;

[0092] A data enhancement module 903, which is used for the obtained ship video image data, uses the Mosaic data enhancement method, and splices four pictures by randomly scaling, cropping and arranging them;

[0093] A model optimization module 904, which is used to optimize the Yolov7 model, introduce an attention mechanism for model training, and enhance the ability to capture ship features and the ability to detect small targets;

[0094] A target detection module 905, which is used to experiment with the trained weight model on the validation set to achieve the target detection of ships;

[0095] A trajectory prediction module 906, which is used to adopt the KCF kernel correlation filtering algorithm, use the Kalman filter to predict the ship motion trajectory, and achieve real-time tracking of moving ships.

[0096] The ship target detection and tracking system based on video image processing provided by the above embodiment of the present application and the ship target detection and tracking method based on video image processing provided by the embodiment of the present application are based on the same inventive concept, and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.

[0097] The embodiment of the present application also provides an electronic device corresponding to the ship target detection and tracking method based on video image processing provided by the foregoing embodiment to execute the ship target detection and tracking method based on video image processing. The embodiment of the present application does not make any limitations.

[0098] Please refer to Figure 10 , which shows a schematic diagram of an electronic device provided by some embodiments of the present application. As Figure 10As shown, the electronic device 20 includes: a processor 200, a memory 201, a bus 202, and a communication interface 203. The processor 200, the communication interface 203, and the memory 201 are connected through the bus 202. A computer program that can run on the processor 200 is stored in the memory 201. When the processor 200 runs the computer program, it executes the ship target detection and tracking method based on video image processing provided in any of the foregoing embodiments of the present application.

[0099] Among them, the memory 201 may include a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 203 (which can be wired or wireless), a communication connection is established between this system network element and at least one other network element, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.

[0100] The bus 202 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 201 is used to store a program. After receiving an execution instruction, the processor 200 executes the program. The ship target detection and tracking method based on video image processing disclosed in any of the foregoing embodiments of the present application can be applied to the processor 200 or implemented by the processor 200.

[0101] The processor 200 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 200 or the instructions in the form of software. The above-mentioned processor 200 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 201, and the processor 200 reads the information in the memory 201 and combines its hardware to complete the steps of the above method.

[0102] The electronic device provided by the embodiments of the present application and the method for ship target detection and tracking based on video image processing provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by them.

[0103] The embodiments of the present application also provide a computer-readable storage medium corresponding to the method for ship target detection and tracking based on video image processing provided in the foregoing embodiments. Please refer to Figure 11 , which shows that the computer-readable storage medium is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by the processor, it will execute the method for ship target detection and tracking based on video image processing provided in any of the foregoing embodiments.

[0104] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated here one by one.

[0105] The computer-readable storage medium provided by the above embodiments of the present application and the method for ship target detection and tracking based on video image processing provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.

[0106] It should be noted that:

[0107] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings provided herein. The structure required to construct such systems will be apparent from the above description. In addition, the present application is not directed to any particular programming language. It should be understood that the content of the present application described herein can be implemented using various programming languages, and the description of the specific language above is to disclose the best mode of the present application.

[0108] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0109] Similarly, it should be understood that, in order to streamline the present application and assist in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed subject matter of the present application requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the preceding disclosed single embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present application.

[0110] Those skilled in the art can understand that the modules in the devices in the embodiments can be adaptively changed and arranged in one or more devices different from those of the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be adopted to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise explicitly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that provides the same, equivalent, or similar purpose.

[0111] In addition, those skilled in the art can understand that although some of the embodiments described herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of this application and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.

[0112] Each component embodiment of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the virtual machine creation system according to the embodiments of the present application. The present application can also be implemented as a device or system program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0113] It should be noted that the above embodiments are illustrative of the present application rather than restrictive of the present application, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several systems, several of these systems can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.

[0114] As described above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the said claims.

Claims

1. A method for ship target detection and tracking based on video image processing, characterized in that, The method includes the steps: S1: Use a high-definition camera to collect channel image data in real time; S2: Use a Gaussian mixture model to construct a dynamic channel background model, and update the channel background in real time to eliminate the interference of water surface fluctuations and illumination changes; S3: For the obtained ship video image data, use the Mosaic data augmentation method to splice four pictures by randomly scaling, cropping and arranging them; S4: Optimize the Yolov7 model, introduce an attention mechanism for model training, and enhance the ability to capture ship features and the detection ability for small targets; S5: Experiment with the trained weight model on the validation set to achieve target detection of ships; S6: Adopt the KCF kernel correlation filtering algorithm, use the Kalman filter to predict the ship's motion trajectory, and achieve real-time tracking of moving ships.

2. The ship target detection and tracking method based on video image processing according to claim 1, characterized in that In step S2, it includes: S201: Model initialization: For the initial image frame extracted from the video stream, for each pixel point x i,t , initialize the Gaussian mixture model; assume that each pixel point is described by K Gaussian distributions, that is where x t represents the feature vector of the pixel at time t, and ω i,t is the weight of the i-th Gaussian distribution at time t, and η(x t , μ i,t , ∑ i,t ω i,t ) is the Gaussian probability density function with mean μ i,t and covariance ∑ i,t ω i,t ; through the statistical analysis of the pixels in the initial image frame, the initial parameters of each Gaussian distribution are estimated; S202: Model update and background modeling: With the input of subsequent image frames, dynamically update the parameters of the Gaussian mixture model; for each pixel point in each frame of image, calculate its Mahalanobis distance from each Gaussian distribution If there exists an i such that d i,t < T, where T is a preset threshold, then it is considered that this pixel point belongs to the background part represented by the i-th Gaussian distribution; at this time, the parameters of this Gaussian distribution are updated respectively, including weight update: ω i,t =(1 - α)ω i,t-1 +α where α is the learning rate, which is used to control the speed of model update; Background update also includes mean update: μ i,t =(1 - ρ)μ i,t-1 +ρx i,t where ρ is the ratio of the learning rate to the weight at the i-th moment. If d i,t ≥ T holds for all i, then the pixel is determined to be foreground, and a new Gaussian distribution is created to describe the pixel; also, the Gaussian distribution with the smallest weight and a relatively small variance is deleted; finally, based on the weight and variance information of each Gaussian distribution, it is determined which Gaussian distributions represent the background.

3. A method for ship target detection and tracking based on video image processing according to claim 1, characterized in that, In step S3, it includes: S301: Mosaic data augmentation; Process the image sequence extracted in step S1 using the Mosaic augmentation method; randomly select four different images from the image sequence; for each image, randomly determine the cropping area, but ensure that the cropping area contains at least part of the ship target or background features related to the ship target; combine the four cropped images into a new image in a random splicing manner; S302: Dataset construction and annotation; Merge the data after Mosaic enhancement processing with the original image sequence data to construct a dataset for model training; for each image in the dataset, if there is a ship target, annotate its position and category information.

4. A method for ship target detection and tracking based on video image processing according to claim 1, characterized in that, In step S3, it includes: For the obtained channel ship video images, write a program to extract video frames at a fixed frequency, and a total of 5000 datasets are obtained, which are divided into a training set and a validation set according to a ratio of 4:

1.

5. A method for ship target detection and tracking based on video image processing according to claim 1, characterized in that, In step S4, it includes: Optimize and improve the Yolov7 model, embed a CBAM module after the C3 module in the backbone network, the structure of the CBAM module contains a dual-branch of channel attention and spatial attention, and introduce an attention mechanism for model training.

6. A method for ship target detection and tracking based on video image processing according to any one of claims 1-5, characterized in that, In step S5, it includes: S501: Validation set selection and model verification; Divide a validation set from the constructed dataset according to a certain ratio, and the validation set contains images with different illumination conditions, different ship types, and different background situations; Load the model weights of different rounds saved during the training process into the optimized Yolov7 model, and perform detection verification on the validation set; input the images in the validation set into the model one by one, and the model outputs the position, category, and confidence information of the detected ship targets; S502: Performance index calculation and model optimization; Compare the detection results of the model with the real annotation information of the images in the validation set, and calculate the detection accuracy index of the model; Adjust the model structure, increase or decrease the number of attention modules, or adjust the parameters of the backbone network and the neck network; or increase the amount of training data by performing data augmentation again or collecting more actual image data to expand the dataset; or optimize the training parameters; retrain and validate until the model achieves the preset detection effect on the validation set.

7. A method for ship target detection and tracking based on video image processing according to claim 1, characterized in that, In step S6, it includes: S601: Kernelized Correlation Filter (KCF) algorithm initialization: After determining the initial position of the ship target in the object detection stage, use the KCF algorithm for object tracking; construct a kernelized correlation filter based on the ship target area detected in the initial frame. S602: Object tracking and update: In subsequent video frames, use the trained kernelized correlation filter to predict the position of the ship target; determine the position of the target in the current frame according to the response value of the filter, and the position with the maximum response value is the predicted target center position; as the ship moves and its appearance changes, continuously update the parameters of the kernelized correlation filter; in each frame, update the filter according to the position and features of the target in the current frame and the filter parameters of the previous frame. When the ship target is partially occluded, the KCF algorithm predicts the position of the target based on the previously learned target features and the image information of the current frame; judge the occlusion situation by comparing the similarity between the features of the predicted target area in the current frame and the previously saved target features; if the similarity is lower than a certain threshold, it is considered that partial occlusion has occurred; during the occlusion, continue to track the target according to the prediction result of the filter and record the duration of the occlusion; when the occlusion is lifted, use the previously saved target features and the information of the current frame to relock the target to ensure the continuity of tracking.

8. A ship target detection and tracking system based on video image processing, characterized in that, It includes: An image acquisition module for using a high-definition camera to collect channel image data in real time. A channel background model construction module for using a mixture of Gaussian models to construct a dynamic channel background model and updating the channel background in real time to eliminate the interference of water surface fluctuations and illumination changes. A data augmentation module for using the Mosaic data augmentation method for the obtained ship video image data, and splicing four pictures by randomly scaling, cropping and arranging them. A model optimization module for optimizing the Yolov7 model, introducing an attention mechanism for model training, and enhancing the ability to capture ship features and the detection ability for small targets. An object detection module for experimenting with the trained weight model on the validation set to achieve object detection of ships. A trajectory prediction module for using the KCF kernelized correlation filter algorithm and using Kalman filtering to predict the ship's motion trajectory to achieve real-time tracking of moving ships.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor runs the computer program to implement the method according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the method according to any one of claims 1-7.

Citation Information

Cited By

  • Ship flow detection and analysis method based on laser radar

    CN121170722A

  • Electric power data acquisition method for submarine optical cable

    CN121559158A

  • A power data acquisition method for a submarine optical cable

    CN121559158B

  • Underwater robot based on three-color camera and visual perception method thereof

    CN122116105A