Maritime target tracking method, device and equipment based on template updating SiameseRPN + + and medium
By adopting the SiameseRPN++ method based on template update in maritime target tracking, the problem of low accuracy of maritime target tracking is solved, and high-precision tracking and stability in complex environments are achieved.
Patent Information
- Application Number
- CN202411994297.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-31
AI Technical Summary
In the prior art, the accuracy of offshore target tracking is low, especially in the case of target posture transformation, occlusion and background interference.
The SiameseRPN++ method based on template update is adopted to extract features in the video frame sequence and track targets through the collaborative work of template branches, detection branches, twin branches and template update branches. The template library stores template images of different states in the target, and determines whether to update the template by similarity.
It improves the accuracy and stability of maritime target tracking, and can maintain high tracking accuracy under target appearance changes, posture changes and background interference, reducing the risk of mis-tracking.
Smart Images

Figure CN119942151A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target tracking, and in particular to a method, device, equipment and medium for tracking a target at sea based on template-updated SiameseRPN++. Background Art
[0002] In the field of computer vision, target tracking is a hot research topic that has attracted much attention. It aims to build a model of the target's appearance and motion information based on the target given in the initial frame in a video or image sequence, and then predict the target's motion state and locate its position in subsequent frames. Marine target tracking, as an application extension of target tracking technology in specific marine environments, is extremely critical to marine environmental monitoring and security in many marine-related fields such as ship tracking, border defense, and fishery supervision.
[0003] Although the research on maritime target tracking has developed rapidly in recent years, there are still many problems that are difficult to solve.
[0004] The posture changes, occlusions, and background interference of targets, such as ships, during motion have a great impact on target tracking algorithms. Related technologies use methods such as data enhancement and data generation to expand the data set, but there are still certain limitations for specific scenarios.
[0005] Therefore, there is an urgent need for a maritime target tracking method that can improve tracking accuracy. Summary of the invention
[0006] Therefore, the technical problem to be solved by the present invention is to overcome the problem of low accuracy in marine target tracking in the related art.
[0007] In order to solve the above technical problems, the present invention provides a method for tracking a target at sea based on template-updated SiameseRPN++, which is applied to a target tracking model. The target tracking model includes a template branch, a detection branch, a twin branch and a template update branch. The method for tracking a target at sea based on template-updated SiameseRPN++ includes:
[0008] Processing the video to be detected to determine a video frame sequence;
[0009] Perform the following tracking operations until the current frame image is the last frame image of the video frame sequence:
[0010] The template branch is used to extract template features according to the target image of this tracking operation; the target image of the first tracking operation is obtained by cropping the first frame image; the target image of this tracking operation is stored as a template image in a template library; the template library includes template images of different states of the target;
[0011] The detection branch is used to extract the current frame features according to the current frame image targeted by this tracking operation; the current frame image targeted by the first tracking operation is the second frame image;
[0012] The twin branches are used to determine the tracking result of the current tracking operation according to the template features extracted by the current tracking operation and the current frame features; the tracking result of the current tracking operation includes the position information of the target in the current frame image of the current tracking operation and the foreground and background classification results;
[0013] According to the tracking result of this tracking operation, the current frame image of this tracking operation is cut to obtain the predicted image of this tracking operation;
[0014] Determine the similarity between the predicted image targeted by the current tracking operation and each template image in the template library by using the template updating branch;
[0015] When it is determined that the current frame image targeted by this tracking operation is not the last frame image, if there is at least one template image whose similarity with the predicted image targeted by this tracking operation falls within a first preset range, the predicted image targeted by this tracking operation is stored in the template library as a template image; the first preset range: greater than a first preset value and less than or equal to a second preset value.
[0016] In an optional implementation, after using the template update branch to determine the similarity between the predicted image targeted by the current tracking operation and each template image in the template library, the method further includes:
[0017] When it is determined that the current frame image targeted by this tracking operation is not the last frame image, if there is at least one template image whose similarity with the predicted image targeted by this tracking operation falls within a second preset range, the template features extracted by this tracking operation are determined as the template features extracted by the next tracking operation; the minimum value of the second preset range is greater than the second preset value.
[0018] In an optional implementation, after using the template update branch to determine the similarity between the predicted image targeted by the current tracking operation and each template image in the template library, the method further includes:
[0019] When it is determined that the current frame image targeted by this tracking operation is not the last frame image, if there is at least one template image whose similarity with the predicted image targeted by this tracking operation falls within a third preset range, the next tracking operation is performed; the maximum value of the third preset range is greater than or equal to the first preset value.
[0020] In an optional implementation, the using the twin branch to determine the tracking result targeted by the current tracking operation according to the template features extracted by the current tracking operation and the current frame features includes:
[0021] Perform cross-correlation calculation on the template features extracted by this tracking operation and the current frame features extracted by this tracking operation to obtain a response map targeted by this tracking operation; the response map includes the foreground and background classification results corresponding to each response point;
[0022] Extracting the response point image whose foreground category probability is greater than a preset threshold in the response image targeted by this tracking operation, and determining it as the sample image targeted by this tracking operation;
[0023] The sample image targeted by this tracking operation is input into the region proposal regression model, and the tracking result targeted by this tracking operation is determined by using the region proposal regression model.
[0024] In an optional implementation, the cross-correlation calculation of the template features extracted by the current tracking operation and the current frame features extracted by the current tracking operation includes:
[0025] Cross-correlation calculation is performed on the shallow features, middle features and deep features in the template features extracted in this tracking operation, and the shallow features, middle features and deep features of the current frame features extracted in this tracking operation.
[0026] In an optional embodiment, the method further includes:
[0027] When it is determined that the current frame image targeted by this tracking operation is the last frame image, the tracking operation is terminated to obtain a target tracking result; the target tracking result includes the position information of the target in each frame image in the video frame sequence and the foreground and background classification results.
[0028] In an optional implementation, after the video to be detected is processed to determine the video frame sequence, the method further includes:
[0029] The video frame sequence is preprocessed, and the preprocessing includes: Wiener filtering processing, cropping processing and linear interpolation processing.
[0030] In a second aspect, the present invention provides a maritime target tracking device based on template updated SiameseRPN++, which is applied to a target tracking model, wherein the target tracking model includes a template branch, a detection branch, a twin branch, and a template update branch; the maritime target tracking device based on template updated SiameseRPN++ includes:
[0031] A first processing module, used for processing the video to be detected to determine a video frame sequence;
[0032] The second processing module is used to perform the following tracking operations until the current frame image is the last frame image of the video frame sequence:
[0033] The second processing module includes a first processing unit, a second processing unit, a third processing unit, a fourth processing unit and a fifth processing unit;
[0034] The first processing unit is used to extract template features according to the target image targeted by the current tracking operation using the template branch; the target image targeted by the first tracking operation is obtained by cropping the first frame image; the target image targeted by the current tracking operation is stored as a template image in a template library; the template library includes template images of different states of the target;
[0035] The second processing unit is used to extract the current frame features according to the current frame image targeted by the current tracking operation by using the detection branch; the current frame image targeted by the first tracking operation is the second frame image;
[0036] The third processing unit is used to determine the tracking result of this tracking operation according to the template features and the current frame features extracted by the twin branches; the tracking result of this tracking operation includes the position information of the target in the current frame image of this tracking operation and the foreground and background classification result;
[0037] The fourth processing unit is used to cut the current frame image targeted by the current tracking operation according to the tracking result targeted by the current tracking operation to obtain the predicted image targeted by the current tracking operation;
[0038] The fifth processing unit is used to determine the similarity between the predicted image targeted by the current tracking operation and each template image in the template library by using the template updating branch;
[0039] When it is determined that the current frame image targeted by this tracking operation is not the last frame image, if there is at least one template image whose similarity with the predicted image targeted by this tracking operation falls within a first preset range, the predicted image targeted by this tracking operation is stored in the template library as a template image; the first preset range: greater than a first preset value and less than or equal to a second preset value.
[0040] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the template-updated SiameseRPN++ maritime target tracking method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.
[0041] In a fourth aspect, the present invention provides a computer-readable storage medium, a single computer-readable storage medium storing computer instructions, the computer instructions being used to enable a computer to execute the maritime target tracking method based on template-updated SiameseRPN++ of the above-mentioned first aspect or any corresponding embodiment thereof.
[0042] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions, wherein the computer instructions are used to enable a computer to execute the maritime target tracking method based on template updating SiameseRPN++ of the above-mentioned first aspect or any corresponding embodiment thereof.
[0043] The technical solution provided by the present invention has the following technical effects:
[0044] Through the collaborative work of the template branch, the detection branch and the twin branch, it is possible to effectively extract features from the video frame sequence and determine the location information of the target in each frame and the foreground and background classification results. In the first tracking operation, the template branch extracts the template features based on the target image of the first frame image, and the detection branch extracts the current frame features based on the second frame image. The twin branch calculates accurate tracking results based on the features of both. Based on the tracking results, the second frame image is cut to obtain a predicted image, and the similarity between the predicted image and the template image in the template library is determined. When the similarity falls within the first preset range, the predicted image is stored as a template image in the template library to update the template. This allows the target to be accurately identified and located in a complex marine background, and even if the target undergoes a certain degree of displacement, rotation, and other changes, it can maintain a high tracking accuracy.
[0045] The template update branch determines whether to update the template based on the similarity between the predicted image and the template image in the template library. When there is a template image whose similarity to the predicted image is within the first preset range (greater than the first preset value and less than or equal to the second preset value), the predicted image is stored in the template library as a new template. This template update strategy enables the model to adapt to changes in the target during the tracking process, such as changes in the target's posture, recovery after partial occlusion, and so on. For example, when a ship is sailing at sea, it may present different angles due to turning. By dynamically updating the template, the model can better adapt to these changes, continuously and accurately track the target, and avoid tracking failures or reduced accuracy due to fixed templates.
[0046] The stability of long-term tracking tasks can be improved. As time goes by, the marine environment is complex and changeable, and the target may experience various complex situations. The technical solution of the present invention can accurately track the target and reduce the risk of mistracking by continuously updating the template, thereby improving the stability of long-term tracking tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the related technologies, the drawings required for use in the specific embodiments or the related technical descriptions will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0048] Figure 1 It is a flow chart of a method for tracking a target at sea based on updating SiameseRPN++ with a template according to an embodiment of the present invention;
[0049] Figure 2 is a schematic diagram of the network structure of a target tracking model according to an embodiment of the present invention;
[0050] Figure 3 is a schematic diagram of a flow chart of a template update strategy according to an embodiment of the present invention;
[0051] Figure 4 is a flow chart of a tracking algorithm according to an embodiment of the present invention;
[0052] Figure 5 is a structural schematic diagram of a marine target tracking device based on template updating SiameseRPN++ according to an embodiment of the present invention;
[0053] Figure 6 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0054] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0055] Target tracking is a hot topic in the field of computer vision. This problem refers to specifying the target according to the initial frame in a video or image sequence, and modeling the target's appearance and motion information, so as to continuously predict the target's motion state and calibrate the target's position in subsequent video frames. As an application and extension of target tracking technology in a maritime environment, maritime target tracking plays an important role in monitoring and protecting the marine environment and safety. It is widely used in ship tracking, border defense, fishery supervision and other fields. Compared with traditional land target tracking, the maritime environment is complex and changeable, including a variety of ship types, occlusion between targets, and interference from the sea background, which makes target tracking face many challenges.
[0056] Maritime target tracking algorithms have made significant progress in recent years, and their main algorithms can be divided into three categories.
[0057] The first category is the tracking algorithm based on correlation filtering, such as the Kalman filter algorithm, Mean-Shift algorithm, MOSSE algorithm, etc. This type of algorithm mainly uses the motion information and feature information of the target between consecutive frames for tracking, converting the solution of the tracker template from complex time domain operations to Fourier domain point multiplication calculations, greatly reducing the amount of calculation and significantly improving the tracker speed. However, this type of algorithm may fail to track when dealing with complex and changeable marine environments.
[0058] The second category is tracking algorithms based on feature matching, such as the optical flow method. This algorithm is a tracking algorithm based on pixel motion, which calculates the target's motion trajectory based on the optical flow information of pixel motion between adjacent frames. In maritime target tracking, the optical flow method can be used to estimate the speed and direction of the ship, and combined with other tracking algorithms to achieve more accurate tracking.
[0059] The third category is the tracking algorithm based on deep learning. The recently developed Siamese series network tracking algorithm has achieved good performance in both accuracy and speed. The earliest SiamFC model used a fully convolutional network to achieve fast matching, but it lacked scale adaptability and had difficulty dealing with occlusion and appearance changes. The subsequent SiamRPN introduced the Region Proposal Network (RPN) based on SiamFC, and enhanced the adaptability to scale changes through the anchor box mechanism, thereby improving tracking accuracy. Further improvements include SiamMask, which combines a semantic segmentation module to enable the algorithm to generate pixel-level segmentation masks, expanding the application scenarios, but increasing the computational complexity. SiamRPN++ introduced deep networks and multi-layer fusion in feature extraction, improving the robustness of the model in complex backgrounds. Ocean innovatively canceled the anchor box design, and improved the model's adaptability to target changes through end-to-end target positioning combined with an online update strategy.
[0060] Although the research on maritime target tracking has developed rapidly in recent years, there are still many problems that are difficult to solve.
[0061] First, the cost of collecting maritime target data is high and the number of samples is small, especially in small sample scenarios. The existing well-performing supervised learning-based target tracking algorithms usually require a large amount of labeled data, and problems such as data distribution differences and expired labeled data will greatly affect the speed and effect of training. Even if the transfer learning method is used, it is difficult to effectively track new categories of targets.
[0062] Second: The target's posture changes, occlusions, and background interference during motion have a great impact on the target tracking algorithm. Related technologies use methods such as data enhancement and data generation to expand the data set, but there are still certain limitations for specific scenarios.
[0063] Third: When the model encounters a target category that has never been seen before, the existing algorithms with better performance can only identify it as one of the existing labels, making it difficult to expand the category and prone to tracking failure. Therefore, it is necessary to design an effective target tracking method for updating the template online to address the above three problems. The present invention aims to provide a marine target tracking method based on template updating SiameseRPN++.
[0064] To this end, the embodiments of the present invention provide a method, device, equipment and medium for tracking maritime targets based on template-updated SiameseRPN++ to solve the above problems.
[0065] According to an embodiment of the present invention, an embodiment of a maritime target tracking method based on template updating SiameseRPN++ is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer device such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in an order different from that shown here.
[0066] Figure 1 It is a flowchart of a method for tracking a maritime target based on updating SiameseRPN++ with a template according to an embodiment of the present invention.
[0067] like Figure 1 As shown, an embodiment of the present invention provides a method for tracking a target at sea based on template-updated SiameseRPN++, which is applied to a target tracking model. The target tracking model includes a template branch, a detection branch, a twin branch, and a template update branch. The network structure of the target tracking model is shown in FIG. Figure 2 shown.
[0068] The maritime target tracking method based on template updating SiameseRPN++ includes:
[0069] S101: Process the video to be detected to determine a video frame sequence.
[0070] In this embodiment, the video to be detected is fully sampled to obtain a video frame sequence, and the video frame sequence includes multiple frame images, for example, a first frame image, a second frame image, a third frame image, and the like.
[0071] In this embodiment, after the video to be detected is processed to determine the video frame sequence, the video frame sequence is preprocessed to obtain a preprocessed video frame sequence. The preprocessing includes: Wiener filtering, cropping and linear interpolation. The preprocessed video frame sequence is applied to the technical solution of S102 described below.
[0072] Each frame image in the video frame sequence is processed with a Wiener filter to eliminate a series of motion blurs generated by rapid movement of marine targets, such as ships, and each frame image after filtering is cropped, and after quadratic linear interpolation, a template image with a first preset size, such as 127*127 size, after eliminating motion blur, and an image with a second preset size, such as 255*255 size, after eliminating motion blur are obtained.
[0073] S102: Perform the following tracking operations until the current frame image is the last frame image of the video frame sequence:
[0074] S1021: Utilize the template branch to extract template features according to the target image for this tracking operation.
[0075] In this embodiment, the target image for the first tracking operation is obtained by cutting the first frame image. The target image for this tracking operation is stored in the template library as a template image. The template library includes template images of different states of the target.
[0076] In this embodiment, for the first tracking operation, the target image for the first tracking operation can be obtained by selecting the target (e.g., a ship) in the first frame image and cutting out the background. The target image includes the target. The target image for the first tracking operation can be stored in the template library as the first state. The template library includes multiple states of a single target. As an example, four states can be set, usually the front state, the left state, the right state, and the back state. For example, the front, left, right, and back of a ship.
[0077] During the first tracking operation, there is only one state sample in the template library. As the video is detected, there will be more and more templates. The same type will be merged and updated, and the different types will be separated, with a maximum of 4 states.
[0078] In this embodiment, the input of the template branch is the target image for this tracking operation, and the output of the template branch is the template feature, specifically the template branch feature map F of each convolutional layer. l (z). The template branch is used to extract features from the input of the template branch. The template branch is specifically a ResNet-50 network. The target image is the template image and will be updated during the tracking process as the algorithm of the template update branch is used.
[0079] S1022: Utilize the detection branch to extract current frame features based on the current frame image targeted by this tracking operation.
[0080] In this embodiment, the current frame image targeted by the first tracking operation is the second frame image.
[0081] In this embodiment, the current frame image targeted by the second tracking operation is the third frame image. The input of the detection branch is the current frame image targeted by this tracking operation, and the output of the detection template branch is the current frame feature, specifically the detection branch feature map F of each convolutional layer. l (x). The detection branch is used to extract features from the search area in the current frame image (the current frame image obtained after preprocessing the search area in the current frame image). The detection branch is specifically a ResNet-50 network. The template branch and the detection branch use a weight-shared ResNet-50 network to extract features to obtain the template branch feature map F l (z) and the detection branch feature map F l (x).
[0082] S1023: Determine the tracking result of this tracking operation according to the template features extracted by this tracking operation and the current frame features using the twin branch.
[0083] In this embodiment, the tracking result of this tracking operation includes the position information of the target in the current frame image of this tracking operation and the foreground and background classification result. The position information can be the coordinate information of the target frame, for example, the coordinate information of four vertices.
[0084] In this embodiment, S1023 uses the twin branch to determine the tracking result of this tracking operation according to the template features extracted by this tracking operation and the current frame features, specifically including:
[0085] The template features extracted by this tracking operation and the current frame features extracted by this tracking operation are cross-correlated to obtain the response map for this tracking operation. The response map includes the foreground and background classification results corresponding to each response point. The foreground and background classification results include the foreground category probability and the background category probability. The foreground is the image including the target, and the background is the background, which is the image not including the target.
[0086] Specifically, a cross-correlation calculation is performed on the template features extracted by this tracking operation and the current frame features extracted by this tracking operation, including: a cross-correlation calculation is performed on the shallow features, middle features and deep features in the template features extracted by this tracking operation, and the shallow features, middle features and deep features of the current frame features extracted by this tracking operation.
[0087] The response point images whose foreground category probabilities in the response graph targeted by this tracking operation are greater than a preset threshold are extracted and determined as sample images targeted by this tracking operation.
[0088] The sample image targeted by this tracking operation is input into the region proposal regression model, and the region proposal regression model is used to determine the tracking result targeted by this tracking operation.
[0089] In this embodiment, the input of the twin branch is the shallow, middle, and deep convolutional layer detection branch feature map and the shallow, middle, and deep convolutional layer template branch feature map. As an example, the shallow, middle, and deep layers may be the 7th, 13th, and 16th convolutional layers. The output of the twin branch is the tracking result for this tracking operation.
[0090] In this embodiment, the twin branches include a preset number of SiameseRPN modules, for example, 3 SiameseRPN modules.
[0091] For a single Siamese RPN module, it has two branches, namely the regression branch and the classification branch, each branch receives and The two feature maps after the convolution layer changes the dimension are used as input. After three layers of modules, a 25×25×2k classification branch response map is generated. And the regression response map of 25×25×4k
[0092]
[0093] Among them, * represents the convolution operation, which is equivalent to using the convolution kernel and Come to and Perform convolution operation.
[0094] For the three Siamese RPN modules, hierarchical aggregation is required. The classification feature maps in these outputs are called S3, S4, and S5, and the regression feature maps are called B3, B4, and B5. The present invention directly uses weighted sum for the RPN output:
[0095]
[0096] Among them, S all is the classification feature map after aggregation, B all is the regression feature map after aggregation, α and β are the corresponding weights respectively.
[0097] In this embodiment, the SiameseRPN module adds a region proposal regression (RPN) model to the twin network. By generating a set of predefined anchor boxes in the search area, predicting whether each anchor box contains the target, and regressing and adjusting the position and size of the anchor box, efficient positioning of the target is achieved. The region proposal regression model includes generating candidate regions, regression optimization, and classification scoring. In anchor selection, an anchor of one size is selected, and four aspect ratios are set, namely [0.33, 0.5, 1, 2, 3]. Regarding the selection of positive and negative samples, similar to the RPN in Faster R-CNN, by setting a threshold, the IOU of the anchor and the GT is compared. If the IOU exceeds the threshold, it is determined to be a positive sample. Similarly, if the IOU is less than a certain threshold, it is determined to be a negative sample. In the present invention, IOU>0.6 is taken as positive and IOU<0.3 is taken as negative. In addition, in order to avoid imbalance between positive and negative samples, the ratio of positive and negative samples is set to about 1:3.
[0098] In this embodiment, each response point in the response graph corresponds to a foreground category probability and a background category probability. If the foreground category probability IOU at a certain response point is greater than a preset threshold (for example, 0.6), it means that the image at the response point has a greater probability of being the location of the target, and it is determined as a positive sample. If the foreground category probability IOU at a certain response point is less than or equal to the preset threshold (or less than another preset threshold, for example, 0.3), it means that the image at the response point is most likely not the location of the target, and it is determined as a negative sample. The preset threshold can be set and modified according to actual needs.
[0099] The response point image whose foreground category probability IOU is greater than the preset threshold is extracted as the sample image for prediction regression, and the target image after regression is obtained by using the region proposal regression model. The target image after regression refers to the target image obtained by cropping the current frame image according to the regression box. Generally, many target images will be obtained. The appropriate target image is screened according to the IOU threshold, and the detection box with the largest probability is taken as the target box.
[0100] The twin branch performs cross-correlation calculation on the image features input by the template branch and the detection branch, and uses three SiameseRPN modules to obtain the target frame coordinates and foreground and background classification results of the target in the video to be detected. The response point images with foreground category probabilities greater than the preset threshold in the response graph are extracted as sample images for prediction and regression, and the predicted target frame coordinates (position information) and foreground and background classification results are obtained using the region proposal regression model.
[0101] S1024: Cutting the current frame image targeted by the current tracking operation according to the tracking result targeted by the current tracking operation to obtain the predicted image targeted by the current tracking operation.
[0102] In this embodiment, the current frame image targeted by the tracking operation is cut according to the position information and the foreground and background classification results in the current frame image targeted by the tracking operation, and the target is cut out to obtain the predicted image targeted by the tracking operation. The predicted image is an image obtained by cutting the current frame image according to the tracking result, and the predicted image includes the target. The foreground and background can be marked by custom classification labels.
[0103] S1025: Using the template update branch, determine the similarity between the predicted image targeted by this tracking operation and each template image in the template library.
[0104] The template update branch consists of three parts: feature modeling, adaptive update, and occlusion relocation. It is used to update the matching template when the target is deformed, the angle changes, or it is occluded, to express the target state in the current time period. The Gaussian mixture model and perceptual hashing algorithm are used to model, cluster, and determine the state of the target, and then a template library is established to judge the prediction results of each frame. When there are similar samples in the template library, the input of the template branch in the network can be directly replaced, allowing the algorithm to select the most reliable target state from previous tracking results to cope with changes in the target's appearance, while avoiding repeated calculation operations. The steps are as follows:
[0105] (1) Feature modeling: Use the Gaussian mixture model to model and cluster the state of the target in each frame, and then establish a template library. Assume that the random variable x is , then the Gaussian mixture model can be expressed as:
[0106]
[0107] Where: N(x; μ k ; I) is the kth type in the Gaussian mixture model. If there are K types that need to be clustered, K Gaussian distributions can be used to represent them, μ k ∈X is the mean, and the covariance matrix is set to the identity matrix I to avoid complex calculations in high-dimensional sample space. k is the mixing coefficient, which is equivalent to the weight of each type and satisfies:
[0108]
[0109] If there are similar templates, the two types k and l are merged into a category sample set n. The merging method is:
[0110] π n =π k +π l ,
[0111] (2) Adaptive update: The perceptual hash algorithm is used to compare the similarity between the current frame prediction target and the template target, so as to determine whether the current frame image is the new state of the target. The perceptual hash (pHash) algorithm performs DCT transformation on the image, obtains the mean value of the DCT coefficient, and realizes the image similarity calculation based on its transform domain characteristics. The algorithm converts the image into a 0 or 1 hash fingerprint of 8×8 size. The similarity of the two images can be determined by counting the hash fingerprints. Generally, if the number of data bits of the same hash fingerprints of the two images exceeds 59 (in terms of proportion, it is greater than 0.92), it can be considered that the two images are highly similar. If it is less than 54 (in terms of proportion, it is less than 0.84), it can be considered that the two images are less similar. The present invention adopts the following discriminant value: when the number of bits of the same bits in the two hash fingerprints n>59, it means that the two images are highly similar and can be considered to be the same state image. When the number of bits of the same bits in the two hash fingerprints is <59 and >54, it means that the two images are somewhat different, but relatively close, indicating that they are new state images of the target. If the number of identical bits in two hash fingerprints is less than 54, it means that the images are far apart and the similarity is low. It can be judged as an occlusion or tracking error, and no update operation is performed.
[0112] (3) Occlusion relocation: When the template update branch determines that the current frame is occluded or lost, the template image before the update is input into the template branch and cross-correlated with the detected video image until the target is found, and then the template update process is restarted.
[0113] like Figure 1 As shown, after using the template update branch to determine the similarity between the predicted image targeted by this tracking operation and each template image in the template library in S1025, the maritime target tracking method based on template update SiameseRPN++ includes: when the similarity falls within the first preset range, specifically, when it is determined that the current frame image targeted by this tracking operation is not the last frame image, if there is at least one template image whose similarity with the predicted image targeted by this tracking operation falls within the first preset range, the predicted image targeted by this tracking operation is stored in the template library as a template image. The template update strategy corresponding to the template update branch is as follows: Figure 3 shown.
[0114] In this embodiment, the first preset range: is greater than the first preset value and less than or equal to the second preset value. The similarity belonging to the first preset range indicates that the target in the predicted image targeted by this tracking operation is in a new state. At this time, the predicted image targeted by this tracking operation is determined as the target image targeted by the next tracking operation, or a Gaussian mixture model is established or updated, and the template branch is used to extract the target template features according to the predicted image targeted by this tracking operation, and the target template features targeted by this tracking operation are determined as the template features extracted by the next tracking operation as the input of the twin branch. This is equivalent to determining the predicted image targeted by this tracking operation as the target image targeted by the next tracking operation, and the target image targeted by the tracking operation is the template image that best suits the current target state. Determining the predicted image targeted by this tracking operation as the target image targeted by the next tracking operation or determining the target template features targeted by this tracking operation as the template features extracted by the next tracking operation as the input of the twin branch can timely adjust the target template according to the dynamic changes of the target.
[0115] In actual scenarios, the appearance of the target may change. For example, in maritime target tracking, the appearance of the ship may change due to different cargoes loaded, or present different visual effects under different lighting conditions. When the predicted image for this tracking operation is determined as the target image for the next tracking operation, the model can capture these appearance changes. Because the predicted image is based on the current tracking results, it reflects the current actual state of the target. If the appearance of the target changes, this change will be reflected in the predicted image, and then in the next tracking operation, the model can use the updated target image (that is, the previous predicted image) as the new target image, thereby adapting to the change in the target appearance and maintaining accurate tracking of the target.
[0116] The change of the target's posture is also a common dynamic situation. For example, a ship at sea may turn or tilt during navigation. By determining the target template features for this tracking operation as the template features extracted for the next tracking operation, the model can respond to this posture change in a timely manner. When the target posture changes, the corresponding target template features will also change. In the next tracking operation, using the updated target template features as the input of the twin branch can enable the model to better match the characteristics of the target in the new posture and improve the accuracy and stability of tracking.
[0117] In complex environments, dynamic changes of targets may cause tracking interruptions in traditional tracking methods. However, the method of timely adjusting the target template according to the dynamic changes of the target can effectively reduce the probability of tracking interruptions. For example, when the target undergoes a large change in appearance or posture in a short period of time, the traditional fixed template may no longer accurately match the target, resulting in tracking failure. However, by continuously updating the target image or target template features, the model can continuously adjust its own tracking strategy as the target changes, maintain the lock on the target, and ensure the continuity of tracking.
[0118] An accurate target template is crucial for locating the target. When the target changes dynamically, a timely updated target template can more accurately reflect the current position and features of the target. Taking the twin branch as an example, it determines the target position by comparing the target template features with the current frame features. If the target template can be updated in time as the target changes dynamically, the twin branch can obtain more accurate information when calculating the target position, thereby improving the accuracy of target positioning and making the tracking result more reliable.
[0119] During long-term tracking, the target may undergo multiple complex dynamic changes, and environmental factors may also change continuously. Through this mechanism of dynamically adjusting the target template, the model can continuously adapt to changes in the target and environment during long-term tracking. For example, when tracking a ship at sea for a long time, the ship may constantly change its state due to factors such as weather and sea conditions, and environmental factors such as lighting and waves may also change. Timely updating of the target template can enable the model to always maintain good tracking performance in such complex long-term tracking scenarios, and improve the robustness and adaptability of the model.
[0120] Conventional SiameseRPN++ selects the target image (e.g., ship N) in the first frame, and uses the target image selected in the first frame as a template to find the most similar target in subsequent video frames to achieve tracking. The present invention adds a template update branch to deal with problems such as target loss, occlusion, and rapid background changes. In order to reduce calculations, Hash fingerprints are used to determine whether to update the template. If updated, a Gaussian mixture model is established for clustering. As an example, the present invention sets 8 types, and clusters of the same type will be merged, that is, there are a maximum of 8 templates in the template library.
[0121] In this embodiment, the predicted image targeted by the current tracking operation is stored in the template library as a template image, and the matching template is updated to express the state of the target in the current time period.
[0122] In this embodiment, the similarity between the predicted image targeted by this tracking operation and each template image in the template library can be calculated, and the similarity can be the number n of identical bits in the Hash fingerprint. Specifically, the similarity can be determined by calculating the number n of identical bits in the Hash fingerprint (generally 64). In this embodiment, the number of identical bits in the Hash fingerprint is calculated in a conventional manner in the art. As an example, the first preset value can be set to 54, the second preset value can be set to 59, and the first preset range is 54<n≤59.
[0123] The template update branch cuts out the predicted image from the current frame image according to the target box coordinates and classification results predicted by the network, calculates the hash fingerprint of the predicted image, obtains the similarity with the template image, establishes or updates the Gaussian mixture model, and obtains the template image that best suits the current target state.
[0124] Target tracking algorithms such as Figure 4 As shown, the process is as follows:
[0125] (1) Detect and crop the predicted image of the video, scale it to 127×127 pixels, and input it into the template update branch.
[0126] (2) A Gaussian mixture model is used to model and cluster the state of the predicted image, and a template library is established to construct k Gaussian distributions N(x; μ k ; I) to represent k types of target states.
[0127] (3) Calculate the hash fingerprints (usually 64) of each frame of the predicted image and all template images, and calculate the number of identical bits n.
[0128] (4) If n>59, it is considered similar to this template feature and is directly used as the template feature to be extracted in the next tracking operation.
[0129] (5) If 54<n≤59, a Gaussian mixture model is established or updated, and the template branch is used to extract this feature as the template feature extracted for the next tracking operation.
[0130] (6) If n≤54, it is considered to be an occlusion or tracking error and the image is not considered.
[0131] In an optional embodiment, if Figure 1As shown, after using the template update branch to determine the similarity between the predicted image targeted by this tracking operation and each template image in the template library in S1025, the maritime target tracking method based on template update SiameseRPN++ also includes: a case where the similarity belongs to a second preset range. Specifically, when it is determined that the current frame image targeted by this tracking operation is not the last frame image, if there is at least one template image whose similarity with the predicted image targeted by this tracking operation belongs to the second preset range, the template features extracted by this tracking operation are determined as the template features extracted by the next tracking operation.
[0132] In this embodiment, the minimum value of the second preset range is greater than the second preset value.
[0133] As an example, the second preset range is n>59, then it is considered to be similar to this template image, and the state is similar or the same. The template features extracted by this tracking operation are directly determined as the template features extracted by the next tracking operation as the input of the twin branch, avoiding repeated calculations and improving target tracking efficiency.
[0134] When the similarity between the predicted image and the template image is within the second preset range (such as the example case of n>59), it means that the predicted image has a high similarity with the existing template image, and the state of the target may not have changed significantly. At this time, the template features extracted by this tracking operation are determined as the template features extracted by the next tracking operation, which can ensure that the twin branch uses stable and effective template information in the subsequent target positioning and classification process. This helps to maintain the continuity and accuracy of tracking when the target state is relatively stable, and avoids errors introduced by unnecessary template updates.
[0135] If the template is updated every time, regardless of whether the target state has changed substantially, the amount of calculation will increase. By setting a second preset range to determine whether to update the template features, when the similarity between the predicted image and the template image is within this range, the current template features are directly used, which can reduce unnecessary feature re-extraction and update operations. This can significantly save computing resources when processing target tracking tasks with a large number of video frames, improve the operating efficiency of the tracking algorithm, enable the algorithm to process each frame of the image faster, and achieve real-time or near real-time target tracking.
[0136] For some targets with periodic state changes, such as in maritime target tracking, some ships may operate according to fixed routes and operating modes, and their appearance and posture may periodically return to a similar state. In this case, when the similarity between the predicted image and the template image is within the second preset range, the current template features can be used to adapt well to this periodic change. The model does not need to relearn a new template every time the target state changes slightly, but can use the existing similar template features for tracking, which enhances the model's tracking adaptability to such targets.
[0137] In an optional embodiment, if Figure 1 As shown, after using the template update branch to determine the similarity between the predicted image targeted by this tracking operation and each template image in the template library in S1025, the maritime target tracking method based on template update SiameseRPN++ also includes: when the similarity belongs to the third preset range, specifically, when it is determined that the current frame image targeted by this tracking operation is not the last frame image, if there is at least one template image whose similarity with the predicted image targeted by this tracking operation belongs to the third preset range, then the next tracking operation is performed. The maximum value of the third preset range is greater than or equal to the first preset value.
[0138] As an example, if the third preset range is n≤54, it is considered to be an occlusion or tracking error, and the image is not considered, and the next tracking operation is performed.
[0139] In an optional implementation, the maritime target tracking method based on template updating SiameseRPN++ further includes:
[0140] When it is determined that the current frame image targeted by this tracking operation is the last frame image, the tracking operation is terminated to obtain the target tracking result, which includes the position information of the target in each frame image in the video frame sequence and the foreground and background classification results.
[0141] The present invention adopts a training mode of pre-training, segmented training and weight freezing, which can use relatively small data to perform transfer learning on offshore scenes and greatly shorten the training time. The training method of the present invention is as follows:
[0142] Step S1: Image preprocessing.
[0143] The video stream is fully sampled to obtain a video frame sequence, and each frame image is filtered through the Wiener filter to eliminate a series of motion blurs caused by the rapid movement of the marine target. Each frame image is scaled to obtain a 255*255 size video frame sequence with motion blur eliminated.
[0144] Step S2: Data enhancement.
[0145] The data set is expanded for the network using data enhancement methods, including geometric transformation, color transformation, noise addition, motion blur elimination, etc., to obtain images with randomly changed RGB channel value arrangements, randomly added Gaussian noise, and four flip angles, namely 0°, 90°, 180°, and 270°, and create corresponding angle labels.
[0146] Step S3: Train the ResNet-50 network.
[0147] Load the pre-trained model of ResNet-50 on the COCO dataset and use the pre-trained model of the COCO dataset as the baseline weight. In order to migrate the COCO pre-trained model to the custom task, modify the output layer and the output number of the specific classification head. Since the pre-trained model of ResNet-50 on COCO is usually targeted at 80-category classification tasks, and the target categories required to be recognized in marine scene tasks may be different, it is necessary to replace the fully connected layer (FC layer) of the network with a classifier adapted to the target task, and freeze the weights of the low-level feature layer (the first 4 layers). Only fine-tune the high-level feature layer (fully connected layer) or a few unfrozen high-level convolutional layers to avoid overfitting and reduce training time. The dataset is expanded by means of data enhancement, and finally the ResNet-50 network fine-tuned for the marine scene is trained.
[0148] Step S4: Train the three SiameseRPN modules in the twin branches.
[0149] The network obtained by fine-tuning the ResNet-50 backbone network (removing the classification head) is used as the feature extraction network of the template branch and the detection branch, and the weight of the feature extraction network is frozen. The output of the two branches is passed to the twin branch, and the detection box regression and foreground and background classification results are input through three SiameseRPN modules, and compared with the true value, the loss function is calculated, and then the weight is updated through reverse gradient calculation. The dataset used in this step is the marine target tracking video.
[0150] During the training process, the loss function is divided into two parts: classification loss and regression loss. The classification loss uses cross entropy, and the regression loss uses normalized smooth L1. Let (Ax, Ay, Aw, Ah) be the position information of an anchor, and (Tx, Ty, Tw, Th) be the position information of GT (real box), then the corresponding normalized distance can be expressed as:
[0151]
[0152] Among them, therefore, The loss function can be expressed as follows:
[0153] σ represents Parameters in the loss function.
[0154] Then the regression loss L reg It can be expressed as:
[0155]
[0156] Classification loss L cls Using softmax, it can be expressed as:
[0157] L cls =[clog(c)+(1-c)log(1-c)], where c represents the classification probability, for example, the foreground class probability.
[0158] Finally, the total loss is:
[0159] loss = L cls +λL reg , λ represents the hyperparameter that controls the regression loss weight.
[0160] Among them, λ is a hyperparameter introduced to balance the two parts of loss.
[0161] For the outputs of the three SiameseRPN modules, multi-layer fusion is required, and the present invention directly performs linear weighting.
[0162]
[0163] Step S5: Train the overall network.
[0164] The weights of the ResNet-50 and SiameseRPN modules in the template branch, detection branch, and twin branch are loosened, and the overall network is trained using the maritime target tracking video dataset until convergence.
[0165] In summary, the present invention discloses a method for tracking targets at sea based on a template-updated SiameseRPN++ network based on an adaptive SiameseRPN++ network, adds a template update branch to cope with various morphological changes or situations where obstacles occur during the tracking process, and proposes an efficient and highly operational training method. First, the target image selected from the first frame image of the video sequence is input into the template branch, and the images other than the first frame image in the video frame sequence are input into the detection branch, and the predicted image in the current video frame is obtained through the SiameseRPN++ network. Then, the predicted image is input into the template update branch, and the Gaussian mixture model and the perceptual hash algorithm are used to determine whether to update the template. Finally, the template is continuously updated or retained using the template update strategy to track the target. In view of the fact that the current SiameseRPN++ model is difficult to solve the problem of target loss or occlusion, the present invention designs a template update branch to improve the stability of the model for long-term tracking tasks and avoid the cumulative drift and accuracy reduction caused by the fixed template. At the same time, in view of the characteristics of small targets and complex backgrounds in marine scenes, the training speed and operability are greatly improved through the pre-training, segmented training and weight freezing training modes of template branches, detection branches and twin branches.
[0166] The technical solution of the present invention is aimed at marine scenes with a wide field of view, small environmental changes and sparse targets. A marine target tracking method based on template updating SiameseRPN++ is used to track specific targets, including: inputting the video to be detected into a target tracking model based on adaptive SiameseRPN++ to obtain a target position regression frame (position information of the target frame) and foreground and background classification results, and then clustering previous prediction results through a Gaussian mixture model and establishing a template library, and then using a template update strategy to replace the input of the template branch in the adaptive SiameseRPN++ network.
[0167] The present invention aims to provide a method for tracking targets at sea based on template-updated SiameseRPN++: a deep twin network is constructed using the SiameseRPN++ architecture, and the video frame features of the template branch and the detection branch are extracted through ResNet-50, and the template features are convolved with the current frame features to obtain the target position. The Gaussian mixture model and template library are used to update and select the most reliable target template and use it to update the matching template of the Siamese network, so that it can adapt to the appearance changes of the target. Finally, a regression model is introduced to further accurately position the target and reduce the impact of the background on network performance.
[0168] Beneficial effects of the present invention:
[0169] The present invention adopts a marine target tracking model based on an adaptive SiameseRPN++ network to track the target in the video to be detected. Through small sample learning, template update strategy and occlusion relocation methods, the target tracking model can stably track the marine target in a small number of samples. The present invention introduces a template update branch in SiameseRPN++ based on an adaptive SiameseRPN++ network, which can significantly enhance the robustness and adaptability of the algorithm. By adaptively updating the template to capture the changes in the appearance, scale and posture of the target, the algorithm can better cope with the changes in illumination, partial occlusion and background interference in complex scenes, and reduce the risk of mistracking. Through the Gaussian mixture model and perceptual hashing algorithm, the template update error rate can be reduced, thereby improving the stability of the model for long-term tracking tasks, and avoiding the cumulative drift and accuracy reduction caused by the fixed template. At the same time, in view of the characteristics of small targets and complex backgrounds in marine scenes, the training speed and operability of the method are greatly improved through the pre-training, segmented training and weight freezing training modes of the template branch, the detection branch and the twin branch.
[0170] It should be noted that the contents not described in detail in the specification of the present invention belong to the common knowledge of those skilled in the art.
[0171] In this embodiment, a maritime target tracking device based on template-updated SiameseRPN++ is also provided. A single device is used to implement the above-mentioned embodiment and optional implementation methods, and the descriptions that have been made will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0172] Figure 5 It is a structural schematic diagram of a maritime target tracking device based on template updating SiameseRPN++ according to an embodiment of the present invention.
[0173] The present invention provides a marine target tracking device based on template updated SiameseRPN++, which is applied to a target tracking model, wherein the target tracking model includes a template branch, a detection branch, a twin branch and a template update branch. Figure 5 As shown, the maritime target tracking device based on template-updated SiameseRPN++ includes:
[0174] The first processing module 11 is used to process the video to be detected to determine a video frame sequence.
[0175] The second processing module 12 is used to perform the following tracking operations until the current frame image is the last frame image of the video frame sequence:
[0176] The second processing module 12 includes a first processing unit 121 , a second processing unit 122 , a third processing unit 123 , a fourth processing unit 124 and a fifth processing unit 125 .
[0177] The first processing unit 121 is used to extract template features according to the target image of this tracking operation using the template branch. The target image of the first tracking operation is obtained by cropping the first frame image. The target image of this tracking operation is stored in the template library as a template image. The template library includes template images of different states of the target.
[0178] The second processing unit 122 is used to extract the current frame features according to the current frame image targeted by the current tracking operation by using the detection branch. The current frame image targeted by the first tracking operation is the second frame image.
[0179] The third processing unit 123 is used to determine the tracking result of this tracking operation based on the template features extracted by this tracking operation and the current frame features using the twin branch. The tracking result of this tracking operation includes the position information of the target in the current frame image of this tracking operation and the foreground and background classification results.
[0180] The fourth processing unit 124 is used to cut the current frame image targeted by the current tracking operation according to the tracking result targeted by the current tracking operation to obtain the predicted image targeted by the current tracking operation.
[0181] The fifth processing unit 125 is configured to use the template updating branch to determine the similarity between the predicted image targeted by the current tracking operation and each template image in the template library.
[0182] When it is determined that the current frame image targeted by the current tracking operation is not the last frame image, if there is at least one template image whose similarity with the predicted image targeted by the current tracking operation falls within a first preset range, the predicted image targeted by the current tracking operation is stored in the template library as a template image. The first preset range is greater than the first preset value and less than or equal to the second preset value.
[0183] In an optional implementation, the fifth processing unit 125 is further configured to, after determining the similarity between the predicted image targeted by the current tracking operation and each template image in the template library by using the template update branch, determine the template features extracted by the current tracking operation as the template features extracted by the next tracking operation if the similarity between at least one template image and the predicted image targeted by the current tracking operation belongs to the second preset range when determining that the current frame image targeted by the current tracking operation is not the last frame image. The minimum value of the second preset range is greater than the second preset value.
[0184] In an optional implementation, the fifth processing unit 125 is further configured to, after determining the similarity between the predicted image targeted by the current tracking operation and each template image in the template library by using the template update branch, perform the next tracking operation if the similarity between at least one template image and the predicted image targeted by the current tracking operation is within a third preset range when it is determined that the current frame image targeted by the current tracking operation is not the last frame image. The maximum value of the third preset range is greater than or equal to the first preset value.
[0185] In an optional implementation, the third processing unit 123 is specifically configured to perform cross-correlation calculation on the template features extracted by the current tracking operation and the current frame features extracted by the current tracking operation to obtain a response map targeted by the current tracking operation. The response map includes the foreground and background classification results corresponding to each response point.
[0186] The response point images whose foreground category probabilities in the response graph targeted by this tracking operation are greater than a preset threshold are extracted and determined as sample images targeted by this tracking operation.
[0187] The sample image targeted by this tracking operation is input into the region proposal regression model, and the region proposal regression model is used to determine the tracking result targeted by this tracking operation.
[0188] In an optional embodiment, the third processing unit 123 is specifically used to perform cross-correlation calculation on the shallow features, middle features and deep features in the template features extracted by this tracking operation, and the shallow features, middle features and deep features of the current frame features extracted by this tracking operation.
[0189] In an optional embodiment, the second processing module 12 includes a sixth processing unit, which is used to end the tracking operation and obtain the target tracking result when it is determined that the current frame image targeted by the tracking operation is the last frame image. The target tracking result includes the position information of the target in each frame image in the video frame sequence and the foreground and background classification results.
[0190] In an optional embodiment, the marine target tracking device based on template updated SiameseRPN++ also includes: a preprocessing module, which is used to preprocess the video frame sequence after processing the video to be detected to determine the video frame sequence, and the preprocessing includes: Wiener filtering processing, cropping processing and linear interpolation processing.
[0191] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0192] The maritime target tracking device based on template updated SiameseRPN++ in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0193] The embodiment of the present invention also provides a computer device having the above Figure 5 The maritime target tracking device shown is based on template updated SiameseRPN++.
[0194] See also Figure 6 , Figure 6 Schematic diagram of the hardware structure of the computer device according to the embodiment of the present invention. Figure 6 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In an optional embodiment, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor device). Figure 6 A processor 10 is taken as an example.
[0195] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0196] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.
[0197] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating device, an application required for at least one function. The data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage devices. In an optional embodiment, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0198] The memory 20 may include a volatile memory, such as a random access memory. The memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive. The memory 20 may also include a combination of the above types of memory.
[0199] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0200] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc. Further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.
[0201] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.
[0202] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A method for tracking a target at sea based on template-updated SiameseRPN++, applied to a target tracking model, wherein the target tracking model comprises a template branch, a detection branch, a twin branch and a template update branch; characterized in that: include: Processing the video to be detected to determine a video frame sequence; Perform the following tracking operations until the current frame image is the last frame image of the video frame sequence: Extracting template features according to the target image targeted by this tracking operation using the template branch; The target image for the first tracking operation is obtained by cutting the first frame image; the target image for this tracking operation is stored in a template library as a template image; the template library includes template images of different states of the target; Utilizing the detection branch to extract current frame features according to the current frame image targeted by this tracking operation; The current frame image targeted by the first tracking operation is the second frame image; Determine the tracking result targeted by this tracking operation by using the twin branches according to the template features extracted by this tracking operation and the current frame features; The tracking result of this tracking operation includes the position information of the target in the current frame image of this tracking operation and the foreground and background classification results; According to the tracking result of the current tracking operation, the current frame image of the current tracking operation is cut to obtain the predicted image of the current tracking operation; Determine the similarity between the predicted image targeted by the current tracking operation and each template image in the template library by using the template updating branch; When it is determined that the current frame image targeted by the current tracking operation is not the last frame image, if there is at least one template image whose similarity with the predicted image targeted by the current tracking operation falls within a first preset range, the predicted image targeted by the current tracking operation is stored in the template library as a template image; The first preset range is greater than the first preset value and less than or equal to the second preset value.
2. The method according to claim 1, characterized in that After using the template update branch to determine the similarity between the predicted image targeted by the current tracking operation and each template image in the template library, the method further includes: When it is determined that the current frame image targeted by this tracking operation is not the last frame image, if there is at least one template image whose similarity with the predicted image targeted by this tracking operation falls within a second preset range, the template features extracted by this tracking operation are determined as the template features extracted by the next tracking operation; the minimum value of the second preset range is greater than the second preset value.
3. The method according to claim 1, characterized in that After using the template update branch to determine the similarity between the predicted image targeted by the current tracking operation and each template image in the template library, the method further includes: When it is determined that the current frame image targeted by this tracking operation is not the last frame image, if there is at least one template image whose similarity with the predicted image targeted by this tracking operation falls within a third preset range, the next tracking operation is performed; the maximum value of the third preset range is greater than or equal to the first preset value.
4. The method according to claim 1, characterized in that The step of using the twin branches to determine the tracking result of the current tracking operation according to the template features extracted by the current tracking operation and the current frame features includes: Perform cross-correlation calculation on the template features extracted by this tracking operation and the current frame features extracted by this tracking operation to obtain a response map targeted by this tracking operation; the response map includes the foreground and background classification results corresponding to each response point; Extracting the response point image whose foreground category probability is greater than a preset threshold in the response image targeted by this tracking operation, and determining it as the sample image targeted by this tracking operation; The sample image targeted by this tracking operation is input into the region proposal regression model, and the tracking result targeted by this tracking operation is determined by using the region proposal regression model.
5. The method according to claim 4, characterized in that The cross-correlation calculation of the template features extracted by the current tracking operation and the current frame features extracted by the current tracking operation includes: Cross-correlation calculation is performed on the shallow features, middle features and deep features in the template features extracted in this tracking operation, and the shallow features, middle features and deep features of the current frame features extracted in this tracking operation.
6. The method according to claim 1, characterized in that The method further comprises: When it is determined that the current frame image targeted by this tracking operation is the last frame image, the tracking operation is terminated to obtain a target tracking result; the target tracking result includes the position information of the target in each frame image in the video frame sequence and the foreground and background classification results.
7. The method according to claim 1, characterized in that After the video to be detected is processed to determine the video frame sequence, the following steps are also included: The video frame sequence is preprocessed, and the preprocessing includes: Wiener filtering processing, cropping processing and linear interpolation processing.
8. A marine target tracking device based on template-updated SiameseRPN++, applied to a target tracking model, wherein the target tracking model includes a template branch, a detection branch, a twin branch and a template update branch; characterized in that: include: A first processing module, used for processing the video to be detected to determine a video frame sequence; The second processing module is used to perform the following tracking operations until the current frame image is the last frame image of the video frame sequence: The second processing module includes a first processing unit, a second processing unit, a third processing unit, a fourth processing unit and a fifth processing unit; The first processing unit is used to extract template features according to the target image of this tracking operation using the template branch; The target image for the first tracking operation is obtained by cutting the first frame image; the target image for this tracking operation is stored in a template library as a template image; the template library includes template images of different states of the target; The second processing unit is used to extract current frame features according to the current frame image targeted by this tracking operation using the detection branch; The current frame image targeted by the first tracking operation is the second frame image; The third processing unit is used to determine the tracking result of the current tracking operation according to the template features extracted by the current tracking operation and the current frame features by using the twin branches; The tracking result of this tracking operation includes the position information of the target in the current frame image of this tracking operation and the foreground and background classification results; The fourth processing unit is used to cut the current frame image targeted by the current tracking operation according to the tracking result targeted by the current tracking operation to obtain the predicted image targeted by the current tracking operation; The fifth processing unit is used to determine the similarity between the predicted image targeted by the current tracking operation and each template image in the template library by using the template updating branch; When it is determined that the current frame image targeted by the current tracking operation is not the last frame image, if there is at least one template image whose similarity with the predicted image targeted by the current tracking operation falls within a first preset range, the predicted image targeted by the current tracking operation is stored in the template library as a template image; The first preset range is greater than the first preset value and less than or equal to the second preset value.
9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the maritime target tracking method based on template-updated SiameseRPN++ as described in any one of claims 1 to 7 by executing the computer instructions.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the maritime target tracking method based on template-updated SiameseRPN++ according to any one of claims 1 to 7.
Citation Information
Patent Citations
Target tracking template updating method and device, computer equipment and storage medium
CN110634153A
Unmanned aerial vehicle adaptive target tracking method based on pseudo twin network
CN113516713A
Face false detection optimization method and device, storage medium, equipment and program product
CN115713796A
Unmanned aerial vehicle ground target tracking method based on improved twin neural network
CN116189019A
Unsupervised target tracking method based on sparse attention updating template features
CN116310971A