A training method and an object tracking method for an anchor-free visual object tracking network

Through the visual target tracking network without anchor design, the network architecture is simplified, real-time tracking performance is improved, and the problems of high computing complexity and poor real-time performance caused by anchor frame design are solved, achieving more efficient target tracking.

CN115205331BActive Publication Date: 2025-07-25CENT SOUTH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210718971.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-23
Publication Date
2025-07-25
Estimated Expiration
2042-06-23

AI Technical Summary

Technical Problem

The existing visual target tracking network based on deep learning depends on anchor box design, resulting in complex hyperparameter adjustment, high computational complexity, poor real-time performance, and difficulty in effectively tracking targets in complex scenarios.

Method used

An anchor-free visual target tracking network is adopted to map feature maps back to the original image receptive field area for prediction, combining feature fusion networks and multi-branch prediction networks to simplify the network architecture and improve real-time tracking performance.

Benefits of technology

Reduces computational volume and storage resources, enhances the network's ability to identify tracking targets and similar targets, improves robustness and real-time tracking performance, and suppresses low-quality prediction boxes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115205331B_ABST
    Figure CN115205331B_ABST
Patent Text Reader

Abstract

The present invention provides a training method for an anchor-free visual object tracking network and an object tracking method based on this network. The anchor-free visual object tracking network in this method is based on an anchor-free design, with a simple framework structure. The feature fusion network is used to enhance the joint response of the target in multi-layer features, improving the network's ability to distinguish between the tracked target and similar targets. In the classification task, the center confidence branch based on the Euclidean distance is used to suppress low-quality prediction boxes. The network maps each point in the feature map back to the receptive field area of the original image as a preselected box, and calculates the distances between this point and the four sides of the target box through regression prediction, so as to predict the width and height, thereby achieving a more concise network architecture and higher real-time tracking performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to a training method for an anchor-free visual object tracking network and an object tracking method based on this network. Background Art

[0002] Visual object tracking is one of the important research fields in computer vision. The main process of visual object tracking is to select a region in a certain frame of a video stream as the tracking target. This region determines the initial size and position of the tracking target. By using the context information of the video or image sequence, the appearance and motion information of the target are modeled. In the subsequent video frames, the visual object tracking algorithm automatically searches for the target using the modeled information, and finally outputs the position information and shape size of the target in a bounding box.

[0003] Visual object tracking has been applied to many fields such as aerospace, transportation, security monitoring, medical health, etc. For example, in the aerospace field, the positioning and tracking of flying targets can be achieved through visual tracking technology; in security monitoring, abnormal activities can be warned by effectively tracking trajectories; in the military field, important targets can be tracked by means of satellites or drones. However, in practical applications, due to the target often being in various complex scenarios such as occlusion, fast movement, illumination change, scale change, and background clutter, visual object tracking is still a challenging problem.

[0004] Currently, most of the common deep learning-based object tracking networks are designed based on anchor boxes. They generate multi-scale anchor boxes at each point in the feature map as preselected boxes, classify the preselected boxes into positive patches and negative patches according to the overlap degree threshold between the preselected box and the target box, and use additional regression offsets to correct the anchor boxes to obtain the accurate target position. Such anchor box designs generate a large number of hyperparameters such as the number of anchors, size, aspect ratio, etc., resulting in the need for hyperparameter adjustment and manual intervention during training, low method efficiency, high training difficulty, and high computational complexity, thus leading to poor tracking real-time performance and difficulty in practical applications. Summary of the Invention

[0005] To solve the above-mentioned defects existing in the prior art, the present invention provides a training method for an anchor-free visual object tracking network and an object tracking method based on this network. In this method, the anchor-free visual object tracking network maps each point in the feature map back to the receptive field area of the original image as a preselected box, calculates the distances between this point and the four sides of the target box through regression prediction, so as to predict the width and height, thereby realizing a more concise network architecture and higher real-time tracking performance.

[0006] According to the first aspect of the present invention, there is provided a training method for an anchor-free visual object tracking network, characterized in that: the anchor-free visual object tracking network includes: a feature extraction backbone network, a feature fusion network, and a multi-branch prediction network; the training method includes:

[0007] Step 101: Preprocess the data set. The data set includes a plurality of original images, and each original image corresponds to a ground truth bounding box, which identifies the rectangular position range where the tracking target is located in the original image.

[0008] After preprocessing, the original images form corresponding image pairs. Each image pair includes a template image and a corresponding search image. The size of the search image is larger than that of the template image, and the tracking target is located at the center position in both the search image and the template image.

[0009] Step 102: Input the template image and the search image of the image pair into the feature extraction backbone network respectively to obtain the corresponding image feature maps of the output image pair.

[0010] Among them, the feature extraction backbone network is a fully convolutional network, which is composed of 1 Stem module and 4 Stage modules connected in series. The Stem module downsamples the input template image and search image. The first Stage module is used to downsample the output of the Stem module, and the latter 3 Stage modules sequentially expand the number of channels. The total stride of the network is 8.

[0011] Step 103: Input the image feature maps into the feature fusion network to obtain the corresponding fused feature maps of the output image pair;

[0012] Step 104: Input the fused feature maps into the multi-branch prediction network to obtain the output classification score map R clc , center confidence score map R cen and distance information map R reg .

[0013] Among them, R reg is 4-channel, and the value R reg (i, j) at any position is (l, t, r, b), representing the predicted distances from the position (i, j) back to the left, top, right, and bottom of the ground truth bounding box in the original image respectively.

[0014] Step 105: Calculate the total loss value L, and train the network parameters of the anchor-free visual object tracking network according to L.

[0015] Among them, L = L clc + λ1L cen + λ2L reg , L clc is the classification loss value, L cenis the central confidence loss value, L reg is the regression loss value, and λ1 and λ2 are preset parameters.

[0016] Among them, the central confidence loss value

[0017] Lable clc (i, j) represents the sample attribution at any position (i, j). If the deviation of (l, t, r, b) at the position (i, j) from the true target box is within the preset range, then the position (i, j) belongs to the positive sample Lable clc (i, j) = 1. If the deviation is not within the preset range, then the position (i, j) belongs to the negative sample Lable clc (i, j) = 0; R cen (i, j) is the value of the central confidence score map at the position (i, j).

[0018] C cen (i, j) is the true centrality at any position (i, j), representing the normalized value of the distance from the position (i, j) mapped back to the original image to the true target center.

[0019] Furthermore, the training method of the anchor-free visual object tracking network provided by the present invention is characterized in that the feature fusion network is composed of a cropping module and a three-dimensional max pooling module.

[0020] Among them, the cropping module crops the input feature map. The image feature map includes multiple groups of corresponding template image feature maps and search image feature maps. The template image feature map is the output result of the template image after expanding the number of channels through one Stage module in the feature extraction backbone network. The search image feature map is the output result of the search image after expanding the number of channels through one Stage module in the feature extraction backbone network. The template image feature map and the search image output by the same Stage module Figure 1 correspond one by one.

[0021] The three-dimensional max pooling module performs three-dimensional maximum pooling processing on the cropped feature maps respectively to obtain multiple groups of corresponding pooled template image feature maps and pooled search image feature maps, then performs relevant operations on each group of pooled template image feature maps and pooled search image feature maps, and finally jointly splices each group to obtain a fused feature map.

[0022] Furthermore, the training method of the anchor-free visual object tracking network provided by the present invention is characterized in that the multi-branch prediction network includes a classification task network and a regression task network, both of which are fully convolutional networks.

[0023] The classification task network generates a 2-channel classification score map R based on the input fused feature map clc and a 1-channel center confidence score map R cen .

[0024] The regression task network generates a distance information map R based on the input fused feature map reg .

[0025] Furthermore, the training method of the anchor-free visual object tracking network provided by the present invention is characterized in that step 105 further includes:

[0026] The calculation method of the classification loss value is: R clc (i, j) is the value of the classification score map at the (i, j) position

[0027] The calculation method of the regression loss value is:

[0028] where (x l , y t ) is the coordinate of the upper left position of the true target box, (x r , y b ) is the coordinate of the lower right position of the true target box, G is the set of pixel points within the true target box obtained from [x l , y t , x r , y b , and P is the set of pixel points within the predicted target box obtained from R reg (i, j).

[0029] According to the second aspect of the present invention, there is provided an object tracking method based on an anchor-free visual object tracking network, characterized in that: the anchor-free visual object tracking network is trained by the training method of the first aspect; the object tracking method includes:

[0030] Step 201: Mark the initial center point and the initial target box of the tracking target in the initial frame image of the continuous video image, process the initial frame image with the initial center point position as the center to obtain a template image, and process the second frame image with the initial center point position as the center to obtain a search image

[0031] Step 202: Input the template image and the search image into the anchor-free visual object tracking network

[0032] Step 203: Obtain the classification score map R clc , the center confidence score map R cen and the distance information map R reg output by the anchor-free visual object tracking network

[0033] Step 204: Determine the tracking center point of the tracking target in the video frame corresponding to the search image according to R clc and R cen Then, determine the horizontal and vertical lengths of the tracking target box in the video image frame corresponding to the search image according to R reg so as to determine the position of the tracking target box.

[0034] Step 205: Process the next frame of the video frame corresponding to the search image with the initial center point position as the center, and use the obtained image as the search image, and repeat steps 202 to 205.

[0035] Furthermore, the object tracking method based on the anchor-free visual object tracking network provided by the present invention is characterized in that step 204 includes:

[0036] According to R clc and R cen calculate the coordinates where λ is a preset parameter, s is a scale change penalty parameter, and w is a cosine window parameter.

[0037] Map the relative displacement of (i0, j0) from the center position of R cen back to the video frame corresponding to the search image to obtain the relative displacement of the tracking center point, and add the relative displacement of the tracking center point to the initial center point position to obtain the position of the tracking center point.

[0038] Obtain the values (l0, t0, r0, b0) of R reg (i0, j0), determine that the horizontal length of the tracking target box is l0 + r0, and predict that the vertical length of the target box is t0 + b0.

[0039] According to the third aspect of the present invention, there is provided a computer device, characterized in that it includes: a memory for storing instructions; and a processor for calling the instructions stored in the memory to execute the training method of the anchor-free visual object tracking network in the first aspect.

[0040] According to the fourth aspect of the present invention, there is provided a computer device, characterized in that it includes: a memory for storing instructions; and a processor for calling the instructions stored in the memory to execute the object tracking method based on the anchor-free visual object tracking network in the second aspect.

[0041] According to the fifth aspect of the present invention, there is provided a computer-readable storage medium, characterized in that it stores instructions, and when the instructions are executed by a processor, the training method of the anchor-free visual object tracking network in the first aspect is executed.

[0042] According to a sixth aspect of the present invention, there is provided a computer-readable storage medium, characterized in that it stores instructions which, when executed by a processor, perform the object tracking method based on the anchor-free visual object tracking network of the second aspect.

[0043] Compared with the prior art, the above technical solution conceived by the present invention has at least the following beneficial effects:

[0044] 1. To reduce the computational amount and storage resources, the present invention is based on an anchor-free design and proposes a tracking network framework with a simple framework structure and powerful performance.

[0045] 2. Aiming at the problem that the features extracted by the Siamese network cannot establish a complete discriminative model for the tracking object and similar objects, the present invention proposes a 3D C-Max fusion feature fusion network to enhance the network's ability to discriminate between the tracking object and similar objects and improve the robustness of the network.

[0046] 3. To improve the tracking performance of the tracker, in the classification task, the present invention proposes a center confidence branch based on the Euclidean distance to suppress low-quality prediction boxes and remove outliers.

[0047] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.

[0049] Figure 1 is a schematic diagram of the structure of an anchor-free visual object tracking network shown according to an exemplary embodiment.

[0050] Figure 2 is a schematic diagram of the structure of a feature extraction backbone network shown according to an exemplary embodiment.

[0051] Figure 3 is a schematic diagram of the structure of a feature fusion network shown according to an exemplary embodiment.

[0052] Figure 4 is a schematic diagram of a training method of an anchor-free visual object tracking network shown according to an exemplary embodiment.

[0053] Figure 5 is a schematic diagram of an object tracking method shown according to an exemplary embodiment.

[0054] Figure 6 is an effect diagram of the object tracking method on an exemplary video sequence shown according to an exemplary embodiment.

[0055] Figure 7 It is the effect diagram of the target tracking method shown according to an exemplary embodiment on another exemplary video sequence.

[0056] Figure 8 It is the performance comparison diagram between the anchor-free visual target tracking network shown according to an exemplary embodiment and the prior art on an exemplary dataset. Detailed implementation manners

[0057] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0058] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings.

[0059] The anchor-free visual target tracking network involved in the present invention, as Figure 1 shown, includes: a feature extraction backbone network, a feature fusion network, and a multi-branch prediction network.

[0060] In some embodiments, as Figure 2 shown, the feature extraction backbone network is a fully convolutional network, which is composed of 1 Stem module and 4 Stage modules connected in series, and the total stride of the network is 8. The Stem module downsamples the input template image and search image, the first Stage module is used to downsample the output of the Stem module, and the latter 3 Stage modules sequentially expand the number of channels.

[0061] Specifically, the feature extraction backbone network can be an improved ConvNext network. The ConvNext network is improved by adjusting the Stem part of ConvNext. The first convolutional module is adjusted from the original convolution kernel kernel = 4 to kernel = 7, the stride Stride = 4 to Stride = 2, adding a padding part padding = 3, and using max pooling Maxpooling for downsampling. The improved ConvNext network of the present invention deletes the downsampling layer (Downsample), and adds a convolutional operation with a convolutional kernel size of 2 and a stride of 2 in Stage1, so that the total downsampling multiple of the entire network is 8, that is, the total stride of the feature extraction network is 8.

[0062] In some embodiments, as Figure 3As shown, the feature fusion network is a 3D C-Max fusion network, which consists of two parts: a crop module and a 3D-Maxpooling module. The crop module crops the size s of the feature map to 7×7 to obtain the cropped feature map X, and then performs 3D-Maxpooling to obtain the output result Y. Any point x in the feature map X i The corresponding value in the output result Y is y i , where:

[0063]

[0064] where p represents the padding size during 3D-Maxpooling, and k represents the width of the region for 3D-Maxpooling, is the region where x i performs one pooling operation.

[0065] After obtaining the output result Y, relevant operations are performed on Y, and then concatenation is carried out to obtain the fused feature map where T is the search image and Z is the template image. The specific calculation method is as follows:

[0066]

[0067] where, Y i represents the result after cropping and 3D-Maxpooling of the feature map output by the i-th Stage in the feature extraction backbone network, F represents the relevant operation, and cat represents concatenation.

[0068] In the present invention, the feature fusion module crops and performs 3D-Maxpooling on the feature maps output by the last 3 Stages in the feature extraction backbone network, and then performs relevant operations on the corresponding Y of the search image feature map and the template image feature map, that is, F(Y i (T), Y i (Z)), and then concatenates the results of the relevant operations and adjusts the number of channels through convolution as the final fused feature map

[0069] In some embodiments, the fused feature map is input into a multi-branch prediction network for final result prediction. The multi-branch prediction network is divided into two branches: a classification task network and a regression task network. Both the classification task network and the regression task network are fully convolutional networks.

[0070] Specifically, the classification task network consists of a classification branch and a center confidence branch. Two convolutions are used to adjust the number of channels of the classification task network, and 2-channel classification score maps R clc and 1-channel center confidence score maps R cen, this central confidence branch is designed based on the Euclidean distance, aiming to suppress low-quality prediction boxes and remove outliers. A convolution is used to adjust the number of channels in the regression task network to generate a 4-channel feature map R reg , which contains the predicted distance values (l, t, r, b) from the pixel points to the four sides of the true target box.

[0071] R clc 、R cen and R reg are the prediction results after the anchor-free visual object tracking network processes the input template image and search image.

[0072] The present invention provides a training method 10 for the above-mentioned anchor-free visual object tracking network, as Figure 4 shown, and its steps are as follows:

[0073] Step 101: Preprocess the data set. The data set includes multiple original images, and each original image corresponds to a true target box, which identifies the rectangular position range where the tracking target is located in the original image. After preprocessing, the original images form corresponding image pairs, and an image pair includes a template image and a corresponding search image, where the size of the search image is larger than that of the template image, and the tracking target is located at the center position of the search image and the template image.

[0074] In some embodiments, it is usually necessary to perform adjustments such as padding, cropping, and scaling on the training data set. For example, the template image is adjusted to a size of 127*127 with the selected target as the image center, and the search image is adjusted to a size of 255*255 with the previous position of the target as the center.

[0075] Step 102: Input the template image and the search image of the image pair into the feature extraction backbone network respectively to obtain the corresponding image feature maps of the output image pair.

[0076] Among them, the feature extraction backbone network is a fully convolutional network, which is composed of 1 Stem module and 4 Stage modules connected in series, and the total network stride is 8.

[0077] In some embodiments, the backbone network of the feature extraction network is composed of an improved ConvNext network.

[0078] Step 103: Input the image feature maps into the feature fusion network to obtain the corresponding fused feature maps of the output image pair. This step is to enhance the network's ability to represent features and suppress the responses of similar targets.

[0079] In some embodiments, the feature fusion network is a 3D C-Max fusion network, which consists of two parts: a crop module and a 3D-Maxpooling module.

[0080] Specifically, the crop module crops the input template feature map to a size of 7×7. The image feature map includes multiple groups of corresponding template image feature maps and search image feature maps. The template image feature map is the output result of the template image processed by a Stage module in the feature extraction backbone network, and the search image feature map is the output result of the search image processed by a Stage module in the feature extraction backbone network. The template image feature map and the search image feature map output by the same Stage module Figure 1 are in one-to-one correspondence.

[0081] The 3D-Maxpooling module performs 3D maximum pooling operations on the search feature map and the cropped template feature map respectively to obtain multiple groups of corresponding pooled template image feature maps and pooled search image feature maps. Then, relevant operations are performed on each group of pooled template image feature maps and pooled search image feature maps, and then the groups are jointly spliced to obtain a fused feature map.

[0082] Step 104: Input the fused feature map into the multi-branch prediction network to obtain the output classification score map R clc , center confidence score map R cen and distance information map R reg .

[0083] In some embodiments, the multi-branch prediction network includes a classification task network and a regression task network, both of which are fully convolutional networks. Among them, the classification task network generates a 2-channel classification score map R clc and a 1-channel center confidence score map R cen according to the input fused feature map; the regression task network generates a 4-channel distance information map R reg . The value R reg at any position R reg (i, j) is (l, t, r, b), representing the predicted distances from the position (i, j) back to the left, top, right, and bottom of the true target box after mapping back to the original image.

[0084] Step 105: Calculate the total loss value L, and train the network parameters of the anchor-free visual object tracking network according to L.

[0085] Among them, L = L clc + λ1L cen + λ2L reg , L clc is the classification loss value, L cen is the center confidence loss value, Lreg is the regression loss value, and λ1 and λ2 are preset parameters.

[0086] Among them, the center confidence loss value

[0087] Lable clc (i, j) represents the sample belonging at any position (i, j). If the deviation between (l, t, r, b) at the position (i, j) and the true target box is within the preset range, then the position (i, j) belongs to the positive sample Lable clc (i, j) = 1. If the deviation is not within the preset range, then the position (i, j) belongs to the negative sample Lable clc (i, j) = 0; R cen (i, j) is the value of the center confidence score map at the position (i, j).

[0088] In some embodiments, this preset range can be set to (1 - α), where α is a coefficient less than 1. According to the horizontal and vertical lengths of the true target box, calculate the horizontal deviation and vertical deviation of (l, t, r, b). If both are within (1 - α) times the horizontal and vertical dimensions of the true target box, then it is a positive sample; otherwise, it is a negative sample. Specifically:

[0089]

[0090]

[0091]

[0092] Among them, (x l , y t ) is the coordinate of the upper left position of the true target box, and (x r , y b ) is the coordinate of the lower right position of the true target box.

[0093] C cen (i, j) is the true centrality at any position (i, j), representing the normalized value of the distance from the position (i, j) mapped back to the original image to the center of the true target.

[0094] The calculation method of the classification loss value is:

[0095] Regression loss value Among them, G is based on [x l , y t , x r , y bThe set of pixel points within the obtained true target box, and P is the set of pixel points within the predicted target box obtained based on R reg (i, j).

[0096] The present invention provides a target tracking method 20 based on the above anchor-free visual target tracking network, wherein the anchor-free visual target tracking network is trained according to the training method 10.

[0097] As Figure 5 shown, the target tracking method 20 includes the following steps:

[0098] Step 201: Mark the initial center point and the initial target box of the tracking target in the initial frame image of the continuous video image, process the initial frame image centered on the position of the initial center point to obtain a template image, and process the second frame image centered on the position of the initial center point to obtain a search image.

[0099] Step 202: Input the template image and the search image into the anchor-free visual target tracking network.

[0100] Step 203: Obtain the classification score map R clc output by the anchor-free visual target tracking network, the center confidence score map R cen and the distance information map R reg .

[0101] Step 204: Determine the tracking center point of the tracking target in the video frame corresponding to the search image according to R clc and R cen , and then determine the horizontal and vertical lengths of the tracking target box in the video image frame corresponding to the search image according to R reg so as to determine the position of the tracking target box.

[0102] Step 205: Process the next frame of the video frame corresponding to the search image centered on the position of the initial center point, and use the obtained image as the search image, and repeat steps 202 to 205.

[0103] In some embodiments, step 204 includes:

[0104] Calculate the coordinates according to R clc and R cen where λ is a preset parameter, s is a scale change penalty parameter, and w is a cosine window parameter.

[0105] Map the relative displacement of (i0, j0) from the center position of R cen back to the video frame corresponding to the search image to obtain the relative displacement of the tracking center point, and add the relative displacement of the tracking center point to the position of the initial center point to obtain the position of the tracking center point.

[0106] ​Obtain R reg The value of (i0, j0) is (l0, t0, r0, b0), and it is determined that the horizontal length of the predicted target box is l0 + r0, and the vertical length of the predicted target box is t0 + b0.

[0107] Based on the same inventive concept, the present invention provides an embodiment of a training method for an anchor-free visual object tracking network.

[0108] This example is implemented based on Pytorch on an Intel(R) Xeon(R) Gold 6248R CPU @ 3.0GHz, 192GB of RAM, and 4 × Tesla V100 NVlink GPUs with 32GB of RAM. During the training process, the network batch size is set to 128, a total of 20 epochs are performed, the stochastic gradient descent method is used, and the warm_up strategy is adopted for training. The initial learning rate is 0.001, reaching a maximum of 0.005 at the 5th epoch, and the learning rate decreases logarithmically to 0.0005 in the subsequent 15 epochs. The backbone network is frozen in the first 10 epochs to train the classification branch and the regression branch, and then the entire network is unfrozen for training. The steps are as follows:

[0109] Step 1: Dataset We use ILSVRC-VID / DET, COCO, and GOT-10k as our basic training sets. Frame pairs are selected at intervals less than 100. As a data augmentation technique, we perform random translation and scaling of the search image with a uniform distribution. Then, it is necessary to perform adjustments such as padding, cropping, and scaling on the dataset. The template image is adjusted to a size of 127 × 127 with the selected target as the center of the image, and the search image is adjusted to a size of 255 × 255 with the previous position of the target as the center.

[0110] Step 2: Input the image pair described in Step 1 into a Siamese network composed of an improved ConvNext network as the backbone network, and extract image features through the backbone network. In this example, the network architecture is improved in terms of network stride and receptive field to obtain a backbone network that can be applied to an anchor-free Siamese tracking network. Adjust the Stem part of ConvNext, change the first convolutional module from the original convolutional kernel kernel = 4 to kernel = 7, the stride Stride = 4 to Stride = 2, add a padding part padding = 3, and use max pooling Maxpooling for downsampling. The improved ConvNext network in this paper chooses to delete the Downsample layer and add a convolutional operation with a convolutional kernel size of 2 and a stride of 2 in Stage1, so that the overall network downsampling multiple is 8. After passing through the feature extraction network, the resolution of the template image output by the last 3 Stages (Stage2, Stage3, Stage4) is 15×15, the resolution of the search image is 31×31, and the number of channels of both the template image and the search image is adjusted to 512 through convolution.

[0111] Step 3: To enhance the network's feature representation ability and suppress the response of similar targets, then in Step 2, input the extracted image features into a 3D C-Max fusion feature fusion module. The feature fusion module fuses the features output by the last 3 Stages in the backbone network, performs a correlation operation on the fused 3 feature maps, splices and combines the results of the correlation operation, and then adjusts the number of channels through convolution as the final feature image.

[0112] The 3D C-Max fusion network consists of crop and 3D max pooling (3D-Maxpooling). Crop the template feature map size to 7×7 to obtain the cropped feature map, and then perform 3D-Maxpooling on the search feature map and the cropped template feature map to obtain the output result.

[0113] After cropping and 3D max pooling, the resolution of the template image is 7×7 and the depth is 256, the resolution of the search image is 31×31 and the number of channels is 256, and the resolution of the finally output feature fusion result is 25×25 and the number of channels is 256.

[0114] Step 4: Input the fused feature map into a multi-branch prediction network to obtain the output classification score map R clc 、center confidence score map R cen and distance information map R reg 。

[0115] The multi-branch prediction network includes a classification task network and a regression task network, both of which are fully convolutional networks. Among them, the classification task network generates a 2-channel classification score map R clc (25×25×2) and a 1-channel center confidence score map R cen (25×25×1) according to the input fused feature map; the regression task network generates a 4-channel distance information map R reg (25×25×4) according to the input fused feature map. The value of R reg at any position R reg (i, j) is (l, t, r, b).

[0116] Step 5: Calculate the total loss value L, and train the network parameters of the anchor-free visual object tracking network according to L.

[0117] Among them, L = L clc + λ1L cen + λ2L reg , L clc is the classification loss value, L cen is the center confidence loss value, L reg is the regression loss value, and λ1 and λ2 are preset parameters.

[0118]

[0119]

[0120]

[0121] Based on the same inventive concept, the present invention provides an embodiment of an object tracking method based on an anchor-free visual object tracking network.

[0122] Tracking is performed through the trained network. The initial frame is input as the template image, and then the next frame is used as the search image. Through this network, R clc , R cen , and R reg containing the predicted box size information are output. A cosine window w is used to penalize the scores of the edge regions farther from the center point. The introduced scale change penalty s is used to reorder the classification scores, and the coordinates (i0, j0) of the maximum value are selected and the relative displacement is calculated, where:

[0123]

[0124] Then, the relative displacement is added to the center point coordinates of the previous frame target to obtain the center point of the current target. According to the predicted maximum value coordinates, the corresponding border size prediction value (l + r, t + b) is selected from R reg .

[0125] The comprehensive performance of the proposed anchor-free visual object tracking network and method is evaluated on the OTB test set, and compared with two other advanced methods such as SiamRPN and SiamFC. The accuracy and success rate of the tracker are evaluated. The accuracy is evaluated by the probability that the estimated center position is within 20 pixel points from the true center position. The success rate is defined as the percentage of the total successful frames in all frames. The intersection of the predicted target box and the true target box of the tracker is defined as the overlap rate. When it is greater than the set threshold, the frame is considered successful. The overlap rate ranges from 0 to 1. Then, the success rate curve graph with the threshold varying from 0 to 1 is drawn. The success graph uses the area under the curve as the evaluation index. Figure 6 , Figure 7 The effect diagram of the network and method in the OTB test set is shown, where the boxes in the figure represent the annotations of the tracking target in the video sequence. Figure 6 It is the result of the Basketball video sequence. Figure 7 It is the result of the BlurBody video sequence. Figure 8 It is verified that the accuracy of the network and method on the OTB test set is 0.87, the success rate is 0.649, and the comprehensive performance is higher than that of the two types of methods, SiamRPN and SiamFC.

[0126] The above OTB test set is OTB100, and the benchmark contains 100 sequences, and simultaneously involves 11 attributes of object tracking, including illumination change, scale change, occlusion, deformation, motion blur, fast motion, in-plane rotation, out-of-plane rotation, out-of-view, background clutter, and low resolution. There are both grayscale images and color images in the sequences, and each video sequence contains at least two attributes.

[0127] Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include known common knowledge or conventional technical means in the technical field not disclosed in this disclosure. The specification and examples are only considered exemplary, and the true scope and spirit of the present invention are pointed out by the following claims.

[0128] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A training method for an anchor-free visual object tracking network, characterized in that: The anchor-free visual object tracking network includes: a feature extraction backbone network, a feature fusion network, and a multi-branch prediction network; the training method includes: Step 101: Preprocess the data set, where the data set includes multiple original images, and each original image corresponds to a ground truth bounding box that identifies the rectangular position range of the tracking object in the original image; After preprocessing, the original image forms a corresponding image pair, which includes a template image and a corresponding search image, where the size of the search image is larger than that of the template image, and the tracking object is located at the center position of the search image and the template image; Step 102: Input the template image and the search image of the image pair into the feature extraction backbone network respectively to obtain the image feature maps corresponding to the output image pair; Among them, the feature extraction backbone network is a fully convolutional network, which is composed of 1 Stem module and 4 Stage modules connected in series. The Stem module downsamples the input template image and search image. The first Stage module is used to downsample the output of the Stem module, and the latter 3 Stage modules sequentially expand the number of channels. The total network stride is 8; Step 103: Input the image feature maps into the feature fusion network to obtain the fused feature maps corresponding to the output image pair; Step 104: Input the fused feature map into a multi-branch prediction network to obtain the output classification score map R clc , center confidence score map R cen and distance information map R reg ; Among them, R reg is 4-channel, and the value R at any position reg (i, j) is (l, t, r, b), representing the predicted distances from the position (i, j) to the left, top, right, and bottom of the true target box respectively after mapping back to the original image; Step 105: Calculate the total loss value L, and train the network parameters of the anchor-free visual object tracking network according to L; where, L = L cls + λ1L cen + λ2L reg , L clc is the classification loss value, L cen is the center confidence loss value, L reg is the regression loss value, and λ1 and λ2 are preset parameters; Among them, the central confidence loss value Lable clc (i, j) represents the sample attribution of any position (i, j). If the deviation between (l, t, r, b) at the position (i, j) and the true target box is within the preset range, the position (i, j) belongs to the positive sample Lable clc (i, j) = 1. If the deviation is not within the preset range, the position (i, j) belongs to the negative sample Lable clc (i, j) = 0; R cen (i, j) is the value of the center confidence score map at the position (i, j); C cen (i, j) is the true centrality of the position (i, j), representing the normalized value of the distance from the position (i, j) after being mapped back to the original image to the true target center.

2. The training method of the anchor-free visual object tracking network according to claim 1, characterized in that The feature fusion network is composed of a cropping module and a 3D max pooling module; The cropping module crops the input feature maps; The image feature maps include multiple groups of corresponding template image feature maps and search image feature maps. The template image feature maps are the output results of the template image after expanding the number of channels by one Stage module in the feature extraction backbone network. The search image feature maps are the output results of the search image after expanding the number of channels by one Stage module in the feature extraction backbone network. The template image feature maps and the search image feature maps output by the same Stage module correspond one by one; The 3D max pooling module performs 3D maximum pooling processing on the cropped feature maps respectively to obtain multiple groups of corresponding pooled template image feature maps and pooled search image feature maps, then performs correlation operations on each group of pooled template image feature maps and pooled search image feature maps, and then combines and splices each group to obtain the fused feature maps.

3. The training method of the anchor-free visual object tracking network according to claim 2, wherein, The multi-branch prediction network includes a classification task network and a regression task network, both of which are fully convolutional networks; The classification task network generates the 2-channel classification score map R based on the input fused feature map clc and the 1-channel center confidence score map R cen ; The regression task network generates the distance information map R based on the input fused feature map reg .

4. The training method of the anchor-free visual object tracking network according to claim 3, characterized in that, Step 105 further includes: The calculation method of the classification loss value is as follows: R clc (i,j) is the value of the classification score map at the (i,j) position; The calculation method of the regression loss value is: Among them, (x l , y t ) is the upper left position coordinate of the true target box, and (x r , y b ) is the lower right position coordinate of the true target box. G is the set of pixel points within the true target box obtained from [x l , y t , x r , y b , and P is the set of pixel points within the predicted target box obtained from R reg (i, j).

5. A target tracking method based on an anchor-free visual target tracking network, characterized in that: The anchor-free visual object tracking network is trained according to the training method of the anchor-free visual object tracking network described in any one of claims 1-4; the object tracking method includes: Step 201: Mark the initial center point and the initial target box of the tracking target in the initial frame image of the continuous video image. Process the initial frame image centered on the position of the initial center point to obtain the template image, and process the second frame image centered on the position of the initial center point to obtain the search image; Step 202: Input the template image and the search image into the anchor-free visual object tracking network; Step 203: Obtain the classification score map R output by the anchor-free visual object tracking network clc , the center confidence score map R cen and the distance information map R reg ; Step 204: Based on the R clc and R cen determine the tracking center point of the tracking target in the video frame corresponding to the search image, and then based on R reg determine the horizontal and vertical lengths of the tracking target box in the video image frame corresponding to the search image, so as to determine the position of the tracking target box; Step 205: Process the next frame of the video frame corresponding to the search image centered on the position of the initial center point, and use the obtained image as the search image. Repeat steps 202 to 205.

6. The object tracking method based on the anchor-free visual object tracking network according to claim 5, characterized in that, The step 204 includes: Calculate coordinates according to R clc and R cen where λ is a preset parameter, s is a scale change penalty parameter, and w is a cosine window parameter; ​ Map the relative displacement between (i0, j0) and R cen back to the video frame corresponding to the search image to obtain the relative displacement of the tracking center point, and add the relative displacement of the tracking center point to the initial center point position to obtain the position of the tracking center point; Obtain R reg (i0, j0) values (l0, t0, r0, b0), determine that the horizontal length of the tracking target box is l0 + r0, and the vertical length of the predicted target box is t0 + b0.

7. A computer device, characterized in that, including: a memory for storing instructions; and a processor for calling the instructions stored in the memory to execute the training method of the anchor-free visual object tracking network according to any one of claims 1-4.

8. A computer device, characterized in that, including: a memory for storing instructions; and a processor for calling the instructions stored in the memory to execute the object tracking method based on the anchor-free visual object tracking network according to claim 5 or 6.

9. A computer-readable storage medium, characterized in that, Stored with instructions that, when executed by a processor, execute the training method of the anchor-free visual object tracking network according to any one of claims 1-4.

10. A computer-readable storage medium, characterized in that, Stored with instructions that, when executed by a processor, execute the object tracking method based on the anchor-free visual object tracking network according to claim 5 or 6.