Image tracking model training and image tracking method and device

By downsampling the target image and search image samples and combining the shunt self-attention and mutual attention mechanism, the twin network model is trained, and the problem of insufficient robustness of scale changes in single-target tracking in videos is solved, achieving a robust adaptation to targets at different scales.

CN120339724APending Publication Date: 2025-07-18CHINA COAL RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510789249.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing single-target tracking technology of video is not robust enough when dealing with target scale changes, making it difficult to effectively adapt to target size changes, occlusion and background interference.

Method used

By downsampling the target image and search image samples, sub-image samples are generated, and the shunt self-attention and shunt mutual attention mechanisms are used to enhance and fusion characteristics, the twin network model is trained to improve adaptability to targets at different scales.

Benefits of technology

Improves the robustness of the model when dealing with different size targets, and enhances the robustness of target scale changes, occlusions, and background interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339724A_ABST
    Figure CN120339724A_ABST
Patent Text Reader

Abstract

The invention provides an image tracking model training and image tracking method and device. The method comprises the following steps: acquiring a training sample set and an initial image tracking model; performing down-sampling on the target image sample and the search image sample of the training sample to generate at least one sub-target image sample and a plurality of sub-search image samples; inputting the sub-target image samples into the initial template feature image model to generate template image features, and inputting the sub-search image samples into the initial search model to generate search image sample features; and generating a prediction result based on the template image features and the search image sample features, and training the initial image tracking model based on the prediction result to generate a target image sample tracking model. By down-sampling the target image sample and the search image sample, the effect of subsequent model training is improved, the adaptability of the model to the targets with different scales is improved, and the model is more stable when processing the targets with different sizes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of image tracking processing, and particularly to an image tracking model training, an image tracking method, and an apparatus therefor. Background Art

[0002] Video single-object tracking, as a fundamental research topic in the field of computer vision, is the basis for tasks such as navigation and guidance, video surveillance, and scene understanding, and can be widely applied in military and civilian fields. The task of video single-object tracking is as follows: for a video sequence, first, initialize the tracker according to the position of the target in the initial frame, then extract the target features and establish a target model, adopt a certain tracking strategy in subsequent frames, estimate the position of the target in the current frame based on the target model, and finally update the target model using the current position and continue tracking the next frame. Summary of the Invention

[0003] The present disclosure aims to solve at least one of the technical problems in the related art to some extent.

[0004] To this end, one object of the present disclosure is to provide an image tracking model training method.

[0005] The second object of the present disclosure is to provide an image tracking method.

[0006] The third object of the present disclosure is to provide an image tracking model training apparatus.

[0007] The fourth object of the present disclosure is to provide an image tracking apparatus.

[0008] The fifth object of the present disclosure is to provide an electronic device.

[0009] The sixth object of the present disclosure is to provide a non-transitory computer-readable storage medium.

[0010] The seventh object of the present disclosure is to provide a computer program product.

[0011] To achieve the above object, an embodiment of the first aspect of the present disclosure proposes an image tracking model training method, including: obtaining a training sample set and an initial image tracking model, where the training sample set includes a plurality of training samples, and each training sample includes a target image sample and a corresponding search image sample, and the initial image tracking model includes an initial template feature image model and an initial search model; for any training sample in the training sample set, downsampling the target image sample of the training sample to generate at least one sub-target image sample, and downsampling the search image sample of the training sample to generate a plurality of sub-search image samples; inputting the sub-target image sample into the initial template feature image model to generate a template image feature, and inputting the sub-search image sample into the initial search model to generate a search image sample feature; generating a prediction result based on the template image feature and the search image sample feature, and training the initial image tracking model based on the prediction result to generate a target image sample tracking model.

[0012] According to an embodiment of the present disclosure, the generating a prediction result based on the template image feature and the search image sample feature includes: combining any template image feature and any search image sample feature to generate a training sample combination feature; generating a candidate prediction result based on the training sample combination feature; and generating the prediction result based on the candidate prediction results of all training sample combination features.

[0013] According to an embodiment of the present disclosure, the generating a candidate prediction result based on the training sample combination feature includes: obtaining the true classification result and the true regression result of the target image sample and the search image sample, and generating a fusion feature based on the template image feature and the search image sample feature of the training sample combination feature; generating a model classification result and a model regression result based on the fusion feature; and generating the prediction error based on the model classification result, the model regression result, the true classification result, and the true regression result as the candidate prediction result.

[0014] According to an embodiment of the present disclosure, the training the initial image tracking model based on the prediction result to generate a target image sample tracking model includes: calculating a loss value based on the prediction error; in response to the loss value being greater than a loss threshold, adjusting the model parameters of the initial image tracking model, and selecting a new training sample set from the training sample set without replacement and inputting it into the adjusted initial image tracking model; repeating the above steps of calculating the loss value based on the prediction error and subsequent steps until the training ends, and outputting the target image sample tracking model.

[0015] According to an embodiment of the present disclosure, generating the fusion feature based on the template image feature and the search image sample feature includes: performing i rounds of attention processing on the template image feature and the search image sample feature to output the fusion feature, where each round of attention processing includes performing shunt self-attention processing and shunt mutual-attention processing in sequence, and the input of the nth round of shunt self-attention processing is the output of the (n - 1)th shunt mutual-attention processing, where n is greater than 1 and less than or equal to i.

[0016] According to an embodiment of the present disclosure, downsampling a candidate image to generate a sub-candidate image, where the candidate image is one of a target image sample and a search image sample, and the sub-candidate image is a sub-target image sample when the candidate image is a target image sample, and the sub-candidate image is a sub-search image sample when the candidate image is a search image sample, includes: obtaining a target sampling frequency; downsampling the candidate image based on the target sampling frequency to generate at least one sub-candidate image.

[0017] To achieve the above object, an embodiment of the second aspect of the present disclosure provides an image tracking method, including: obtaining a target image and a search image; inputting the target image and the search image into a target image sample tracking model to output relevant operation features, and tracking the target image based on the relevant operation features, where the target image sample tracking model is trained by the image tracking model training method as described in the embodiment of the first aspect.

[0018] To achieve the above object, an embodiment of the third aspect of the present disclosure provides an image tracking model training device, including: an obtaining module, configured to obtain a training sample set and an initial image tracking model, where the training sample set includes a plurality of training samples, and each training sample includes a target image sample and a corresponding search image sample, and the initial image tracking model includes an initial template feature image model and an initial search model; a generating module, configured to, for any training sample in the training sample set, downsample the target image sample of the training sample to generate at least one sub-target image sample, and downsample the search image sample of the training sample to generate a plurality of sub-search image samples; an input module, configured to input the sub-target image sample into the initial template feature image model to generate a template image feature, and input the sub-search image sample into the initial search model to generate a search image sample feature; a training module, configured to generate a prediction result based on the template image feature and the search image sample feature, and train the initial image tracking model based on the prediction result to generate a target image sample tracking model.

[0019] To achieve the above object, an embodiment of the fourth aspect of the present disclosure provides an image tracking device, including: an acquisition module, configured to acquire a target image and a search image; a tracking module, configured to input the target image and the search image into a target image sample tracking model to output relevant operation features, and track the target image based on the relevant operation features, where the target image sample tracking model is trained by the image tracking model training method as described in the embodiment of the first aspect.

[0020] To achieve the above object, an embodiment of the fifth aspect of the present disclosure provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to implement the image tracking model training method as described in the embodiment of the first aspect of the present disclosure, or to implement the image tracking method as described in the embodiment of the second aspect.

[0021] To achieve the above object, an embodiment of the sixth aspect of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to implement the image tracking model training method as described in the embodiment of the first aspect of the present disclosure, or to implement the image tracking method as described in the embodiment of the second aspect.

[0022] To achieve the above object, an embodiment of the seventh aspect of the present disclosure provides a computer program product, including a computer program, where the computer program, when executed by a processor, is used to implement the image tracking model training method as described in the embodiment of the first aspect of the present disclosure, or to implement the image tracking method as described in the embodiment of the second aspect.

[0023] Therefore, by downsampling the target image samples and the search image samples to generate at least one sub-target image sample and multiple sub-search image samples, and then randomly combining the at least one sub-target image sample and the multiple sub-search image samples, or generating a training sample with more content, the effect of subsequent model training can be improved, and at the same time, the adaptability of the model to targets of different scales can be enhanced, which makes the model more robust when processing targets of different sizes. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is a schematic diagram of an image tracking model training method according to an embodiment of the present disclosure; Figure 2 is a schematic diagram of another image tracking model training method according to an embodiment of the present disclosure; Figure 3 is a schematic diagram of generating a fusion feature based on the template image feature and the search image sample feature according to an embodiment of the present disclosure; Figure 4 is a schematic diagram of the specific implementation of the self-enhancement module based on shunted self-attention according to an embodiment of the present disclosure ; Figure 5 is a schematic flowchart of the shunted mutual attention processing according to an embodiment of the present disclosure Figure 6 is a schematic diagram of an image tracking method according to an embodiment of the present disclosure Figure 7 is a schematic diagram of the implementation effect of a target image sample tracking model according to an embodiment of the present disclosure Figure 8 is a schematic diagram of an image tracking model training device according to an embodiment of the present disclosure Figure 9 is a schematic diagram of an image tracking device according to an embodiment of the present disclosure Figure 10 is a schematic diagram of an electronic device according to an embodiment of the present disclosure Specific Embodiments

[0025] The embodiments of the present disclosure will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present disclosure, but should not be construed as limiting the present disclosure

[0026] In the technical solution of the present disclosure, the acquisition, storage, use, processing, etc. of data all comply with the relevant regulations of relevant laws and regulations

[0027] It should be noted that in the embodiments of the present application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary. The purpose is only to illustrate the feasibility in the implementation of the technical solution of the present application, but it does not mean that the applicant has already or necessarily used this solution

[0028] Figure 1 is a schematic diagram of an image tracking model training method according to an embodiment of the present disclosure. As Figure 1 shown, the image tracking model training method includes the following steps S101. Obtain a training sample set and an initial image tracking model. The training sample set includes a plurality of training samples, each training sample includes a target image sample and a corresponding search image sample, and the initial image tracking model includes an initial template feature image model and an initial search model

[0029] The image tracking model training method of the embodiments of the present application can be applied to the scenario of image tracking and recognition. The execution subject of the image tracking model training of the embodiments of the present application can be the image tracking model training device of the embodiments of the present application, and this image tracking model training device can be set on an electronic device.

[0030] It should be noted that the initial image tracking model in the embodiments of the present disclosure is a model adopting a siamese network structure. The initial image tracking model includes an initial template feature image model and an initial search model.

[0031] The siamese network object tracking algorithm uses a neural network to learn more expressive features through training on a dataset, and applies these features to target localization in the search area. The input of the siamese network for single-object tracking is two pictures, one is the template image of the target to be tracked, and the other is the image of a slightly larger search area. The output is the classification result of the target category in the predicted search area and the regression result of the bounding box. Therefore, the network needs to learn from the data how to distinguish whether there is a target in the search image. If there is, the target needs to be marked with a bounding box.

[0032] The structure of the network is divided into three parts: the backbone network, which is generally modified from the network structure that has achieved excellent performance in the image classification task, and is used to extract the template image features and search image features from the template image and the search image; the fusion network, which is used to further enhance and fuse the template image features and search image features to obtain fused features; the tracking head network, which generally uses some convolutional layers or multi-layer perceptrons to obtain the classification result and the regression result.

[0033] Under the condition of adopting the same backbone network and tracking head network, the currently excellent trackers adopt the Transformer network to design the fusion network. The self-attention mechanism therein is used to enhance the feature representation of the template image features and search image features, and the cross-attention mechanism is used to effectively fuse the two. However, the attention operator participating in the attention mechanism calculation during the enhancement and fusion process is of a single scale, so it cannot effectively capture features of different scales, thus affecting the robustness of the tracker when the scale of the tracked target changes.

[0034] It should be noted that the target image sample refers to the image of the object to be recognized or located. It can be a complete picture or a specific area cropped from the picture. For example, in the target tracking task, the target image sample may be the target object marked in the first frame of the video sequence; in the image retrieval task, it may be the query picture provided by the user.

[0035] A search image sample refers to a set of images in which a target object is searched for. This can be a dataset containing a large number of unlabeled pictures, or a video stream obtained in real time. In practical applications, the search image is the place where the model performance is tested or actual search operations are carried out.

[0036] S102. For any training sample in the training sample set, downsample the target image sample of the training sample to generate at least one sub-target image sample, and downsample the search image sample of the training sample to generate multiple sub-search image samples.

[0037] In actual operations, there may be problems such as a single training sample and a small number of samples, which will affect the training effect and make the finally generated model unable to achieve the expected effect.

[0038] In the embodiments of the present disclosure, by downsampling the target image sample and the search image sample to generate at least one sub-target image sample and multiple sub-search image samples, and then randomly combining at least one sub-target image sample and multiple sub-search image samples, or generating a training sample with more content, the effect of subsequent model training is improved.

[0039] At the same time, by downsampling the target image sample and the search image sample to generate sub-image samples with different resolutions, the adaptability of the model to targets of different scales can be improved, which makes the model more robust when processing targets of different sizes.

[0040] In a possible implementation, different sub-image samples can capture the features of the target at different scales, which helps the model better cope with problems such as changes in target size, occlusion, and background interference.

[0041] S103. Input the sub-target image sample into the initial template feature image model to generate template image features, and input the sub-search image sample into the initial search model to generate search image sample features.

[0042] S104. Generate a prediction result based on the template image features and the search image sample features, and train the initial image tracking model based on the prediction result to generate a target image sample tracking model.

[0043] It should be noted that each training sample corresponds to a true prediction result, which is used to indicate whether the target image sample matches the search image sample, or whether the target image sample exists in the search image sample.

[0044] After obtaining the prediction result generated by the template image features and the search image sample features, the current prediction result can be compared with the true prediction result to determine whether the current prediction result is accurate.

[0045] In a possible way, since the target image sample of the training sample is downsampled to generate at least one sub-target image sample, and the search image sample of the training sample is downsampled to generate multiple sub-search image samples, the prediction results generated based on the template image features and the search image sample features can be multiple. Due to different sampling frequencies or sampling methods, the results of comparing each prediction result with the true prediction result can be multiple, and the multiple comparison results can also be different. Therefore, the richness of the training sample can be further improved, and the training effect of the model can be improved.

[0046] By designing an initial image tracking model based on a shunt attention mechanism fusion network, the multi-scale enhancement and fusion capabilities of the template image features and the search image features are improved. The network consists of an autonomous enhancement module based on shunt self-attention and a cross-fusion module based on shunt mutual-attention. Finally, robust tracking adapting to target scale changes is achieved.

[0047] In the embodiment of the present disclosure, first, a training sample set is obtained, and an initial image tracking model is obtained. The training sample set includes multiple training samples, each training sample includes a target image sample and a corresponding search image sample, and the initial image tracking model includes an initial template feature image model and an initial search model. Then, for any training sample in the training sample set, the target image sample of the training sample is downsampled to generate at least one sub-target image sample, and the search image sample of the training sample is downsampled to generate multiple sub-search image samples. Then, the sub-target image sample is input into the initial template feature image model to generate template image features, and the sub-search image sample is input into the initial search model to generate search image sample features. Finally, a prediction result is generated based on the template image features and the search image sample features, and the initial image tracking model is trained based on the prediction result to generate a target image sample tracking model. By downsampling the target image sample and the search image sample to generate at least one sub-target image sample and multiple sub-search image samples, and then randomly combining the at least one sub-target image sample and the multiple sub-search image samples, or generating a training sample with more content, the training effect of the subsequent model is improved, and at the same time, the adaptability of the model to targets of different scales is improved, which makes the model more robust when processing targets of different sizes.

[0048] In the embodiments of the present disclosure, downsampling is performed on a candidate image to generate a sub-candidate image, where the candidate image is one of a target image sample and a search image sample. When the candidate image is the target image sample, the sub-candidate image is a sub-target image sample, and when the candidate image is the search image sample, the sub-candidate image is a sub-search image sample. First, a target sampling frequency can be obtained, and then the candidate image is downsampled based on the target sampling frequency to generate at least one sub-candidate image.

[0049] It should be noted that the sampling frequency can be pre-designed or randomly generated, and no limitation is made here. The sampling frequency can be multiple or one, and can be specifically limited according to actual design needs.

[0050] In the above embodiments, a prediction result is generated based on the template image feature and the search image sample feature, and it can also be Figure 2 Further explained, the method includes: S201, combining any template image feature and any search image sample feature to generate a training sample combined feature.

[0051] In a possible implementation manner, the target image sample of the training sample is downsampled to generate n sub-target image samples, and the search image sample of the training sample is downsampled to generate m sub-search image samples. Thus, based on any template image feature and any search image sample feature, m×n training sample combined features can be generated.

[0052] S202, generating a candidate prediction result based on the training sample combined feature.

[0053] S203, generating a prediction result based on the candidate prediction results of all training sample combined features.

[0054] It should be noted that to generate a candidate prediction result based on the training sample combined feature, the true classification result and the true regression result of the target image sample and the search image sample can be obtained first, and a fusion feature is generated based on the template image feature and the search image sample feature of the training sample combined feature. Then, a model classification result and a model regression result are generated based on the fusion feature. Finally, a prediction error is generated based on the model classification result, the model regression result, the true classification result, and the true regression result as the candidate prediction result.

[0055] It should be noted that the regression result is data used to identify the corresponding candidate box, and the candidate box is used to frame the object to be recognized or the object to be tracked. It should be noted that the regression result may include the position, size, coordinates, etc. of the candidate box.

[0056] The classification result refers to the result of assigning data points to one or more of the predefined categories. The goal of the classification problem is to predict the output labels based on the input features, which are usually discrete values or categories.

[0057] For example, taking the ResNet50 network, which achieves good performance in the image classification task, as the initial image tracking model, denoted as , and removing the last fully connected layer of the ResNet50 network to meet the requirements of the single-object tracking task. At the same time, in order to make the stride of the backbone network 8, the operations after the last downsampling convolutional layer in this network are removed, and the downsampling of this convolutional layer is cancelled by setting the stride to 1, and it is modified to a dilated convolution to ensure that the size of the output feature map remains unchanged.

[0058] The target image samples and the search image samples extracted from the training set of the single-object tracking dataset are respectively input into two ResNet50 networks with the same structure and shared parameters for feature extraction, and the template image features and the search image features extracted using ResNet50 are output, that is:

[0059] where is a tensor of size , and are the height and width, is the number of feature channels; is a tensor of size , and are the height and width.

[0060] The template image features and the search image features are input into the fusion network based on the shunt attention mechanism, and the relevant operation features after enhancing and fusing the two are output, and the relevant operation network is denoted as , that is:

[0061] where is a tensor of size , and are the height and width, is the number of feature channels.

[0062] The relevant operation features The input tracking head network, denoted as , outputs the candidate prediction results of the entire Siamese network for the target to be tracked in the search image, including the classification result and the regression result .

[0063]

[0064] Among them, is a tensor with a size of , is a tensor with a size of , and are the height and width.

[0065] The classification result predicted by the tracking head network and the regression result are used to calculate the error with the true classification result and the regression result of the target to be tracked in the search image.

[0066] The binary cross-entropy loss function (Binary Cross Entropy Loss) is used to calculate the error of the classification result, and the intersection over union loss function (Intersection Over Union Loss) is used to calculate the error of the regression result. The former is denoted as , and the latter is denoted as . Then the total error is calculated as follows:

[0067] Among them, and are two weight factors, both set to 1 in this disclosure.

[0068] The calculated error is backpropagated using the stochastic gradient descent method to optimize the parameters in the entire Siamese network.

[0069] In the embodiments of this disclosure, the initial image tracking model is trained based on the prediction results to generate the target image sample tracking model. First, the loss value can be calculated based on the prediction error. In response to the loss value being greater than the loss threshold, the model parameters of the initial image tracking model are adjusted, and a new training sample set is selected without replacement from the training sample set and input into the adjusted initial image tracking model. The above steps of calculating the loss value based on the prediction error and subsequent steps are repeated until the training ends, and the target image sample tracking model is output.

[0070] It can be understood that the training of the model is a repetitive iterative process. The model is trained by continuously adjusting the network parameters of the model until the value of the overall loss function of the model is less than the preset value, or the value of the overall loss function of the model no longer changes or changes slowly, and the model converges to obtain a trained model.

[0071] In the embodiments of the present disclosure, as Figure 3 shown, a fusion feature is generated based on the template image feature and the search image sample feature. The template image feature and the search image sample feature can be subjected to i rounds of attention processing to output a fusion feature, where each round of attention processing includes performing shunted self-attention processing and shunted mutual-attention processing in sequence. The input of the nth round of shunted self-attention processing is the output of the (n - 1)th shunted mutual-attention processing, where n is greater than 1 and less than or equal to i.

[0072] In the embodiments of the present disclosure, the process based on shunted self-attention processing includes: Input the template image feature into the self-enhancing module based on shunted self-attention (Self-EnhancingModule), denoted as , to obtain the self-enhanced template image feature, still denoted as ; input the search image feature into the same self-enhancing module based on shunted self-attention , to obtain the self-enhanced search image feature, still denoted as , where the superscript indicates the th self-enhancing module based on shunted self-attention. The formula is expressed as:

[0073] where the specific implementation of the self-enhancing module based on shunted self-attention is as follows, and the implementation details are as Figure 4 shown: A. Query operator calculation for template shunting: The shunted attention mechanism shunts the feature scale but shares the query operator of the same scale. For the template image feature obtained after dimensionality reduction, first deform it to obtain , and then perform feature transformation on the -dimensional feature using a linear layer:

[0074] Here, it is set that the output feature dimension of the linear layer is the same as the input feature dimension, both being . The query operator of the obtained template image feature Considering that the number of heads is 2, perform deformation to obtain:

[0075] where represents the dimension divided by 2 and rounded down.

[0076] Split the query operator of the above template image features along the second dimension for 2 shunts to perform self-attention calculations respectively:

[0077] where, is the splitting operation, , is the template feature query operator of the first shunt, is the template feature query operator of the second shunt.

[0078] B. Calculation of key-value operators for template shunts: The shunt attention mechanism shunts the feature scales and uses key operators and value operators of different scales. For the template image features obtained after dimensionality reduction perform shunt processing: ① The first shunt: First, use a convolutional layer to extract scale features:

[0079] where, is the two-dimensional convolutional layer of scale features for the first shunt, , that is, the scale of this scale is the same as the scale of the query operator of the template feature.

[0080] For the convenience of linear transformation, perform deformation operations on to obtain , and then use a linear layer to perform feature transformation on the -dimensional features:

[0081] Here, set the output feature dimension of the linear layer to be the same as the input feature dimension, both being . Obtain the key-value operator of the template image features of the first shunt.

[0082] Considering that the number of heads is 2, perform deformation to obtain:

[0083] Split the key-value operator of the template image features of the first branch shunt along the first dimension to obtain the key operator and value operator of the first branch shunt:

[0084] Among them, is the splitting operation, , is the template feature key operator of the first branch shunt, is the template feature value operator of the first branch shunt.

[0085] ② Second branch shunt: Similarly, the key-value operator of the second branch shunt can be obtained , is the template feature key operator of the second branch shunt, is the template feature value operator of the second branch shunt.

[0086] C. Calculation of the multi-head self-attention mechanism for template shunt: Perform the multi-head self-attention mechanism calculation on the query operator, key operator, and value operator of the two branch shunts respectively: ① First branch shunt:

[0087] Among them, , is the Softmax function, is the scaling factor, usually set to of the dimension.

[0088] The calculation result of the self-attention mechanism of the first branch shunt is:

[0089] ② Second branch shunt:

[0090] Among them, .

[0091] The calculation result of the second branch shunt is:

[0092] ③ Confluence of shunts: Concatenate and along the feature dimension:

[0093] Among them, is the concatenation operation, 。

[0094] Finally, deform to obtain the multi-scale feature enhancement result of the template image:

[0095] The shunt and confluence complete the feature enhancement of different scales of the template image.

[0096] D. As Figure 4 shown, by analogy with steps ABC, the multi-scale feature enhancement result of the search image can be obtained:

[0097] The shunt and confluence complete the feature enhancement of different scales of the search image.

[0098] In the embodiment of the present disclosure, the process of shunt mutual attention processing can be as Figure 5 shown, including: A. Calculation of the query operator for template shunt: The shunt attention mechanism shunts the feature scales but shares the query operator of the same scale. For the template image features obtained after dimensionality reduction , first deform it to obtain , and then perform feature transformation on the -dimensional features using a linear layer:

[0099] Here, the output feature dimension of the linear layer is set to be the same as the input feature dimension, both being . The query operator of the obtained template image features. Considering that the number of heads is 2, is deformed to obtain:

[0100] where represents dividing the dimension by 2 and rounding down.

[0101] Split the above query operator of the template image features along the second dimension for the two shunts to perform self-attention calculations respectively:

[0102] Among them, is the splitting operation, , is the query operator of the template features of the first shunt, is the query operator of the template features of the second shunt.

[0103] B. Key-value operator calculation for search shunting: The shunting attention mechanism shunts the feature scales and uses key operators and value operators of different scales. For the search image features after dimensionality reduction perform shunting processing: ① The first shunt: First, use a convolutional layer to extract scale features:

[0104] Among them, is the two-dimensional convolutional layer of the scale features of the first shunt, , that is, the scale is the same as the query operator scale of the search feature.

[0105] To facilitate linear transformation, perform a deformation operation on to obtain , and then use a linear layer to perform feature transformation on the -dimensional features:

[0106] Here, set the output feature dimension of the linear layer to be the same as the input feature dimension, both being . The key-value operator of the search image features of the first shunt obtained .

[0107] Considering that the number of heads is 2, perform a deformation on to obtain:

[0108] Split the key-value operator of the search image features of the above first shunt along the first dimension to obtain the key operator and value operator of the first shunt:

[0109] Among them, is the splitting operation, , is the search feature key operator of the first shunt, is the search feature value operator of the first shunt.

[0110] ② The second shunt: Similarly, the key-value operator of the second shunt can be obtained, is the search feature key operator of the second shunt, is the search feature value operator of the second shunt.

[0111] C. Calculation of multi-head mutual attention mechanism of “template-search” split: The multi-head mutual attention mechanism is calculated for the template feature query operator, the search feature key operator, and the search feature value operator of the two branches respectively: ①The first branch:

[0112] in, , is the Softmax function, is the scaling factor, usually set to Dimension.

[0113] The calculation result of the mutual attention mechanism of the first branch is:

[0114] ② The second branch:

[0115] in, .

[0116] The calculation result of the second branch flow is:

[0117] ③Diversion and confluence: Will and Concatenate in the feature dimension:

[0118] in, It is a splicing operation. .

[0119] Finally The template image features and search image features are deformed to obtain multi-scale feature fusion results:

[0120] The split-stream fusion completes the feature fusion of the template image and the search image at different scales.

[0121] D. Figure 5 As shown, by analogy with steps ABC, we can obtain the multi-scale feature fusion result of the search image features and the template image features:

[0122] The split-stream fusion completes the feature fusion of the search image and the template image at different scales.

[0123] In the present invention, a shunt self-attention process is first performed, and then a shunt mutual-attention process is performed to obtain the cross-fused search image features output in the shunt mutual-attention process as the relevant operation features for output:

[0124] wherein, is an assignment operation.

[0125] Figure 6 is a schematic diagram of an image tracking method according to an embodiment of the present disclosure, as Figure 1 shown, the image tracking method includes the following steps: S601, obtaining a target image and a search image.

[0126] It should be noted that the search image in the embodiments of the present disclosure can be one or several images, or a continuous video frame, etc., and no limitation is made here.

[0127] S602, inputting the target image and the search image into the target image sample tracking model to output relevant operation features, and tracking the target image based on the relevant operation features.

[0128] It should be noted that the target image sample tracking model in the embodiments of the present disclosure is trained by the image tracking model training method as in Figures 1-4 the embodiment.

[0129] Through experiments, it is found that the image tracking method designed by the present disclosure can perform high-precision single-object tracking on a 3090 graphics card after training and converging on a single-object tracking dataset, and the method shows strong robustness when the target undergoes scale changes, such as deformation or occlusion. For example, as Figure 7 shown in the image, through the target image sample tracking model in the embodiments of the present disclosure, a deformed bus can be tracked and recognized, and an occluded athlete can also be tracked and recognized.

[0130] Corresponding to the image tracking model training methods provided in the above several embodiments, an embodiment of the present disclosure also provides an image tracking model training device. Since the image tracking model training device provided in the embodiments of the present disclosure corresponds to the image tracking model training methods provided in the above several embodiments, the implementation manners of the above image tracking model training methods are also applicable to the image tracking model training device provided in the embodiments of the present disclosure, and will not be described in detail in the following embodiments.

[0131] Figure 8It is a schematic diagram of an image tracking model training device according to an embodiment of the present disclosure. As shown in FIG. 8, the image tracking model training device 800 includes: an acquisition module 810, a generation module 820, an input module 830, and a training module 840.

[0132] The acquisition module 810 is configured to acquire a training sample set and an initial image tracking model. The training sample set includes a plurality of training samples, and each training sample includes a target image sample and a corresponding search image sample. The initial image tracking model includes an initial template feature image model and an initial search model.

[0133] The generation module 820 is configured to, for any training sample in the training sample set, downsample the target image sample of the training sample to generate at least one sub-target image sample, and downsample the search image sample of the training sample to generate a plurality of sub-search image samples.

[0134] The input module 830 is configured to input the sub-target image sample into the initial template feature image model to generate a template image feature, and input the sub-search image sample into the initial search model to generate a search image sample feature.

[0135] The training module 840 is configured to generate a prediction result based on the template image feature and the search image sample feature, and train the initial image tracking model based on the prediction result to generate a target image sample tracking model.

[0136] According to an embodiment of the present disclosure, generating a prediction result based on the template image feature and the search image sample feature includes: combining any template image feature and any search image sample feature to generate a training sample combination feature; generating a candidate prediction result based on the training sample combination feature; and generating a prediction result based on the candidate prediction results of all training sample combination features.

[0137] According to an embodiment of the present disclosure, generating a candidate prediction result based on the training sample combination feature includes: obtaining the true classification result and the true regression result of the target image sample and the search image sample, and generating a fusion feature based on the template image feature and the search image sample feature of the training sample combination feature; generating a model classification result and a model regression result based on the fusion feature; and generating a prediction error based on the model classification result, the model regression result, the true classification result, and the true regression result as the candidate prediction result.

[0138] According to an embodiment of the present disclosure, training an initial image tracking model based on prediction results to generate a target image sample tracking model includes: calculating a loss value based on a prediction error; in response to the loss value being greater than a loss threshold, adjusting model parameters of the initial image tracking model, and selecting a new training sample set without replacement from a training sample set and inputting the new training sample set into the adjusted initial image tracking model; repeating the above steps of calculating the loss value based on the prediction error and subsequent steps until the training ends, and outputting the target image sample tracking model.

[0139] According to an embodiment of the present disclosure, generating a fusion feature based on a template image feature and a search image sample feature includes: performing i rounds of attention processing on the template image feature and the search image sample feature to output a fusion feature, where each round of attention processing includes successively performing shunt self-attention processing and shunt mutual-attention processing, and the input of the nth round of shunt self-attention processing is the output of the (n - 1)th shunt mutual-attention processing, where n is greater than 1 and less than or equal to i.

[0140] According to an embodiment of the present disclosure, downsampling a candidate image to generate a sub-candidate image, where the candidate image is one of a target image sample and a search image sample, and the sub-candidate image is a sub-target image sample when the candidate image is the target image sample, and the sub-candidate image is a sub-search image sample when the candidate image is the search image sample, includes: obtaining a target sampling frequency; downsampling the candidate image based on the target sampling frequency to generate at least one sub-candidate image.

[0141] Thus, by downsampling the target image sample and the search image sample to generate at least one sub-target image sample and multiple sub-search image samples, and then randomly combining the at least one sub-target image sample and the multiple sub-search image samples, or generating a training sample with more content, the effect of subsequent model training can be improved, and at the same time, the adaptability of the model to targets of different scales can be enhanced, which makes the model more robust when processing targets of different sizes.

[0142] Corresponding to the image tracking methods provided in the above several embodiments, an embodiment of the present disclosure also provides an image tracking device. Since the image tracking device provided in the embodiment of the present disclosure corresponds to the image tracking methods provided in the above several embodiments, the implementation manners of the above image tracking methods are also applicable to the image tracking device provided in the embodiment of the present disclosure and will not be described in detail in the following embodiments.

[0143] Figure 9 is a schematic diagram of an image tracking device according to an embodiment of the present disclosure, as Figure 9 shown, the image tracking device 900 includes: An acquisition module 910, configured to acquire a target image and a search image.

[0144] A tracking module 920 is configured to input a target image and a search image into a target image sample tracking model to output relevant operation features, and track the target image based on the relevant operation features.

[0145] To implement the above embodiments, an electronic device 1000 is further proposed in an embodiment of the present disclosure. Figure 10 It is a schematic diagram of an electronic device according to an embodiment of the present disclosure, as Figure 10 shown. The electronic device 1000 includes: a processor 1001 and a memory 1002 communicatively connected to the processor. The memory 1002 stores instructions executable by at least one processor. The instructions are executed by at least one processor 1001 to implement, as in the present disclosure Figures 1-5 the image tracking model training method of the embodiment, or as in Figure 6 the embodiment and Figure 7 the image tracking method of the embodiment.

[0146] To implement the above embodiments, a non-transitory computer-readable storage medium storing computer instructions is further proposed in an embodiment of the present disclosure. The computer instructions are used to cause a computer to implement, as in the present disclosure Figures 1-5 the image tracking model training method of the embodiment, or as in Figure 6 the embodiment and Figure 7 the image tracking method of the embodiment.

[0147] To implement the above embodiments, a computer program product is further proposed in an embodiment of the present disclosure, including a computer program. When the computer program is executed by a processor, it implements, as in the present disclosure Figures 1-5 the image tracking model training method of the embodiment, or as in Figure 6 the embodiment and Figure 7 the image tracking method of the embodiment.

[0148] It should be noted that personal information from users should be collected for legal and reasonable purposes and not shared or sold outside of these legal uses. In addition, such collection / sharing should be carried out after obtaining the informed consent of the user, including but not limited to notifying the user to read the user agreement / user notice before the user uses the function and signing an agreement / authorization including authorizing the relevant user information. In addition, any necessary steps should be taken to protect and safeguard access to such personal information data and ensure that others with access to the personal information data comply with their privacy policies and procedures.

[0149] This application is expected to provide an implementation scheme for users to selectively prevent the use or access of personal information data. That is, the present disclosure is expected to provide hardware and / or software to prevent or block access to such personal information data. Once the personal information data is no longer needed, the risk can be minimized by restricting data collection and deleting the data. In addition, when applicable, personal identifiers are removed from such personal information to protect the privacy of users.

[0150] In the description of the foregoing embodiments, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0151] In addition, the terms "first" and "second" are used only for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0152] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred implementation of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present application belong.

[0153] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered as a definitional sequence list of executable instructions for implementing logical functions, and can be embodied specifically in any computer-readable medium for use by an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions, or in connection with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that contains, stores, communicates, propagates, or transports a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.

[0154] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or combinations thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0155] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of the above-described embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0156] In addition, each functional unit in various embodiments of the present application may be integrated into a processing module, may exist independently physically for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0157] The above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A method for training an image tracking model, characterized in that Including: Obtain a training sample set and an initial image tracking model. The training sample set includes a plurality of training samples, each training sample including a target image sample and a corresponding search image sample. The initial image tracking model includes an initial template feature image model and an initial search model; For any training sample in the training sample set, downsample the target image sample of the training sample to generate at least one sub-target image sample, and downsample the search image sample of the training sample to generate multiple sub-search image samples; Input the sub-target image sample into the initial template feature image model to generate a template image feature, and input the sub-search image sample into the initial search model to generate a search image sample feature; Generate a prediction result based on the template image feature and the search image sample feature, and train the initial image tracking model based on the prediction result to generate a target image sample tracking model.

2. The method according to claim 1, wherein The generating the prediction result based on the template image feature and the search image sample feature includes: Combine any template image feature and any search image sample feature to generate a training sample combined feature; Generate a candidate prediction result based on the training sample combined feature; Generate the prediction result based on the candidate prediction results of all training sample combined features.

3. The method according to claim 2, wherein The generating the candidate prediction result based on the training sample combined feature includes: Obtain the true classification result and the true regression result of the target image sample and the search image sample, and generate a fusion feature based on the template image feature and the search image sample feature of the training sample combined feature; Generate a model classification result and a model regression result based on the fusion feature; Generate the prediction error based on the model classification result, the model regression result, the true classification result, and the true regression result as the candidate prediction result.

4. The method according to claim 3, wherein The training the initial image tracking model based on the prediction result to generate a target image sample tracking model includes: Calculate a loss value based on the prediction error; In response to the loss value being greater than a loss threshold, adjust the model parameters of the initial image tracking model, and select a new training sample set from the training sample set without replacement and input it into the adjusted initial image tracking model; Repeat the above steps of calculating the loss value based on the prediction error and subsequent steps until the training ends, and output the target image sample tracking model.

5. The method according to claim 3, wherein The generating the fusion feature based on the template image feature and the search image sample feature of the training sample combined feature includes: Perform i rounds of attention processing on the template image feature and the search image sample feature to output the fusion feature, where each round of attention processing includes performing split self-attention processing and split mutual-attention processing in sequence. The input of the nth round of split self-attention processing is the output of the (n - 1)th split mutual-attention processing, and n is greater than 1 and less than or equal to i.

6. The method according to claim 1, characterized in that, Downsample the candidate image to generate a sub-candidate image, where the candidate image is one of a target image sample and a search image sample, the sub-candidate image is a sub-target image sample when the candidate image is a target image sample, and the sub-candidate image is a sub-search image sample when the candidate image is a search image sample, including: Obtain a target sampling frequency; Downsample the candidate image based on the target sampling frequency to generate at least one sub-candidate image.

7. An image tracking method, characterized in that, Including: Obtain a target image and a search image; Input the target image and the search image into a target image sample tracking model to output relevant operation features, and track the target image based on the relevant operation features, where the target image sample tracking model is trained by the image tracking model training method described in any one of claims 1-6.

8. An image tracking model training device, characterized in that, Including: An acquisition module for acquiring a training sample set and an initial image tracking model, where the training sample set includes multiple training samples, each training sample includes a target image sample and a corresponding search image sample, and the initial image tracking model includes an initial template feature image model and an initial search model; A generation module for, for any training sample in the training sample set, downsampling the target image sample of the training sample to generate at least one sub-target image sample, and downsampling the search image sample of the training sample to generate multiple sub-search image samples; An input module for inputting the sub-target image sample into the initial template feature image model to generate template image features, and inputting the sub-search image sample into the initial search model to generate search image sample features; A training module for generating a prediction result based on the template image features and the search image sample features, and training the initial image tracking model based on the prediction result to generate a target image sample tracking model.

9. An image tracking device, characterized in that, Including: An acquisition module for acquiring a target image and a search image; A tracking module for inputting the target image and the search image into a target image sample tracking model to output relevant operation features, and tracking the target image based on the relevant operation features, where the target image sample tracking model is trained by the image tracking model training method described in any one of claims 1-6.

10. An electronic device, characterized in that, Including a memory and a processor; Wherein, the processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to be used to implement the image tracking model training method described in any one of claims 1-6, or implement the image tracking method described in claim 7.

Citation Information

Patent Citations

  • Target tracking method and system and storage medium

    CN115100235A

  • Unmanned aerial vehicle target tracking method based on space-time memory network

    CN116630369A