A target tracking method, device, electronic device, program product and medium

By fusing text features with image features and processing them with pre-trained decoder and feature interactor, the problem that visual image features are susceptible to noise is solved, and the target tracking performance is significantly improved.

CN119625746BActive Publication Date: 2025-06-20LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510162628.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-20
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

In existing end-to-end detection solutions, visual image features are easily affected by noise such as scale changes, deformations, and lighting changes, resulting in reduced tracking performance.

Method used

By fusing text features with image features, the pre-trained decoder and feature interactor are used to decode and enhance the global image features, enhanced detection sequence vectors and tracking sequence vectors to improve the target tracking performance.

Benefits of technology

It significantly improves the overall performance of multi-objective tracking, enhances the robustness of the category and attribute descriptions of the tracked objects, and reduces the impact of noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625746B_ABST
    Figure CN119625746B_ABST
Patent Text Reader

Abstract

The present invention provides a target tracking method, device, electronic device, program product and medium, which relates to the field of artificial intelligence. In this method, first, a detection sequence vector and a tracking sequence vector used in multi-target tracking can be obtained. The detection sequence vector records the detection information of the object to be tracked and is used to detect the object to be tracked; the tracking sequence vector records the image features of the tracked object and is used to continue tracking the tracked object. Subsequently, the detection sequence vector can be enhanced by using the class description text of the object to be tracked, and the tracking sequence vector can be enhanced by using the attribute description text of the object to be tracked, so as to incorporate text semantic features into the above sequence vectors. Considering that the text semantic features have strong high-order semantic description capabilities and are completely unaffected by noises such as scale changes, deformations, and illumination changes, the performance of multi-target tracking can be improved after the text semantic enhancement of the detection sequence vector and the tracking sequence vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly relates to a target tracking method, device, electronic device, program product, and medium. Background Art

[0002] In related technologies, end-to-end detection is a basic paradigm for solving the multi-target tracking problem. This method can integrate detection and tracking, and can learn to model the long-term temporal changes of target objects, implicitly perform temporal association, which is beneficial to the improvement of the overall performance of the tracking system. However, in related end-to-end detection schemes, visual image features are generally used for detection and tracking, and visual image features are easily affected by noise such as scale changes, deformations, and illumination changes, which may easily lead to a decrease in tracking performance. Summary of the Invention

[0003] The purpose of the present invention is to provide a target tracking method, device, electronic device, program product, and medium, which can improve tracking detection by fusing text features and image features to enhance target tracking performance.

[0004] To solve the above technical problems, the present invention provides a target tracking method, including:

[0005] Obtain the global image feature of the image to be detected, the category description text corresponding to the object to be tracked, the detection sequence vector, and the tracking sequence vector; the detection sequence vector records the detection information of the object to be tracked, and the tracking sequence vector records the image features of the tracked object;

[0006] Fuse the category description text and the detection sequence vector into an enhanced detection sequence vector;

[0007] Use a pre-trained decoder to jointly decode the global image feature, the enhanced detection sequence vector, and the tracking sequence vector to obtain the tracking information of the object to be tracked in the image to be detected; the tracking information includes at least position information and image features;

[0008] Update the tracking sequence vector using the tracking information, and use a pre-trained feature interactor and the attribute description text corresponding to the object to be tracked to enhance the updated tracking sequence vector, so as to continue target tracking using the enhanced tracking sequence vector.

[0009] Optionally, fusing the category description text and the detection sequence vector into an enhanced detection sequence vector includes:

[0010] Encode the category description text to obtain a category description text vector;

[0011] Concatenate the category description text vector and the detection sequence vector, and perform linear transformation processing and normalization processing on the concatenation result in sequence to obtain an enhanced detection sequence vector.

[0012] Optionally, the pre-trained feature interactor includes a first mapping unit, a second mapping unit, pre-trained text features, and pre-trained image features;

[0013] Use the pre-trained feature interactor and the attribute description text corresponding to the object to be tracked to enhance the updated tracking sequence vector, including:

[0014] Obtain text features, and extract the image features of the tracked object from the tracking sequence vector; among them, the text features are generated using the attribute description text;

[0015] Use the first mapping unit to project the pre-trained text features into mapped image features, and use the second mapping unit to project the pre-trained image features into mapped text features;

[0016] Fuse the text features, pre-trained text features, and mapped text features to obtain fused text features, and use the pre-trained text-image contrast learning model to encode the fused text features to obtain encoded text features;

[0017] Fuse the image features, pre-trained image features, and mapped image features to obtain fused image features, and use the pre-trained text-image contrast learning model to encode the fused image features to obtain encoded image features;

[0018] Perform cross-attention mechanism processing on the encoded text features and the encoded image features to obtain an enhanced tracking sequence vector.

[0019] Optionally, both the first mapping unit and the second mapping unit are multi-layer perceptrons.

[0020] Optionally, obtaining text features includes:

[0021] According to the category description text, obtain the text features of the category corresponding to the category description text in the tracking text prompt library; the tracking text prompt library contains the text features corresponding to various objects to be tracked.

[0022] Optionally, it further includes:

[0023] Obtain the attribute description text corresponding to the object to be tracked;

[0024] Encode the attribute description text to obtain text features;

[0025] Establish a corresponding relationship between the text features and the category of the object to be tracked, and save the corresponding relationship to the tracking text prompt library.

[0026] Optionally, the pre-trained feature interactor includes a first mapping unit and pre-trained text features;

[0027] Enhance the updated tracking sequence vector by using the pre-trained feature interactor and the attribute description text corresponding to the object to be tracked, including:

[0028] Obtain text features and extract the image features of the tracked object from the tracking sequence vector; among them, the text features are generated using the attribute description text;

[0029] Project the pre-trained text features into mapped image features using the first mapping unit;

[0030] Fuse the text features and the pre-trained text features to obtain fused text features, and encode the fused text features using the pre-trained text-image contrast learning model to obtain encoded text features;

[0031] Fuse the image features and the mapped image features to obtain fused image features, and encode the fused image features using the pre-trained text-image contrast learning model to obtain encoded image features;

[0032] Perform cross-attention mechanism processing on the encoded text features and the encoded image features to obtain an enhanced tracking sequence vector.

[0033] Optionally, the pre-trained feature interactor includes a second mapping unit and pre-trained image features;

[0034] Enhance the updated tracking sequence vector by using the pre-trained feature interactor and the attribute description text corresponding to the object to be tracked, including:

[0035] Obtain text features and extract the image features of the tracked object from the tracking sequence vector; among them, the text features are generated using the attribute description text;

[0036] Project the pre-trained image features into mapped text features using the second mapping unit;

[0037] Fuse the text features and the mapped text features to obtain fused text features, and encode the fused text features using the pre-trained text-image contrast learning model to obtain encoded text features;

[0038] Fuse the image features and the pre-trained image features to obtain fused image features, and encode the fused image features using the pre-trained text-image contrast learning model to obtain encoded image features;

[0039] Perform cross-attention mechanism processing on the encoded text features and the encoded image features to obtain an enhanced tracking sequence vector.

[0040] Optionally, obtain the global image features of the image to be detected, including:

[0041] Use a pre-trained convolutional neural network model to extract features from the image to be detected to obtain primary image features;

[0042] According to the spatial position corresponding to the primary image features in the image to be detected, perform spatial position encoding on the primary image features to obtain encoded primary image features;

[0043] Use the encoded primary image features as the first key value and the first query value, use the primary image features as the first numerical value, and use the pre-trained image feature encoder to perform self-attention mechanism processing on the first key value, the first query value, and the first numerical value to obtain a first processing result, and use the pre-trained image feature encoder to perform forward propagation calculation on the first processing result to obtain global image features.

[0044] Optionally, use a pre-trained decoder to jointly decode the global image features, the enhanced detection sequence vector, and the tracking sequence vector to obtain the tracking information of the object to be tracked in the image to be detected, including:

[0045] Concatenate the enhanced detection sequence vector and the tracking sequence vector to obtain an object sequence vector;

[0046] According to the spatial position corresponding to the global image features in the image to be detected, perform spatial position encoding on the global image features to obtain encoded global image features;

[0047] Use the pre-trained decoder to perform self-attention mechanism processing on the object sequence vector to obtain a second processing result;

[0048] Use the second processing result, the encoded global image features, and the global image features as the second query value, the second key value, and the second numerical value in sequence, use the pre-trained decoder to perform attention mechanism processing on the second query value, the second key value, and the second numerical value to obtain a third processing result, and use the pre-trained decoder to perform forward propagation calculation on the third processing result to obtain tracking information.

[0049] Optionally, the tracking information further includes the tracking status category of the object to be tracked and the confidence corresponding to the tracking status category;

[0050] Use the tracking information to update the tracking sequence vector, including:

[0051] If it is determined that the object to be tracked belongs to a newly added tracked object according to the tracking status category, when the confidence of the object to be tracked is greater than the first preset threshold, add the image features of the object to be tracked to the tracking sequence vector;

[0052] If it is determined that the object to be tracked belongs to the tracked object according to the tracking status category, when the confidence of the object to be tracked is less than the second preset threshold, the image features of the object to be tracked are removed from the tracking sequence vector, or when the confidence of the object to be tracked is not less than the second preset threshold, the tracking sequence vector is updated using the image features of the object to be tracked.

[0053] Optionally, it further includes:

[0054] Obtain multiple training images, the category description text corresponding to the training tracking object, the initial detection sequence vector and the tracking sequence vector, and extract the training image features of the training images; the true position information of the training tracking object is marked in the training images, and the initial tracking sequence vector is empty;

[0055] Randomly initialize the initial detection sequence vector, and fuse the category description text with the initial detection sequence vector into the initial enhanced detection sequence vector;

[0056] Use the initial enhanced detection sequence vector, the tracking sequence vector, the training image features, the initial decoder and the initial feature interactor to determine the predicted position information of the training tracking object in each training image;

[0057] Use the preset loss function to calculate the loss between the predicted position information and the true position information to obtain the loss value;

[0058] Use the loss value to update the parameters of the initial detection sequence vector, the initial decoder, and the initial feature interactor to obtain the detection sequence vector, the pre-trained decoder, and the pre-trained feature interactor.

[0059] The present invention also provides an object tracking device, including:

[0060] An acquisition module, configured to acquire the global image features of the image to be detected, the category description text corresponding to the object to be tracked, the detection sequence vector and the tracking sequence vector; the detection sequence vector records the detection information of the object to be tracked, and the tracking sequence vector records the image features of the tracked object;

[0061] A detection sequence enhancement module, configured to fuse the category description text with the detection sequence vector into an enhanced detection sequence vector;

[0062] A tracking and detection module, configured to jointly decode the global image features, the enhanced detection sequence vector, and the tracking sequence vector using the pre-trained decoder to obtain the tracking information of the object to be tracked in the image to be detected; the tracking information includes at least position information and image features;

[0063] A tracking sequence enhancement module, which is used to update the tracking sequence vector by using tracking information, and enhance the updated tracking sequence vector by using a pre-trained feature interactor and the attribute description text corresponding to the object to be tracked, so as to continue target tracking by using the enhanced tracking sequence vector.

[0064] The present invention also provides an electronic device, including:

[0065] A memory for storing a computer program;

[0066] A processor for implementing the above target tracking method when executing the computer program.

[0067] The present invention also provides a computer program product, including a computer program or instruction, and the computer program or instruction implements the above target tracking method when executed by a processor.

[0068] The present invention also provides a non-volatile computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are loaded and executed by a processor, the above target tracking method is implemented.

[0069] The present invention provides a target tracking method, including: obtaining the global image feature of the image to be detected, the category description text corresponding to the object to be tracked, the detection sequence vector and the tracking sequence vector; the detection sequence vector records the detection information of the object to be tracked, and the tracking sequence vector records the image feature of the tracked object; fusing the category description text with the detection sequence vector into an enhanced detection sequence vector; using a pre-trained decoder to jointly decode the global image feature, the enhanced detection sequence vector and the tracking sequence vector to obtain the tracking information of the object to be tracked in the image to be detected; the tracking information at least includes position information and image feature; updating the tracking sequence vector by using the tracking information, and enhancing the updated tracking sequence vector by using a pre-trained feature interactor and the attribute description text corresponding to the object to be tracked, so as to continue target tracking by using the enhanced tracking sequence vector.

[0070] The beneficial effects of the present invention are as follows: First, the present invention can obtain the global image features of the image to be detected, the category description text corresponding to the object to be tracked, the detection sequence vector, and the tracking sequence vector. Among them, the category description text is used to describe the category of the object to be tracked; the detection sequence vector records the detection information of the object to be tracked and is used to detect the object to be tracked; the tracking sequence vector records the image features of the tracked object and is used to continue tracking the tracked object. Subsequently, the present invention can fuse the category description text with the detection sequence vector into an enhanced detection sequence vector, that is, it can use the text semantic features to enhance the detection sequence vector. Subsequently, the present invention can use a pre-trained decoder to jointly decode the global image features, the enhanced detection sequence vector, and the tracking sequence vector to obtain the tracking information of the object to be tracked in the image to be detected, and the tracking information at least includes position information and image features. Subsequently, the present invention can use the tracking information to update the tracking sequence vector, and use the pre-trained feature interaction device and the attribute description text corresponding to the object to be tracked to enhance the updated tracking sequence vector, that is, it can also use the text semantic features to enhance the tracking sequence vector to continue target tracking using the enhanced tracking sequence vector. Simply put, the present invention can use the category description text of the object to be tracked to enhance the detection sequence vector, and use the attribute description text of the object to be tracked to enhance the tracking sequence vector. Considering that the text semantic features have strong high-order semantic description capabilities and are completely unaffected by noise such as scale changes, deformations, and illumination changes, the overall performance of multi-object tracking can be significantly enhanced. The present invention also provides an object tracking device, an electronic device, a computer-readable storage medium, and a computer program product, which have the above beneficial effects. Description of the Drawings

[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0072] Figure 1 It is a flowchart of an object tracking method provided by an embodiment of the present invention;

[0073] Figure 2 It is a schematic diagram of a pre-trained image feature encoder provided by an embodiment of the present invention;

[0074] Figure 3 It is a schematic diagram of the generation process of an enhanced detection sequence vector provided by an embodiment of the present invention;

[0075] Figure 4Schematic diagram of a pre-trained decoder provided by an embodiment of the present invention;

[0076] Figure 5 Schematic diagram of a method for enhancing tracking sequence vectors based on text-image bidirectional prompt learning provided by an embodiment of the present invention;

[0077] Figure 6 Schematic diagram of a method for enhancing tracking sequence vectors based on text-image unidirectional prompt learning provided by an embodiment of the present invention;

[0078] Figure 7 Schematic diagram of a method for enhancing tracking sequence vectors based on image-text unidirectional prompt learning provided by an embodiment of the present invention;

[0079] Figure 8 Schematic diagram of a system provided by an embodiment of the present invention;

[0080] Figure 9 Block diagram of the structure of a target tracking device provided by an embodiment of the present invention;

[0081] Figure 10 Block diagram of the structure of an electronic device provided by an embodiment of the present invention. Specific embodiments

[0082] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0083] In related technologies, end-to-end detection is a basic paradigm for solving the multi-object tracking problem. This method can integrate detection and tracking, and can learn and model the long-term changes of target objects, implicitly perform temporal association, which is beneficial to improving the overall performance of the tracking system. For example, the MOTR framework (End-to-End Multiple-object Tracking with Transformer) is an end-to-end multi-object tracking framework based on the Transformer. However, in related solutions, visual image features are generally used for tracking detection, and visual image features are easily affected by noise such as scale changes, deformations, and illumination changes, which may easily lead to a decrease in tracking performance.

[0084] In view of this, regarding the technical problem of how to improve the performance of multi-object tracking, the present invention takes into account that text semantic features have strong high-order semantic description capabilities and are completely unaffected by noises such as scale changes, deformations, and illumination changes. Therefore, a target tracking method can be provided, which can improve tracking detection by fusing text features and image features to enhance the target tracking performance.

[0085] The target tracking method provided by the present invention will be introduced below. For ease of understanding, please refer to Figure 1 , Figure 1 which is a flowchart of a target tracking method provided by an embodiment of the present invention. This method may include:

[0086] S101. Obtain the global image features of the image to be detected, the category description text corresponding to the object to be tracked, the detection sequence vector, and the tracking sequence vector; the detection sequence vector records the detection information of the object to be tracked, and the tracking sequence vector records the image features of the object that has been tracked.

[0087] In this step, the global image features, the category description text, the detection sequence vector, and the tracking sequence vector can be obtained first to use these data as model inputs. The global image features, the category description text, the detection sequence vector, and the tracking sequence vector will be introduced below.

[0088] The image to be detected is an image that needs to be subjected to target tracking detection, and it can come from an image sequence, such as a video stream. The image to be detected may contain object objects, and these object objects may be the objects to be tracked. For example, the present invention can perform tracking detection on objects such as people, animals, and vehicles in the image to be detected. To achieve the tracking effect, it is necessary to extract the global image features from the image to be detected and use the global image features for tracking detection. Among them, the global image features are the overall image features of the image to be detected, which cover all the contents of the image to be detected.

[0089] The category description text is the text content used to describe the category of the object to be tracked. For example, when the object to be tracked is a person, the category description text may be "Person"; when the object to be tracked is a vehicle, the category description text may be "Car".

[0090] The detection sequence vector (detect query) is a learnable vector used to detect the object to be tracked, and it can contain multiple vector parameters, and these multiple vector parameters can be used to record the detection information of multiple objects to be tracked. For example, one vector parameter corresponds to a potential category of the object to be tracked and records the detection information of the object to be tracked. Belonging to the category of learnable vectors means that the detection sequence vector needs to be updated with parameters during the model training process to meet the tracking detection requirements.

[0091] The track query is a vector used to continue tracking a tracked object, which records the image features of the tracked object. The tracked object refers to an object that has been detected and tracked in a previous image. When performing the first tracking detection, the track query is empty.

[0092] It should be noted that this embodiment does not limit the extraction method of global image features, which can be set according to actual application requirements. For example, a traditional convolutional neural network model (CNN, Convolutional Neural Networks) can be used for feature extraction. To enhance the correlation between image features, on the basis of using a convolutional neural network model, an image feature encoder containing a self-attention mechanism module (such as a multi-head attention module) can be introduced, and the primary image features output by the convolutional neural network model can be further processed based on the self-attention mechanism (Self-Attention) in the image feature encoder.

[0093] Based on this, obtaining the global image features of the image to be detected may include:

[0094] Step 11: Use a pre-trained convolutional neural network model to extract features from the image to be detected to obtain primary image features.

[0095] In this step, the pre-trained convolutional neural network model refers to a convolutional neural network model that has been trained. This embodiment does not limit the specific type of the pre-trained convolutional neural network model. For example, it can be a RestNet series model, a Darknet series model, etc.

[0096] Step 12: Perform spatial position encoding on the primary image features according to the spatial positions corresponding to the primary image features in the image to be detected to obtain encoded primary image features.

[0097] In this step, considering that each primary image feature is generated according to the image content at each spatial position in the image to be detected; and there is a correlation between the image contents at different spatial positions. For example, the image contents at multiple spatial positions in multiple images to be detected form a complete object to be tracked, that is, there is also a correlation between the primary image features at different spatial positions. Therefore, to ensure that the image feature encoder can learn the correlation relationship between the primary image features, the primary image features can be spatially position-encoded according to the spatial positions corresponding to the primary image features in the image to be detected to obtain encoded primary image features.

[0098] It should be noted that this embodiment does not limit the specific process of spatial position encoding, and relevant technologies in artificial intelligence can be referred to.

[0099] Step 13: Use the pre-trained image feature encoder to perform self-attention mechanism processing on the first key value, the first query value, and the first numerical value with the encoded primary image features as the first key value and the first query value, and the primary image features as the first numerical value, and obtain the first processing result. Then, use the pre-trained image feature encoder to perform forward propagation calculation on the first processing result to obtain the global image features.

[0100] In this step, the pre-trained image feature encoder refers to an image feature encoder that has been trained and is at least provided with a self-attention mechanism module and a forward propagation module (Feedforward Neural Network), which can perform self-attention mechanism processing and forward propagation processing on the encoded primary image features. Specifically, the encoded primary image features can be used as the first key value (K, Key) and the first query value (Q, Query), and the primary image features can be used as the first numerical value (V, Value), and the pre-trained image feature encoder is used to perform self-attention mechanism processing on the first key value, the first query value, and the first numerical value to obtain the first processing result. Subsequently, the pre-trained image feature encoder can be used to perform forward propagation calculation on the first processing result to obtain the global image features.

[0101] For ease of understanding, please refer to Figure 2 , Figure 2 which is a schematic diagram of a pre-trained image feature encoder provided by an embodiment of the present invention. In this schematic diagram, the structure outlined by the dashed box is the Transformer encoder in the pre-trained image feature encoder. The Transformer encoder can include a multi-head attention calculation module (MHA, Multi-head Attention, that is, the self-attention mechanism module), an addition normalization module, and a forward propagation module. The pre-trained image feature encoder can be formed by connecting N Transformer encoders in series, where N≥1. The high-level image features are the global image features.

[0102] S102: Fuse the category description text with the detection sequence vector to obtain an enhanced detection sequence vector.

[0103] In this step, the category description text can be fused with the detection sequence vector to incorporate the text semantic information corresponding to the category of the object to be tracked into the detection sequence vector, thereby obtaining an enhanced detection sequence vector. In this way, when using the enhanced detection sequence vector for detection and tracking, the model can well perceive the category information of the object to be tracked, thereby improving the detection and tracking performance.

[0104] It should be noted that this embodiment does not limit how to fuse the category description text and the detection sequence vector into an enhanced detection sequence vector, which can be determined according to actual application requirements. For example, first, the category description text can be encoded. For example, the WordPiece tokenization method can be used to encode the category description text to obtain a category description text vector. Subsequently, the category description text vector and the detection sequence vector can be concatenated, and the concatenated result can be sequentially subjected to a linear transformation process and a normalization process to obtain a fused detection sequence vector, so as to ensure better fusion of the category description text vector and the detection sequence vector.

[0105] Based on this, fusing the category description text and the detection sequence vector into an enhanced detection sequence vector may include:

[0106] Step 21: Encode the category description text to obtain a category description text vector;

[0107] Step 22: Concatenate the category description text vector and the detection sequence vector, and sequentially perform a linear transformation process and a normalization process on the concatenated result to obtain an enhanced detection sequence vector.

[0108] It should be noted that this embodiment does not limit the specific processes of the above-mentioned encoding process, linear transformation process, and normalization process, and relevant technologies in the field of artificial intelligence can be referred to. For easy understanding, please refer to Figure 3 , Figure 3 which is a schematic diagram of the generation process of an enhanced detection sequence vector provided by an embodiment of the present invention. Among them, "text description" is the attribute description text, "detect query" is the detection sequence vector, and "detect query based on text prompt" is the enhanced detection sequence vector.

[0109] S103. Use a pre-trained decoder to jointly decode the global image feature, the enhanced detection sequence vector, and the tracking sequence vector to obtain tracking information of the object to be tracked in the image to be detected; the tracking information includes at least position information and image features.

[0110] In this step, a pre-trained decoder can be used to jointly decode the global image feature, the enhanced detection sequence vector, and the tracking sequence vector to obtain tracking information of the object to be tracked in the image to be detected. The tracking information includes at least the position information and image features corresponding to the object to be tracked in the image to be detected. That is, the pre-trained decoder is used to perform target tracking detection according to the input.

[0111] It should be noted that this embodiment does not limit the specific process of jointly decoding the global image features, enhanced detection sequence vectors, and tracking sequence vectors using the pre-trained decoder. For example, first, the enhanced detection sequence vector and the tracking sequence vector can be concatenated to obtain an object sequence vector (object query), and the pre-trained decoder can be used to perform self-attention mechanism processing on the object sequence vector to obtain a second processing result. Subsequently, the pre-trained decoder can be used to continue performing attention mechanism processing on the above second processing result and the global image features to obtain a third processing result, and forward propagation calculation can be performed on the third processing result to obtain tracking information.

[0112] Based on this, jointly decoding the global image features, enhanced detection sequence vectors, and tracking sequence vectors using the pre-trained decoder to obtain the tracking information of the object to be tracked in the image to be detected may include:

[0113] Step 31: Concatenate the enhanced detection sequence vector and the tracking sequence vector to obtain an object sequence vector.

[0114] Step 32: According to the spatial position corresponding to the global image features in the image to be detected, perform spatial position encoding on the global image features to obtain encoded global image features.

[0115] In this step, the description of performing spatial position encoding on the global image features is similar to the description of performing spatial position encoding on the processed image features, and will not be elaborated here.

[0116] Step 33: Use the pre-trained decoder to perform self-attention mechanism processing on the object sequence vector to obtain a second processing result.

[0117] In this step, self-attention mechanism processing will be first performed on the object sequence vector to obtain a second processing result to perform feature interaction on the enhanced detection sequence vector and the tracking sequence vector.

[0118] Step 34: Use the second processing result, the encoded global image features, and the global image features as the second query value, the second key value, and the second numerical value in sequence, use the pre-trained decoder to perform attention mechanism processing on the second query value, the second key value, and the second numerical value to obtain a third processing result, and use the pre-trained decoder to perform forward propagation calculation on the third processing result to obtain tracking information.

[0119] In this step, attention mechanism processing will be continued on the second processing result and the global image features to obtain a third processing result, and forward propagation calculation will be performed on the third processing result to perform feature interaction on the second processing result and the global image features.

[0120] For ease of understanding, please refer to Figure 4 , Figure 4A schematic diagram of a pre-trained decoder provided by an embodiment of the present invention is shown. In this schematic diagram, the structure outlined by the dashed box is the Transformer decoder unit in the pre-trained decoder. The Transformer decoder unit may include a multi-head attention calculation module (MHA, Multi-head Attention), an addition normalization module, and a forward propagation module. The pre-trained decoder may be formed by connecting N Transformer decoder units in series, where N≥1.

[0121] S104. Update the tracking sequence vector using the tracking information, and enhance the updated tracking sequence vector using the pre-trained feature interactor and the attribute description text corresponding to the object to be tracked, so as to continue target tracking using the enhanced tracking sequence vector.

[0122] In this step, first, the tracking sequence vector can be updated using the tracking information obtained in the previous step. For example, the image features of the newly detected object to be tracked can be added to the tracking sequence vector, the tracking sequence vector can be updated using the new image features of the tracked object, and the image features of the tracked object that have disappeared can be removed from the tracking sequence vector.

[0123] To facilitate the tracking sequence vector, in addition to outputting the position information and image features of the object to be tracked in the image to be detected, the pre-trained decoder can also output the tracking status category of the object to be tracked and the confidence corresponding to this tracking status category. That is, in short, the pre-trained decoder can also determine the tracking situation of the object to be tracked. The tracking status category can specifically be a newborn tracking status or a terminated tracking status. The newborn tracking status indicates that the object to be tracked belongs to a newly added tracking object in the current image; the terminated tracking status indicates that the object to be tracked belongs to a tracked object in the historical image. The confidence refers to the degree of credibility that the object to be tracked belongs to this status. The following introduces the specific method of updating the tracking sequence vector using the image features, tracking status category, and confidence.

[0124] Based on this, the tracking information further includes the tracking status category of the object to be tracked and the confidence corresponding to the tracking status category; updating the tracking sequence vector using the tracking information may include:

[0125] Step 41: If it is determined that the object to be tracked belongs to a newly added tracking object according to the tracking status category, when the confidence of the object to be tracked is greater than the first preset threshold, add the image features of the object to be tracked to the tracking sequence vector.

[0126] In this step, if the object to be tracked belongs to a newly added tracked object, it is necessary to determine whether the confidence level of the object is greater than the first preset threshold. If it is greater, it can be reliably determined that the object to be tracked is a newly added object in the current image, and then the image features of the object to be tracked can be added to the tracking sequence vector.

[0127] Step 42: If it is determined that the object to be tracked belongs to a tracked object according to the tracking status category, when the confidence level of the object to be tracked is less than the second preset threshold, the image features of the object to be tracked are removed from the tracking sequence vector, or when the confidence level of the object to be tracked is not less than the second preset threshold, the tracking sequence vector is updated using the image features of the object to be tracked.

[0128] In this step, if the object to be tracked belongs to a tracked object, it is necessary to make a discrimination process based on the confidence level of the object. For example, when it is determined that the confidence level of the object to be tracked is less than the second preset threshold, it is no longer possible to reliably determine that the object to be tracked is in the tracked state. At this time, the image features of the object to be tracked can be removed from the tracking sequence vector. Another example is that when it is determined that the confidence level of the object to be tracked is not less than the second preset threshold, it can be reliably determined that the object to be tracked is in the tracked state. At this time, the tracking sequence vector can be updated using the image features of the object to be tracked to add the latest image features of the tracked object to the tracking sequence vector.

[0129] It should be noted that the specific values of the first preset threshold and the second preset threshold are not limited in this embodiment and can be set according to actual application requirements.

[0130] Furthermore, in step S104, the updated tracking sequence vector also needs to be enhanced using a pre-trained feature interaction module and the attribute description text corresponding to the object to be tracked. The attribute description text is used to describe the attributes of the object to be tracked. For example, when the tracked object is a pedestrian, during the pedestrian tracking process, the commonly used pedestrian attributes for recognition are: gender (male, female), age (adult, child), clothing color, accessories (bag, bicycle, etc.). By performing feature interaction between the attribute description text and the tracking sequence vector, it can be ensured that the model can pay attention to other attributes of the object to be tracked, thereby improving the tracking and detection effect.

[0131] It should be noted that this embodiment does not limit how to enhance the updated tracking sequence vector using the pre-trained feature interactor and the attribute description text corresponding to the object to be tracked. For example, the attribute description text can be used to perform cross-attention mechanism processing on the tracking sequence vector to achieve feature interaction. Taking into account that the attribute description text and the tracking sequence vector belong to features of two different modalities, in order to enhance the degree of interaction between cross-modal features, this embodiment can also map the features of the above two modalities to each other, and can use the text-image contrast learning model to narrow the distance between the features of the above two modalities. Please refer to the description in the subsequent embodiments for details.

[0132] Based on the above embodiments, the present invention can first obtain the global image features of the image to be detected, the category description text corresponding to the object to be tracked, the detection sequence vector and the tracking sequence vector. Among them, the category description text is used to describe the category of the object to be tracked; the detection sequence vector records the detection information of the object to be tracked, which is used to detect the object to be tracked; the tracking sequence vector records the image features of the tracked object, which is used to continue to track the tracked object. Subsequently, the present invention can fuse the category description text with the detection sequence vector into an enhanced detection sequence vector, that is, the detection sequence vector can be enhanced using the text semantic features. Subsequently, the present invention can use the pre-trained decoder to decode the global image features, the enhanced detection sequence vector, and the tracking sequence vector together to obtain the tracking information of the object to be tracked in the image to be detected, and the tracking information at least includes position information and image features. Subsequently, the present invention can use the tracking information to update the tracking sequence vector, and use the pre-trained feature interactor and the attribute description text corresponding to the object to be tracked to enhance the updated tracking sequence vector, that is, the tracking sequence vector can also be enhanced using the text semantic features to continue to track the target using the enhanced tracking sequence vector. Simply put, the present invention can use the category description text of the object to be tracked to enhance the detection sequence vector, and use the attribute description text of the object to be tracked to enhance the tracking sequence vector. Considering that the text semantic features have strong high-order semantic description capabilities and are completely unaffected by noise such as scale changes, deformation, and illumination changes, the overall performance of multi-target tracking can be significantly enhanced.

[0133] Based on the above embodiments, the following introduces the specific process of enhancing the updated tracking sequence vector by using the pre-trained feature interactor and the attribute description text corresponding to the object to be tracked. In a possible case, the pre-trained feature interactor may include a first mapping unit, a second mapping unit, pre-trained text features, and pre-trained image features. Both the first mapping unit and the second mapping unit are trainable module units; the pre-trained text features and the pre-trained image features are both learnable vectors and need to optimize parameters during the model training phase. The above process of enhancing the updated tracking sequence vector by using the pre-trained feature interactor and the attribute description text corresponding to the object to be tracked may include:

[0134] S201. Obtain text features and extract the image features of the tracked object from the tracking sequence vector; wherein, the text features are generated using the attribute description text.

[0135] In this step, the text features corresponding to the above attribute description text can be obtained, and the image features of the tracked object can be obtained from the tracking sequence vector. It should be noted that, to improve the interaction effect, when obtaining the text features, all text features related to the attributes of the object to be tracked can be obtained.

[0136] Furthermore, for the convenience of obtaining text features, this embodiment may preset a tracking text prompt library, which contains the text features corresponding to the attribute description texts of various objects to be tracked. Then, this embodiment can obtain the text features corresponding to the category of the category description text in the tracking text prompt library according to the category description text.

[0137] Based on this, obtaining text features may include:

[0138] Step 51: According to the category description text, obtain the text features corresponding to the category of the category description text in the tracking text prompt library; the tracking text prompt library contains the text features corresponding to various objects to be tracked.

[0139] It should be noted that this embodiment does not limit the specific process of converting the attribute description text into text features. For example, the attribute description text can be encoded, such as encoding the attribute description text using the Wordpiece Tokenization method to obtain the text features.

[0140] Based on this, this method may further include:

[0141] Step 61: Obtain the attribute description text corresponding to the object to be tracked.

[0142] Step 62: Encode the attribute description text to obtain the text features.

[0143] Step 63: Establish a correspondence between the text feature and the category of the object to be tracked, and save the correspondence to the tracking text prompt library.

[0144] It should be noted that this embodiment does not limit the specific attribute description text, and can be set according to actual application requirements.

[0145] S202: Use a first mapping unit to project the pre-trained text features into mapped image features, and use a second mapping unit to project the pre-trained image features into mapped text features.

[0146] In this step, the pre-trained text features can be projected into mapped image features using the first mapping unit, and the pre-trained image features can be projected into mapped text features using the second mapping unit. The purpose of setting the pre-trained text features and pre-trained image features is that in terms of text semantics, although the prior attribute description text can cover most of the attributes of the object to be tracked, the actual tracking detection scene is dynamic and changeable. On the one hand, there are attributes that the attribute description text cannot cover. On the other hand, the attribute description text only statically describes the attributes of the object to be tracked, and cannot adapt well to the dynamic changes of the object to be tracked itself. Therefore, pre-trained text features can be introduced to use pre-trained text features to complete the information of static text features through dynamic training. Similarly, pre-trained image features can be used to complete the information of image features.

[0147] Furthermore, the first mapping unit is used to project the pre-trained text features into mapped image features, and the second mapping unit is used to project the pre-trained image features into mapped text features. The core purpose is to enhance the degree of interaction between text features and image features. It should be noted that this embodiment does not limit the specific types and structures of the first mapping unit and the second mapping unit, and can be set according to actual application requirements. For example, the first mapping unit and the second mapping unit can both be multi-layer perceptrons.

[0148] S203, fusing the text features, pre-trained text features, and mapped text features to obtain fused text features, and encoding the fused text features using a pre-trained text-image contrast learning model to obtain encoded text features.

[0149] In this step, the text features, pre-trained text features, and mapped text features can be fused to obtain fused text features to enrich the text semantic information in the fused text features. Subsequently, the fused text features can be encoded using the pre-trained text-image contrast learning model to obtain encoded text features, thereby narrowing the distance between text features and image features for better feature interaction.

[0150] It should be noted that this embodiment does not limit the specific type of the pre-trained text-image contrast learning model, for example, it can be a Clip model.

[0151] S204, fusing the image features, the pre-trained image features, and the mapped image features to obtain fused image features, and encoding the fused image features using a pre-trained text-image contrast learning model to obtain encoded image features.

[0152] In this step, the image features, pre-trained image features, and mapped image features can be fused to obtain fused image features to enrich the image information in the fused image features. Subsequently, the fused image features can be encoded using the pre-trained text-image contrast learning model to obtain encoded image features, thereby narrowing the distance between text features and image features for better feature interaction.

[0153] S205. Perform cross-attention mechanism processing on the encoded text features and the encoded image features to obtain an enhanced tracking sequence vector.

[0154] In this step, the encoded text features and the encoded image features can be processed by a cross-attention mechanism, and the processing results of the cross-attention mechanism can be used to form an enhanced tracking sequence vector, thereby improving the detection and tracking performance of the tracking sequence vector.

[0155] For easier understanding, please refer to Figure 5 , Figure 5 A schematic diagram of a tracking sequence vector enhancement method based on text-image bidirectional prompt learning provided by an embodiment of the present invention. In this schematic diagram, the gray part constitutes a pre-trained feature interactor, the "learnable text feature vector" is a pre-trained text feature, the "learnable image feature vector" is a pre-trained image feature, the "text-image projection calculation" uses the first mapping unit, the "image-text projection calculation" uses the second mapping unit, and the "text encoding" and "image encoding" use the text-image contrast learning model.

[0156] Furthermore, the above embodiment introduces a bidirectional mapping of text to image, but any unidirectional mapping of text to image or image to text is also a way of multimodal learning. The following introduces a tracking sequence vector enhancement method based on unidirectional mapping of text to image.

[0157] Based on this, the pre-trained feature interactor includes a first mapping unit and a pre-trained text feature. The updated tracking sequence vector is enhanced using the pre-trained feature interactor and the attribute description text corresponding to the object to be tracked, including:

[0158] S301. Obtain text features and extract the image features of the tracked object from the tracking sequence vector; wherein, the text features are generated using the attribute description text.

[0159] S302. Use the first mapping unit to project the pre-trained text features into mapped image features.

[0160] Different from the above embodiments, in this embodiment, the pre-trained feature interactor only includes the first mapping unit and the pre-trained text features. Therefore, it is only necessary to map the pre-trained text features into mapped image features using the first mapping unit.

[0161] S303. Fuse the text features and the pre-trained text features to obtain fused text features, and use the pre-trained text-image contrast learning model to encode the fused text features to obtain encoded text features.

[0162] S304. Fuse the image features and the mapped image features to obtain fused image features, and use the pre-trained text-image contrast learning model to encode the fused image features to obtain encoded image features.

[0163] In steps S303 and S304, different from steps S203 and S204, since only one-way mapping between text and image is performed in this embodiment, the fused text features are only obtained by fusing the text features and the pre-trained text features, and the fused image features are only obtained by fusing the image features and the mapped image features.

[0164] S305. Perform cross-attention mechanism processing on the encoded text features and the encoded image features to obtain an enhanced tracking sequence vector.

[0165] For ease of understanding, please refer to Figure 6 , Figure 6 is a schematic diagram of a method for enhancing the tracking sequence vector based on one-way text-image prompt learning provided by an embodiment of the present invention. In this schematic diagram, the gray part constitutes the pre-trained feature interactor, the "learnable text feature vector" is the pre-trained text feature, the "text-image projection calculation" uses the first mapping unit, and the "text encoding" and "image encoding" use the pre-trained text-image contrast learning model.

[0166] The following introduces the method for enhancing the tracking sequence vector based on one-way mapping between image and text. Based on this, the pre-trained feature interactor includes a second mapping unit and pre-trained image features. Using the pre-trained feature interactor and the attribute description text corresponding to the object to be tracked to enhance the updated tracking sequence vector may include:

[0167] S401. Obtain text features and extract the image features of the tracked object from the tracking sequence vector; wherein, the text features are generated using the attribute description text.

[0168] S402. Project the pre-trained image features into mapped text features using the second mapping unit.

[0169] Different from the above embodiments, in this embodiment, the pre-trained feature interactor only includes the second mapping unit and the pre-trained image features. Therefore, it is only necessary to map the pre-trained image features into mapped text features using the second mapping unit.

[0170] S403. Fuse the text features and the mapped text features to obtain fused text features, and encode the fused text features using the pre-trained text-image contrastive learning model to obtain encoded text features.

[0171] S404. Fuse the image features and the pre-trained image features to obtain fused image features, and encode the fused image features using the pre-trained text-image contrastive learning model to obtain encoded image features.

[0172] In steps S403 and S404, different from steps S203 and S204, since only image-text one-way mapping is performed in this embodiment, the fused text features are only obtained by fusing the text features and the mapped text features, and the fused image features are only obtained by fusing the image features and the pre-trained image features.

[0173] S405. Perform cross-attention mechanism processing on the encoded text features and the encoded image features to obtain an enhanced tracking sequence vector.

[0174] For easy understanding, please refer to Figure 7 , Figure 7 is a schematic diagram of a method for enhancing the tracking sequence vector based on image-text one-way prompt learning provided by an embodiment of the present invention. In this schematic diagram, the gray part constitutes the pre-trained feature interactor, the "learnable image feature vector" is the pre-trained image feature, the "image-text projection calculation" uses the second mapping unit, and the "text encoding" and "image encoding" use the text-image contrastive learning model.

[0175] Based on the above embodiments, the training methods of the detection sequence vector, the pre-trained decoder, and the pre-trained feature interactor are introduced below. Based on this, this method may further include:

[0176] S501. Obtain multiple training images, the category description text corresponding to the training tracking object, the initial detection sequence vector and the tracking sequence vector, and extract the training image features of the training images; the true position information of the training tracking object is marked in the training images, and the initial tracking sequence vector is empty.

[0177] In this step, first, multiple training images can be obtained, and the true position information of the training tracking object has been labeled for these training images. Subsequently, feature extraction can be performed on these training images to obtain training image features.

[0178] In addition, in this step, the category description text, the initial detection sequence vector, and the tracking sequence vector corresponding to the training tracking object can also be obtained. Among them, the initial detection sequence vector refers to the detection sequence vector that has not been trained and can be initialized by random initialization.

[0179] S502. Randomly initialize the initial detection sequence vector, and fuse the category description text with the initial detection sequence vector into an initial enhanced detection sequence vector.

[0180] S503. Use the initial enhanced detection sequence vector, the tracking sequence vector, the training image features, the initial decoder, and the initial feature interactor to determine the predicted position information of the training tracking object in each training image.

[0181] In this step, the initial enhanced detection sequence vector, the tracking sequence vector, the training image features, the initial decoder, and the initial feature interactor can be used to determine the predicted position information of the training tracking object in each training image. The specific training process can refer to the relevant descriptions in steps S103 and S104.

[0182] S504. Use a preset loss function to calculate the loss between the predicted position information and the true position information to obtain a loss value.

[0183] In this step, to ensure that the model can reliably determine the position of the object to be tracked in the image after training, it is necessary to use a preset loss function to calculate the loss value between the predicted position information and the true position information, and use the loss value to update the parameters of the initial detection sequence vector, the initial decoder, and the initial feature interactor.

[0184] It should be noted that the specific preset loss function is not limited in this embodiment and can be set according to actual application requirements. For example, a possible preset loss function can be:

[0185] ;

[0186] Among them, represents the predicted value, represents the true value, N is the number of training images, represents the total number of objects in N training images, represents the predicted value of the tracked object generated in the i-th training image based on the tracking sequence vector ( , track query), denote the ground truth of the tracked object in the i-th training image, denote the predicted value of the newly added tracked object generated in the i-th training image based on the initial detection sequence vector ( , detect query), denote the ground truth of the newly added tracked object in the i-th training image. The loss value between the predicted value and the ground truth in a single-frame image is:

[0187] ;

[0188] where denote the classification loss between the predicted value and the ground truth, denote the L1 loss between the predicted value and the ground truth, denote the Generalized IoU loss between the predicted value and the ground truth, , , are all hyperparameters.

[0189] S505. Update the parameters of the initial detection sequence vector, the initial decoder, and the initial feature interactor using the loss value to obtain the detection sequence vector, the pre-trained decoder, and the pre-trained feature interactor.

[0190] It should be noted that in this embodiment, steps S503 to S505 can be executed multiple times according to the training situation until the model converges.

[0191] Based on the above embodiments, the above object tracking method will be introduced in full below based on a specific schematic diagram. Please refer to Figure 8 , Figure 8 which is a schematic diagram of a system provided by an embodiment of the present invention. For the input video image frame t, first, the primary image features are extracted through the backbone network. Here, the backbone network is a CNN convolutional neural network, which can be any convolutional neural network, such as Resnet50, Darknet, etc. Then, the primary image features are input into the image feature encoder for high-level image feature encoding to obtain the global image features. Next, the global image features, the concatenated detection sequence vector, and the tracking sequence vector are input into the decoder for feature decoding to obtain the tracking information of the target, such as position information, tracking id information, etc. Then, according to the tracking information of the target, the tracking trajectory is parsed and updated. Finally, according to the target trajectory information, the tracking sequence vector is calculated, and the calculated tracking sequence vector is used to update the historical tracking sequence vector for the parsing of the next frame of video image.

[0192] In the present invention, the detection sequence vector "detect query" is a set of special learnable parameters, and each vector parameter represents a potential target object, which is optimized through backpropagation during the training process. For the enhancement process of the detection sequence vector, please refer to Figure 3 , first, according to the tracking category, input the text to describe the object. For example, when the tracking object is "person", the text description is "person", and then convert "person" into a vector with a fixed dimension. For example, the WordPiece tokenization method can be used for text vector conversion. Subsequently, the initialized detection sequence vector is usually initialized to random values, and the vector dimension is the same as the text description encoding dimension. Then, the text vector encoding and the initialized detection sequence vector are concatenated, and through linear transformation fusion calculation, the detection sequence vector based on the text prompt is learned. Compared with the detection sequence vector learned solely based on visual features, in the present invention, text prompts are introduced to enhance the object detection expression ability of the detect query.

[0193] In the present invention, the image feature encoding can adopt the Transformer Encoder structure for advanced image feature learning. The input is the primary image feature vector and the spatial position encoding vector, and the output is the encoded advanced image feature vector. During the feature calculation process, the self-attention learning encoding mode is adopted to facilitate the network to better extract the relationships between different objects at different positions. The calculation process is as Figure 2 shown, where N represents repeating the calculation inside the dotted line box with repeated colors N times, and N = 6.

[0194] In the present invention, feature decoding is used to parse the object to be tracked in the current image frame, and obtain the position information, category information, confidence score of the target, and the trajectory ID information of the target. In the present invention, the Transformer Decoder can be used for feature decoding during the feature decoding process, and the calculation process is as Figure 4 shown. First, the self-attention mechanism is used to perform relevant calculations on the concatenated detection sequence vector and the tracking sequence vector, and then based on the position encoding vector and the advanced image feature vector, the cross-attention mechanism is used to perform relevant calculations. This process is repeated N times, and finally, the relevant information of the decoded tracking target is output. In Figure 4 , N represents repeating the calculation inside the gray dotted line box N times, and N = 6.

[0195] After feature decoding, relevant information of the target object is obtained, and the tracking trajectory is updated according to the relevant information of the target object. In the present invention, the tracking trajectory update means that, according to the target object obtained by feature decoding in the current frame, the trajectory state of the target object is judged, which is divided into the following three cases:

[0196] Trajectory new construction: When the target object is not assigned a trajectory ID number and the target confidence score is greater than a preset threshold, it indicates that the target object is a newly detected target, and a new trajectory ID number is assigned to the target object.

[0197] Trajectory disappearance: When the detection confidence score of the existing trajectory object in the current frame is less than the preset threshold and the time less than the preset threshold is greater than the preset time, it indicates that the target trajectory has left the monitoring range, and the target trajectory will be removed.

[0198] Trajectory continuous tracking: When the detection confidence score of the existing trajectory object in the current frame is greater than the preset threshold, the trajectory object is retained and continued to be tracked.

[0199] After obtaining the latest tracking trajectory information, the system will update the tracking sequence vector based on the tracking result. In the present invention, in order to enhance the tracking feature expression ability of the tracking sequence vector, the present invention designs a calculation method for the tracking sequence vector based on text-image bidirectional prompt learning, and the calculation process is as Figure 5 shown.

[0200] In the present invention, a tracking text prompt library is introduced. The tracking text prompt library, as a kind of prior knowledge, describes the tracking target object through text information. The advantage of describing the target object with text information is that, as a more advanced semantic feature, the text description can effectively cope with noise effects such as scale change, deformation, and illumination change, and can more robustly extract the semantic features of the tracking object for use in tracking trajectory recognition. When establishing the tracking text prompt library, description texts beneficial to tracking recognition can be designed according to the prior knowledge of the tracking category. For example, when the tracking object is a pedestrian, during the pedestrian tracking process, the commonly used pedestrian attributes for recognition are: gender (male, female), age (adult, child), clothing color, accessories (bag, bicycle, etc.).

[0201] In this multi-object tracking system, the tracking sequence vector is jointly calculated by fusing text features and image features. During the model training process, the text feature vector T consists of three parts: the text features obtained after encoding the tracking text prompt library , the learnable text feature vector , and the mapped text feature vector . The image feature vector V also consists of three parts: the image feature vector obtained from the tracking target , the learnable image feature vector , the mapped image feature vector . In the following formula, [] represents concatenate.

[0202] ;

[0203] ;

[0204] wherein, the learnable text feature vector and the learnable image feature vector can set the dimension size according to the actual situation. The learnable text feature vector and the learnable image feature vector are initialized through Gaussian distribution and optimized and learned through backpropagation during the training process.

[0205] The mapped text feature vector is calculated from the learnable image feature vector through the coupling function . In the present invention, the coupling function can be composed of a multi-layer perceptron (MLP) structure. Similarly, the mapped image feature vector is calculated from the learnable text feature vector through the coupling function . In the present invention, the coupling function is composed of a multi-layer perceptron (MLP) structure.

[0206] ;

[0207] ;

[0208] ;

[0209] Finally, the image feature vector V and the text feature vector T are respectively input into the vision encoder (vision encoder) and the text encoder (text encoder) to generate new feature vectors, and a new tracking query vector is calculated through the cross-attention mechanism crossAttn (cross attention) for use in the object tracking of the next frame. In the present invention, the vision encoder and the text encoder can directly calculate using the pre-trained CLIP model parameters.

[0210] Therefore, during the training process of the tracking model, for the update module of the track query, the model parameters that need to be learned and trained are: the learnable text feature vector 、Learnable image feature vectors 、Text-image projection coupling function Parameter, image-text projection coupling function Parameter and all parameters in cross-attention learning. The remaining model parameters can be obtained by freezing the pre-trained model parameters.

[0211] In the present invention, a text-image bidirectional prompt learning mechanism is adopted to update the track query. The coupling functions and serve as bridges to effectively connect text-image features. By means of the multi-modal learning mode of CLIP, the representative features between text and image are drawn closer, enhancing the feature expression ability of the tracking target from both semantic and visual aspects, so as to improve the generalization ability of the tracking system in different scenarios, which is beneficial to the deployment of the multi-object tracking system in different monitoring scenarios.

[0212] Next, the target tracking device, electronic device, computer-readable storage medium, and computer program product provided by the embodiments of the present invention will be introduced. The target tracking device, electronic device, computer-readable storage medium, and computer program product described below can be correspondingly referred to the target tracking method described above.

[0213] Please refer to Figure 9 , Figure 9 which is the structural block diagram of a target tracking device provided by an embodiment of the present invention. The device may include:

[0214] An acquisition module 901, configured to acquire the global image feature of the image to be detected, the category description text corresponding to the object to be tracked, the detection sequence vector, and the tracking sequence vector; the detection sequence vector records the detection information of the object to be tracked, and the tracking sequence vector records the image features of the tracked object;

[0215] A detection sequence enhancement module 902, configured to fuse the category description text with the detection sequence vector into an enhanced detection sequence vector;

[0216] A tracking detection module 903, configured to jointly decode the global image feature, the enhanced detection sequence vector, and the tracking sequence vector by using a pre-trained decoder to obtain the tracking information of the object to be tracked in the image to be detected; the tracking information includes at least position information and image features;

[0217] A tracking sequence enhancement module 904, configured to update the tracking sequence vector by using the tracking information, and enhance the updated tracking sequence vector by using a pre-trained feature interaction device and the attribute description text corresponding to the object to be tracked, so as to continue target tracking by using the enhanced tracking sequence vector.

[0218] Optionally, the detection sequence enhancement module 902 may include:

[0219] A text encoding sub-module for encoding the category description text to obtain a category description text vector;

[0220] An enhancement sub-module for concatenating the category description text vector and the detection sequence vector, and sequentially performing linear transformation processing and normalization processing on the concatenation result to obtain an enhanced detection sequence vector.

[0221] Optionally, the pre-trained feature interactor includes a first mapping unit, a second mapping unit, pre-trained text features, and pre-trained image features; the tracking sequence enhancement module 904 may include:

[0222] A first acquisition sub-module for acquiring text features and extracting the image features of the tracked object from the tracking sequence vector; wherein, the text features are generated using the attribute description text.

[0223] A first mapping sub-module for projecting the pre-trained text features into mapped image features using the first mapping unit, and projecting the pre-trained image features into mapped text features using the second mapping unit;

[0224] A first fusion sub-module for fusing the text features, pre-trained text features, and mapped text features to obtain fused text features, and encoding the fused text features using the pre-trained text-image contrast learning model to obtain encoded text features;

[0225] A second fusion sub-module for fusing the image features, pre-trained image features, and mapped image features to obtain fused image features, and encoding the fused image features using the pre-trained text-image contrast learning model to obtain encoded image features;

[0226] A first enhancement sub-module for performing cross-attention mechanism processing on the encoded text features and the encoded image features to obtain an enhanced tracking sequence vector.

[0227] Optionally, both the first mapping unit and the second mapping unit are multi-layer perceptrons.

[0228] Optionally, the acquisition sub-module may include:

[0229] A text feature acquisition unit for obtaining the text features corresponding to the category of the category description text in the tracking text prompt library; the tracking text prompt library contains the text features corresponding to various objects to be tracked.

[0230] Optionally, the device may further include:

[0231] An attribute description text acquisition module, configured to acquire an attribute description text corresponding to an object to be tracked;

[0232] An attribute description text encoding module, configured to encode the attribute description text to obtain a text feature;

[0233] A prompt library storage module, configured to establish a correspondence between the text feature and the category of the object to be tracked, and store the correspondence in a tracking text prompt library.

[0234] Optionally, the pre-trained feature interactors include a first mapping unit and pre-trained text features; the tracking sequence enhancement module 904 may include:

[0235] A second acquisition sub-module, configured to acquire the text feature and extract an image feature of the tracked object from the tracking sequence vector; wherein, the text feature is generated using the attribute description text;

[0236] A second mapping sub-module, configured to project the pre-trained text feature into a mapped image feature using the first mapping unit;

[0237] A third fusion sub-module, configured to fuse the text feature and the pre-trained text feature to obtain a fused text feature, and encode the fused text feature using a pre-trained text-image contrastive learning model to obtain an encoded text feature;

[0238] A fourth fusion sub-module, configured to fuse the image feature and the mapped image feature to obtain a fused image feature, and encode the fused image feature using a pre-trained text-image contrastive learning model to obtain an encoded image feature;

[0239] A second enhancement sub-module, configured to perform a cross-attention mechanism process on the encoded text feature and the encoded image feature to obtain an enhanced tracking sequence vector.

[0240] Optionally, the pre-trained feature interactors include a second mapping unit and pre-trained image features; the tracking sequence enhancement module 904 may include:

[0241] A third acquisition sub-module, configured to acquire the text feature and extract an image feature of the tracked object from the tracking sequence vector; wherein, the text feature is generated using the attribute description text;

[0242] A third mapping sub-module, configured to project the pre-trained image feature into a mapped text feature using the second mapping unit;

[0243] A fifth fusion sub-module, configured to fuse the text feature and the mapped text feature to obtain a fused text feature, and encode the fused text feature using a pre-trained text-image contrastive learning model to obtain an encoded text feature;

[0244] The sixth fusion sub-module is used to fuse the image features and the pre-trained image features to obtain fused image features, and encode the fused image features using the pre-trained text-image contrastive learning model to obtain encoded image features;

[0245] The third enhancement sub-module is used to perform cross-attention mechanism processing on the encoded text features and the encoded image features to obtain enhanced tracking sequence vectors.

[0246] Optionally, the acquisition module 901 may include:

[0247] The primary feature extraction sub-module is used to extract features from the image to be detected using the pre-trained convolutional neural network model to obtain primary image features;

[0248] The first spatial encoding sub-module is used to perform spatial position encoding on the primary image features according to the spatial positions corresponding to the primary image features in the image to be detected to obtain encoded primary image features;

[0249] The first processing sub-module is used to use the encoded primary image features as the first key value and the first query value, the primary image features as the first numerical value, and perform self-attention mechanism processing on the first key value, the first query value, and the first numerical value using the pre-trained image feature encoder to obtain a first processing result, and perform forward propagation calculation on the first processing result using the pre-trained image feature encoder to obtain global image features.

[0250] Optionally, the tracking detection module 903 may include:

[0251] The splicing sub-module is used to splice the enhanced detection sequence vector and the tracking sequence vector to obtain an object sequence vector;

[0252] The second spatial encoding sub-module is used to perform spatial position encoding on the global image features according to the spatial positions corresponding to the global image features in the image to be detected to obtain encoded global image features;

[0253] The third processing sub-module is used to perform self-attention mechanism processing on the object sequence vector using the pre-trained decoder to obtain a second processing result;

[0254] The fourth processing sub-module is used to use the second processing result, the encoded global image features, and the global image features as the second query value, the second key value, and the second numerical value in sequence, and perform attention mechanism processing on the second query value, the second key value, and the second numerical value using the pre-trained decoder to obtain a third processing result, and perform forward propagation calculation on the third processing result using the pre-trained decoder to obtain tracking information.

[0255] Optionally, the tracking information further includes the tracking status category of the object to be tracked and the confidence level corresponding to the tracking status category; the tracking sequence enhancement module 904 includes:

[0256] A new processing sub-module, configured to, if it is determined according to the tracking status category that the object to be tracked belongs to a newly added tracking object, add the image feature of the object to be tracked to the tracking sequence vector when the confidence level of the object to be tracked is greater than a first preset threshold;

[0257] An update and exit processing sub-module, configured to, if it is determined according to the tracking status category that the object to be tracked belongs to a tracked object, remove the image feature of the object to be tracked from the tracking sequence vector when the confidence level of the object to be tracked is less than a second preset threshold, or update the tracking sequence vector by using the image feature of the object to be tracked when the confidence level of the object to be tracked is not less than the second preset threshold.

[0258] Optionally, the apparatus may further include:

[0259] An initialization module, configured to obtain multiple training images, category description texts corresponding to the training tracking objects, an initial detection sequence vector, and a tracking sequence vector, and extract the training image features of the training images; the real position information of the training tracking objects is marked in the training images, and the initial tracking sequence vector is empty;

[0260] An initial detection sequence enhancement module, configured to randomly initialize the initial detection sequence vector and fuse the category description text with the initial detection sequence vector to obtain an initial enhanced detection sequence vector;

[0261] A training module, configured to determine the predicted position information of the training tracking objects in each training image by using the initial enhanced detection sequence vector, the tracking sequence vector, the training image features, an initial decoder, and an initial feature interaction unit;

[0262] A loss value calculation module, configured to calculate the loss between the predicted position information and the real position information by using a preset loss function to obtain a loss value;

[0263] A parameter update module, configured to update the parameters of the initial detection sequence vector, the initial decoder, and the initial feature interaction unit by using the loss value to obtain a detection sequence vector, a pre-trained decoder, and a pre-trained feature interaction unit.

[0264] Please refer to Figure 10 , Figure 10 which is a structural block diagram of an electronic device provided by an embodiment of the present invention. An embodiment of the present invention provides an electronic device 10, including a processor 11 and a memory 12; wherein, the memory 12 is used to store a computer program; the processor 11 is configured to execute the target tracking method provided by the foregoing embodiment when executing the computer program.

[0265] For the specific process of the above-mentioned target tracking method, reference can be made to the corresponding content provided in the foregoing embodiments, and details will not be elaborated herein.

[0266] Moreover, as a carrier for resource storage, the memory 12 can be a read-only memory, a random access memory, a magnetic disk, an optical disk, etc., and the storage method can be temporary storage or permanent storage.

[0267] In addition, the electronic device 10 further includes a power supply 13, a communication interface 14, an input / output interface 15, and a communication bus 16; among them, the power supply 13 is used to provide operating voltage for each hardware device on the electronic device 10; the communication interface 14 can create a data transmission channel between the electronic device 10 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present invention, and no specific limitation is imposed thereon here; the input / output interface 15 is used to obtain external input data or output data to the outside, and the specific interface type can be selected according to specific application needs, and no specific limitation is imposed here.

[0268] The embodiment of the present invention further provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the target tracking method provided in the above embodiments is implemented.

[0269] Since the embodiments of the computer program product part correspond to the embodiments of the target tracking method part, for the embodiments of this part, please refer to the description of the embodiments of the target tracking method part, and details will not be elaborated herein.

[0270] The embodiment of the present invention further provides a non-volatile computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the target tracking method provided in the above embodiments is implemented.

[0271] Since the embodiments of the non-volatile computer-readable storage medium part correspond to the embodiments of the target tracking method part, for the embodiments of this part, please refer to the description of the embodiments of the target tracking method part, and details will not be elaborated herein.

[0272] The embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description in the method part.

[0273] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.

[0274] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0275] The above has introduced in detail a target tracking method, device, electronic device, program product, and medium provided by the present invention. Specific examples are used herein to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.

Claims

1. A target tracking method, characterized in that: include: Obtaining global image features of the image to be detected, category description text corresponding to the object to be tracked, a detection sequence vector, and a tracking sequence vector; The detection sequence vector records the detection information of the object to be tracked, and the tracking sequence vector records the image features of the tracked object; Merging the category description text with the detection sequence vector into an enhanced detection sequence vector; Using a pre-trained decoder to decode the global image features, the enhanced detection sequence vector, and the tracking sequence vector, to obtain tracking information of the object to be tracked in the image to be detected; the tracking information includes at least position information and image features; The tracking sequence vector is updated using the tracking information, and the updated tracking sequence vector is enhanced using a pre-trained feature interactor and an attribute description text corresponding to the object to be tracked, so as to continue to track the target using the enhanced tracking sequence vector; The pre-trained feature interactor comprises a first mapping unit, a second mapping unit, a pre-trained text feature, and a pre-trained image feature; Performing enhancement processing on the updated tracking sequence vector using the pre-trained feature interactor and the attribute description text corresponding to the object to be tracked, including: Acquire text features, and extract image features of the tracked object from the tracking sequence vector; wherein the text features are generated using the attribute description text; Using the first mapping unit to project the pre-trained text features into mapped image features, and using the second mapping unit to project the pre-trained image features into mapped text features; The text feature, the pre-trained text feature, and the mapped text feature are fused to obtain a fused text feature, and the fused text feature is encoded using a pre-trained text-image contrast learning model to obtain an encoded text feature; The image features, pre-trained image features, and mapped image features are fused to obtain fused image features, and the fused image features are encoded using a pre-trained text-image contrast learning model to obtain encoded image features; The encoded text features and the encoded image features are processed through a cross-attention mechanism to obtain an enhanced tracking sequence vector.

2. The target tracking method according to claim 1, characterized in that: The method further comprises: fusing the category description text with the detection sequence vector into an enhanced detection sequence vector, comprising: Encoding the category description text to obtain a category description text vector; The category description text vector is concatenated with the detection sequence vector, and the concatenated result is sequentially subjected to linear transformation processing and normalization processing to obtain the enhanced detection sequence vector.

3. The target tracking method according to claim 1, characterized in that: The first mapping unit and the second mapping unit are both multi-layer perceptrons.

4. The target tracking method according to claim 1, characterized in that: Get text features, including: According to the category description text, a text feature of a category corresponding to the category description text is obtained in a tracking text prompt library; the tracking text prompt library contains text features corresponding to various objects to be tracked.

5. The target tracking method according to claim 4, characterized in that: Also includes: Obtaining the attribute description text corresponding to the object to be tracked; Encoding the attribute description text to obtain the text feature; A correspondence is established between the text feature and the category of the object to be tracked, and the correspondence is saved in the tracking text prompt library.

6. The target tracking method according to claim 1, characterized in that: Obtain the global image features of the image to be detected, including: Using a pre-trained convolutional neural network model to extract features from the image to be detected, to obtain primary image features; According to the spatial position corresponding to the primary image feature in the image to be detected, the primary image feature is spatially encoded to obtain an encoded primary image feature; The encoded primary image feature is used as a first key value and a first query value, and the primary image feature is used as a first numerical value. A pre-trained image feature encoder is used to perform a self-attention mechanism on the first key value, the first query value and the first numerical value to obtain a first processing result, and the pre-trained image feature encoder is used to perform a forward propagation calculation on the first processing result to obtain the global image feature.

7. The target tracking method according to claim 1, characterized in that: The global image feature, the enhanced detection sequence vector, and the tracking sequence vector are decoded together by using a pre-trained decoder to obtain tracking information of the object to be tracked in the image to be detected, including: Concatenating the enhanced detection sequence vector and the tracking sequence vector to obtain an object sequence vector; According to the spatial position corresponding to the global image feature in the image to be detected, the global image feature is spatially encoded to obtain an encoded global image feature; Using the pre-trained decoder to perform self-attention mechanism processing on the object sequence vector to obtain a second processing result; The second processing result, the encoded global image feature and the global image feature are used as the second query value, the second key value and the second numerical value in turn, the pre-trained decoder is used to perform an attention mechanism on the second query value, the second key value and the second numerical value to obtain a third processing result, and the pre-trained decoder is used to perform forward propagation calculation on the third processing result to obtain the tracking information.

8. The target tracking method according to claim 1, characterized in that: The tracking information also includes the tracking state category of the object to be tracked and the confidence level corresponding to the tracking state category; Updating the tracking sequence vector using the tracking information includes: If it is determined according to the tracking state category that the object to be tracked is a newly added tracking object, then when the confidence of the object to be tracked is greater than a first preset threshold, adding the image feature of the object to be tracked to the tracking sequence vector; If it is determined according to the tracking state category that the object to be tracked belongs to a tracked object, then when the confidence of the object to be tracked is less than a second preset threshold, the image features of the object to be tracked are moved out of the tracking sequence vector, or when the confidence of the object to be tracked is not less than the second preset threshold, the tracking sequence vector is updated using the image features of the object to be tracked.

9. The target tracking method according to any one of claims 1 to 8, characterized in that: Also includes: Acquire multiple training images, category description texts corresponding to training tracking objects, initial detection sequence vectors and tracking sequence vectors, and extract training image features of the training images; The training image is annotated with the real position information of the training tracking object, and the initial tracking sequence vector is empty; Randomly initializing an initial detection sequence vector, and fusing the category description text with the initial detection sequence vector into an initial enhanced detection sequence vector; Determine the predicted position information of the training tracking object in each of the training images using the initial enhanced detection sequence vector, the tracking sequence vector, the training image features, the initial decoder and the initial feature interactor; Calculating the loss between the predicted position information and the actual position information using a preset loss function to obtain a loss value; The loss value is used to update the parameters of the initial detection sequence vector, the initial decoder, and the initial feature interactor to obtain the detection sequence vector, the pre-trained decoder, and the pre-trained feature interactor.

10. A target tracking device, characterized in that: include: An acquisition module, used to acquire global image features of an image to be detected, a category description text corresponding to an object to be tracked, a detection sequence vector and a tracking sequence vector; the detection sequence vector records detection information of the object to be tracked, and the tracking sequence vector records image features of a tracked object; A detection sequence enhancement module, used for fusing the category description text with the detection sequence vector into an enhanced detection sequence vector; A tracking detection module, configured to decode the global image features, the enhanced detection sequence vectors, and the tracking sequence vectors using a pre-trained decoder to obtain tracking information of the object to be tracked in the image to be detected; the tracking information includes at least position information and image features; A tracking sequence enhancement module, used to update the tracking sequence vector using the tracking information, and enhance the updated tracking sequence vector using a pre-trained feature interactor and an attribute description text corresponding to the object to be tracked, so as to continue to track the target using the enhanced tracking sequence vector; The pre-trained feature interactor comprises a first mapping unit, a second mapping unit, a pre-trained text feature, and a pre-trained image feature; The tracking sequence enhancement module comprises: A first acquisition submodule is used to acquire text features and extract image features of the tracked object from the tracking sequence vector; wherein the text features are generated using the attribute description text; A first mapping submodule, configured to project the pre-trained text features into mapped image features using the first mapping unit, and project the pre-trained image features into mapped text features using the second mapping unit; A first fusion submodule is used to fuse the text features, pre-trained text features, and mapped text features to obtain fused text features, and encode the fused text features using a pre-trained text-image contrast learning model to obtain encoded text features; A second fusion submodule is used to fuse the image features, the pre-trained image features, and the mapped image features to obtain fused image features, and encode the fused image features using a pre-trained text-image contrast learning model to obtain encoded image features; The first enhancement submodule is used to perform cross-attention mechanism processing on the encoded text features and the encoded image features to obtain an enhanced tracking sequence vector.

11. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the target tracking method as claimed in any one of claims 1 to 9 when executing the computer program.

12. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the target tracking method according to any one of claims 1 to 9 is implemented.

13. A non-volatile computer-readable storage medium, characterized in that: The non-volatile computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are loaded and executed by the processor, the target tracking method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Semantic feature selection and attention fusion-oriented video description generation method and system

    CN117789076A

  • Target tracking method and system, computer equipment and storage medium

    CN118379515A