Target tracking method, and storage medium and vehicle

By using a pre-trained timing attention model to compress the feature information of the target object and output the weight feature information, the problem of long-term target tracking in the prior art is solved, fast and accurate target tracking is achieved, and the real-time response capability of the autonomous driving system is improved.

WO2025139384A1PCT designated stage expired Publication Date: 2025-07-03ANHUI NIO AUTONOMOUS DRIVING TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/130394
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-11-07
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The existing target tracking scheme has complex algorithms, long system computing, and cannot obtain tracking results in real time, and cannot meet the scenario requirements such as autonomous driving that require real-time decision-making.

Method used

The pre-trained timing attention model is used to compress the feature information of the target object, and the feature information with weights is output for target tracking.

Benefits of technology

It realizes fast and accurate target object tracking, meets the real-time tracking needs, simplifies the calculation process, reduces the calculation amount, and improves the response speed and safety of the autonomous driving system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024130394_03072025_PF_FP_ABST
    Figure CN2024130394_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of target tracking. Particularly provided are a target tracking method, and a storage medium and a vehicle. The method comprises: acquiring feature information of a target object within a predetermined time period, and using the feature information as input data; inputting the input data into a pre-trained temporal attention model, such that the temporal attention model performs compression processing on the input data to obtain compressed data, and outputs, on the basis of the compressed data, feature information to which a weight has been applied, wherein the feature information to which the weight has been applied can reflect the probability of the target object being at a predetermined position at a predetermined moment; and on the basis of the feature information to which the weight has been applied, tracking the target object. By means of the technical solution provided in the present application, the target object can be tracked more quickly and accurately, thereby meeting the present real-time tracking requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Target tracking method, storage medium and vehicle

[0001] This application claims priority to Chinese patent application No. 202311867213.X filed on December 28, 2023, entitled “Target Tracking Method, Storage Medium and Vehicle”. The entire contents of the above Chinese patent application are incorporated into this application by reference. Technical Field

[0002] The present application relates to the field of target tracking technology, and specifically provides a target tracking method, a storage medium, and a vehicle. Background Art

[0003] Currently, target tracking technology is used in many fields. For example, in the field of autonomous driving, target tracking technology can continuously track target objects around the vehicle to help the system predict the behavior of other vehicles or pedestrians, and then assist the system in making vehicle driving decisions.

[0004] Most existing target tracking solutions have complex algorithms, long system computation times, and are unable to obtain real-time tracking results. This makes them inadequate for scenarios requiring real-time decision-making based on real-time tracking results, such as autonomous driving. Therefore, a new target tracking solution is needed to address these technical issues.

[0005] Summary of the Invention

[0006] In order to solve the above technical problems, the present application proposes a target tracking method, storage medium and vehicle, which can track the target object more quickly and accurately, thereby meeting the current real-time tracking needs.

[0007] To achieve the above objectives, the technical solution of this application is implemented as follows:

[0008] In a first aspect, the present application provides a target tracking method, the method comprising:

[0009] Acquire characteristic information of a target object within a predetermined time period as input data;

[0010] Inputting the input data into a pre-trained temporal attention model so that the temporal attention model compresses the input data to obtain compressed data, and outputs weighted feature information based on the compressed data; wherein the weighted feature information can reflect the probability of the target object being at a predetermined position at a predetermined time;

[0011] The target object is tracked based on the weighted feature information.

[0012] In some embodiments, obtaining characteristic information of a target object within a predetermined time period as input data includes:

[0013] Acquire each frame of image within a time period from the current moment to a predetermined historical moment; wherein the image is acquired based on an image acquisition device installed at a predetermined position and with a predetermined angle;

[0014] Acquire the target objects in each frame of the image; wherein the number of the target objects in each frame of the image is not greater than a preset number threshold;

[0015] Based on the target object in each frame of the image, obtaining a feature matrix of the target object corresponding to each frame of the image;

[0016] The input data is obtained based on a feature matrix of the target object corresponding to each frame of the image.

[0017] In some embodiments, acquiring the target object in each frame of the image includes:

[0018] For each frame of the image, the following operations are performed: an object of a predetermined type in the frame of the image is acquired as the target object in the frame of the image.

[0019] In some embodiments, obtaining a feature matrix of the target object corresponding to each frame of the image based on the target object in each frame of the image includes:

[0020] For each frame of the image, the following operations are performed: feature extraction is performed on each target object in the frame image to obtain a feature vector of each target object in the frame image; based on the feature vector of each target object in the frame image, a feature matrix of the target object corresponding to the frame image is obtained.

[0021] In some embodiments, the feature vector of each target object in the frame image is represented by a row vector, and each row vector has a predetermined dimension; obtaining a feature matrix of the target object corresponding to the frame image based on the feature vector of each target object in the frame image includes:

[0022] Splicing the feature vectors of each target object in the frame image in the vertical direction to obtain a splicing matrix of the frame image;

[0023] When the number of rows of the splicing matrix of the frame image is equal to the preset number threshold, using the splicing matrix of the frame image as the feature matrix of the target object corresponding to the frame image;

[0024] When the number of rows of the stitching matrix of the frame image is less than the preset number threshold, the preset row vectors are sequentially spliced ​​in the last row of the stitching matrix of the frame image until the number of rows of the stitching matrix of the frame image is equal to the preset number threshold, so as to obtain the feature matrix of the target object corresponding to the frame image; wherein the preset row vector has the predetermined dimension.

[0025] In some embodiments, each value of the preset row vector is 0.

[0026] In some embodiments, the method further comprises:

[0027] The position information of the preset row vector in the feature matrix of the target object corresponding to the frame image is recorded.

[0028] In some embodiments, the feature matrix of the target object corresponding to each frame of the image has a predetermined number of rows and a predetermined number of columns; obtaining the input data based on the feature matrix of the target object corresponding to each frame of the image includes:

[0029] For each frame of the image, the last row of the feature matrix of the target object corresponding to the frame of the image is spliced ​​with the feature matrix of the target object corresponding to the next frame of the image to obtain the input data.

[0030] In some embodiments, the input data includes an input matrix; the input matrix includes at least one preset row vector; the temporal attention model compresses the input data in the following manner to obtain the compressed data:

[0031] Performing a linear transformation on the input matrix using a self-attention mechanism to obtain a query matrix, a key matrix, and a value matrix;

[0032] performing a first removal process on the preset row vectors contained in the query matrix, and concatenating the remaining row vectors in the query matrix according to the arrangement order before the first removal process, to obtain a compressed query matrix;

[0033] performing a second removal process on the preset row vectors contained in the key matrix, and concatenating the remaining row vectors in the key matrix according to the arrangement order before the second removal process, to obtain a compressed key matrix;

[0034] performing a third removal process on the preset row vectors contained in the value matrix, and splicing the remaining row vectors in the value matrix according to the arrangement order before the third removal process, to obtain a compressed value matrix;

[0035] The compressed query matrix, the compressed key matrix, and the compressed value matrix are used as the compressed data.

[0036] In some embodiments, the preset row vector has a predetermined dimension; the temporal attention model outputs the weighted feature information in the following manner:

[0037] Calculating a product of the compressed query matrix and the transpose of the compressed key matrix as a first matrix;

[0038] calculating the arithmetic square root of the predetermined dimension;

[0039] calculating a quotient of the first matrix and the arithmetic square root;

[0040] Normalizing the quotient to obtain a normalized matrix;

[0041] The product of the normalized matrix and the compressed value matrix is ​​calculated as a second matrix, and the second matrix is ​​used as the weighted feature information.

[0042] In some embodiments, the temporal attention model further outputs the weighted feature information in the following manner:

[0043] Inserting the preset row vector at at least one predetermined row position of the second matrix so that the number of rows of the second matrix is ​​equal to the predetermined number of rows, thereby obtaining a padded second matrix;

[0044] The filled second matrix is ​​used as the weighted feature information.

[0045] In a second aspect, the present application provides a computer-readable storage medium, which stores a plurality of program codes, wherein the program codes are suitable for being loaded and run by a processor to execute the target tracking method described in any one of the technical solutions in the above-mentioned first aspect.

[0046] In a third aspect, the present application provides a vehicle comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein a computer program is stored in the memory, and when the computer program is executed by the at least one processor, the target tracking method described in any one of the technical solutions in the first aspect is implemented.

[0047] The target tracking method, storage medium, and vehicle provided in the embodiments of the present application obtain feature information of a target object within a predetermined time period as input data, input the input data into a pre-trained temporal attention model, so that the temporal attention model compresses the input data and outputs weighted feature information based on the compressed data. Based on the weighted feature information, the target object is tracked. This allows the present application to obtain information for target tracking based solely on the feature information of the target object and the temporal attention model. Because the temporal attention model is a pre-trained model, the output of the model can be obtained more quickly and accurately. Moreover, because the weighted feature information output by the model can reflect the probability of the target object being at a predetermined location at a predetermined time, the target object can be tracked more accurately based on the output. Furthermore, the temporal attention model first compresses the input data before performing the corresponding processing, which can further simplify the calculation process, reduce the amount of calculation, and thus obtain the output result more quickly. It can be seen that the technical solution provided by the present application can track the target object more quickly and accurately, thereby meeting the current real-time tracking needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The disclosure of this application will become more easily understood with reference to the accompanying drawings. Those skilled in the art will readily appreciate that these drawings are merely for the purpose of illustrating this application and are not intended to limit the scope of protection of this application. Furthermore, similar numbers in the drawings represent similar components, where:

[0049] FIG1 is a flowchart of the main steps of the target tracking method provided in an embodiment of the present application;

[0050] FIG2 is a schematic diagram of the process of obtaining an input matrix in an embodiment of the present application;

[0051] FIG3 is a schematic diagram of the calculation process of the temporal attention model in an embodiment of the present application;

[0052] Figure 4 is an algorithm flow chart of the temporal attention model in an embodiment of the present application. DETAILED DESCRIPTION

[0053] Some embodiments of the present application are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application and are not intended to limit the scope of protection of the present application.

[0054] In the description of this application, "module" and "processor" may include hardware, software, or a combination of both. A module may include hardware circuitry, various suitable sensors, communication ports, and memory. It may also include software components, such as program code, or a combination of software and hardware. A processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other suitable processor. A processor has data and / or signal processing capabilities. A processor may be implemented in software, hardware, or a combination of both. Non-transitory computer-readable storage media include any suitable medium capable of storing program code, such as magnetic disks, hard disks, optical disks, flash memory, read-only memory, random access memory, etc. The term "A and / or B" refers to all possible combinations of A and B, such as only A, only B, or both A and B. The terms "at least one of A or B" or "at least one of A and B" have similar meanings to "A and / or B" and may include only A, only B, or both A and B. The singular forms "one" and "the" may also include the plural forms.

[0055] Most existing target tracking solutions have complex algorithms, long system calculation times, and cannot obtain tracking results in real time. This cannot meet current application requirements for some scenarios that require real-time decision-making based on real-time tracking results.

[0056] Taking autonomous driving as an example, advanced driver assistance features are gaining increasing attention from users. With advances in sensor and information technology, the user experience is also improving. In autonomous driving systems, tracking objects at different times helps the system predict the behavior of other vehicles or pedestrians and make further decisions. However, existing target tracking solutions consume a lot of computation time, resulting in long response times for autonomous driving tasks and failing to meet current real-time response requirements.

[0057] In order to solve the above technical problems, the present application provides a target tracking method. As shown in FIG1 , the target tracking method in an embodiment of the present application mainly includes the following steps S101 to S103 .

[0058] Step S101, acquiring characteristic information of a target object within a predetermined time period as input data;

[0059] In this embodiment, the predetermined time period is preferably a historical time period, that is, the feature information of the target object in a certain historical time period is used as input data for the subsequent temporal attention model.

[0060] In this embodiment, the target objects at each moment within the above-mentioned predetermined time period may change. For example, the target objects at moment t1 are A, B, and C, and the target objects at moment t2 are A, B, C, and D. The target objects described in this embodiment refer to all target objects within the above-mentioned predetermined time period.

[0061] In order to accurately obtain the characteristic information of the target object and thus obtain more accurate target tracking results, the method of obtaining the characteristic information of the target object within a predetermined time period as input data described in this embodiment includes: obtaining each frame of image within a time period from the current moment to a predetermined historical moment; wherein the image is obtained based on an image acquisition device installed at a predetermined position and with a predetermined angle; obtaining the target object in each frame of the image; wherein the number of the target objects in each frame of the image is not greater than a preset number threshold; based on the target objects in each frame of the image, obtaining a characteristic matrix of the target object corresponding to each frame of the image; and obtaining the input data based on the characteristic matrix of the target object corresponding to each frame of the image.

[0062] Specifically, using the autonomous driving scenario as an example, the image acquisition device can be a camera mounted on the vehicle body, used to capture images of the scene around the vehicle. The target objects can be other vehicles, pedestrians, etc. within a predetermined range around the vehicle. While the vehicle is driving, the camera continuously captures images of the scene around the vehicle. Based on these images, the system can identify and extract features of the target objects in each frame of the scene image, thereby obtaining the input data required for subsequent operations.

[0063] It should be noted that since the system itself specifies the number of target objects that can be processed for each frame of image, the number of target objects obtained when identifying target objects in each frame of image cannot exceed a preset number threshold. For example, the number threshold can be set to 200. If the number of target objects in a frame of image exceeds this number, the target objects farthest from the vehicle can be discarded according to the distance rule to ensure that the number of target objects in the frame of image is below the preset number threshold.

[0064] In order to accurately obtain the target object that the system needs to pay attention to or track for each frame of image, the acquisition of the target object in each frame of image described in this embodiment includes: for each frame of image, performing the following operations: obtaining an object of a predetermined type in the frame of image as the target object in the frame of image.

[0065] Specifically, the type of target object can be pre-set. For example, the system sets the target object as the vehicles located around the vehicle. Then, for each frame image mentioned above, all vehicles located around the vehicle in the frame image are regarded as the target object corresponding to the frame image.

[0066] In order to more accurately obtain the feature matrix of the target object corresponding to each frame of image, the present embodiment obtains the feature matrix of the target object corresponding to each frame of image based on the target object in each frame of image, including: for each frame of image, performing the following operations: extracting features of each target object in the frame of image to obtain a feature vector of each target object in the frame of image; and obtaining the feature matrix of the target object corresponding to the frame of image based on the feature vector of each target object in the frame of image.

[0067] In this embodiment, the feature information of a target object refers to the spatial information of the target object, including its spatial shape and position information, such as its length, width, center coordinates, and current position. Feature extraction of a target object involves extracting the aforementioned spatial information and converting it into feature information that can be used for computation. This feature information can be expressed as a feature vector. Specifically, for each target object in the frame image, feature extraction yields a feature vector reflecting the spatial information of the target object. By combining the feature vectors of each target object, a feature matrix reflecting the spatial information of all target objects in the frame image can be obtained.

[0068] In practical applications, in order to more accurately and comprehensively express the spatial information of a target object, and thus more accurately and comprehensively express the actual motion state of the target object, this embodiment can also convert the extracted feature vectors into higher-dimensional features. For example, a neural network can be used to convert the feature vectors into 64-dimensional features, so that the subsequent temporal attention model can output more accurate feature information.

[0069] In this embodiment, the feature vector of each target object in the frame image is represented by a row vector, and each row vector has a predetermined dimension. Under this premise, in order to accurately obtain a feature matrix that is conducive to subsequent rapid calculations, the feature matrix of the target object corresponding to the frame image is obtained based on the feature vector of each target object in the frame image, as described in this embodiment, including: splicing the feature vectors of each target object in the frame image in the vertical direction to obtain the splicing matrix of the frame image; when the number of rows of the splicing matrix of the frame image is equal to the preset number threshold, the splicing matrix of the frame image is used as the feature matrix of the target object corresponding to the frame image; when the number of rows of the splicing matrix of the frame image is less than the preset number threshold, the preset row vectors are sequentially spliced ​​in the last row of the splicing matrix of the frame image until the number of rows of the splicing matrix of the frame image is equal to the preset number threshold, so as to obtain the feature matrix of the target object corresponding to the frame image; wherein, the preset row vector has the predetermined dimension.

[0070] In this embodiment, the feature vectors of all target objects in each frame are vertically concatenated, and the concatenated matrix is ​​expanded to a predetermined size to obtain a feature matrix corresponding to that frame. This feature matrix is ​​used to reflect the spatial feature information of all target objects in that frame. This operation is repeated for each frame to obtain a feature matrix corresponding to each frame.

[0071] In this embodiment, the number of rows in the feature matrix for each image frame physically represents the number of target objects detected in each frame. This is a non-fixed value, meaning the number of detected target objects in each frame is constantly changing. The number of columns in the feature matrix for each image frame physically represents the characteristic dimension of each target object. This is a preset fixed value that generally remains unchanged once specified in the system. The system extracts target object features based on this dimension.

[0072] Specifically, as shown in Figure 2, assume that three target objects object0, object1, and object2 are detected in the frame image at time t0, and their respective feature vectors are generated. The dimension of these feature vectors is n_channel, for example, the dimension is 7. Assuming that the above-mentioned preset number threshold is n_max, for example, the threshold is also a value of 7, then after the three feature vectors are spliced ​​into a matrix of 3 rows and 7 columns, since the number of rows of the spliced ​​matrix is ​​less than the above-mentioned preset number threshold, the preset row vectors are sequentially spliced ​​in the last row of the spliced ​​matrix until the number of rows of the spliced ​​matrix also reaches 7. In this way, the feature matrix corresponding to the frame image at time t0 can be obtained.

[0073] The processing method of the frame images at time t1 and other times within the above-mentioned preset time period is the same as that at time t0, and will not be repeated here.

[0074] In order to simplify the subsequent calculation process, each value of the preset row vector described in this embodiment is 0.

[0075] That is, in this embodiment, the above-mentioned splicing matrix is ​​padded with the value 0 so that the number of rows of the splicing matrix is ​​equal to the above-mentioned preset number threshold. The purpose of doing this is that the number of target objects appearing in the perception range at each moment is not fixed, and the training framework of the temporal attention model requires a fixed matrix size for each input to improve the model training speed and the model's prediction speed for actual input data. Therefore, this embodiment needs to fill the above-mentioned feature matrix with a fixed size so that the input matrix subsequently input to the temporal attention model meets the model requirements.

[0076] In this embodiment, the fixed size specifically refers to the number of rows and columns of the input matrix. The number of columns in the input matrix corresponds to the dimension of each target object's features; the number of rows in the input matrix depends on the number of image frames within a preset time period and the number of rows in the feature matrix corresponding to each image frame. These parameters can be pre-set. In actual applications, the system collects and processes the target object's feature data according to these pre-set parameters to obtain an input matrix of fixed size.

[0077] In order to enable the system to quickly identify fill data and accurately distinguish fill data from non-fill data, the method described in this embodiment also includes: recording the position information of the preset row vector in the feature matrix of the target object corresponding to the frame image.

[0078] It will be appreciated that in practical applications, the feature vectors for each target object in each image frame can also be represented by column vectors, with each column vector having a predetermined dimension. In this case, when concatenating the feature vectors, the feature vectors of each target object are concatenated horizontally to obtain a concatenated matrix. For concatenated matrices with fewer than a preset number of columns, concatenation is further performed using preset column vectors to ensure that the input matrix meets the required dimensions. The specific details of the above implementation process are the same as those described above for row vectors and are not further elaborated here.

[0079] Specifically, as shown in Figure 2, a mask matrix is ​​used to record the filling position of the above-mentioned preset row vector, so that the temporal attention model can quickly determine which data is filling data and which data is valid feature data when compressing the input matrix in the subsequent data, thereby simplifying the subsequent calculation process.

[0080] As described above, the feature matrix of the target object corresponding to each frame of the image described in this embodiment has a predetermined number of rows and a predetermined number of columns. Under this premise, in order to quickly and accurately obtain the input data of the temporal attention model, the feature matrix of the target object corresponding to each frame of the image described in this embodiment is used to obtain the input data, including: for each frame of the image, the feature matrix of the target object corresponding to the frame of the image is spliced ​​with the feature matrix of the target object corresponding to the next frame of the image at the last row of the feature matrix of the target object corresponding to the frame of the image to obtain the input data.

[0081] That is, in this embodiment, the feature matrices corresponding to each frame of the image are spliced ​​vertically to obtain the input data of the temporal attention model that meets the requirements.

[0082] Specifically, as shown in Figure 2, the system encodes the target frame information of the target object detected in each frame of the previous T frames into higher-dimensional features and fills it to a specific size n_max. For example, the 10 frames before the current moment are selected and recorded as t1 to t10. Assuming that n1 target objects are detected at time t1, each target object is marked with a detection frame. The spatial information of each target object is encoded into a higher-dimensional feature, and the feature vectors of each target frame are spliced ​​together in sequence. The size / dimensionality of this feature is recorded as n_channel, and the actual effective feature size is (n1, n_channel). Afterwards, the feature is padded to (n_max, n_channel), and a mask is recorded to record which positions are filled. Finally, the feature matrices of multiple frames are spliced ​​into a feature matrix of (n_max*T, n_channel), which is the input matrix of the temporal attention model.

[0083] Step S102: Inputting the input data into a pre-trained temporal attention model, so that the temporal attention model compresses the input data to obtain compressed data, and outputs weighted feature information based on the compressed data; wherein the weighted feature information can reflect the probability of the target object being at a predetermined position at a predetermined time;

[0084] This embodiment converts the spatial information of target objects at multiple historical moments (such as center point coordinates, object length and width, etc.) into higher-dimensional features, and uses the self-attention method to comprehensively utilize the information of historical moments to adjust the feature weights of each target object in the historical frame, so as to make subsequent trajectory tracking and behavior prediction more accurate.

[0085] In this embodiment, the input data includes an input matrix, and the input matrix includes at least one preset row vector. That is, in this embodiment, the feature information of each target object is represented in the form of a feature matrix, and the feature matrix is ​​a feature matrix filled with preset row vectors. Under this premise, in order to quickly and effectively compress the input data, the temporal attention model described in this embodiment compresses the input data in the following manner to obtain the compressed data: a self-attention mechanism is used to perform a linear transformation on the input matrix to obtain a query matrix, a key matrix and a value matrix; a first removal process is performed on the preset row vectors contained in the query matrix, and the remaining row vectors in the query matrix are spliced ​​according to the arrangement order before the first removal process to obtain a compressed query matrix; a second removal process is performed on the preset row vectors contained in the key matrix, and the remaining row vectors in the key matrix are spliced ​​according to the arrangement order before the second removal process to obtain a compressed key matrix; a third removal process is performed on the preset row vectors contained in the value matrix, and the remaining row vectors in the value matrix are spliced ​​according to the arrangement order before the third removal process to obtain a compressed value matrix; the compressed query matrix, the compressed key matrix and the compressed value matrix are used as the compressed data.

[0086] In this embodiment, after obtaining the feature matrix and the mask for recording the fill data through the method of step S101, the feature matrix and the mask are input into the temporal attention model. The temporal attention model processes the feature matrix to obtain a feature adjusted by the attention module, which is input into the subsequent trajectory tracking or behavior prediction model. In this embodiment, the temporal attention model uses the self-attention mechanism to perform a linear transformation on the input matrix to obtain a query matrix, a key matrix, and a value matrix. Specifically, the number of rows n_max*T of the input matrix is ​​recorded as n_seq_len, that is, a feature matrix of size (n_seq_len, n_channel) is input. The temporal attention model uses three pre-trained transformation matrices of size (n_channel, n_channel) to multiply the input matrix respectively, and performs three different linear transformations on the input matrix to obtain a query matrix query (Q), a key matrix key (K), and a value matrix value (V) of size (n_seq_len, n_channel). The above linear transformation can improve the fitting ability of the temporal attention model. Then, based on the mask obtained in step S101, all padding data is removed from the query matrix (Q), key matrix (K), and value matrix (V) obtained through linear transformation. The remaining eigenvectors are reassembled in their original order to obtain the compressed query matrix, key matrix, and value matrix. Using these compressed query matrix, key matrix, and value matrix for subsequent calculations can greatly simplify the computational process, reduce the amount of computation, and shorten the computational time, thereby meeting the current demand for real-time tracking of target objects.

[0087] As described above, the preset row vector described in this embodiment has a predetermined dimension. Under this premise, in order to accurately obtain the feature information with weights applied to accurately reflect the probability of a target object being at a predetermined position at a predetermined time, the temporal attention model described in this embodiment outputs the feature information with weights applied in the following manner: calculating the product of the compressed query matrix and the transpose of the compressed key matrix as a first matrix; calculating the arithmetic square root of the predetermined dimension; calculating the quotient of the first matrix and the arithmetic square root; normalizing the quotient to obtain a normalized matrix; calculating the product of the normalized matrix and the compressed value matrix as a second matrix, and using the second matrix as the feature information with weights applied.

[0088] In this embodiment, the product of the transpose of the compressed query matrix and the compressed key matrix is ​​calculated, that is, the similarity between the compressed query matrix and the compressed key matrix is ​​calculated; the quotient of the first matrix and the arithmetic square root is calculated to balance the overall distribution of the above-mentioned normalized matrix.

[0089] Specifically, the temporal attention model in this embodiment uses the following formula to calculate and obtain the weighted feature information:

[0090] Among them, result is the feature information with weights applied, that is, the feature matrix with weights applied; Q0 is the compressed query matrix; K0 is the compressed key matrix; n_channel is the predetermined dimension; V0 is the compressed value matrix.

[0091] The softmax function in the above formula can be expressed as:

[0092] Softmax(x) represents the process of normalizing x.

[0093] In order to make the weighted feature information meet the predetermined size required for subsequent target tracking calculations to further speed up the calculations, the temporal attention model described in this embodiment also outputs the weighted feature information in the following manner: inserting the preset row vector at at least one predetermined row position of the second matrix so that the number of rows of the second matrix is ​​equal to the predetermined number of rows, and obtaining a filled second matrix; and using the filled second matrix as the weighted feature information.

[0094] In this embodiment, the predetermined number of rows is equal to the number of rows of the input matrix. That is, in this embodiment, the size of the output matrix is ​​equal to the size of the input matrix, both being (n_max*T, n_channel).

[0095] Specifically, based on the mask mentioned above, the second matrix is ​​rearranged according to the mask to obtain the filled second matrix. This process is inverse to the process of compressing the matrix. The system first generates a mapping table based on the mask. This generation process can be considered as a process of prefix summing the mask. This table records the row number correspondence between the query matrix Q and the compressed query matrix reduced_Q. For example, it is recorded that the 10th row in Q is recorded in the 5th row of reduced_Q. Then, a reverse search is performed on this table. For example, if the system wants to query the position of the 5th row of reduced_Q in the original Q, then through this reverse search, the system can know that it is in the 10th row of Q. Therefore, the eigenvectors are rearranged in the end, that is, each row of the second matrix is ​​inserted into the corresponding row of the query result according to the query result, and finally the number 0 is filled in the parts that have not been inserted.

[0096] The above calculation process of the temporal attention model is shown in Figures 3 and 4.

[0097] Step S103: Tracking the target object based on the weighted feature information.

[0098] Taking the autonomous driving scenario as an example, in this embodiment, the weighted feature information is referred to as attention-weighted feature information. This means that when the system is tracking a target vehicle at the current moment, looking back at historical moments, the system tends to assume that the vehicle was likely not very far away at the previous moment. Therefore, the system multiplies vehicle features at different distances by different weights to represent the "probability of the vehicle being at this location." This is the "attention-applying process." In addition to this spatial attention, in practical applications, temporal attention can also be applied, such that the system considers feature information closer to the current moment more important.

[0099] Since target tracking essentially involves matching the probability of vehicles being at a certain location at different times based on their speed and position, the goal of target tracking in this embodiment is to accurately match the location information of the same vehicle at different times. Applying attention weights improves the accuracy of this matching, effectively filtering out unreasonable matches. For example, a vehicle that was 100 meters ahead of the vehicle at one moment is unlikely to be 100 meters behind it at the next moment. Therefore, this weighted feature information allows for more accurate matching of the locations of each vehicle at different times, leading to more accurate tracking of each vehicle.

[0100] This application extracts features of the target object in each frame image of a historical moment, obtains feature vectors, splices the feature vectors, obtains the feature matrix of the frame image, fills the feature matrix to a fixed size, and uses a mask to record whether each position is actual valid feature data or filled in. This filled feature matrix is ​​used as the input of the temporal attention model, and the input matrix is ​​multiplied by three transformation matrices to obtain query, key, and value matrices; according to the mask, the query, key, and value matrices are compressed. The compressed query matrix, the compressed key matrix, and the compressed value matrix are fused with matrix multiplication, softmax, matrix multiplication and other operations, and finally a feature matrix with attention weights is obtained as the output.

[0101] Based on the above steps S101-S103, the present application can solve the technical problem in the prior art that tracking results cannot be obtained in real time due to long calculation time.

[0102] The technical solution provided by the embodiment of the present application obtains the feature information of the target object within a predetermined time period as input data, inputs the input data into a pre-trained temporal attention model, and causes the temporal attention model to compress the input data and output weighted feature information based on the compressed data. Based on the weighted feature information, the target object is tracked, so that the present application can obtain information for target tracking based solely on the feature information of the target object and the temporal attention model. Because the temporal attention model is a pre-trained model, the output of the model can be obtained more quickly and accurately. Since the weighted feature information output by the model can reflect the probability of the target object being at a predetermined location at a predetermined time, the target object can be tracked more accurately based on the output. In addition, the temporal attention model first compresses the input data and then performs corresponding processing, which can further simplify the calculation process, reduce the amount of calculation, and thus obtain the output result more quickly. It can be seen that the technical solution provided by the present application can track the target object more quickly and accurately, thereby meeting the current real-time tracking needs.

[0103] The technical solution provided by the embodiment of the present application is particularly suitable for application in autonomous driving scenarios on highways. It is a low-latency implementation method of a temporal attention model for target tracking in autonomous driving highway scenarios. Since the number of targets that a vehicle needs to track in a highway scenario is usually small (usually no more than 100 vehicles), the system obtains fewer feature vectors and the mask has a high sparsity. The method proposed in the present application uses a compressed query matrix, a compressed key matrix, and a compressed value matrix to calculate a weighted feature matrix, which can significantly reduce the computational complexity of the temporal attention model and solve the problem that the prior art cannot track the vehicles in the scene in a timely manner due to the long computation time, thereby causing potential safety hazards. The present application can shorten the response time of autonomous driving tasks such as target object recognition and tracking, reduce response delays, and improve the safety of autonomous driving systems in high-speed scenarios.

[0104] The technical solution provided in the embodiments of the present application designs a new computing solution based on the characteristics of fewer targets and higher latency requirements in autonomous driving highway scenarios, which can be used for real-time target tracking / target tracing tasks.

[0105] It should be pointed out that although the various steps in the above embodiments are described in a specific order, those skilled in the art will understand that in order to achieve the effect of the present application, different steps do not have to be performed in such an order. They can be performed simultaneously (in parallel) or in other orders. These changes are within the scope of protection of the present application.

[0106] Furthermore, another aspect of the present application also provides a computer-readable storage medium. In a computer-readable storage medium embodiment according to the present application, the computer-readable storage medium can be configured to store a program for executing the target tracking method of the above-mentioned method embodiment, and the program can be loaded and run by the processor to implement the above-mentioned target tracking method. For ease of explanation, only the parts related to the embodiment of the present application are shown. For specific technical details not disclosed, please refer to the method part of the embodiment of the present application. The computer-readable storage medium can be a storage device formed by various electronic devices. Optionally, the computer-readable storage medium in the embodiment of the present application is a non-transitory computer-readable storage medium.

[0107] Furthermore, another aspect of the present application provides a vehicle, which may include at least one processor and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program, and when executed by the at least one processor, the computer program implements the target tracking method described in any of the above embodiments. The vehicle described in the present application may include a driving device, an intelligent vehicle, a robot, or other devices.

[0108] In some embodiments of the present application, the vehicle further comprises at least one sensor configured to sense information. The sensor is communicatively coupled to any of the processors described herein. Optionally, the vehicle further comprises an autonomous driving system configured to guide the vehicle's autonomous driving or assist in driving. The processor communicates with the sensor and / or autonomous driving system to implement the target tracking method described in any of the above embodiments.

[0109] Automated Driving Systems (ADS) refers to systems that continuously perform all dynamic driving tasks (DDT) within their operational domain design (ODD). Specifically, the system is only allowed to fully assume the responsibility of autonomous vehicle control under specified driving scenarios. When the vehicle meets the ODD conditions, the system is activated, replacing the human driver as the vehicle's primary driver. The DDT refers to the continuous lateral (left and right steering) and longitudinal motion control (acceleration, deceleration, and constant speed) of the vehicle, as well as the detection and response to objects and events in the vehicle's driving environment. The ODD refers to the conditions under which the automated driving system can safely operate. These conditions can include geographic location, road type, speed range, weather, time of day, and national and local traffic laws and regulations.

[0110] The program for executing the target tracking method of the above-mentioned method embodiment can be divided into multiple subroutines, each of which can be loaded and run by a processor to execute different steps of the target tracking method of the above-mentioned method embodiment. Specifically, each subroutine can be stored in different memories, and each processor can be configured to execute the programs in one or more memories to jointly implement the target tracking method of the above-mentioned method embodiment, that is, each processor executes different steps of the above-mentioned method embodiment to jointly implement the target tracking method of the above-mentioned method embodiment.

[0111] The aforementioned multiple processors may be processors deployed on the same device. For example, the aforementioned vehicle may be composed of multiple processors, and the aforementioned multiple processors may be processors configured on the vehicle. Furthermore, the aforementioned multiple processors may be processors deployed on different devices. For example, the aforementioned vehicle may be a server cluster, and the aforementioned multiple processors may be processors on different servers in the server cluster.

[0112] It will be understood by those skilled in the art that all or part of the processes in the method for implementing the above embodiment of the present application can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium can include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electric carrier signal, telecommunication signal and software distribution medium that can carry the computer program code. It should be noted that the content contained in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable storage media do not include electric carrier signals and telecommunication signals.

[0113] Thus far, the technical solutions of the present application have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is readily understood by those skilled in the art that the scope of protection of the present application is obviously not limited to these specific embodiments. Without departing from the principles of the present application, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present application.

Claims

1. A target tracking method, characterized in that, The method includes: Obtaining the feature information of the target object within a predetermined time period as input data; Inputting the input data into a pre-trained temporal attention model, so that the temporal attention model performs compression processing on the input data to obtain compressed data, and outputs weighted feature information based on the compressed data; wherein, the weighted feature information can reflect the probability of the target object at a predetermined position at a predetermined moment; Tracking the target object based on the weighted feature information.

2. The target tracking method according to claim 1, characterized in that The obtaining the feature information of the target object within a predetermined time period as input data includes: Obtaining each frame of image within the time period from the current moment to a predetermined historical moment; wherein, the image is acquired by an image acquisition device installed at a predetermined position and having a predetermined angle; Obtaining the target object in each frame of the image; wherein, the number of target objects in each frame of the image is not greater than a preset number threshold; Based on the target object in each frame of the image, obtaining the feature matrix of the target object corresponding to each frame of the image; Based on the feature matrix of the target object corresponding to each frame of the image, obtaining the input data.

3. The target tracking method according to claim 2, wherein The obtaining the target object in each frame of the image includes: For each frame of the image, perform the following operation: Obtaining the object with a predetermined type in the frame of the image as the target object in the frame of the image.

4. The target tracking method according to claim 2, wherein The obtaining the feature matrix of the target object corresponding to each frame of the image based on the target object in each frame of the image includes: For each frame of the image, perform the following operation: For each target object in the frame of the image, perform feature extraction to obtain the feature vector of each target object in the frame of the image; Based on the feature vectors of each target object in the frame of the image, obtaining the feature matrix of the target object corresponding to the frame of the image.

5. The target tracking method according to claim 4, wherein, The feature vector of each target object in the frame of the image is represented by a row vector, and each row vector has a predetermined dimension; the obtaining the feature matrix of the target object corresponding to the frame of the image based on the feature vectors of each target object in the frame of the image includes: Concatenating the feature vectors of each target object in the frame of the image in the vertical direction to obtain the concatenated matrix of the frame of the image; When the number of rows of the concatenated matrix of the frame of the image is equal to the preset number threshold, using the concatenated matrix of the frame of the image as the feature matrix of the target object corresponding to the frame of the image; When the number of rows of the concatenated matrix of the frame of the image is less than the preset number threshold, sequentially concatenating preset row vectors at the last row of the concatenated matrix of the frame of the image until the number of rows of the concatenated matrix of the frame of the image is equal to the preset number threshold, so as to obtain the feature matrix of the target object corresponding to the frame of the image; wherein, the preset row vector has the predetermined dimension.

6. The target tracking method according to claim 5, wherein, Each numerical value of the preset row vector is 0.

7. The target tracking method according to claim 5, wherein The method further includes: Recording the position information of the preset row vector in the feature matrix of the target object corresponding to the frame of the image.

8. The target tracking method according to claim 2, wherein The feature matrix of the target object corresponding to each frame of the image has a predetermined number of rows and a predetermined number of columns; obtaining the input data based on the feature matrix of the target object corresponding to each frame of the image includes: For each frame of the image, concatenate the feature matrix of the target object corresponding to the next frame at the last row of the feature matrix of the target object corresponding to this frame of the image to obtain the input data. The input data includes an input matrix; the input matrix contains at least one preset row vector; the temporal attention model compresses the input data in the following manner to obtain the compressed data:

9. The target tracking method according to claim 1, wherein Perform a linear transformation on the input matrix using a self-attention mechanism to obtain a query matrix, a key matrix, and a value matrix; Perform a first removal process on the preset row vector contained in the query matrix, and concatenate the remaining row vectors in the query matrix in the arrangement order before the first removal process to obtain a compressed query matrix; Perform a second removal process on the preset row vector contained in the key matrix, and concatenate the remaining row vectors in the key matrix in the arrangement order before the second removal process to obtain a compressed key matrix; Perform a third removal process on the preset row vector contained in the value matrix, and concatenate the remaining row vectors in the value matrix in the arrangement order before the third removal process to obtain a compressed value matrix; Use the compressed query matrix, the compressed key matrix, and the compressed value matrix as the compressed data. The preset row vector has a predetermined dimension; the temporal attention model outputs the weighted feature information in the following manner:

10. The target tracking method according to claim 9, wherein Calculate the product of the compressed query matrix and the transpose of the compressed key matrix as the first matrix; Calculate the arithmetic square root of the predetermined dimension; Calculate the quotient of the first matrix and the arithmetic square root; Normalize the quotient to obtain a normalized matrix; Calculate the product of the normalized matrix and the compressed value matrix as the second matrix, Use the second matrix as the weighted feature information. The temporal attention model also outputs the weighted feature information in the following manner:

11. The target tracking method according to claim 10, characterized in that, Insert the preset row vector at at least one predetermined row position of the second matrix so that the number of rows of the second matrix is equal to the predetermined number of rows to obtain a filled second matrix; Use the filled second matrix as the weighted feature information. The program code is adapted to be loaded and run by a processor to execute the target tracking method according to any one of claims 1 to 11.

12. A computer-readable storage medium storing multiple program codes, characterized in that, Includes:

13. A vehicle, characterized in that, At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores a computer program, and when the computer program is executed by the at least one processor, the target tracking method according to any one of claims 1 to 11 is implemented. ​

Citation Information

Patent Citations

  • Target tracking method and device, storage medium and electronic equipment

    CN110555405A

  • Multi-target tracking method and device, electronic equipment and storage medium

    CN112529934A

  • Traffic video structured data generation method and device and medium

    CN116069801A

  • Image data processing method and device, equipment and medium

    CN116977663A

  • Target tracking method, storage medium and vehicle

    CN117765030A