A data association method and system based on a cross-Transformer network architecture

By adopting a data association method based on the cross-transformer network architecture in multi-objective tracking technology, the dependence on real posterior density and high computing storage requirements in the prior art are solved, and more accurate and efficient data association is achieved.

CN119272809BActive Publication Date: 2025-06-24HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411289608.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-14
Publication Date
2025-06-24
Estimated Expiration
2044-09-14

AI Technical Summary

Technical Problem

Existing multi-objective tracking techniques rely on prior information of true posterior density, resulting in inaccurate data correlation and high computation and storage requirements.

Method used

The data association method based on the cross Transformer network architecture is adopted, and data standardization is carried out by defining the relative distance feature matrix and data filling block technology, and the row-crossing Transformer network and 3-channel cross-transformer network are constructed, and the association matrix is ​​transformed using the Kuhn-Munkres algorithm.

Benefits of technology

In the presence of false positives and missed detections, multiple targets are effectively processed, reducing dependence on real posterior density, improving the accuracy of data associations, and reducing computing and storage requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119272809B_ABST
    Figure CN119272809B_ABST
Patent Text Reader

Abstract

The present invention provides a data association method and system based on a cross Transformer network architecture, belonging to the technical field of artificial intelligence. In order to solve the problem that existing multi-object tracking has false alarms and missed detections in the network architecture, resulting in a decrease in the accuracy of data association results. The present invention defines a relative distance feature matrix, which allows data standardization by adopting data filling and block technology; then constructs a data association network based on cross Transformer, which effectively extracts features from the distance matrix corresponding to consecutive time instances; in addition, a data post-processing algorithm is used to refine the association results obtained from the CTDA network, so as to ensure accurate target tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and more particularly, to a data association method and system based on a cross Transformer network architecture. Background Art

[0002] Multi-target tracking (MTT) or multi-object tracking (MOT) plays a crucial role in various fields such as surveillance, autonomous driving, and robotics. The MTT problem refers to associating the relative measurements obtained by sensors over time with multiple moving targets (objects) and estimating their trajectories and other attributes (speed, acceleration, etc.). A key problem in MTT is to solve the unknown correspondence between objects and measurements, and this task is called data association. In scenarios with occlusion, sensor noise, clutter, and the possibility of objects entering or leaving the scene, the MTT task is highly challenging.

[0003] To address these challenges, many algorithms and methods for multi-target tracking have been proposed, which can generally be classified into maximum likelihood data association methods and Bayesian data association methods. Maximum likelihood data association methods are based on the likelihood ratio of observed data for association. Classical algorithms include the joint maximum likelihood algorithm, integer programming, generalized association methods, etc. Maximum likelihood data association methods are not applicable to cases with crossing objects, maneuvering objects, and dense multi-object scenarios, thus limiting their application in practice. Bayesian data association methods are based on the Bayesian criterion and mainly include the nearest neighbor (NN) algorithm, probabilistic data association (PDA) algorithm, joint probabilistic data association (JPDA) algorithm, and multiple hypothesis tracking (MHT) method, etc. These methods, such as JPDA, require the assumption of a constant number of objects. The computational complexity of the MHT method increases rapidly with the increase in the number of objects. Another class of MTT methods is based on random finite sets (RFS), which uses multi-object conjugate priors to derive a closed-form expression for the multi-object posterior and obtains the Bayesian optimal estimate. Typical algorithms include the probability hypothesis density (PHD) filter, cardinalized probability hypothesis density (CPHD) filter, generalized labeled multi-Bernoulli (GLMB) tracker. Methods based on RFS can transmit more information between different time steps, thus improving the tracking results. However, these methods usually rely on prior information about the true posterior density. The approximation of the true posterior density may lead to inaccurate data association. In addition, due to the use of high-dimensional sampling processes, the computational and storage requirements of these methods have increased. Summary of the Invention

[0004] The technical problem to be solved by the present invention is:

[0005] To solve the problems of existing multi-object tracking that rely on the prior information of the true posterior density in the network architecture, resulting in inaccurate data association, as well as high computational and storage requirements.

[0006] The technical solution adopted by the present invention to solve the above technical problems:

[0007] The present invention provides a data association method based on a cross Transformer network architecture, including the following steps:

[0008] S100. Define the relative distance feature matrix in the x, y, and z directions to ensure the generalization of the normalized features of data at different positions;

[0009] S200. Data preprocessing, introducing a chunking method and a padding method, filling the data in the dataset D1 and the dataset D2 to be associated with additional data of an appropriate dimension and performing chunking processing, and using the enhanced matrix containing relative distance features as the input of the network;

[0010] S300. Construct a data association network based on cross Transformer, including constructing a Transformer encoder and a Transformer decoder, and then obtaining an output matrix using a row-column cross Transformer network and a 3-channel cross Transformer network;

[0011] S400. Each element of the matrix obtained from the data association network based on cross Transformer in step S300 represents the association probability of a given input data pair. Combine the Kuhn-Munkres algorithm to find the optimal assignment, and then convert the network output into an association matrix.

[0012] Further, in step S100, it includes:

[0013] Let And D1 and D2 respectively represent two groups of position data;

[0014] Define m×n matrices of relative distances in the x, y, and z directions respectively, that is, relative distance feature matrices, for characterizing the association relationship features between the dataset D1 and the dataset D2;

[0015] Among them, for the x-axis direction,

[0016] Let The relative distance feature matrix for x-direction association is as follows,

[0017]

[0018] Wherein, And are the x-direction data from dataset D1 and dataset D2 respectively; x bias is a small quantity used to avoid data errors caused by division by zero; x mid is the geometric median of the absolute values of all relative distances between the two datasets,

[0019]

[0020] For the relative distance feature matrices in the y-axis direction and z-axis direction, they are respectively denoted as and

[0021] Furthermore, in step S200, it includes:

[0022] Let N pad be a fixed value, and N pad is less than or equal to m and n; Let and be the integers of the rows and columns of the scaled padding matrix respectively, is the ceiling operation;

[0023] Let be the padding matrix associated with x after partitioning,

[0024]

[0025] Partition as follows,

[0026]

[0027] wherein, and are the number of partitions of the rows and columns respectively; each partition sub-matrix is an N pad ×N pad matrix;

[0028] Using the partition defined in formula (4), the scaled padding matrix will be reshaped into a batch matrix as follows,

[0029]

[0030] The padding matrices related to directions y and z are defined in the same way as the direction x; subsequently, the enhanced matrix containing relative distance features is used as the network input.

[0031] Furthermore, in step S300, it includes:

[0032] S310. Construct a Transformer encoder, which includes three integrated modules, namely, multi-head attention, feed-forward network, and normalization network;

[0033] The self-attention neural network uses Query to retrieve the attention values corresponding to Keys and assigns them to the corresponding Values; then redistributes the Value according to the attention weights; let Attention denote the attention processing,

[0034]

[0035] where, and represent the Query and Key matrices respectively; represents the y matrix value, d is the dimension of Q, K, and V in the self-attention neural network, and Q, K, and V represent the same feature encoding;

[0036] Use matrices W Q , W K and W V in combination with the multi-head attention network to extract different types of relationships from the data,

[0037] MultiHead(Q, K, V) = Concat(head1, …, head n )W O (7)

[0038] where, is each attention network head, and W O are trainable parameters, which are equivalent to the Q, K, and V tensors having the same shape;

[0039] S320. Construct a Transformer decoder, which includes two multi-head attention networks, set a mask layer in the first multi-head attention network; use Q, K, and V in the second multi-head attention network;

[0040] Use the decoder input as the query feature, and the encoder input as the Key and Value; combine the resulting weighted input encoding with the decoder input, and pass through normalization and a feed-forward neural network to generate the decoder output;

[0041] S330. Combine the Transformer encoder in step S310 and the Transformer decoder in step S320 to construct a data association network based on cross-Transformer, including,

[0042] S331. Construct a row-column cross Transformer network, which is applied to three channels of the input matrix; each channel is processed separately by a row Transformer encoder and a column Transformer encoder, regarding the input rows and columns as feature vectors respectively; then the outputs of the encoders are input into the row and column Transformer decoders in a cross pattern.

[0043] Introduce an element-wise method.

[0044]

[0045] Where he is the output of the mixing matrix, Reshape(·) represents converting the dimensions of the matrix to a specified number, [·] represents the concatenation operation, W and b are trainable weights and biases, and h1 and h2 are decoders.

[0046] S332. Construct a 3-channel cross Transformer network. The output of the element-wise mixing processing module composed of 3-channel data is guided into three Transformer encoders to obtain encoded features; then, the output of each encoder corresponding to each channel is simultaneously fed into the first MHA module of the decoders of other channels; each output is simultaneously used as the key and value inputs of the second MHA module in the corresponding decoder of the channel; then the output matrix is obtained through element-wise mixing.

[0047]

[0048] Where O1, O2, and O3 are the outputs of the decoders corresponding to the three channels.

[0049] Furthermore, in step S300, it also includes using binary cross-entropy loss as the cost to calculate the relative association probability of the network output.

[0050]

[0051] Where and are the elements of the i-th row and j-th column predicted by the association matrix O pr and the true association matrix O gt respectively.

[0052] Furthermore, it can be used for multi-object tracking or multi-object tracing in monitoring, autonomous driving, and robotics to establish associations between measurement values and multiple targets, or to solve the unknown correspondence between targets and measurement values.

[0053] A data association system based on a cross-Transformer network architecture, the system having program modules corresponding to the above steps and executing the steps in the above data association method based on the cross-Transformer network architecture when running.

[0054] A computer-readable storage medium storing a computer program configured to implement the steps of the data association method based on the cross-Transformer network architecture when called by a processor.

[0055] Compared with the prior art, the beneficial effects of the present invention are:

[0056] The present invention relates to a data association method and system based on a cross-Transformer network architecture, which is used to process multiple targets in the presence of false alarms and missed detections; to address the challenges of different data quantities and sizes in different application tasks, a relative distance feature matrix is defined, which allows data normalization by adopting data filling and chunking techniques; then a data association network based on Cross Transformer is constructed, which effectively extracts features from the distance matrix corresponding to consecutive time instances; in addition, a data post-processing algorithm is used to refine the association results obtained from the CTDA network, thereby ensuring accurate target tracking.

[0057] The data filling technique of the present invention has a clearer physical meaning, because it is equivalent to generating multiple false alarms at an infinite distance from the original data; therefore, during the entire filling and partitioning process, no additional information is introduced into the RDF matrix, ensuring that the basic characteristics of the matrix are not affected after the filling process. The cross-Transformer module constructed by the present invention is relatively simple to operate compared with the existing row-column Transformer module, and the number of Transformers in the cross structure applied between each pair of channels will vary with the dimension of the data, and the number of encoders and decoders can also be different, with a wider scope of application. Description of the Drawings

[0058] Figure 1 It is a flowchart of a data association method based on a cross-Transformer network architecture in an embodiment of the present invention;

[0059] Figure 2 It is a flowchart of data filling and chunking during data preprocessing in an embodiment of the present invention;

[0060] Figure 3 It is a structural diagram of an encoder network based on Transformer in an embodiment of the present invention;

[0061] Figure 4This is the structural diagram of the decoder network based on Transformer in the embodiments of the present invention;

[0062] Figure 5 This is the data association network diagram based on cross Transformer in the embodiments of the present invention;

[0063] Figure 6 This is the processing method diagram of the Kuhn - Munkres (KM) algorithm in the embodiments of the present invention;

[0064] Figure 7 This is the test result diagram of this algorithm in the embodiments of the present invention, where (a), (b), (c) and (d) respectively correspond to the test result diagrams of four cases.

[0065] Explanation of reference numerals: Detailed implementation manners

[0066] To make the above - mentioned objects, features and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be given with reference to the accompanying drawings.

[0067] The data association process aims to solve the problem of determining whether a set of data (which can be measurement data or estimated data) originates from the same target. In the MTT process, new target detection requires "measurement - measurement" data association across multiple sampling periods. To update the track and maintain the tracking estimate, "measurement - track" association is needed to identify new measurement data for track correction. The correlation between these associations, including "measurement - measurement" and "measurement - track" data associations, is described by an association matrix. Given two sets of data:

[0068] and To be associated, the association matrix of the two sets of data is expressed as

[0069]

[0070] where

[0071]

[0072] It should be noted that each column and each row of the association matrix allows at most one non - zero element, that is, 1, and all other elements are 0. The purpose of this algorithm is to estimate the states of different numbers of targets by using the position measurements of noise - based radars.

[0073] Specific implementation manner 1: As shown in Figures 1 to 5 The present invention provides a data association method based on a cross - Transformer network architecture, including the following steps:

[0074] S100. Define the Relative Distance Feature (RDF) matrix to ensure the generalization of the normalized features of data at different positions, including:

[0075] Let and D1 and D2 respectively represent two sets of position data, which can be measurements, estimates, and predictions from real targets or false alarms.

[0076] To characterize the correlation relationship features between datasets D1 and D2, m×n matrices regarding the relative distances in the x, y, and z directions are respectively defined, which are called the Relative Distance Feature (RDF) matrices.

[0077] Among them, for the x-axis direction,

[0078] Let The RDF matrix associated with the x direction is as follows:

[0079]

[0080] Among them, and are respectively the x-direction data from dataset D1 and dataset D2; x bias is a small quantity with an appropriate size to avoid data errors caused by division by zero; x mid is the geometric median of the absolute values of all relative distances between the two datasets.

[0081]

[0082] For the RDF matrices in the y-axis direction and z-axis direction, they are respectively expressed as and and can be constructed in the same way.

[0083] In the RDF matrix, values close to 1 indicate that the corresponding data points in datasets D1 and D2 are similar and are likely to be correlated with each other; conversely, values close to 0 indicate that the data points are significantly different and do not belong to the same target.

[0084] S200. Data preprocessing. Introduce the chunking method and the padding method to fill the data in the datasets D1 and D2 to be associated with additional data of appropriate dimensions and perform chunking processing, including:

[0085] In many tasks, there are significant variations in the number and dimensions of the acquired measurement values; in addition, even in a single task, the number of measurements at different times may also vary.

[0086] In the case of tracking a limited number of targets, appropriate maximum padding values can be pre-allocated; however, when faced with a large number of targets and clutter, overly large padding values become computationally infeasible; in order to make the data association algorithm proposed by the present invention applicable to data sets of different sizes, a superior data partitioning technique needs to be introduced.

[0087] Let N pad be a fixed value, and N pad be less than or equal to m and n; let and be the integers of the rows and columns of the scaled padding matrix respectively, be the ceiling operation;

[0088] Let be the padding matrix associated with x after partitioning,

[0089]

[0090] In order to adapt to the input dimension of the CTDA network, further partitioning is performed as follows,

[0091]

[0092] wherein, and are the number of row and column partitions respectively; each partition sub-matrix is an N pad ×N pad matrix;

[0093] Using the partitioning defined in formula (6), the scaled padding matrix will be reshaped into a batch matrix as follows,

[0094]

[0095] The padding matrices related to directions y and z can be defined in the same way as the direction x; subsequently, the enhanced matrix containing relative distance features is used as the input of the CTDA network;

[0096] Compared with the prior art, the data padding technique of the present invention has a clearer physical meaning, because it is equivalent to generating multiple false alarms at an infinite distance from the original data; therefore, during the entire padding and partitioning process, no additional information is introduced into the RDF matrix; it ensures that the basic characteristics of the matrix after the padding process are not affected;

[0097] S300. Construct a data association network based on cross-Transformer, including,

[0098] S310. Construct a Transformer encoder, which includes three integrated modules, namely, multi-head attention (MHA), feed-forward network (FFN), and normalization network. As shown in Figure 3 , among which multi-head attention, that is, the attention module, is particularly important. It is the cornerstone of the MHA backbone;

[0099] The main objective of the self-attention neural network is to use the Query to retrieve the attention values corresponding to the Keys and assign them to the corresponding Values; then redistribute the Value according to the attention weights. Let Attention represent the attention processing, and its form is as follows.

[0100]

[0101] Among them, and represent the Query and Key matrices respectively; represents the y matrix value, d is the dimension of Q, K, and V in the self-attention neural network, and Q, K, and V represent the same feature encoding;

[0102] In addition, by using the matrices W Q , W K , and W V combined with the multi-head attention network, different types of relationships can be extracted from the data.

[0103] MultiHead(Q, K, V) = Concat(head1, …, head n )W O (9)

[0104] Among them, is each attention network head, and W O are trainable parameters, which are equivalent to the Q, K, and V tensors having the same shape;

[0105] S320. Construct a Transformer decoder, which includes two multi-head attention networks (MHA). Set a mask layer in the first multi-head attention network. The mask layer can be configured as a matrix with all elements equal to 1, so that it is the same as the standard multi-head attention network; use Q, K, and V in the second multi-head attention network.

[0106] Specifically, the decoder input is used as the query feature, while the encoder input is used as the Key and Value; this arrangement enables the decoder to utilize the encoded features obtained through multi-head attention as queries to retrieve information from the encoder output; subsequently, the resulting weighted input encoding is combined with the decoder input and passed through normalization and a feed-forward neural network to produce the decoder output;

[0107] S330. Construct a cross-Transformer data association network (CTDA) by combining the Transformer encoder in step S310 and the Transformer decoder in step S320, including:

[0108] S331. Construct a row-column cross Transformer network, which is applied to three channels of the input matrix; each channel is processed separately by a row Transformer encoder and a column Transformer encoder, considering the input rows and columns as feature vectors respectively; it should be noted that the input of each channel is processed independently; then, the output of the encoder is input to the row and column Transformer decoders in a cross mode;

[0109] Since the Transformer decoder includes two MHA modules, in the context of the row-column cross Transformer network, the outputs of the row and column encoders serve as the Key and Value inputs of the second MHA module in the row and column Transformer decoders respectively; meanwhile, the output of each encoder is cross-fed into the other Transformer decoder as the input of the first MHA module layer; the outputs of the row and column decoders (each channel consists of two matrices) are combined using a processing module; to meet different requirements, an element-wise method is introduced to improve the efficiency of this algorithm,

[0110]

[0111] where h e is the output of the mixing matrix, Reshape(·) represents converting the dimension of the matrix to a specified number, [·] represents the concatenation operation, W and b are trainable weights and biases, and h1 and h2 are the decoders;

[0112] S332. Build a 3-channel cross Transformer network, which implements a cross Transformer network structure for the 3-channel encoded data obtained through bidirectional processing. The output of the element-wise mixing processing module composed of 3-channel data is initially directed to three Transformer encoders to obtain encoded features. Then, the output of each encoder corresponding to each channel is simultaneously fed into the first MHA module of the decoders of other channels. In addition, each output is simultaneously used as the key and value input of the second MHA module in the corresponding decoder of the channel. Then, an output matrix is obtained through element-wise mixing.

[0113]

[0114] where O1, O2, and O3 are the outputs of the decoders corresponding to the three channels.

[0115] The cross Transformer adopted in the 3-channel cross Transformer network is different from the corresponding cross Transformer in the row-column cross Transformer. First, compared with the row-column Transformer module, the 3-channel cross Transformer network contains fewer layers. This adjustment is made by observing that the information fusion between different channels is often relatively less complex than the operations required by the row-column cross Transformer module, and using a deeper network in this specific case may lead to overfitting. Second, the number of Transformers in the cross structure applied between each pair of channels varies with the dimension of the data, rather than being fixed. In addition, the number of encoders and decoders can be different.

[0116] S340. Since the data association matrix exhibits sparsity and is composed of binary values (0-1 matrix), the binary cross-entropy loss is used to calculate the relative association probability of the network output as the cost.

[0117]

[0118] where and are the elements of the i-th row and j-th column predicted for the association matrix O pr and the true association matrix O gt respectively.

[0119] S400. Each element of the matrix obtained from the cross Transformer-based data association network represents the association probability corresponding to the given input data. To further convert the network output into an association matrix, the Kuhn-Munkres (KM) algorithm, also known as the Hungarian matching algorithm, is combined to find the optimal assignment. The specific steps are as followsFigure 6 as shown;

[0120] It should be noted that the KM algorithm is designed specifically to solve complex situations and is very useful in dealing with more challenging scenarios; in most cases, it is sufficient to directly apply a threshold to the network output to generate the association matrix.

[0121] Specific Embodiment 2: The present invention provides a data association system based on a cross-Transformer network architecture, which system has program modules corresponding to the above steps and, when running, executes the steps in the above-described data association method based on a cross-Transformer network architecture.

[0122] The other combinations and connection relationships in this embodiment are the same as those in Specific Embodiment 1.

[0123] Specific Embodiment 3: The present invention provides a computer-readable storage medium storing a computer program configured to implement the steps of the data association method based on a cross-Transformer network architecture when called by a processor.

[0124] The other combinations and connection relationships in this embodiment are the same as those in Specific Embodiment 1.

[0125] Simulation Experiment

[0126] A customized 3D dataset is used to evaluate the association accuracy of the CTDA algorithm.

[0127] Evaluation Metrics: When evaluating the association between two datasets, accuracy is usually used as an evaluation metric. However, since the two sets of data do not always correspond one-to-one considering false alarms and missed detections, a more comprehensive accuracy evaluation metric is adopted as follows:

[0128]

[0129] where a ij and are the elements in row i and column j of the correlation matrix A m×n and the true correlation matrix respectively;

[0130]

[0131] Set up the dataset and training. This algorithm can associate two-dimensional or three-dimensional trajectory data. The advantage is that it does not require prior quantitative information, such as false alarm probability and miss detection probability. To train the network in this algorithm, a large amount of data is needed to enhance the generalization ability of the network and prevent overfitting. However, obtaining real trajectory data can be challenging, and the available data is usually limited. To overcome these limitations, we developed a comprehensive multi-objective simulation platform to generate a large amount of simulated data for trajectory measurement for network training.

[0132] First, establish a 3D dynamic model and a measurement model to generate training data. Select the state vector as x = [x, y, z, v, ζ, ψ] T and the measurement z = [r, θ, φ],

[0133]

[0134] where x, y, and z represent the position of the target; v, ζ, and ψ represent the velocity, pitch angle, and yaw angle respectively; θ and φ are the azimuth angle and elevation angle respectively; u D , u L and u G are the controls of the velocity, pitch angle, and yaw angle respectively; η is the process noise and measurement noise vector.

[0135] The parameters of the customized multi-objective dataset are shown in Table 1 below,

[0136] Table 1

[0137]

[0138] Using these parameters, a total of 2000 trajectories were generated. Among them, 1800 trajectories were used for training, and the remaining 200 trajectories were reserved for testing the network performance.

[0139] The network was trained for 8 epochs on the training dataset with a batch size of 256. The Adam optimizer was adopted with a learning rate of 10 -4 . Since the training process was relatively short, learning rate decay was not used.

[0140] In the CTDA network of this algorithm, the selected data dimension is 40, which corresponds to the size of the input and output matrices. The multi-head attention (MHA) module in CTDA uses 4 heads. A dropout rate of 0.3 is selected for regularization. During testing and filtering, a threshold of 0.1 (λ a = 0.1) is selected.

[0141] To prove the accuracy of the CTDA network, consider the following four cases:

[0142] Case 1: The number of clutter is zero, the target number ranges from 0 to 100, and the measurement noise is zero;

[0143] Case 2: The target number is set to 20, the clutter number ranges from 0 to 100, and the measurement noise is zero;

[0144] Case 3: The number of clutter is 20, and the target number ranges from 0 to 100;

[0145] Case 4: The number of clutter is 20, the target number is 10, and the false negative rate ranges from 0% to 100%.

[0146] The test results of the CTDA network are as Figure 7 shown. First, compare the association accuracy between the method proposed in the present invention and the KM assignment algorithm in Case 1; as Figure 7 shown in (a) of , it can be seen that only the relative distance accuracy of the KM algorithm drops rapidly with respect to the target number. By combining the KM algorithm and the normalization technique given in formula (3), the accuracy becomes very close to 1 in most cases. For the CTDA method, the accuracy rate exceeds 90% in the test scenario, and can reach 95% in most cases, but it is slightly lower than the normalized KM algorithm. This is because the training data set of the network contains false negatives and false positives, and in the scenario where there are only targets without false positives and false negatives, the recognition accuracy rate will reduce the accuracy rate.

[0147] As Figure 7 shown in (b) of , as the number of clutter (i.e., false alarms) increases, the KM algorithm becomes invalid. In contrast, the CTDA method demonstrates reliable association performance and can achieve an accuracy rate of more than 75% even in the presence of a large number of false alarms.

[0148] The association results of Case 3 and Case 4 are respectively as Figure 7 shown in (c) of and Figure 7 shown in (d) of . The CTDA network proposed in the present invention shows the adaptability to handle unknown and different numbers of targets to be tracked and scenarios involving false negatives.

[0149] Although the present invention is disclosed as above, the protection scope of the present invention is not limited thereto. Those skilled in the art of the present invention can make various changes and modifications without departing from the spirit and scope of the present invention, and these changes and modifications will all fall within the protection scope of the present invention.

Claims

1. A data association method based on a cross Transformer network architecture, characterized in that: The following steps are involved: S100, defining a relative distance feature matrix about the x, y and z directions to ensure the generalization of normalized features of data at different positions; S200, data preprocessing, introducing a block method and a filling method, filling the data in the to-be-associated data set D1 and the data set D2 with additional data of appropriate dimensions and performing block processing, and taking the enhanced matrix containing relative distance features as the input of the network; S300, constructing a data association network based on a cross Transformer, including constructing a Transformer encoder and a Transformer decoder, and then using a row-column cross Transformer network and a 3-channel cross Transformer network to obtain an output matrix; Among them, S330, combining the Transformer encoder and the Transformer decoder to build a cross-Transformer-based data association network, including, S331. Construct a row-column crossover Transformer network, which is applied to the three channels of the input matrix; each channel is processed separately by a row Transformer encoder and a column Transformer encoder, which treat the rows and columns of the input as feature vectors respectively; the output of the encoder is input to the row and column Transformer decoders in a crossover mode; Introducing the element-by-element method, Among them, h e is the output of the mixing matrix, Reshape(·) means converting the dimension of the matrix to a specified number, [·] means the concatenation operation, W and b are trainable weights and biases, and h1 and h2 are decoders; S332, construct a 3-channel cross Transformer network, the output of the element-by-element mixing processing module consisting of 3-channel data is guided to three Transformer encoders to obtain the encoding features; then, the output of each encoder corresponding to each channel is simultaneously fed to the first MHA module of the decoder of other channels; each output is simultaneously used as the key and value input of the second MHA module in the corresponding decoder of the channel; then the output matrix is ​​obtained by element-by-element mixing, Among them, O1, O2 and O3 are the outputs of the decoders corresponding to the three channels; S400, each element of the matrix obtained from the cross-Transformer-based data association network in step S300 represents the association probability of a given input data pair, and the Kuhn-Munkres algorithm is combined to find the best allocation, and then the network output is converted into an association matrix.

2. According to claim 1, a data association method based on a cross Transformer network architecture is characterized in that: In step S100, it includes: set up and D1 and D2 represent two sets of position data respectively; Define m×n matrices of relative distances in the x, y, and z directions, namely, relative distance feature matrices, respectively, to characterize the association relationship characteristics between data set D1 and data set D2; Among them, for the x-axis direction, set up The relative distance feature matrix associated with the x direction is as follows, in, and are the x-direction data from dataset D1 and dataset D2 respectively; bias is a small amount used to avoid data errors caused by division by zero; x mid is the geometric median of the absolute values ​​of all relative distances between two data sets, For the relative distance feature matrices in the y-axis direction and the z-axis direction, they are expressed as and 3. According to claim 2, a data association method based on a cross Transformer network architecture is characterized in that: In step S200, it includes: Let N pad is a fixed value, and N pad Less than or equal to m and n; let and are integers for the rows and columns of the scaled padding matrix, is the round-up operation; make is the fill matrix associated with x after partitioning, Division as follows, in, and are the number of partitions for rows and columns respectively; each partition submatrix is an N pad ×N pad matrix; Using the partitioning defined in Equation (4), the scaled padding matrix is ​​reshaped into a batch matrix as follows, The padding matrices associated with directions y and z are defined in the same way as for direction x; subsequently, the augmented matrix containing the relative distance features is used as the network input.

4. According to claim 3, a data association method based on a cross Transformer network architecture is characterized in that: In step S300, it includes: S310, build the Transformer encoder, including three integrated modules, namely multi-head attention, feedforward network and normalization network; The self-attention neural network uses Query to retrieve the attention values ​​corresponding to Keys and assign them to the corresponding Values; then redistributes Value according to the attention weights; let Attention represent the attention processing, in, and Represent Query and Key matrices respectively; represents the y matrix value, d is the dimension of Q, K and V in the self-attention neural network, Q, K and V represent the same feature encoding; Using the matrix W Q , W K and W V Combined with a multi-head attention network, different types of relationships can be extracted from the data. MultiHead(Q,K,V)=Concat(head1,…,head n )W O (7) in, is each attention network head, and W O It is a trainable parameter, which is equivalent to the Q, K and V tensors having the same shape; S320. Build a Transformer decoder, including two multi-head attention networks, set a mask layer in the first multi-head attention network, and use Q, K, and V in the second multi-head attention network. The decoder input is used as the query feature and the encoder input is used as the key and value; the resulting weighted input encoding is combined with the decoder input and passed through normalization and a feed-forward neural network to produce the decoder output.

5. The data association method based on the cross Transformer network architecture according to claim 4 is characterized in that: In step S300, it also includes using binary cross entropy loss as a cost to calculate the relative association probability of the network output, in, and The incidence matrix O pr and the true incidence matrix O gt Predict the element at row i and column j.

6. The data association method based on the cross Transformer network architecture according to claim 5, characterized in that: It can be used for multi-target tracking or multi-target tracking of monitoring, autonomous driving and robots to associate measurement values ​​with multiple targets, or to solve unknown correspondences between targets and measurement values.

7. A data association system based on a cross Transformer network architecture, characterized in that: The system has a program module corresponding to the steps of any one of claims 1 to 6 above, and executes the steps in the above-mentioned data association method based on the cross Transformer network architecture when running.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the data association method based on the cross Transformer network architecture described in any one of claims 1 to 6 when called by a processor.

Citation Information

Patent Citations

  • Image registration method based on cross neighborhood attention enhancement mechanism

    CN118447061A

  • Multi-modal feature fusion image classification method and application in humanoid robot

    CN118628802A