Dynamic Target Recognition Method Based on Multimodal Information Fusion

By constructing a dual-space representation model DSST based on Swin Transformer and introducing a joint loss function, the problems of poor generalization ability of traditional target recognition methods and independent modal signal processing are solved, and more efficient and accurate automatic driving target recognition is achieved.

CN116310676BActive Publication Date: 2025-07-25TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310165738.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-23
Publication Date
2025-07-25
Estimated Expiration
2043-02-23

AI Technical Summary

Technical Problem

The existing target recognition methods rely on manual design features, have poor generalization capabilities, and different modal signal processing methods independently lead to distortion of the recognition results, and have poor accuracy.

Method used

A dual-space representation model DSST based on Swin Transformer is constructed, including a modal invariant first subspace and a modal specific second subspace. Through modal representation learning and fusion, a joint loss function is introduced to construct an autonomous driving target recognition model RD-DSST.

Benefits of technology

Effectively reduce modal gaps, enhance modal interaction, improve identification accuracy and efficiency, and improve the operability of autonomous driving target recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116310676B_ABST
    Figure CN116310676B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of driverless technology, and specifically relates to a dynamic target recognition method based on multi-modal information fusion, including: S1, obtaining multi-modal information data and dividing it into a training set and a test set; S2, constructing a dual-space representation model DSST based on Swin Transformer for the multi-modal information data, and the dual-space representation model DSST performs modal representation learning and modal fusion in sequence; S3, introducing a joint loss to construct an autonomous driving target recognition model RD-DSST; S4, training the autonomous driving target recognition model RD-DSST on the training set and saving the converged model; S5, calling the converged model in step S4 to perform target recognition on the test set and automatically generating recognition results. The dynamic target recognition method provided by the present invention can not only reduce the distributed modal gap caused by the heterogeneity of different signals, but also fully consider modal differences and utilize modal correlations, providing a reliable guarantee for the decision-making planning and control of driverless driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of unmanned driving technology, and in particular to a dynamic target recognition method based on multi-modal information fusion. Background Art

[0002] In recent years, as Internet companies, new car manufacturers and traditional car companies have invested in the autonomous driving market, the field of autonomous driving has become hot. As the most important part of the perception system of autonomous driving vehicles, object recognition directly determines whether the autonomous driving vehicle can operate safely as expected, which makes object recognition one of the most active research areas in computer vision.

[0003] Traditional object recognition methods mainly consist of the following steps: first, a candidate box is generated on the image using a sliding window and the features corresponding to the image in the candidate box are extracted; then the candidate box is classified using a classifier such as a support vector machine; and finally, non-maximum suppression is performed to output the result. This method has poor generalization ability due to its reliance on manually designed features, and the sliding window global search also leads to high time complexity.

[0004] In order to overcome the above defects of traditional target recognition methods, recognition methods based on deep learning have emerged. However, no matter which recognition method is used, it uses a completely independent method to process individual modal signals or treats all modal signals equally in the same way, which leads to distortion of target recognition results and poor accuracy. Summary of the invention

[0005] In order to overcome the technical defect of poor accuracy of existing target recognition, the present invention provides a dynamic target recognition method based on multimodal information fusion.

[0006] The present invention provides a dynamic target recognition method based on multimodal information fusion, comprising the following steps:

[0007] S1. Obtain multimodal information data and divide it into training set and test set;

[0008] S2. constructing a dual-space representation model DSST based on Swin Transformer for multimodal information data, wherein the dual-space representation model DSST includes a first subspace and a second subspace, wherein the first subspace modality is unchanged and the second subspace modality is specific, and the dual-space representation model DSST sequentially performs modality representation learning and modality fusion;

[0009] S3, introduce joint loss and build the autonomous driving target recognition model RD-DSST;

[0010] S4. Train the autonomous driving target recognition model RD-DSST on the training set and save the converged model.

[0011] S5. Call the model converged in step S4 to perform target recognition on the test set and automatically generate recognition results.

[0012] Optionally, step S2 is divided into the following sub-steps:

[0013] S21. Use the video data from the lidar sensor and the visible light camera as the corpus, given the sequence U m ∈{R, v}, and each corpus is respectively represented as and Each corpus sequence is mapped to a vector of a fixed size The dual-space representation model DSST adopts a hierarchical design, gradually reducing the resolution of the input feature map to expand the receptive field. Finally, the hidden layer state representation is added with a fully connected dense layer to obtain The hidden modal invariance vector representation of the first subspace is The modal-specific vector representation of the second subspace is and Implemented through the encoding function:

[0014]

[0015]

[0016] Then the hidden modal invariance vector and the modal-specific vector of the radar sensor are as follows:

[0017]

[0018]

[0019] The hidden modal invariance vector and the modal-specific vector of the video data are as follows:

[0020]

[0021]

[0022] S22. After obtaining the above four vectors and project these modalities into their respective representations and fuse them into a joint vector.

[0023] Optionally, step S22 includes the following sub-steps:

[0024] S221. Based on Transformer, using the attention module, for the above four vectors and perform concatenation, and the scaling function defined by it is as follows:

[0025]

[0026] Among them, Q, K, and V are the query matrix, key matrix, and value matrix. Transformer calculates multiple such parallel attentions, and the output of each attention is called a head. The calculation method of the i-th head is as follows:

[0027] head i = Attention(QW i q , KW i k , VW i v )

[0028] Among them, W i q , W i k and W i v are the parameters for the head, used to linearly project the matrix into the local space;

[0029] S222. Represent the modality as a matrix stacked. For the attention mechanism, set Transformer generates a new matrix Finally, using the output of Transformer, construct a joint vector by the concatenation method Task prediction is through the generation function

[0030] Optionally, step S3 includes the following sub-steps:

[0031] S31. Based on the RGB and depth modalities in the video data of the visible light camera, construct a loss function:

[0032]

[0033] Among them, W j represents the matrix weight for single-modal learning, W J represents the weight matrix for joint learning, y i represents the sample;

[0034] S32. Calculate the matching score values of RGB-RGB and Depth-Depth using the FC1024 features of RGB and the FC1024 features of Depth respectively, and then obtain the final score through weighted fusion:

[0035]

[0036] where p1 and p2 are the matching precisions of the video data observed separately in each mode.

[0037] Optionally, in step S5, based on the automatic recognition results of accuracy and recall rate, the accuracy rate refers to the proportion of the correct number in a certain type of detection results to the total number of all detection results, and the recall rate refers to the proportion of the number of correct positive samples in a certain type of detection results to the total number of all positive samples in the test set. The accuracy rate The recall rate where T and F represent correct and incorrect detection results respectively, P and N represent true values as positive samples and negative samples respectively, TP represents that the detection result is a positive sample and the true value is also a positive sample; FP represents that the detection result is a positive sample and the true value is a negative sample; TN represents that the detection result is a negative sample and the true value is also a negative sample; FN represents that the detection result is a negative sample and the true value is a positive sample.

[0038] Optionally, the accuracy rate and the recall rate are combined through the average precision. The average precision represents the average value of the accuracy rates at different recall rates for a certain category. The average precision where the recall rate is equally spaced within the range of [0-1], r i represents the i-th recall rate, N represents the total number of recall rate values, and p(r i ) represents the accuracy rate at the i-th recall rate.

[0039] Optionally, for targets of different categories, the mean average precision needs to be calculated. The mean average precision is the average value of the average precisions of all categories. The mean average precision where Ap(j) represents the average precision of the j-th category target, and M represents the total number of categories.

[0040] Optionally, F1-Measure is adopted as one of the evaluation indicators.

[0041] The technical solution provided by the present invention has the following advantages compared with the prior art:

[0042] The dynamic target recognition method based on multi-modal information fusion provided by the present invention constructs a dual-space representation model DSST for multi-modal information data. The first subspace is modality-invariant and cross-modal, learning their commonalities through representation learning to narrow the modality gap. The second subspace is modality-specific and private to each modality, ensuring their features. In this way, the distributed modality gap caused by the heterogeneity of different signals is effectively solved. In addition, this method introduces a joint loss to enhance the interaction between modalities, fully considering modality differences and utilizing modality correlations, and then proposes the RD-DSST model, which can learn the common features among multiple different modalities. Based on the above advantages, this method can intuitively display the recognition results, ensure the accuracy of the recognition results, make the recognition faster and more efficient, greatly improve the operability of target recognition in the field of autonomous driving, and improve work efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present invention and, together with the specification, are used to explain the principles of the present invention.

[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0045] Figure 1 It is the overall flowchart of the dynamic target recognition method based on multi-modal information fusion according to the embodiment of the present invention;

[0046] Figure 2 It is the structural diagram of the dual-space representation model DSST according to the embodiment of the present invention;

[0047] Figure 3 It is the structural diagram of the RGB and Depth information fusion in step S22 according to the embodiment of the present invention;

[0048] Figure 4 It is the structural diagram of the autonomous driving target recognition model RD-DSST according to the embodiment of the present invention;

[0049] Figure 5 It is the network structural diagram of the Swin Transformer in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] In order to be able to more clearly understand the above objects, features and advantages of the present invention, the following will further describe the solution of the present invention. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.

[0051] Numerous specific details are set forth in the following description to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present invention, rather than all the embodiments.

[0052] The following specifically describes the embodiments of the present invention with reference to the accompanying drawings.

[0053] In one embodiment, as Figure 1 shown, the dynamic target recognition method based on multimodal information fusion includes the following steps:

[0054] S1. Obtain multimodal information data and divide it into a training set and a test set;

[0055] S2. Construct a dual-space representation model DSST based on Swin Transformer for the multimodal information data. The dual-space representation model DSST includes a first subspace and a second subspace. The first subspace is modality-invariant, and the second subspace is modality-specific. The dual-space representation model DSST performs modality representation learning and modality fusion in sequence;

[0056] S3. Introduce a joint loss and construct an autonomous driving target recognition model RD-DSST;

[0057] S4. Train the autonomous driving target recognition model RD-DSST on the training set and save the converged model;

[0058] S5. Call the converged model in step S4 and perform target recognition on the test set to automatically generate recognition results.

[0059] It should be noted that the first subspace is modality-invariant, where cross-modal representation learning captures their commonalities and reduces the modality gap. The second subspace is modality-specific, private to each modality, and preserves their features. Then modality fusion is performed for downstream prediction.

[0060] Specifically, the autonomous driving target recognition model RD-DSST is for RGB and Depth modality fusion based on DSST, and by introducing a joint loss, the interaction between modalities is enhanced.

[0061] Specifically, in step S4, the Adam optimizer is used in the training process, and the loss function is the joint loss function. In actual training, the momentum is set to 0.843, the learning rate is set to 0.0032, and the model is saved after convergence.

[0062] The dynamic target recognition method based on multimodal information fusion in this embodiment constructs a dual-space representation model DSST for multimodal information data. The first subspace is modality-invariant and cross-modal, learning their commonalities through representation learning and narrowing the modality gap. The second subspace is modality-specific and private to each modality, ensuring their features. In this way, the distributed modality gap caused by the heterogeneity of different signals is effectively solved. In addition, this method introduces a joint loss to enhance the interaction between modalities, fully considering modality differences and leveraging modality correlations, and then proposes the RD-DSST model, which can learn the common features between multiple different modalities. Based on the above advantages, this method can intuitively display the recognition results and accuracy, making the recognition faster and more efficient, greatly improving the operability of target recognition in the field of autonomous driving and enhancing the work efficiency.

[0063] In a preferred embodiment, step S2 is divided into the following sub-steps:

[0064] S21. Use the video data of the lidar sensor and the visible light camera as the corpus, and given the sequence U m ∈ {R, v}, each corpus is respectively represented as and Each corpus sequence is mapped to a vector of a fixed size The dual-space representation model DSST adopts a hierarchical design, gradually reducing the resolution of the input feature map to expand the receptive field. Finally, a fully connected dense layer is added to the hidden layer state representation to obtain The hidden modality-invariant vector representation of the first subspace is denoted as The modality-specific vector representation of the second subspace is denoted as and Implemented through the encoding function:

[0065]

[0066]

[0067] Then the hidden modality-invariant vector and the modality-specific vector of the radar sensor are as follows:

[0068]

[0069]

[0070] The hidden modality-invariant vector and the modality-specific vector of the video data are as follows:

[0071]

[0072] Specifically, an HDL-64E lidar sensor is selected. This sensor has 64 channels, a measurement range of up to 120 m, a 360° horizontal field of view and a 26.9° vertical field of view, an azimuth resolution of 0.08°, and a vertical resolution of 0.04°. The laser scanner rotates at a speed of 10 frames per second and captures approximately 100k points per cycle.

[0073] Specifically, the camera of the visible light camera is installed at a position approximately horizontal to the ground. The camera image is cropped to a size of 1382 * 512 pixels using the format mode of libdc, and the corrected image will become smaller. The laser scanner (facing forward) triggers the camera at a speed of 10 frames per second and dynamically adjusts the shutter time (the maximum shutter time is 2 ms).

[0074] Specifically, as Figure 5 shown, the dual-space representation model DSST adopts a hierarchical design and consists of a total of 4 Stages. Each stage reduces the resolution of the input feature map and expands the receptive field layer by layer like a CNN.

[0075] It is easy to understand that the first subspace is modality-invariant, and a shared representation is learned in a common subspace with distribution similarity. This constraint helps to minimize the heterogeneity gap, which is an ideal characteristic of multimodal fusion. The second subspace is modality-specific and captures the unique features of that modality.

[0076] S22, obtaining the above four vectors and After that, project these modalities into their respective representations and fuse them into a joint vector.

[0077] This embodiment provides a preferred operation method for step S2. In other embodiments, other common method steps can also be used to complete the establishment, representation, and fusion of the dual-space representation model DSST.

[0078] More specifically, step S22 includes the following sub-steps:

[0079] S221. Based on Transformer, using the attention module, concatenate the above four vectors and The scaling function defined by it is as follows:

[0080]

[0081] Among them, Q, K, and V are the query matrix, key matrix, and value matrix respectively. The transformer calculates multiple such parallel attentions, and the output of each attention is called a head. The calculation method of the i-th head is as follows:

[0082] head i =Attention(QW i q ,KW i k ,VW i v )

[0083] Among them, W i q , W i k and W i v are the parameters for the head, used to linearly project the matrix into the local space;

[0084] S222. Represent the modality as a matrix stacked. For the attention mechanism, set The transformer generates a new matrix Finally, using the output of the transformer, a joint vector is constructed by concatenation Task prediction is through the generation function The model structure is as Figure 2 shown.

[0085] In a preferred embodiment, step S3 includes the following sub-steps:

[0086] S31. Based on the RGB and depth modalities in the video data of the visible light camera, construct a loss function:

[0087]

[0088] Among them, W j represents the matrix weight for single-modal learning, W J represents the weight matrix for joint learning, and y i represents the sample;

[0089] S32. Use the FC1024 features of RGB and the FC1024 features of Depth to calculate the RGB-RGB matching score value and the Depth-Depth matching score value respectively, and then obtain the final score through weighted fusion:

[0090]

[0091] where p1 and p2 are the matching accuracies of the video data observed separately for each mode.

[0092] The structural diagram of RGB and Depth information fusion is as Figure 3 shown, and the implementation details of the entire network are as Figure 4 shown.

[0093] It is easy to understand that in the data with the video data of the visible light camera for autonomous driving as the corpus, since RGB and depth respectively describe the texture and shape information of vehicles and pedestrians in the street scene, these two modalities should be relevant and complementary. By introducing a joint loss to enhance the interaction between modalities, a new autonomous driving target recognition model RD-DSSTRD-DSST is constructed.

[0094] This embodiment provides a preferred operation method for step S3. In other embodiments, other common method steps can also be used to introduce the joint loss.

[0095] In some embodiments, in step S5, based on the automatic recognition results of accuracy and recall rate, the accuracy rate refers to the proportion of the number of correct detections in a certain type of detection result to the total number of all detection results, and the recall rate refers to the proportion of the number of correct positive samples in a certain type of detection result to the total number of all positive samples in the test set. The accuracy recall rate where T and F respectively represent correct and incorrect detection results, P and N respectively represent true values as positive samples and negative samples, TP represents the detection result as a positive sample and the true value is also a positive sample; FP represents the detection result as a positive sample and the true value is a negative sample; TN represents the detection result as a negative sample and the true value is also a negative sample; FN represents the detection result as a negative sample and the true value is a positive sample.

[0096] Furthermore, the accuracy and recall rate are combined through the average precision. The average precision represents the average value of the accuracy rates at different recall rates for a certain category. The average precision where the recall rate is evenly sampled in the range of [0 - 1], r i represents the i-th recall rate, N represents the total number of recall rate samples, and p(ri) represents the accuracy rate at the i-th recall rate.

[0097] Furthermore, for targets of different categories, the mean average precision needs to be calculated. The mean average precision is the average value of the average precisions of all categories. The mean average precision where Ap(j) represents the average precision of the j-th category target, and M represents the total number of categories.

[0098] Furthermore, F1-Measure is adopted as one of the evaluation metrics.

[0099] It should be noted that F1-Measure is an evaluation metric that is often used in information retrieval and natural language processing. F1-Measure is a comprehensive evaluation metric given based on both Precision and Recall.

[0100] As described above, a dynamic target recognition method based on multi-modal information fusion in this application inputs multi-modal information data into the RD-DSST model for target recognition, and then automatically obtains the recognition result. It can not only reduce the distributed modal gap caused by the heterogeneity of different signals, but also fully consider modal differences and utilize modal correlations, providing a reliable guarantee for the decision-making planning and control of unmanned driving.

[0101] The above are only specific implementation manners of the present invention, enabling those skilled in the art to understand or implement the present invention. Although the foregoing embodiments have been described in detail, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments, and they should all be covered by the protection scope of the claims.

Claims

1. A dynamic target recognition method based on multi-modal information fusion, characterized in that, It includes the following steps: S1. Obtain multi-modal information data and divide it into a training set and a test set; S2. Construct a dual-space representation model DSST based on Swin Transformer for the multi-modal information data. The dual-space representation model DSST includes a first subspace and a second subspace. The first subspace is modality-invariant, and the second subspace is modality-specific. The dual-space representation model DSST performs modality representation learning and modality fusion in sequence, which is divided into the following sub-steps: S21. Use the video data of lidar sensors and visible light cameras as the corpus, and given a sequence , each corpus is respectively represented as and . Each corpus sequence is mapped to a vector of a fixed size . The dual-space representation model DSST adopts a hierarchical design, gradually reducing the resolution of the input feature map layer by layer to expand the receptive field. Finally, a fully connected dense layer is added to the hidden layer state representation to obtain . The hidden mode invariance vector of the first subspace is represented as . The mode-specific vector of the second subspace is represented as . and are implemented through an encoding function: ; Then the hidden mode invariance vector of the radar sensor and the mode-specific vector are as follows: ; Video data hiding modal invariance vector and modal-specific vectors are as follows: ; S22. Obtain the above four vectors , , and . After that, first concatenate the four vectors through the Transformer multi-head attention mechanism for feature interaction, then stack the modal features into a matrix and generate a new representation through self-attention, and finally splice and construct a joint vector containing multi-modal information; S3. Introduce a joint loss and construct an autonomous driving target recognition model RD-DSST; S4. Train the autonomous driving target recognition model RD-DSST on the training set and save the converged model; S5. Call the model converged in step S4 and perform target recognition on the test set to automatically generate recognition results.

2. The dynamic target recognition method based on multimodal information fusion according to claim 1, wherein Step S22 includes the following sub-steps: S221. Based on Transformer, using the attention module, concatenate the above four vectors , , and . The scaling function defined by it is as follows: ; Among them, Q, K, and V are the query matrix, key matrix, and value matrix. The transformer calculates multiple such parallel attentions, and the output of each attention is called a head. The calculation method of the i-th head is as follows: ; Among them, , and are parameters for the head, used to linearly project the matrix into the local space; S222. Represent the modality as a matrix Stacked. For the attention mechanism, set , the transformer generates a new matrix . Finally, using the output of the transformer, a joint vector is constructed by the concatenation method . The task prediction is through the generation function .

3. The dynamic target recognition method based on multi-modal information fusion according to claim 2, characterized in that Step S3 includes the following sub-steps: S31. Based on the RGB and depth modalities in the video data of the visible light camera, construct a loss function: ; Among them, represents the matrix weight for single-modal learning, represents the weight matrix for joint learning, represents the sample; S32. Use the FC1024 features of RGB and the FC1024 features of Depth to calculate the matching score values of RGB-RGB and Depth-Depth respectively, and then obtain the final score through weighted fusion: ; Among them, p1 and p2 are the matching accuracies of the video data observed separately for each mode.

4. The dynamic target recognition method based on multi-modal information fusion according to claim 3, characterized in that In step S5, based on the automatic recognition results of precision and recall rate, the precision rate refers to the proportion of the number of correct detections in a certain type of detection results to the total number of all detection results, and the recall rate refers to the proportion of the number of correctly detected positive samples in a certain type of detection results to the total number of all positive samples in the test set. The precision rate , and the recall rate , where T and F represent correct and incorrect detection results respectively, P and N represent true positive and true negative samples respectively, TP represents that the detection result is a positive sample and the true value is also a positive sample; FP represents that the detection result is a positive sample and the true value is a negative sample; TN represents that the detection result is a negative sample and the true value is also a negative sample; FN represents that the detection result is a negative sample and the true value is a positive sample.

5. The dynamic target recognition method based on multi-modal information fusion according to claim 4, characterized in that The accuracy rate and the recall rate are comprehensively evaluated by the average precision, where the average precision represents the average value of the accuracy rates of a certain category at different recall rates. The average precision , where the recall rate is evenly sampled in the range of [0 - 1], represents the i-th recall rate, N represents the total number of samples of the recall rate, represents the accuracy rate at the i-th recall rate.

6. The dynamic target recognition method based on multi-modal information fusion according to claim 5, characterized in that For different categories of targets, the mean average precision needs to be calculated. The mean average precision is the average of the average precisions of all categories. The mean average precision , where represents the average precision of the j-th category of targets, and M represents the total number of categories.

7. The dynamic target recognition method based on multi-modal information fusion according to claim 5, wherein Adopt F1-Measure as one of the evaluation metrics, .

Citation Information

Patent Citations

  • RGBT target tracking method based on cross-modal attention mechanism and twin structure

    CN113628249A

  • Space-time self-learning target tracking method based on anti-occlusion mechanism

    CN115601568A