Method and system for sensing scene aircraft by guiding vehicle based on thunder-vision fusion

By adopting the Razor-vision fusion technology in airport guide vehicles, combining radar and image data, and using SMCA model and uncertainty-optimized loss function, the shortcomings of pure radar perception in detecting small-scale or long-range aircraft are solved, and a higher precision aircraft perception is achieved.

CN120070577APending Publication Date: 2025-05-30ANHUI KELI INFORMATION IND +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510212360.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When using pure radar sensing in airport guide vehicles, it is difficult to effectively detect aircraft of small scale or long distances, and the depth estimation error of the camera is large and the light sensitivity is high, which cannot meet the requirements of precise monitoring of aircraft distance and attitude.

Method used

The perception method based on lightning vision fusion is adopted, and the fusion of radar point cloud data and image data is used to fusion across modal features using the SMCA model. Through the uncertainty-optimized loss function training, the contribution weight of the loss function is dynamically adjusted to output the real-time distance and attitude of the aircraft relative to the guide vehicle.

Benefits of technology

The airport guided vehicle's perception accuracy of scene aircraft is improved, combined with the radar's ranging accuracy and high image resolution, overcome the shortcomings when using radar or camera alone, and achieve a more accurate and stable perception effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070577A_ABST
    Figure CN120070577A_ABST
Patent Text Reader

Abstract

The invention discloses a method and system for sensing a scene aircraft by a guide vehicle based on thunder-vision fusion, and the method comprises the steps: S1, carrying out the initialization query based on a central thermodynamic diagram and a classification alertness mechanism: predicting the thermodynamic diagram based on an aerial view generated based on radar point cloud data, selecting a candidate object from the thermodynamic diagram, and carrying out the recognition of the candidate object; fusing the category information of the candidate objects into the features of the query object through a category embedding method; s2, Leiye fusion based on an SMCA model: performing cross attention calculation on the initialized features of the query object and the features of the image data to complete cross-modal feature fusion; s3, loss function training based on uncertainty optimization: dynamically adjusting the contribution weight of a loss function according to a confidence coefficient parameter; and S4, outputting a sensing result of the guide vehicle to the scene aircraft. The advantages of radar sensing and high-resolution images are combined, the defects existing when the radar sensing and the high-resolution images are independently used are overcome, and the sensing precision of the airport guide vehicle to the scene aircraft is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving perception, and in particular, to a method and system for a ground guidance vehicle to perceive surface aircraft based on radar-vision fusion. Background Art

[0002] In an autonomous driving perception system, 3D object detection technology aims to provide the spatial coordinates of an object in 3D space and determine its category.

[0003] When a ground guidance vehicle is operating, it is necessary to ensure a relatively safe distance between the aircraft and the guidance vehicle and monitor the driving attitude of the aircraft to ensure that the aircraft safely enters the corresponding parking position.

[0004] It can be seen that the real-time monitoring of the distance between the aircraft and the guidance vehicle and the monitoring of the position and attitude of the aircraft are the key points in the perception of the ground guidance vehicle and surface aircraft.

[0005] Currently, although the method of only using radar perception has the advantages of high ranging accuracy and strong anti-interference ability, it is limited by the sparsity of point clouds and is prone to missed detection when detecting small-scale targets or long-distance targets.

[0006] However, for the high-resolution images of cameras, the deficiencies of the above pure radar perception scheme do not exist, but it has inherent defects such as large depth estimation errors and high light sensitivity.

[0007] Therefore, how to simultaneously utilize the fusion of visual information of radar and cameras to perceive the real-time distance between the aircraft and the guidance vehicle and the position and attitude of the aircraft is a research topic worthy of study. Summary of the Invention

[0008] In the first aspect of the present invention, in order to solve the above technical problems, the present invention provides a method for a ground guidance vehicle to perceive surface aircraft based on radar-vision fusion, and the method includes the following steps: S1. Initialization query based on the center heat map and classification alert mechanism: Predict a heat map based on the bird's-eye view generated from radar point cloud data, select candidate objects from the heat map, and incorporate the category information of the candidate objects into the features of the query object through the method of category embedding; S2. Radar-vision fusion based on the SMCA model: Perform cross-attention calculation according to the features of the query object initialized in S1 and the features of the image data to complete cross-modal feature fusion; S3. Training of the loss function based on uncertainty optimization: Construct a loss function and dynamically adjust the contribution weight of the loss function according to the confidence parameter; S4. Output the perception result of the ground guidance vehicle for the surface aircraft: Analyze the bird's-eye view after radar-vision fusion and output the real-time distance between the aircraft and the guidance vehicle and the attitude of the aircraft.

[0009] Further, predicting a heat map from the bird's-eye view generated based on the radar point cloud data, and selecting candidate objects from the heat map includes: Bird's-eye view , where represents the size of the bird's-eye view, represents the dimension of the bird's-eye view, and the heat map , where represents the number of classification categories; Select the top N candidate objects with the largest heat values from the heat map as the initialization objects for query, where the initialization objects include object location information and instance encoding information.

[0010] Further, integrating the category information of the candidate objects into the features of the query object by the method of category embedding includes: Denote the feature of the query object as , where represents the location where the query object is located, represents the category of the query object; Project the feature of the query object and the one-hot category vector after category encoding onto the vector in the space and add them to complete category embedding.

[0011] Further, the heat map is generated by a Gaussian kernel function, and the peak of the heat value corresponds to the center coordinates of the aircraft.

[0012] Further, the S2 further includes: Extract geometric features from the radar point cloud data through a first decoding layer, and extract semantic features from the image data acquired by the camera through a second decoding layer; Define the first input sequence as the Query sequence (Q), the second input sequence as the Key sequence (K), and the Value sequence (V); Then the attention weights of the comprehensive Query sequence (Q), Key sequence (K), and Value sequence (V) satisfy the expression:

[0013] where the softmax function normalizes the attention scores to obtain the normalized attention weights, represents the dimension of the Key sequence, represents the transpose, represents the transpose of K; Based on the SMCA model, the mask matrix Multiply with the attention weight points to generate an aerial view after radar-vision fusion, where the mask matrix satisfies the expression:

[0014] where represents the element position number of M, and represent the two-dimensional center point calculated by projecting the query prediction onto the image plane, represents the minimum circumferential radius of the projection angle of the three-dimensional bounding box; represents the hyperparameter for adjusting the Gaussian distribution.

[0015] Furthermore, the loss function satisfies the expression:

[0016] where is the binary cross-entropy loss, is the L1 norm difference between the center of the predicted aerial view and the actual center of the aerial view, are the cross-entropy loss coefficient, the norm difference coefficient, and the intersection over union loss coefficient calculated based on uncertainty respectively, is the confidence parameter, is the predicted three-dimensional bounding box, is the true three-dimensional bounding box.

[0017] In the second aspect of the present invention, a system for implementing the method for a guided vehicle to perceive surface aircraft based on radar-vision fusion includes: A radar module for collecting three-dimensional radar point cloud data of surface aircraft; A vision module including a camera for collecting image data of surface aircraft; A processing unit integrated with an SMCA model and a feedforward neural network for performing effective fusion of the radar point cloud data and the image data and outputting the real-time distance of the surface aircraft relative to the airport guided vehicle and the attitude information of the aircraft; A control module for receiving the output parameters of the processing unit and controlling the driving strategy of the guided vehicle.

[0018] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: By effectively fusing the radar point cloud data and the image data, the present invention combines the advantages of radar perception and high-resolution images and offsets the deficiencies existing when they are used alone, greatly improving the perception accuracy of the airport guided vehicle for surface aircraft. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required in the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0020] Figure 1 It is a flowchart disclosed in the embodiments of the present invention. Specific implementation manners

[0021] To enable those skilled in the art of this technology to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0022] "Radar-vision fusion" refers to the data fusion of radar (radar point cloud features) and vision (image features).

[0023] The present invention aims to provide a method for an airport guidance vehicle to perceive surface aircraft based on radar-vision fusion, improve the alignment accuracy of radar point cloud data features and image features based on the SMCA model, and optimize the training stability through an uncertainty loss function.

[0024] Please refer to Figure 1 , the method for an airport guidance vehicle to perceive surface aircraft based on radar-vision fusion mainly includes the following steps: Step 1: Initial query based on the center heat map and classification alert mechanism Predict the heat map based on the bird's eye view generated from the radar point cloud data, select candidate objects from the heat map, and integrate the category information of the candidate objects into the features of the query object through the method of category embedding.

[0025] After obtaining the radar point cloud data, in order to identify the detection target and extract information such as its distance and pose, it is necessary to query and classify the target objects in the radar point cloud data.

[0026] Preferably, this embodiment uses an initialization method based on the center heat map to improve the classification performance of the model.

[0027] Specifically, a d-dimensional radar bird's eye view (BEV) of a surface aircraft is given, where the bird's eye view , where represents the size of the bird's eye view, Indicates the dimension of the bird's-eye view.

[0028] Based on this bird's-eye view Predict a special heat map , where Indicates the number of classification categories.

[0029] Next, select the top N candidate objects with the largest heat values from the heat map as the initialization objects for the query. Among them, the initialization objects include object location information and instance encoding information.

[0030] To further explain, the object location information is used to describe the location of the object; the instance encoding information is used to describe the size and direction of the query box.

[0031] To further explain, the heat map is generated by the Gaussian kernel function, and the peak of the heat value corresponds to the center coordinates of the aircraft.

[0032] On this basis, use the classification alert mechanism to implement category embedding for each query object.

[0033] Denote the feature of the query object as , where represents the location of the query object, represents the category of the query object.

[0034] Project the feature of the query object and the one-hot category vector after category encoding onto the vector in the space and add them together to complete the category embedding. The purpose of this step is to integrate the category information into the feature of the query object, so that the model can better perform classification and help the model more accurately identify and classify different objects.

[0035] Step 2, Radar-vision fusion based on the SMCA model Based on the multi-head attention mechanism, fuse the radar point cloud data and the image data. That is, perform cross-attention calculation according to the features of the query object initialized in Step 1 and the features of the image data to complete cross-modal feature fusion.

[0036] First, extract geometric features from the radar point cloud data through the first decoding layer, and extract semantic features from the image data obtained by the camera through the second decoding layer.

[0037] Then, define the first input sequence as the Query sequence (Q), and the second input sequence as the Key sequence (K) and the Value sequence (V).

[0038] Calculate the attention score A, which satisfies the expression:

[0039] Among them, Q is the matrix of the Query sequence, and K is the matrix of the Key sequence. represents the transpose. represents the transpose of K.

[0040] Scale the attention scores A and use the softmax function to obtain the normalized attention weights. Among them, the attention weights satisfy the expression:

[0041] Among them, is the dimension of the Key sequence. The softmax function normalizes the attention scores to obtain the normalized attention weights. Specifically, the softmax function converts each attention score into a probability distribution, ensuring that their sum is equal to 1.

[0042] Finally, use the normalized attention weights to perform a weighted sum on the Value sequence. Then, the attention weights that integrate the Query sequence (Q), Key sequence (K), and Value sequence (V) satisfy the expression:

[0043] Among them, V is the matrix of the Value sequence.

[0044] However, due to the different sources of radar point cloud data and image data, it is difficult for the algorithm to perform rapid and accurate matching during radar-vision fusion. Based on this, the SMCA (spatially modulated cross attention) model is introduced. The SMCA model weighs the cross-attention degree through a two-dimensional circular Gaussian mask centered on the projected two-dimensional center of each query object.

[0045] Further explained, based on the SMCA model, multiply the mask matrix by the attention weights to generate a bird's-eye view after radar-vision fusion. Among them, the mask matrix satisfies the expression:

[0046] Among them, represents the element position number of M. , represent the two-dimensional center point calculated by projecting the query prediction onto the image plane. represents the minimum circumferential radius of the projection angle of the three-dimensional bounding box. represents the hyperparameter that adjusts the Gaussian distribution.

[0047] Introducing a Gaussian mask to dynamically adjust the cross-modal attention weights enhances the spatial alignment ability of radar and visual features, improves the fusion effect in complex scenarios (such as metal reflection interference), and enables the model to better and faster learn to select the positions of image features based on the input lidar features.

[0048] Step 3: Training with a loss function based on uncertainty optimization Construct a loss function and dynamically adjust the contribution weights of the loss function according to the confidence parameter.

[0049] After performing the cross-attention mechanism, use the FFN (feed-forward neural network) to generate the final bounding box prediction results.

[0050] Among them, the loss function for evaluation consists of three parts:

[0051]

[0052]

[0053] Among them, is the binary cross-entropy loss, is the L1 norm difference between the predicted center of the bird's-eye view and the actual center of the bird's-eye view, is based on uncertainty optimization metric, are the cross-entropy loss coefficient, the norm difference coefficient, and the intersection over union loss coefficient calculated based on uncertainty respectively, is the confidence parameter, is the predicted 3D bounding box, is the true 3D bounding box.

[0054] When the model has a high confidence in the previous steps, , it has no impact on the loss function.

[0055] When the model has a low confidence in the previous steps, will reduce the impact on the loss function.

[0056] This design enables the model to jointly optimize uncertainty and 3D matching degree, and through dynamically balance the model confidence and matching accuracy, avoid the negative impact of low-confidence predictions on training, and improve the model robustness.

[0057] Step 4: Output the perception results of the airport ground guidance vehicle for the surface aircraft After successfully fusing the perception information from the radar and the camera, parsing the bird's-eye view after radar-vision fusion can output the real-time distance between the aircraft on the scene and the airport guide vehicle and the attitude of the aircraft.

[0058] It should be noted that the distance between the aircraft on the scene and the airport guide vehicle and the attitude of the aircraft are included in the radar-vision fusion and feature matching in Step 2.

[0059] The present invention also protects a perception system for an airport guide vehicle to an aircraft on the scene based on radar-vision fusion, which mainly includes a radar module, a vision module, a processing unit, and a control module, wherein: The radar module is used to collect three-dimensional radar point cloud data of the aircraft on the scene.

[0060] The vision module includes a camera and is used to collect image data of the aircraft on the scene.

[0061] The processing unit is integrated with an SMCA model and a feedforward neural network, and is used to perform effective fusion of the radar point cloud data and the image data, and output the real-time distance between the aircraft on the scene and the airport guide vehicle and the attitude information of the aircraft.

[0062] The control module receives the output parameters of the processing unit and controls the driving strategy of the guide vehicle.

[0063] By introducing the SMCA model, the present invention can effectively solve the problem that it is difficult for the radar-vision fusion algorithm to quickly and accurately match due to the different sources of radar point cloud features and image features, thereby improving the performance and accuracy of radar-vision fusion, and realizing high-precision and full-time perception of the airport guide vehicle to the aircraft on the scene.

[0064] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for sensing aircraft on the ground by a guide vehicle based on radar and vision fusion, characterized in that: The method comprises the following steps: S1. Initialization query based on central heat map and classification alert mechanism: predict the heat map based on the bird's-eye view generated by radar point cloud data, select candidate objects from the heat map, and integrate the category information of the candidate objects into the features of the query object through the category embedding method; S2, radar-visual fusion based on SMCA model: performing cross-attention calculation according to the features of the query object and the features of the image data initialized in S1, and completing cross-modal feature fusion; S3. Loss function training based on uncertainty optimization: construct a loss function and dynamically adjust the contribution weight of the loss function according to the confidence parameter; S4. Output the perception result of the guide vehicle on the scene aircraft: Analyze the bird's-eye view after radar and vision fusion, and output the real-time distance of the aircraft relative to the guide vehicle and the aircraft's attitude.

2. The method for sensing a surface aircraft by a guide vehicle based on radar and visual fusion according to claim 1 is characterized in that: The step of predicting a heat map based on a bird's-eye view generated from radar point cloud data and selecting a candidate object from the heat map comprises: Bird's eye view ,in, Indicates the size of the bird's-eye view. Dimensions representing a bird’s-eye view, heat map ,in, Indicates the number of classification categories; The first N candidate objects with the largest heat values ​​are selected from the heat map as initialization objects for query, wherein the initialization objects include object location information and instance encoding information.

3. The method for sensing a surface aircraft by a guide vehicle based on radar and visual fusion according to claim 2 is characterized in that: The method of integrating the category information of the candidate object into the features of the query object by category embedding includes: The characteristics of the query object are ,in, Indicates the location of the query object, Indicates the category of the query object; The characteristics of the query object Project the one-hot category vector after category encoding to The vectors in the space are added to complete the category embedding.

4. The method for sensing a surface aircraft by a guide vehicle based on radar and visual fusion according to claim 2 is characterized in that: The heat map Generated by a Gaussian kernel function, the peak value of the thermal value corresponds to the coordinates of the center of the aircraft.

5. The method for sensing a surface aircraft by a guide vehicle based on radar and visual fusion according to claim 1, characterized in that: The S2 further includes: The first decoding layer is used to extract geometric features from radar point cloud data, and the second decoding layer is used to extract semantic features from image data acquired by the camera; Define the first input sequence as the Query sequence (Q), and the second input sequence as the Key sequence (K) and the Value sequence (V); Then the attention weights of the comprehensive Query sequence (Q), Key sequence (K) and Value sequence (V) satisfy the expression: Among them, the softmax function normalizes the attention score to obtain the standardized attention weight. Indicates the dimension of the Key sequence, represents transpose, represents the transpose of K; Based on the SMCA model, the mask matrix Multiply it with the attention weight to generate a bird's-eye view after the radar and vision are fused, where the mask matrix Satisfies the expression: in, represents the element position number of M, , represents the two-dimensional center point calculated by projecting the query prediction onto the image plane, The minimum circumference radius of the projection angle of the 3D bounding box; represents the hyperparameters that tune the Gaussian distribution.

6. The method for sensing a surface aircraft by a guide vehicle based on radar and visual fusion according to claim 1, characterized in that: The loss function satisfies the expression: in, is the binary cross entropy loss, is the L1 norm difference between the predicted bird's-eye view center and the actual bird's-eye view center, They are the cross entropy loss coefficient, the norm difference coefficient, and the intersection-over-union loss coefficient based on uncertainty calculation. is the confidence parameter, is the predicted 3D bounding box, is the true 3D bounding box.

7. A system for realizing the method for sensing the surface aircraft by the guide vehicle based on radar and visual fusion as described in any one of claims 1 to 6, characterized in that: include: Radar module, used to collect 3D radar point cloud data of aircraft on the scene; A vision module, including a camera, for collecting image data of aircraft on the scene; A processing unit, integrated with an SMCA model and a feedforward neural network, for performing effective fusion of the radar point cloud data and the image data, and outputting the real-time distance of the surface aircraft relative to the airport guidance vehicle and the attitude information of the aircraft; The control module receives the output parameters of the processing unit and controls the driving strategy of the guide vehicle.

Citation Information

Cited By

  • Three-dimensional target detection method for improving multi-modal fusion

    CN120877050A