A multi-modal place recognition method based on a multi-scale self-attention mechanism

By fusing multimodal sensor data through a multi-scale self-attention mechanism, generating multimodal global descriptors and reordering them, the robustness problem of location recognition technology under environmental and perspective changes is solved, and the recognition accuracy and generalization ability are improved.

CN119625667BActive Publication Date: 2025-11-21BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411444267.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-16
Publication Date
2025-11-21
Estimated Expiration
2044-10-16

AI Technical Summary

Technical Problem

Existing location recognition technologies suffer from reduced recognition performance when faced with changes in environmental factors such as lighting and weather, and are not robust to changes in vehicle perspective. Furthermore, their multimodal feature fusion is insufficient, making it difficult to achieve stable recognition under different environmental conditions.

Method used

Using panoramic multi-view RGB images, LiDAR environmental point clouds, and millimeter-wave radar environmental point clouds as inputs, multi-modal feature fusion is performed through a multi-scale self-attention mechanism to generate a multi-modal global descriptor (LCGD). The descriptor is then reordered by combining Euclidean distance and the probability density function of the scattering cross section of the millimeter-wave radar point cloud to improve recognition accuracy and generalization ability.

Benefits of technology

This enhances the accuracy and generalization of location identification technology under different environmental conditions, reduces the impact of single sensor failure, and achieves higher accuracy location identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625667B_ABST
    Figure CN119625667B_ABST
Patent Text Reader

Abstract

The application discloses a multi-modal place recognition method based on a multi-scale self-attention mechanism, and belongs to the field of place recognition.The application is implemented in the following manner: simultaneously taking a look-around multi-view RGB image, a laser radar environment point cloud and a millimeter wave radar environment point cloud as input, performing multi-modal feature fusion on multiple feature scales through a self-attention mechanism, generating a multi-modal global LCGD descriptor, and screening k candidate samples according to the LCGD descriptor in terms of an Euclidean distance; extracting a point cloud scattering cross-sectional area probability density function from the millimeter wave radar environment point cloud, calculating a reordering total distance, reordering the k candidate samples, and taking the coordinates of a candidate sample with the minimum reordering distance as the current time coordinates of the vehicle.The application can be used to solve the problems that the existing place recognition technology has poor robustness to changes in environmental factors such as light and weather and changes in the vehicle's own perspective, and to enhance the recognition accuracy of the place recognition method and the generalization of the place recognition method under different environmental conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of location recognition technology, and in particular relates to a multimodal location recognition method based on a multi-scale self-attention mechanism. Background Technology

[0002] Location recognition is a key technology in intelligent vehicle navigation and positioning. It uses the vehicle's onboard sensors to detect and match locations within its environment. During the vehicle's movement, location recognition technology extracts a global location descriptor, expressed as a vector or matrix, for each location. When the vehicle needs to obtain its own location information, it compares the current location descriptor with descriptors in a database to determine the vehicle's current position on the global map. It can also determine if the vehicle has reached the vicinity of historical locations in the database. Location recognition technology can solve the problem of rapid global positioning for intelligent vehicles in long-term, repetitive driving environments. It can also be applied to Simultaneous Localization and Mapping (SLAM) technology to solve loop closure detection problems and eliminate accumulated positioning errors in SLAM.

[0003] Location recognition technology relies on sensor observation, but each sensor has its inherent shortcomings and degradation scenarios. For example, although cameras are rich in semantic information and inexpensive, they are easily affected by lighting conditions. LiDAR is not affected by lighting conditions and can capture three-dimensional structural information of the environment, but its data is sparse compared to camera data and is still affected by extreme weather. Millimeter-wave radar is not affected by extreme weather and lighting conditions, but it is even sparser than LiDAR. Therefore, location recognition algorithms based on single-modal sensor data are limited by the inherent shortcomings of the sensors and cannot achieve stable recognition results under various environmental conditions. Multimodal fusion technology is an effective means to overcome the shortcomings of single sensors. Therefore, introducing environmental observation data of multiple modalities into location recognition tasks can enhance the generalization ability of location recognition technology under different environmental conditions. Although existing technologies have made progress in research based on multi-sensor fusion, the following problems still exist: (1) It is difficult to apply the fusion of only cameras and LiDAR in weather conditions such as rain and fog; (2) It is not conducive to multimodal feature fusion if only single-view camera data is used, and different modal data cannot be aligned; (3) The influence of vehicle heading angle on location recognition has not been given sufficient attention and cannot cope with changes in viewing angle. Summary of the Invention

[0004] To address the aforementioned issues, this invention aims to provide a multimodal location recognition method based on a multi-scale self-attention mechanism. This method simultaneously uses panoramic multi-view RGB images, LiDAR environmental point clouds, and millimeter-wave radar environmental point clouds as inputs. It fuses multimodal features at multiple feature scales using a self-attention mechanism to generate a LiDAR-Camera Global Descriptor (LCGD). Then, k candidate samples are selected based on Euclidean distance using the LCGD descriptor. Next, the probability density function of the point cloud scattering cross-section is extracted from the millimeter-wave radar environmental point cloud, and the total reordering distance is calculated. The k candidate samples are then reordered, and the coordinates of the candidate sample with the smallest reordering distance are used as the current coordinates of the intelligent vehicle. This method addresses the shortcomings of existing location recognition technologies, such as poor robustness to changes in environmental factors like lighting and weather, as well as changes in the vehicle's own perspective. This enhances the recognition accuracy and generalization ability of location recognition technology under different environmental conditions.

[0005] The objective of this invention is mainly achieved through the following technical solutions:

[0006] A multimodal location recognition method based on a multi-scale self-attention mechanism selects a surround-view multi-view camera, a multi-line lidar, and a millimeter-wave radar as environmental perception sensors. The multimodal location recognition method includes the following steps:

[0007] Step S1: Using an intelligent vehicle equipped with a surround-view multi-view camera, multi-line LiDAR, and millimeter-wave radar, collect multiple sets of environmental observation samples to form a dataset. The environmental observation samples include the coordinates of the intelligent vehicle at the time of collection, the vehicle's odometer information, RGB images of the surround-view multi-view, LiDAR environmental point cloud, and multi-view millimeter-wave radar environmental point cloud. The surround-view multi-view RGB images include six RGB images: front view, left front view, left rear view, rear view, right rear view, and right front view. The multi-view millimeter-wave radar environmental point cloud includes front view, left front view, left rear view, right rear view, and right front view.

[0008] Step S2: Obtain multiple lidar depth images by spherical projection of the lidar environmental point cloud from S1;

[0009] Step S3: Back-project the multi-view RGB image from S1 onto the lidar coordinate system to obtain an RGB point cloud that simultaneously contains RGB values ​​and three-dimensional coordinates; then stitch the RGB point clouds obtained from multiple views together and use spherical projection to obtain a panoramic RGB image from the lidar view.

[0010] Step S4: Group the multiple LiDAR depth images obtained in S2 and the panoramic RGB images obtained in S3 to construct a three-dimensional data set; the three-dimensional data set contains multiple sample groups, each sample group contains a query sample, a positive sample and multiple negative samples; use the three-dimensional data set as the training dataset;

[0011] Step S5: Use a multimodal descriptor extraction network to extract and aggregate features from the panoramic RGB image obtained in S3 and the LiDAR depth map obtained in S2 to obtain the LCGD descriptor; the multimodal descriptor extraction network is composed of a multimodal feature fusion module and a global feature aggregation module.

[0012] The panoramic RGB image obtained in S3 and the LiDAR depth map obtained in S2 are used to extract features through the multimodal feature fusion module to obtain visual feature maps and LiDAR feature maps. Then, the visual feature maps and LiDAR feature maps are used to obtain visual sub-descriptors and LiDAR sub-descriptors through the global feature aggregation module. Finally, the two are concatenated to obtain the LCGD descriptor.

[0013] Step S6: Using Triplet Loss, the multimodal descriptor extraction network in S5 is trained using the training dataset obtained in step S4, so that the descriptors of the query samples are close to the positive sample descriptors, while avoiding the negative sample descriptors; thus obtaining the trained multimodal descriptor extraction network.

[0014] Step S7: Using the multimodal descriptor extraction network trained in step S6, process each sample in the data subset db in step S4 according to the method in step S5 to obtain the descriptor database.

[0015] Step S8: In the intelligent vehicle localization stage, perform steps S2, S3, and S5 on the RGB images and LiDAR environmental point clouds collected at the current time to obtain the LCGD descriptor at the current time; compare the LCGD descriptor at the current time with all descriptors in the descriptor database of S7 according to the Euclidean distance, select the k with the smallest distance from the descriptor database as candidate descriptors, and use their corresponding samples as candidate samples;

[0016] Step S9: Perform the following operations on the samples in the multi-view millimeter-wave radar environment point cloud acquired at the current moment: Project the millimeter-wave radar environment point clouds of all views in the current frame sample to the forward-looking millimeter-wave radar coordinate system to obtain the millimeter-wave radar environment point cloud in the forward-looking coordinate system; Using vehicle odometer information, project the millimeter-wave radar environment point clouds in the forward-looking coordinate system of the previous 3 frames of the current frame sample to the current sample coordinate system to obtain the dense millimeter-wave radar environment point cloud at the current moment;

[0017] Step S10: Remove the dynamic points from the dense millimeter-wave radar environment point cloud obtained in S9 at the current moment to obtain the static point cloud; calculate the RCS probability density function of the static point cloud at the current moment.

[0018] Step S11: Using the methods of steps S9 and S10, calculate the RCS probability density function of the point cloud of the k candidate samples obtained in S8.

[0019] Step S12: Using the RCS probability density function of the current point cloud obtained in step S10 and the RCS probability density function of the point cloud of the k candidate samples obtained in step S11, calculate the RCS probability distribution distance; combine the RCS probability distribution distance with the Euclidean distance to obtain the total reordering distance; the Euclidean distance is the Euclidean distance between the current LCGD descriptor and the k candidate descriptors obtained in step S8; reorder the k candidate samples using the total reordering distance;

[0020] Step S13: Using the re-sorting result of step S12, select the sample with the smallest re-sorting distance as the matching sample, and take the coordinates of the sample as the coordinates of the intelligent vehicle at the current moment, so as to complete the location identification with higher accuracy.

[0021] Furthermore, in step S1, the panoramic multi-view RGB image can cover a 360° field of view, the lidar point cloud contains the spatial location of the point cloud, and each frame of data in the millimeter-wave radar observation sequence contains the spatial location of the point cloud, radial velocity, and RCS value.

[0022] Furthermore, the process of transforming the spherical projection of the lidar point cloud into a depth map in step S2 is given by the following formula:

[0023]

[0024] in Let r = ||p||2 be the pixel coordinates of the depth map, and let p = (xyz) be the ranging value of the laser point. T Let h and w represent the coordinates of the laser point, and h and w represent the width and height of the target depth map, respectively. f = f up +f down Indicates the field of view of the lidar;

[0025] Furthermore, the specific steps for obtaining the panoramic RGB image from the lidar viewpoint in step S3 are as follows:

[0026] Step S31: Perform depth estimation on the RGB images from multiple viewing angles to obtain the absolute depth value of each pixel in each image at the real scale.

[0027] Step S32: Using the absolute depth value obtained in step S31 and the intrinsic and extrinsic parameter matrices of each camera, project the RGB value of each pixel in the image from a certain viewpoint from the pixel coordinate system to the lidar coordinate system to obtain the RGB point cloud.

[0028] Step S33: Perform projection processing on each view image using step S32, and stitch together all the obtained results to obtain an RGB point cloud containing a 360° view in the lidar coordinate system.

[0029] Step S34: Perform spherical projection on the RGB point cloud containing a 360° view obtained in step S33 to obtain a panoramic RGB image from the perspective of the lidar.

[0030] Furthermore, the specific steps for constructing the training dataset in step S4 are as follows:

[0031] Step S41: In the dataset of step S1, based on the distance threshold dis_thres, select environmental observation samples at intervals according to the coordinates in the environmental observation samples. The selected environmental observation samples are called data subset db, and the unselected environmental observation samples are called data subset query, thus obtaining the query sample group.

[0032] Step S42: In the data subset query of S41, select environmental observation samples with sample timestamps less than the threshold date_thres as the data subset train_query, and the remaining samples as the data subset val_query;

[0033] Step S43: For each query sample in the data subset query, find the negative sample group corresponding to the query sample whose distance is greater than the threshold neg_thres in the db subset; find the positive sample group corresponding to the query sample whose distance is less than the threshold pos_thres in the db subset, and randomly select a positive sample from them as the positive sample corresponding to the query sample.

[0034] The training set of triples is obtained by combining each query sample in the data subset train_query with its corresponding positive sample and multiple negative samples; the validation set of triples is obtained by combining each query sample in the data subset val_query with its corresponding positive sample and multiple negative samples; the training set of triples and the validation set of triples together constitute the training dataset.

[0035] Furthermore, the specific implementation steps of step S5 are as follows:

[0036] Step S51: Denote the panoramic RGB image obtained in S3 as having a size of B×C. I ×H I ×W IThe tensor, B represents the size of the computation batch, H I and W I C represents the height and width of the RGB image, respectively. I This represents the number of RGB image channels; the LiDAR depth map obtained in S2 is denoted as having a size of B×C. L ×H L ×W L The tensor, H L and W L C represents the height and width of the LiDAR depth map, respectively. L This indicates the number of channels in the lidar depth map;

[0037] The i-th layer features extracted from the panoramic RGB image and the LiDAR depth map are denoted as follows: and Will and The features from different modalities are input together into the multimodal feature fusion module for fusion; let the fused visual features and LiDAR features be respectively... and Both have the same shape as the input feature tensor; and and The modal fusion features are obtained by adding them separately. and and

[0038] Step S52, Modal fusion feature I F and L F Each feature aggregation module performs its own feature aggregation to obtain visual sub-descriptors. and lidar sub-descriptor

[0039] Step S53: Take the I obtained in step S52 D and L D By concatenating the data along the channel dimension and then performing L2 normalization, the LCGD descriptor is obtained.

[0040] Furthermore, the formula for global feature aggregation in step S52 is described as follows:

[0041]

[0042] Where N is the number of feature vectors, K is the number of clusters, and w k b k c kAll are parameters to be learned, x i (j) represents the eigenvector x i The j-th component;

[0043] Further, the step S51 of... and The specific steps for fusing different modal features in the multimodal feature fusion module are as follows:

[0044] Step S511, the input feature map of the multimodal feature fusion module is denoted as... and Corresponding to visual features and LiDAR features respectively; through two vertically compressed VC layers, respectively... and Compressed to a size of and Horizontal characteristics;

[0045] The VC layer consists of a 1×1 Conv layer, a Sigmoid activation layer, a 1×1 Conv layer, a MaxPooling layer, and four 1-D Conv layers in sequence; the MaxPooling layer is used for feature downsampling; a skip connection is used between the first and third 1-D Conv layers to accelerate convergence;

[0046] Step S512: Concatenate the outputs of the VC layer in the width direction to obtain the multimodal feature embedding. Then, multi-head self-attention is used to perform multimodal feature fusion to obtain F′. e The size of the feature map does not change before and after this step;

[0047] Step S513, F′ e Available in two sizes and The tensor is then used to copy the feature map along the height dimension, resulting in the fused features output by the multimodal feature fusion engine. and

[0048] Furthermore, the triplet loss function in step S6 is calculated as follows:

[0049]

[0050] Where q and p q These represent the query sample and the corresponding positive sample, respectively. This represents the i-th negative sample corresponding to the query sample, [...] + Represents the hinge loss function [...] + =max(0,…), d(·) represents the Euclidean distance, m is a constant representing the distance threshold of the descriptor, and Nneg It is the number of negative samples;

[0051] Furthermore, the specific process of removing dynamic points and calculating the probability density function of the scattering cross section (RCS) in step S10 is as follows:

[0052] Step S101: For all measurement points of the dense millimeter-wave radar in the current frame, calculate the radial velocity of each point relative to the millimeter-wave radar:

[0053]

[0054] in Let v represent the radial velocity at point i. i Let θ represent the velocity of the i-th point in the millimeter-wave radar coordinate system. i Let θ represent the azimuth angle of the i-th point. vi This represents the velocity measurement angle at the i-th point;

[0055] Step S102: Since the radial velocity of most measurement points is 0 when the vehicle is stationary, if more than 60% of the measurement points in the millimeter-wave radar point cloud have a radial velocity of less than 1 m / s, the vehicle in that frame is determined to be stationary; otherwise, the vehicle in that frame is determined to be moving.

[0056] Step S103: If the vehicle in the current frame is stationary, the measurement points with a radial velocity greater than 1 m / s are determined to be dynamic points; if the vehicle in the current frame is moving, the stationary measurement points satisfy the following equation:

[0057]

[0058] in Let v represent the radial velocity at point i. s The speed of the millimeter-wave radar is α, the vehicle's heading angle is θ. i Let be the azimuth angle of the i-th point; the dynamically measured point is a discrete point for the above formula, therefore, based on the above formula, let α and v s As a solution parameter, the Random Sampling Consensus Algorithm (RANSAC) is used to calculate outliers, and outliers are identified as dynamic measurement points and then filtered out.

[0059] Step S104: Normalize the RCS value of the millimeter-wave radar measurement point to the range of (0,1) using the maximum and minimum values ​​of the RCS measurement value of this frame;

[0060] Step S105, using bin RCS To determine the interval size, calculate the RCS histogram and normalize it across all intervals to eliminate the influence of different point cloud quantities in different frames on the histogram, thus obtaining the RCS probability density function.

[0061] Furthermore, the specific steps for calculating the total reordered distance in step S12 are as follows:

[0062] Step S121: Calculate the probability distribution distance between the query RCS probability density function and the candidate RCS probability density function based on KL divergence:

[0063]

[0064] Where H(R1) and H(R2) represent the RCS probability density functions obtained in step S10, and k is the index of the discrete interval of the probability density function;

[0065] Step S122: Calculate the total reordering distance using the RCS probability distribution distance and the LCGD Euclidean distance of the multimodal descriptor:

[0066] d rerank (q,c)=αd R (H(q),H(c))+(1-α)d E (LCGD(q),LCGD(c))

[0067] Where q and c represent the query sample and candidate sample, respectively, f(·) represents the LCGD descriptor extracted from the input sample according to step S5, and d E (·) represents Euclidean distance, and α is a set weight used to adjust the proportion of the two distance measurements;

[0068] Step S123: Calculate the total reordering distance d between the query sample and each candidate sample. rerank and with d rerank Sort the k candidate samples in descending order.

[0069] Beneficial effects:

[0070] 1. This invention discloses a multimodal location recognition method based on a multi-scale self-attention mechanism. By simultaneously using surround-view multi-view RGB images, LiDAR environmental point clouds, and millimeter-wave radar environmental point clouds as inputs, and fusing multimodal features at multiple feature scales through a self-attention mechanism, it can leverage the advantages of different sensors, reduce the impact of single sensor failure on location recognition performance, and solve the problems of decreased recognition performance when facing changes in environmental factors such as lighting and weather, as well as the poor robustness of existing technologies to changes in the vehicle's own perspective. This enhances the recognition accuracy and generalization of location recognition technology under different environmental conditions.

[0071] 2. This invention discloses a multimodal location recognition method based on a multi-scale self-attention mechanism. It simultaneously uses multi-view RGB images, LiDAR environmental point clouds, and millimeter-wave radar environmental point clouds as inputs to achieve multimodal data acquisition and data spatial alignment. A training set is constructed by acquiring triplet data sets, and multimodal feature fusion is performed at multiple feature scales using a self-attention mechanism to generate a multimodal global descriptor (LCGD). The LCGD descriptors of all samples in the database are extracted according to Euclidean distance, and a multimodal descriptor extraction network is trained using triplet loss. The LCGD descriptor is extracted based on the current observation, and the k nearest neighbor candidate samples are queried in the database. This invention achieves more robust location recognition through multi-sensor fusion, leveraging the advantages of different sensors, reducing the impact of single sensor failure on location recognition performance, and improving the robustness of multimodal location recognition.

[0072] 3. The present invention discloses a multimodal location recognition method based on a multi-scale self-attention mechanism. Based on the beneficial effect 2, dynamic points are removed from the millimeter-wave radar point cloud, and the probability density function of the scattering cross section of the millimeter-wave radar point cloud is calculated. The total reordering distance is calculated by combining the Euclidean distance of the LCGD descriptor, and the k candidate LCGD descriptors are reordered. The coordinates of the sample with the closest total reordering distance are used as the coordinates of the current intelligent vehicle, thereby achieving higher accuracy location recognition.

[0073] 4. This invention discloses a multi-modal location identification method based on a multi-scale self-attention mechanism. In the process of removing dynamic points and calculating the RCS probability density function, the Random Sampling Consensus (RANSAC) algorithm is used to calculate outliers, and these outliers are identified as dynamic measurement points and filtered out. The RCS values ​​of millimeter-wave radar measurement points are normalized to the (0,1) range using the maximum and minimum values ​​of the RCS measurement values ​​in that frame. (The last sentence appears to be incomplete and possibly refers to a different method.) RCS To determine the interval size, an RCS histogram is calculated and normalized across all intervals to eliminate the influence of varying point cloud quantities in different frames on the histogram, resulting in the RCS probability density function. The probability distribution distance between the query RCS probability density function and the candidate RCS probability density function is then calculated based on KL divergence. The total reordering distance is calculated using the RCS probability distribution distance and the multimodal descriptor LCGD Euclidean distance, which improves the accuracy of candidate sample retrieval and enhances the recognition accuracy and generalization ability of location identification technology under different environmental conditions. Attached Figure Description

[0074] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts.

[0075] Figure 1 A flowchart illustrating multimodal location recognition based on a multi-scale self-attention mechanism is shown.

[0076] Figure 2 A schematic diagram illustrating the principle of multimodal location recognition based on a multi-scale self-attention mechanism is shown.

[0077] Figure 3 A flowchart illustrating multimodal data spatial alignment is shown schematically.

[0078] Figure 4 The flowchart illustrating the construction process of the database sample set and query sample set is shown in the diagram.

[0079] Figure 5 A flowchart illustrating the construction process of the triplet training dataset is shown.

[0080] Figure 6 A schematic diagram of the multimodal descriptor extraction network structure is shown.

[0081] Figure 7 The schematic diagram illustrates the structure of the multimodal feature fusion module. Detailed Implementation

[0082] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.

[0083] One specific embodiment of the present invention discloses a multimodal location identification method based on a multi-scale self-attention mechanism. This method fuses multimodal information through a multi-head self-attention mechanism to generate LCGD descriptors, producing candidate samples. The candidate samples are then reordered using a millimeter-wave radar post-fusion approach, thereby achieving robust location identification. The method principle is described in [link to method description]. Figure 2 .

[0084] This invention provides a multimodal location identification method based on a multi-scale self-attention mechanism. Location identification is achieved by extracting LCGD descriptors from input multimodal information using a deep neural network and then fusing and reordering them after millimeter-wave radar analysis. The LCGD descriptor extraction module consists of three parts: multimodal data spatial alignment, multimodal feature extraction and fusion, and multimodal feature aggregation. The millimeter-wave radar post-fusion and reordering includes modules for dynamic point filtering, RCS probability density function calculation, and reordering of candidate descriptors.

[0085] like Figure 1 As shown in the figure, this embodiment discloses a multimodal location recognition method based on a multi-scale self-attention mechanism, and the specific implementation steps are as follows:

[0086] Step S1: Using an intelligent vehicle equipped with a surround-view multi-view camera, multi-line lidar, and millimeter-wave radar, collect multiple sets of environmental observation samples in the environment to form a dataset.

[0087] Specifically, when intelligent vehicles collect environmental data, preferably, the lidar beamwidth should be greater than or equal to 32 to achieve better recognition results. The millimeter-wave radar used should be a frequency-modulated continuous wave (FMCW) radar, capable of simultaneously measuring distance and speed. During environmental data collection, it is necessary to ensure that the recorded road and building structures remain largely unchanged. The intelligent vehicle is equipped with an inertial navigation system to remove invalid points, points outside the region of interest, and correct motion distortion from the collected raw point cloud data. The surround-view multi-view RGB images include: front view, left front view, left rear view, rear view, right rear view, and right front view, totaling six RGB images. The multi-view millimeter-wave radar environmental point cloud includes: front view, left front view, left rear view, right rear view, and right front view.

[0088] Step S2: Obtain multiple lidar depth images by spherical projection of the lidar environmental point cloud from S1;

[0089] Specifically, when calculating the lidar depth map from the lidar point cloud using spherical projection, the coordinate transformation formula from the three-dimensional coordinates of the lidar points to the following:

[0090]

[0091] in Let r = ||p||2 be the pixel coordinates of the depth map, and let p = (xyz) be the ranging value of the laser point. T Let h and w represent the coordinates of the laser point, and h and w represent the width and height of the target depth map, respectively. f = f up +f down This represents the field of view of the lidar; if multiple point clouds are projected onto the same depth map coordinate using the above formula, the minimum depth of these point clouds is selected as the depth value at that coordinate on the depth map.

[0092] Through spherical projection, the originally disordered and massive sparse point cloud data is transformed into ordered, reduced-volume, and more suitable deep neural network processing data.

[0093] Step S2: Obtain multiple lidar depth images by spherical projection of the lidar environmental point cloud from S1;

[0094] Step S3: Back-project the multi-view RGB image from S1 onto the lidar coordinate system to obtain an RGB point cloud that simultaneously contains RGB values ​​and three-dimensional coordinates; then stitch the RGB point clouds obtained from multiple views together and use spherical projection to obtain a panoramic RGB image from the lidar view.

[0095] like Figure 3 As shown, depth estimation is first performed on the multi-view RGB images to obtain the absolute depth value of each pixel in each image at the real scale. Next, using the estimated absolute depth value d and the intrinsic and extrinsic parameters of each camera, the RGB values ​​of each pixel in the image at a certain viewpoint are projected from the pixel coordinate system to the LiDAR coordinate system to obtain an RGB point cloud. The RGB point cloud data simultaneously contains the three-dimensional coordinates and RGB values ​​of the point cloud. For each viewpoint image, the process described in step S32 is used to calculate and stitch the results to obtain an RGB point cloud in the LiDAR coordinate system containing a 360° viewpoint.

[0096] Then, the RGB point cloud containing a 360° view is projected onto the spherical surface described in step S2 to obtain a panoramic RGB image from the perspective of the lidar, thereby completing cross-modal data space alignment.

[0097] Step S4: Group the multiple LiDAR depth images obtained in S2 and the panoramic RGB images obtained in S3 to construct a three-dimensional data set. Each three-dimensional data set contains multiple sample groups, and each sample group contains one query sample, one positive sample, and multiple negative samples. Use the three-dimensional data set as the training dataset. The construction process is as follows: Figure 4 and Figure 5 As shown;

[0098] First, from the dataset collected in step S1, a database sample set and a query sample set are divided. Data sets `db` and `query` are defined as the database sample set and the query sample set, respectively. For each sample, the distance between that sample and all samples in the `db` set is calculated. If the minimum distance is greater than the threshold `dis_thres` or the `db` set is empty, the sample is added to the `db` set; otherwise, it is added to the `query` set. Through these operations, the database set is essentially obtained by sampling based on coordinate intervals, using `dis_thres` as the distance threshold.

[0099] After obtaining the database sample set db and the query sample set query, according to the sample timestamp, the samples in the query set whose timestamps are less than a threshold date_thres are added to the train_query set for training, and the remaining samples are added to the val_query set for validation;

[0100] For each query sample in the `train_query` and `val_query` subsets, find samples in the `db` subset whose distance to that sample is less than the threshold `pos_thres` as the corresponding positive sample group, and randomly select one positive sample from them; for each sample in the `query` subset, find samples in the `db` subset whose distance to that sample is greater than the threshold `neg_thres` as the corresponding negative sample group; combine each query sample in the `train_query` subset with its corresponding positive sample and multiple negative samples to obtain the triplet training set; combine each query sample in the `val_query` subset with its corresponding positive sample and multiple negative samples to obtain the triplet validation set; the triplet training set and the triplet validation set together constitute the training dataset;

[0101] Step S5, using Figure 6 The multimodal descriptor extraction network shown extracts and aggregates features from the panoramic RGB image obtained in S3 and the LiDAR depth map obtained in S2 to obtain the LCGD descriptor; the multimodal descriptor extraction network is composed of a multimodal feature fusion module and a global feature aggregation module.

[0102] First, denote the input panoramic RGB image as having dimensions B×C. I ×H I ×W I The tensor, where B represents the batch size of the data, H I and W I C represents the height and width of the RGB image, respectively. I This represents the number of RGB image channels; the input LiDAR depth map is denoted as a B×C map. L ×H L ×W L The tensor, H L and W L C represents the height and width of the LiDAR depth map, respectively. L This represents the number of channels in the LiDAR depth map. A ResNet residual feature extraction network with pre-trained weights is used to extract features from the multi-view RGB image and the LiDAR depth map, respectively. The extracted i-th layer features are denoted as... and Will and The visual features and LiDAR features are input together into the multimodal feature fusion module for fusing different modal features. Let the fused visual features and LiDAR features be denoted as follows: and Both have the same shape as the input feature tensor; and and The modal fusion features are obtained by adding them separately. and After that and The data is fed into the next layer of the ResNet network, and the above process is repeated for feature extraction and fusion. Then, the modal fusion features are processed. and Each feature aggregation module performs its own feature aggregation to obtain visual sub-descriptors. and lidar sub-descriptor Finally, the sub-descriptor I obtained from the above steps is... D and L D By concatenating the data along the channel dimension and then performing L2 normalization, a multimodal global descriptor is obtained.

[0103] Furthermore, the structure of the multimodal feature fusion module in step S5 is as follows: Figure 7 As shown. The input feature map of this module is denoted as... and These correspond to the vision branch and the LiDAR branch, respectively, where i represents the number of layers in the feature extractor. For ease of writing, the subscript i is no longer used to specify the number of layers in this module description. First, two vertically compressed (VC) layers are used to extract the features... and Compressed to a size of and The horizontal features are described. The VC layer consists of a 1×1 Conv layer, a Sigmoid activation layer, a 1×1 Conv layer, a MaxPooling layer, and four 1-D Conv layers. The MaxPooling layer is used for feature downsampling. Skip connections are used between the first and third 1-D Conv layers to accelerate convergence. The outputs of the VC layers are then concatenated in the width direction to obtain the multimodal feature embedding. Then, multi-head self-attention is used to perform multimodal feature fusion to obtain F′. e The feature map size remains unchanged before and after this step. F′ e Available in two sizes and The tensor is then used to copy the feature map along the height dimension, restoring it to the same height as the input. and With the same size, the fused features of the i-th layer multimodal feature fusion machine are finally obtained. and

[0104] Furthermore, the formula for feature aggregation using the differentiable feature aggregation module in step S5 is described as follows:

[0105]

[0106] Where N is the number of feature vectors, K is the number of clusters, and w k b k c k All are parameters to be learned, x i (j) represents the eigenvector x i The j-th component;

[0107] Step S6: Using Triplet Loss, the multimodal descriptor extraction network in S5 is trained using the training dataset obtained in step S4, so that the descriptors of the query samples are close to the positive sample descriptors, while avoiding the negative sample descriptors; thus obtaining the trained multimodal descriptor extraction network.

[0108] Furthermore, the triplet loss used in training deep neural networks is calculated as follows:

[0109]

[0110] Where q and p q These represent the query sample and the corresponding positive sample, respectively. This represents the i-th negative sample corresponding to the query sample, [...] + Represents the hinge loss function [...] + =max(0,…), d(·) represents the Euclidean distance, m is a constant representing the distance threshold of the descriptor, and N neg This represents the number of negative samples. The purpose of this loss function is to learn to reduce the distance between the LCGD descriptor of the query sample and the LCGD descriptor of the positive sample, while increasing the distance between the LCGD descriptors of the query sample and the negative sample.

[0111] Step S7: Using the multimodal descriptor extraction network trained in step S6, process each sample in the data subset db in step S4 according to the method in step S5 to obtain the descriptor database.

[0112] Step S8: In the intelligent vehicle localization stage, perform steps 2, 3, and 5 on the RGB images and LiDAR environmental point clouds collected at the current time to obtain the LCGD descriptor at the current time; compare the LCGD descriptor at the current time with all descriptors in the descriptor database of S7 according to the Euclidean distance, select the k with the smallest distance from the descriptor database as candidate descriptors, and use their corresponding samples as candidate samples;

[0113] Step S9: Perform the following operations on the samples in the multi-view millimeter-wave radar environment point cloud acquired at the current moment: Project the millimeter-wave radar environment point clouds of all views in the current frame sample to the forward-looking millimeter-wave radar coordinate system to obtain the millimeter-wave radar environment point cloud in the forward-looking coordinate system; Using vehicle odometer information, project the millimeter-wave radar environment point clouds in the forward-looking coordinate system of the previous 3 frames of the current frame sample to the current sample coordinate system to obtain the dense millimeter-wave radar environment point cloud at the current moment;

[0114] Step S9: Remove the dynamic points from the dense millimeter-wave radar environment point cloud obtained in S9 at the current moment to obtain the static point cloud; calculate the probability density function of the scattering cross section (RCS) of the static point cloud at the current moment.

[0115] Different dynamic point filtering strategies are required for the vehicle's moving and stationary states. First, for all measurement points of the dense millimeter-wave radar in the current frame, the radial velocity of each point relative to the millimeter-wave radar is calculated:

[0116]

[0117] in Let v represent the radial velocity at point i. i Let θ represent the velocity of the i-th point in the millimeter-wave radar coordinate system. i Let θ represent the azimuth angle of the i-th point. vi This represents the velocity measurement angle at the i-th point.

[0118] Since the radial velocity of most measurement points is 0 when the vehicle is stationary, if more than 60% of the measurement points in a frame of millimeter-wave radar point cloud have a radial velocity less than the threshold v_thres, the vehicle in that frame is determined to be stationary; otherwise, the vehicle in that frame is determined to be moving. If the vehicle in the current frame is stationary, measurement points with a radial velocity greater than 1 m / s are determined to be dynamic points. It should be noted that during data acquisition, the real-time movement speed of the vehicle itself can also be recorded, and the vehicle's movement status can be directly determined using the real-time speed data.

[0119] If the vehicle is moving in the current frame, then stationary measurement points in the environment satisfy the following equation:

[0120]

[0121] in Let v represent the radial velocity at point i. s The speed of the millimeter-wave radar is α, the vehicle's heading angle is θ. iLet be the azimuth angle of the i-th point. Static measurement points in the environment should conform to the above formula, while dynamic measurement points, due to their arbitrary velocities, will not strictly conform to the above formula. Therefore, dynamic measurement points in the environment are discrete points according to the above formula. Thus, based on the above formula, α and v... s As a solution parameter, the Random Sampling Consensus Algorithm (RANSAC) is used to calculate outliers, which are then identified as dynamic measurement points. Based on the identification results, dynamic points are filtered out.

[0122] The RCS values ​​of the millimeter-wave radar measurement points were then normalized to the (0,1) range using the maximum and minimum values ​​of the RCS measurements in that frame. (bin) RCS Given the interval size, calculate the RCS histogram and normalize it across all intervals to eliminate the influence of different point cloud quantities in different frames on the histogram, thus obtaining the RCS probability density function.

[0123] Step S11: Using the methods of steps S9 and S10, calculate the RCS probability density function of the point cloud of the k candidate samples obtained in S8.

[0124] Step S12: Using the RCS probability density function of the current point cloud obtained in step S10 and the RCS probability density function of the point cloud of the k candidate samples obtained in step S11, calculate the RCS probability distribution distance; combine the RCS probability distribution distance with the Euclidean distance to obtain the total reordering distance; the Euclidean distance is the Euclidean distance between the current LCGD descriptor and the k candidate descriptors obtained in step S8; reorder the k candidate samples using the total reordering distance;

[0125] First, the probability distribution distance between the query RCS probability density function and the candidate RCS probability density function is calculated based on KL divergence:

[0126]

[0127] Where P(R) q ) and P(R candidate ) represents the RCS probability density function obtained in step S9, corresponding to the query sample and the candidate sample respectively, and k is the discrete interval index of the probability density function.

[0128] The total reordering distance is calculated using the RCS probability distribution distance and the LCGD Euclidean distance of the multimodal descriptor:

[0129] d rerank (q,c)=αd R (H(q),H(c))+(1-α)d E (LCGD(q),LCGD(c))

[0130] Where q and c represent the query sample and candidate sample, respectively, LCGD(·) represents the extraction of multimodal descriptors for the input sample according to step S5, and d E (·) represents Euclidean distance, and α is the set weight used to adjust the proportion of the two distance measurements.

[0131] Step S13: Using the re-sorting result of step S12, select the sample with the smallest re-sorting distance as the matching sample, and take the coordinates of the sample as the coordinates of the intelligent vehicle at the current moment, so as to complete the location identification with higher accuracy.

[0132] By employing the above steps and simultaneously using measurement information from three sensors—a surround-view multi-angle camera, LiDAR, and millimeter-wave radar—as data sources, and based on a multi-head self-attention mechanism, multimodal feature fusion is performed at multiple scales. Furthermore, millimeter-wave radar is used for post-fusion to reorder candidate locations. This approach addresses the shortcomings of existing location recognition technologies, such as poor robustness to changes in lighting and weather conditions, as well as changes in the vehicle's own perspective. It provides a novel location recognition scheme that enhances the accuracy and generalization ability of location recognition technology under different environmental conditions.

[0133] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by hardware related to computer program instructions, and the program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.

[0134] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A multimodal location recognition method based on a multi-scale self-attention mechanism, characterized in that: Includes the following steps, Step S1: Using an intelligent vehicle equipped with a surround-view multi-view camera, multi-line LiDAR, and millimeter-wave radar, multiple sets of environmental observation samples are collected in the environment to form a dataset. The environmental observation samples include the coordinates of the intelligent vehicle at the time of collection, the vehicle's odometer information, surround-view multi-view RGB images, LiDAR environmental point clouds, and multi-view millimeter-wave radar environmental point clouds. The surround-view multi-view RGB images include six RGB images: front view, left front view, left rear view, rear view, right rear view, and right front view. The multi-view millimeter-wave radar environmental point clouds include front view, left front view, left rear view, right rear view, and right front view. Step S2: Obtain multiple lidar depth images by spherical projection of the lidar environmental point cloud from S1; Step S3: Back-project the multi-view RGB image from S1 onto the lidar coordinate system to obtain an RGB point cloud that simultaneously contains RGB values ​​and three-dimensional coordinates; then stitch the RGB point clouds obtained from multiple views together and use spherical projection to obtain a panoramic RGB image from the lidar view. Step S4: Group the multiple LiDAR depth images obtained in S2 and the panoramic RGB images obtained in S3 to construct a three-dimensional data set; the three-dimensional data set contains multiple sample groups, each sample group contains a query sample, a positive sample and multiple negative samples; use the three-dimensional data set as the training dataset; Step S5: Use a multimodal descriptor extraction network to extract and aggregate features from the panoramic RGB image obtained in S3 and the LiDAR depth map obtained in S2 to obtain the LCGD descriptor; the multimodal descriptor extraction network is composed of a multimodal feature fusion module and a global feature aggregation module. The panoramic RGB image obtained in S3 and the LiDAR depth map obtained in S2 are used to extract features through the multimodal feature fusion module to obtain visual feature maps and LiDAR feature maps. Then, the visual feature maps and LiDAR feature maps are used to obtain visual sub-descriptors and LiDAR sub-descriptors through the global feature aggregation module. Finally, the visual sub-descriptors and LiDAR sub-descriptors are concatenated to obtain the LCGD descriptor. Step S6: Using Triplet Loss, the multimodal descriptor extraction network in S5 is trained using the training dataset obtained in Step S4, so that the descriptors of the query samples are close to the positive sample descriptors and far away from the negative sample descriptors; thus obtaining the trained multimodal descriptor extraction network. Step S7: Using the multimodal descriptor extraction network trained in step S6, process each sample in the data subset db in step S4 according to the method in step S5 to obtain the descriptor database. Step S8: In the intelligent vehicle localization stage, perform steps S2, S3, and S5 on the RGB images and LiDAR environmental point clouds collected at the current time to obtain the LCGD descriptor at the current time; compare the LCGD descriptor at the current time with all descriptors in the descriptor database of S7 according to the Euclidean distance, select the k with the smallest distance from the descriptor database as candidate descriptors, and use their corresponding samples as candidate samples; Step S9: Perform the following operations on the samples in the multi-view millimeter-wave radar environment point cloud acquired at the current moment: Project the millimeter-wave radar environment point clouds of all views in the current frame sample to the forward-looking millimeter-wave radar coordinate system to obtain the millimeter-wave radar environment point cloud in the forward-looking coordinate system; Using vehicle odometer information, project the millimeter-wave radar environment point clouds in the forward-looking coordinate system of the previous 3 frames of the current frame sample to the current sample coordinate system to obtain the dense millimeter-wave radar environment point cloud at the current moment; Step S10: Remove the dynamic points from the dense millimeter-wave radar environment point cloud obtained in S9 at the current moment to obtain the static point cloud; calculate the RCS probability density function of the static point cloud at the current moment. Step S11: Using the methods of steps S9 and S10, calculate the RCS probability density function of the point cloud of the k candidate samples obtained in S8. Step S12: Using the RCS probability density function of the current point cloud obtained in step S10 and the RCS probability density function of the point cloud of the k candidate samples obtained in step S11, calculate the RCS probability distribution distance; combine the RCS probability distribution distance with the Euclidean distance to obtain the total reordering distance; the Euclidean distance is the Euclidean distance between the current LCGD descriptor and the k candidate descriptors obtained in step S8; reorder the k candidate samples using the total reordering distance; Step S13: Using the re-sorting results from step S12, select the sample with the smallest re-sorting distance as the matching sample, and take the coordinates of this sample as the coordinates of the intelligent vehicle at the current moment, thereby completing a more accurate location identification.

2. The multimodal location recognition method based on a multi-scale self-attention mechanism as described in claim 1, characterized in that: The specific implementation method of step S3 is as follows: Step S31: Perform depth estimation on the RGB images from multiple viewing angles to obtain the absolute depth value of each pixel in each image at the real scale. Step S32: Using the absolute depth value obtained in step S31 and the intrinsic and extrinsic parameter matrices of each camera, project the RGB value of each pixel in the image from a certain viewpoint from the pixel coordinate system to the lidar coordinate system to obtain the RGB point cloud. Step S33: Perform projection processing on each view image using step S32, and stitch together all the obtained results to obtain an RGB point cloud containing a 360° view in the lidar coordinate system. Step S34: Perform spherical projection on the RGB point cloud containing a 360° view obtained in step S33 to obtain a panoramic RGB image from the perspective of the lidar.

3. The multimodal location recognition method based on a multi-scale self-attention mechanism as described in claim 2, characterized in that: The specific implementation method of step S4 is as follows: Step S41: In the dataset of step S1, based on the distance threshold dis_thres, select environmental observation samples at intervals according to the coordinates in the environmental observation samples. The selected environmental observation samples are called data subset db, and the unselected environmental observation samples are called data subset query, thus obtaining the query sample set. Step S42: In the data subset query of S41, select environmental observation samples with sample timestamps less than the threshold date_thres as the data subset train_query, and the remaining samples as the data subset val_query; Step S43: For each query sample in the data subset query, find the negative sample group corresponding to the query sample whose distance is greater than the threshold neg_thres in the db subset; find the positive sample group corresponding to the query sample whose distance is less than the threshold pos_thres in the db subset, and randomly select a positive sample from them as the positive sample corresponding to the query sample. The training set of triples is obtained by combining each query sample in the data subset train_query with its corresponding positive sample and multiple negative samples; the validation set of triples is obtained by combining each query sample in the data subset val_query with its corresponding positive sample and multiple negative samples; the training set of triples and the validation set of triples together constitute the training dataset.

4. The multimodal location recognition method based on a multi-scale self-attention mechanism as described in claim 3, characterized in that: The specific implementation method of step S5 is as follows: Step S51: Denote the panoramic RGB image obtained in S3 as having a size of B×C. I ×H I ×W I The tensor, B represents the size of the computation batch, H I and W I C represents the height and width of the RGB image, respectively. I This represents the number of RGB image channels; the LiDAR depth map obtained in S2 is denoted as having a size of B×C. L ×H L ×W L The tensor, H L and W L C represents the height and width of the LiDAR depth map, respectively. L This indicates the number of channels in the lidar depth map; The i-th layer features extracted from the panoramic RGB image and the LiDAR depth map are denoted as follows: and Will and The features from different modalities are input together into the multimodal feature fusion module for fusion; let the fused visual features and LiDAR features be respectively... and Both have the same shape as the input feature tensor; and and The modal fusion features are obtained by adding them separately. and and Step S52, Modal fusion feature I F and L F Each feature aggregation module performs its own feature aggregation to obtain visual sub-descriptors. and lidar sub-descriptor Step S53: Take the I obtained in step S52 D and L D By concatenating the data along the channel dimension and then performing L2 normalization, the LCGD descriptor is obtained.

5. The multimodal location recognition method based on a multi-scale self-attention mechanism as described in claim 4, characterized in that: Step S51 describes the following: and The specific method for fusing different modal features in the multimodal feature fusion module is as follows: Step S511, the input feature map of the multimodal feature fusion module is denoted as... and These correspond to visual features and LiDAR features, respectively. By using two vertically compressed VC layers respectively and Compressed to a size of and Horizontal characteristics; The VC layer consists of a 1×1 Conv layer, a Sigmoid activation layer, a 1×1 Conv layer, a MaxPooling layer, and four 1-D Conv layers. The MaxPooling layer is used for feature downsampling. Skip connections are used between the first and third 1-D Conv layers to accelerate convergence. Step S512: Concatenate the outputs of the VC layer in the width direction to obtain the multimodal feature embedding. Then, multi-head self-attention is used to perform multimodal feature fusion to obtain F'. e The size of the feature map does not change before and after this step; Step S513, F' e Available in two sizes and The tensor is then used to copy the feature map along the height dimension, resulting in the fused features output by the multimodal feature fusion engine. and 6. The multimodal location recognition method based on a multi-scale self-attention mechanism as described in claim 5, characterized in that: The process of transforming the spherical projection of the lidar point cloud into a depth map in step S2 is given by the following formula: in Let r = ||p||2 be the pixel coordinates of the depth map, and p = (xyz) be the ranging value of the laser point. T Let h and w represent the coordinates of the laser point, and h and w represent the width and height of the target depth map, respectively. f = f up +f down Indicates the field of view of the lidar; The triplet loss function in step S6 is calculated as follows: Where q and p q These represent the query sample and the corresponding positive sample, respectively. This represents the i-th negative sample corresponding to the query sample, [...] + Represents the hinge loss function [...] + =max(0,…), d(·) represents the Euclidean distance, m is a constant representing the distance threshold of the descriptor, and N neg It represents the number of negative samples.

7. The multimodal location recognition method based on a multi-scale self-attention mechanism as described in claim 6, characterized in that: The specific method for removing dynamic points and calculating the probability density function of the scattering cross section (RCS) in step S10 is as follows: Step S101: For all measurement points of the dense millimeter-wave radar in the current frame, calculate the radial velocity of each point relative to the millimeter-wave radar: in Let v represent the radial velocity at point i. i Let θ represent the velocity of the i-th point in the millimeter-wave radar coordinate system. i Let θ represent the azimuth angle of the i-th point. vi This represents the velocity measurement angle at the i-th point; Step S102: Since the radial velocity of most measurement points is 0 when the vehicle is stationary, if more than 60% of the measurement points in the millimeter-wave radar point cloud have a radial velocity of less than 1 m / s, it is determined that the vehicle in that frame is stationary. Otherwise, the vehicle in that frame is determined to be in motion. Step S103: If the vehicle in the current frame is stationary, the measurement points with a radial velocity greater than 1 m / s are determined to be dynamic points; if the vehicle in the current frame is moving, the stationary measurement points satisfy the following equation: in Let v represent the radial velocity at point i. s The speed of the millimeter-wave radar is α, the vehicle's heading angle is θ. i Let be the azimuth angle of the i-th point; the dynamically measured point is a discrete point for the above formula, therefore, based on the above formula, let α and v s As the solution parameters, the Random Sampling Consensus (RANSAC) algorithm is used to calculate outliers, and outliers are identified as dynamic measurement points and then filtered out. Step S104: Normalize the RCS value of the millimeter-wave radar measurement point to the range of (0,1) using the maximum and minimum values ​​of the RCS measurement value of this frame; Step S105, using bin RCS Given the interval size, calculate the RCS histogram and normalize it across all intervals to eliminate the influence of different point cloud quantities in different frames on the histogram, thus obtaining the RCS probability density function.

8. The multimodal location recognition method based on a multi-scale self-attention mechanism as described in claim 7, characterized in that: The specific steps for calculating the total reordering distance in step S12 are as follows: Step S121: Calculate the probability distribution distance between the query RCS probability density function and the candidate RCS probability density function based on KL divergence: Where H(R1) and H(R2) represent the RCS probability density functions obtained in step S10, and k is the index of the discrete interval of the probability density function; Step S122: Calculate the total reordering distance using the RCS probability distribution distance and the LCGD Euclidean distance of the multimodal descriptor: d rerank (q,c)=αd R (H(q),H(c))+(1-α)d E (LCGD(q),LCGD(c)) Where q and c represent the query sample and candidate sample, respectively, LCGD(·) represents the extraction of LCGD descriptors for the input sample according to step S5, and d E (·) represents Euclidean distance, and α is a set weight used to adjust the proportion of the two distance measurements; Step S123: Calculate the total reordering distance d between the query sample and each candidate sample. rerank and with d rerank Sort the k candidate samples in descending order.

Citation Information

Patent Citations

  • Algorithm detection method, computer equipment and readable storage medium

    CN116842352A

  • Cross-modal multi-task environment sensing method and system

    CN117237895A