Vehicle autonomous navigation site identification method based on multi-modal space and frequency domain fusion

By introducing frequency domain feature analysis and spatial domain and frequency domain feature fusion into vehicle autonomous navigation location identification method, the problem of insufficient vehicle location identification accuracy in GPS-free areas is solved, and high robustness and accurate positioning are achieved in complex environments.

CN122244621BActive Publication Date: 2026-07-21SHANDONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANDONG UNIV OF SCI & TECH
Filing Date
2026-04-30
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing vehicle location identification methods lack positioning accuracy in areas without GPS, and are prone to error accumulation, especially in complex environments. Furthermore, existing technologies fail to effectively utilize frequency domain features for data processing, resulting in insufficient robustness.

Method used

A vehicle autonomous navigation location identification method based on multimodal spatial and frequency domain fusion is adopted. By building an on-board multimodal perception model, frequency domain feature analysis is introduced, and combined with an effective fusion mechanism of spatial and frequency domain features, multimodal representations are extracted from image and LiDAR data, including a backbone network, a global feature aggregation module, and a global descriptor output module, to achieve deep interaction and feature extraction of image and LiDAR data.

Benefits of technology

It enables more reliable and accurate vehicle location identification in complex and dynamic real-world scenarios, improving the robustness and positioning accuracy of autonomous vehicle navigation and adapting to the high-reliability positioning requirements under complex road conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122244621B_ABST
    Figure CN122244621B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of vehicle positioning, and specifically discloses a vehicle autonomous navigation site identification method based on multi-modal space and frequency domain fusion. The method builds a vehicle-mounted multi-modal perception model based on the deep interaction of spatial and frequency domain features, which includes a backbone network, a global feature aggregation module and a global descriptor output module. The backbone network introduces frequency domain feature analysis, and on this basis, proposes an effective fusion mechanism of spatial and frequency domain features, which can extract more discriminative multi-modal representations from image and laser radar data, so as to realize more reliable and accurate vehicle position identification in complex and dynamic real scenes. The present application well meets the high-reliability positioning needs of automatic driving in complex road conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of vehicle positioning technology, and specifically relates to a method for vehicle autonomous navigation location identification based on multimodal spatial and frequency domain fusion. Background Technology

[0002] When vehicles enter areas without GPS (such as densely populated urban areas, forests, and mountainous regions), the accuracy of satellite positioning technology decreases significantly. Simultaneous localization and mapping (SLAM) becomes crucial for enabling autonomous navigation in GPS-free areas. Even with highly accurate pose estimation at each step in SLAM, small errors accumulate over time, causing the estimated vehicle and the constructed map to drift, ultimately resulting in a significant discrepancy with real-time data. Therefore, vehicle position recognition effectively mitigates this problem by continuously correcting the vehicle's trajectory based on its current location, reducing accumulated errors.

[0003] Existing location identification methods fall into three main categories: visual location identification, radar location identification, and multimodal location identification. Multimodal location identification, by fusing complementary features from different data sources (images and LiDAR), generally outperforms single-mode location identification technologies in most cases. However, current location identification technologies focus on feature processing in the explicit spatial domain, neglecting the implicit frequency domain and failing to analyze and extract existing data features in the frequency domain.

[0004] For image data, most existing methods extract and represent image features in the spatial domain. However, these features are highly sensitive to changes in illumination and scene appearance. When environmental conditions such as variations in illumination intensity, color shifts, shadow occlusion, seasonal changes, and day-night cycles occur, the spatial domain features of the image are prone to significant fluctuations, thus affecting the stability of scene recognition and long-term localization. In contrast, frequency domain features can mitigate the impact of illumination changes on feature representation to some extent and enhance the representation of the scene's inherent reflectivity and stable structural information. Therefore, they are more suitable for robust feature extraction in long-term autonomous localization scenarios.

[0005] Furthermore, the representation of images differs significantly between the spatial and frequency domains. Spatial domain features primarily reflect texture, edge, and appearance information within the local neighborhood of a pixel, while frequency domain features mainly reflect the energy distribution of the image at different frequencies and directions. For different objects or scene structures, their spectral distribution typically exhibits different peak characteristics and directional response properties. Compared to local features in the spatial domain, these frequency domain features generally demonstrate better stability under conditions such as viewpoint changes and local occlusion. On the other hand, on dynamic platforms or platforms with limited computing resources, images are also susceptible to motion blur and compression artifacts, leading to degradation of spatial domain structural information. In the frequency domain, motion blur manifests as attenuation of spectral energy and changes in specific frequency responses. By analyzing and processing frequency domain features, structural information in degraded images can be recovered to some extent, thereby improving the system's adaptability under complex degradation conditions.

[0006] For LiDAR data, existing methods primarily rely on spatial domain features to model point clouds. However, during the acquisition process, the point density of LiDAR point clouds fluctuates with changes in detection distance, observation angle, and environmental conditions. Furthermore, it is affected by measurement noise, local missing data, and uneven discrete sampling, thus limiting the stability of spatial domain features. In contrast, frequency domain analysis focuses more on the spectral energy distribution of the overall shape of the target or scene, exhibiting stronger tolerance to discrete sampling perturbations and local missing data. During frequency domain feature extraction, multi-scale geometric structure representation of the target or scene can be performed. High-frequency components mainly reflect local edges, detail changes, and surface undulations, while low-frequency components mainly reflect the overall contour and global structural information. Therefore, feature modeling of LiDAR data from a frequency domain perspective is beneficial for improving its structural representation capabilities and recognition stability in complex environments.

[0007] In summary, thanks to the globality, decoupling, compressibility, and analyzability of frequency features, this invention extracts features from multimodal location recognition technology from a frequency domain perspective, thereby discovering data features from the frequency domain. However, this does not mean that this invention abandons the spatial domain. Benefiting from the ability of the spatial domain to extract data features and the maturity of its technology, this invention effectively fuses features from the spatial and frequency domains to enable effective processing and extraction of feature data.

[0008] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art. Summary of the Invention

[0009] The purpose of this invention is to propose a vehicle autonomous navigation location identification method based on multimodal spatial and frequency domain fusion. In the vehicle-mounted multimodal perception model built in the method, by introducing frequency domain feature analysis and combining it with the proposed effective fusion mechanism based on spatial and frequency domain features, more discriminative multimodal representations can be extracted from image and LiDAR data, thereby achieving more reliable and accurate vehicle location identification in complex and dynamic real-world scenarios.

[0010] To achieve the above objectives, the present invention adopts the following technical solution: A location identification method for autonomous vehicle navigation based on multimodal spatial and frequency domain fusion includes the following steps: Step 1. Acquire multimodal raw data collected by the vehicle at the same time, including surround view image data and LiDAR point cloud data; preprocess the multimodal raw data and construct a training dataset; Step 2. Construct an in-vehicle multimodal perception model based on deep interaction of spatial and frequency domain features; this in-vehicle multimodal perception model includes a backbone network, a global feature aggregation module, and a global descriptor output module; The backbone network consists of multiple feature extraction layers connected in series, and each feature extraction layer is equipped with a dual-branch feature extraction unit and a fusion unit; the fusion unit includes an intra-modal multi-frequency fusion unit and a cross-modal space and frequency domain interaction unit; The dual-branch feature extraction unit simultaneously receives image and lidar modal data and performs spatial domain feature extraction separately. The intramodal multi-frequency fusion unit is used to perform frequency domain decomposition on the image modal and lidar modal features extracted from the spatial domain features, and to fuse and enhance the low-frequency and high-frequency features within each modality. The cross-modal space and frequency domain interaction unit is used to realize feature alignment, information exchange and joint representation generation between image modalities and lidar modalities, and obtain the multimodal features after interaction; The global feature aggregation module encodes the multimodal features after interaction into global representations of the corresponding modalities; the global descriptor output module connects the aggregated global representations of each modality to form a global descriptor for location recognition. Step 3. Train the vehicle-mounted multimodal perception model built in Step 2 based on the training dataset in Step 1, and use the trained model to obtain the location recognition result corresponding to the current position of the vehicle.

[0011] The present invention has the following advantages: As described above, this method constructs an in-vehicle multimodal perception model, MSFF-Net, based on deep interaction of spatial and frequency domain features. This model includes a backbone network, a global feature aggregation module, and a global descriptor output module. The backbone network incorporates frequency domain feature analysis, and based on this, innovatively proposes an effective fusion mechanism for spatial and frequency domain features. This scheme can extract more discriminative multimodal representations from image and LiDAR data, thereby achieving more reliable and accurate vehicle location recognition in complex and dynamic real-world scenarios. This invention utilizes frequency domain structural information to enhance the robustness of vehicle multimodal location recognition, effectively meeting the high-reliability positioning requirements of autonomous driving in complex road conditions. Attached Figure Description

[0012] Figure 1 This is a structural diagram of the vehicle-mounted multimodal perception model based on deep interaction of spatial and frequency domain features in an embodiment of the present invention; Figure 2 This is a network structure diagram of the fusion unit in an embodiment of the present invention; Figure 3 This is a network structure diagram of the image adaptive wavelet transform unit in an embodiment of the present invention; Figure 4 This is a network structure diagram of the update / prediction unit in the image adaptive wavelet transform unit of this invention embodiment; Figure 5 This is a network structure diagram of the adaptive wavelet transform unit of the lidar in an embodiment of the present invention; Figure 6 This is a network structure diagram of the update / prediction unit in the adaptive wavelet transform unit of the lidar according to an embodiment of the present invention; Figure 7 This is a network structure diagram of the image mode / LiDAR mode multi-frequency fusion unit in an embodiment of the present invention. Detailed Implementation

[0013] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments: Example 1 To address the problem of sensitivity to changes in illumination, viewing angle differences, and sensor noise caused by relying solely on spatial domain features in existing vehicle positioning technologies, this invention describes a vehicle autonomous navigation location identification method based on multimodal spatial and frequency domain fusion. This method includes the following steps: Step 1. Acquire multimodal raw data collected by the vehicle at the same time, including surround view image data and lidar point cloud data; preprocess the multimodal raw data and construct a training dataset.

[0014] In this embodiment, the surround view image data is collected by multiple cameras mounted on the vehicle.

[0015] Specifically, for image data acquisition, this embodiment abandons the conventional single-view perception scheme and instead adopts a 360° panoramic perception array composed of 6 cameras distributed horizontally. Through multi-dimensional perspective collaboration, the redundancy and semantic richness of environmental information are greatly improved, effectively eliminating the visual blind spots existing in the traditional single-view perception scheme.

[0016] In the data preprocessing stage, the present invention stitches together images captured simultaneously by multiple cameras at the same time to obtain a complete image, providing global continuous surround view environmental features for subsequent steps.

[0017] For LiDAR data input in multimodal scenarios, given the massive data volume, disordered and dense spatial distribution of the original 3D point cloud, direct feature extraction in 3D space would result in extremely high computational costs and latency. To optimize data processing efficiency and meet the real-time requirements of onboard computing platforms, this invention employs projection transformation to perform dimensionality reduction preprocessing on the original 3D point cloud, mapping the input 3D LiDAR data into a 2D range view (RV). It should be noted that generating the RV image through projection transformation is a fairly standard method and will not be elaborated upon here.

[0018] This invention uses a subset of the Boston Harbor dataset from the nuScenes dataset as the source of multimodal training data. The training input sample data consists of six camera images and LiDAR point cloud data collected by the vehicle platform. The training output sample data is the global descriptor corresponding to the vehicle's position at the time of collection. During data preparation, collected data within a preset time range is first selected as training query samples; then, historical collected data is sampled at intervals based on the spatial location of the vehicle trajectory to construct a database sample; simultaneously, validation and test samples are divided from the remaining collected data. This yields a query set, database set, validation set, and test set suitable for location recognition tasks.

[0019] Step 2. Construct an in-vehicle multimodal perception model based on deep interaction of spatial and frequency domain features, such as... Figure 1 As shown, the vehicle-mounted multimodal perception model includes a backbone network, a global feature aggregation module, and a global descriptor output module.

[0020] The backbone network preferably adopts a multi-stage residual network as the backbone feature extraction structure. In one preferred approach, the present invention uses the ResNet-18 architecture as the feature extraction backbone and utilizes the residual module to alleviate the gradient vanishing problem in deep network training.

[0021] An adaptive wavelet transform module is embedded in each of the four residual stages (Stages 1-4) of ResNet-18 to enable multi-level frequency domain analysis of multimodal information. The output feature map size of each stage is downsampled proportionally with the increase of the layer depth.

[0022] In this embodiment, the backbone network includes multiple feature extraction layers connected in series (for example, the number of stages N is set to 4, i.e., 4 feature extraction layers), and each feature extraction layer is provided with a dual-branch feature extraction unit and a fusion unit.

[0023] The dual-branch feature extraction unit includes a lidar feature extraction unit and an image feature extraction unit, which are used to simultaneously receive image and lidar modal data, and extract spatial domain features from the data of the two modalities respectively.

[0024] The image feature extraction unit is used to receive image modal data and extract spatial domain features from it; the lidar feature extraction unit is used to receive lidar modal data and extract spatial domain features from it.

[0025] The fusion unit includes an intramodal multi-frequency fusion unit and a cross-modal space and frequency domain interaction unit.

[0026] The intramodal multi-frequency fusion unit is used to perform frequency domain decomposition on the image modal and lidar modal features extracted from the spatial domain, and to fuse and enhance the low-frequency and high-frequency features within each modality.

[0027] The cross-modal space and frequency domain interaction unit is used to realize feature alignment, information exchange and joint representation generation between image modalities and lidar modalities, and obtain multimodal features after interaction, including image modalities and lidar modal features.

[0028] The multimodal features output by the cross-modal space and frequency domain interaction unit include image modal features and lidar modal features, and are respectively residually connected with the feature extraction branch results of the corresponding modality on the feature extraction layer.

[0029] The image modal features and lidar modal features obtained after residual concatenation are used as the multimodal feature inputs for the next feature extraction layer; the inputs for the first feature extraction layer are the original stitched overall image and the two-dimensional distance view.

[0030] The global feature aggregation module encodes the multimodal features after interaction into global representations of the corresponding modalities; the aggregated global representations of each modality are then connected to form a global descriptor for location recognition.

[0031] To address the issue of insufficient robustness of existing location recognition MPR methods in dynamic and complex environments, this invention innovatively proposes an MSFF-Net model. By constructing a deep interactive logic between spatial domain scene features and frequency domain adaptive wavelet transform, it achieves nonlinear fusion of multimodal data in multi-scale frequency space. This approach effectively overcomes the limitations of traditional single spatial domain coding and theoretically greatly approximates and improves the upper limit of multimodal fusion gain.

[0032] like Figure 2 As shown, in this embodiment, the intramodal multi-frequency fusion unit includes an image adaptive wavelet transform unit, a lidar adaptive wavelet transform unit, an image modal multi-frequency fusion unit, and a lidar modal multi-frequency fusion unit.

[0033] The image adaptive wavelet transform unit is used to perform frequency domain decomposition on image modal features, dividing the input image features into low-frequency features and high-frequency features. The low-frequency features obtained from the frequency domain decomposition of image modal features are used to characterize the overall outline, main structure, and stable semantic information of the scene; the high-frequency features are used to characterize the edge, texture, and detail changes of the scene in different directions.

[0034] like Figure 3 The network structure of the image adaptive wavelet transform unit is shown.

[0035] In frequency feature encoding, image modal data is input into the image adaptive wavelet transform unit, which uses two-dimensional discrete wavelet transform (2D DWT) as the frequency domain decomposition tool. By performing high-pass and low-pass filtering along the horizontal and vertical directions, the stitched panoramic image is decomposed into four frequency sub-bands: the low-frequency sub-band (LL) retains the main contour information, while the high-frequency sub-bands (LH, HL, HH) capture vertical, horizontal, and diagonal texture details, respectively. For low-frequency branch optimization, considering the susceptibility of low-frequency components in the panoramic stitched image to environmental noise, a channel attention mechanism is introduced at the end of the low-frequency feature extraction branch. This adaptively adjusts the response intensity of each channel to suppress irrelevant noise and enhance key structural information. For high-frequency branch optimization, to capture complex high-frequency fine-grained features from the panoramic viewpoint, a multi-head attention mechanism is introduced at the end of the high-frequency branch. This constructs a parallel attention subspace, enabling multi-scale modeling of image details.

[0036] like Figure 3 As shown, the image adaptive wavelet transform unit in this embodiment includes a two-dimensional discrete wavelet transform segmentation unit, a low-frequency branch, a high-frequency branch, an update unit, a prediction unit, and an attention enhancement unit.

[0037] A two-dimensional discrete wavelet transform segmentation unit is used to perform frequency domain decomposition on the input image features, dividing them into a low-frequency subband LL and high-frequency subbands LH, HL, and HH. The low-frequency subband LL mainly represents the overall contour, main structure, and stable semantic information of the scene; the high-frequency subbands LH, HL, and HH mainly represent the edge, texture, and detail changes of the scene in different directions. By dividing the input features into low-frequency and high-frequency components, this invention can decouple the globally stable information and locally sensitive information in the image, providing a foundation for subsequent branch enhancement and fusion processing.

[0038] The low-frequency subband LL output by the two-dimensional discrete wavelet transform segmentation unit enters the low-frequency branch. The high-frequency subbands LH, HL, and HH output by the two-dimensional discrete wavelet transform segmentation unit enter the high-frequency branch after convergence processing.

[0039] Specifically, the high-frequency subbands LH, HL, and HH are first averaged to obtain a unified high-frequency representation, which reduces the influence of local fluctuations between high-frequency responses in different directions and improves the stability and compactness of high-frequency features in subsequent processing.

[0040] Thus, the low-frequency branch preserves global contour features, while the high-frequency branch preserves local detail features, forming a complementary dual-branch frequency domain representation structure. Furthermore, to enhance information synergy between the low-frequency and high-frequency branches, the image adaptive wavelet transform unit in this embodiment also includes an update unit and a prediction unit.

[0041] The update unit supplements and corrects the low-frequency branches using detailed information from the high-frequency branches, thereby enhancing the low-frequency branches' ability to perceive local changes and edge structures. The prediction unit uses the overall structural information from the low-frequency branches to constrain and guide the high-frequency branches, thereby improving the high-frequency branches' responsiveness to effective texture details and suppressing noise interference.

[0042] Through the bidirectional interaction between the update unit and the prediction unit, an information feedback mechanism can be established between low-frequency features and high-frequency features, thereby achieving adaptive reconstruction and optimization of image frequency domain features.

[0043] The update unit and the prediction unit are implemented using the same or similar network structures. For example... Figure 4 The network structure diagram of the update / prediction unit in the image adaptive wavelet transform unit is shown. Both the update unit and the prediction unit include a reflection filling subunit, a convolutional feature extraction subunit, an activation function subunit, a channel mapping subunit, and an output constraint subunit.

[0044] The input features are first expanded by a reflection-filling subunit to mitigate feature truncation and information loss caused by convolution operations in the boundary region. Then, local neighborhood features are extracted by a 3×3 convolutional feature extraction subunit. Next, non-linear enhancement is performed by an activation function subunit. Following this, cross-channel feature mapping and compression are performed by a 1×1 convolutional channel mapping subunit. Finally, the output is constrained by the Tanh function in the output constraint subunit to generate guiding features for updating or prediction. The reflection-filling subunit is a relatively conventional boundary processing structure used to improve the local feature modeling effect of the update and prediction units in the boundary region.

[0045] Through the above structural design, the update unit and prediction unit in the image adaptive wavelet transform unit can not only retain the spatial correlation information within the local receptive field, but also achieve adaptive adjustment of the features of different channels.

[0046] After the update and prediction are completed, the low-frequency branch and the high-frequency branch enter the corresponding attention enhancement unit respectively; the attention enhancement unit includes the channel attention mechanism unit and the multi-head attention mechanism unit.

[0047] The low-frequency branch access channel attention mechanism unit assigns weights based on the importance of different channels to the overall scene structure representation, thereby enhancing low-frequency semantic information relevant to the location recognition task.

[0048] High-frequency branches are connected to multi-head attention mechanism units to perform parallel modeling of local texture details in different directions and scales, thereby improving the ability of high-frequency features to express edge structures, fine-grained textures and local discriminative cues.

[0049] This invention improves the representation quality of dual-branch frequency domain features by introducing different forms of attention mechanisms in the low-frequency and high-frequency branches, thereby achieving differentiated enhancement for different frequency band characteristics.

[0050] This invention is based on the image adaptive wavelet transform (IAWT) and multi-frequency fusion (MFF) mechanism of two-dimensional discrete wavelet transform. By embedding two-dimensional discrete wavelet transform operators in the feature encoding process, multi-resolution frequency domain decomposition is performed on the image to be processed, thereby achieving adaptive capture of low-frequency information covering the global contour and multi-scale high-frequency features covering edge texture.

[0051] Building upon this foundation, an innovative multi-frequency fusion MFF mechanism is introduced, which incorporates a frequency-domain-guided dynamic weight adjustment mechanism. Low-frequency semantic features are used as the query benchmark, and high-frequency details are weighted, reconstructed, and semantically aligned through cross-attention logic. This achieves deep intramodal fusion while effectively suppressing spectral aliasing and scale feature interference between heterogeneous domain data, significantly enhancing the model's feature representation stability and discrimination robustness in complex dynamic environments.

[0052] Specifically, such as Figure 5 As shown, the adaptive wavelet transform unit of the LiDAR includes a vertical segmentation unit, an update unit, a prediction unit, and an attention enhancement unit. The vertical segmentation unit decomposes the input LiDAR range view features along the vertical direction to obtain interrelated low-frequency and high-frequency branch features.

[0053] The input features are divided into two sub-features based on their sampling positions along the vertical dimension, which are used to construct the low-frequency branch and the high-frequency branch, respectively. The low-frequency branch features mainly represent the relatively stable overall geometric contours and structural distribution information in the scene, while the high-frequency branch features mainly represent local geometric changes, edge abrupt changes, and detailed undulations.

[0054] This invention decouples structural information at different scales in the lidar range view by performing segmentation operations along the vertical direction, thus providing a foundation for subsequent updates and predictions. The reason for performing frequency domain feature processing in the vertical direction of the lidar range view is based on a comprehensive consideration of the multimodal information distribution characteristics and the imaging features of the vehicle-mounted sensor.

[0055] On the one hand, in this invention, the image modality, stitched together horizontally from multiple cameras, already provides relatively sufficient horizontal field-of-view coverage and lateral scene semantic information. On the other hand, due to the influence of the vehicle-mounted camera's pitch field of view, the distribution of LiDAR scan lines, and the distance-view imaging mechanism, the structural information in the vertical direction is relatively more limited and more prone to exhibiting characteristics such as sparsity, discreteness, and discontinuity. Therefore, this invention focuses on the vertical direction in the frequency domain modeling of the LiDAR modality to enhance the ability to express height changes, facade structures, obstacle outlines, and geometric hierarchical relationships. This complements the relatively sufficient horizontal semantic information already acquired in the image modality, improving the completeness and robustness of the multimodal joint representation.

[0056] The update unit and the prediction unit are used to establish a two-way information exchange mechanism between low-frequency branches and high-frequency branches.

[0057] The update unit supplements and corrects the overall contour information in the low-frequency branch using local geometric details from the high-frequency branch, thereby enhancing the low-frequency branch's ability to perceive local structural changes. The prediction unit guides and constrains the detailed response in the high-frequency branch using the overall structural information from the low-frequency branch, reducing the high-frequency branch's sensitivity to noise and random disturbances. Through the cooperation of the update and prediction units, this invention enables adaptive feedback between the low-frequency and high-frequency branches, improving the discriminative power and stability of the lidar's frequency domain features.

[0058] like Figure 6 As shown, similar to the update unit and prediction unit in the above image adaptive wavelet transform unit, the update unit and prediction unit of the lidar adaptive wavelet transform unit can adopt the same or similar network structure.

[0059] In this embodiment, the update unit / prediction unit of the adaptive wavelet transform unit of the lidar can include a reflection filling subunit, a convolution subunit, an activation function subunit, a channel mapping subunit, and an output constraint subunit.

[0060] Specifically, the input features are first filled with reflection to reduce feature loss caused by boundary information truncation; then, local neighborhood structural features are extracted along the vertical direction through a convolutional layer with a kernel size of 3×1; then, the nonlinear expressive power is enhanced by an activation function; then, cross-channel mapping and information compression are completed through a convolutional layer with a kernel size of 1×1; finally, the output results are constrained by the Tanh function to generate guiding features for updating or prediction.

[0061] Because the convolutional kernel size has a directional receptive field in the vertical dimension, it can more effectively adapt to the frequency domain modeling requirements of this invention in the vertical direction. After updating and predicting, the low-frequency branch and high-frequency branch features are respectively input into the corresponding attention enhancement units. The attention enhancement units are used to weight and enhance the updated and predicted low-frequency branch features and high-frequency branch features respectively, so as to output the frequency domain enhancement features of the LiDAR mode. Through the above structure, frequency domain modeling and adaptive enhancement of vertical structural information can be achieved within the LiDAR mode.

[0062] Specifically, the attention enhancement unit adopts a channel attention mechanism and is used to weight and enhance the updated and predicted low-frequency branch features and high-frequency branch features respectively, so as to output the frequency domain enhancement features of the lidar mode.

[0063] The attention enhancement unit assigns different weights to the vertical structural features represented by different channels, thereby highlighting key geometric information relevant to the location recognition task and suppressing irrelevant responses caused by sparse sampling, measurement noise or dynamic interference.

[0064] The attention-enhanced low-frequency branch features and high-frequency branch features are used as frequency domain enhancement features of the lidar mode, and are further input into the subsequent multi-frequency fusion unit of the corresponding mode for joint modeling.

[0065] The image modality multi-frequency fusion unit and the lidar modality multi-frequency fusion unit are used to jointly model the low-frequency and high-frequency features extracted within the same modality (image modality and lidar modality), respectively. Through information filtering, weight allocation and feature reconstruction, fusion features with both global stability and local discriminativeness are obtained.

[0066] like Figure 7 As shown, the image modal multi-frequency fusion unit / LiDAR modal multi-frequency fusion unit includes a high-frequency feature input terminal, a low-frequency feature input terminal, a linear mapping unit, a correlation calculation unit, an activation unit, a weighted fusion unit, and an output unit.

[0067] The high-frequency feature input terminal is used to input high-frequency features obtained from frequency domain decomposition (e.g., Figure 2 High-frequency features shown , The low-frequency feature input terminal is used to input low-frequency features (e.g., features) obtained from frequency domain decomposition. , ).

[0068] in , These are the high-frequency and low-frequency features obtained from image modal data decomposition, respectively. , These are the high-frequency and low-frequency features obtained from the decomposition of lidar mode data, respectively.

[0069] The linear mapping unit is used to perform mapping transformations on high-frequency features and low-frequency features respectively to generate query vector Q, key vector K, and value vector V for subsequent interactive computation.

[0070] The relevance calculation unit is used to calculate feature similarity based on the association between query vector Q and key vector K; the activation unit is used to normalize or non-linearly enhance the similarity to obtain attention weights.

[0071] The weighted fusion unit is used to reconstruct the value vector V based on the attention weights and fuse it with low-frequency features; the output unit is used to output the fused multi-frequency feature representation.

[0072] This invention enables adaptive interaction and joint enhancement between low-frequency and high-frequency information within the same modality.

[0073] Each modal multi-frequency fusion unit uses low-frequency features as semantic guidance information and high-frequency features as detailed supplementary information.

[0074] In a preferred embodiment, the input high-frequency features are transformed into value vector V and key vector K by the first linear mapping unit and the second linear mapping unit, respectively, and the input low-frequency features are transformed into query vector Q by the third linear mapping unit.

[0075] The query vector Q is used to characterize the overall structural semantics in low-frequency features, while the key vector K and the value vector V are used to characterize the local detail information in high-frequency features.

[0076] By using the above settings, the relatively stable global information in the low-frequency features can be used to selectively guide the local details in the high-frequency features, thereby avoiding the direct accumulation of irrelevant disturbances and noise in the high-frequency features.

[0077] The relevance calculation unit calculates the relevance matrix between the query vector Q and the key vector K. The relevance matrix is ​​obtained through inner product operations and then input into the activation unit for processing to generate the attention weight matrix.

[0078] The activation unit preferably uses the Softmax function, normalization function or other nonlinear mapping method to highlight high-frequency detail regions related to low-frequency semantic information and weaken the response of irrelevant regions.

[0079] This invention, through the correlation calculation and activation process, can establish an explicit dependency between low-frequency features and high-frequency features, enabling the reconstruction process of high-frequency features to have semantic selectivity.

[0080] The weighted fusion unit performs matrix multiplication on the value vector V based on the attention weight matrix to generate high-frequency enhanced features guided by low-frequency semantics. Subsequently, the high-frequency enhanced features are fused with the low-frequency features to obtain multi-frequency fused features that combine globally stable structural information and local discriminative details.

[0081] In a preferred embodiment, the fusion process is implemented by element-wise summation, that is, the weighted reconstructed high-frequency enhanced features are added to the original low-frequency features, and then the input is used to perform nonlinear transformation of the activation function to output the final multi-frequency fused features. Figure 7 The circle symbol in the middle represents the element-wise summation operation, and ReLU is used to enhance the nonlinearity of the fusion result.

[0082] Through the above multi-frequency fusion process, the multi-frequency fusion unit of the corresponding modality can use low-frequency features to perform semantic constraints and detail filtering on high-frequency features, thereby suppressing the adverse effects of spectral aliasing, local noise and invalid texture on feature expression.

[0083] Compared to simply concatenating or linearly superimposing low-frequency features with high-frequency features, the image modality / LiDAR modality multi-frequency fusion unit described in this invention introduces a correlation interaction mechanism between the query vector Q, key vector K, and value vector V, which enables low-frequency structural information to dynamically guide high-frequency detail information. This allows the output features to simultaneously possess global consistency and local discriminability, thereby improving the feature quality in subsequent cross-modal interaction and global aggregation processes.

[0084] This invention is based on the LiDAR adaptive wavelet transform (LAWT) and multi-frequency fusion (MFF) mechanism. By performing split, update, and predict sequence operations along the vertical dimension of the point cloud projection image, the original point cloud signal is decomposed into odd and even sub-signals with independent time-frequency characteristics, thereby achieving high-precision representation of multi-scale vertical structural features in the point cloud. Furthermore, by combining the multi-frequency fusion (MFF) mechanism, a dynamic adaptive weight adjustment mechanism based on vertical distribution is introduced. By performing deep fusion and semantic completion on the approximation coefficients and detail coefficients within the mode, the intra-modal noise interference and geometric deformation deviation generated by the LiDAR point cloud during dynamic scanning are effectively suppressed, significantly improving the representation consistency of three-dimensional spatial structural features under complex vehicle navigation conditions.

[0085] The global feature aggregation module is used to globally encode the features output by the cross-modal space and frequency domain dual-domain interaction unit.

[0086] Specifically, in this embodiment, there are two global feature aggregation modules, and each global feature aggregation module performs global aggregation on image modal features and lidar modal features respectively to obtain the global representation of the corresponding modality.

[0087] Subsequently, the global representations of each modality are concatenated to form a unified global descriptor. This global descriptor can be used for similarity matching with historical location features in the database to output the final location identification result.

[0088] By adopting the above-mentioned modal aggregation and final joint output method, it is possible to improve the overall descriptor expression completeness and retrieval accuracy while retaining the discrimination characteristics of each modality (image modality and lidar modality).

[0089] This invention achieves a hierarchical expansion from the overall structure to internal functional units through a layered processing flow of "spatial domain feature encoding - intra-modal multi-frequency fusion - cross-modal spatial and frequency domain dual-domain interaction - global feature aggregation - global descriptor output". This network can not only perform phased, multi-level feature extraction for image and LiDAR modalities, but also achieve fusion enhancement of low-frequency and high-frequency features within a modality, and realize joint interaction between spatial and frequency domains between modalities, ultimately forming a highly robust global descriptor suitable for location recognition tasks in autonomous vehicle navigation scenarios.

[0090] Step 3. Train the vehicle-mounted multimodal perception model built in Step 2 based on the training dataset in Step 1, and use the trained model to obtain the location recognition result corresponding to the current position of the vehicle.

[0091] During model training, an adaptive wavelet transform module is first used to extract intra-modal multi-scale frequency features from the raw images and point cloud data (which need to be converted into a 2D distance view) acquired by vehicle-mounted sensors (such as cameras and LiDAR). Then, under a multi-domain guided fusion strategy, a decoupling mechanism is used to separate domain-variable and domain-invariant information to eliminate perception aliasing in dynamic traffic scenarios such as urban roads. Finally, a Transformer architecture is used to achieve cross-modal deep fusion of the spatial and frequency domains, and end-to-end optimization is performed based on a loss function (a relatively conventional one, such as a triplet loss function). During inference, the optimized spatial and frequency domain fusion weights are retained. When facing complex conditions such as dynamic occlusion (such as pedestrians and other vehicles) and viewpoint changes common in vehicle operation, the translation and rotation invariance of frequency domain features is used to assist spatial domain features in accurate discrimination. The method of this invention, while adapting to the computing power of the vehicle computing platform and ensuring inference efficiency, significantly improves the recall rate and F1max index of closed-loop detection, providing a highly robust global spatial cognition for simultaneous localization and mapping (SLAM) of autonomous vehicles, and effectively guiding the vehicle's global path planning and safe autonomous navigation.

[0092] In addition, to verify the effectiveness of the MSFF-Net in-vehicle multimodal perception model with deep interaction between spatial and frequency domain features proposed in this invention, the following experiments were conducted. Five comparison methods were introduced in the comparative experiments: NetVLAD method, Patch-NetVLAD method, PointNetVLAD method, Overlap Transformer method, and LCPR method.

[0093] The NetVLAD method comes from: R. Arandjelovic, P. Gronat, A. Torii, T. Pajdla andJ. Sivic, "NetVLAD: CNN Architecture for Weakly Supervised PlaceRecognition," 2016 IEEE Conference on Computer Vision and Pattern Recognition(CVPR), Las Vegas, NV, USA, 2016, pp. 5297-5307, doi: 10.1109 / CVPR.2016.572.

[0094] The Patch-NetVLAD method comes from: S. Hausler, S. Garg, M. Xu, M. Milford and T.Fischer, "Patch-NetVLAD: Multi-Scale Fusion of Locally-Global Descriptors for Place Recognition," 2021 IEEE / CVF Conference on Computer Vision and PatternRecognition (CVPR), Nashville, TN, USA, 2021, pp. 14136-14147, doi: 10.1109 / CVPR46437.2021.01392.

[0095] The PointNetVLAD method originates from: MA Uy and GH Lee, “PointNetVLAD: Deeppoint cloud based retrieval for large-scale PR,” in Proc. Conf. Comput. Vis.Pattern Recognit., 2018, pp. 4470–4479.

[0096] The Overlap Transformer method comes from: J. Ma, J. Zhang, J. Xu, R. Ai, W. Gu, and X. Chen, "OverlapTransformer: An efficient and yaw-angle-invarianttransformer network for LiDAR-based place recognition," IEEE Robot. Automat.Lett., vol. 7, no. 3, pp. 6958–6965, Jul. 2022.

[0097] The LCPR method comes from: Zijie Zhou, Jingyi Xu, Guangming Xiong, and JunyiMa.Lcpr: A multi-scale attention-based lidar-camera fusion network for placerecognition. IEEE Robotics and Automation Letters, 9(2):1342–1349, 2024.

[0098] This embodiment uses the authoritative NuScenes public dataset in the field of vehicle navigation as the basic verification environment. To evaluate the robustness of the proposed network under real-world autonomous driving conditions, this embodiment conducts training and generalization capability tests on the Boston Seaport (BS) subset and the SG-OneNorth (SON) subset, respectively. Addressing the issues of uneven sample distribution and multi-sensor cross-domain deployment in complex urban road environments, a standardized data organization logic is strictly followed. The BS subset (total of 22,103 samples) and the SON subset (total of 8,104 samples) are finely segmented to construct a large-scale evaluation array covering 8,407 training query sets, 1,002 validation query sets, and a total of 6,316 test query sets. By performing in-depth optimization of model parameters on the BS subset and then transferring the model across domains to the SON subset for validation, the model's generalization ability and environmental adaptability to different geographical regions, urban layouts, and dynamic interference conditions are effectively improved by utilizing a large-scale sample database (Ndatabase of 9,686 and 4,796 respectively) while maintaining the consistency of the original spatial-frequency domain features of the sensor.

[0099] To effectively and intuitively demonstrate the network's advantages, a range of widely accepted metrics are used to evaluate location recognition PR performance, including precision-recall curves (PR curves), Recall@1, Recall@5, and F1max.

[0100] , , .

[0101] Where TP represents the number of samples correctly predicted as positive, FP represents the number of samples incorrectly predicted as positive, and FN represents the number of positive samples incorrectly predicted as negative.

[0102] Table 1 Comparison results of different methods

[0103] As can be seen from Table 1 above, the method of this invention performs excellently in all evaluation metrics on both the BS and SON datasets, achieving the best results. This demonstrates that the method of this invention achieves the best accuracy in robust location recognition tasks, significantly outperforming all other methods compared, thus verifying the effectiveness of the method. This invention significantly improves the accuracy of loop closure detection in complex environments. Experimental data shows that the Recall@1 metric of this technical solution reaches 98.44% on the large authoritative NuScenes dataset. This technical effect confirms that in actual vehicle autonomous navigation, this invention can provide SLAM systems with near-zero error global location accuracy. Figure 1 Consistency assurance effectively reduces the risk of positioning drift.

[0104] This invention introduces a spatial-frequency dual-domain deep interaction architecture inspired by human visual cognition. It utilizes adaptive wavelet transform to effectively decouple and extract domain-invariant information (DII) with cross-domain stability, significantly improving the model's robustness in loop closure detection under complex conditions. Simultaneously, by performing distance view (RV) dimensionality reduction projection on the original 3D point cloud, combined with a lightweight backbone network and a multi-frequency fusion (MFF) mechanism, it ensures omnidirectional coverage of global information while significantly reducing the memory footprint and computational latency of the onboard computing platform. The resulting global descriptor greatly optimizes the retrieval efficiency of large-scene map databases. Furthermore, the spatial and frequency dual-domain alignment strategy proposed in this invention effectively offsets perceptual aliasing and regional feature differences between heterogeneous modalities, enhancing the model's generalization and transfer capabilities across different geographical regions and sensor configurations. This provides crucial technical support for the long-term positioning consistency and reliable path planning of autonomous vehicles in complex urban environments.

[0105] Example 2 This embodiment 2 describes a computer device, which includes a memory and one or more processors. An in-vehicle camera and a LiDAR are both connected to the computer device and are used to send acquired image modal data and LiDAR modal data to the computer device. Executable code is stored in the memory. When the processor executes the executable code, it implements the steps of the vehicle autonomous navigation location recognition method based on multimodal spatial and frequency domain fusion described in embodiment 1 above. In this embodiment, the computer device can be any device or apparatus with data processing capabilities, and will not be described in detail here.

[0106] Example 3 This embodiment 3 describes a computer-readable storage medium storing a program that, when executed by a processor, is used to implement the steps of the vehicle autonomous navigation location identification method based on multimodal spatial and frequency domain fusion in embodiment 1 above.

[0107] The computer-readable storage medium can be an internal storage unit of any device or apparatus with data processing capabilities, such as a hard disk or memory, or an external storage device of any device with data processing capabilities, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc.

[0108] Of course, the above description is only a preferred embodiment of the present invention. The present invention is not limited to the above-described embodiments. It should be noted that any equivalent substitutions or obvious modifications made by those skilled in the art under the guidance of this specification fall within the scope of this specification and should be protected by the present invention.

Claims

1. A method for location identification in autonomous vehicle navigation based on multimodal spatial and frequency domain fusion, characterized in that, Includes the following steps: Step 1. Acquire multimodal raw data collected by the vehicle at the same time, including surround view image data and LiDAR point cloud data; preprocess the multimodal raw data and construct a training dataset; Step 2. Construct an in-vehicle multimodal perception model based on deep interaction of spatial and frequency domain features, which includes a backbone network, a global feature aggregation module, and a global descriptor output module; The backbone network consists of multiple feature extraction layers connected in series, and each feature extraction layer is equipped with a dual-branch feature extraction unit and a fusion unit; the fusion unit includes an intra-modal multi-frequency fusion unit and a cross-modal space and frequency domain interaction unit; The dual-branch feature extraction unit simultaneously receives image and lidar modal data and performs spatial domain feature extraction separately. The intramodal multi-frequency fusion unit is used to perform frequency domain decomposition on the image modal and lidar modal features extracted from the spatial domain features, and to fuse and enhance the low-frequency and high-frequency features within each modality. The cross-modal space and frequency domain interaction unit is used to realize feature alignment, information exchange and joint representation generation between image modalities and lidar modalities, and obtain the multimodal features after interaction; The global feature aggregation module encodes the multimodal features after interaction into global representations of the corresponding modalities; the global descriptor output module connects the aggregated global representations of each modality to form a global descriptor for location recognition. The cross-modal space and frequency domain interaction unit is set after the intramodal multi-frequency fusion unit. The cross-modal space and frequency domain interaction unit includes a feature connection unit, a Transformer interaction unit, and a joint representation output unit. The feature connection unit is used to connect or combine the fusion features of the image mode output from the intramodal multi-frequency fusion unit and the fusion features of the lidar mode in a preset dimension to construct cross-modal input features; The Transformer interaction unit is used to receive cross-modal input features and uses a global attention mechanism to model the relationship between the image modality and the lidar modality, enabling bidirectional interaction and alignment between the semantic information in the image modality and the geometric information in the lidar modality. The joint characterization output unit is used to output the image modal features and lidar modal features after cross-modal interaction is completed; Step 3. Train the vehicle-mounted multimodal perception model based on the training dataset, and use the trained model to obtain the location recognition result corresponding to the current position of the vehicle.

2. The vehicle autonomous navigation location identification method based on multimodal spatial and frequency domain fusion according to claim 1, characterized in that, In step 1, the surround view image data is obtained by multiple cameras mounted on the vehicle. The preprocessing process for the panoramic image data and LiDAR point cloud data is as follows: First, the images captured by multiple cameras are stitched together to obtain a complete image; then, the collected LiDAR point cloud data is projected and transformed to form a two-dimensional distance view.

3. The vehicle autonomous navigation location identification method based on multimodal spatial and frequency domain fusion according to claim 1, characterized in that, The dual-branch feature extraction unit includes a lidar feature extraction unit and an image feature extraction unit; The image feature extraction unit is used to receive image modal data and extract spatial domain features from it; the lidar feature extraction unit is used to receive lidar modal data and extract spatial domain features from it. The multimodal features output by the cross-modal space and frequency domain interaction unit include image modal features and lidar modal features, and are respectively residually connected with the feature extraction branch results of the corresponding modality on the feature extraction layer; The image modal features and lidar modal features obtained after residual concatenation are used as the multimodal feature inputs for the next feature extraction layer; The input to the first feature extraction layer is the original stitched overall image and a two-dimensional distance view.

4. The vehicle autonomous navigation location identification method based on multimodal spatial and frequency domain fusion according to claim 1, characterized in that, The intramodal multi-frequency fusion unit includes an image adaptive wavelet transform unit, a lidar adaptive wavelet transform unit, an image modal multi-frequency fusion unit, and a lidar modal multi-frequency fusion unit; The image adaptive wavelet transform unit is used to perform frequency domain decomposition on the image modal features, dividing the input image features into low-frequency features and high-frequency features; the lidar adaptive wavelet transform unit is used to perform vertical frequency domain decomposition on the lidar modal features, dividing the input range view features into low-frequency features and high-frequency features along the vertical direction. The low-frequency features obtained from the frequency domain decomposition of image modal features are used to characterize the overall outline, main structure and stable semantic information of the scene; the high-frequency features are used to characterize the edge, texture and detail changes of the scene in different directions. Low-frequency features obtained by frequency domain decomposition of lidar modal features are used to characterize the relatively stable overall geometric contour and structural distribution information in the scene, while high-frequency features are used to characterize local geometric changes, edge abrupt changes and detail undulation information. Image modal multi-frequency fusion unit or lidar modal multi-frequency fusion unit is used to jointly model low-frequency and high-frequency features extracted within the same modality to obtain fused features that have both global stability and local discriminativeness.

5. The vehicle autonomous navigation location identification method based on multimodal spatial and frequency domain fusion according to claim 4, characterized in that, The image adaptive wavelet transform unit includes a two-dimensional discrete wavelet transform segmentation unit, a low-frequency branch, a high-frequency branch, an update unit, a prediction unit, and an attention enhancement unit; The two-dimensional discrete wavelet transform segmentation unit is used to perform frequency domain decomposition on the input image features, decomposing the input image features into low-frequency subband LL and high-frequency subbands LH, HL and HH; the low-frequency subband LL enters the low-frequency branch, and the high-frequency subbands LH, HL and HH are first mean-calculated to obtain a unified high-frequency representation, and then enter the high-frequency branch after convergence processing. The update unit uses detailed information from high-frequency branches to supplement and correct low-frequency branches, thereby enhancing the perception of local changes and edge structures in low-frequency branches. The prediction unit uses overall structural information from low-frequency branches to constrain and guide high-frequency branches, thereby improving the response of high-frequency branches to effective texture details and suppressing noise interference. After the update and prediction are completed, the low-frequency branch and the high-frequency branch are respectively entered into the corresponding attention enhancement unit; The attention enhancement unit includes a channel attention mechanism unit and a multi-head attention mechanism unit; The low-frequency branch access channel attention mechanism unit assigns weights according to the importance of different channels to the overall scene structure representation, thereby enhancing the low-frequency semantic information related to the place recognition task. High-frequency branches are connected to multi-head attention mechanism units to perform parallel modeling of local texture details in different directions and at different scales, thereby improving the ability of high-frequency features to express edge structures, fine-grained textures and local discriminative cues.

6. The vehicle autonomous navigation location identification method based on multimodal spatial and frequency domain fusion according to claim 5, characterized in that, Each update unit / prediction unit includes a reflection filling subunit, a convolutional feature extraction subunit, an activation function subunit, a channel mapping subunit, and an output constraint subunit. The input features are first expanded by the reflection filling subunit to mitigate feature truncation and information loss caused by convolution operations in the boundary region. Subsequently, local neighborhood features are extracted by the convolutional feature extraction subunit with a kernel size of 3×3. Then, nonlinear enhancement is performed through the activation function subunit; Then, cross-channel feature mapping and compression are performed through a channel mapping subunit with a kernel size of 1×1; finally, the output result is constrained by the Tanh function in the output constraint subunit to generate guiding features for updating or prediction.

7. The vehicle autonomous navigation location identification method based on multimodal spatial and frequency domain fusion according to claim 4, characterized in that, The adaptive wavelet transform unit of the lidar includes a vertical segmentation unit, an update unit, a prediction unit, and an attention enhancement unit; wherein the vertical segmentation unit is used to decompose the input lidar range view features along the vertical direction to obtain interrelated low-frequency branch features and high-frequency branch features. The update unit is used to supplement and correct the overall contour information in the low-frequency branch using local geometric details in the high-frequency branch, so as to enhance the perception ability of the low-frequency branch to local structural changes; the prediction unit is used to guide and constrain the detailed response in the high-frequency branch using the overall structural information in the low-frequency branch, so as to reduce the sensitivity of the high-frequency branch to noise and random disturbances. After the update and prediction are completed, the low-frequency branch and high-frequency branch features are respectively input into the corresponding attention enhancement units; The attention enhancement unit employs a channel attention mechanism and is used to weight and enhance the updated and predicted low-frequency branch features and high-frequency branch features respectively, so as to output the frequency domain enhancement features of the lidar mode.

8. The vehicle autonomous navigation location identification method based on multimodal spatial and frequency domain fusion according to claim 4, characterized in that, The image modal multi-frequency fusion unit or the lidar modal multi-frequency fusion unit have the same structure, and both include a high-frequency feature input terminal, a low-frequency feature input terminal, a linear mapping unit, a correlation calculation unit, an activation unit, a weighted fusion unit, and an output unit. The high-frequency feature input terminal is used to input the high-frequency features obtained by frequency domain decomposition, and the low-frequency feature input terminal is used to input the low-frequency features obtained by frequency domain decomposition. The linear mapping unit is used to perform mapping transformations on high-frequency features and low-frequency features respectively to generate query vector Q, key vector K and value vector V for subsequent interactive computation; The relevance calculation unit is used to calculate feature similarity based on the association between query vector Q and key vector K; Activation units are used to normalize or non-linearly enhance similarity to obtain attention weights; The weighted fusion unit is used to reconstruct the value vector V based on the attention weights and fuse it with low-frequency features; the output unit is used to output the multi-frequency feature representation of the fused image mode or lidar mode.

9. The vehicle autonomous navigation location identification method based on multimodal spatial and frequency domain fusion according to claim 1, characterized in that, There are two global feature aggregation modules, and each global feature aggregation module is used to perform global aggregation of image modal features and lidar modal features to obtain the corresponding modal-level global representation; Subsequently, the global descriptor output module connects the global representations of each modality to form a unified global descriptor. The global descriptor is used to perform similarity matching with historical location features in the database to output the final location identification result.