Multi-modal descriptor location identification method based on Mamba mechanism

By adopting a multimodal descriptor location recognition method based on the Mamba mechanism in the field of computer vision, the calculation complexity and real-time problems when the camera and lidar data are fusion in the prior art are solved, and a more efficient location recognition effect is achieved.

CN120071064AActive Publication Date: 2025-05-30HEBEI UNIV OF ENG
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510129052.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-05-30
Estimated Expiration
2045-02-05

AI Technical Summary

Technical Problem

The prior art is more prominent when fusing cameras with lidar data, computational complexity and real-time problems are particularly prominent, especially when processing long sequence and high-dimensional features, resulting in delays in system response time.

Method used

The location recognition method of multimodal descriptor based on the Mamba mechanism is adopted, and the multi-view image and distance image are extracted and fused through the Mamba Fusion module to generate multimodal descriptors to improve the accuracy and robustness of site recognition.

Benefits of technology

The Mamba mechanism is used to realize sequence modeling of linear time complexity, which improves the computing efficiency, reduces calculation delays, and enhances the accuracy and robustness of site recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071064A_ABST
    Figure CN120071064A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal descriptor location identification method based on a Mama mechanism, and the method comprises the steps: carrying out the feature extraction of a multi-view image and a distance image through a plurality of levels of fusion modules, and obtaining image branch features and laser radar branch features; wherein the fusion module comprises a Mama Fusion module, the Mama Fusion module processes the input features after convolution processing to obtain split features, splices the input features and the split features to obtain output features of the current fusion module as input features of the next-level fusion module, compresses image branch features and laser radar branch features respectively, and obtains the input features of the next-level fusion module; image branch panoramic features and laser radar branch panoramic features are obtained; and carrying out perception extraction and fusion on the image branch panoramic features and the laser radar branch panoramic features to obtain a global multi-mode descriptor, and searching according to the global multi-mode descriptor of a site to obtain a site identification result so as to achieve a positioning function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a multi-modal descriptor location recognition method based on the Mamba mechanism. Background Art

[0002] Location recognition is an important task in the field of computer vision and has a wide range of practical applications, such as the closed loop of SLAM or autonomous driving. Location recognition is usually a retrieval problem. By comparing the current sensor observation information with a map or a database of previous observations, the closest observation value is retrieved to determine whether the sensor is currently located at a previously visited location. It is an important branch of simultaneous localization and mapping (SLAM), which can find the re-visited locations, thereby reducing the cumulative error.

[0003] The essence of multi-modal fusion lies in mining the complementary information between different modalities of different sensors. Using a single sensor for location recognition is affected by different environments. For example, using only an RGB camera is affected by different weather and lighting changes; using only a lidar cannot capture the fine details of the observed scene. To improve the location recognition problem and enhance the robustness, the respective advantages of lidar and camera are utilized to enhance the location recognition ability in a fused manner.

[0004] The traditional method of fusing a camera and a lidar is to splice and weighted average the camera and lidar data in the way of Tansformer and CNN. However, this will increase the computational complexity. Especially when dealing with long sequences and high-dimensional features, the computational and memory requirements increase significantly; there is also the problem of real-time performance. The computational delay may affect the response time of the system, especially in a dynamic environment, and the above problems are more prominent. Summary of the Invention

[0005] To solve the above technical problems, the present invention proposes a multi-modal descriptor location recognition method based on the Mamba mechanism to solve the problems existing in the above prior art.

[0006] To achieve the above object, the present invention provides a multi-modal descriptor location recognition method based on the Mamba mechanism, including:

[0007] Obtain multi-view images and distance images;

[0008] Extract features from the multi-view images and the distance image through several levels of fusion modules to obtain image branch features and lidar branch features. Among them, the fusion module includes a Mamba Fusion module. In the fusion module, two convolutional layers are connected to the input end of the Mamba Fusion module, which are respectively used to process the input features and then input them into the Mamba Fusion module. The Mamba Fusion module processes the input features after convolutional processing to obtain multi-view image split features and lidar split features, and performs corresponding splicing on the input features, multi-view image split features, and lidar split features to obtain the output features of the current-level fusion module as the input features of the next-level fusion module. The input features of the first-level fusion module are multi-view images and distance images, and the output of the last-level fusion module is image branch features and lidar branch features.

[0009] Perform compression processing on the image branch features and lidar branch features respectively to obtain image branch panoramic features and lidar branch panoramic features.

[0010] Perform perceptual extraction and fusion on the image branch panoramic features and lidar branch panoramic features to obtain a global multi-modal descriptor.

[0011] Optionally, the process of obtaining the distance image includes:

[0012] Obtain point cloud data, and use spherical projection on the point cloud data to obtain a distance image.

[0013] Optionally, in the Mamba Fusion module, reshape and splice the input features, and input the spliced features into the Mamba module. The Mamba module processes the spliced features, performs feature fusion and split reshaping on the features processed by the Mamba module, so that the size of the split reshaped features is the same as that of the input features, and output features are obtained. The output features include image encoding features and lidar encoding features, and the image encoding features and lidar encoding features are used as the input features of the next-level fusion module.

[0014] Optionally, in the Mamba module, the Mamba module includes a first linear layer and a second linear layer connected in parallel. The first linear layer and the second linear layer are used to perform linear processing on the spliced features. The output of the first linear layer is directly used as the A matrix of the SSM. The X matrix, B matrix, and C matrix of the SSM are obtained by calculating the output of the first linear layer processed by the convolutional layer through different activation functions. The output of the SMM is weighted and spliced with the calculation result of the output of the second linear layer processed by the convolutional layer through the activation function and then normalized. The normalized result is processed by the third linear layer and normalized again to obtain the features processed by the Mamba module.

[0015] Optionally, the image branch features and lidar branch features are respectively compressed through the VC layer.

[0016] Optionally, the perception module performs perception extraction on the panoramic features of the image branch and the panoramic features of the lidar branch, where the perception module includes an MLP layer, a NetVLAD, and an MLP layer connected in sequence.

[0017] Optionally, the panoramic features of the image branch and the panoramic features of the lidar branch are perceptually extracted to obtain an image branch sub-descriptor and a lidar sub-descriptor, and the image branch sub-descriptor and the lidar sub-descriptor are fused through a splicing operation to obtain a global multi-modal descriptor.

[0018] Optionally, the process of obtaining the location recognition result includes: comparing and querying the location global multi-modal descriptor with the data stored in the database to obtain the descriptor most similar to the location global multi-modal descriptor, and using the location data corresponding to the most similar descriptor as the location recognition result.

[0019] Compared with the prior art, the present invention has the following advantages and technical effects:

[0020] The present invention provides a multi-modal descriptor location recognition method and system based on the Mamba mechanism for camera and radar fusion, which is used for location recognition in the global positioning and SLAM closed-loop of intelligent vehicle autonomous driving. The present invention considers that when the self-attention mechanism based on the Transformer model is working, it will generate quadratic complexity, resulting in problems such as slow fusion speed and low efficiency. To solve the above defects, the fusion module adopts the Mamba mechanism, which is based on the state space model (SSM), enabling it to capture the global context information of the input data, rather than just local features. Compared with traditional Transformer-based methods, the Mamba Fusion module maintains a linear computational complexity. This global perception ability helps to retain and integrate more extensive spatial information during the location recognition process, thereby improving the accuracy and robustness of recognition. Through multi-scale extraction and fusion, a multi-modal descriptor is finally formed. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0022] Figure 1 It is a schematic flow chart of the multi-modal descriptor location recognition method based on the Mamba mechanism according to the embodiment of the present invention;

[0023] Figure 2 This is the fusion flowchart of the Mamba Fusion module according to an embodiment of the present invention. Detailed implementation manners

[0024] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0025] It should be noted that the steps shown in the flowchart of the drawings may be executed in a computer system such as a set of computer executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.

[0026] Aiming at the deficiencies of the prior art, the present invention proposes a multi-modal descriptor location recognition method based on the Mamba mechanism, which adopts the structure of a deep learning network and uses the Mamba mechanism model to fuse the camera and lidar to generate a multi-modal descriptor, so as to improve the accuracy, robustness and recognition speed of location recognition.

[0027] The Mamba model realizes sequence modeling with linear time complexity by introducing the concept of a selective state space. When processing long sequence data, its computational complexity increases linearly with the sequence length, rather than growing quadratically like the Transformer, thus significantly improving the operation efficiency. Secondly, the Mamba model adopts a hardware-aware parallel scan algorithm, which calculates through a recursive mode, optimizes the calculation process, and reduces the IO access between GPU memory levels, thereby achieving an improvement in the calculation speed on an RTX3090Ti GPU, up to 3 times at most. The Mamba model introduces a novel selective mechanism, enabling the model to dynamically adjust its behavior according to the input. This mechanism enables the model to effectively filter out irrelevant information and strengthen task-related information, similar to introducing a gating mechanism in an RNN, but more flexibly applied to the model within the framework of the SSM. And due to its linear time complexity characteristics, the Mamba model is more efficient when processing long sequence data. When processing multi-modal data, Mamba can make full use of its ability to capture long-range dependencies, giving play to the efficiency advantages in the learning and reasoning processes, making the Mamba model have a more efficient ability to process multi-modal data. Therefore, we propose to apply the Mamba mechanism for camera and lidar data fusion to overcome the problems of computational complexity and real-time performance, generate multi-modal descriptors, and improve the robustness, accuracy and recognition speed in location recognition.

[0028] Such as Figure 1As shown in the figure, in this embodiment, a multi-modal descriptor location recognition method based on the Mamba mechanism is provided, including the following steps:

[0029] First, obtain point cloud data through a LIDAR radar and multi-view images through cameras with different perspectives. To solve the unordered arrangement of point cloud data, the point cloud data is converted into a 2D distance image through relevant preprocessing, and the distance image is used as the input for extracting descriptors of point cloud features.

[0030] After that, the input convolutional layer Input conv is used to perform preliminary feature extraction on the multi-view images (RGB images) and the distance images corresponding to the point cloud data obtained by the LIDAR respectively. (For multi-view images, the main information extracted is the color and texture information of the environment, which helps to identify semantic content in the environment; the LiDAR point cloud is projected into a distance image, and this distance image contains the distance information of each point in the point cloud data to the sensor, and this information is used as the input for feature extraction). By extracting the information contained in the multi-view images and the LIDAR distance images, the features of the multi-view images and the distance images at different resolutions are obtained respectively.

[0031] After that, the features of the multi-view images and the features of the distance images collected are sent to the Mamba Fusion module for multi-scale fusion. Since the multi-view images carry color and texture information and mainly lack depth information, while the distance images carry depth information, after being processed by the Mamba Fusion module, the features from the image branch and the lidar branch, that is, the features of the multi-view images and the features of the distance images, are fused together. At this time, the fused features are obtained, and these fused features carry both texture information and depth information. Then, the fused features are segmented and reshaped into the original size and sent back to the original branch to be fused with the corresponding original features. This is to extract feature information from different scales.

[0032] After being processed by the Mamba module four times, the information features of the multi-view images and the distance images with different extraction degrees are fused and extracted, making full use of the environmental panoramic views of different sensor modalities and their mutual relationships to generate discriminative location recognition descriptors respectively, including the image branch features and the lidar branch features.

[0033] Afterwards, the outputs of the two encoding branches (image branch features and LiDAR branch features) are respectively input into the VC layer for compression processing to obtain more dense information. The VC layer can effectively reduce the spatial dimension of the features while maintaining the feature expression ability. Then, it successively passes through the MLP layer, NetVLAD, and MLP layer to further aggregate and compress the features, thereby generating two sub-descriptors. Finally, the sub-descriptors from the two encoding branches are connected through the descriptor aggregation module to generate a multi-modal global descriptor. The multi-modal global descriptor is then used for database creation and fast location retrieval. By comparing and querying the features with the data in the database using the multi-modal global descriptor, the descriptor closest to this descriptor is found to achieve place recognition and reach the positioning function.

[0034] A detailed description of the above technical solution is as follows:

[0035] S1. First, point cloud data is obtained through a LIDAR radar, and multi-view images are obtained through cameras from different perspectives. The multi-view images are RGB images. Afterwards, the point cloud data is preprocessed to generate a distance image. Specifically, a spherical projection is used on the point cloud data obtained by the LIDAR radar to generate a distance image. Each three-dimensional point P=(x, y, z) in the point cloud data is converted into the corresponding image two-dimensional R=(u, v) through formula (1), where x, y, z represent the numerical values of different dimensions in the coordinate system, and u, v represent the horizontal and vertical coordinates of the pixels corresponding to the point cloud in the distance image:

[0036]

[0037] where r is the distance, f = f up + f down is the vertical field of view of the sensor, f up represents the maximum field of view, f down represents the minimum field of view, and w, h are the width and height of the distance image R. The point cloud binary file of the dataset is converted into a depth image, that is, a distance image. The size of the distance image is 1×32×1056, where 1 is the depth of the distance image, 32 is the height of the distance image, and 1056 is the width of the distance image.

[0038] S2. The multi-view images and the distance images are respectively sent into the camera and LiDAR encoding branches. The size of the input RGB image is 3×704×256.

[0039] S3 extracts blocks through the InputConv input convolutional layer, extracts intermediate features of the multi-view image and the distance image, obtains multi-view image features and distance image features, and inputs the multi-view image features and distance image features into the MambaFusion module. The MambaFusion module reshapes and fuses the multi-view image and LIDAR distance image features, and then processes and splits and reshapes the fused intermediate features through MambaFusion to obtain intermediate features that match the corresponding branch sizes and sends them back to the original branches. This process is repeated 4 times through MambaFusion processing.

[0040] Among them, as Figure 2 shown, in the MambaFusion module, first, the multi-view image features and distance image features are reshaped and concatenated. The concatenated features are input into the Mamba module. The Mamba module processes the concatenated features, performs feature fusion and split reshaping on the processed features, reshapes them to the original size, obtains multi-view image split features and lidar split features, concatenates the multi-view image split features and multi-view image features, and concatenates the distance image split features and distance image features to obtain encoded features for each branch, including image encoded features and lidar encoded features. The encoded features of each branch are processed through a convolutional layer and then input into the MambaFusion module again. The above process is repeated until the maximum number of processing times of the MambaFusion module is reached, and the output features of the last Mamba Fusion module are output to obtain image branch features and lidar branch features.

[0041] In the Mamba module, the first linear layer and the second linear layer are included to perform linear processing on the features after concatenating the image encoded features and lidar encoded features respectively. The output of the first linear layer is directly used as the A matrix of the SSM. The X matrix, B matrix, and C matrix of the SSM are obtained by calculating the output of the first linear layer processed by the convolutional layer through different activation functions. The output of the SMM is weighted, concatenated with the output of the second linear layer processed by the convolutional layer through an activation function, and extended through a state expansion factor. The normalized result is processed through the third linear layer and then normalized again as the final output.

[0042] In the Mamba module, the data operation process of the SSM model is as follows: First, data rearrangement is performed, and the input sequence x(t), matrices A, B, and C are divided into several blocks according to chunk_size;

[0043] In different blocks, intra-block calculations are then performed to calculate the output \(Y_{diag}\) within each block, which is obtained through the matrix multiplication of \(C\), \(B\), \(L\) (the cumulative effect of the state transition matrix calculated by the segsum function), and the input \(x\), i.e., \(Y_{diag}=C\times B\times L\times x\). The final state \(states\) within each block is calculated through the matrix multiplication of \(B\), \(decay\_states\) (calculating the decay of the state), and \(x\), i.e.,

[0044] \(states = B\times decay\_states\times x\), representing the state within the block;

[0045] Then an intra-block loop is carried out. The state is initialized using \(initial\_states\), or the default zero vector. \(decay\_chunk\) calculates the inter-block decay. \(new\_states\) is obtained through the matrix multiplication of \(decay\_chunk\) and \(states\), i.e., \(new\_states = decay\_chunk\times states\), representing the updated state. Then, for state output conversion, \(state\_decay\_out\) calculates the impact of state decay on the output. \(Y_{off}\) is obtained through the matrix multiplication of \(C\), \(states\), and \(state\_decay\_out\), i.e., \(Y_{off}=C\times states\times state\_decay\_out\), representing the contribution of inter-block information transfer to the output. Finally, output merging is performed by adding the intra-block output \(Y_{diag}\) and the inter-block output \(Y_{off}\) to obtain the final output \(Y\).

[0046] In the proposed MambaFusion fusion module, the Mamab model can replace the SSM model by introducing the structured state space sequence model (S4), preprocess the input data through the input selection mechanism, and use the hardware-aware parallel algorithm to set the running environment of the model. The state space duality (SSD) is introduced, and the multi-head attention (MHA) is integrated into the SSM model or S4 model to optimize the framework. When data enters the Mamab model, data preprocessing is first performed. The multi-view image features and the range image features of the lidar are normalized and scaled, mapped to the same dimension to ensure the consistency of the data range. Under the action of the input selection mechanism, the multi-view image and range image features are processed by linear transformation, attention mechanism, gated unit or state-related function, and then screened. This screening method can be random, fixed or dynamic; finally, the screening result s[k] related to the task is generated; the state is implicitly updated by the HiPPO algorithm, local features are constructed, and local features are output. Under the action of the multi-head attention mechanism (MHA), the attention weights are calculated through the Q, K, V mechanism to enhance the fine-grained relationship global features, and the enhanced local features are output. Then, the fusion and splitting of the two features are performed. The state space duality (SSD) is mainly responsible for controlling the feature evolution and update, and accelerating the screening efficiency and processing efficiency. It is the underlying implementation logic of the Mamba module.

[0047] S4 is an extension of the SSM (state space model). S4 is a structured state space sequence model. It is inspired by a specific continuous system and maps one-dimensional sequences through implicit latent states, which is equivalent to simply writing the SSM in the form of matrix multiplication. The S4 model constructs matrix A by introducing HiPPO (High-order Polynomial Projection Operators). This matrix A can capture the nearest tokens and decay the old tokens, thus better memorizing historical information.

[0048] In the Mamba module, the SSM (state space model) and the structured state space model (S4) are inspired by continuous systems and are represented by linear differential equations in mathematics:

[0049]

[0050] A represents the state transition matrix; B represents the input matrix; C represents the output matrix; D represents the direct transfer matrix; x(t), where X in the attached figure represents the input sequence; y(t) represents the output sequence; h(t) represents the hidden state, and h'(t) represents the derivative of the hidden state h(t) with respect to time t. and If the system is affected by past states and they are introduced, A i (t) is the state matrix related to time delay, i is the time step; τ is the unit time delay; C i (t) is the output matrix related to time delay, N is the hidden state size, and M is the output sequence size.

[0051] Due to the continuous-time model, integrating SSM and S4 into deep learning algorithms faces great challenges, so zero-order hold is used to discretize it:

[0052]

[0053] y t = h t

[0054] By simplifying the structures of A, B, and C, for example, the structure of matrix A is further simplified from diagonal to scalar multiplied by the identity matrix structure, it can be represented by a scalar, where the superscript - represents the scalar representation of the matrix, and the state space duality SSD improves the original Mamba performance.

[0055] S4 sends the finally obtained image branch features and lidar branch features into the corresponding VC layers respectively. Through the VC layers, the image branch features and lidar branch features are compressed to obtain more dense image branch panoramic features and lidar branch panoramic features.

[0056] S5 inputs the panoramic features processed by the VC layer into the perception module respectively. The perception module sequentially includes connected MLP layers, NetVLAD, and MLP layers, and generates image branch sub-descriptors and lidar sub-descriptors through the perception module.

[0057] S6 splices the image branch sub-descriptors and lidar sub-descriptors through the descriptor aggregation module to generate a global multi-modal descriptor for querying or referencing in the database.

[0058] For the above camera and lidar encoding branches, under different branches, as an improvement, the input convolutional layer is improved through different adaptive structures to adaptively extract features from the initial data of the above-mentioned camera and lidar through different structures, providing relevant effective data for the subsequent fusion model and ensuring the improvement of the accuracy of subsequent fusion and data recognition. Among them, the input convolutional layer under the camera encoding branch can choose to use the backbone network of YOLOv8n to extract the structure of the image through a structure with better feature extraction effect to achieve the extraction of each tiny feature of the multi-view image of the camera.

[0059] Meanwhile, the output end of the backbone network of YOLOv8n is further optimized for feature extraction using the MCS module based on multi-channel attention and multi-dimensional weighting to ensure clearer features in multi-view images.

[0060] Specifically, the MCS module combines the MEW multi-axis external weight module with the CBAM attention. The MCS module effectively extracts global and local information through multi-dimensional weighting, thus better capturing the subtle features of different targets in multi-view images. On this basis, the CBAM mechanism makes precise adjustments and optimizations, enabling the model to more concentratedly focus on the features that can reflect the location, significantly reducing background and image object noise interference, and improving detection accuracy. After inserting the MCS module into the backbone network of YOLOv8n, that is, setting the MCS module at the output end of the backbone network, the features processed by the MCS module are used as the final output, enabling the algorithm to better focus on the key features that reflect the location. The formula for the specific process of MEW is as follows:

[0061] x 1 ,x 2 ,x 3 ,x 4 =Split(X)

[0062] x i(I,J) =W (I,J) ⊙F (I,J) [x i

[0063]

[0064] x 4 '=DW(x 4 )

[0065] Y=Concat(x 1 ',x 2 ',x 3 ',x 4 ')+X

[0066] Where X represents the output features of the corresponding backbone network, x represents the features after splitting the output features, split represents the separation function, i = 1, 2, 3 respectively represent the first three branches corresponding to the MEW multi-axis external weight module, i = 4 represents the last branch corresponding to the MEW multi-axis external weight module, and the apostrophe on x represents the corresponding output. W (I,J) and F (I,J) ​respectively represent the learnable external weights of the corresponding axes and the two-dimensional discrete Fourier transform. ⊙ represents element-wise multiplication. When i = 1, (I, J) represents the Height-Width axis; when i = 2, (I, J) represents the Channel-Width axis; when i = 3, (I, J) represents the Channel-Height axis. represents the two-dimensional inverse discrete Fourier transform. DW represents depthwise separable convolution. Y represents the output feature of MEW. The output feature of MEW is input into the CBAM attention for relevant processing to obtain the initial feature of the camera encoding branch for subsequent relevant fusion processing.

[0067] In the input convolutional layer under the lidar encoding branch, a spatial attention mechanism can be optionally added after the convolutional layer for relevant feature extraction. The spatial attention mechanism weights the spatial positions containing the main location features in the distance image, thereby highlighting the regions most contributing to the task and suppressing irrelevant or redundant regions to ensure the effectiveness of feature extraction from the distance image, where

[0068] the spatial attention calculation formula is as follows:

[0069]

[0070] where F represents the feature after convolutional processing of the distance image, σ represents the activation function, f represents convolutional processing, AvgPool represents average pooling, MaxPool represents max pooling, Ms(F) represents the spatial attention feature, and the subscripts avg and max of the F feature represent the corresponding processing through average pooling and max pooling. The feature after spatial attention processing is used as the initial feature of the lidar score for subsequent relevant fusion processing.

[0071] Regarding the above MambaFusion fusion module, the specific content introduced or improved is as follows:

[0072] In the present invention, the Mamba mechanism adopted by the fusion module. For this Mamba mechanism, a structured state space sequence model (S4) can be introduced. This model modifies the A matrix in the SSM into a low-rank corrected conditional matrix. An input selection mechanism and a hardware-aware parallel algorithm are introduced to improve the performance of S4 and significantly improve the computational efficiency. State space duality (SSD) is introduced to enhance the original mamba, and a design integrates multi-head attention (MHA) into the SSM to optimize the framework, thereby promoting the fusion of sensor data.

[0073] S4 is an extension of SSM. S4 is a Structured State Space Sequence Model, which is inspired by a specific continuous system. It maps a one-dimensional sequence through an implicit latent state, equivalent to simply writing the SSM in the form of matrix multiplication. The S4 model constructs matrix A by introducing HiPPO (High-order Polynomial Projection Operators). This matrix A can capture the most recent tokens and attenuate the old tokens, thus better memorizing historical information.

[0074] The matrix A of HiPPO is derived by projecting the past input sequence. A is constructed based on the orthogonal basis of Legendre polynomials.

[0075]

[0076] The input selection mechanism allows the model to selectively propagate or forget information based on the current token. This is achieved by making the SSM (State Space Model) parameters a function of the input, thus addressing the weakness of SSM in processing discrete modal data. It mainly includes projection and selection. The projection part is responsible for mapping the input sequence at the input end of the SSM model to a high-dimensional space suitable for the state space model for preprocessing, ensuring that the input data can effectively interact with the state space model. The projection can be a linear or non-linear transformation, usually completed by a neural network layer. Its goal is to transform the input data into a specific representation required by the model so that subsequent state update and output generation steps can be effectively carried out. Selection mechanism: The selection mechanism dynamically adjusts the parameters of the state space model according to the current input sequence, mainly the input matrix B that affects state transition and the output matrix C that affects the output. Dynamically adjusting parameters: This means that the selection mechanism allows the model to selectively propagate or forget information according to the characteristics of the specific input. For example, when processing different types of input data, the model can optimize the state update and output generation process by dynamically adjusting B and C. The selection mechanism can be based on the attention mechanism, gating mechanism or other methods of dynamically adjusting parameters. In the selective state space model, this flexibility is particularly important because it enables the model to adapt to various complex input data patterns.

[0077] The hardware-aware parallel algorithm is a hardware-aware parallel recursive pattern algorithm. This algorithm utilizes the parallelism of modern accelerators (such as GPUs and TPUs) to run the above model and performs the calculations of the selective SSM in a memory-efficient manner. It can also be said that the hardware-aware algorithm adopts the Parallel Scan Algorithm, which is an algorithm that utilizes the properties of linear associative calculations. It achieves efficient calculations by constructing a balanced binary tree and sweeping from the leaves to the root.

[0078] The SSD algorithm, which is a sequence processing method based on the state space model, decomposes the sequence into several blocks and conducts efficient information transfer within and between the blocks. SSD utilizes the low-rank decomposition and exponential decay characteristics of matrices to transform the complex sequence modeling problem into a series of efficient matrix multiplication operations, thereby significantly reducing the computational complexity.

[0079] An MHA mechanism is set at the output end of the SSM model. The input of MHA includes three vectors: the query vector, the key vector, and the value vector. For a given query vector, MHA will perform a weighted sum on the key vector, where the weights are calculated from the similarity between the query vector and the key vector, and then multiply the obtained weighted sum by the value vector for output. The multi-head mechanism of MHA can effectively improve the expressive power of the model and also enable the model to learn more diverse and complex features. Under the multi-head mechanism, the input sequence data is divided into multiple heads, and each head performs independent calculations to obtain different outputs. These outputs are finally concatenated together to form the final output.

[0080] The training module uses the training set to train the deep neural network model with lidar point cloud data and multi-view camera image data, aiming to improve the accuracy and robustness of place recognition in low-texture and appearance-similar environments. During training, data preprocessing includes uniformly resizing the input camera image data (resize) and applying standard normalization processing. Specifically, the color channels of the image are normalized, with the mean set to [0.485, 0.456, 0.406] and the standard deviation set to [0.229, 0.224, 0.225]. For LiDAR point cloud data, appropriate format conversion and normalization processing are applied to ensure that the point cloud data can be effectively fused with the image data in the subsequent network. Data augmentation operations, including random cropping, rotation, flipping, etc., are applied to the image data to improve the generalization ability of the model and avoid overfitting.

[0081] The training dataset contains three types of data: training query data, database query data, and validation query data. The dataset is trained using a triplet loss function to ensure that the model can learn to distinguish the distance relationship between positive and negative samples. The training set uses the TripletDataset, where each triplet consists of a query image, a positive sample image, and a negative sample image. The training set is constructed according to distance thresholds (nonTrivPosDistThres, posDistThr) and the number of negative samples (nNeg). The full dataset includes the entire database dataset (DatabaseQueryDataset) for model evaluation to ensure that the model can effectively perform place recognition on all database samples. Data loading is batch-processed through the DataLoader provided by PyTorch, and multi-process (num_workers) is used to accelerate the data loading process.

[0082] The present invention uses a deep convolutional neural network (CNN) and a custom multi-modal fusion module (MambaFusion) to perform place recognition tasks by fusing the features of image and point cloud data.

[0083] During the training process, triplet loss is adopted to optimize the performance of the model. The triplet loss function realizes effective discrimination in place recognition tasks by minimizing the distance between the positive sample and the query sample while maximizing the distance between the negative sample and the query sample. The formula for triplet loss L:

[0084] L = max(d(query, positive) - d(query, negative) + margin, 0)

[0085] Where d(query, positive) and d(query, negative) represent the Euclidean distances between the query sample and the positive and negative samples respectively, and margin is a hyperparameter that controls the distance difference between positive and negative samples. The positive sample is defined as the sample located no more than 9 meters away from the capture location of the query. The negative sample is defined as being at least 18 meters away from the capture location of the query.

[0086] During training, the Adam optimizer is used for model training, and the decay factor of the learning rate scheduler is 0.8, with an adjustment interval of 5.

[0087] The training process includes an initial stage, a training stage, and a validation stage. In the initial stage, the model parameters are initialized, and the pre-trained weights are loaded as the initial parameters of the image feature extraction network. In the training stage, in each training cycle, the triplets in the training set are input into the model for forward propagation, the loss is calculated, and the model parameters are updated through backpropagation. The Adam optimizer and the learning rate decay strategy are used to optimize the model. In the validation stage, the data of the entire database is used to evaluate the model, and the recognition accuracy and recall rate are calculated. The training strategy is dynamically adjusted according to the performance of the validation set.

[0088] The above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A multimodal descriptor location recognition method based on Mamba mechanism, characterized in that: include: Acquire multi-view images and range images; The multi-view image and the range image are subjected to feature extraction through several levels of fusion modules to obtain image branch features and laser radar branch features; wherein the fusion module includes a Mamba Fusion module, in which the input end of the Mamba Fusion module is connected to two convolution layers for processing the input features and then inputting them into the Mamba Fusion module, the Mamba Fusion module processes the input features after the convolution processing to obtain multi-view image splitting features and laser radar splitting features, and the input features and the multi-view image splitting features and the laser radar splitting features are correspondingly spliced ​​to obtain the output features of the current level fusion module as the input features of the next level fusion module, wherein the input features of the first level fusion module are the multi-view image and the range image, and the output of the last level fusion module is the image branch features and the laser radar branch features; The image branch features and the laser radar branch features are compressed respectively to obtain the image branch panoramic features and the laser radar branch panoramic features; The image branch panoramic features and the lidar branch panoramic features are sensed, extracted and fused to obtain a global multimodal descriptor of the location. The location recognition result is obtained by searching based on the global multimodal descriptor of the location.

2. The method according to claim 1, characterized in that: The process of acquiring the distance image includes: Point cloud data is acquired, and spherical projection is applied to the point cloud data to obtain a range image.

3. The method according to claim 1, characterized in that In the Mamba Fusion module, the input features are reshaped and spliced, and the spliced ​​features are input into the Mamba module. The spliced ​​features are processed by the Mamba module, and the features processed by the Mamba module are fused and split and reshaped so that the split and reshaped features have the same size as the input features, and the output features are obtained, wherein the output features include image coding features and lidar coding features, and the image coding features and lidar coding features are used as input features of the next-level fusion module.

4. The method according to claim 3, characterized in that: In the Mamba module, the Mamba module includes a first linear layer and a second linear layer connected in parallel, wherein the first linear layer and the second linear layer are used to perform linear processing on the spliced ​​features, and the output of the first linear layer is directly used as the A matrix of the SSM. The X matrix, B matrix and C matrix of the SSM are obtained by calculating the output of the first linear layer after the convolution layer processing through different activation functions. The output of the SMM is spliced ​​and normalized by weighted union and the calculation result of the output of the second linear layer after the convolution layer processing through the activation function, and the normalized result is processed by the third linear layer and normalized again to obtain the features processed by the Mamba module.

5. The method according to claim 1, characterized in that: The image branch features and lidar branch features are compressed respectively through the VC layer.

6. The method according to claim 1, characterized in that The image branch panoramic features and the lidar branch panoramic features are perceived and extracted through a perception module, wherein the perception module includes an MLP layer, a NetVLAD and an MLP layer connected in sequence.

7. The method according to claim 1, characterized in that The image branch panoramic features and the lidar branch panoramic features are perceived and extracted to obtain the image branch sub-descriptors and the lidar sub-descriptors, which are fused with the lidar sub-descriptors through splicing operations to obtain the global multimodal descriptor of the location.

8. The method according to claim 1, characterized in that: The process of obtaining location recognition results includes: The global multimodal descriptor of the location is compared and queried with the data stored in the database to obtain the descriptor that is most similar to the global multimodal descriptor of the location, and the location data corresponding to the most similar descriptor is used as the location recognition result.

Citation Information

Patent Citations

  • Feature-adaptive mutual-guiding multi-source information fusion classification method and system

    CN114187526A

  • Multi-modal descriptor location identification method and system based on camera and radar fusion

    CN117392629A

  • Mama-based point cloud semantic segmentation method and device, equipment and medium

    CN119006814A

  • Unmanned vehicle robust position identification method based on look-around image

    CN119131740A

  • Multisource remote sensing image semantic segmentation method based on Transform, Mama and diffusion model

    CN119152205A