Loopback detection repositioning method and device based on laser radar, inertial measurement unit and large model semantic reasoning, electronic equipment, storage medium and program product

By employing a loop closure detection method that integrates lidar, inertial measurement unit, and large model fusion, the robustness and accuracy issues of lidar in complex indoor environments are resolved, enabling efficient and reliable repositioning and supporting the stable operation and path planning of autonomous navigation systems.

CN121453050APending Publication Date: 2026-02-03BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511564792.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing lidar loop closure detection methods lack robustness in complex indoor environments, making it difficult to balance accuracy and efficiency, resulting in inaccurate repositioning of autonomous mobile robots in environments lacking GPS signals.

Method used

By fusing data from lidar and inertial measurement units and combining it with large-scale semantic reasoning, multi-layered geometric and semantic fusion is achieved. By constructing efficient loop closure detection and global optimization, positioning accuracy and robustness are improved.

Benefits of technology

Achieving high-precision and robust repositioning in complex indoor environments significantly reduces the mismatch rate, improves system stability and reliability, and supports the long-term operation and path planning of autonomous navigation systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121453050A_ABST
    Figure CN121453050A_ABST
Patent Text Reader

Abstract

The invention discloses a loopback detection repositioning method and device based on a laser radar, an inertial measurement unit and large model semantic reasoning, electronic equipment, a storage medium and a program product, and belongs to the technical field of positioning and navigation. The method comprises the following steps: collecting a laser radar point cloud sequence and motion data of an inertial measurement unit in real time; constructing a local geometric feature and a global descriptor based on the point cloud data, and optimizing the matching efficiency in combination with inertial measurement information; after preliminary geometric matching is completed, a large model subjected to multi-modal pre-training is introduced to execute semantic consistency verification and confidence coefficient correction, and a geometric-semantic joint similarity index is generated to suppress false loopback; when effective loopback is detected, geometric constraints, inertial constraints and semantic consistency constraints are jointly introduced into factor graph optimization, and global pose correction is achieved. According to the method, a'geometry-inertia-semantics' three-layer fused closed-loop detection system is constructed, the positioning robustness, continuity and generalization performance in a complex dynamic environment are effectively improved, and high-precision and high-reliability repositioning capability is provided for scenes such as an autonomous mobile robot, intelligent inspection and storage navigation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of positioning and navigation, in particular to a loop detection and repositioning method and device based on a laser radar, an inertial measurement unit and a large model semantic inference, an electronic device, a storage medium and a program product. The core of the method is that in an indoor environment where global signals such as GPS are missing, the point cloud sequence generated by the laser radar and the motion estimation provided by the inertial measurement unit are fused to realize efficient loop detection. When the loop is detected, the system globally optimizes and corrects the accumulated dead reckoning error, thereby realizing accurate repositioning. The present application is specially skilled in the fusion scheme of radar and inertial measurement unit to ensure the robustness and reliability in complex scenes such as light changes, visual feature missing or wireless signal interference. BACKGROUND

[0002] With the rapid development of intelligent technology, intelligent terminals such as autonomous mobile robots and unmanned vehicles are increasingly widely used in large indoor scenes. The stable operation of these systems fundamentally depends on their accurate and robust repositioning ability in GPS signal missing environments. In particular, after a long period of operation or after a tracking failure, whether the system can correct the accumulated dead reckoning error by identifying the visited place (i.e. loop detection) has become a key indicator to measure its intelligent level and reliability.

[0003] In the field of laser SLAM, loop detection is the core link to achieve this capability. For this purpose, researchers have proposed a variety of loop detection methods based on laser radar point cloud data. Some early schemes try to use point cloud matching or global descriptor-based methods directly, but often face the balance problem between computational efficiency and discrimination in large-scale environments. Subsequently, some algorithms based on local feature description and matching were developed, which construct scene representation by extracting key points and feature vectors in point cloud, but their stability and recall rate still need to be improved when facing dramatic changes in viewing angle, scene dynamic interference or long-term environmental changes. In addition, although some schemes draw on the idea of visual Bag-of-Words to accelerate the retrieval process to some extent, there are still bottlenecks in the uniqueness and noise resistance of point cloud feature expression.

[0004] The common dilemma faced by these existing laser radar loop detection schemes is that they either lack robustness in complex scenes and are prone to false matching, or while ensuring accuracy, the algorithm complexity is too high to meet the real-time requirements. This contradiction between performance and efficiency limits the deployment effect of many high-end applications in real, dynamic and large-scale indoor environments.

[0005] Under the background that the existing laser radar loop detection method generally faces the problem that robustness, precision and efficiency are difficult to be considered, how to realize higher reliability and higher precision repositioning in a complex indoor environment has become a technical bottleneck to be broken through in the field. The improvement of this capability is crucial for realizing continuous, stable and reliable indoor autonomous navigation system - accurate and reliable closed-loop correction can not only effectively suppress the cumulative error of long-term operation of the system and improve the consistency of the global trajectory, but also provide a solid pose basis for upper functions such as path planning, scene understanding and decision control, thereby overall promoting the large-scale landing application of service robots, intelligent warehousing and unmanned systems in complex indoor scenes, and is one of the core problems to be solved for the mature and practical high-precision positioning and navigation technology. SUMMARY

[0006] Therefore, the purpose of the present application is to provide a loop detection repositioning method, device, electronic equipment, storage medium and program product based on laser radar, inertial measurement unit and large model semantic reasoning, so as to realize high-precision and high-robustness global repositioning in a complex indoor environment with missing GPS signals. Based on the traditional fusion architecture of laser radar and inertial measurement unit, the present application further introduces the high-level semantic understanding ability of the large model, thereby realizing the transition from geometric perception to semantic understanding, and effectively solving the key problems such as insufficient recognition accuracy, poor scene adaptability, unsatisfactory energy efficiency control and privacy protection in the prior art.

[0007] I. Overall method concept

[0008] Based on the above purpose, the first aspect of the exemplary embodiments of the present application provides a loop detection repositioning method, comprising:

[0009] Real-time acquisition of raw data of a laser radar and an inertial measurement unit carried by a robot;

[0010] Constructing a point cloud descriptor based on the laser radar data, and optimizing the matching efficiency in combination with the inertial measurement unit information, so as to improve the overall speed of loop detection;

[0011] Constructing an environment map based on the point cloud sequence and the inertial measurement data, and realizing efficient loop detection;

[0012] When an effective loop is detected, globally optimizing and correcting the accumulated pose error of the system, so as to realize accurate and robust repositioning.

[0013] II. Semantic enhancement mechanism driven by large model

[0014] On this basis, the application further proposes a repositioning auxiliary mechanism based on a large model (LLM). The mechanism fully utilizes the capabilities of the large model in cross-modal perception, semantic reasoning, and spatio-temporal consistency understanding, and performs deep fusion on the laser radar point cloud and inertial measurement unit data, thereby increasing semantic layer verification on the basis of geometric constraints.

[0015] Specifically, the large model performs global semantic consistency verification on the candidate loopback frame by virtue of its pre-training capability in large-scale space and motion knowledge, and performs semantic weighted correction on the traditional geometric matching confidence, thereby significantly reducing the false matching rate in a dynamic environment.

[0016] In addition, the large model exhibits unique advantages in long-term trajectory consistency modeling. By combining the historical semantic features and inertial drift patterns accumulated by the robot during long-term operation, the application realizes "memory-based repositioning" based on semantic graph priors. This mechanism enables the system to restore the global pose under environmental changes such as sudden changes in illumination, local occlusion, or structural reconstruction, realizing cross-domain fusion from geometric loopback detection to semantic loopback detection.

[0017] III. Multi-source fusion and optimization design

[0018] The application comprehensively utilizes the geometric structure perception capability of the laser radar, the motion constraint capability of the inertial measurement unit, and the high-level semantic reasoning capability of the large model, and constructs an end-to-end, multi-source fusion, and semantic-enhanced loopback detection and repositioning system.

[0019] The system realizes a "geometric-inertial-semantic" three-layer fusion closed-loop correction mechanism, significantly improving the positioning accuracy, stability, and system generalization performance in complex dynamic scenes.

[0020] Further, the application proposes the following innovative designs at the system architecture and implementation level:

[0021] (1) Adaptive multi-modal fusion strategy: dynamically adjust the fusion weight of laser radar and IMU data according to the sensor signal-to-noise ratio, self-motion acceleration, and environmental complexity, thereby maintaining optimal robustness under different working conditions;

[0022] (2) Lightweight local inference architecture: adopt modular parallel computing and sparse matrix optimization strategies to realize full-flow local inference on the robot end, meeting the real-time and energy consumption constraints of power-limited devices;

[0023] (3) Explainability optimization and online self-learning mechanism: By recording the semantic confidence and matching results of the large model in loop detection, an "explainability memory bank" is constructed to realize self-adaptive error correction and online optimization of the system during long-term operation;

[0024] (4) Global consistency correction based on semantic topology graph: Introduce semantic topology constraints in the factor graph optimization stage, so that the relocation not only depends on geometric information, but also uses semantic structure for global consistency propagation, further improving the positioning reliability of the system in large-scale complex space.

[0025] Therefore, the "geometric-inertial-semantic" three-layer fusion framework proposed by the present application not only realizes breakthrough in positioning accuracy and robustness, but also has good scalability and adaptive characteristics. The sensor modalities and model sizes can be flexibly added or reduced according to the task scene, and it has the potential for deployment in multiple scenarios and multiple platforms.

[0026] IV. Device and implementation

[0027] Based on the same inventive concept, the second aspect of the exemplary embodiments of the present application provides a loop detection relocation device, the overall structure of which is used to realize the main steps of the method on a robot or an autonomous mobile device. The device includes the following modules:

[0028] Multi-source data acquisition module. Used for real-time acquisition of raw data from laser radar and inertial measurement unit (IMU). This module can synchronously acquire point cloud sequence, acceleration, angular velocity and other multi-source sensor information, and perform timestamp alignment and data integrity check to ensure the consistency of the sensor data in the time and space dimensions. The acquisition module can be deployed on the robot body or its edge computing unit to provide high-quality input data for subsequent feature extraction and semantic reasoning.

[0029] Feature extraction module. Used for constructing high-distinguishable spatial features based on preprocessed point cloud data. The module extracts robust point cloud structure features through local geometric analysis, polar coordinate grid division and covariance feature decomposition, and generates descriptor vectors with rotation invariance. This module ensures the robustness of feature expression to attitude changes and environmental noise while maintaining computational efficiency.

[0030] Loop matching and semantic verification module. It is the core innovative part of the present application. This module combines IMU motion prior information and large model (LLM) semantic understanding ability to realize multi-level loop detection and verification.

[0031] Specifically, it includes two levels of processes: (1) geometric level matching layer: based on the pose prior of IMU estimation, the candidate scene is quickly filtered in the historical key frame, and the similarity is calculated by using the point cloud descriptor to complete the preliminary loop matching; (2) semantic level verification layer: the inference engine of the fine-tuned large model is called to verify the semantic consistency of the candidate loop frame and correct the confidence, realizing the cross-modal fusion from geometric matching to semantic confirmation.

[0032] Through the above-mentioned double-layer mechanism, the system can effectively reduce the false matching rate in a complex environment and improve the adaptability to dynamic scenes and long-term drift.

[0033] A global optimization module is used to globally optimize and correct the overall trajectory of the system after detecting an effective loop. Based on the factor graph optimization framework, the module introduces the newly identified loop constraint into the global graph, constructs a comprehensive optimization model containing the odometer constraint and the semantic loop constraint, and solves the global pose sequence by minimizing the weighted residual function, so as to realize the consistency reconstruction and error suppression of the trajectory. The optimization result can be fed back to the navigation and path planning system in real time to form a closed-loop control.

[0034] In the above-mentioned device, each module cooperates with each other through data flow and logical dependency relationship:

[0035] The multi-source data acquisition module provides perception input for the system; the feature extraction module structurally expresses the original data; the loop matching and semantic verification module is responsible for loop identification and semantic confirmation; and the global optimization module completes the final pose fusion and trajectory correction.

[0036] Such hierarchical decoupling design not only ensures the modular implementation of the system, but also has good expansibility and portability. In actual application, each module can be flexibly combined according to the hardware conditions and task scenes, and the deployment in an embedded platform, an edge computing unit or a cloud collaborative system is supported.

[0037] Based on the same inventive concept, the third aspect of the exemplary embodiments of the present application provides an electronic device including a processor, a memory and a communication interface and the like hardware units, and the computer program instructions stored in the memory are executable on the processor, and are used to execute all or part of the steps of the loop detection and repositioning method.

[0038] The fourth aspect of the exemplary embodiments of the present application provides a non-transitory computer readable storage medium for storing computer program instructions, when the instructions are executed by a processor, all or part of the functions of the above-mentioned method are realized.

[0039] In addition, the fifth aspect of the exemplary embodiments of the present application provides a computer program product, which comprises computer program instructions, when the program runs on a computer or a robot terminal, makes the computer execute the loop detection repositioning method, so as to realize high-precision repositioning based on laser radar, inertial measurement unit and large model fusion.

[0040] V. Overall innovation summary

[0041] In summary, the loop detection repositioning method based on laser radar, inertial measurement unit and large model fusion provided by the present application realizes high-robustness repositioning in complex scenes through multi-source data acquisition, feature extraction, semantic enhancement loop matching and global optimization, breaks through the robustness bottleneck of traditional SLAM in dynamic environment, and provides a feasible high-precision positioning solution for intelligent inspection, indoor navigation, warehouse logistics and autonomous mobile system.

[0042] The present application provides a new theoretical framework and engineering implementation path for constructing an autonomous positioning and cognitive system for complex environments by fusing traditional sensor positioning technology and intelligent semantic reasoning mechanism. BRIEF DESCRIPTION OF DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the present disclosure or the related art, the following will briefly introduce the drawings needed to be used in the embodiments or related art descriptions. Obviously, the drawings in the following description are only embodiments of the present disclosure, and other drawings can also be obtained by those skilled in the art without creative labor.

[0044] Figure 1 An application scenario diagram of the loop detection repositioning method provided by the embodiments of the present disclosure;

[0045] Figure 2 A flow diagram of the loop detection repositioning method provided by the embodiments of the present disclosure;

[0046] Figure 3 A structure diagram of the loop detection repositioning method device provided by the embodiments of the present disclosure;

[0047] Figure 4 An electronic device hardware structure diagram of the embodiments of the present disclosure; DETAILED DESCRIPTION

[0048] It can be understood that before using the technical solutions disclosed in the embodiments of the present application, the type, use range, use scene, etc. of the personal information involved in the present application should be informed to the user and the authorization of the user should be obtained in accordance with relevant laws and regulations.

[0049] For example, in response to receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed by the user will need to acquire and use personal information of the user. Thus, the user can autonomously select whether to provide personal information to the software or hardware such as an electronic device, an application program, a server or a storage medium, etc. performing the operation of the technical solution of the present application according to the prompt information.

[0050] As an optional but non-limiting implementation manner, in response to receiving an active request of a user, the manner of sending a prompt information to the user may, for example, be a pop-up window manner, in which the prompt information can be presented in a text manner. In addition, the pop-up window can also carry a selection control for the user to select “agree” or “disagree” to provide personal information to the electronic device.

[0051] It can be understood that the above notification and acquisition of user authorization process is only illustrative and does not limit the implementation manner of the present application, and other manners meeting relevant laws and regulations can also be applied to the implementation manner of the present application.

[0052] It can be understood that the data (including but not limited to the data itself, acquisition or use of the data) involved in the technical solution should comply with the requirements of relevant laws and regulations and relevant provisions.

[0053] In order to make the purpose, technical solutions and advantages of the present disclosure clearer, the principles and spirits of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are only given to enable those skilled in the art to better understand and implement the present disclosure, and do not limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.

[0054] In this document, it should be understood that any number of elements in the drawings are used for illustration and not limitation, and any naming is only for distinction and does not have any limiting meaning.

[0055] It should be noted that, unless otherwise defined, technical terms or scientific terms used in the embodiments of the present disclosure shall have the common meaning understood by one of ordinary skill in the art to which the present disclosure belongs. The terms "first", "second", and similar terms used in the embodiments of the present disclosure do not denote any order, quantity, or importance, but are used to distinguish different components. The terms "include", "contain", and similar terms mean that the elements or objects before the terms encompass the elements or objects listed after the terms and their equivalents, and do not exclude other elements or objects. The terms "connect" or "connected" or similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms "upper", "lower", "left", "right", and the like are used only to represent relative positional relationships, and when the absolute positions of the described objects are changed, the relative positional relationships can also be changed accordingly. The article "a" or "an" before an element does not exclude the presence of multiple such elements.

[0056] The principles and spirits of the present disclosure will be explained in detail below with reference to several representative embodiments of the present disclosure.

[0057] In the background that the existing laser radar loop detection method generally faces the difficulty of balancing robustness, precision and efficiency, how to realize higher reliability and higher precision repositioning in complex indoor environment has become a technical bottleneck that needs to be broken through in this field. The improvement of this ability is crucial for realizing continuous, stable and reliable indoor autonomous navigation system - accurate and reliable closed-loop correction can not only effectively suppress the cumulative error of long-term operation of the system and improve the consistency of the global trajectory, but also can provide a solid pose foundation for upper functions such as path planning, scene understanding and decision control, thereby overall promoting the large-scale landing application of service robots, intelligent warehousing and unmanned systems in complex indoor scenes. It is one of the core problems that must be solved for the mature and practical high-precision positioning and navigation technology.

[0058] In the dead reckoning system based on the inertial measurement unit, the random walk error inherent in the sensor will accumulate over time, causing the positioning trajectory to drift significantly, which is the core challenge faced in realizing long-term accurate positioning in complex indoor environments. The present invention recognizes this common problem, and to overcome this difficulty, the inventors propose to fuse the multi-source data of laser radar and inertial measurement unit, and perform efficient loop detection on the robot, which can effectively identify and correct such accumulated errors. The system makes full use of the laser radar and inertial sensors equipped on the robot, and collects real-time high-precision three-dimensional point cloud and motion inertia information in the environment. By constructing an environment descriptor with strong discrimination and performing fast loop matching, the system can fully cover indoor scenes of different structures and layouts, and deeply fuse environmental context information, thereby globally optimizing the system pose when the loop is closed, and ultimately significantly improving the accuracy, robustness and long-term stability of the positioning system.

[0059] In order to overcome the limitations in the prior art, the present invention provides a loop detection and repositioning method based on laser radar, inertial measurement unit and large model semantic reasoning.

[0060] The system can be completely independently run locally on the robot without relying on external network or cloud computing power. The terminal device collects real-time multi-source sensor data such as laser radar and inertial measurement unit, and automatically completes a series of processes such as data preprocessing, feature extraction, descriptor construction, loop detection, post-processing optimization and result output in the on-board computing unit. The system effectively improves the accuracy and stability of the identification process through multi-sensor data cross-validation, loop closure confirmation based on geometric consistency, and error matching elimination mechanisms. Finally, the system outputs structured loop detection result information, including loop frame ID pair, relative pose transformation matrix, matching confidence and optimized global pose, which can be used for warehouse logistics robot global positioning correction, service robot long-term autonomous navigation, industrial factory scene inspection path planning and other practical application scenarios, significantly enhancing the operation reliability of intelligent terminals in complex environments.

[0061] Reference Figure 1 Fig. 1 shows an application scenario of the loop detection and repositioning method provided by the exemplary embodiments of the present invention. The scenario mainly includes a robot 101 and a server 102. The robot 101 is responsible for carrying and running the repositioning system described in the present invention, and performs real-time data collection and loop detection in the indoor environment; the server 102 can be used for task planning and system management and other interactive operations. The robot 101 can be completely locally run during positioning and navigation, forming an autonomous and reliable closed-loop system.

[0062] It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principles of the present application, and the embodiments of the present application are not limited in this respect. In actual application, the system can be flexibly deployed according to different building environments and terminal device types to meet the diversified needs of intelligent building management, indoor navigation and behavior analysis, etc.

[0063] Reference Figure 2 , the method comprises the following steps:

[0064] Step S201: end-side data acquisition

[0065] In the actual use stage, the robot system acquires multi-source sensor data in real time for comprehensively perceiving its own motion state and the information of the environment in which it is located. The acquired data includes point cloud sequences of the laser radar and motion sensor information such as angular velocity and acceleration output by the inertial measurement unit. The system has the functions of multi-source data time sequence alignment, online feature extraction and descriptor matching, can stably identify loops in challenging scenes such as changes in illumination and lack of visual features, realizes high robustness and high accuracy of loop detection and relocation, and ensures that complete and usable data basis can be obtained in different environments.

[0066] Step S202: data preprocessing

[0067] At the terminal device, the original multi-source sensor data acquired is subjected to multi-level preprocessing to improve the data quality and lay a foundation for subsequent feature extraction. Specifically, first, the time stamps of all sensor data are aligned to ensure the synchronization of multi-source data in the same time window; second, a variety of filtering algorithms (such as low-pass filtering, mean filtering, median filtering, Kalman filtering, etc.) are used to remove noise and outliers in the data and retain real and valid motion and environmental signals; finally, the preprocessed data is subjected to normalization processing to eliminate the differences in dimensions and amplitudes of different sensors, and the data stream is segmented based on a sliding window to provide stable and continuous data segments for subsequent feature calculation.

[0068] Step S203: point cloud feature extraction

[0069] Based on the preprocessed laser radar point cloud data, the terminal device locally executes a point cloud feature extraction process. This process aims to extract structural features with strong discriminability from the original point cloud to construct a robust data representation for the core loop detection link.

[0070] Specifically, the system first projects the three-dimensional point cloud into a polar coordinate system, and divides the point cloud space into uniform grid cells according to the preset radial and angular resolution. During the division process, the system ensures that each grid cell has sufficient point cloud density, usually requiring at least 10 or more valid points in each cell to ensure the stability of subsequent feature calculation.

[0071] Subsequently, the system performs geometric distribution analysis on the three-dimensional point set in each grid cell that meets the density requirement, and quantitatively describes the shape characteristics of the local area by calculating its covariance matrix. Among the three types of features, point, surface, and volume, surface features are proven to be the most discriminative and stable choice: on the one hand, there are a large number of planar structures (such as walls, floors, and tabletops) in indoor environments, and surface features have good universality; on the other hand, relative to point features that are easily affected by noise and volume features that have low discriminability, surface features achieve the best balance between computational complexity and discriminability.

[0072] For a grid cell containing N three-dimensional points {p1, p2,..., pn} and meeting the density requirement, the calculation process of its covariance matrix C is as follows: first, calculate the centroid μ of all points in the grid:

[0073]

[0074] Then, based on the offset vectors of each point relative to the centroid, calculate the 3x3 covariance matrix C:

[0075]

[0076] By eigenvalue decomposition, the system extracts surface features that satisfy the condition λ1≈λ2>>λ3 first, and records the corresponding feature vectors to represent the normal direction of the plane. For grid cells with insufficient points, the system will ignore the cell.

[0077] Finally, the system generates a geometric structure description based on the covariance analysis results of each grid cell, mainly using surface features. This grid-based surface feature extraction method based on density guarantee not only ensures the explicitness and stability of the features, but also provides high-quality input for subsequent descriptor generation. The stable features extracted through the above steps lay a reliable data foundation for subsequent efficient and accurate loop detection, significantly enhancing the system's perception and recognition ability of complex scene structures.

[0078] Step S204: Descriptor generation

[0079] Based on the point cloud features extracted in step S203, the system performs a descriptor generation process to construct a compact scene representation with rotation invariance. Specifically, the system uniformly divides the 360° scanning space into K sectors in the polar coordinate system with a preset angular resolution Δθ, where K = 360 / Δθ. Taking a common 2° resolution as an example, the system will generate K = 180 sectors, and divide them according to a preset radial distance resolution (such as 0.3 meters) and height interval (such as limited to the Z-axis -0.3 meters to 0.5 meters).

[0080] During the descriptor construction process, the system starts from the point cluster closest to the origin and processes each sector in turn in a clockwise direction. For each sector i, the system selects the closest valid point cluster and extracts its covariance matrix C i The eigenvalues (λ1 (i) , λ2 (i) , λ3 (i) ) and the radial distance d from the origin are used as the feature representation of the sector. The construction formula of the descriptor matrix D is as follows:

[0081] D = [f(θ1), f(θ2), …, f(θ K )] T

[0082] Where represents the feature vector of the i-th sector

[0083] Since each sector extracts a vector containing three eigenvalues and a radial distance instead of a single numerical value, the final global scene descriptor is a structured K x 4 matrix.

[0084] The construction of the descriptor ensures rotation invariance through two levels: first, the circumferential discretization of the polar coordinate system decouples the descriptor structure from the robot orientation; second, the encoding method based on the covariance eigenvalues itself has rotation invariance, further enhancing the system's adaptability to changes in viewing angle.

[0085] The K x 4 global scene descriptor generated by the above method not only retains the geometric structure information of the environment, but also provides more rich scene differentiation through multi-dimensional feature representation, laying a stable and comparable data foundation for efficient similarity matching in subsequent loop detection.

[0086] Step S205: Fast coarse screening based on IMU

[0087] Based on the motion information provided by the inertial measurement unit and the position prior in the historical map, the system performs a fast coarse screening process to significantly improve the matching efficiency of loop detection. The specific implementation process is as follows:

[0088] The system first calculates the relative motion change between adjacent key frames through IMU pre-integration. The pre-integrated observation values Δp, Δv, ΔR in the time interval [t k ,t k+1 ] are calculated by the following formula in the body coordinate system:

[0089]

[0090] where a t and ω t are accelerometer and gyroscope measurement values, b a and b g are corresponding biases, η a and η g are measurement noises, and ΔR represents a rotation matrix for converting a vector in the body coordinate system at the current time to the body coordinate system at the starting time.

[0091] Considering the random walk characteristics of the IMU bias, the system estimates the position uncertainty through error state covariance propagation. The covariance matrix P of the position error can be propagated in the following manner:

[0092]

[0093] where Φ is the error state transition matrix, and Q is the noise covariance matrix.

[0094] Based on the current position ^p t estimated by IMU pre-integration and its covariance ellipsoid P t , the system combines the stored key frame position information {p1, p2,..., p n} in the historical map to define the search region in the time and spatial dimensions:

[0095]

[0096] where d max is determined according to the position uncertainty, and Δt min is used to avoid matching key frames that are too close in time.

[0097] Through this screening mechanism based on IMU motion prior, the system reduces the number of candidate key frames from hundreds to tens of magnitude, greatly improving the efficiency of the subsequent descriptor matching stage while ensuring a high recall rate of true loop frames.

[0098] Step S206: Descriptor-based loop detection

[0099] Based on the scene descriptor generated in step S204, the system performs a loop detection process. The core advantage of this process is its inherent rotational invariance, which is directly derived from the unique design adopted during the construction of the descriptor: the effective point cluster closest to the origin of the coordinate system is taken as the starting reference point, and the features of each sector are encoded in a clockwise direction. This ordered construction method ensures that the structure of the generated descriptor remains consistent regardless of the orientation of the robot passing through the same location, fundamentally solving the problem of matching failure caused by rotational perspective.

[0100] In the implementation process, the system calculates the similarity between the Kx4-dimensional descriptor matrix of the current frame and the descriptor of the candidate historical key frame determined in step S205. The system uses the following distance measurement formula to evaluate the difference between the descriptors:

[0101]

[0102] where A and B represent the current frame descriptor matrix and the candidate historical frame descriptor matrix respectively, A i and B i represent the feature vectors corresponding to the i-th sector, ||·||2 represents the L2 norm of the vector, and K is the total number of sectors.

[0103] By calculating the average Euclidean distance between the descriptor matrices using the above formula, the system quantitatively evaluates the matching degree of the current scene and the historical scene. When the distance score is lower than the preset threshold, the candidate frame is determined as a potential loop.

[0104] This matching mechanism based on descriptor distance calculation, combined with the rotational invariance of the descriptor itself, can effectively deal with dynamic changes in the environment, observation noise, and changes in the orientation of the robot, providing high-quality candidate pairs for subsequent geometric verification. Through the above design, the system realizes a significant reduction in the risk of false matching caused by different directions of robot travel while maintaining high recall rate, ensuring the reliability of loop detection in various practical application scenarios.

[0105] Step S207: Semantic consistency verification and repositioning enhancement based on large models

[0106] After completing the IMU coarse screening and point cloud descriptor matching, the system further introduces a large model inference module to perform semantic verification and confidence correction on the candidate loop frames.

[0107] Specifically, the system inputs the point cloud descriptors, IMU time series features, and historical scene semantic labels of the current frame and the candidate frame into a large model (such as a fine-tuned multi-modal Transformer or visual language model). Based on the spatial relationship and semantic distribution features learned by the model, the semantic consistency score between the two frames is generated.

[0108] The score and the geometric matching confidence jointly constitute a joint similarity index:

[0109] S joint = a · S geo + (1 - a) · S sem

[0110] where S geo is the geometric matching score, S sem is the semantic consistency score output by the large model, and a is an adjustable weight coefficient.

[0111] Through this geometric-semantic joint evaluation mechanism, the system can effectively suppress the false loop closure caused by factors such as illumination change, structural occlusion, and scene dynamic change while maintaining high-precision matching, thereby improving the reliability and robustness of loop detection in complex environments.

[0112] In addition, the large model module can also participate in the subsequent factor graph global optimization as a semantic memory unit (Semantic Memory Unit), providing semantic prior and environmental association information for long-term unvisited areas, and further improving the positioning continuity and recoverability of the system in cross-period task execution.

[0113] Step S208: Add loop constraint to factor graph

[0114] Based on the reliable loop candidate pair confirmed in step S207, the system performs a constraint update process of the factor graph optimization framework. Specifically, the system first calculates the relative pose transformation matrix T ij between the current frame and the loop candidate frame through a point cloud registration algorithm. The transformation matrix contains the relative rotation R ij and translation t ij :

[0115]

[0116] Subsequently, the system models the relative pose transformation as a binary pose constraint edge and adds it to the global factor graph. The constraint edge connects the current robot pose node T i and the historical pose node T j corresponding to the loop key frame, forming a complete loop closure.

[0117] In the factor graph optimization framework, the error function e {ij} corresponding to the loop constraint is defined as:

[0118]

[0119] where log(·) is the Lie algebra mapping operation, which maps the transformation matrix to the tangent space. The covariance matrix ∑ij Uncertainty determination by point cloud matching.

[0120] The system solves the optimal pose by minimizing a weighted least squares cost function of all constraints:

[0121] X = argmin [∑‖e odom ‖ 2 +∑‖e loop ‖ 2 ]

[0122] where e odom represents the odometry constraint error, e loop represents the loop closure constraint error, Σ odom and∑ loop are the corresponding information matrices, respectively.

[0123] By introducing such loop closure constraints, the system constructs a closed optimization loop in the factor graph, providing key spatial correlation information for subsequent global optimization. This step effectively establishes geometric relationships between discrete-time pose estimates, significantly improving the global consistency of the system and laying the necessary mathematical foundation for eliminating odometry cumulative errors.

[0124] Step S209: output consistent trajectory

[0125] After the system completes the factor graph optimization solution locally on the terminal device, it outputs the robot motion trajectory after global correction. This trajectory integrates lidar observations, inertial measurement unit data, and loop detection constraints, significantly suppressing cumulative errors and ensuring the global consistency of the trajectory. The output includes the optimized robot pose sequence, trajectory confidence, and key frame association information, etc. The recognition results are packaged in a standardized data structure and can be directly used for robot global positioning correction, high-precision map construction and updating, including warehouse logistics autonomous navigation, service robot path planning, etc. scenarios, achieving efficient, real-time precise positioning and motion tracking.

[0126] With reference to Figure 3 , the loop detection repositioning device of the present application comprises:

[0127] A multi-source data acquisition module 410 is used to control the robot system to collect multi-source sensor data in real time during the actual use stage, so as to comprehensively perceive the motion state and the environment information. The raw data collected by the module includes the point cloud sequence of the lidar and the angular velocity and acceleration of the inertial measurement unit, etc. The motion sensor information provides a complete and usable multi-source data basis for the subsequent processing flow.

[0128] The data preprocessing module 420 is configured to perform multi-level preprocessing on the collected original multi-source sensor data. The module first implements timestamp alignment of various types of sensor data to ensure synchronization of multi-source data; then applies various filtering algorithms (including low-pass filtering, mean filtering, median filtering, Kalman filtering, etc.) to remove noise and outliers, and retains true motion and environmental signals; finally, the data is normalized and segmented based on a sliding window to provide stable and continuous data segments for subsequent feature extraction. The data preprocessing module 420 can dynamically adjust the preprocessing parameters according to different sensor characteristics and environmental changes to improve data quality and the robustness of subsequent analysis.

[0129] The feature extraction module 430 is configured to extract structural features in the point cloud based on the preprocessed laser radar point cloud data locally on the terminal device. The module projects the three-dimensional point cloud into the polar coordinate system, identifies the distribution of the point cloud in the polar coordinate space using a clustering algorithm, calculates the covariance matrix parameters of each distribution, and then generates statistical features representing local geometric structures to provide feature data with strong discrimination for subsequent descriptor generation.

[0130] The descriptor generation module 440 is configured to construct a scene descriptor with rotational invariance based on the point cloud features generated by the feature extraction module. The module divides the scanning space into K sectors under the polar coordinate system at a specified resolution, sequentially selects the nearest point cluster in each sector and encodes its radial distance and covariance features, and finally generates a structured Kx4-dimensional global scene descriptor to provide stable and reliable data representation for subsequent loop detection.

[0131] The loop detection module 450 is configured to achieve efficient loop detection based on inertial measurement unit information and scene descriptors. The module first defines the candidate search range in the time and space dimensions based on the motion prior information provided by the IMU and the historical map data, achieving fast coarse screening; then within the range, the scene descriptor similarity between the current frame and the candidate historical frame is compared to complete fine loop matching and recognition; finally, the inference capability of the large model is integrated to verify the semantic consistency and correct the confidence of the candidate results, effectively improving the detection reliability in complex scenarios. This multi-level detection mechanism not only ensures the operation efficiency of the system, but also significantly enhances the accuracy and robustness of loop detection.

[0132] The result output module 460 is configured to structure and output the pose information optimized by loop detection. The module encapsulates the optimized robot trajectory, key frame pose and their correlation into a standard data format, and outputs complete positioning results including global pose sequence, trajectory confidence and map matching information, providing real-time and reliable pose reference for upper-layer applications such as robot navigation, path planning and map construction.

[0133] It should be noted that the above functional modules are only representative embodiments of the present application, which can be flexibly adjusted or combined according to different device types and scene requirements in actual application, and any equivalent replacement or deformation made by those skilled in the art based on the content of the present application shall be covered within the protection scope of the present patent.

[0134] Figure 4 A more specific electronic device hardware structure schematic diagram provided by the embodiment is shown, which can include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040 and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030 and the communication interface 1040 are connected with each other through the bus 1050 for internal communication.

[0135] The processor 1010 can be implemented by a general CPU (Central Processing Unit), a microprocessor, an ASIC (Application Specific Integrated Circuit) or one or more integrated circuits, etc., for executing related programs to implement the technical solutions provided by the embodiments of the present application.

[0136] The memory 1020 can be implemented by a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs, and when the technical solutions provided by the embodiments of the present application are implemented by software or firmware, the related program codes are stored in the memory 1020 and called and executed by the processor 1010.

[0137] The input / output interface 1030 is used to connect input / output modules to realize information input and output. The input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. The input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.

[0138] The communication interface 1040 is used to connect a communication module (not shown in the figure) to realize the communication interaction between the device and other devices. The communication module can realize communication through a wired manner (such as USB, network cable, etc.) or a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).

[0139] The bus 1050 includes a path for transferring information between the various components (for example, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040) of the device.

[0140] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device can also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device can also only contain the components necessary to implement the embodiments of the present specification, and does not have to contain all the components shown in the figure.

[0141] The electronic device of the above embodiment is used to implement the corresponding loopback detection relocation method in any of the preceding embodiments, and has the beneficial effects of the corresponding method embodiments, which are not described here.

[0142] Based on the same inventive concept, the disclosure also provides a non-transitory computer readable storage medium corresponding to any of the above method embodiments, the non-transitory computer readable storage medium stores computer instructions for causing the computer to execute the loopback detection relocation method according to any of the above embodiments.

[0143] The computer readable medium of the present embodiment includes permanent and non-permanent, removable and non-removable media, which can be realized by any method or technology to store information. The information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0144] The above non-transitory computer readable storage medium can be any available medium or data storage device that can be accessed by a computer, including but not limited to magnetic storage (such as floppy disk, hard disk, magnetic tape, magneto-optical disk (MO) and the like), optical storage (such as CD, DVD, BD, HVD and the like), and semiconductor memory (such as ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid state disk (SSD)) and the like.

[0145] The computer program product of the above embodiments is configured to cause the computer and / or the processor to perform the loop detection relocation method according to any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0146] Based on the same inventive concept, the disclosure also provides a computer program product comprising computer program instructions corresponding to the loop detection relocation method according to any of the above embodiments. In some embodiments, the computer program instructions can be executed by one or more processors of a computer to cause the computer and / or the processor to perform the loop detection relocation method. Corresponding to the execution subject of each step in each embodiment of the loop detection relocation method, the processor performing the corresponding step can belong to the corresponding execution subject.

[0147] The computer program product of the above embodiments is configured to cause the computer and / or the processor to perform the loop detection relocation method according to any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0148] Those skilled in the art will understand that the embodiments of the disclosure can be implemented as a system, a method or a computer program product. Therefore, the disclosure can be embodied in a form of entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, which is generally referred to as "circuitry", "module" or "system". In addition, in some embodiments, the disclosure can also be embodied in the form of a computer program product in one or more computer readable media, which contains computer readable program codes.

[0149] Any combination of one or more computer readable medium can be employed. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any suitable combination of the above. More specific examples (non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus or device.

[0150] A computer readable signal medium can include a propagated data signal with computer executable instructions. A propagated signal can be an electromagnetic signal, an optical signal, and / or any other suitable type of signal. A computer readable medium can include any suitable medium that is accessible by a computer. Examples of a computer readable medium include a random access memory (RAM), a read-only memory (ROM), a compact disk (CD-ROM), a floppy disk, a hard disk, an optical disk, a magnetic tape, and / or another suitable medium. Combinations of the above should also be included within the scope of computer readable media.

[0151] The program code embodied on a computer readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire line, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0152] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Python, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0153] It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0154] These computer program instructions can also be stored in a computer readable medium that can direct a computer, a programmable data processing apparatus, and / or other

[0155] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0156] Further, although the operations of the method of the present disclosure are described in a particular, sequential order, this order is not meant to be a limitation. For example, some operations can be performed in an order different than that described. Further, some operations can be performed in parallel, in combination with, or in place of, one another. In addition, some operations can be omitted. Moreover, where appropriate, aspects of the disclosure can be implemented by various means, for example, hardware, software, firmware, or a combination thereof. In one embodiment, one or more computer programs might be embodied in machine-executable code, e.g., code for execution by a processor or a controller.

[0157] The flow diagrams and the block diagrams in the drawings are illustrative of possible architectures, functions, and operations for systems, methods, and computer program products according to various embodiments of the present disclosure. It will be understood that each block of the flow diagrams and / or block diagrams, and combinations of blocks in the flow diagrams and / or block diagrams, can be implemented by various means, such as hardware, software, firmware, or a combination thereof. Also, it will be appreciated that each block and / or combination of blocks can be implemented by special purpose hardware-based computer systems which are specifically programmed, constructed, and / or configured to carry out one or more computer processes. In this regard, each block in the flow diagrams and / or block diagrams can represent a module, segment, or portion of code which comprises one or more executable instructions to implement the specified logical function(s). It should also be noted that the functions of one or more blocks in the flow diagrams and / or block diagrams, and combinations of blocks in the flow diagrams and / or block diagrams, can also be implemented by special purpose hardware-based computer systems which are specifically programmed, constructed, and / or configured to carry out one or more computer processes.

[0158] It should be noted that, although the foregoing details refer to several modules or units of the device for action execution, such a division is not mandatory. Indeed, according to an embodiment of the application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into several modules or units embodied.

[0159] Those skilled in the art should understand that the above discussion of any embodiment is merely exemplary and is not intended to be limiting of the scope of the application (including the claims) which is intended to be limited only by the claims. The above embodiments or technical features among different embodiments can also be combined, steps can be implemented in any order, and there are many other variations of the aspects of the embodiments of the application as described above, which are not provided in detail in order to be brief. The embodiments of the application are not limited in scope by the sum of the embodiments disclosed because the embodiments of the application include any combination of the embodiments.

[0160] In addition, to simplify the description and discussion, and so as not to make the embodiments of the application difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components can or can not be shown in the provided drawings. Furthermore, devices can be shown in block diagram form in order to avoid making the embodiments of the application difficult to understand, and this also takes into account the fact that the details regarding the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the application are to be implemented (i.e., these details should be well within the understanding of those skilled in the art). Where specific details (e.g., circuitry) are set forth in order to describe an illustrative embodiment of the application, it should be apparent to those skilled in the art that the embodiment of the application can be practiced without these specific details or with an equivalent arrangement. Therefore, the description should not be construed as limiting, but merely as illustrative.

[0161] Although the application has been described in conjunction with specific embodiments thereof, numerous alternatives, modifications, and variations will be readily apparent to those skilled in the art. For example, other memory architectures (e.g., dynamic RAM (DRAM)) can use the embodiments discussed.

[0162] The embodiments of the application are intended to cover all such alternatives, modifications, and variations which fall within the broad scope of the appended claims. Accordingly, any one or more of the omitted, modified, equivalently replaced, improved, etc., should be included within the scope of the application.

[0163] While the principles of the disclosure have been described above in connection with specific embodiments, it is to be understood that this disclosure is not limited to the disclosed embodiments, but is intended to encompass various modifications and equivalent arrangements within the scope of the appended claims. The scope of the claims should be accorded the broadest interpretation so as to encompass all such modifications and equivalent arrangements as is permitted under the law.

Claims

1. A loop detection and relocalization method based on laser radar, inertial measurement unit and large model semantic inference, characterized in that, Comprise the following steps: (1) Multi-source data acquisition: Real-time acquisition of laser radar point cloud data and inertial measurement unit (IMU) motion data carried by mobile robots, and time synchronization, coordinate unification and noise filtering of the two types of data are performed to form a multi-source perception sequence that can be fused; (2) Feature extraction and fusion modeling: Local geometric features and global point cloud descriptors are constructed based on laser radar data, and feature correction is performed in combination with IMU motion constraint information; At the same time, based on the multi-modal weight adaptive mechanism, the fusion proportion of laser radar and IMU information is dynamically adjusted according to the environmental complexity, signal-to-noise ratio and self-motion acceleration to obtain stable and efficient feature representation; (3) Geometric loop detection: Calculate the point cloud descriptor similarity between the current frame and the historical key frame, screen the candidate loop frame and perform preliminary geometric matching to obtain the geometric matching confidence S geo ; (4) Large Model driven semantic consistency verification and relocalization enhancement: after completing geometric matching, the point cloud descriptors, IMU time sequence features and historical semantic labels of the candidate frame and the current frame are input into a large model (LLM) pre-trained by multiple modalities and fine-tuned by tasks, and the model performs cross-modal semantic reasoning to generate a semantic consistency score S between the two frames sem ; Comprehensive calculation of joint similarity index: S joint = a · S geo + (1 - a) · S sem Wherein, α is a dynamic adjustable weight coefficient; Through the index, the loop relationship is jointly judged to suppress false loops caused by changes in light, structural shielding or scene reconstruction, and to enhance the robustness and generalization of the system in complex environments; In addition, the large model constructs a semantic association embedding in the inference process, which is used as a semantic memory unit (Semantic Memory Unit) to provide semantic prior knowledge for long-term unvisited areas and support memory-based relocalization (Memory-based Relocalization); (5) Global semantic topology optimization and loop correction: When an effective loop is detected, geometric constraints, IMU constraints and semantic consistency constraints output by the large model are jointly introduced into the factor graph optimization framework; By adding semantic topological relationship constraints, geometric-inertial-semantic three-layer fusion loop optimization is realized to obtain a globally consistent pose solution; The optimized results are fed back to the navigation system in real time to realize adaptive error correction and long-term continuous positioning.

2. The method of claim 1, wherein, According to claim 1, wherein the large model is a Transformer structure model pre-trained on a multi-modal scene data set, which has the ability of spatial semantic understanding, semantic matching and temporal consistency modeling.

3. The method according to claim 1 or 2, characterized in that, The multi-modal fusion weight α is adaptively adjusted according to the environmental complexity, self-motion state and sensor signal-to-noise ratio to balance the influence of geometric matching and semantic matching on the final similarity.

4. The method according to any one of claims 1 to 3, characterized in that, The global optimization adopts a nonlinear least squares optimization method based on a factor graph, and simultaneously introduces semantic topological constraints to maintain the geometric consistency and semantic consistency of the global trajectory.

5. The method according to any one of claims 1 to 4, characterized in that, The loop detection and relocalization method further includes an explainability optimization and online self-learning mechanism, which forms a semantic memory bank by recording the semantic confidence of the large model and the historical matching results to realize adaptive error correction of the system in subsequent operation.

6. A loop detection relocation apparatus based on the method of any one of claims 1 to 5, characterized by Comprise: Multi-source data acquisition module: for synchronous acquisition of laser radar point cloud and IMU motion data; Feature extraction module: for generating point cloud descriptors and motion constraint features; Loop matching and semantic verification module: for performing geometric matching and semantic consistency verification based on the large model; Global optimization module: for fusing loop constraints and semantic priors, performing global factor graph optimization, and realizing pose correction.

7. An electronic device, comprising: A computer program product comprising a computer readable medium having stored thereon computer program means, the program being arranged such that, when it is executed by a processor, it causes the processor to carry out the loop detection relocation method according to any one of claims 1 to 5.

8. A non-transitory computer readable storage medium having stored thereon a computer program, the program, when executed by a computer, causing the computer to carry out the loop detection relocation method according to any one of claims 1 to 5.

9. A computer program product, characterised in that, A computer program product comprising computer program instructions arranged to cause a computer to carry out the loop detection relocation method according to any one of claims 1 to 5 when said instructions are run on a computer or robot terminal.