Autonomous lifelong SLAM method and system based on visual language model hidden space representation

The introspective lifelong SLAM method based on the latent space representation of the visual language model solves the problems of insufficient semantic understanding of SLAM technology in dynamic environments and real-time reconstruction of NeRF, and realizes high-precision, low-energy dynamic scene map construction and autonomous navigation.

CN120599495APending Publication Date: 2025-09-05TONGJI UNIV

Patent Information

Application Number
CN202510687988.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing SLAM technology lacks semantic understanding in dynamic open environments, resulting in interference from dynamic objects and cumulative errors. In addition, NeRF high-precision reconstruction is difficult to deploy in real time and cannot effectively combine semantic information with geometric optimization.

Method used

Through the latent space representation of the visual language model, a scene map with joint geometry and semantics is generated. Key frames are filtered using dynamic masks, and layered NeRF rendering is performed. Combined with latent space difference introspection optimization, cross-modal semantic alignment and layered rendering are achieved to construct a robust dynamic scene map.

Benefits of technology

Realize high-precision semantic map construction and continuous positioning in dynamic and complex environments, suppress cumulative drift, ensure long-term stability and adaptability, and provide high-precision and low-energy autonomous navigation support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599495A_ABST
    Figure CN120599495A_ABST
Patent Text Reader

Abstract

The invention relates to an introspection lifelong SLAM method and system based on visual language model hidden space representation, and the method comprises the steps: extracting a semantic tag based on an RGB-D image through a semantic encoder, and generating a scene map and a semantic topological graph based on the RGB-D image and the semantic tag; generating a dynamic mask based on the scene map, obtaining a dynamic mask coverage rate, and screening key frames with high static confidence values based on the coverage rate; calculating camera pose estimation corresponding to the key frame in real time, sampling the key frame to realize layering of the key frame, and performing layering rendering by using a NeRF model to obtain a virtual view; the hidden space difference degree of the virtual view and the corresponding real image is calculated, whether error introspection needs to be carried out or not is judged based on the hidden space difference degree, and the system is used for achieving the method. Compared with the prior art, the method has the advantages that open semantic reasoning of VLM, high-precision reconstruction of NeRF and real-time positioning of SLAM are combined, and positioning and mapping accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous robot navigation and three-dimensional environment understanding, and in particular to a self-reflective lifelong SLAM method and system based on latent space representation of a visual language model. Background Art

[0002] Simultaneous localization and mapping (SLAM) technology, which uses sensor data to achieve real-time positioning and environmental modeling, has made significant progress in static, structured scenes. However, when faced with dynamic, open environments and the need for long-term continuous operation, the limitations of traditional SLAM frameworks are becoming increasingly prominent. Traditional methods rely on geometric feature matching and probabilistic graph optimization, lacking a deep understanding of scene semantics. This makes it difficult for the constructed metric maps to support high-level task decisions. For example, dynamic objects can introduce feature matching outliers, causing pose estimation drift; while changes in illumination or scene structure during long-term operation can lead to cumulative errors, causing the map to gradually become ineffective. Existing improvements attempt to enhance robustness through semantic SLAM or dynamic object detection. However, the introduction of semantic labels is limited by a closed vocabulary and cannot adapt to unknown objects in open environments. Dynamic detection algorithms also struggle to balance efficiency and accuracy, and are particularly prone to false rejection or missed detection in complex dynamic scenes.

[0003] In recent years, Neural Radiance Fields (NeRF) and Visual Language Models (VLM) have provided new approaches to overcome these bottlenecks. NeRF achieves high-fidelity 3D reconstruction through implicit neural representations, but its high computational overhead and static scene assumptions limit its real-time application in SLAM. While VLM possesses cross-modal semantic alignment capabilities, it has not yet been effectively integrated into the SLAM geometry optimization process. Current research has yet to address the collaborative issues of multimodal semantic-geometric representation fusion, dynamic scene layered rendering, and lifelong adaptive learning. For example, while NeRF-based SLAM methods (such as iMAP) can construct dense maps, they cannot distinguish dynamic objects from static backgrounds and are difficult to deploy in a lightweight manner. Semantic SLAM, for example, uses Mask-SLAM. For example, the Chinese patent application "CN111402336A" provides a method that combines semantic SLAM technology with RGB-D cameras and deep neural networks to accurately estimate camera pose and construct semantic maps in dynamic environments. While the introduction of object-level semantics improves map accuracy to a certain extent, its closed labeling system leads to rigid scene understanding.

[0004] Therefore, how to combine the open semantic reasoning of VLM, the high-precision reconstruction of NeRF and the real-time positioning of SLAM is a technical problem that needs to be solved. Summary of the Invention

[0005] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and to provide an introspective lifelong SLAM method and system based on the latent space representation of the visual language model, and to provide a mapping method that combines the open semantic reasoning of VLM, the high-precision reconstruction of NeRF and the real-time positioning of SLAM.

[0006] The purpose of the present invention can be achieved by the following technical solutions:

[0007] According to a first aspect of the present invention, there is provided an introspective lifelong SLAM method based on latent space representation of a visual language model, comprising:

[0008] Acquire an RGB-D image of a target scene, extract semantic labels based on the RGB-D image using a semantic encoder, generate a scene map based on the RGB-D image and the semantic labels, and generate a semantic topology map based on the scene map;

[0009] Based on the scene map, a dynamic mask is generated using a latent space feature similarity measure and a moving average filter dynamic detection algorithm, a dynamic mask coverage is obtained, and key frames with high static confidence values ​​are selected based on the coverage;

[0010] Calculating the camera pose estimation corresponding to the key frames in real time, layering the key frames by sampling the key frames, and performing layered rendering using the NeRF model to obtain a virtual view; after the layering, a static base layer and a dynamic object layer are obtained;

[0011] The latent space difference between the virtual view and the corresponding real image is calculated, and based on the latent space difference, it is determined whether error introspection is required. If so, the scene map is optimized and corrected, otherwise no operation is performed.

[0012] As a preferred technical solution, the method for generating a scene map includes:

[0013] The RGB-D image includes an image sequence and a depth map, the image sequence is normalized, and the depth map is converted into a disparity map;

[0014] Extracting semantic labels based on the normalized image sequence and the disparity map;

[0015] Extracting image features of the normalized image sequence using the CLIP-ViT visual encoder, mapping the disparity map into a geometric code using variant residual convolution, and generating a semantic code based on the semantic label;

[0016] The cross-attention mechanism is used to fuse the image features, geometric encoding and semantic encoding to obtain the scene map, which is expressed as:

[0017]

[0018] Among them, f geo represents the geometric encoding; Q(·) represents the query; f sem represents semantic encoding; K(·) represents key; f rgb represents the image feature; V(·) represents the value; and d represents the scaling factor.

[0019] As a preferred technical solution, the method for generating a dynamic mask includes:

[0020] Get the scene maps of adjacent frames and calculate the cosine similarity matrix between them;

[0021] Performing sliding window mean filtering on the cosine similarity matrix to generate a dynamic probability map;

[0022] A mask threshold is set, and a binarization operation is performed on the dynamic probability map based on the mask threshold to obtain a dynamic mask; the mask threshold is updated in real time according to the dynamics of the scene.

[0023] As a preferred technical solution, the method for obtaining the key frame includes:

[0024] Obtaining an inlier ratio of feature matching between a current frame and a plurality of key frames closest to it, and if the inlier ratio is less than a first preset value, determining that the current frame is an unstable frame, otherwise it is a stable frame;

[0025] Calculating semantic significance weights based on semantic coding generated by the semantic tags;

[0026] It is determined whether the coverage is less than a second preset value, and stable frames whose coverage is less than the second preset value and whose semantic significance weight is greater than a third preset value are retained as the key frames.

[0027] As a preferred technical solution, the layering method includes:

[0028] Dividing the key frame into an N×N grid, and performing high-resolution sampling on the grid area marked by the dynamic mask to obtain a dynamic object layer;

[0029] The grid area except the area marked by the dynamic mask is sampled at a low resolution to obtain a static base layer.

[0030] As a preferred technical solution, the layered rendering method includes:

[0031] For the static base layer:

[0032] The static base layer is represented by a HashGrid-encoded voxel grid, wherein each voxel stores radiation field parameters of the static base layer;

[0033] Rendering is performed using a NeRF model based on the static base layer and the pose estimation to obtain a static rendered image; in each iterative optimization update of the NeRF model, sparsely sampling a preset proportion of voxels in the static base layer to optimize the radiation field parameters of the static rendered image;

[0034] For the dynamic object layer:

[0035] Generate a dynamic object proxy based on the semantic coding, and use the NeRF model to render and generate a dynamic rendering image in combination with the pose estimation; the proxy includes motion trajectory parameters and appearance coding;

[0036] When rendering using the NeRF model, the loss function is:

[0037]

[0038] Among them, r represents the sampling light; R represents the set of sampling rays in space; represents the color of the rendered light; C(r) represents the real pixel color; λ tv represents the regularization coefficient; represents the total variation regularization term used to suppress voxel noise.

[0039] Get virtual views based on static rendered images and dynamic renderings.

[0040] As a preferred technical solution, the layered rendering method further includes:

[0041] Calculate the Fisher information matrix of historical key frames;

[0042] When adding scene map data, the changes of important parameters are constrained based on the Fisher information matrix. The loss function of the constraint is:

[0043]

[0044] Among them, λ represents the weight adjustment factor; F i (·) represents the Fisher information matrix; θ i represents the parameters used to train new scenes; θ i,old represents the optimal parameters for constructing the old scene;

[0045] And in the constraint process, the learning rates of the new and old scenes are dynamically updated, and the learning rate of the new scene is greater than the learning rate of the old scene.

[0046] As a preferred technical solution, the error introspection includes:

[0047] Encoding the virtual view and the corresponding real image to obtain a virtual latent space encoding and a real latent space encoding;

[0048] The latent space difference is calculated based on the virtual latent space encoding and the real latent space encoding, and its expression is:

[0049]

[0050] Among them, Δ z represents the latent space difference; N represents the number of grids; represents the virtual latent space encoding of the i-th grid in the virtual view; represents the true latent space encoding of the i-th grid in the real image;

[0051] If the latent space difference is greater than the mask threshold in the dynamic mask, an introspection signal is triggered and local optimization is performed. The local optimization includes: matching key frames within a preset range based on the semantic topology graph and screening loop candidates, and dynamically updating the dynamic object layer;

[0052] If multiple introspection signals are triggered continuously, global optimization is performed, and the global optimization includes: resetting BA optimization parameters and performing global relocalization based on the loop candidates;

[0053] If the latent space difference is less than or equal to the mask threshold in the dynamic mask, no operation is performed.

[0054] As a preferred technical solution, the global relocation method includes:

[0055] For frames whose latent space difference is greater than the mask threshold in the dynamic mask, obtaining corresponding semantic codes;

[0056] Obtain a corresponding semantic topology map based on the corresponding semantic coding, and use the VLM to generate a scene text description based on the semantic topology map;

[0057] Determine the similarity between the scene description and each loop candidate, select the topological point with the highest similarity as the relocation point, and feed back the semantic encoder based on the scene text description.

[0058] According to a second aspect of the present invention, a self-reflective lifelong SLAM system based on latent space representation of a visual language model is provided to implement the above method.

[0059] Compared with the existing technology, the present invention integrates semantic information into the map constructed based on RGB-D images to form a joint representation, adopts layered rendering based on the joint representation to realize progressive map modeling, and realizes robust perception and long-term reliable modeling of dynamic scenes through cross-modal semantic alignment and layered NeRF collaborative optimization; and in the process of progressive map modeling, high-precision reconstruction of NeRF and real-time positioning of SLAM are realized by constructing a lightweight lifelong learning mechanism and a dynamic introspection detection mechanism; more specifically, the latent space difference is introduced as the trigger signal of error introspection when constructing the dynamic introspection detection, and the joint representation is optimized at the semantic level, thereby forming a closed-loop feedback mechanism of inverse rendering, guiding NeRF to achieve higher-precision reconstruction, realizing open reasoning of the visual language model while suppressing the cumulative drift in long-term operation; introducing the Fisher information matrix of historical key frames for the changes of important parameters in the reconstruction process, preventing the old scene data from being forgotten when receiving new scene data, avoiding repeated modeling, and improving the adaptive robustness of sensor degradation. In summary, the collaborative design of the present invention, based on multimodal tight coupling, semantically driven introspection and lightweight lifelong learning, solves the problem in the existing technology that the open semantic reasoning of VLM, the high-precision reconstruction of NeRF and the real-time positioning of SLAM cannot be effectively combined, ensuring that NeRF can achieve long-term, accurate and stable map reconstruction in dynamic and complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 is a flow chart of the present invention;

[0061] Figure 2 Schematic diagram of the cross-modal latent space alignment framework of the present invention;

[0062] Figure 3 A flowchart of the lifetime modeling of the layered differentiable radiation field of the present invention;

[0063] Figure 4 This is a flowchart of the error self-reflection correction of the present invention. DETAILED DESCRIPTION

[0064] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0065] Unless otherwise defined, the technical or scientific terms used in this application should have the ordinary meaning understood by a person of ordinary skill in the technical field to which this application belongs. The words "one", "a", "the" and the like used in this application do not indicate a limit on quantity and may indicate the singular or plural. The terms "include", "comprise", "have" and any variations thereof used in this application are intended to cover non-exclusive inclusions; for example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units that are not listed, or may also include other steps or units that are inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The word "multiple" used in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0066] Example 1

[0067] Traditional SLAM technology relies on geometric feature matching for positioning and mapping, but faces significant bottlenecks in dynamic and complex scenes. On the one hand, the lack of semantic understanding makes it difficult to effectively remove interference from dynamic objects, and long-term operation can easily lead to cumulative errors due to changes in illumination and evolving scene structures. On the other hand, existing semantic SLAM relies on closed-vocabulary detection models, which are not suitable for identifying unknown objects in open environments. High-precision reconstruction methods based on NeRF are difficult to deploy in real time due to excessive computational load and do not address the coupled optimization problem of dynamic objects and static scenes. Achieving the coordinated optimization of semantic perception, dynamic adaptation, and lifelong learning under resource-constrained conditions has become a key challenge in improving the robustness and practicality of SLAM.

[0068] In order to solve the above-mentioned technical problems, the present invention provides a self-reflective lifelong SLAM method based on the latent space representation of the visual language model, which realizes robust perception and long-term reliable modeling of dynamic scenes through cross-modal semantic alignment and hierarchical NeRF collaborative optimization. The system uses a pre-trained visual language model to map RGB-D data to a unified latent space, generates a joint representation that integrates geometry and semantics, and combines dynamic masks to screen high-confidence static key frames; through lightweight hierarchical NeRF reverse rendering of virtual views, it triggers self-reflective error correction driven by latent space difference and suppresses cumulative drift in long-term operation; adopts an elastic weight solidification strategy to constrain parameter updates, and combines semantic topology maps to achieve cross-modal relocation of low-texture areas. This method realizes high-precision semantic map construction and continuous positioning in dynamic and complex environments, ensures long-term stability and adaptability to scene changes through a self-reflective lifelong learning mechanism, and provides high-precision, low-energy technical support for autonomous navigation in complex scenes.

[0069] The detailed process is as follows Figure 1 Shown, including:

[0070] S1. Generate semantic topology graph.

[0071] S11. Collect RGB-D images of the target scene through sensors deployed on the UAV. In detail, the RGB-D images include image sequences and depth maps.

[0072] S12. Normalize the image sequence. Specifically, resize it to 224×224 and normalize the pixel values ​​to [-1, 1]. Convert the depth map to a disparity map to enhance the sensitivity of geometric features.

[0073] S13, based on the normalized image sequence and disparity map, use the lightweight target detection model to extract the object-level semantic label L t .

[0074] S14. Use CLIP-ViT’s visual encoder to extract image features of the normalized image sequence Mapping disparity maps to geometric codes using ResNet-18 variant residual convolution Generate semantic encoding based on semantic tags

[0075] S15. Use the cross-attention mechanism to fuse image features, geometric encoding, and semantic encoding to obtain a joint representation, namely the scene map. The process is as follows: Figure 2 As shown, the query is obtained based on image features, the key is obtained based on geometric encoding, and the value is obtained based on semantic encoding. The fusion expression is:

[0076]

[0077] Among them, f geo represents the geometric encoding; Q(·) represents the query; f sem represents semantic encoding; K(·) represents key; f rgb represents the image feature; V(·) represents the value; d represents the scaling factor, and in this embodiment, d=512 is set; and Q(·), K(·) and V(·) are all learnable linear projections.

[0078] S16. Generate a scene map based on the RGB-D image and the semantic labels, and generate a semantic topology map based on the scene map.

[0079] S2. Based on the scene map, a dynamic mask is generated using the latent space feature similarity measurement and the moving average filter dynamic detection algorithm. The dynamic mask coverage is obtained and key frames with high static confidence values ​​are selected based on the coverage.

[0080] S21. Dynamic mask generation:

[0081] S212: Obtain scene maps of adjacent frames and calculate the cosine similarity matrix between the two, which is expressed as:

[0082]

[0083] Among them, Z t Represents the scene map corresponding to the t-th frame image, Z t-1 Represents the scene map corresponding to the t-1th frame image, where H and W are resolutions.

[0084] S212. Perform sliding window mean filtering on the cosine similarity matrix to generate a dynamic probability map, where the size of the sliding window is 3×3, and its expression is:

[0085] P d =1-Sigmoid(S t,t-1 ),

[0086] Among them, S t,t-1 Represents the cosine similarity matrix of adjacent frame maps.

[0087] S213, setting mask threshold θ d ,Based on the mask threshold, the dynamic probability map is binarized to obtain a dynamic mask, and the mask threshold is updated in real time with the dynamics of the scene.

[0088] S22. Key frame selection:

[0089] S221. Obtain the inlier ratio of the feature matching between the current frame and its most adjacent key frames (in this embodiment, five adjacent key frames are selected). If the inlier ratio is less than 60%, the current frame is determined to be an unstable frame, otherwise it is a stable frame.

[0090] S222. Calculate semantic saliency weight based on the semantic coding generated by the semantic tag.

[0091] S223. Determine whether the coverage is less than 10%, retain stable frames with a coverage less than 10% and a semantic significance weight greater than 90%, and the inlier ratio of the stable frame is greater than 70%, which is a key frame.

[0092] S3. Calculate the camera pose estimation corresponding to the key frames in real time, implement key frame layering by sampling the key frames, and use the NeRF model for layered rendering to obtain the virtual view.

[0093] S31. Divide the key frame into 8×8 grids, and perform high-resolution (512×512) sampling on the grid area marked by the dynamic mask to obtain a dynamic object layer.

[0094] S32. Perform low-resolution (128×128) sampling on the grid area except the area marked by the dynamic mask to obtain a static base layer.

[0095] S33, layered rendering, its process is as follows Figure 3 Shown, including:

[0096] For a static base layer:

[0097] The static base layer is represented by a HashGrid-encoded voxel grid, and each voxel stores the radiation field parameters of the static base layer.

[0098] The NeRF model is used to render based on the static base layer and pose estimation to obtain a static rendered image; in each iterative optimization update of the NeRF model, 10% of the voxels in the static base layer are sparsely sampled to optimize the radiation field parameters of the static rendered image.

[0099] For dynamic object layers:

[0100] Generate dynamic object proxies based on semantic coding, and use NeRF model to render dynamic renderings based on pose estimation. The proxies include motion trajectory parameters and appearance coding.

[0101] When using the NeRF model for rendering, the loss function is:

[0102]

[0103] Among them, r represents the sampling light; R represents the set of sampling rays in space; represents the color of the rendered light; C(r) represents the real pixel color; λ tv represents the regularization coefficient; represents the total variation regularization term used to suppress voxel noise.

[0104] In addition, if new scene map data is added, the Fisher information matrix of the historical keyframes is calculated, and the changes of important parameters are constrained based on the Fisher information matrix. The constraint loss function is:

[0105]

[0106] Among them, λ represents the weight adjustment factor; F i (·) represents the Fisher information matrix; θ i represents the parameters used to train new scenes; θ i,old represents the optimal parameters for constructing the old scene;

[0107] In the constraint process, the learning rates of the new and old scenes are dynamically updated, and the learning rate of the new scene is greater than the learning rate of the old scene. In this embodiment, the learning rate of the new scene is set to 1e-3 and the learning rate of the old scene is set to 1e-5.

[0108] S34. Obtain a virtual view based on the static rendering image and the dynamic rendering image.

[0109] The construction of the virtual view is achieved through the above process. In addition, during the construction process, the error introspection provided in step S4 is used for optimization, so that the layered rendering process forms a closed-loop feedback of reverse rendering, and only the dynamic object layer is optimized in the reverse rendering optimization process.

[0110] S4. Calculate the latent space difference between the virtual view and the corresponding real image, and determine whether error introspection is needed based on the latent space difference. If so, optimize and correct the scene map, otherwise no operation is performed. The process is as follows: Figure 4 shown.

[0111] S41. Encode the virtual view and the corresponding real image to obtain a virtual latent space code and a real latent space code.

[0112] S42. Calculate the latent space difference based on the virtual latent space coding and the real latent space coding, and its expression is:

[0113]

[0114] Among them, Δ z represents the latent space difference; N represents the number of grids; represents the virtual latent space encoding of the i-th grid in the virtual view; represents the true latent space encoding of the i-th grid in the real image.

[0115] S43. If the latent space difference is greater than the mask threshold in the dynamic mask, that is, Δz>θ z, then a self-reflection signal is triggered and local optimization is performed, that is, key frames within a radius of 5m are matched based on the semantic topology graph and loop candidates are screened out, and the dynamic object layer is dynamically updated; otherwise, that is, Δz≤θ z Execute step S45.

[0116] S44, if the introspection signal is triggered multiple times continuously, global optimization is performed, which includes: resetting the BA optimization parameters and performing global relocation based on the loop candidate; otherwise, Δz≤θ z Execute step S45.

[0117] Among them, global relocation includes:

[0118] S441. For frames whose latent space difference is greater than the mask threshold in the dynamic mask, obtain corresponding semantic codes.

[0119] S442: Obtain a corresponding semantic topology map based on the corresponding semantic code, and generate a scene text description using the semantic topology map using the VLM.

[0120] S443: Determine the similarity between the scene description and each loop candidate, select the topological point with the highest similarity as the relocation point, and feed back the semantic encoder based on the scene text description.

[0121] Through the above global optimization, global relocation of low-texture areas is achieved in the form of cross-modal retrieval, which improves the accuracy of NeRF model reconstruction.

[0122] S45. No operation.

[0123] This embodiment also provides an introspective lifelong SLAM system based on latent space representation of a visual language model. The system is used to implement the above method, and its application object is a drone. The system includes a sensor, and the sensor includes at least a high-precision RGB-D camera. Therefore, the information collected by the system includes at least an RGB-D image.

[0124] In addition, the system provided by the present invention also includes a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.

[0125] Many components in a device are connected to the I / O interface, including: input units, such as a keyboard and mouse; output units, such as various types of displays and speakers; storage units, such as magnetic disks and optical disks; and communication units, such as network cards, modems, and wireless communication transceivers. The communication unit allows the device to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks.

[0126] The processing unit performs the various methods and processes described above, such as methods S1 to S4. For example, in some embodiments, methods S1 to S4 can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed on the device via a ROM and / or a communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of methods S1 to S4 described above can be performed. Alternatively, in other embodiments, the CPU can be configured to execute methods S1 to S4 by any other appropriate means (for example, by means of firmware).

[0127] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0128] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0129] In the context of the present invention, machine-readable medium can be a tangible medium that can contain or store a program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0130] Example 2

[0131] In this embodiment, taking the detection of underground mine tunnels by drones as an example, the underground mine tunnels are reconstructed using the method and system provided in the embodiment.

[0132] In detail, the high-precision RGB-D camera onboard the drone captures color images and depth information of the environment in real time, while a lightweight target detection model is used to extract semantic labels for objects in the image, such as "conveyor belt" and "ore pile." These multimodal data are fed into a pre-trained visual language model, where a cross-modal attention mechanism is used to map geometric depth features and open-lexicon semantic features to a unified latent space, generating a joint representation vector that integrates geometric structure and semantic attributes. For example, the "metal support" in the mine tunnel is not only encoded as three-dimensional point cloud coordinates, but its functional attributes as a "rigid support structure" are also explicitly expressed through the semantic encoding layer, providing interpretable feature primitives for subsequent dynamic reasoning and map construction.

[0133] When a drone encounters interference from mobile devices in a mine tunnel, the system uses temporal consistency analysis based on the joint representation of the latent space to calculate the similarity differences between feature vectors in adjacent frames, identify dynamic regions, and generate binary masks. For example, the motion trajectory of a transport vehicle is represented in the latent space as regions of wildly fluctuating feature similarity. The system dynamically suppresses the feature matching weights in these regions accordingly. Furthermore, by combining the stability scores of static infrastructure such as tracks and lighting in the semantic encoding, the system selects frames with stable backgrounds and only feeds these frames into the backend optimization module, avoiding pose estimation drift caused by dynamic interference.

[0134] To account for changes in tunnel structure over time, the system decomposes keyframes into a static base layer and a dynamic object layer through sampling. This layer performs hierarchical modeling and combines the modeling results to create a virtual view. The static layer uses a high-resolution voxel grid to describe permanent structures, while the dynamic layer uses latent space semantic encoding to drive a deformable radiation field. When new scene map data is added, historical scene knowledge is retained through a weighted solidification strategy. For example, the support structure parameters of an old tunnel are locked, while the geometric and semantic features of newly added landslide areas are incrementally encoded into the dynamic layer, enabling the continuous evolution of the map without losing historical information.

[0135] If the error introspection mechanism is triggered during reconstruction—that is, if the difference between the latent space encoding of the virtual view of the mine tunnel rendered using the lightweight layered NeRF model and the real-world image exceeds a preset threshold during continuous drone operation—accumulated error is detected, triggering a local loop detection mechanism. The system then rapidly matches historical keyframes based on node relationships in the semantic topology graph, corrects the current pose using a graph optimization algorithm, and simultaneously updates the dynamic object layer. If the abnormal state persists for a long time, the local voxel parameters of the NeRF model, known as the build-up optimization parameters, are updated, and global relocalization is performed. At this point, if traditional geometric feature matching fails in low-texture areas of the mine tunnel, the system uses the latent space semantic encoding to match it against the pre-stored semantic topology graph. For example, the relative position of the drone is inferred by comparing the semantic relationship of "ventilation duct on the left" in the current frame with the node connection rules of "ventilation duct-main tunnel" in the topology graph. If the introspection module continuously reports positioning anomalies, the system activates a cross-modal relocalization mode: the VLM generates a textual description of the current scene, retrieves matching historical map nodes using the text-semantic graph, and combines sparse geometry verification to restore global positioning.

[0136] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A self-reflective lifelong SLAM method based on latent space representation of visual language model, characterized by: The method includes: Acquire an RGB-D image of a target scene, extract semantic labels based on the RGB-D image using a semantic encoder, generate a scene map based on the RGB-D image and the semantic labels, and generate a semantic topology map based on the scene map; Based on the scene map, a dynamic mask is generated using a latent space feature similarity measure and a moving average filter dynamic detection algorithm, a dynamic mask coverage is obtained, and key frames with high static confidence values ​​are selected based on the coverage; Calculating the camera pose estimation corresponding to the key frames in real time, layering the key frames by sampling the key frames, and performing layered rendering using the NeRF model to obtain a virtual view; after the layering, a static base layer and a dynamic object layer are obtained; The latent space difference between the virtual view and the corresponding real image is calculated, and based on the latent space difference, it is determined whether error introspection is required. If so, the scene map is optimized and corrected, otherwise no operation is performed.

2. The introspective lifelong SLAM method based on latent space representation of a visual language model according to claim 1, characterized in that The method for generating a scene map includes: The RGB-D image includes an image sequence and a depth map, the image sequence is normalized, and the depth map is converted into a disparity map; Extracting semantic labels based on the normalized image sequence and the disparity map; Extracting image features of the normalized image sequence using the CLIP-ViT visual encoder, mapping the disparity map into a geometric code using variant residual convolution, and generating a semantic code based on the semantic label; The cross-attention mechanism is used to fuse the image features, geometric encoding and semantic encoding to obtain the scene map, which is expressed as: Among them, f geo represents the geometric encoding; Q(·) represents the query; f sem represents semantic encoding; K(·) represents key; f rgb represents the image feature; V(·) represents the value; and d represents the scaling factor.

3. The introspective lifelong SLAM method based on latent space representation of a visual language model according to claim 1, characterized in that The method for generating a dynamic mask includes: Get the scene maps of adjacent frames and calculate the cosine similarity matrix between them; Performing sliding window mean filtering on the cosine similarity matrix to generate a dynamic probability map; A mask threshold is set, and a binarization operation is performed on the dynamic probability map based on the mask threshold to obtain a dynamic mask; the mask threshold is updated in real time according to the dynamics of the scene.

4. The introspective lifelong SLAM method based on latent space representation of a visual language model according to claim 1, characterized in that The method for obtaining the key frame includes: Obtaining an inlier ratio of feature matching between a current frame and a plurality of key frames closest to it, and if the inlier ratio is less than a first preset value, determining that the current frame is an unstable frame, otherwise it is a stable frame; Calculating semantic significance weights based on semantic coding generated by the semantic tags; It is determined whether the coverage is less than a second preset value, and stable frames whose coverage is less than the second preset value and whose semantic significance weight is greater than a third preset value are retained as the key frames.

5. The introspective lifelong SLAM method based on latent space representation of a visual language model according to claim 1, characterized in that The layering method includes: Dividing the key frame into an N×N grid, and performing high-resolution sampling on the grid area marked by the dynamic mask to obtain a dynamic object layer; The grid area except the area marked by the dynamic mask is sampled at a low resolution to obtain a static base layer.

6. The introspective lifelong SLAM method based on latent space representation of a visual language model according to claim 5, characterized in that The layered rendering method includes: For the static base layer: The static base layer is represented by a HashGrid-encoded voxel grid, wherein each voxel stores radiation field parameters of the static base layer; Rendering is performed using a NeRF model based on the static base layer and the pose estimation to obtain a static rendered image; in each iterative optimization update of the NeRF model, sparsely sampling a preset proportion of voxels in the static base layer to optimize the radiation field parameters of the static rendered image; For the dynamic object layer: Generate a dynamic object proxy based on the semantic coding, and use the NeRF model to render and generate a dynamic rendering image in combination with the pose estimation; the proxy includes motion trajectory parameters and appearance coding; When rendering using the NeRF model, the loss function is: Among them, r represents the sampling light; R represents the set of sampling rays in space; represents the color of the rendered light; C(r) represents the real pixel color; λ tv represents the regularization coefficient; represents the total variation regularization term used to suppress voxel noise. Get virtual views based on static rendered images and dynamic renderings.

7. The introspective lifelong SLAM method based on latent space representation of a visual language model according to claim 6, characterized in that The layered rendering method further includes: Calculate the Fisher information matrix of historical key frames; When adding scene map data, the changes of important parameters are constrained based on the Fisher information matrix. The loss function of the constraint is: Among them, λ represents the weight adjustment factor; F i (·) represents the Fisher information matrix; θ i represents the parameters used to train new scenes; θ i,old represents the optimal parameters for constructing the old scene; And in the constraint process, the learning rates of the new and old scenes are dynamically updated, and the learning rate of the new scene is greater than the learning rate of the old scene.

8. The introspective lifelong SLAM method based on latent space representation of a visual language model according to claim 1, characterized in that The error introspection includes: Encoding the virtual view and the corresponding real image to obtain a virtual latent space encoding and a real latent space encoding; The latent space difference is calculated based on the virtual latent space encoding and the real latent space encoding, and its expression is: Among them, Δ z represents the latent space difference; N represents the number of grids; represents the virtual latent space encoding of the i-th grid in the virtual view; represents the true latent space encoding of the i-th grid in the real image; If the latent space difference is greater than the mask threshold in the dynamic mask, an introspection signal is triggered and local optimization is performed. The local optimization includes: matching key frames within a preset range based on the semantic topology graph and screening loop candidates, and dynamically updating the dynamic object layer; If multiple introspection signals are triggered continuously, global optimization is performed, and the global optimization includes: resetting BA optimization parameters and performing global relocalization based on the loop candidates; If the latent space difference is less than or equal to the mask threshold in the dynamic mask, no operation is performed.

9. The introspective lifelong SLAM method based on latent space representation of a visual language model according to claim 8, characterized in that The global relocation method includes: For frames whose latent space difference is greater than the mask threshold in the dynamic mask, obtaining corresponding semantic codes; Obtain a corresponding semantic topology map based on the corresponding semantic coding, and use the VLM to generate a scene text description based on the semantic topology map; Determine the similarity between the scene description and each loop candidate, select the topological point with the highest similarity as the relocation point, and feed back the semantic encoder based on the scene text description.

10. A self-reflective lifelong SLAM system based on latent space representation of visual language model, characterized by: The system is used to implement the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Dynamic environment camera pose estimation and semantic map construction method based on semantic SLAM

    CN111402336A

Cited By

  • Navigation map processing method and device, electronic equipment and computer storage medium

    CN121274940A

  • Virtual-real combined scene space interaction self-adaptive online processing method

    CN121564296A

  • Mixed reality scene space interaction adaptive online processing method

    CN121564296B

  • Real-time dynamic map construction method and system based on neural radiation field and medium

    CN121861224A

  • Industrial defect generation method based on physical constraint and language prompt

    CN121962793A