Methods and systems for generating object node information in 3D scene atlases

By generating multi-level 3D scene maps and using the fusion of visual language models and large language models to correct the description of observation frames, the problem of inaccurate node information is solved, and the environmental perception and autonomous navigation capabilities of the intelligent agent are improved.

CN118587379BActive Publication Date: 2025-10-31SHANDONG NEW GENERATION INFORMATION IND TECH RES INST CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411017196.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2025-10-31
Estimated Expiration
2044-07-29

AI Technical Summary

Technical Problem

Existing technologies for constructing 3D scene maps suffer from insufficient richness and accuracy of node information, leading to errors or incompleteness in descriptions and affecting the expressive power and application scope of scene maps.

Method used

By mining spatiotemporal clues in the scene perception information of intelligent agents, multi-level 3D scene maps are generated using visual language models and large language models. The corrected observation frame descriptions are then fused to generate accurate object node information.

Benefits of technology

It significantly improves the quality and application value of node information in 3D scene maps, enhances the environmental perception, autonomous navigation and intelligent decision-making capabilities of intelligent agents, and reduces the impact of uncertainty and illusion problems in the information provided by large models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118587379B_ABST
    Figure CN118587379B_ABST
Patent Text Reader

Abstract

This invention provides a method and system for generating object node information in a 3D scene atlas, belonging to the field of intelligent agent scene perception. The method includes: generating attribute description information for each observation frame related to object nodes using a visual language model; obtaining the text encoding vector of each frame description and the image encoding vector of the current frame image; determining the cosine similarity between the text encoding vector and the image encoding vector, and selecting the frame description with the highest similarity as the corrected description for the current frame; fusing the corrected observation frame descriptions and obtaining the node information of the object nodes using a large language model. By mining spatiotemporal clues in the intelligent agent's scene perception information, the node information in the 3D scene atlas is generated and improved, reducing the impact of erroneous descriptions and helping the intelligent agent achieve accurate scene perception, thereby enhancing the intelligent agent's capabilities in environmental perception, autonomous navigation, intelligent decision-making, and adaptation to environmental changes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent agent scene perception technology, specifically to a method and system for generating object node information in a three-dimensional scene map. Background Technology

[0002] With the rapid development of mobile robot technology and the fast expansion of the industry, people's work and lifestyles are undergoing tremendous changes. Currently, mobile robots are ubiquitous, from industrial production lines to commercial buildings and living spaces, their presence increasingly prominent. These robots perform diverse functions, including but not limited to providing guidance services, delivering goods, cleaning, and conducting security patrols. They not only provide a powerful impetus for economic and social progress but also bring unprecedented convenience to every detail of life. These intelligent mobile robots are gradually becoming an indispensable part of life and work.

[0003] To enable mobile robots to navigate and perceive their environment independently, cutting-edge perception and navigation algorithms and sensor technologies are essential. By using sensors such as LiDAR, various types of cameras (e.g., monocular, binocular, RGB-D), and inertial measurement units, combined with SLAM (Simultaneous Localization and Mapping) technology, a metric map of the scene can be created. This allows the robot to have a basic understanding of its surroundings, at least enabling it to identify the locations of obstacles and navigable areas.

[0004] However, in the construction of 3D scene atlases, the richness and accuracy of node information are crucial to the application of the entire atlas, directly affecting its expressive power and application scope. Accurate node information can provide more precise search results and enhance the depth of scene understanding. Existing technologies rely on single-frame images to generate node descriptions, which ignores the dynamic information and temporal correlations in the intelligent perception process. Furthermore, limitations inherent in large-scale model technology lead to potentially erroneous or incomplete descriptions, resulting in poor interactivity and low levels of intelligence. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for generating object node information in a three-dimensional scene atlas, which can solve all or at least part of the technical problems existing in the prior art.

[0006] To achieve the above objectives, embodiments of the present invention provide a method for generating object node information in a three-dimensional scene atlas. This method constructs a three-dimensional scene atlas containing multiple levels for a target environment. The three-dimensional scene atlas is composed of nodes and edges across multiple levels, and each node contains attribute description information. The method for generating object node information in the three-dimensional scene atlas includes:

[0007] For observation frames related to object nodes, attribute description information for each frame is generated using a visual language model;

[0008] Obtain the text encoding vector of each frame description and the image encoding vector of the current frame image respectively;

[0009] Determine the cosine similarity between the text encoding vector and the image encoding vector, and select the frame description with the highest similarity as the corrected description for the current frame;

[0010] The observed frame descriptions are fused and corrected, and the node information of the object nodes is obtained using a large language model.

[0011] Optionally, a multi-level 3D scene atlas can be constructed for the target environment, including:

[0012] Identify target entities in the scene, and add descriptive information to the target entities using a large language model and a visual language model;

[0013] Extract the spatial information of the target entity, and construct a hierarchical three-dimensional scene map based on the target entity, its descriptive information, and the spatial information.

[0014] Optionally, each edge in the three-dimensional scene graph connects two nodes or one node in the same layer with nodes in the upper or lower layers of its own layer.

[0015] Optionally, attribute description information for each frame can be generated according to the following formula:

[0016] , ;

[0017] In the formula, Represents a collection of attribute descriptions. Representing a visual language model, Represents the relationship between object nodes All relevant observation frames, This indicates a prompt word for this node.

[0018] Optionally, the text encoding vector describing each frame and the image encoding vector of the current frame image are obtained separately, including:

[0019] Using a text encoder and an image encoder, the text encoding vector describing each frame and the image encoding vector of the current frame image are obtained respectively.

[0020] Optionally, the cosine similarity between the text encoding vector and the image encoding vector is determined according to the following formula, and the frame description with the highest similarity is selected as the corrected description for the current frame:

[0021] ;

[0022] ;

[0023] In the formula, Represents the calculation of cosine similarity. Represents the image encoding vector. Represents a text encoding vector. This indicates the corrected description for the current frame. This indicates the description of the frame with the highest similarity.

[0024] Optionally, the node information of the object nodes can be obtained according to the following formula:

[0025]

[0026] In the formula, Node information representing object nodes. This indicates the corrected description of the observation frame. It guides the large model based on Each observation frame contains prompts describing the node information of the object's nodes.

[0027] On the other hand, the present invention also provides a system for generating object node information in a three-dimensional scene atlas, which constructs a three-dimensional scene atlas containing multiple levels for a target environment. The three-dimensional scene atlas is composed of nodes and edges at multiple levels, and each node contains attribute description information. The system for generating object node information in the three-dimensional scene atlas includes:

[0028] The generation unit is used to generate attribute description information for each observation frame associated with an object node using a visual language model. The attribute description information of the object node includes attribute description and 3D point cloud.

[0029] The encoding vector acquisition unit is used to acquire the text encoding vector describing each frame and the image encoding vector of the current frame image, respectively.

[0030] The cosine similarity calculation unit is used to determine the cosine similarity between the text encoding vector and the image encoding vector, and select the frame description with the highest similarity as the corrected description of the current frame.

[0031] The fusion unit is used to fuse the corrected observation frame descriptions and obtain the node information of object nodes using a large language model.

[0032] On the other hand, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method for generating object node information in a three-dimensional scene atlas described above.

[0033] On the other hand, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the steps of the method for generating object node information in a three-dimensional scene atlas as described above.

[0034] By mining spatiotemporal cues from the scene perception information of intelligent agents, node information in 3D scene atlases is generated and improved, reducing the impact of erroneous descriptions and significantly enhancing the information quality and application value of nodes in the 3D scene atlas. This helps intelligent agents achieve deep and accurate scene perception, thereby improving their capabilities in environmental perception, autonomous navigation, intelligent decision-making, and adaptation to environmental changes. It effectively reduces the uncertainty of information provided by large models and the impact of illusion problems on the accuracy of node information in 3D scene atlases, ensuring the correctness of the information contained in the atlas, better handling of dynamically changing complex scenes, and ensuring accurate and efficient task planning and execution.

[0035] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description section. Attached Figure Description

[0036] The accompanying drawings are provided to further illustrate embodiments of the present invention and form part of the specification. They are used together with the following detailed description to explain the embodiments of the present invention, but do not constitute a limitation thereof. In the drawings:

[0037] Figure 1 This is a flowchart illustrating the implementation of a method for generating object node information in a three-dimensional scene atlas, as provided in an embodiment of the present invention.

[0038] Figure 2 This invention provides a hierarchical three-dimensional scene atlas.

[0039] Figure 3 This is a block diagram of a three-dimensional scene atlas object node information generation technology provided by an embodiment of the present invention;

[0040] Figure 4 This is a schematic diagram of the structure of a system for generating object node information in a three-dimensional scene atlas, provided in an embodiment of the present invention. Detailed Implementation

[0041] To move beyond basic maps containing only metric information, there is an urgent need to develop richer 3D scene atlases. This process first involves extracting multi-dimensional and multi-level visual features. For example, deep learning algorithms are used for object recognition to detect and label objects in the scene, while Optical Character Recognition (OCR) technology is used to recognize text within the scene. Then, large language models (LLMs) and large visual language models (LVLMs) are used to add descriptive information to entities such as objects and text in the atlas, and entities and their relationships are automatically identified from extensive text data, ensuring that entity relationships in the map are both accurate and comprehensive. Furthermore, on top of basic entities such as objects and text, higher-level spatial information, such as "rooms" and "floors," is extracted, and the categories of these spaces are identified and labeled, constructing a hierarchical and comprehensive scene perception framework. This 3D scene atlas not only has good scalability in large-scale scenes but also empowers agents to perform navigation tasks driven by natural language commands, while improving the efficiency of semantic search and path planning. This method enables the creation of a more detailed and dynamic environmental model, which not only provides richer spatial information but also supports more advanced interactive and automated tasks, greatly enhancing the mobile robot's environmental understanding and task execution capabilities.

[0042] However, the applicant discovered that in the construction of 3D scene atlases, the richness and accuracy of node information are crucial to the application of the entire atlas, directly affecting its expressive power and application scope. Accurate node information can provide more precise search results and enhance the depth of scene understanding. Existing technologies rely on single-frame images to generate node descriptions, which ignores dynamic information and temporal correlations in the intelligent perception process. Furthermore, limitations inherent in large-scale model technology lead to potential errors or incompleteness in the descriptions.

[0043] Therefore, this application aims to provide a method and system for generating object node information in a 3D scene atlas. By mining spatiotemporal clues in the scene perception information of intelligent agents, node information in the 3D scene atlas is generated and improved, reducing the impact of erroneous descriptions and significantly improving the quality and application potential of node information in the 3D scene atlas.

[0044] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention.

[0045] See Figure 1 The diagram shown is an implementation flowchart of a method for generating object node information in a three-dimensional scene atlas provided by an embodiment of the present invention. A three-dimensional scene atlas containing multiple levels is constructed for the target environment. The three-dimensional scene atlas is composed of nodes and edges contained in multiple levels, and each node contains attribute description information.

[0046] Specifically, constructing a multi-level 3D scene atlas for the target environment includes: identifying target entities in the scene and adding descriptive information to the target entities using a large language model and a visual language model; extracting the spatial information of the target entities and constructing a hierarchical 3D scene atlas based on the target entities, their descriptive information, and the spatial information.

[0047] It should be noted that in the three-dimensional scene graph, each edge connects two nodes or one node in the same layer with nodes in the upper or lower layers of the same layer.

[0048] In some implementations, it is possible to build such a system for unknown environments. Figure 2 The included Three-dimensional scene atlas at each level ,in express The set of all nodes at each level. As a hierarchy of A set of nodes:

[0049]

[0050] node ( , ) contains a collection of attribute descriptions The attribute descriptions of nodes cover multiple angles and dimensions, containing all the information that can serve as the basis for intelligent agents to plan and execute navigation and operational tasks. For object nodes, a distinction is made between "operable" and "non-operable" objects. The former includes various small objects such as water cups and food; the latter includes objects that intelligent agents generally cannot move, such as furniture and home appliances. Unlike "room" and "building" nodes, object layer nodes... The descriptive information also includes 3D point clouds. and semantic feature vectors obtained based on visual models CLIP (Constastive Language-Image Pretraining) or DINO (Self-DIstillation with NOlabels). These are used for representing the geometric positions of nodes and for identifying and comparing objects, respectively. The set of all edges in the 3D scene graph is used... This indicates that each edge can only connect two nodes at the same level or one node to a node at its upper or lower level. Layer nodes The initial edge can only connect The nodes in the set. Indicates connection hierarchy and hierarchy The set of all edges of a node (e.g., Figure 2 middle Indicates connection hierarchy and hierarchy The set of all edges of a node Indicates connection hierarchy and hierarchy The set of all edges of a node Indicates connection hierarchy and hierarchy The set of all edges of a node. This indicates the connection hierarchy. The set of all edges between two nodes. The edges connecting two nodes at the same level represent the spatial relationship between the nodes and other associations inferred by the large language model and the visual language model, such as a cup on a table, a cup can be put in a cabinet, etc.; the edges connecting two nodes at different levels represent the containment relationship between the nodes, such as the bathroom node at the room level containing the toilet and bathtub nodes at the object level.

[0051] Specifically, the implementation flowchart for the method of generating object node information in a 3D scene atlas includes the following execution steps:

[0052] Step 100: For the observation frames related to the object nodes, use the visual language model to generate attribute description information for each frame.

[0053] The attribute description information of the object node includes attribute description and 3D point cloud.

[0054] For details, please refer to Figure 3 As shown, for object nodes All related Observation frames , Using visual language models Generate a description for each frame. , :

[0055] , ;

[0056] In the formula, Represents a collection of attribute descriptions. Representing a visual language model, Represents the relationship between object nodes All relevant observation frames, This indicates a prompt word for this node.

[0057] Step 101: Obtain the text encoding vector of each frame description and the image encoding vector of the current frame image respectively.

[0058] In some implementations, due to limitations of single-frame image information, the illusion problem inherent in large model techniques, and the uncertainties contained in the generated responses, the descriptive information... This information may be inaccurate, leading to incorrect node information and misleading downstream tasks such as navigation planning, ultimately causing task failure. Therefore, please refer to... Figure 3 As shown, using a text encoder and image encoder Each frame's text encoding vector is obtained separately. , and the image encoding vector of the current frame image .

[0059] Specifically, a text encoder and an image encoder are used to obtain the text encoding vector describing each frame and the image encoding vector of the current frame image, respectively.

[0060] Step 102: Determine the cosine similarity between the text encoding vector and the image encoding vector, and select the frame description with the highest similarity as the corrected description for the current frame.

[0061] For details, please refer to Figure 3 As shown, the cosine similarity between the text encoding vector and the image encoding vector is calculated. The description of the frame with the highest similarity is selected as the corrected description for the current frame.

[0062] ;

[0063] ;

[0064] In the formula, Represents the calculation of cosine similarity. Represents the image encoding vector. Represents a text encoding vector. This indicates the corrected description for the current frame. This indicates the description of the frame with the highest similarity.

[0065] By using the aforementioned node information correction mechanism, multiple frames with a certain time span and important correlations are integrated to solve the problems of incomplete node information and misleading content obtained from a single frame.

[0066] Step 103: Fuse the corrected observation frame descriptions and use the large language model to obtain the node information of the object nodes.

[0067] For details, please refer to Figure 3 As shown, summary A revised description of the observation frames , Utilizing large language models Summarize the object nodes Node information:

[0068]

[0069] In the formula, Node information representing object nodes. This indicates the corrected description of the observation frame. It guides the large model based on Each observation frame contains prompts describing the node information of the object's nodes.

[0070] The technical effects achieved by this application are as follows:

[0071] By mining spatiotemporal cues from the scene perception information of intelligent agents, node information in 3D scene atlases is generated and improved, reducing the impact of erroneous descriptions and significantly enhancing the information quality and application value of nodes in the 3D scene atlas. This helps intelligent agents achieve deep and accurate perception of the scene, more effectively understand human instructions and expectations, and thus improve the agent's capabilities in environmental perception, autonomous navigation, intelligent decision-making, and adaptation to environmental changes. Furthermore, it effectively reduces the uncertainty of information provided by large models and the impact of illusion problems on the accuracy of node information in the 3D scene atlas, ensuring the correctness of the information contained in the atlas. This better addresses dynamically changing and complex scenes, ensuring accurate and efficient task planning and execution.

[0072] See Figure 4 The diagram shown is a structural schematic of a system for generating object node information in a 3D scene atlas according to an embodiment of the present invention. It constructs a 3D scene atlas containing multiple levels for a target environment. The 3D scene atlas is composed of nodes and edges across multiple levels, and each node contains attribute description information. The system for generating object node information in the 3D scene atlas includes:

[0073] The generation unit 400 is used to generate attribute description information for each observation frame related to the object node using a visual language model, wherein the attribute description information of the object node includes attribute description and three-dimensional point cloud.

[0074] The encoding vector acquisition unit 401 is used to acquire the text encoding vector of each frame description and the image encoding vector of the current frame image, respectively.

[0075] The cosine similarity calculation unit 402 is used to determine the cosine similarity between the text encoding vector and the image encoding vector, and select the frame description with the highest similarity as the corrected description of the current frame.

[0076] The fusion unit 403 is used to fuse the corrected observation frame description and obtain the node information of the object nodes using the large language model.

[0077] On the other hand, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method for generating object node information in a three-dimensional scene atlas as described in any of the above embodiments.

[0078] On the other hand, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the steps of the method for generating object node information in a three-dimensional scene atlas as described in any of the above embodiments.

[0079] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0080] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1A device that provides the functions specified in one or more boxes.

[0081] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0082] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0083] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0084] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, like read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0085] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0086] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0087] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for generating object node information in a three-dimensional scene atlas, characterized in that, A three-dimensional scene atlas containing multiple levels is constructed for the target environment. The three-dimensional scene atlas consists of nodes and edges across multiple levels, and each node contains attribute description information. The method for generating object node information in the three-dimensional scene atlas includes: For multiple observation frames with a time span related to object nodes, attribute description information for each frame is generated using a visual language model. Obtain the text encoding vector of each frame description and the image encoding vector of the current frame image respectively; Determine the cosine similarity between the text encoding vector and the image encoding vector, and select the frame description with the highest similarity as the corrected description for the current frame; The modified observation frame descriptions are fused together, and the node information of the object nodes is obtained using a large language model. The attribute description information for each frame is generated according to the following formula: , ; In the formula, Represents a collection of attribute descriptions. Representing a visual language model, Represents the relationship between object nodes All relevant observation frames, This indicates a prompt word for this node; Specifically, the cosine similarity between the text encoding vector and the image encoding vector is determined according to the following formula, and the frame description with the highest similarity is selected as the corrected description for the current frame: ; ; In the formula, Represents the calculation of cosine similarity. Represents the image encoding vector. Represents a text encoding vector. This indicates the corrected description for the current frame. This indicates the description of the frame with the highest similarity. The node information of the object nodes can be obtained using the following formula: In the formula, Node information representing object nodes. This indicates the corrected description of the observation frame. It guides the large model based on Each observation frame contains prompts describing the node information of the object's nodes. This includes constructing a multi-layered 3D scene atlas for the target environment, including: Identify target entities in the scene, and add descriptive information to the target entities using a large language model and a visual language model; Extract the spatial information of the target entity, and construct a hierarchical three-dimensional scene map based on the target entity, its descriptive information, and the spatial information.

2. The method for generating object node information in a three-dimensional scene atlas according to claim 1, characterized in that, In the three-dimensional scene graph, each edge connects two nodes in the same layer or one node to a node in the upper or lower layer of its own layer.

3. The method for generating object node information in a three-dimensional scene atlas according to claim 1, characterized in that, Obtain the text encoding vector of each frame description and the image encoding vector of the current frame image, including: Using a text encoder and an image encoder, the text encoding vector describing each frame and the image encoding vector of the current frame image are obtained respectively.

4. A system for generating object node information in a three-dimensional scene atlas, applicable to the method for generating object node information in a three-dimensional scene atlas according to any one of claims 1-3, characterized in that, A multi-level 3D scene atlas is constructed for the target environment. The 3D scene atlas consists of nodes and edges across multiple levels, and each node contains attribute description information. The system for generating object node information in the 3D scene atlas includes: The generation unit is used to generate attribute description information for each frame of observation frames related to object nodes using a visual language model. The encoding vector acquisition unit is used to acquire the text encoding vector describing each frame and the image encoding vector of the current frame image, respectively. The cosine similarity calculation unit is used to determine the cosine similarity between the text encoding vector and the image encoding vector, and select the frame description with the highest similarity as the corrected description of the current frame. The fusion unit is used to fuse the corrected observation frame descriptions and obtain the node information of object nodes using a large language model.

5. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method for generating object node information in a three-dimensional scene atlas as described in any one of claims 1-3.

6. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for generating object node information in a three-dimensional scene atlas as described in any one of claims 1-3.

Citation Information

Patent Citations

  • Image fine granularity identification method and device, storage medium and computer equipment

    CN116664857A

  • Visual language navigation method combining image description and text generation image

    CN117571014A

  • Space-aware map marking method and system, electronic equipment and storage medium

    CN118314293A