Dynamic scene construction method and system based on topological map
Through the dynamic scene construction method based on topological maps, the problem of inaccurate positioning in dynamic scenes of traditional robot navigation is solved, and the topological map is efficiently constructed and updated in a dynamic environment to maintain information integrity and positioning accuracy.
Patent Information
- Application Number
- CN202510471446.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional robot navigation methods have low positioning accuracy in dynamic scenarios, and existing SLAM algorithms cannot effectively process dynamic objects, resulting in mismatch between scene construction and reality, making it difficult to identify new and old targets and update the scene map.
A dynamic scene construction method based on topology map is adopted, and the node sleep mechanism is triggered through the time threshold, a dynamic topology map is built, historical topology relationships are preserved, path planning continuity is supported, and local updates are performed when the environment changes.
Effectively reflect dynamic elements of the environment, reduce information loss, maintain the stability and information integrity of the graph, improve positioning accuracy and robustness, and adapt to changes in dynamic scenarios.
Smart Images

Figure CN120252732A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of robot navigation, and particularly relates to a method and system for constructing a dynamic scene based on a topological map. Background Art
[0002] The statements in this part only provide background technical information related to the present invention, and do not necessarily constitute prior art.
[0003] Traditional robot navigation methods construct a scene model through three-dimensional scale space constraints. Once the scene changes, the correspondence between the actual scene and the objects in the scene model changes, greatly reducing the accuracy of navigation and positioning. SLAM algorithms are all based on the assumption of a static scene (the scene model does not change over time), but there are a large number of dynamic objects in the actual life scene, making excellent SLAM algorithms such as FastSLAM and ORBSLAM3 unable to meet the application requirements of real dynamic scenes.
[0004] Most of the dynamic SLAM methods proposed in recent years are dedicated to eliminating the interference of dynamic objects on the construction of static scenes. For example, by identifying prior dynamic objects and removing dynamic feature points to establish a static scene map, combining dynamic multi-object tracking with ORBSLAM3 to accurately segment dynamic objects, and using Kalman filtering to distinguish dynamic and semi-static landmarks to make full use of semi-static landmarks and pure static landmarks to improve mapping performance. These methods do not consider the handling of dynamic objects in scene construction and how to handle the changing scene structure. The constructed model does not match the actual running scene, which has a great impact on the positioning result. In addition, there are many difficulties in updating the three-dimensional scene map constructed by traditional methods. It is impossible to identify new and old targets in the scene, nor to achieve matching and updating of new and old targets in the scene. Summary of the Invention
[0005] In order to solve the technical problems existing in the above background art, the present invention provides a method and system for constructing a dynamic scene based on a topological map. The topological map is used to represent the dynamic scene, so as to effectively reflect the dynamic elements in the environment. When updating the scene map, the node sleep mechanism is triggered by a time threshold to ensure that even when the node is in a sleep state or there are other changes in the environment, the map still retains the original structure and information as much as possible, which helps to maintain the overall stability and information integrity of the map and minimizes the information loss caused by dynamic changes.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] The first aspect of the present invention provides a method for constructing a dynamic scene based on a topological map, which includes:
[0008] Obtain a number of consecutive frames of images;
[0009] For each frame of image, extract semantic nodes and feature nodes, and merge the semantic nodes and feature nodes respectively in the feature dimension and the spatial dimension to obtain a new node set. Using all the nodes in the new node set as central nodes, link other nodes to obtain a new edge set for the current frame of image;
[0010] For each frame of image, perform a union operation on the new node set and the node set of the previous frame of image to obtain the node set of the current frame of image, and perform a union operation on the new edge set and the edge set of the previous frame of image to obtain the edge set of the current frame of image, and construct a dynamic topological map to obtain a scene graph;
[0011] For the scene graph, trigger the node sleep mechanism through a time threshold, mark the nodes and edges that have not been observed for a long time as the sleep state. When positioning, the edges and nodes in the sleep state are ignored, and the nodes in the sleep state retain the historical topological relationship to support the continuity of path planning, and the state can be activated through re-observation.
[0012] Further, the steps of merging the semantic nodes in the feature dimension include:
[0013] For two semantic entities A and B, calculate the similarity between the two semantic entities. When the similarity is greater than the threshold, merge the two semantic entities into one semantic entity e new ;
[0014] Obtain the respective attribute sets P A and P B of semantic entities A and B, find the common attribute P AB of A and B, the unique attribute P A-B of A, and the unique attribute P B-A of B, and calculate to obtain the attribute set of semantic entity e new : P new =(P A-B ∪P B-A )∪{p:v|p∈P AB}, where p and v are the attribute name and attribute value respectively.
[0015] Further, the steps of merging the feature nodes in the feature dimension include:
[0016] For two feature nodes C and D, calculate the feature density: where v match represents the number of features successfully matched between feature node D and feature node C, and v ob represents the total number of features of feature node C;
[0017] If the feature density is within the set range, merge the feature node C and the D node.
[0018] Further, the step of merging the semantic node or the feature node in the spatial dimension includes:
[0019] Let p i and p j be the spatial positions representing two nodes respectively. When the spatial relationship between them satisfies p i = p i ∩ p j or p j = p i ∩ p j , then the two nodes belong to the same space, perform a union operation, and synthesize them into one node.
[0020] The second aspect of the present invention provides a dynamic scene construction system based on a topological map, which includes:
[0021] An image acquisition module, which is configured to: acquire a continuous number of frames of images;
[0022] A node extraction module, which is configured to: for each frame of image, extract semantic nodes and feature nodes, and merge the semantic nodes and feature nodes respectively in the feature dimension and the spatial dimension to obtain a new node set, and use all the nodes in the new node set as central nodes to link other nodes to obtain a new edge set of the current frame of image;
[0023] A graph construction module, which is configured to: for each frame of image, perform a union operation on the new node set and the node set of the previous frame of image to obtain the node set of the current frame of image, perform a union operation on the new edge set and the edge set of the previous frame of image to obtain the edge set of the current frame of image, construct a dynamic topological map, and obtain a scene graph;
[0024] A graph update module, which is configured to: for the scene graph, trigger a node sleep mechanism through a time threshold, mark the nodes and edges that have not been observed for a timeout as the sleep state, the edges and nodes in the sleep state are ignored during positioning, the nodes in the sleep state retain the historical topological relationship to support the continuity of path planning, and the state can be activated through re-observation.
[0025] Further, the step of merging the semantic nodes in the feature dimension includes:
[0026] For two semantic entities A and B, calculate the similarity between the two semantic entities. When the similarity is greater than the threshold, merge the two semantic entities into one semantic entity e new ;
[0027] Obtain the attribute sets P A and P B, find the common attribute P of A and B AB , the unique attribute P of A A-B and the unique attribute P of B B-A , and calculate to obtain the semantic entity e new 's attribute set: P new = (P A-B ∪ P B-A ) ∪ {p: v|p ∈ P AB}, where p and v are the attribute name and attribute value respectively.
[0028] Furthermore, the steps of merging the feature nodes in the feature dimension include:
[0029] For two feature nodes C and D, calculate the feature density: where v match represents the number of features successfully matched between the feature node D and the feature node C, and v ob represents the total number of features of the feature node C;
[0030] If the feature density is within the set range, then merge the feature nodes C and D.
[0031] Furthermore, the steps of merging the semantic nodes or feature nodes in the spatial dimension include:
[0032] Let p i and p j be the spatial positions representing two nodes respectively. When the spatial relationship between them satisfies p i = p i ∩ p j or p j = p i ∩ p j , then the two nodes belong to the same space and are processed by taking the union and combined into one node.
[0033] The third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in a method for constructing a dynamic scene based on a topological map as described above.
[0034] The fourth aspect of the present invention provides a computer device, including a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor. When the processor executes the program, it implements the steps in a method for constructing a dynamic scene based on a topological map as described above.
[0035] Compared with the prior art, the beneficial effects of the present invention are:
[0036] The present invention uses a topological map to represent a dynamic scene. Each time a new image frame arrives, new nodes and edges are updated or added based on the content of the current frame, so as to effectively reflect the dynamic elements in the environment and record their information in the graph structure, which can overcome the influence of dynamic objects in the environment and record and save the information of dynamic objects.
[0037] When updating the scene graph, the present invention distinguishes short-term and long-term observation scenes through timestamps, maintains the activation positioning and associated update of high-frequency dynamic nodes; timed-out nodes automatically go to sleep, and new topological relationships are added according to observations; removed and moved nodes are distinguished, the former permanently sleeps after a time delay, and the latter updates the subgraph structure and activates associated nodes; sleeping nodes retain historical topology to support planning continuity and can be reactivated, realizing lightweight map maintenance adaptable to the environment, ensuring that even when nodes are removed or there are other changes in the environment, the graph still retains the original structure and information as much as possible. This strategy helps to maintain the overall stability and information integrity of the graph, minimizing information loss caused by dynamic changes. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation to the invention.
[0039] Figure 1 is a flowchart of a method for constructing a dynamic scene based on a topological map according to Embodiment 1 of the present invention;
[0040] Figure 2 is a schematic diagram of the node structure of the global topological map within a short operating cycle according to Embodiment 1 of the present invention;
[0041] Figure 3 is a schematic diagram of the change in the node structure of the global topological map within a long operating cycle according to Embodiment 1 of the present invention;
[0042] Figure 4 is a schematic diagram of a subgraph related to dynamic node 5 according to Embodiment 1 of the present invention;
[0043] Figure 5 is a schematic diagram of the structure of a computer device according to Embodiment 4 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0045] It should be noted that the following detailed description is illustrative and aims to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention pertains.
[0046] Embodiment 1
[0047] This embodiment provides a method for constructing a dynamic scene based on a topological map.
[0048] The method for constructing a dynamic scene based on a topological map provided in this embodiment does not rely on mathematical scale constraints, realizes the construction of an adaptive updated dynamic topological map, improves the accuracy and robustness of positioning in a dynamic scene, and provides a feasible idea and method for effectively constructing a dynamic scene and reasonably processing dynamic objects.
[0049] The method for constructing a dynamic scene based on a topological map provided in this embodiment, first, creates nodes of a dynamic topological structure, obtains environmental data by a multi-modal sensor, extracts high-dimensional information from the environmental data, constructs semantic entities and feature entities, performs vector encoding on the extracted entities, and calculates the similarity of the entities to merge entities with a similarity higher than a threshold. Secondly, uses spatial consistency checking to determine whether entity nodes and feature nodes from different extractors represent the same node in space, and merges the nodes representing the same node into one node for node creation. Then, establishes edges of one-step length proximity relationships for different node relationships in the same frame of image, and forms a one-step length node scene subgraph of different nodes. During subsequent continuous scene transitions, creates new nodes, adds edges between different nodes, performs union processing on the topological maps of associated frames, and stitches together an expanding topological map from the scene graphs. Finally, for the impact on the topological map caused by possible node changes in a dynamic environment, uses a graph update method of node activation-sleep.
[0050] The method for constructing a dynamic scene based on a topological map provided in this embodiment, as Figure 1 shown, includes the following steps:
[0051] Step 1, Node creation.
[0052] Step 101, Extraction and merging of semantic entities (semantic nodes).
[0053] For the creation of nodes in a dynamic topological structure, high-dimensional information is extracted from environmental data. Information extraction models including but not limited to YOLO, mask-RCNN, OCR, Apriltag, etc. are used to extract information including but not limited to semantics, encoding, text, texture, function, shape, and structure. Semantic entities with entity meanings and feature entities lacking semantics are constructed, and text_to_vector vector encoding is performed on the semantic entities and feature entities respectively.
[0054] It should be noted that considering that all objects existing in the scene have semantics, due to the capabilities of the information extraction model itself and the observation perspective problem, semantics cannot be provided for all objects in the observed image. Therefore, some objects will need to be memorized at the image level. This type is classified as feature entities, and those that can provide actual semantics are classified as semantic entities. When feature entities obtain specific semantics, they will be classified as entity semantics.
[0055] For two semantic entities A and B, let (A, B) be the feature vector representation of the two semantic entities. Methods including but not limited to cosine similarity are used to calculate their similarity S, as shown in Equation (1.1). A threshold α is set to determine whether the two semantic entities are similar enough to be regarded as the same semantic entity. If S > α, then semantic entities A and B are merged into the same semantic entity e new .
[0056]
[0057] Among them, A i is the i-th feature of semantic entity A, B i is the i-th feature of semantic entity B, and n is the dimension of the feature vector.
[0058] For the two semantic entities A and B merged into one semantic entity e new , semantic entities A and B have their respective attribute sets P A and P B . Among them, each attribute is represented in the form of a key-value pair key:value (attribute name: attribute value). Then the attribute sets of A and B can be represented as P A ={p A1 :v A1 ,p A2 :v A2 ,...} and P B ={p B1 :v B1 ,p B2 :v B2 ,...}.
[0059] The common attribute P of A and B can be found in the attribute setAB All properties P of A and B A∪B Properties P unique to A A-B And properties P unique to B B-A Are respectively shown in formulas (1.2) to (1.5) as follows:
[0060] P AB = P A ∩ P B (1.2)
[0061] P A ∪ B = P A ∪ P B (1.3)
[0062] P A-B = P A - P B (1.4)
[0063] P B-A = P B - P A (1.5)
[0064] For the property P shared by A and B AB , if their property values are the same, then retain the property; for the properties unique to A and B, directly merge them into the new entity e new In, finally obtain the property set of the semantic entity e new :
[0065] P new = (P A-B ∪ P B-A ) ∪ {p:v | p ∈ P AB} (1.6)
[0066] The new entity e new Will have the integrated property set P new , that is:
[0067] e new = M(A, B) = (P new ) (1.7)
[0068] Among them, p and v are respectively the property name and the property value.
[0069] Step 102, Feature node creation.
[0070] By performing ORB and XFeat feature extraction on the feature entity, construct regional feature nodes; record the newly added features in the same region as the incremental representation of the feature nodes, and match the feature nodes through feature density:
[0071] For the feature node C in a certain area, its incremental feature is expressed as where refers to the feature vector obtained by performing ORB feature extraction on the feature entity, refers to the feature vector obtained by performing XFeat feature extraction on the feature entity; for the query node D, calculate the feature density of the matching feature subset between it and the feature node C under the current observation:
[0072]
[0073] where the query node is also a feature entity; v match represents the number of features successfully matched between the query node D and the feature node C; v ob represents the total number of features of the feature node C under the current observation; ρ f represents the density between the matching feature subset and the feature node under the current observation, and is used to measure the tightness of feature point matching.
[0074] When the feature density is within an appropriate range, the C and D feature nodes are merged, and the remaining unmatched features in this observation are added to the incremental feature set F of the feature node C; if C and D are still different nodes under the current observation, the C and D feature nodes are respectively tracked and recorded.
[0075] Step 103: Perform spatial consistent node merging on semantic nodes or feature nodes respectively to obtain a new node set.
[0076] For each node (semantic node or feature node) obtained in the above steps, perform spatial consistency check using, including but not limited to, Euclidean distance calculation. If the nodes belong to the same space, they are considered to be entities representing the same object, and they are merged into one node for node creation.
[0077] Let and be the spatial positions of two nodes node_A and node_B representing the feature space of the object of the table respectively. When the spatial relationship between the two satisfies Equation (1.9), the two nodes belong to the same space, and a union operation is performed on them to synthesize a node representing a certain object (1.10).
[0078] p i = p i ∩ p j or p j = p i ∩ p j (1.9)
[0079] node_desk = node_A ∪ node_B (1.10)
[0080] In an image containing a set of objects \(E'\), each object \(e\) i \(\in E'\) has a corresponding \(n\)-dimensional feature vector representation \(\nu\) ei \(\in \mathbb{R}^n\). n Create a set of nodes \(V\), where each node \(v\) i \(\in V\) corresponds to an object \(e\) i \(\in E'\). That is, there is a one-to-one mapping \(f: E' \to V\) such that each object \(e\) i can uniquely determine a node \(v\) i .
[0081] Step 2, Topological map synthesis.
[0082] Step 201, Topological sub-scene construction.
[0083] Taking all the nodes in the current observation (the new set of nodes in the image at time \(t\)) as central nodes respectively, link other nodes to construct a one-step topological sub-scene graph belonging to that node; repeating this step, a new edge set \(E\) of this frame of image can be obtained, forming a scene topological graph (the scene topological graph of the current image frame) \(G=(V, E)\).
[0084] Among them, the one-step topological sub-scene graph refers to the sub-graph formed by all other nodes directly linked to the central node through only one edge, which describes the scene composition within the current field of view.
[0085] It should be noted that the nodes in the same frame can be directly connected, but this does not mean that there will be an edge between every two points. For each node, it can be regarded as the center to establish its connections with all other nodes in the current observation, forming a sub-scene graph centered on that node. For example, if there are four semantic nodes of a table, a computer, a cup, and a book in the current frame of image, taking the table as the central node, three one-step long edges of table-computer, table-cup, and table-book can be directly established, thus constituting the sub-scene graph of the table.
[0086] For example, for nodes 1, 2, 3, 4, 5 in the same frame, the sub-scene centered on 1 is 1-(2, 3, 4, 5), the sub-scene centered on 2 is 2-(1, 3, 4, 5), the sub-scene centered on 3 is 3-(1, 2, 4, 5), …, the sub-scene centered on 5 is 5-(1, 2, 3, 4). However, in the sub-scene centered on 1, the edges connecting nodes 2-3, 3-4, 4-5, and 3-5 are not included.
[0087] Step 202, Topological graph synthesis.
[0088] Set a sequence \(S = \{I_1, I_2, \ldots, I\) T} represents consecutive image frames, where I t is the image at time t. The node set V t and the edge set E t represent the set of nodes and the set of edges at time t respectively. For each frame I t , a set of nodes V t′ will be created using the method in Step 1, where each node v i′ ∈V t′ can represent an object detected in the image at time t. Then, use the method in Step 201 to determine which nodes in V t′ should be connected to generate the edge set E t′ .
[0089] At the initial time t = 1, create the node set V1 = f(I1) and the edge set E1 = g(V1, V1) from I1. At this time, the graph G1 = (V1, E1) represents the content of the first frame image. For each subsequent frame, create new nodes V t from the current frame I t′ = f(I t ), create a new edge set E t′ = g(V t′ , V t′ , V t′ ) in the new node set V t′ ; perform the union operation on the new node set V t-1 and the node set V t at the previous time to obtain the complete node set V t-1 = V t′ ∪V t′ up to the current time; take the union of the new edges E t′ created by the nodes in V
[0090] E t = E t-1 ∪E t′ (2.1)
[0091] In continuous scene transitions, new nodes are created using each frame of the image, edges between different nodes are added to synthesize a new topological graph, and the union operation is performed on the node graphs with associated nodes in consecutive frames to continuously expand the node relationship graph for dynamic topological map construction.
[0092] Step 203, Graph update.
[0093] The graph update process is based on Steps 201 and 202, and the graph update process is described as follows:
[0094] (1) For the global topological map structure within a short time process, such as Figure 2As shown, record this topological graph in the data structure shown in Table 1. Consider a dynamic node 5, whose construction process satisfies that described in steps 201-202. As Figure 4 shown, the time differences of the first three observations in (a)-(c) are relatively short, occurring at 00:03 on January 1, 2024, 00:04 on January 1, 2024, and 01:00 on January 1, 2024 respectively. They are sequentially recorded into the scene subgraph centered on node 5, which is reflected in Figure 2 as the connections are established among frame3, frame4, and frame n around node 5 during these three observations. At this time, all the nodes in the global map are in the active state, and all the relationships between nodes and edges will participate in the calculation to determine the current position during positioning. (2) For the global topological map structure of a long-term process, when there appears Figure 4 the observation result in (d) in (at this time, the observation time is far from the previous observation time, occurring 5 days later), supplement the new node relationships belonging to dynamic node 5 into the subgraph centered on node 5, such as 109, and then update the connection information between node 5 and the original node edges, such as 102 and 108. These nodes are updated to the active state. For nodes and the relationships between nodes that have not been observed for too long, such as 4, 6, 7, 8, 107, these nodes will gradually enter the extinction process and turn into the dormant state, Figure 3 which reflects the update process of the global topological map. Figure 3 In, the degree of node dormancy is represented by the gradually fading node color. During positioning, the relationships of these dormant edges and nodes will be ignored. These nodes will turn from dormant to active when they are observed again next time.
[0095] Generally speaking, in the short-term high-frequency observation scenario, the dynamic node and its associated subgraph maintain the active state through continuous observations. All nodes and edges participate in the positioning calculation in real time and are continuously updated according to the observations. For the long-term discontinuous observation scenario, the node dormancy mechanism is triggered by the time threshold, and the relationships of nodes and edges that have not been updated for a long time are marked as the dormant state. At the same time, newly observed nodes are supplemented and their connection information is updated.
[0096] Generally speaking, the disappearance of dynamic nodes is divided into two cases: removal and movement. In the case of node removal, the node is never activated again. After such a node disappears, it will turn from an active node to a dormant node over time and then turn into permanent dormancy after a long time. In the case of node movement, such a node will be discovered during a new observation. At this time, the topological structure of the scene subgraph centered on this node will change, that is, the moving node reconstructs the subgraph topology based on the new observation results and synchronously adjusts the relationships of adjacent nodes, such as Figure 4 the situation shown in (d) in. The new structure of node 5 becomes 5-(102, 107, 108). InFigure 3 Among them, due to the sleep of nodes, the topological structure between activated nodes is discontinuous. For the continuity of planning, the sleeping nodes can participate in the planning process. For example, during the process from node 1 to node 5, the link of node 1-3-4-5 is temporarily in a sleeping state, but it can still be used as the basis for planning.
[0097] Table 1. Global topological map data structure
[0098]
[0099]
[0100] A method for constructing a dynamic scene based on a topological map provided in this embodiment. The construction of the scene topological map does not depend on the scale space, and the map can be updated according to subsequent observations, and the method for coping with dynamic scene changes is more flexible.
[0101] In the existing solution, the feature points belonging to dynamic objects are removed before estimating the pose. However, the dynamic objects may contain some important information, which is completely ignored during the map construction process. A method for constructing a dynamic scene based on a topological map provided in this embodiment, because nodes are created in step 1 and the topological map is synthesized in steps 201 and 202 to represent the environmental information, and the dynamic changes in the environment are captured and represented by using the node information and the update of the edges. Every time a new frame arrives, new nodes and edges are updated or added based on the content of the current frame, so as to effectively reflect the dynamic elements in the environment and record their information in the graph structure, which can overcome the influence of dynamic objects in the environment and record and save the information of dynamic objects.
[0102] The influence of dynamic objects is compared with traditional visual SLAM. When building a map in SLAM, due to the existence of dynamic objects, problems occur in the matching process at the front end of SLAM (SLAM infers its own motion by referring to objects to build a scene map. When there are dynamic objects in the scene, the result is incorrect, just like calculating one's own motion speed with a moving car on the road as a reference and getting almost 0). The solution to such problems is to pick out all the objects that may move in the environment and not let them participate in the SLAM environment modeling process. However, this also leads to another problem. These objects do exist, but the robot does not record relevant information in the scene, and it will perform inefficiently due to the lack of information in some complex tasks and inferences (for example, people are definitely removed in the SLAM method, and information such as a person appearing in room 204 cannot be provided to the inference task). Therefore, using the semantic topology method itself does not have the above problems. It does not estimate the motion by referring to a certain moving object during the process of constructing the topological map, but records the relationship between them and other nodes when they appear in the scene.
[0103] A method for constructing a dynamic scene based on a topological map provided in this embodiment enables a robot to update its cognitive model of the environment in real time and adapt to a changing scene by constructing a topological map containing dynamic object information. Due to the graph update method in step 203, only when there are changes in the environment does the graph need to be adjusted locally, such as adding or updating specific nodes and edges, rather than reconstructing the entire map every time. For dynamic elements, only the part where changes occur needs to be updated instead of reconstructing the entire map, greatly reducing the consumption of computing resources.
[0104] In traditional SLAM, the solution to the dynamic scene update problem is to reconstruct the entire scene map, and it is impossible to update the existing method based on the original map. To achieve both constructing an environmental model and updating the original scene structure, this embodiment chooses to use a topological structure to implement. Updating the scene graph is to update the nodes (including adding new nodes, associating nodes, node dormancy, etc.), so there is no need to delete the original graph and reconstruct a new one. Only updating on the basis of the original graph is sufficient, which is the natural advantage of the topological structure.
[0105] A method for constructing a dynamic scene based on a topological map provided in this embodiment can integrate multi-dimensional information in nodes. Compared with the traditional method that can only record the scene scale or semantics, in the method provided in this embodiment, during the creation process of the feature nodes in step 102, the environmental features are continuously observed and added to the feature nodes describing the scene information, so that the information of the original scene is as fully contained in the topological map as possible. Due to the graph update method in step 203, through the differential topological update strategy of node activation-dormancy adjustment, node removal / movement under short / long-term observed scenes, and the mechanism for retaining the historical relationships of dormant nodes, it is ensured that even when the environment changes, the graph still retains the original structure and information as much as possible. This strategy helps to maintain the overall stability and information integrity of the graph, minimizing the information loss caused by dynamic changes; it can retain the information of the original scene to the greatest extent, providing more data for the robot to perform complex tasks.
[0106] Embodiment Two
[0107] This embodiment provides a dynamic scene construction system based on a topological map, which specifically includes:
[0108] An image acquisition module, which is configured to: acquire a continuous number of frames of images;
[0109] The node extraction module is configured to: for each frame of image, extract semantic nodes and feature nodes, and merge the semantic nodes and feature nodes respectively in the feature dimension and the spatial dimension to obtain a new node set. Using all the nodes in the new node set as central nodes, link other nodes to obtain a new edge set of the current frame of image;
[0110] The graph construction module is configured to: for each frame of image, perform a union operation on the new node set and the node set of the previous frame of image to obtain the node set of the current frame of image, and perform a union operation on the new edge set and the edge set of the previous frame of image to obtain the edge set of the current frame of image, and construct a dynamic topological map to obtain a scene graph;
[0111] The graph update module is configured to: for the scene graph, trigger a node sleep mechanism through a time threshold, mark the nodes and edges that have not been observed for a timeout as in a sleep state. During localization, the edges and nodes in the sleep state are ignored, and the nodes in the sleep state retain the historical topological relationship to support the continuity of path planning, and can be activated through re-observation.
[0112] Furthermore, the steps of merging the semantic nodes in the feature dimension include:
[0113] For two semantic entities A and B, calculate the similarity between the two semantic entities. When the similarity is greater than the threshold, merge the two semantic entities into one semantic entity e new ;
[0114] Obtain the respective attribute sets P A and P B of semantic entities A and B, find the common attribute P AB of A and B, the unique attribute P A-B of A, and the unique attribute P B-A of B, and calculate to obtain the attribute set of semantic entity e new : P new =(P A-B ∪P B-A )∪{p:v|p∈P AB}, where p and v are the attribute name and attribute value respectively.
[0115] Furthermore, the steps of merging the feature nodes in the feature dimension include:
[0116] For two feature nodes C and D, calculate the feature density: where v match represents the number of features successfully matched between feature node D and feature node C, and v ob represents the total number of features of feature node C;
[0117] If the feature density is within the set range, merge the feature node C and the D node.
[0118] Further, the steps of merging the semantic nodes or feature nodes in the spatial dimension include:
[0119] Let p i and p j be the spatial positions representing two nodes respectively. When the spatial relationship between them satisfies p i = p i ∩p j or p j = p i ∩p j , then the two nodes belong to the same space, perform a union operation, and synthesize them into one node.
[0120] It should be noted here that each module in this embodiment corresponds to each step in Embodiment 1 one by one, and its specific implementation process is the same, so it will not be repeated here.
[0121] Embodiment 3
[0122] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in a method for constructing a dynamic scene based on a topological map as described in Embodiment 1 above.
[0123] Embodiment 4
[0124] This embodiment provides a computer device, as Figure 5 shown, including a display device, an input device, a computer-readable storage medium (volatile memory and non-volatile storage medium), a processor, a communication interface (i.e., a network interface), and a computer program stored on the computer-readable storage medium and executable on the processor. Among them, the processor, the communication interface, and the computer-readable storage medium can be connected through a bus or other means. Among them, the communication interface is used to receive and send data, and when the processor executes the program, it implements the steps in a method for constructing a dynamic scene based on a topological map as described in Embodiment 1 above.
[0125] Among them, any reference to a memory, storage, database, or other medium provided in this application and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0126] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks.
[0127] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks.
[0128] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks.
[0129] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for constructing a dynamic scene based on a topological map, characterized in that Including: Obtain a continuous number of frames of images; For each frame of image, extract semantic nodes and feature nodes, and merge the semantic nodes and feature nodes respectively in the feature dimension and the spatial dimension to obtain a new node set. Using all the nodes in the new node set as central nodes, link other nodes to obtain a new edge set of the current frame of image; For each frame of image, perform a union operation on the new node set and the node set of the previous frame of image to obtain the node set of the current frame of image, perform a union operation on the new edge set and the edge set of the previous frame of image to obtain the edge set of the current frame of image, construct a dynamic topological map, and obtain a scene graph; For the scene graph, trigger a node sleep mechanism through a time threshold, mark the nodes and edges that have not been observed for a long time as in the sleep state. When positioning, the edges and nodes in the sleep state are ignored. The nodes in the sleep state retain the historical topological relationship to support the continuity of path planning, and the state can be activated through re-observation.
2. The dynamic scene construction method based on a topological map according to claim 1, wherein The steps of merging the semantic nodes in the feature dimension include: For two semantic entities A and B, calculate the similarity between the two semantic entities. When the similarity is greater than the threshold, merge the two semantic entities into one semantic entity e new ; Obtain the attribute sets \(P\) of semantic entities \(A\) and \(B\) respectively A and \(P\) B , find the common attribute \(P\) of \(A\) and \(B\) AB , the unique attribute \(P\) of \(A\) A-B and the unique attribute \(P\) of \(B\) B-A , and calculate to obtain the attribute set of semantic entity \(e\) new : \(P\) new = \((P\) A-B \(\cup P\) B-A )\(\cup\{p:v|p\in P\) AB \}, where \(p\) and \(v\) are the attribute name and attribute value respectively 3. A method for constructing a dynamic scene based on a topological map according to claim 1, characterized in that, The steps of merging the feature nodes in the feature dimension include: For two feature nodes C and D, calculate the feature density: where v match represents the number of features successfully matched between feature node D and feature node C, and v ob represents the total number of features of feature node C; If the feature density is within the set range, merge feature nodes C and D.
4. A method for constructing a dynamic scene based on a topological map according to claim 1, characterized in that, The steps of merging the semantic nodes or feature nodes in the spatial dimension include: Let p i and p j represent the spatial positions of two nodes respectively. When the spatial relationship between them satisfies p i = p i ∩ p j or p j = p i ∩ p j , then the two nodes belong to the same space, and union processing is performed to combine them into one node.
5. A dynamic scene construction system based on a topological map, characterized in that, Including: An image acquisition module configured to: obtain a continuous number of frames of images; A node extraction module configured to: for each frame of image, extract semantic nodes and feature nodes, and merge the semantic nodes and feature nodes respectively in the feature dimension and the spatial dimension to obtain a new node set. Using all the nodes in the new node set as central nodes, link other nodes to obtain a new edge set of the current frame of image; A graph construction module configured to: for each frame of image, perform a union operation on the new node set and the node set of the previous frame of image to obtain the node set of the current frame of image, perform a union operation on the new edge set and the edge set of the previous frame of image to obtain the edge set of the current frame of image, construct a dynamic topological map, and obtain a scene graph; A graph update module configured to: for the scene graph, trigger a node sleep mechanism through a time threshold, mark the nodes and edges that have not been observed for a long time as in the sleep state. When positioning, the edges and nodes in the sleep state are ignored. The nodes in the sleep state retain the historical topological relationship to support the continuity of path planning, and the state can be activated through re-observation.
6. A dynamic scene construction system based on a topological map according to claim 5, characterized in that, The steps of merging the semantic nodes in the feature dimension include: For two semantic entities A and B, calculate the similarity between the two semantic entities. When the similarity is greater than the threshold, merge the two semantic entities into one semantic entity e new ; Obtain the attribute sets P of semantic entities A and B respectively A and P B , find the common attribute P of A and B AB , the unique attribute P of A A-B and the unique attribute P of B B-A , and calculate to obtain the attribute set of semantic entity e new : P new =(P A-B ∪P B-A )∪{p:v|p∈P AB}, where p and v are the attribute name and attribute value respectively 7. A dynamic scene construction system based on a topological map according to claim 5, characterized in that, The steps of merging the feature nodes in the feature dimension include: For two feature nodes C and D, calculate the feature density: where v match represents the number of features successfully matched between feature node D and feature node C, and v ob represents the total number of features of feature node C; If the feature density is within the set range, merge feature nodes C and D.
8. A dynamic scene construction system based on a topological map according to claim 5, characterized in that, The steps of merging the semantic nodes or feature nodes in the spatial dimension include: Let p i and p j represent the spatial positions of two nodes respectively. When the spatial relationship between them satisfies p i = p i ∩ p j or p j = p i ∩ p j , the two nodes belong to the same space, and union processing is performed to combine them into one node.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in a method for constructing a dynamic scene based on a topological map according to any one of claims 1-4.
10. A computer device, comprising a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in a method for constructing a dynamic scene based on a topological map according to any one of claims 1-4.
Citation Information
Cited By
Low-speed unmanned vehicle artificial intelligence decision and performance evaluation system
CN121187274A