A real-time environment three-dimensional reconstruction system based on semantic interaction

A real-time environmental 3D reconstruction system with semantic interaction, combined with multi-sensor fusion and energy function optimization, generates a high-fidelity continuous 3D mesh model with semantic interaction capabilities. This solves the shortcomings of existing technologies in terms of geometry, semantics, and real-time interactivity, and enables support for advanced human-computer interaction and complex tasks.

CN122244369APending Publication Date: 2026-06-19BEIJING INST OF TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING INST OF TECH
Filing Date
2026-01-21
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing technologies struggle to seamlessly and synchronously generate accurate geometric structures, high-fidelity photorealistic textures, and rich interactive semantics in a continuous 3D mesh representation within a real-time system. Traditional methods often compromise between geometry, semantics, and real-time interactivity, failing to meet the demands of advanced human-computer interaction and complex tasks.

Method used

A real-time environmental 3D reconstruction system based on semantic interaction is adopted. The system acquires natural language target query commands and 3D scene information through the data acquisition module. Combined with semantic segmentation, semantic constraint and parallel refinement modules, a continuous 3D mesh model is generated and optimized. Multi-sensor fusion technology and energy function are used to optimize the geometric and semantic boundary alignment of the semantic mesh to achieve high-fidelity texture and semantic interaction.

Benefits of technology

It generates high-fidelity, textured, and semantically interactive 3D mesh models that can respond to user intent in real time, enabling a direct workflow from natural language queries to highlighting, and providing a dynamically queryable interactive interface that supports advanced robotic tasks and immersive applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122244369A_ABST
    Figure CN122244369A_ABST
Patent Text Reader

Abstract

This invention provides a real-time 3D environmental reconstruction system based on semantic interaction. It integrates semantic prior information into the mesh optimization process, enabling clear differentiation between adjacent but semantically different objects, thus generating a mesh model that is more topologically accurate and geometrically precise. By employing an energy function that includes a geometric term, a semantic consistency term, and a semantically modulated smoothness term to optimize the vertex positions of the initial semantic mesh, it can actively reduce smoothness constraints at the boundaries of different semantic objects while maintaining surface smoothness, effectively preventing the blurring of object edges and generating object representations with clear outlines and sharp boundaries. Therefore, this invention transforms the traditional passive mapping process of "perception-processing-mapping" into an intention-driven active mapping process of "query-perception-processing-highlighting on the map," generating high-fidelity, textured 3D mesh models with semantic interaction capabilities in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of computer vision, robot localization and 3D reconstruction technology, and particularly relates to a real-time environment 3D reconstruction system based on semantic interaction. Background Technology

[0002] In the journey towards embodied intelligence and general artificial intelligence, traditional maps, as geometric replicas of the physical world, are no longer sufficient to meet the needs of autonomous interaction by intelligent agents. To empower the next generation of AI, the concept of a "living map" that transcends static representation becomes crucial. The realization of such maps relies on the real-time, seamless unification of three core attributes: accurate geometry, rich interactive semantics, and realistic high-fidelity textures. However, existing technological methods often compromise between these key attributes, hindering the realization of truly "living maps."

[0003] Currently, the mainstream 3D reconstruction and mapping technologies mainly include the following approaches, but each has its inherent technical shortcomings:

[0004] Solution 1: The literature (J. Ruan, B. Li, Y. Wang and Y. Sun, "SLAMesh: Real-timeLiDAR Simultaneous Localization and Meshing," 2023 IEEE International Conference on Robotics and Automation (ICRA), London, United Kingdom, 2023.) proposes the SLAMesh algorithm, which generates high-precision 3D mesh models using a pure geometric reconstruction framework. However, its core flaw lies in its lack of understanding of scene content, failing to capture necessary semantic and textural information, ultimately producing geometrically accurate but information-poor "sterile digital copies." Furthermore, when dealing with structurally similar or spatially adjacent objects (e.g., vehicles parked on the ground), the pure geometric method is prone to generating incorrect topological connections, causing objects to adhere to or deform with the environment, failing to accurately represent the physical realism of the scene.

[0005] Option 2: The literature (X. Chen, A. Milioto, E. Palazzolo, P. Giguère, J. Behley and C. Stachniss, "SuMa++: Efficient LiDAR-based Semantic SLAM," 2019 IEEE / RSJ International Conference on Intelligent Robots and Systems (IROS), China, 2019.) proposes the SuMa++ algorithm. Semantic SLAM systems based on the SuMa++ algorithm aim to integrate semantic labels into 3D maps, thereby endowing the map with the ability to understand the scene. However, this integration often comes at the cost of geometric accuracy, texture fidelity, or real-time interactivity. More importantly, the semantic generation process of these systems is usually "passive," that is, the scene is statically labeled according to a pre-trained model during the mapping process, generating a pre-labeled static map. In this mode, the agent cannot interact with and query the map in real time according to its own immediate intentions (such as through natural language commands), which greatly limits its application potential in advanced human-computer interaction and complex task execution.

[0006] Solution 3: The literature (C. Zheng et al., "FAST-LIVO2: Fast, Direct LiDAR–Inertial–Visual Odometry," in IEEE Transactions on Robotics, vol. 41, pp. 326-346, 2025.) proposes the FAST-LIVO2 algorithm. SLAM systems based on LiDAR-vision-inertial fusion, represented by FAST-LIVO2, can generate dense, colored 3D point cloud maps in real time through tight coupling of multiple sensors, which are visually appealing. However, its fundamental limitation lies in the fact that the final output is a discrete point cloud representation. Although the effect is striking from a distance, the inherent sparsity of the point cloud causes it to "disintegrate" when viewed up close, failing to form a continuous surface. This discontinuous representation is a fatal flaw for applications requiring fine interaction (such as physical simulation, virtual object placement) and high-fidelity rendering, as it lacks the continuous surface foundation necessary for effective rendering and interaction.

[0007] In summary, a significant technological gap exists in the existing technology: currently, no real-time system can seamlessly and synchronously unify precise geometric mesh structures, high-fidelity photorealistic textures, and rich interactive semantics driven by user intent into a continuous 3D mesh representation. Therefore, there is an urgent need to develop a new technical solution to overcome the aforementioned deficiencies in the existing technology. Summary of the Invention

[0008] To address the aforementioned problems, this invention provides a real-time environmental 3D reconstruction system based on semantic interaction, which can generate continuous 3D mesh models with accurate geometric structure, high-fidelity photorealistic textures, and rich interactive semantics in real time and synchronously.

[0009] A real-time environmental 3D reconstruction system based on semantic interaction includes a data acquisition module, a semantic segmentation module, a semantic constraint module, a parallel refinement module, and a semantic rendering module. The data acquisition module is used to acquire the user-defined natural language target query command, the current frame two-dimensional image of the three-dimensional scene, and the current frame lidar point cloud; The semantic segmentation module is used to obtain the two-dimensional semantic mask corresponding to the two-dimensional image of the current frame according to the natural language target query command. Each pixel in the two-dimensional semantic mask represents the probability that each pixel in the two-dimensional image belongs to a certain semantic category. Simultaneously, the semantic segmentation module divides the three-dimensional space of the three-dimensional scene into multiple three-dimensional voxels, and then determines the three-dimensional voxel to which each point in the current frame's LiDAR point cloud belongs based on the spatial mapping relationship. The semantic segmentation module determines whether the number of points contained in each three-dimensional voxel in the current frame has changed compared to the number of points contained in the previous frame. For three-dimensional voxels with a negative result, the semantic segmentation module does not perform subsequent operations on them, and receives the next frame of LiDAR point cloud after all voxels have been processed. Three-dimensional voxels with a positive result are recorded as changed three-dimensional voxels. The semantic segmentation module then determines the semantic probability distribution vector of each changed three-dimensional voxel based on the spatial mapping relationship between the two-dimensional image and the LiDAR point cloud. ,in, This represents the probability that the spatial location of the point contained in the changed 3D voxel belongs to each semantic category; finally, the semantic segmentation module performs Delaunay triangulation operation on the point contained in each changed 3D voxel to obtain the initial semantic mesh corresponding to each changed 3D voxel. The semantic constraint module uses a set energy function to optimize each initial semantic mesh, so that the geometric boundary of each initial semantic mesh is aligned with the semantic boundary of the 3D scene, and the optimized semantic mesh corresponding to each changed 3D voxel is obtained. The parallel refinement module determines whether the reliability of each vertex in each optimized semantic mesh is greater than a set threshold, and performs texture and semantic updates on the vertices with a yes result to obtain the final semantic mesh corresponding to each modified 3D voxel, and obtains the corresponding semantic probability distribution vector based on the final semantic mesh. ; The semantic rendering module is used to render based on each semantic probability distribution vector. Each final semantic grid is rendered separately to achieve highlighting of the target object instance specified in the natural language target query command.

[0010] Furthermore, the semantic segmentation module determines the semantic probability distribution vector of any changed 3D voxel based on the spatial mapping relationship between the 2D image and the LiDAR point cloud. The method is as follows: Get the current number Confidence weight of the frame 2D image relative to the currently changed 3D voxel as follows:

[0011] in, For the current number The confidence score of the two-dimensional semantic mask corresponding to the frame two-dimensional image. For the cosine term of the current changing three-dimensional voxel observation view, To address the current depth uncertainty of the modified 3D voxels, This is a scaling factor used to adjust the intensity of the depth penalty; Based on confidence weight Get the currently changed 3D voxel relative to the current 3D voxel. semantic probability distribution vector of a frame 2D image as follows:

[0012] in, For the current change of three-dimensional voxels relative to the first The semantic probability distribution vector of a two-dimensional image frame. For the current number The semantic probability distribution vector of the two-dimensional semantic mask corresponding to the frame two-dimensional image. For the first The cumulative confidence weight value corresponding to the frame. For the current number The cumulative confidence weight value corresponding to the frame, where, The calculation formula is: .

[0013] Furthermore, the semantic segmentation module uses the following method to obtain the two-dimensional semantic mask corresponding to the two-dimensional image of the current frame based on the natural language target query instruction: The Grounding DINO model is used to convert natural language target query instructions into detection boxes of targets contained in the current frame's two-dimensional image; The Segment Anything Model is used to generate a two-dimensional semantic mask based on the target's bounding box.

[0014] Furthermore, the semantic constraint module uses a predefined energy function when optimizing any initial semantic grid. as follows:

[0015] in, Here is a planar metric function used to constrain the fit between the initial semantic mesh surface and the lidar point cloud. This is a semantic consistency term used to constrain mesh edges of objects with different semantic meanings. For semantic consistency items The corresponding weights This is a smoothness term used to ensure that the normal vectors of adjacent mesh patches remain consistent. For smoothness term The corresponding weights.

[0016] Furthermore, semantic consistency items The calculation method is as follows:

[0017] in, For the current initial semantic grid, the first One vertex, For the current initial semantic grid, the first One vertex, As vertices The semantic probability distribution As vertices The semantic probability distribution Let the set of edges between any two vertices in the current initial semantic grid be defined. for and The divergence between them As vertices With vertex The semantic weights between them, and the semantic weights The calculation method is as follows:

[0018] in, As vertices With vertex The length of the side between them As vertices In the current number Semantic confidence of a frame As the vertex In the current number Semantic confidence of the frame; simultaneously, semantic confidence The calculation method is as follows:

[0019] in, To set the scaling factor, As vertices The cumulative weight of the voxel.

[0020] Furthermore, the smoothness term The calculation method is as follows:

[0021] in, For the current initial semantic grid, the first One vertex, As vertices The set of neighboring nodes, For the set of neighbor nodes The first in One vertex, As vertices Surface normal vector, As vertices Surface normal vector, As vertices With vertex Smoothing weights between them, and The calculation method is as follows:

[0022] in, As vertices With vertex Geometric weights between them As vertices With vertex The semantic similarity between them, and have

[0023] in, As vertices The semantic probability distribution As vertices The semantic probability distribution for and The divergence between them To set the coefficients.

[0024] Furthermore, the semantic constraint module obtains the optimized semantic mesh corresponding to any changed 3D voxel as follows: The energy function of the initial semantic grid corresponding to the current modified 3D voxel is nonlinearly iterated using the Gauss-Seidel method until the set upper limit of the number of iterations is reached. The semantic grid corresponding to the energy function obtained in the last iteration is then used as the optimized semantic grid corresponding to the current modified 3D voxel.

[0025] Furthermore, the parallel refinement module obtains the reliability of any vertex in any optimized semantic mesh. The method is as follows:

[0026] in, The photometric noise at the current vertex. For the geometric noise of the current vertex, The luminosity residual at the current vertex. This represents the geometric residual of the current vertex.

[0027] Furthermore, the photometric residual at the current vertex and geometric residuals The calculation method is as follows:

[0028]

[0029] in, The projection coordinates of the current vertex onto the current frame's 2D image. The pixel color at that location. The projection coordinates of the current vertex across all 2D images in previous frames. The average pixel color at that location. The actual projected depth of the current vertex. This is the reference depth value for the current vertex. A constant is set to prevent division by zero.

[0030] Furthermore, the parallel refinement module performs texture and semantic updates for any vertex whose judgment result is yes, as follows:

[0031] in, For the vertex whose current judgment result is yes, in the current i-th The corresponding pixel in the two-dimensional semantic mask of the frame two-dimensional image belongs to the category. The unnormalized cumulative log probability, The vertex whose current judgment result is yes is at the th The corresponding pixel in the two-dimensional semantic mask of the frame two-dimensional image belongs to the category. The unnormalized cumulative log probability, For the current number The two-dimensional semantic mask corresponding to the frame two-dimensional image. The vertex whose current judgment result is yes is in Belongs to the category The log-likelihood function measured by reliability, and The calculation method is as follows:

[0032] in, For the vertex whose current judgment result is yes, in the current i-th Frame reliability, The vertex whose current judgment result is yes is in Belongs to the category The primitive log-likelihood function.

[0033] Beneficial effects: 1. This invention provides a real-time environmental 3D reconstruction system based on semantic interaction. First, semantic prior information is integrated into the mesh optimization process, solving the inherent geometric ambiguity problem in pure geometric methods. For example, it can clearly distinguish between adjacent but semantically different objects, such as vehicles and the ground, thereby generating a mesh model that is more topologically correct and geometrically more accurate. Then, an energy function containing a geometric term, a semantic consistency term, and a semantically modulated smoothness term is used to optimize the vertex positions of the initial semantic mesh. This can maintain surface smoothness while actively reducing smoothness constraints at the boundaries of different semantic objects, thereby effectively preventing the blurring of object edges and generating object representations with clear outlines and sharp boundaries. Thus, this invention can utilize multi-sensor fusion technology to transform the traditional passive mapping of "perception-processing-mapping" into an intention-driven active mapping of "query-perception-processing-highlighting on the map," which can generate high-fidelity, textured 3D mesh models with semantic interaction capabilities in real time.

[0034] 2. This invention provides a real-time 3D environmental reconstruction system based on semantic interaction, establishing a direct workflow from natural language queries to real-time highlighting of target objects on a high-fidelity map. This "semantic brush" function transforms the map from a static dataset into a dynamic, queryable, and interactive interface, realizing a paradigm shift in the way humans interact with intelligent agents.

[0035] 3. This invention provides a real-time environmental 3D reconstruction system based on semantic interaction, which can generate a single and consistent continuous mesh model that achieves high fidelity in the three dimensions of geometry, semantics and texture. It is not only a geometrically accurate digital twin, but also a semantically rich and textured "active map", providing a solid foundation for advanced robotic tasks and immersive applications.

[0036] 4. This invention provides a real-time environmental 3D reconstruction system based on semantic interaction. It adopts an incremental meshing strategy and only updates "active" voxels whose number of LiDAR points changes, thereby avoiding the huge computational overhead caused by global reconstruction of the entire scene. Attached Figure Description

[0037] Figure 1 The flowchart of the method described in this invention illustrates the two main stages: semantically guided reconstruction and parallel optimization.

[0038] Figure 2 This is a reconstruction rendering of the invention in a large-scale real-world scene, showcasing accurate geometry, photorealistic textures, and interactive effects of highlighting vehicles using the "semantic brush" function.

[0039] Figure 3 This is a qualitative comparison chart of the geometric reconstruction quality between the present invention and existing technologies.

[0040] Figure 4 This is a qualitative comparison chart of the texture mapping quality between the present invention and existing technologies.

[0041] Figure 5 This invention enables the method to distinguish different types of objects and achieve highlight effects in the same reconstruction. Detailed Implementation

[0042] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0043] Reference Figure 1This invention proposes a real-time 3D environmental reconstruction system based on semantic interaction, comprising a data acquisition module, a semantic segmentation module, a semantic constraint module, a parallel refinement module, and a semantic rendering module. The core of this system is a methodology called "semantic brush," which deeply integrates semantic understanding into the entire 3D reconstruction process through a two-stage pipeline from coarse to fine. This design transforms the traditional passive mapping (i.e., "perception-processing-mapping") into a proactive, intent-driven paradigm (i.e., "query-perception-processing-highlighting on the map"). This proactive interactive characteristic is the fundamental innovation that distinguishes this invention from existing passive semantic SLAM systems, providing key technical support for advanced human-computer interaction, augmented reality, and other application scenarios.

[0044] In general, the overall workflow of the real-time environmental 3D reconstruction system of this invention is divided into two main stages: Phase 1: Semantic-Guided Reconstruction. The goal of this phase is to construct an initial mesh model embedded with semantic information. This phase begins with a semantic segmentation module, which translates the user's natural language query (e.g., "car") into a semantic mask on a 2D image in real time. Subsequently, this semantic information is deeply fused with geometric point cloud information from LiDAR at the voxel level. Next, an incremental meshing process efficiently generates the initial semantic mesh. Finally, a semantic constraint module optimizes this initial semantic mesh using a unified energy function, aligning its geometric boundaries with semantic boundaries.

[0045] Phase Two: Parallel Optimization. This phase aims to improve the accuracy of various attributes of the initial semantic mesh, including texture and semantic labels. To concentrate computational resources, this phase first performs frustum culling on the initial semantic mesh. Then, a unified reliability evaluation metric calculates the observation quality of each visible vertex. Guided by this unified reliability metric, semantic optimization and texture mapping work in parallel, collaboratively optimizing the semantic labels and high-fidelity textures for each vertex.

[0046] Ultimately, the real-time environment 3D reconstruction system outputs a photorealistic, textured, optimized semantic mesh. The semantic rendering module utilizes this optimized semantic mesh to highlight target objects in real-time based on the user's initial query, thus completing a full closed loop from natural language input to 3D model interaction.

[0047] The functions of each module are described in detail below.

[0048] I. Data Acquisition Module The data acquisition module is used to acquire the user-defined natural language target query command, the current frame two-dimensional image of the three-dimensional scene, and the current frame lidar point cloud; It should be noted that the present invention uses a multi-sensor data acquisition module to simultaneously capture LiDAR scans, inertial measurement unit (IMU) data, and camera images.

[0049] II. Semantic Segmentation Module The semantic segmentation module is used to obtain the two-dimensional semantic mask corresponding to the two-dimensional image of the current frame according to the natural language target query command. Each pixel in the two-dimensional semantic mask represents the probability that each pixel in the two-dimensional image belongs to a certain semantic category. Simultaneously, the semantic segmentation module divides the three-dimensional space of the three-dimensional scene into multiple three-dimensional voxels, and then determines the three-dimensional voxel to which each point in the current frame's LiDAR point cloud belongs based on the spatial mapping relationship. The semantic segmentation module determines whether the number of points contained in each three-dimensional voxel in the current frame has changed compared to the number of points contained in the previous frame. For three-dimensional voxels with a negative result, the semantic segmentation module does not perform subsequent operations on them, and receives the next frame of LiDAR point cloud after all voxels have been processed. Three-dimensional voxels with a positive result are recorded as changed three-dimensional voxels. The semantic segmentation module then determines the semantic probability distribution vector of each changed three-dimensional voxel based on the spatial mapping relationship between the two-dimensional image and the LiDAR point cloud. ,in, This represents the probability that the spatial location of the point contained in the changed 3D voxel belongs to each semantic category; finally, the semantic segmentation module performs Delaunay triangulation operation on the point contained in each changed 3D voxel to obtain the initial semantic mesh corresponding to each changed 3D voxel. For example, when the system receives a natural language query (e.g., the text prompt "car"), the semantic segmentation module is activated. To achieve open set segmentation of any object, this invention preferably employs a zero-shot segmentation model. A specific embodiment uses the Grounded-Segment-Anything (G-SAM) model. G-SAM achieves its function by combining the Grounding DINO model and the Segment Anything Model (SAM): the former is responsible for converting the text prompt into a detection box in the image, while the latter generates a high-precision two-dimensional segmentation mask based on the detection box. To meet real-time requirements (e.g., in a 10Hz LiDAR system, the processing time per frame needs to be less than 100ms), this invention also provides an alternative: employing the computationally more efficient FastSAM model. FastSAM is a convolutional neural network (CNN) based architecture that decouples the segmentation task into two stages: full instance segmentation and prompt-guided selection, which can significantly reduce computational resource requirements while maintaining competitive segmentation performance. This strategic model choice between accuracy and efficiency ensures the practicality of this invention across different hardware platforms and application scenarios.

[0050] Furthermore, the semantic segmentation module determines the semantic probability distribution vector of any changed 3D voxel based on the spatial mapping relationship between the 2D image and the LiDAR point cloud. The method is as follows: Get the current number Confidence weight of the frame 2D image relative to the currently changed 3D voxel as follows:

[0051] in, For the current number The confidence score of the two-dimensional semantic mask corresponding to the frame two-dimensional image reflects the reliability of the segmentation result; For the current change of the observation perspective cosine term of the three-dimensional voxel, where It is the angle between the direction of the camera ray and the surface normal. This term is used to reduce the weight of observations at grazing angles. To address the current depth uncertainty of the modified 3D voxels, The term penalizes observations with significant depth uncertainty; This is a scaling factor used to adjust the intensity of the depth penalty; in this way, geometric information (viewpoint, depth) and semantic information (segmentation confidence) are tightly coupled at the most basic data level—the voxel level.

[0052] Based on confidence weight Get the currently changed 3D voxel relative to the current 3D voxel. semantic probability distribution vector of a frame 2D image as follows:

[0053] in, For the current change of three-dimensional voxels relative to the first The semantic probability distribution vector of a two-dimensional image frame. For the current number The semantic probability distribution vector of the two-dimensional semantic mask corresponding to the frame two-dimensional image. For the first The cumulative confidence weight value corresponding to the frame. For the current number The cumulative confidence weight value corresponding to the frame, where, The calculation formula is: .

[0054] Therefore, the generated two-dimensional semantic mask is fused into a three-dimensional voxel. Each three-dimensional voxel maintains a semantic probability distribution vector. , This represents the probability that the spatial location of a 3D voxel belongs to each semantic category. When the number of laser point clouds in a 3D voxel changes, the semantic probability distribution of that 3D voxel will be updated using a confidence-weighted average.

[0055] It should be noted that, in order to construct a continuous surface model without sacrificing real-time performance, this invention employs an incremental meshing strategy. This strategy updates only the "active" voxels where the number of LiDAR points changes, thereby avoiding the enormous computational overhead of globally reconstructing the entire scene.

[0056] Within an active voxel, the system first performs local planar fitting on the contained LiDAR point cloud. Then, these 3D points are projected onto the fitted 2D plane, and 2D Delaunay triangulation is performed. Delaunay triangulation is a triangulation method that maximizes the minimum angle, generating well-shaped meshes that avoid elongated triangles, providing a good foundation for subsequent optimization and rendering. After 2D triangulation, the resulting connectivity is lifted back into 3D space, forming new mesh patches (triangular faces).

[0057] A crucial step in this process is "semantic inheritance." When a new mesh vertex is created, it directly and completely inherits the semantic probability distribution vector of its parent voxel. This design ensures that the generated mesh carries preliminary semantic information from its inception, rather than undergoing post-processing semantic annotation after geometric reconstruction. This allows the integration of geometry and semantics throughout the entire mesh generation process.

[0058] III. Semantic Constraint Module The semantic constraint module uses a set energy function to optimize each initial semantic mesh, so that the geometric boundary of each initial semantic mesh is aligned with the semantic boundary of the 3D scene, and the optimized semantic mesh corresponding to each changed 3D voxel is obtained. Specifically, the semantic constraint module obtains the optimized semantic grid corresponding to any changed 3D voxel by performing nonlinear iteration on the energy function of the initial semantic grid corresponding to the current changed 3D voxel using the Gauss-Seidel method until the set upper limit of the number of iterations is reached. The semantic grid corresponding to the energy function obtained in the last iteration is then used as the optimized semantic grid corresponding to the current changed 3D voxel.

[0059] Furthermore, the semantic constraint module uses a predefined energy function when optimizing any initial semantic grid. as follows:

[0060] in, Let be the planar metric function used to constrain the fit between the initial semantic mesh surface and the lidar point cloud, that is, It is a standard point-to-plane metric used to ensure that the mesh surface fits the original LiDAR point cloud, guaranteeing geometric fidelity; This is a semantic consistency term used to constrain mesh edges that connect different semantic objects, and its purpose is to penalize mesh edges that connect different semantic objects. For semantic consistency items The corresponding weights A smoothness term is used to constrain the normal vectors of adjacent mesh patches to maintain consistency, in order to produce a smooth surface; For smoothness term The corresponding weights. Therefore, it can be seen that... and It is a weighting coefficient used to balance the effects of various factors.

[0061] Furthermore, semantic consistency items The calculation method is as follows:

[0062] in, For the current initial semantic grid, the first One vertex, For the current initial semantic grid, the first One vertex, As vertices The semantic probability distribution As vertices The semantic probability distribution Let the set of edges between any two vertices in the current initial semantic grid be defined. for and The divergence between them As vertices With vertex The semantic weights between vertices are related to the semantic confidence of the vertices and the edge length, and the semantic weights... The calculation method is as follows:

[0063] in, As vertices With vertex The length of the side between them As vertices In the current number Semantic confidence of a frame As the vertex In the current number The semantic confidence of a frame is determined by the accumulated observation weights of its parent voxels. The specific calculation method for exporting is as follows:

[0064] in, To set the scaling factor, As vertices The cumulative weight of the voxel.

[0065] It should be noted that semantic consistency items The goal is to sum all edges in the grid. This invention uses the Jensen-Shannon (JS) divergence to measure the divergence between two vertices. and semantic probability distribution and The difference between them. The reason for choosing JS divergence is that it is a symmetric and bounded measure (range is...). ), which is well-suited for comparing probability distributions. Compared to the asymmetric and potentially unbounded Körbek-Leibler (KL) divergence, the JS divergence is numerically more stable and physically more intuitive, and can robustly quantify the semantic “distance” between two vertices.

[0066] Furthermore, unlike traditional methods, the smoothness term of this invention... Subject to dynamic modulation of semantic information, the specific calculation method is as follows:

[0067] in, For the current initial semantic grid, the first One vertex, As vertices The set of neighboring nodes, For the set of neighbor nodes The first in One vertex, As vertices Surface normal vector, As vertices Surface normal vector, As vertices With vertex Smoothing weights between them, and The calculation method is as follows:

[0068] in, As vertices With vertex The geometric weights between them (e.g., inversely proportional to the side length). As vertices With vertex The semantic similarity between them, and have

[0069] in, As vertices The semantic probability distribution As vertices The semantic probability distribution for and The divergence between them To set the coefficients.

[0070] The underlying logic of this design is that when two vertices are semantically similar (e.g., both belong to the "wall"), The divergence approaches 0. The value tends to 1, indicating a strong smoothing constraint; however, when the semantics of two vertices differ significantly (e.g., one belongs to "car" and the other to "ground"), the result is different. The divergence is relatively large. As the boundary approaches zero, the smoothness constraint is significantly weakened. This mechanism is the core of this invention's ability to maintain clear object boundaries. It directly uses semantic information as a control variable in the geometric optimization process, achieving a deep fusion of semantics and geometry, which is fundamentally different from existing technologies (which only use semantics as post-processing labels).

[0071] IV. Parallel Refinement Module The parallel refinement module determines whether the reliability of each vertex in each optimized semantic mesh is greater than a set threshold, and performs texture and semantic updates on the vertices with a yes result to obtain the final semantic mesh corresponding to each modified 3D voxel, and obtains the corresponding semantic probability distribution vector based on the final semantic mesh. ; It should be noted that after optimizing the semantic mesh construction, the system enters a parallel optimization phase to further improve the quality of textures and semantic labels. To improve efficiency, the system first performs view frustum culling to ensure that only vertices visible in the current camera frame are considered for optimization.

[0072] For each visible vertex, the system evaluates the quality of its projection from 3D space to a 2D image. This evaluation is performed by calculating a uniform reliability metric. This metric is designed to robustly penalize observations that are inconsistent in both photometric and geometrical aspects. Specifically, the parallel refinement module obtains the reliability of any vertex in any optimized semantic mesh. The method is as follows:

[0073] in, The photometric noise at the current vertex. For the geometric noise of the current vertex, The luminosity residual at the current vertex. This represents the geometric residual of the current vertex.

[0074] Specifically, the photometric residual of the current vertex and geometric residuals The calculation method is as follows:

[0075]

[0076] in, The projection coordinates of the current vertex onto the current frame's 2D image. The pixel color at that location. The projection coordinates of the current vertex across all 2D images in previous frames. The average pixel color at that location. The actual projected depth of the current vertex. This is the reference depth value for the current vertex. A constant is set to prevent division by zero.

[0077] It should be noted that this unified reliability metric The design is a critical engineering decision. It integrates errors from different physical modalities (color and depth) into a single, principled metric. Only those... Only observations with high values ​​(i.e., high consistency in both luminosity and geometry) are considered reliable and passed on to subsequent update steps. This rigorous data filtering mechanism effectively prevents noisy data introduced by factors such as motion blur, drastic lighting changes, occlusion, or inaccurate depth estimation, while also preventing contamination of texture and semantic information, thus ensuring the overall high fidelity of the final map.

[0078] Based on this, in a unified reliability metric Guided by this, the parallel refinement module performs parallel Bayesian updates on texture and semantics. For texture mapping, each vertex maintains a Gaussian color state. This state is in a state determined by Updates are performed within a modulated Bayesian framework. High-reliability color observations contribute greater weight to the final color of the vertices. For semantic optimization, the semantic probability distribution of each vertex is enhanced in the logarithmic probability domain through a recursive Bayesian update process.

[0079] Specifically, the parallel refinement module performs texture and semantic updates for any vertex whose judgment result is yes, as follows:

[0080] in, For the vertex whose current judgment result is yes, in the current i-th The corresponding pixel in the two-dimensional semantic mask of the frame two-dimensional image belongs to the category. The unnormalized cumulative log probability, The vertex whose current judgment result is yes is at the th The corresponding pixel in the two-dimensional semantic mask of the frame two-dimensional image belongs to the category. The unnormalized cumulative log probability, For the current number The two-dimensional semantic mask corresponding to the frame two-dimensional image. The vertex whose current judgment result is yes is in Belongs to the category The log-likelihood function measured by reliability, and The calculation method is as follows:

[0081] in, For the vertex whose current judgment result is yes, in the current i-th Frame reliability, The vertex whose current judgment result is yes is in Belongs to the category The primitive log-likelihood function.

[0082] This design will This is interpreted as a parameter controlling the sharpness of the likelihood distribution. A highly reliable observation ( This will produce a highly confident update, while an unreliable observation () will produce a more confident update. This will produce a nearly uniform likelihood distribution, with negligible impact on the existing probability distribution. Through direct reuse... This invention establishes a tight, principled coupling between texture updates and semantic updates, ensuring that information from unreliable viewpoints is consistently and synchronously discounted when updating both attributes.

[0083] V. Semantic Rendering Module The semantic rendering module is used to render based on each semantic probability distribution vector. Each final semantic grid is rendered separately to achieve highlighting of the target object instance specified in the natural language target query command.

[0084] It should be noted that after completing texture and semantic updates, this invention first uses the softmax function to calculate the unnormalized cumulative log probability. Restored to semantic probability distribution vector The semantic rendering module utilizes this final, refined semantic distribution to highlight target object instances based on the user's initial natural language query when rendering the final textured mesh (e.g., rendering all mesh faces identified as "cars" in a specific color). Figure 2 As shown in (c1) and (c2) in the diagram, this completes the full closed loop from user intent input to real-time interaction with the 3D scene.

[0085] Furthermore, this invention is not limited to a specific hardware configuration, but can be implemented by a specific hardware system. A preferred hardware embodiment is a handheld data acquisition device. This device integrates a LiDAR (e.g., Livox Avia), a global shutter RGB camera, an inertial measurement unit (IMU), and an onboard computing unit (e.g., Intel NUC). All sensors undergo rigorous intrinsic and extrinsic parameter and time synchronization calibration to ensure data stream alignment. This portable device provides a concrete, non-limiting physical carrier for deploying the method of this invention, enabling its application to real-time mapping tasks in various indoor and outdoor scenarios.

[0086] Qualitative comparison results as follows Figure 3 , Figure 4 and Figure 5 As shown, the superiority of the present invention is further demonstrated intuitively. Figure 3 The geometric reconstruction comparison shows that the present invention (top-down) can clearly separate the vehicle from the ground, avoiding the adhesion and deformation artifacts that occur in other methods (middle and down-down). Figure 4 The texture mapping comparison shows that the continuous mesh texture of the present invention (second column) maintains clear and sharp details when viewed at close range, while the point cloud-based coloring method (third and fourth columns) suffers from disintegration and distortion due to its discreteness. Figure 5 This demonstrates how the present invention distinguishes different types of objects and creates highlight effects in the same reconstruction.

[0087] In summary, this invention discloses a real-time 3D environment reconstruction system based on semantic interaction, and the data processing flow can be summarized as follows: a) Receive natural language queries and generate corresponding two-dimensional semantic masks for synchronized image frames; b) The two-dimensional semantic mask and the three-dimensional point cloud data are fused within a voxel grid, and the semantic probability distribution of each voxel is updated using a confidence-weighted average method. c) Perform incremental meshing within the active voxel to generate an initial semantic mesh, where the mesh vertices inherit the semantic probability distribution of their parent voxel; d) Optimize the vertex positions of the initial semantic grid by minimizing a uniform energy function, which includes a geometric term, a semantic consistency term, and a semantically modulated smoothness term.

[0088] e) Calculate a unified reliability metric for visible mesh vertices based on the photometric and geometric consistency between the 3D mesh and the 2D image; f) Guided by the unified reliability metric, perform parallel Bayesian updates to simultaneously optimize the texture graph properties and per-vertex semantic label properties of the mesh.

[0089] Therefore, this invention, through its unique "semantic brush" methodology, deeply integrates semantic understanding into the entire process from data fusion, mesh generation, geometry optimization to parallel attribute optimization, successfully unifying high-precision geometry, high-fidelity texture, and high-level interactive semantics within a real-time framework. Experimental results demonstrate that this invention achieves a leading level in terms of geometric reconstruction accuracy, texture fidelity, and semantic correctness compared to existing technologies.

[0090] Of course, the present invention may have other various embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and modifications according to the present invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims.

Claims

1. A real-time 3D environment reconstruction system based on semantic interaction, characterized in that, It includes a data acquisition module, a semantic segmentation module, a semantic constraint module, a parallel refinement module, and a semantic rendering module; The data acquisition module is used to acquire the user-defined natural language target query command, the current frame two-dimensional image of the three-dimensional scene, and the current frame lidar point cloud; The semantic segmentation module is used to obtain the two-dimensional semantic mask corresponding to the two-dimensional image of the current frame according to the natural language target query command. Each pixel in the two-dimensional semantic mask represents the probability that each pixel in the two-dimensional image belongs to a certain semantic category. Simultaneously, the semantic segmentation module divides the three-dimensional space of the three-dimensional scene into multiple three-dimensional voxels, and then determines the three-dimensional voxel to which each point in the current frame's LiDAR point cloud belongs based on the spatial mapping relationship. The semantic segmentation module determines whether the number of points contained in each three-dimensional voxel in the current frame has changed compared to the number of points contained in the previous frame. For three-dimensional voxels with a negative result, the semantic segmentation module does not perform subsequent operations on them, and receives the next frame of LiDAR point cloud after all voxels have been processed. Three-dimensional voxels with a positive result are recorded as changed three-dimensional voxels. The semantic segmentation module then determines the semantic probability distribution vector of each changed three-dimensional voxel based on the spatial mapping relationship between the two-dimensional image and the LiDAR point cloud. ,in, This represents the probability that the spatial location of the point contained in the changed 3D voxel belongs to each semantic category; finally, the semantic segmentation module performs Delaunay triangulation operation on the point contained in each changed 3D voxel to obtain the initial semantic mesh corresponding to each changed 3D voxel. The semantic constraint module uses a set energy function to optimize each initial semantic mesh, so that the geometric boundary of each initial semantic mesh is aligned with the semantic boundary of the 3D scene, and the optimized semantic mesh corresponding to each changed 3D voxel is obtained. The parallel refinement module determines whether the reliability of each vertex in each optimized semantic mesh is greater than a set threshold, and performs texture and semantic updates on the vertices with a yes result to obtain the final semantic mesh corresponding to each modified 3D voxel, and obtains the corresponding semantic probability distribution vector based on the final semantic mesh. ; The semantic rendering module is used to render based on each semantic probability distribution vector. Each final semantic grid is rendered separately to achieve highlighting of the target object instance specified in the natural language target query command.

2. The real-time environmental 3D reconstruction system based on semantic interaction as described in claim 1, characterized in that, The semantic segmentation module determines the semantic probability distribution vector of any changed 3D voxel based on the spatial mapping relationship between the 2D image and the LiDAR point cloud. The method is as follows: Get the current number Confidence weight of the frame 2D image relative to the currently changed 3D voxel as follows: in, For the current number The confidence score of the two-dimensional semantic mask corresponding to the frame two-dimensional image. For the cosine term of the current changing three-dimensional voxel observation view, To address the current depth uncertainty of the modified 3D voxels, This is a scaling factor used to adjust the intensity of the depth penalty; Based on confidence weight Get the currently changed 3D voxel relative to the current 3D voxel. semantic probability distribution vector of a frame 2D image as follows: in, For the current change of three-dimensional voxels relative to the first The semantic probability distribution vector of a two-dimensional image frame. For the current number The semantic probability distribution vector of the two-dimensional semantic mask corresponding to the frame two-dimensional image. For the first The cumulative confidence weight value corresponding to the frame. For the current number The cumulative confidence weight value corresponding to the frame, where, The calculation formula is: .

3. The real-time environmental 3D reconstruction system based on semantic interaction as described in claim 1, characterized in that, The semantic segmentation module uses the following method to obtain the two-dimensional semantic mask corresponding to the two-dimensional image of the current frame based on the natural language target query instruction: The Grounding DINO model is used to convert natural language target query instructions into detection boxes of targets contained in the current frame's two-dimensional image; The Segment Anything Model is used to generate a two-dimensional semantic mask based on the target's bounding box.

4. The real-time environmental 3D reconstruction system based on semantic interaction as described in claim 1, characterized in that, The semantic constraint module uses a predefined energy function when optimizing any initial semantic grid. as follows: in, Here is a planar metric function used to constrain the fit between the initial semantic mesh surface and the lidar point cloud. This is a semantic consistency term used to constrain mesh edges of objects with different semantic meanings. For semantic consistency items The corresponding weights This is a smoothness term used to ensure that the normal vectors of adjacent mesh patches remain consistent. For smoothness term The corresponding weights.

5. The real-time environmental 3D reconstruction system based on semantic interaction as described in claim 4, characterized in that, Semantic consistency items The calculation method is as follows: in, For the current initial semantic grid, the first One vertex, For the current initial semantic grid, the first One vertex, As vertices The semantic probability distribution As vertices The semantic probability distribution Let the set of edges between any two vertices in the current initial semantic grid be defined. for and The divergence between them As vertices With vertex The semantic weights between them, and the semantic weights The calculation method is as follows: in, As vertices With vertex The length of the side between them As vertices In the current number Semantic confidence of a frame. As the vertex In the current number Semantic confidence of the frame; simultaneously, semantic confidence The calculation method is as follows: in, To set the scaling factor, As vertices The cumulative weight of the voxel.

6. The real-time 3D environment reconstruction system based on semantic interaction as described in claim 4, characterized in that, Smoothness term The calculation method is as follows: in, For the current initial semantic grid, the first One vertex, As vertices The set of neighboring nodes, For the set of neighboring nodes The first in One vertex, As vertices Surface normal vector, As vertices Surface normal vector, As vertices With vertex Smoothing weights between them, and The calculation method is as follows: in, As vertices With vertex Geometric weights between them As vertices With vertex The semantic similarity between them, and have in, As vertices The semantic probability distribution As vertices The semantic probability distribution for and The divergence between them To set the coefficients.

7. The real-time environmental 3D reconstruction system based on semantic interaction as described in claim 1, characterized in that, The semantic constraint module obtains the optimized semantic mesh corresponding to any changed 3D voxel as follows: The energy function of the initial semantic grid corresponding to the current modified 3D voxel is nonlinearly iterated using the Gauss-Seidel method until the set upper limit of the number of iterations is reached. The semantic grid corresponding to the energy function obtained in the last iteration is then used as the optimized semantic grid corresponding to the current modified 3D voxel.

8. The real-time environmental 3D reconstruction system based on semantic interaction as described in claim 1, characterized in that, The parallel refinement module obtains the reliability of any vertex in any optimized semantic mesh. The method is as follows: in, The photometric noise at the current vertex. For the geometric noise of the current vertex, The luminosity residual at the current vertex. This represents the geometric residual of the current vertex.

9. The real-time environmental 3D reconstruction system based on semantic interaction as described in claim 8, characterized in that, Photometric residual at current vertex and geometric residuals The calculation method is as follows: in, The projection coordinates of the current vertex onto the current frame's 2D image. The pixel color at that location. The projection coordinates of the current vertex across all 2D images in previous frames. The average pixel color at that location. The actual projected depth of the current vertex. This is the reference depth value for the current vertex. A constant is set to prevent division by zero.

10. The real-time environmental 3D reconstruction system based on semantic interaction as described in claim 1, characterized in that, The parallel refinement module performs texture and semantic updates on any vertex for which the judgment result is yes, as follows: in, For the vertex whose current judgment result is yes, in the current i-th The corresponding pixel in the two-dimensional semantic mask of the frame two-dimensional image belongs to the category. The unnormalized cumulative log probability, The vertex whose current judgment result is yes is at the th position. The corresponding pixel in the two-dimensional semantic mask of the frame two-dimensional image belongs to the category. The unnormalized cumulative log probability, For the current number The two-dimensional semantic mask corresponding to the frame two-dimensional image. The vertex whose current judgment result is yes is in Belongs to the category The log-likelihood function measured by reliability, and The calculation method is as follows: in, For the vertex whose current judgment result is yes, in the current i-th Frame reliability, The vertex whose current judgment result is yes is in Belongs to the category The primitive log-likelihood function.