Three-dimensional real scene reconstruction method based on autonomous cruise of robot and related equipment

Through robot autonomous cruise and 3D Gaussian splashing technology, the problems of inconsistent data acquisition and inefficient computing caused by manual operations in the existing technology are solved, and high-precision, dynamic update indoor three-dimensional reconstruction is achieved.

CN120339522AInactive Publication Date: 2025-07-18SHENZHEN QIHANG TERRITORY TECH CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510589530.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing indoor inspection and three-dimensional modeling methods rely on manual operations, resulting in inconsistent data acquisition, large calculation volume and low efficiency, making it difficult to meet the needs of high-efficiency and high-precision three-dimensional real-life reconstruction, especially in real-time modeling in complex indoor environments.

Method used

Using the method of autonomous cruise of robots, data acquisition and time synchronization and calibration are carried out by carrying multiple sensors, environment maps are built and paths are planned, and 3D Gaussian splashing technology is used to represent and render three-dimensional scenes in the cloud, realizing high-precision three-dimensional reconstruction.

Benefits of technology

It improves data acquisition efficiency and reconstruction accuracy, realizes efficient three-dimensional reconstruction of indoor environments, and supports dynamic updates and interactive three-dimensional model display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339522A_ABST
    Figure CN120339522A_ABST
Patent Text Reader

Abstract

The invention provides a three-dimensional real scene reconstruction method based on autonomous cruise of a robot and related equipment, and the method comprises the steps: carrying out the primary collection of the data of an indoor environment through a plurality of sensors carried on the robot, and carrying out the time synchronization and calibration processing of the collected data, and obtaining a first multi-modal data set; according to the first multi-modal data set, an environment map is constructed, path planning is carried out, and a robot inspection path is obtained; controlling the robot to autonomously cruise in the indoor environment according to the inspection path, and performing secondary data acquisition on the indoor environment through multiple sensors in the cruising process to obtain a second multi-modal data set; and according to the second multi-modal data set, performing three-dimensional scene representation and rendering processing at the cloud through a 3D Gaussian splashing technology to obtain an indoor three-dimensional real scene model. According to the invention, high-precision three-dimensional reconstruction of the indoor environment is realized through autonomous cruise of the robot and by using an efficient 3D Gaussian splashing technology, and the data acquisition efficiency and the reconstruction precision are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional modeling, and particularly to a three-dimensional real-scene reconstruction method and related devices based on autonomous robot cruising. Background Art

[0002] Existing indoor inspection and three-dimensional modeling methods mainly rely on manual operation, and collect environmental data by using hand-held scanning devices or simple automation devices. These traditional methods usually require professional personnel to perform on-site operations. The data collection process is time-consuming, and the data quality is greatly affected by the experience and operation techniques of the operators, making it difficult to achieve standardization and consistency of data collection.

[0003] In addition, traditional three-dimensional reconstruction technologies face problems of large computational amount and low efficiency when dealing with large-scale indoor environments. Existing point cloud reconstruction methods often require a large amount of computing resources, have a slow reconstruction speed, and are difficult to meet the actual requirements of the engineering construction field for high-efficiency and high-precision three-dimensional real-scene reconstruction, especially there are obvious technical bottlenecks in real-time modeling in complex indoor environments. Summary of the Invention

[0004] The main purpose of the present invention is to solve the technical problems that existing indoor inspections rely on manual operations resulting in inconsistent data collection, and traditional three-dimensional reconstruction technologies have a large computational amount and low efficiency; The first aspect of the present invention provides a three-dimensional real-scene reconstruction method based on autonomous robot cruising, and the three-dimensional real-scene reconstruction method based on autonomous robot cruising includes: Performing primary data collection on the indoor environment through multi-sensors mounted on the robot, and performing time synchronization and calibration processing on the primary collected data to obtain a first multi-modal data set; Constructing an environmental map and performing path planning according to the first multi-modal data set to obtain a robot inspection path; Controlling the robot to perform autonomous cruising in the indoor environment according to the inspection path, and performing secondary data collection on the indoor environment through the multi-sensors during the cruising process to obtain a second multi-modal data set; Performing three-dimensional scene representation and rendering processing on the second multi-modal data set in the cloud through 3D Gaussian splashing technology to obtain an indoor three-dimensional real-scene model.

[0005] Optionally, in the first implementation manner of the first aspect of the present invention, the performing primary data collection on the indoor environment through multi-sensors mounted on the robot, and performing time synchronization and calibration processing on the primary collected data to obtain a first multi-modal data set includes: Collecting point cloud data of the indoor environment through a lidar mounted on the robot to obtain original point cloud data; Collect image data of the indoor environment through an RGB camera mounted on the robot to obtain the original image data; Measure the pose of the robot through an inertial measurement unit mounted on the robot to obtain the original pose data; Perform time alignment processing on the original point cloud data, the original image data, and the original pose data according to the time synchronization protocol to obtain time-synchronized data; Perform error calibration on the time-synchronized data according to the automatic calibration and error compensation algorithm to obtain the first multi-modal data set.

[0006] Optionally, in the second implementation manner of the first aspect of the present invention, the constructing an environment map and performing path planning according to the first multi-modal data set to obtain the robot patrol path includes: Perform environment mapping processing on the first multi-modal data set through the SLAM algorithm to obtain a point cloud map of the indoor environment; Analyze and process the point cloud map of the indoor environment to determine the areas and collection points that need to be focused on for collection, and obtain a collection task list; According to the point cloud map of the indoor environment and the collection task list, perform global path planning through the A* search algorithm to obtain a global patrol path; According to the global patrol path and the point cloud map of the indoor environment, perform local path optimization through the dynamic window algorithm, and perform obstacle avoidance processing on potential obstacles to obtain the robot patrol path.

[0007] Optionally, in the third implementation manner of the first aspect of the present invention, the controlling the robot to perform autonomous cruising in the indoor environment according to the patrol path, and performing secondary data collection of the indoor environment through the multi-sensor during the cruising process to obtain the second multi-modal data set includes: Set robot control parameters according to the robot patrol path, regulate the movement of the robot through a robot gait controller, and drive the robot to perform autonomous cruising according to the patrol path; During the autonomous cruising of the robot, collect environmental point cloud data in real time through the lidar, collect environmental color image data through the RGB camera, and collect robot pose information through the inertial measurement unit to obtain the original secondary collection data; Perform Voxel Grid filtering processing on the point cloud data in the original secondary collection data to obtain optimized point cloud data, and accurately align the image data and pose information in the original secondary collection data through a sensor fusion algorithm to obtain visual positioning data; Perform accurate matching and stitching on the optimized point cloud data and the visual positioning data through the iterative closest point algorithm to obtain the second multi-modal data set.

[0008] Optionally, in the fourth implementation of the first aspect of the present invention, the three-dimensional scene representation and rendering process is performed on the cloud through the 3D Gaussian splashing technology according to the second multimodal dataset to obtain the indoor three-dimensional real scene model, including: Perform data preprocessing on the optimized point cloud data and visual positioning data in the second multimodal dataset, and perform data optimization and compression through an edge computing device to obtain transmission-optimized data; Upload the transmission-optimized data to a cloud server, and perform Gaussian kernel distribution modeling on the point cloud in the transmission-optimized data through the 3D Gaussian splashing technology to obtain a three-dimensional Gaussian point cloud representation; According to the three-dimensional Gaussian point cloud representation and the image in the transmission-optimized data, perform scene fitting processing through a differentiable optimization framework, adjust the Gaussian parameters, and obtain an optimized three-dimensional scene representation; Perform real-time update processing of the three-dimensional scene according to the optimized three-dimensional scene representation through an incremental model update algorithm to obtain a dynamically updated three-dimensional model; Render the dynamically updated three-dimensional model through a WebGL rendering engine to generate the interactive indoor three-dimensional real scene model.

[0009] Optionally, in the fifth implementation of the first aspect of the present invention, the performing Gaussian kernel distribution modeling on the point cloud in the transmission-optimized data to obtain a three-dimensional Gaussian point cloud representation includes: Calculate the semantic gradient statistical value of the point cloud in the transmission-optimized data, and determine the boundary blurred Gaussian points in the scene by analyzing the semantic gradient statistical value; Perform boundary adaptive Gaussian splitting processing on the boundary blurred Gaussian points, replace a single large-size Gaussian point with multiple small-size Gaussian points arranged along the object boundary, and obtain a boundary-refined Gaussian point cloud; Apply a Gaussian volume control algorithm to the boundary-refined Gaussian point cloud, redefine the spatial distribution of Gaussian points by adjusting the scale parameter and rotation parameter of Gaussian points, and obtain semantically separated point cloud data; Calculate the mask label value of each Gaussian point in the semantically separated point cloud data by minimizing the difference between the pixel-level mask prediction and the true mask, and classify the Gaussian points as foreground or background according to the set threshold and mask label value to obtain a three-dimensional Gaussian point cloud representation with semantic labels.

[0010] Optionally, in the sixth implementation of the first aspect of the present invention, the performing real-time update processing of the three-dimensional scene according to the optimized three-dimensional scene representation through an incremental model update algorithm to obtain a dynamically updated three-dimensional model includes: Identify Gaussian points in the boundary regions from the optimized three-dimensional scene representation, and filter out Gaussian points with semantic gradient statistical values lower than a preset threshold from the Gaussian points in the boundary regions to obtain a set of tiny fuzzy Gaussian points; Apply a mask consistency test algorithm to the set of tiny fuzzy Gaussian points, calculate the projection mask values of the Gaussian points from each perspective, and remove Gaussian points with projection mask values fluctuating more than a preset fluctuation threshold between different perspectives to obtain a three-dimensional scene representation with optimized boundaries; Perform alternating optimization processing based on the three-dimensional scene representation with optimized boundaries and the newly acquired data in the transmission-optimized data to obtain a three-dimensional scene representation with co-optimized parameters; Process the newly acquired data according to the three-dimensional scene representation with co-optimized parameters through an incremental modeling algorithm, identify the changed regions in the scene, and only update the Gaussian point parameters in the changed regions to obtain a dynamically updated three-dimensional model.

[0011] A second aspect of the present invention provides a three-dimensional real-scene reconstruction device based on robot autonomous cruising. The three-dimensional real-scene reconstruction device based on robot autonomous cruising includes: A data acquisition module, configured to perform primary data acquisition on the indoor environment through multiple sensors mounted on the robot, perform time synchronization and calibration processing on the primary acquired data, and obtain a first multi-modal data set; A path planning module, configured to construct an environmental map and perform path planning based on the first multi-modal data set to obtain a robot inspection path; An autonomous cruising module, configured to control the robot to perform autonomous cruising in the indoor environment according to the inspection path, and perform secondary data acquisition on the indoor environment through the multiple sensors during the cruising process to obtain a second multi-modal data set; A three-dimensional modeling module, configured to perform three-dimensional scene representation and rendering processing through 3D Gaussian splashing technology in the cloud based on the second multi-modal data set to obtain an indoor three-dimensional real-scene model.

[0012] A third aspect of the present invention provides a three-dimensional real-scene reconstruction device based on robot autonomous cruising, including: a memory and at least one processor. Instructions are stored in the memory, and the memory and the at least one processor are interconnected by a line; the at least one processor invokes the instructions in the memory so that the three-dimensional real-scene reconstruction device based on robot autonomous cruising executes the steps of the above-mentioned three-dimensional real-scene reconstruction method based on robot autonomous cruising.

[0013] A fourth aspect of the present invention provides a computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When it runs on a computer, it causes the computer to execute the steps of the above-mentioned three-dimensional real-scene reconstruction method based on robot autonomous cruising.

[0014] The above three-dimensional real-scene reconstruction method and related devices based on robot autonomous cruising perform one-time data acquisition of the indoor environment through multiple sensors mounted on the robot, perform time synchronization and calibration processing on the acquired data to obtain a first multi-modal data set; construct an environment map and perform path planning based on the first multi-modal data set to obtain a robot inspection path; control the robot to perform autonomous cruising in the indoor environment according to the inspection path, and perform secondary data acquisition of the indoor environment through multiple sensors during the cruising process to obtain a second multi-modal data set; perform three-dimensional scene representation and rendering processing through the 3D Gaussian splashing technology in the cloud according to the second multi-modal data set to obtain an indoor three-dimensional real-scene model. The present invention realizes high-precision three-dimensional reconstruction of the indoor environment through robot autonomous cruising and the use of the efficient 3D Gaussian splashing technology, improving the data acquisition efficiency and reconstruction accuracy.

[0015] Other features and advantages of the present invention will be described in the following specification, and part of them will become obvious from the specification, or be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims and drawings.

[0016] To make the above objectives, features and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic diagram of the first embodiment of the three-dimensional real-scene reconstruction method based on robot autonomous cruising in the embodiment of the present invention; Figure 2 It is a schematic diagram of an embodiment of the three-dimensional real-scene reconstruction device based on robot autonomous cruising in the embodiment of the present invention; Figure 3 It is a schematic diagram of an embodiment of the three-dimensional real-scene reconstruction device based on robot autonomous cruising in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0019] As used in the embodiments of the present invention, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include other unlisted steps or units, or may optionally further include other steps or units inherent to these processes, methods, products, or devices.

[0020] For ease of understanding of this embodiment, first, a three-dimensional real-scene reconstruction method based on robot autonomous cruising disclosed in the embodiments of the present invention will be introduced in detail. The multi-agent includes a problem rewriting agent, a document selection agent, an answer generation agent, and a retrieval agent. As Figure 1 shown, this method includes the following steps: 101. Perform primary data acquisition on the indoor environment through multiple sensors mounted on the robot, and perform time synchronization and calibration processing on the primary acquired data to obtain a first multi-modal dataset; In an embodiment of the present invention, the performing primary data acquisition on the indoor environment through multiple sensors mounted on the robot, and performing time synchronization and calibration processing on the primary acquired data to obtain a first multi-modal dataset includes: acquiring point cloud data of the indoor environment through a lidar mounted on the robot to obtain raw point cloud data; acquiring image data of the indoor environment through an RGB camera mounted on the robot to obtain raw image data; measuring the pose of the robot through an inertial measurement unit mounted on the robot to obtain raw pose data; performing time alignment processing on the raw point cloud data, the raw image data, and the raw pose data according to a time synchronization protocol to obtain time-synchronized data; and performing error calibration on the time-synchronized data according to an automatic calibration and error compensation algorithm to obtain a first multi-modal dataset.

[0021] Specifically, in the process of one-time data acquisition, the lidar mounted on the robot first collects point cloud data of the indoor environment to obtain the original point cloud data. The lidar uses a scanning laser rangefinder based on the time-of-flight (TOF) principle. By emitting laser light and receiving the laser signal reflected from the surface of environmental objects, it measures the round-trip time of the light to calculate the distance, thereby obtaining the three-dimensional spatial point cloud information of the indoor environment. The lidar usually rotates and scans at a certain frequency (for example, 10 Hz). Each scan can generate point cloud data of about 300,000 points, which accurately record the geometric shape and spatial position information of indoor objects. The original point cloud data contains the spatial coordinates (x, y, z) of each point and the reflection intensity information, and different material surfaces can also be identified through the analysis of the point density. In a complex indoor environment, the density and quality of the point cloud data are affected by factors such as the material of the object surface, environmental lighting, and the laser incident angle. Therefore, during the acquisition process, the robot will appropriately adjust the emission power and reception sensitivity of the lidar to adapt to different environmental conditions.

[0022] Specifically, the RGB camera mounted on the robot collects image data of the indoor environment to obtain the original image data. The RGB camera is configured with a wide-angle lens (field of view angle of about 90°) and collects color images of the environment at an appropriate frame rate (such as 30 fps), recording visual information such as the texture and color of the environment. The original image data is usually stored in the form of a high-resolution (such as 2048×1536) color image, containing pixel information of the red, green, and blue color channels. During the acquisition process, the exposure time and image gain of the camera are automatically adjusted according to the environmental lighting conditions to ensure clear images are obtained in different lighting environments. The RGB camera is also equipped with automatic white balance and high dynamic range (HDR) technology to cope with uneven indoor lighting conditions and enhance the detail performance of the image in the bright-dark transition area. For low-light environments, the camera will automatically extend the exposure time and increase the ISO sensitivity to ensure an image with sufficient brightness is obtained; while in strong light environments, the exposure time will be shortened to avoid overexposure, so that the collected image data has appropriate brightness and contrast under various lighting conditions.

[0023] Specifically, the robot pose is measured simultaneously by an inertial measurement unit (IMU) mounted on the robot to obtain the raw pose data. The IMU integrates sensors such as a three-axis accelerometer, a three-axis gyroscope, and a three-axis magnetometer. By measuring physical quantities such as the acceleration, angular velocity, and magnetic field direction of the robot, it calculates the position, direction, and attitude information of the robot in space. The raw pose data represents the rotation state of the robot in the form of quaternions or Euler angles and describes the complete spatial pose of the robot in combination with the displacement information. The sampling frequency is usually 200 Hz, much higher than that of other sensors, to capture the rapid motion changes of the robot. The IMU uses temperature compensation technology to reduce sensor drift and suppresses high-frequency noise through a built-in digital filter. For a quadruped robot platform, pose measurement also needs to consider the influence of body vibration on sensor readings and reduces such interference through mechanical damping and algorithm compensation. The raw pose data also includes an estimate of measurement uncertainty, which is used to evaluate the reliability of the pose information and serves as a weight reference in subsequent data fusion.

[0024] Specifically, the raw point cloud data, raw image data, and raw pose data are time-aligned according to the time synchronization protocol to obtain time-synchronized data. The time synchronization protocol uses the Precision Time Protocol (PTP) to provide a unified time reference for all sensors. Through a master-slave clock architecture, this protocol synchronizes the timestamps of each sensor with sub-microsecond precision, solving the time deviation problem caused by inconsistent sampling frequencies of different sensors. In specific implementation, the high-frequency clock of the inertial measurement unit is used as the master clock, and other sensors are used as slave clocks, and the timestamps are synchronized by exchanging network information. The time alignment process uses an interpolation algorithm to map sensor data with different sampling rates onto a unified time axis, including downsampling high-frequency sensor data and interpolating low-frequency sensor data in time. For sudden sensor delays, the time synchronization module also implements a dynamic adjustment window mechanism to adaptively adjust the time matching strategy. In a strong interference environment, the synchronization protocol can also detect and correct clock drift to ensure the time synchronization accuracy during long-term data acquisition, and finally generate a multi-sensor data sequence that logically belongs to the same time point.

[0025] Specifically, error calibration is performed on the time synchronization data according to the automatic calibration and error compensation algorithm to obtain the first multi-modal data set. The automatic calibration process uses an optimization-based method to determine the spatial transformation relationship between sensors, including the rotation matrix and translation vector, to solve the installation error problem between sensors. In the calibration process, the corresponding relationship between the observation data of different sensors is established by identifying feature points (such as corner points and planes) in the environment first, and then the optimal transformation parameters are solved through non-linear optimization to align the sensor data in a unified coordinate system. The error compensation algorithm fuses the inertial data from the IMU and the position estimation of the lidar through a Kalman filter to eliminate the influence of sensor noise and cumulative error. The algorithm performs online estimation and correction of the zero bias and scale factor of the IMU to reduce the drift error during long-term data acquisition. For the RGB camera, polynomial model-based distortion correction is applied to eliminate the radial and tangential distortion caused by the lens. For the lidar, the point cloud distortion caused by the robot's movement is compensated. The first multi-modal data set after error calibration has high consistency in both spatial and temporal dimensions.

[0026] 102. According to the first multi-modal data set, construct an environmental map and perform path planning to obtain the robot patrol path; In an embodiment of the present invention, the constructing an environmental map and performing path planning according to the first multi-modal data set to obtain the robot patrol path includes: performing environmental mapping processing on the first multi-modal data set by the SLAM algorithm to obtain an indoor environmental point cloud map; analyzing and processing the indoor environmental point cloud map to determine the areas and collection points that need to be focused on for collection, obtaining a collection task list; according to the indoor environmental point cloud map and the collection task list, performing global path planning by the A* search algorithm to obtain a global patrol path; according to the global patrol path and the indoor environmental point cloud map, performing local path optimization by the dynamic window algorithm and performing obstacle avoidance processing on potential obstacles to obtain the robot patrol path.

[0027] Specifically, first, the SLAM algorithm is used to perform environmental mapping processing based on the first multi-modal dataset to obtain an indoor environmental point cloud map. This step adopts the SLAM technology based on factor graph optimization to fuse the point cloud data from the lidar with the pose information of the inertial measurement unit. In the specific implementation, the SLAM algorithm identifies geometric features such as planes, corners, and line segments from the point cloud through feature extraction, and establishes the correspondence between the current frame and historical key frames based on these features. During the matching process, the algorithm uses the iterative closest point (ICP) method to calculate the rigid body transformation between point clouds, and combines the pose prior constraints provided by the IMU data to construct a factor graph containing pose nodes and observation edges. Subsequently, the factor graph is solved through sparse non-linear optimization technology to obtain a globally consistent robot trajectory and environmental map. The loop detection technology is also applied during the optimization process. When the robot revisits a known area, additional constraints are added through scene recognition to eliminate cumulative drift. The finally generated indoor environmental point cloud map is stored in a voxelized form, and each voxel contains occupancy probability information. The resolution is usually set to 5 cm, which not only ensures the map accuracy but also controls the data volume, enabling the map to accurately represent the indoor environmental structures such as walls, doors, windows, and furniture.

[0028] Specifically, the indoor environmental point cloud map is analyzed and processed to determine the areas and collection points that need to be focused on, and a collection task list is obtained. This step first applies a region segmentation algorithm to divide the point cloud map into different semantic regions, such as rooms, corridors, and open spaces. The segmentation is achieved based on point cloud density clustering and plane detection, identifying the main structural elements such as walls, floors, and ceilings, and dividing independent room units. The information entropy is calculated for each region to evaluate the integrity and uncertainty of the point cloud data. Regions with high information entropy indicate insufficient data quality and need to be focused on for collection. At the same time, key points in the environment are identified based on structural feature analysis, such as door frames, windows, corners, and equipment positions. The view point quality assessment algorithm is used to evaluate the information gain of each potential collection point from multiple angles, considering factors such as visible range, occlusion degree, and viewing angle. Based on these analysis results, the system generates a collection task list, including the boundary coordinates of each area that needs to be focused on for collection, the positions of the collection points, and the collection priorities. These collection points are distributed at key positions in the environment to ensure that the robot can obtain complete and high-quality environmental data.

[0029] Specifically, based on the indoor environmental point cloud map and the acquisition task list, global path planning is performed through the A* search algorithm to obtain the global inspection path. The A* algorithm is based on heuristic search and combines the actual distance cost and heuristic estimation to find the optimal path. In implementation, first, the point cloud map is converted into a two-dimensional grid map, and each grid is marked as free space or an obstacle. To improve navigation safety, the obstacles are dilated, and the dilation range is determined according to the actual size of the robot, usually set to 1.5 times the radius of the robot. The path planning takes the current position of the robot as the starting point, and sequentially takes the acquisition points in the acquisition task list as the target points, and calculates the shortest path connecting all the acquisition points. The heuristic function of the A* algorithm uses the Euclidean distance, and a priority queue is used to manage the nodes to be expanded during the search process, and the nodes with the smallest total cost estimation are preferentially expanded. The planning process also considers the factor of energy efficiency, and preferentially selects flat areas among similar paths to avoid frequent climbing to save energy. The finally generated global inspection path consists of a series of path points, and these points connect all the acquisition points in the optimal order to form a complete inspection trajectory covering all key areas.

[0030] Specifically, based on the global inspection path and the indoor environmental point cloud map, local path optimization is performed through the dynamic window algorithm, obstacle avoidance processing is performed on potential obstacles, and the robot inspection path is obtained. The dynamic window algorithm considers the kinematic constraints of the robot and searches for the optimal control command in the velocity-time space. During the implementation process, the algorithm first samples and generates multiple sets of feasible velocity control instructions based on the current velocity of the robot, and each set of instructions contains two components: linear velocity and angular velocity. For each set of velocity instructions, the forward simulation is used to predict the motion trajectory of the robot in a short time, and the safety, target orientation, and smoothness of these trajectories are evaluated. The safety evaluation calculates the minimum distance between the trajectory and the obstacle; the target orientation measures the degree to which the trajectory is oriented towards the target point; and the smoothness considers the continuity of the velocity change. By comprehensively weighting the scores of these three aspects, the velocity instruction with the highest score is selected as the current control output. The algorithm continuously executes at a frequency of 10Hz to respond to environmental changes in real time. For suddenly emerging dynamic obstacles, the system uses the front sensors on the robot to detect them in time and immediately adjusts the motion trajectory to avoid the obstacles. The inspection path optimized by the dynamic window algorithm not only follows the overall direction of the global plan but also can flexibly respond to local environmental changes, ensuring that the robot can complete the inspection task safely and efficiently in a complex indoor environment.

[0031] 103. Control the robot to perform autonomous cruising in the indoor environment according to the inspection path, and during the cruising process, perform secondary data acquisition on the indoor environment through multiple sensors to obtain the second multi-modal data set; In one embodiment of the present invention, the robot is controlled to perform autonomous cruising in an indoor environment according to the inspection path. During the cruising process, the multi-sensors are used to perform secondary data collection on the indoor environment, and a second multi-modal data set is obtained, including: setting robot control parameters according to the robot inspection path, regulating the movement of the robot through a robot gait controller, and driving the robot to perform autonomous cruising according to the inspection path; during the autonomous cruising of the robot, real-time collection of environmental point cloud data is performed through the lidar, collection of environmental color image data is performed through the RGB camera, and collection of robot pose information is performed through the inertial measurement unit to obtain original secondary collection data; performing Voxel Grid filtering processing on the point cloud data in the original secondary collection data to obtain optimized point cloud data, and precisely aligning the image data and pose information in the original secondary collection data through a sensor fusion algorithm to obtain visual positioning data; performing precise matching and stitching on the optimized point cloud data and the visual positioning data through the iterative closest point algorithm to obtain the second multi-modal data set.

[0032] Specifically, first, robot control parameters are set according to the robot inspection path, and the movement of the robot is regulated through a robot gait controller to drive the robot to perform autonomous cruising according to the inspection path. This process utilizes the gait control system of the quadruped robot to convert the trajectory points generated by path planning into drive commands for each joint of the robot. The gait controller adopts a hierarchical control architecture, with the top layer performing path tracking, the middle layer generating a gait sequence, and the bottom layer implementing joint motion control. The path tracking controller receives the waypoint sequence of the inspection path, calculates the deviation between the robot and the predetermined trajectory, and generates speed and steering commands. The gait planning layer automatically selects an appropriate gait pattern according to the environmental complexity, adopting a fast gait (such as a trot gait, moving about 0.8 meters per second) in flat areas and switching to a stable gait (such as a diagonal gait, moving about 0.4 meters per second) in complex terrain areas. The joint control layer converts the gait planning into specific joint angle and torque commands, and ensures precise execution through the PID control algorithm. The robot foot-end contact force sensor provides ground feedback information, enabling the control system to adapt to different ground materials, such as carpets, tiles, or wooden floors, etc., to achieve precise path tracking while maintaining stability, with the average trajectory tracking error controlled within ±5 centimeters.

[0033] Specifically, during the autonomous cruise of the robot, the lidar is used to collect environmental point cloud data in real time, the RGB camera is used to collect environmental color image data, and the inertial measurement unit is used to collect the pose information of the robot, obtaining the original secondary acquisition data. The difference between this acquisition process and the primary acquisition is that the secondary acquisition is carried out along the optimized inspection path, and the data acquisition is more targeted and systematic. The lidar collects the three-dimensional point cloud of the environment at a scanning frequency of 15 Hz. Each scan covers a 360° horizontal view angle and a 30° vertical view angle, generating approximately 400,000 measurement points, recording more comprehensive environmental structure information. At the same time, the RGB camera collects environmental images with a resolution of 2K at a frequency of 30 Hz, and adopts the HDR mode to cope with the complex indoor lighting conditions, ensuring clear images can be obtained even in the light and dark boundary areas. The image acquisition angle is carefully designed so that adjacent images have an overlap area of approximately 60%, facilitating subsequent image stitching and 3D reconstruction. The inertial measurement unit records the acceleration, angular velocity, and pose information of the robot at a high frequency of 200 Hz, and adopts the Zero Velocity Update Technology (ZUPT) to correct the integration drift regularly. The data collected by these three sensors are associated through high-precision timestamps, forming the original secondary acquisition data, which is superior to the primary acquisition in terms of spatial coverage and detail capture.

[0034] Specifically, the point cloud data in the original secondary acquisition data is processed by Voxel Grid filtering to obtain optimized point cloud data, and the image data and pose information in the original secondary acquisition data are accurately aligned through a sensor fusion algorithm to obtain visual positioning data. The Voxel Grid filtering process first divides the three-dimensional space into cube grids (voxels) of the same size. The typical voxel size is 5 cm, and then the centroid of all points falling within each voxel is calculated as the representative point of the voxel. This processing method not only reduces the data volume (usually the compression ratio reaches 10:1), but also retains the geometric structure characteristics of the environment. For sparse regions, an adaptive voxel size strategy is adopted, using larger voxels to reduce the influence of noise; for regions with rich details such as edges and corners, smaller voxels are used to retain the structural details. At the same time, the sensor fusion algorithm processes the RGB image and pose information, extracts feature points (using ORB feature descriptors) from the image sequence through visual odometry technology, and tracks the movement of these feature points between consecutive frames. The algorithm synthesizes the pose prior and the visual feature matching results, and estimates the precise pose of the camera through a tightly coupled optimization framework, solving the problems of image deformation and motion blur. The visual positioning data generated by this fusion processing contains accurate internal and external camera parameters and their uncertainty estimates, establishing an accurate mapping relationship between each frame of image and the corresponding three-dimensional position.

[0035] Specifically, the iterative closest point (ICP) algorithm is used to precisely match and stitch the optimized point cloud data and visual positioning data to obtain the second multi-modal dataset. The ICP algorithm finds the best rigid body transformation between two point clouds through iterative optimization. In the implementation, a point-to-plane ICP variant is adopted. In each iteration, the correspondence between the source point cloud and the target point cloud is first established, then the local plane features are extracted, the distance from the point to the plane is calculated as the error metric, and the transformation matrix is solved by minimizing this error. To improve the matching efficiency and robustness, the algorithm introduces a multi-resolution strategy. Initially, the low-resolution point cloud is used to quickly estimate a rough transformation, and then it is refined on the high-resolution point cloud. At the same time, the RANSAC framework is integrated to handle outliers (incorrect matches), enabling the algorithm to work stably in an environment with partial occlusion and dynamic objects. For adjacent scans with low overlap, the algorithm additionally utilizes the constraint conditions provided by the visual positioning data to improve the matching results. The loop detection and global consistency optimization are also applied during the stitching process to eliminate long-distance cumulative errors. The finally generated second multi-modal dataset is a highly consistent data set, including accurately aligned high-quality point cloud data, image data, and their corresponding accurate pose information, which completely records the geometric and visual features of the indoor environment.

[0036] 104. According to the second multi-modal dataset, a three-dimensional scene representation and rendering process is performed in the cloud through 3D Gaussian splashing technology to obtain an indoor three-dimensional real scene model.

[0037] In an embodiment of the present invention, the performing a three-dimensional scene representation and rendering process in the cloud through 3D Gaussian splashing technology according to the second multi-modal dataset to obtain an indoor three-dimensional real scene model includes: performing data preprocessing on the optimized point cloud data and visual positioning data in the second multi-modal dataset, performing data optimization and compression through an edge computing device to obtain transmission-optimized data; uploading the transmission-optimized data to a cloud server, performing Gaussian kernel distribution modeling on the point cloud in the transmission-optimized data through 3D Gaussian splashing technology to obtain a three-dimensional Gaussian point cloud representation; performing scene fitting processing on the basis of the three-dimensional Gaussian point cloud representation and the image in the transmission-optimized data through a differentiable optimization framework, adjusting the Gaussian parameters to obtain an optimized three-dimensional scene representation; performing real-time update processing of the three-dimensional scene according to the optimized three-dimensional scene representation through an incremental model update algorithm to obtain a dynamically updated three-dimensional model; and performing rendering processing on the dynamically updated three-dimensional model through a WebGL rendering engine to generate the interactive indoor three-dimensional real scene model.

[0038] Specifically, first, preprocess the optimized point cloud data and visual positioning data in the second multi-modal dataset. Through the edge computing device, optimize and compress the data to obtain the transmission-optimized data. This process is executed on the edge computing unit carried by the robot, using a multi-core processor and a dedicated neural network acceleration chip for local computing. The data preprocessing first performs semantic segmentation on the point cloud, classifying the environment into categories such as walls, floors, ceilings, and objects, and adopting different compression strategies for different categories. For planar regions (such as walls and floors), apply the plane fitting algorithm to extract geometric parameters, and only retain the plane equation and boundary information. For complex objects, use the octree encoding structure to store the point cloud, and assign different precision levels according to visual importance, retaining higher precision in areas with rich details. For image data, apply the video compression technology based on H.265 encoding, keep the key frames of high quality, and record the corresponding relationship of feature points between key frames. The pose information is represented sparsely, only retaining the complete poses of key frames, and the poses of intermediate frames are restored by interpolation. The entire preprocessing process also includes data screening, removing low-quality or redundant point clouds and images. The finally generated transmission-optimized data volume is significantly smaller than the original data, while retaining the main geometric and visual features of the environment, enabling the data to be efficiently transmitted to the cloud through a limited bandwidth.

[0039] Specifically, after uploading the transmission-optimized data to the cloud server, perform Gaussian kernel distribution modeling on the point cloud in the transmission-optimized data through the 3D Gaussian splatter technology to obtain a three-dimensional Gaussian point cloud representation. This step is executed on the cloud server equipped with GPU acceleration, using the 3D Gaussian splatter algorithm to convert the discrete point cloud into a continuous probability distribution representation. In the specific implementation, first perform density analysis on the point cloud data to determine the initial parameters of the Gaussian kernel for each region. The algorithm assigns 9 parameters to each Gaussian kernel: three-dimensional spatial position coordinates (x, y, z), scale parameters (sx, sy, sz), and rotation quaternion (q). The placement of the initial Gaussian points is based on an adaptive sampling strategy, distributing more dense Gaussian points in regions with rich geometric details (such as edges and corners), and using fewer Gaussian points in flat regions. Each Gaussian point is also associated with transparency and color attributes to describe the visual characteristics of the point. After initialization, the algorithm uses an adaptive split and merge mechanism to dynamically adjust the number of Gaussian points, controlling the model complexity while maintaining the scene representation ability. Overlapping Gaussian points are merged, and new Gaussian points are added to areas with insufficient representation. This representation method converts the traditional discrete point cloud into a continuous probability distribution representation, retaining both the geometric details of the environment and providing higher representation efficiency and rendering performance.

[0040] Specifically, based on the three-dimensional Gaussian point cloud representation and the images in the optimized data transmission, scene fitting processing is performed through a differentiable optimization framework to adjust the Gaussian parameters and obtain an optimized three-dimensional scene representation. The differentiable optimization framework uses an energy minimization-based method and defines a rendering loss function to measure the difference between the synthesized image and the real image. This framework uses the stochastic gradient descent algorithm to iteratively optimize the Gaussian parameters. Each iteration consists of four main steps: rendering a synthetic view from the current Gaussian parameters; calculating the loss between the synthetic view and the real image; computing the gradient of each Gaussian kernel parameter through backpropagation; and updating the Gaussian parameters according to the gradient. The loss function combines the L1 color loss, structural similarity loss, and depth consistency constraint to ensure geometric accuracy while guaranteeing visual quality. During the optimization process, an adaptive learning rate strategy is adopted, using a larger learning rate at the beginning to converge quickly and reducing the learning rate later for fine-tuning. The system also introduces a regularization term to control the shape and distribution of the Gaussian kernel and prevent overfitting. The entire optimization is executed in parallel on a cloud GPU cluster, and the parameter adjustment is completed through multiple iterations. The optimized three-dimensional scene representation consists of precisely calibrated Gaussian parameters, accurately capturing the geometric structure and visual appearance of the environment and supporting high-quality rendering from any perspective.

[0041] Specifically, through the incremental model update algorithm, real-time update processing of the three-dimensional scene is performed based on the optimized three-dimensional scene representation to obtain a dynamically updated three-dimensional model. The incremental update algorithm adopts a difference detection and local optimization strategy, integrating newly acquired data while keeping the modeled area stable. The algorithm first registers the newly acquired data and aligns it with the existing model, and then calculates the geometric and appearance differences between the two. For the detected changed areas (such as object movement, addition, or deletion), the system recalculates the Gaussian parameters of that area while keeping the parameters of the unchanged areas. The update process uses a hierarchical strategy, first quickly detecting large-scale changes at a low resolution and then precisely updating the details at a high resolution. To handle occlusion and lighting changes in dynamic scenes, the algorithm introduces a temporal consistency constraint, separating permanent changes and temporary changes by analyzing observational data at multiple time points. The model update supports multiple modes: the regular automatic update mode performs a full-scene scan and update at fixed intervals; the on-demand update mode focuses on updating specified areas by the user; and the real-time stream update mode dynamically updates the model during continuous robot patrol. The incremental update significantly reduces the consumption of computing resources by only processing the changed parts of the scene, enabling the three-dimensional model to continuously reflect the real-time state of the physical environment.

[0042] Specifically, the dynamically updated 3D model is rendered by the WebGL rendering engine to generate an interactive indoor 3D real scene model. The WebGL rendering engine is implemented based on HTML5 and JavaScript, uses GPU acceleration to achieve high-performance 3D graphics rendering, and can run in mainstream browsers without installing additional plugins. The rendering process adopts the hierarchical LOD (Level of Detail) technology, automatically adjusts the number and accuracy of Gaussian points according to the viewing distance, uses a simplified representation in the distance, and displays full-resolution details nearby to ensure smooth operation on various devices. The engine implements a shader program dedicated to Gaussian splashing, including functions such as Gaussian kernel projection, α blending, and lighting calculation, supports real-time soft shadows and global illumination approximation, and enhances the rendering realism. The user interaction layer provides navigation control functions, supports basic operations such as panning, rotating, and zooming, as well as advanced functions such as path roaming, first-person view, and measurement annotation. The system also integrates the results of spatial semantic analysis, supports object-based interaction, and users can select specific objects to view detailed information or perform virtual operations. To adapt to different network environments, the engine implements adaptive streaming loading, preferentially loads the content within the field of view, and gradually refines the details. The final indoor 3D real scene model is presented in the form of a web application, and users can access and interact through various terminal devices such as computers, tablets, or mobile phones to experience the virtual indoor environment, providing an intuitive 3D visualization tool for engineering construction, design planning, and remote collaboration.

[0043] Further, the Gaussian kernel distribution modeling of the point cloud in the transmission-optimized data to obtain a three-dimensional Gaussian point cloud representation includes: calculating the semantic gradient statistical value of the point cloud in the transmission-optimized data, and determining the boundary-blurred Gaussian points in the scene by analyzing the semantic gradient statistical value; performing boundary-adaptive Gaussian splitting processing on the boundary-blurred Gaussian points, replacing a single large-size Gaussian point with multiple small-size Gaussian points arranged along the object boundary to obtain a boundary-refined Gaussian point cloud; applying the Gaussian volume control algorithm to the boundary-refined Gaussian point cloud, redefining the spatial distribution of Gaussian points by adjusting the scale parameter and rotation parameter of the Gaussian points to obtain semantically separated point cloud data; calculating the mask label value of each Gaussian point in the semantically separated point cloud data by minimizing the difference between the pixel-level mask prediction and the real mask, and classifying the Gaussian points into foreground or background according to the set threshold and mask label value to obtain a three-dimensional Gaussian point cloud representation with semantic labels.

[0044] Specifically, first, calculate the semantic gradient statistical values for the points in the transmission-optimized data, and determine the boundary-blurred Gaussian points in the scene by analyzing the semantic gradient statistical values. This process uses the spatial distribution characteristics and color information of the point cloud data to identify the Gaussian points in the object boundary region. In specific implementation, the system first divides the point cloud into regular voxel grids and calculates the attribute distribution of points within each voxel. The calculation of the semantic gradient statistical values uses a multi-scale gradient operator to analyze the change rates of color and geometric features between adjacent voxels. At the object boundary, these features usually show significant changes, such as color mutations, depth discontinuities, or changes in the normal vector direction. For each Gaussian point, the algorithm comprehensively considers the gradient magnitude, direction consistency, and spatial coherence in its surrounding area to generate the semantic gradient statistic. This statistic is calculated through a weighted formula, and different weights are assigned to the color gradient, depth gradient, and normal gradient respectively to adapt to different types of object boundaries. The identification of boundary-blurred Gaussian points uses an adaptive threshold technique, and the Gaussian points with semantic gradient statistical values within a specific range are marked as boundary-blurred points. This method particularly focuses on those Gaussian points located at the junction of different semantic regions, which usually cover multiple objects and cause blurred boundaries during rendering. The system also combines the size information of the Gaussian points and preferentially marks those points with larger volumes that span the boundary, as they have the most significant impact on the boundary quality.

[0045] Specifically, perform boundary-adaptive Gaussian splitting on the boundary-blurred Gaussian points, replace a single large-size Gaussian point with multiple small-size Gaussian points arranged along the object boundary, and obtain a Gaussian point cloud with refined boundaries. This step uses an iterative splitting strategy to analyze each Gaussian point marked as boundary-blurred and determine the optimal splitting direction and the number of splits. The calculation of the splitting direction is based on the analysis of the local gradient field, aligning the splitting line with the object boundary. The system first estimates the boundary direction by obtaining the principal component analysis of the principal gradient vectors in the local area. Subsequently, perform Gaussian splitting perpendicular to the boundary direction to ensure that the newly generated Gaussian points can fit the object edge more accurately. The splitting process follows the principle of energy conservation, distributing the attributes such as color and transparency of the original Gaussian point to the sub-Gaussian points proportionally, and at the same time adjusting the spatial distribution and scale parameters of the sub-Gaussian points. For complex boundary regions, use a multi-level splitting strategy, first perform a rough split, and then perform refined splitting according to local details. To prevent the increase in computational burden caused by excessive splitting, the system sets a minimum Gaussian size threshold and a maximum number of split constraints. The splitting process also considers the distribution of adjacent Gaussian points to avoid generating overly dense or sparse regions. Through this adaptive splitting process, the large-size Gaussian points that originally covered multiple objects are decomposed into a more refined representation, and each newly generated small-size Gaussian point is more focused on expressing the surface of a single object, thereby making the object boundary sharper and clearer in the three-dimensional representation.

[0046] Specifically, the Gaussian volume control algorithm is applied to the Gaussian point cloud with refined boundaries. By adjusting the scale parameters and rotation parameters of the Gaussian points, the spatial distribution of the Gaussian points is redefined to obtain point cloud data with semantic separation. The Gaussian volume control algorithm aims to optimize the spatial coverage of each Gaussian point and avoid a single Gaussian point from simultaneously representing regions belonging to different semantic classes. In implementation, first, a neighbor graph of the Gaussian points is constructed, and the adjacency relationship is determined based on spatial distance and attribute similarity. For each Gaussian point, the semantic consistency of its neighboring points is analyzed. When semantic inconsistency is detected (such as simultaneously neighboring foreground and background points), the shape parameters of this point are adjusted. The specific adjustments include: compressing the scale parameter of the Gaussian point along the semantic boundary direction, causing the Gaussian point to "shrink" visually away from the boundary; rotating the Gaussian ellipsoid so that its major axis is parallel to the object surface to reduce the volume crossing the boundary; for particularly complex regions, the system also introduces anisotropic scaling so that the Gaussian points have different scales in different directions. The adjustment process uses an iterative optimization method, and after each iteration, the boundary clarity is evaluated through rendering tests until the preset quality standard is reached. Volume control also considers the density balance of the Gaussian points to ensure that the adjusted Gaussian point distribution can accurately represent the object boundary without being overly dense or sparse in some regions. This process ultimately generates point cloud data with semantic separation, where the spatial influence range of each Gaussian point is strictly controlled within the same semantic region, laying the foundation for subsequent semantic labeling.

[0047] Specifically, the mask label value of each Gaussian point in the semantically separated point cloud data is calculated by minimizing the difference between the pixel-level mask prediction and the ground truth mask, and the Gaussian points are classified as foreground or background according to the set threshold and the mask label value, obtaining a three-dimensional Gaussian point cloud representation with semantic labels. This step uses a rendering-based optimization method to project the three-dimensional scene onto a two-dimensional image plane and compare the difference between the rendering result and the reference mask image. In specific implementation, the system first assigns an initial mask label value to each Gaussian point, usually making a rough estimate based on spatial position or color information. Then, a mask rendering pipeline is constructed, and the mask labels of each Gaussian point are projected onto the image plane using the alpha blending principle to generate a mask prediction map. The optimization process adopts a gradient-based method, defining a loss function to measure the difference between the predicted mask and the ground truth mask, including cross-entropy loss and boundary-aware loss. The gradient of each Gaussian point mask label is calculated through backpropagation, and then the label value is updated using the gradient descent method. This process is repeated from multiple perspectives to ensure three-dimensional consistency. To handle the inconsistency between perspectives, a multi-view fusion strategy based on confidence weighting is introduced, and the supervision signals of more reliable perspectives are given higher weights. After optimization, the system classifies the Gaussian points as foreground or background according to the set threshold (usually 0.5). The threshold selection takes into account the characteristics of the application scenario and can be biased towards precision or recall according to needs. The final three-dimensional Gaussian point cloud representation with semantic labels contains geometric, appearance, and semantic information at the same time. Each Gaussian point is clearly classified into a specific semantic category, enabling the entire three-dimensional scene to have the ability of semantic understanding while being visually represented.

[0048] Further, the real-time update process of the three-dimensional scene according to the optimized three-dimensional scene representation by the incremental model update algorithm to obtain a dynamically updated three-dimensional model includes: identifying the Gaussian points in the boundary region from the optimized three-dimensional scene representation, and screening out the Gaussian points with semantic gradient statistical values lower than the preset threshold from the Gaussian points in the boundary region to obtain a set of tiny fuzzy Gaussian points; applying a mask consistency test algorithm to the set of tiny fuzzy Gaussian points, calculating the projected mask values of the Gaussian points under each perspective, and removing the Gaussian points with projected mask values fluctuating greater than the preset fluctuation threshold between different perspectives to obtain a boundary-optimized three-dimensional scene representation; performing an alternating optimization process according to the boundary-optimized three-dimensional scene representation and the newly acquired data in the transmission-optimized data to obtain a three-dimensional scene representation with parameter co-optimization; processing the newly acquired data through an incremental modeling algorithm according to the three-dimensional scene representation with parameter co-optimization, identifying the changed regions in the scene, and only updating the Gaussian point parameters in the changed regions to obtain a dynamically updated three-dimensional model.

[0049] Specifically, in this embodiment, first, Gaussian points in the boundary region are identified from the optimized three-dimensional scene representation, and Gaussian points with semantic gradient statistical values lower than a preset threshold are screened out from the Gaussian points in the boundary region to obtain a set of tiny fuzzy Gaussian points. This process first uses a spatial clustering algorithm to identify the boundary regions in the scene, which are usually located at the junctions of different semantic labels. The boundary region identification adopts a spatial gradient calculation method, constructs a grid in the three-dimensional space, and analyzes the change rate of the semantic labels of adjacent Gaussian points. After identifying the boundary regions, the system calculates the semantic gradient statistical value of each Gaussian point in the region, which characterizes the certainty degree of the semantic attributes of the Gaussian point. The calculation method is to analyze the mask label consistency of the Gaussian point from multiple perspectives and its semantic separation from the surrounding Gaussian points. The system sets a preset threshold (usually in the range of 0.2 - 0.3), and screens out the Gaussian points with semantic gradient statistical values lower than this threshold. These points usually appear as Gaussian points with small sizes and uncertain semantics, and they are often the modeling residues caused by inaccurate masks or complex object boundaries. The screening process also considers the scale information of the Gaussian points, especially focuses on those tiny Gaussian points with scales smaller than the pixel projection scale. These points contribute less visually but may interfere with the semantic boundaries. The finally obtained set of tiny fuzzy Gaussian points contains those small-sized Gaussian points located at the object boundaries but with unclear semantic attribution, and these points need to be further processed to improve the boundary clarity.

[0050] Specifically, the mask consistency test algorithm is applied to the set of tiny fuzzy Gaussian points, the projection mask values of the Gaussian points from each perspective are calculated, and the Gaussian points with projection mask values fluctuating more than a preset fluctuation threshold between different perspectives are removed to obtain a three-dimensional scene representation with optimized boundaries. The mask consistency test algorithm aims to evaluate the semantic performance consistency of each tiny fuzzy Gaussian point from different perspectives. In implementation, the system selects multiple representative perspectives (usually 8 - 12 uniformly distributed perspectives), projects each Gaussian point onto the image planes of these perspectives, and calculates its corresponding mask value. The projection process uses the Gaussian splatter rendering technique, considering the spatial position, scale, and rotation parameters of the Gaussian points to accurately simulate their performance from different perspectives. For each Gaussian point, the system calculates the standard deviation of its mask values from all perspectives as a measure of mask fluctuation. The fluctuation threshold is set based on experimental analysis and is usually 0.15. A value higher than this threshold indicates that there is a significant inconsistency in the semantic attribution of the Gaussian point from different perspectives. This inconsistency is usually caused by mask prediction errors or inaccurate Gaussian point positions. For the Gaussian points with fluctuation values exceeding the threshold, the system directly removes them instead of trying to repair them because these points are often artifacts caused by noise or systematic errors. After the mask consistency test, the remaining Gaussian points have stable semantic attribution from multiple perspectives, forming a three-dimensional scene representation with optimized boundaries, which has clearer and more consistent object boundaries.

[0051] Specifically, based on the newly acquired data in the three-dimensional scene representation optimized according to the boundary and the transmission-optimized data, an alternating optimization process is performed to obtain a three-dimensional scene representation with co-optimized parameters. Alternating optimization is an iterative improvement method that switches back and forth between different parameter sets for optimization to achieve the goal of overall quality improvement. In implementation, the system first aligns the newly acquired data with the existing three-dimensional scene and uses rigid body transformation to convert both to the same coordinate system. Then the alternating optimization process begins: In the first stage, the geometric parameters (position, scale, rotation) of the Gaussian points are fixed, and only the appearance parameters (color, transparency) are optimized. The optimization is driven by the rendering loss, comparing the synthetic image with the real image and adjusting the appearance parameters to minimize the difference between the two. In the second stage, the appearance parameters are fixed, and the geometric parameters are optimized. This stage pays particular attention to the Gaussian points in the boundary region and finely adjusts their spatial distribution to better fit the object contour. In the third stage, the semantic labels and appearance parameters are jointly optimized to make the visual performance consistent with the semantic segmentation. The system iterates cyclically among these three stages, with each stage running a fixed number of steps (usually 50 - 100 steps). The optimization process uses an adaptive learning rate strategy, adopting different learning rates for different parameter types and gradually decreasing them as the iteration progresses. Through this alternating optimization method, the system can improve the visual quality while maintaining geometric stability, and at the same time ensure the coordination between the semantic labels and the visual performance, finally obtaining a three-dimensional scene representation with co-optimized parameters.

[0052] Specifically, through the incremental modeling algorithm, based on the three-dimensional scene representation optimized collaboratively according to parameters, the newly acquired data is processed to identify the changed areas in the scene, and only the Gaussian point parameters in the changed areas are updated to obtain a dynamically updated three-dimensional model. The incremental modeling algorithm adopts a difference detection strategy to avoid reconstructing the entire scene. In implementation, the system first constructs a registration mapping between the old and new data to accurately align the newly acquired data with the existing model. Then, three-layer difference detection is performed: at the macroscopic level, the changed areas of the spatial structure are quickly identified through voxel occupancy comparison; at the mesoscopic level, the changes at the object level are analyzed through feature matching; at the microscopic level, the changes in surface attributes such as color and texture are evaluated. For the detected changed areas, the system classifies them into three categories: newly added areas, deleted areas, and modified areas. For newly added areas, the point cloud is extracted from the newly acquired data, new Gaussian points are initialized, and they are incorporated into the existing model through local optimization. For deleted areas, the corresponding Gaussian points are directly removed. For modified areas, the spatial distribution structure of the Gaussian points is retained, and only their appearance parameters and scale parameters are updated. The update process adopts an incremental training method, and the parameters are fine-tuned using a small learning rate to avoid destroying the existing optimization results. The system also implements a change credibility evaluation mechanism. For areas with slight changes or insufficient evidence, a conservative update strategy is adopted to avoid introducing errors. For dynamically changing objects that change continuously, the system marks their attributes and gives special treatment, such as using timestamps to record their state history. Through this selective update method, the incremental modeling algorithm significantly improves the model maintenance efficiency, enables the three-dimensional scene to be continuously updated with the environmental changes, maintains consistency with the physical world, and finally obtains a dynamically updated three-dimensional model that can reflect the real-time state of the indoor environment.

[0053] In this embodiment, the multi-sensors mounted on the robot are used to collect data from the indoor environment once, and the collected data is subjected to time synchronization and calibration processing to obtain the first multi-modal data set; according to the first multi-modal data set, an environmental map is constructed and path planning is performed to obtain the robot patrol path; the robot is controlled to autonomously cruise in the indoor environment according to the patrol path, and during the cruise, the multi-sensors are used to collect data from the indoor environment a second time to obtain the second multi-modal data set; according to the second multi-modal data set, 3D Gaussian splashing technology is used in the cloud for three-dimensional scene representation and rendering processing to obtain a three-dimensional real scene model of the indoor environment. The present invention realizes high-precision three-dimensional reconstruction of the indoor environment through the robot's autonomous cruise and the use of efficient 3D Gaussian splashing technology, improving the data collection efficiency and reconstruction accuracy.

[0054] The above describes the three-dimensional real scene reconstruction method based on the robot's autonomous cruise in the embodiments of the present invention. Next, the three-dimensional real scene reconstruction device based on the robot's autonomous cruise in the embodiments of the present invention will be described. The three-dimensional real scene reconstruction device based on the robot's autonomous cruise is shown in Figure 2, an embodiment of the three-dimensional real-scene reconstruction device based on robot autonomous cruise in the embodiments of the present invention includes: A data acquisition module 201, configured to perform primary data acquisition on the indoor environment through multiple sensors mounted on the robot, perform time synchronization and calibration processing on the primary acquired data, and obtain a first multi-modal data set; A path planning module 202, configured to construct an environment map and perform path planning according to the first multi-modal data set, and obtain a robot inspection path; An autonomous cruise module 203, configured to control the robot to perform autonomous cruise in the indoor environment according to the inspection path, perform secondary data acquisition on the indoor environment through the multiple sensors during the cruise process, and obtain a second multi-modal data set; A three-dimensional modeling module 204, configured to perform three-dimensional scene representation and rendering processing through the 3D Gaussian splashing technology in the cloud according to the second multi-modal data set, and obtain an indoor three-dimensional real-scene model.

[0055] In the embodiments of the present invention, the three-dimensional real-scene reconstruction device based on robot autonomous cruise runs the above-mentioned three-dimensional real-scene reconstruction method based on robot autonomous cruise. The three-dimensional real-scene reconstruction device based on robot autonomous cruise performs primary data acquisition on the indoor environment through multiple sensors mounted on the robot, performs time synchronization and calibration processing on the acquired data, and obtains a first multi-modal data set; constructs an environment map and performs path planning according to the first multi-modal data set, and obtains a robot inspection path; controls the robot to perform autonomous cruise in the indoor environment according to the inspection path, performs secondary data acquisition on the indoor environment through multiple sensors during the cruise process, and obtains a second multi-modal data set; performs three-dimensional scene representation and rendering processing through the 3D Gaussian splashing technology in the cloud according to the second multi-modal data set, and obtains an indoor three-dimensional real-scene model. The present invention realizes high-precision three-dimensional reconstruction of the indoor environment through robot autonomous cruise and utilizes the efficient 3D Gaussian splashing technology, improving the data acquisition efficiency and reconstruction accuracy.

[0056] Above Figure 2 The three-dimensional real-scene reconstruction device based on robot autonomous cruise in the embodiments of the present invention is described in detail from the perspective of modular functional entities. Next, the three-dimensional real-scene reconstruction device based on robot autonomous cruise in the embodiments of the present invention is described in detail from the perspective of hardware processing.

[0057] Figure 3FIG. 0 is a schematic structural diagram of a three-dimensional real-scene reconstruction device based on robot autonomous cruising provided by an embodiment of the present invention. The three-dimensional real-scene reconstruction device 300 based on robot autonomous cruising may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPU) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 for storing application programs 333 or data 332 (for example, one or more mass storage device terminals). Among them, the memory 320 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the three-dimensional real-scene reconstruction device 300 based on robot autonomous cruising. Further, the processor 310 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the three-dimensional real-scene reconstruction device 300 based on robot autonomous cruising to implement the steps of the above-mentioned three-dimensional real-scene reconstruction method based on robot autonomous cruising.

[0058] The three-dimensional real-scene reconstruction device 300 based on robot autonomous cruising may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 331, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art can understand that Figure 3 The shown structure of the three-dimensional real-scene reconstruction device based on robot autonomous cruising does not limit the three-dimensional real-scene reconstruction device based on robot autonomous cruising provided by the present invention, and may include more or fewer components than shown, or combine some components, or have different component arrangements.

[0059] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium, and when the instructions are run on a computer, the computer is caused to execute the steps of the three-dimensional real-scene reconstruction method based on robot autonomous cruising.

[0060] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, device, or unit can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0061] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0062] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of the present invention.

Claims

1. A three-dimensional real-scene reconstruction method based on autonomous robot cruising, characterized in that, The 3D real - scene reconstruction method based on robot autonomous cruise includes: Using multi - sensors mounted on the robot to perform primary data collection on the indoor environment, and performing time synchronization and calibration processing on the primary - collected data to obtain the first multi - modal data set; Based on the first multi - modal data set, constructing an environmental map and performing path planning to obtain the robot inspection path; Controlling the robot to perform autonomous cruise in the indoor environment according to the inspection path, and during the cruise, using the multi - sensors to perform secondary data collection on the indoor environment to obtain the second multi - modal data set; Based on the second multi - modal data set, performing 3D scene representation and rendering processing through 3D Gaussian splashing technology in the cloud to obtain the indoor 3D real - scene model.

2. The three-dimensional real-scene reconstruction method based on autonomous robot cruising according to claim 1, wherein The step of using multi - sensors mounted on the robot to perform primary data collection on the indoor environment, and performing time synchronization and calibration processing on the primary - collected data to obtain the first multi - modal data set includes: Using the lidar mounted on the robot to collect point - cloud data of the indoor environment to obtain the original point - cloud data; Using the RGB camera mounted on the robot to collect image data of the indoor environment to obtain the original image data; Using the inertial measurement unit mounted on the robot to measure the robot pose to obtain the original pose data; Performing time alignment processing on the original point - cloud data, the original image data, and the original pose data according to the time synchronization protocol to obtain the time - synchronized data; Performing error calibration on the time - synchronized data according to the automatic calibration and error compensation algorithm to obtain the first multi - modal data set.

3. The three-dimensional real-scene reconstruction method based on autonomous robot cruising according to claim 1, wherein The step of based on the first multi - modal data set, constructing an environmental map and performing path planning to obtain the robot inspection path includes: Performing environmental mapping processing on the first multi - modal data set through the SLAM algorithm to obtain the indoor environmental point - cloud map; Analyzing and processing the indoor environmental point - cloud map to determine the areas and collection points that need to be focused on for collection to obtain the collection task list; According to the indoor environmental point - cloud map and the collection task list, performing global path planning through the A* search algorithm to obtain the global inspection path; According to the global inspection path and the indoor environmental point - cloud map, performing local path optimization through the dynamic window algorithm and performing obstacle avoidance processing on potential obstacles to obtain the robot inspection path.

4. The three-dimensional real scene reconstruction method based on robot autonomous cruise according to claim 2, wherein, The step of controlling the robot to perform autonomous cruise in the indoor environment according to the inspection path, and during the cruise, using the multi - sensors to perform secondary data collection on the indoor environment to obtain the second multi - modal data set includes: Setting robot control parameters according to the robot inspection path, regulating the robot movement through the robot gait controller, and driving the robot to perform autonomous cruise according to the inspection path; During the robot autonomous cruise process, using the lidar to collect environmental point - cloud data in real - time, using the RGB camera to collect environmental color image data, and using the inertial measurement unit to collect robot pose information to obtain the original secondary - collected data; Perform Voxel Grid filtering on the point cloud data in the original secondary acquisition data to obtain optimized point cloud data, and accurately align the image data and pose information in the original secondary acquisition data through a sensor fusion algorithm to obtain visual positioning data; Perform accurate matching and stitching on the optimized point cloud data and the visual positioning data through the Iterative Closest Point algorithm to obtain the second multi-modal dataset.

5. The three-dimensional real-scene reconstruction method based on robot autonomous cruise according to claim 1, wherein According to the second multi-modal dataset, perform three-dimensional scene representation and rendering processing through 3D Gaussian splashing technology in the cloud to obtain an indoor three-dimensional real scene model, including: Perform data preprocessing on the optimized point cloud data and visual positioning data in the second multi-modal dataset, and perform data optimization and compression through an edge computing device to obtain transmission-optimized data; Upload the transmission-optimized data to a cloud server, and perform Gaussian kernel distribution modeling on the point cloud in the transmission-optimized data through 3D Gaussian splashing technology to obtain a three-dimensional Gaussian point cloud representation; According to the three-dimensional Gaussian point cloud representation and the image in the transmission-optimized data, perform scene fitting processing through a differentiable optimization framework, and adjust the Gaussian parameters to obtain an optimized three-dimensional scene representation; Perform real-time update processing of the three-dimensional scene according to the optimized three-dimensional scene representation through an incremental model update algorithm to obtain a dynamically updated three-dimensional model; Perform rendering processing on the dynamically updated three-dimensional model through a WebGL rendering engine to generate the interactive indoor three-dimensional real scene model.

6. The three-dimensional real-scene reconstruction method based on robot autonomous cruise according to claim 5, wherein The performing Gaussian kernel distribution modeling on the point cloud in the transmission-optimized data to obtain a three-dimensional Gaussian point cloud representation includes: Calculate the semantic gradient statistical value of the point cloud in the transmission-optimized data, and determine the boundary fuzzy Gaussian points in the scene by analyzing the semantic gradient statistical value; Perform boundary adaptive Gaussian splitting on the boundary fuzzy Gaussian points, and replace a single large-size Gaussian point with multiple small-size Gaussian points arranged along the object boundary to obtain a boundary-refined Gaussian point cloud; Apply a Gaussian volume control algorithm to the boundary-refined Gaussian point cloud, and redefine the spatial distribution of Gaussian points by adjusting the scale parameter and rotation parameter of Gaussian points to obtain semantically separated point cloud data; Calculate the mask label value of each Gaussian point in the semantically separated point cloud data by minimizing the difference between the pixel-level mask prediction and the true mask, and classify the Gaussian points as foreground or background according to the set threshold and mask label value to obtain a three-dimensional Gaussian point cloud representation with semantic labels.

7. The three-dimensional real-scene reconstruction method based on autonomous robot cruising according to claim 6, wherein, The performing real-time update processing of the three-dimensional scene according to the optimized three-dimensional scene representation through an incremental model update algorithm to obtain a dynamically updated three-dimensional model includes: Identify the Gaussian points in the boundary region from the optimized three-dimensional scene representation, and screen out the Gaussian points with semantic gradient statistical values lower than a preset threshold from the Gaussian points in the boundary region to obtain a set of tiny fuzzy Gaussian points; Apply the mask consistency test algorithm to the set of tiny fuzzy Gaussian points, calculate the projection mask values of the Gaussian points from each perspective, and remove the Gaussian points whose projection mask values fluctuate more than a preset fluctuation threshold between different perspectives to obtain a three-dimensional scene representation with optimized boundaries; Perform alternating optimization processing based on the three-dimensional scene representation with optimized boundaries and the newly acquired data in the transmission-optimized data to obtain a three-dimensional scene representation with co-optimized parameters; Process the newly acquired data according to the three-dimensional scene representation with co-optimized parameters through an incremental modeling algorithm, identify the changed areas in the scene, and only update the Gaussian point parameters in the changed areas to obtain a dynamically updated three-dimensional model.

8. A three-dimensional real-scene reconstruction device based on autonomous robot cruising, characterized in that, The three-dimensional real-scene reconstruction device based on robot autonomous cruise includes: A data acquisition module, configured to perform primary data acquisition on the indoor environment through multiple sensors mounted on the robot, and perform time synchronization and calibration processing on the primary acquired data to obtain a first multi-modal data set; A path planning module, configured to construct an environmental map and perform path planning based on the first multi-modal data set to obtain a robot inspection path; An autonomous cruise module, configured to control the robot to perform autonomous cruise in the indoor environment according to the inspection path, and perform secondary data acquisition on the indoor environment through the multiple sensors during the cruise to obtain a second multi-modal data set; A three-dimensional modeling module, configured to perform three-dimensional scene representation and rendering processing on the second multi-modal data set in the cloud through 3D Gaussian splashing technology to obtain an indoor three-dimensional real-scene model.

9. A three-dimensional real-scene reconstruction device based on autonomous robot cruising, characterized in that, The three-dimensional real-scene reconstruction device based on robot autonomous cruise includes: a memory and at least one processor, and instructions are stored in the memory; The at least one processor invokes the instructions in the memory so that the three-dimensional real-scene reconstruction device based on robot autonomous cruise executes the steps of the three-dimensional real-scene reconstruction method based on robot autonomous cruise according to any one of claims 1-7.

10. A computer-readable storage medium having instructions stored thereon, characterized in that, When the instructions are executed by the processor, the steps of the three-dimensional real-scene reconstruction method based on robot autonomous cruise according to any one of claims 1-7 are implemented.

Citation Information

Cited By

  • Scene perception model training method and device, robot control method and robot

    CN120635678A

  • Point cloud completion method based on octree and multi-granularity grid fusion

    CN120997408A

  • An octree and multi-granularity grid fusion-based point cloud completion method

    CN120997408B

  • Self-adaptive 4D Gaussian splashing high-precision three-dimensional reconstruction system and method

    CN121033280A

  • Real-time parking guiding method and system based on three-dimensional digital twinning and AR

    CN121281310A