A digital twin method, device and medium based on radar-visual fusion and primitive segmentation
Through the digital twin technology of radar-vision fusion and primitive segmentation, combined with lidar and optical imaging sensors, the problems of error and distortion in digital twin technology are solved, high precision, real-time performance and user experience are improved, and an efficient three-dimensional map loading system is built.
Patent Information
- Application Number
- CN202510586178.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-05-08
AI Technical Summary
Existing digital twin technology has errors and distortions when constructing three-dimensional maps, slow loading speeds, and stuck operations. It cannot meet high-precision and real-time requirements, and traditional methods lead to poor user experience.
The method of radar-visual fusion and primitive segmentation is adopted to collect data through lidar and optical imaging sensors, perform data preprocessing, radar-visual fusion, and primitive segmentation to generate a 2D base map, and use LOD technology to dynamically load data blocks of different resolutions to build a digital twin map.
It improves the accuracy and real-time performance of 3D maps, enhances loading speed and user experience, reduces GPU computing burden, and supports efficient point cloud data loading and map rendering.
Smart Images

Figure CN120101776B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a digital twin method, device and medium based on radar-visual fusion and primitive segmentation. The radar-visual fusion refers to the effective fusion of data acquired by laser radar and optical imaging sensors to improve the system's perception ability of the target, and belongs to the field of digital twin technology. Background Art
[0002] Inspection robots, which integrate multiple sensors and can simulate manual operations, are replacing traditional manual inspections. They are increasingly being used in workplaces like substations, computer rooms, sewage treatment plants, pipeline corridors, and tunnels, where daily inspections are demanding and pose significant risks to personnel. To provide users with a more realistic, intuitive, and convenient way to use the inspection robot's control system, digital twin technology is often used to bridge the real world and the digital virtual world, creating a three-dimensional map of the workspace, enabling the inspection robot's positioning and monitoring.
[0003] However, the existing digital twin technology has limitations. Specifically, the common method to implement digital twin technology is manual modeling and adding visual interaction. Although automatic synchronization of virtual space to physical space is achieved, the constructed three-dimensional map not only has certain errors and distortions, but also generally adopts overall loading, which can easily lead to slow loading speed, stuck operation, and poor user experience.
[0004] Moreover, as users' demands for perception data become increasingly higher, coupled with the fact that high-precision digital twin and other scenario applications have also put forward higher requirements for the accuracy and real-time performance of perception, in order to break through the limitations of existing perception equipment and detection equipment, the emerging radar and vision fusion technology can achieve continuous tracking of inspection robot information by integrating lidar and optical imaging sensors, forming a full-process holographic perception of target characteristics, position, speed, trajectory, behavior and other information.
[0005] Therefore, the present invention urgently needs to develop a digital twin technology based on radar-visual fusion and primitive segmentation to break through the current technical limitations. Summary of the Invention
[0006] In response to the above-mentioned existing technical problems, the present invention provides a digital twin method, device and medium based on radar-visual fusion and primitive segmentation, which completes digital twin modeling through radar-visual fusion and primitive segmentation, and dynamically loads data blocks of different resolutions according to the user's perspective and needs to improve loading efficiency and user experience.
[0007] To achieve the above objectives, the present invention provides a digital twin method based on radar-visual fusion and primitive segmentation, which utilizes a digital twin system mainly composed of mapping equipment, DDT, MinIO, and robots, and includes the following steps:
[0008] Through the mapping equipment equipped with lidar, optical imaging sensor, inertial measurement unit and global satellite positioning system, point cloud data, image data, IMU data and GPS data of the surrounding environment are collected to obtain SLAM data set, which is uploaded to MinIO through the front-end and back-end of DDT in turn;
[0009] The DDT backend performs data preprocessing, radar-visual fusion, primitive segmentation, and 2D base map generation on the SLAM dataset to obtain primitive data. The primitive data is then stored in MinIO in Potree format to build a primitive database for the 3D map.
[0010] Edit the scene metadata through the DDT front-end, trigger the DDT back-end to generate new metadata, and update the metadata database in MinIO;
[0011] The DDT backend groups and manages the metadata of the graphics elements to form data blocks of different resolutions. Each data block includes point cloud data of one or more graphics elements, location information, height information, color information related to the point cloud data, and metadata information used to identify the resolution level of the data block and its location in the 3D map.
[0012] When viewing the digital twin map of the robot's environment through the DDT front-end, the corresponding data blocks are organized and loaded from the graphic database through the octree structure, and the LOD technology is used to dynamically switch data blocks of different resolutions according to the position and viewing angle of the virtual camera. The loaded data blocks are then rendered in three dimensions to obtain the current digital twin map.
[0013] The method of the present invention further comprises: the SLAM data set including: an original file and a description file;
[0014] The original files include point cloud data, image data, IMU data, and GPS data;
[0015] The description file includes external parameters from the laser radar to the optical imaging sensor, and external parameters from the laser radar to the inertial measurement unit.
[0016] The method of the present invention further provides that the data preprocessing includes: cleaning, denoising, deduplication and dedistortion processing of the SLAM data set.
[0017] The method of the present invention further includes: uniformly converting the three-dimensional spatial coordinates of the point cloud data and the image data through the displacement relationship and rotation angle relationship between the laser radar and the optical imaging sensor, and converting and mapping the point cloud data that are not uniform in time and space through the acceleration a and angular velocity ω in the IMU data to obtain a three-dimensional map of the surrounding environment.
[0018] The method of the present invention further comprises: dividing the three-dimensional map into 50m×50m primitives to obtain primitive data of the three-dimensional map; and each primitive is composed of a number of point clouds, each point cloud including relevant position information, height information, and color information.
[0019] The method of the present invention further includes: generating the 2D base map based on the bird's-eye view of the three-dimensional map after the segmentation of the graphic elements, and obtaining graphic element data of the 2D base map.
[0020] Furthermore, the present invention provides that the scene editing includes: copying, deleting and moving point cloud data of the same three-dimensional map, and performing map splicing, map expansion, map alignment, map semantic definition and map navigation definition on different three-dimensional maps.
[0021] The present invention further utilizes LOD technology to dynamically switch data blocks of different resolutions according to the position of the virtual camera and the distance of the viewing angle, including: when the viewing angle is far away, loading low-resolution data blocks according to the LOD technology; when the viewing angle is close, loading high-resolution data blocks according to the LOD technology.
[0022] Secondly, the present invention provides a digital twin device based on radar and visual fusion and primitive segmentation, comprising: a memory and at least one processor, wherein the memory and the at least one processor are interconnected via a line;
[0023] Wherein, a computer program is stored in the memory; and the at least one processor calls the computer program in the memory to execute the steps of the digital twin method based on radar-visual fusion and primitive segmentation.
[0024] Third, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the digital twin method based on radar-visual fusion and primitive segmentation are implemented.
[0025] In summary, the present invention proposes a digital twin technology based on radar and visual fusion and primitive segmentation, which has the following technical advantages:
[0026] 1. The present invention uses a laser radar to scan the target area to collect point cloud data, and at the same time combines it with an optical imaging sensor to enrich the image information of objects around the target area. First, because the laser radar has strong penetration and anti-interference capabilities, and can rotate 360° as the operator moves when collecting data, the mapping efficiency of the present invention is greatly improved. Secondly, because the optical imaging sensor can provide rich target image information and supports 360° panoramic shooting, the target features obtained by the present invention are restored with high fidelity and efficiency. Furthermore, the combination of laser radar and optical imaging sensors can construct a complete and high-precision map, improve real-time performance and accuracy, and enhance robustness.
[0027] 2. The present invention divides point cloud data into primitives for storage, supports dynamic loading of primitives for three-dimensional models, and after dividing the large model, only small batches of primitives within the field of view need to be loaded, thereby improving loading speed and operation smoothness.
[0028] 3. The three-dimensional map constructed by the present invention is easy to update. It only requires replacing some updated primitives, which enables flexible modification of the three-dimensional model and reduces work intensity.
[0029] 4. The three-dimensional map constructed by the present invention not only supports easy export in URL or file format, but can also be quickly embedded in the DDT front end for display, achieving seamless integration and efficient presentation, and also has the ability to conduct rich interactions with other terminals.
[0030] 5. The present invention applies block loading and LOD technology to digital twin maps, enabling the system to dynamically load data blocks of different resolutions according to the user's perspective and needs, and perform three-dimensional rendering on the loaded data blocks, achieving efficient point cloud data loading and map rendering, significantly reducing the computing burden of the GPU, and improving loading efficiency and user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0032] Figure 1 A schematic diagram of the steps provided for an embodiment of the method of the present invention;
[0033] Figure 2 A data flow timing diagram provided for an embodiment of the method of the present invention;
[0034] Figure 3A block diagram of the hardware composition principles of an embodiment of the present invention is provided. DETAILED DESCRIPTION
[0035] In the following description, specific details such as specific system structures and technologies are provided for illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details.
[0036] Example 1: The present invention is based on a digital twin method of radar-visual fusion and primitive segmentation.
[0037] like Figure 1 As shown, this embodiment provides a digital twin method based on radar-visual fusion and primitive segmentation, using a digital twin system composed of mapping equipment, robots, DDT, and MinIO, which is specifically introduced as follows.
[0038] S1, such as Figure 2 As shown in the mapping stage in Figure 1, a mapping device equipped with a lidar, optical imaging sensor, inertial measurement unit, and global satellite positioning system collects point cloud data, image data, IMU data, and GPS data of the surrounding environment to obtain a SLAM (Simultaneous Localization and Mapping) dataset of the surrounding environment. The dataset is then uploaded to MinIO via the DDT front-end and DDT back-end. The specific steps are as follows.
[0039] S1-1. Use mapping equipment to perceive the surrounding environment and obtain a SLAM dataset of the surrounding environment.
[0040] In practice, to construct a comprehensive and accurate SLAM (Simultaneous Localization and Mapping) dataset, we use mapping equipment equipped with lidar, optical imaging sensors, inertial measurement units (IMUs), and global positioning systems. This equipment provides a comprehensive understanding of the surrounding environment and collects the raw data required to build the SLAM dataset.
[0041] It should be noted that these sensors each have their own functions: (1) LiDAR accurately measures the reflection time of the laser beam to obtain accurate distance information from the environment, laying a solid foundation for the robot's positioning and map construction. (2) Optical imaging sensors capture 360° images in all directions, providing rich environmental visual information, helping the robot to deeply understand and perceive the environment. (3) The inertial measurement unit (IMU) integrates accelerometers and gyroscopes to accurately measure linear acceleration and angular velocity, and calculates the robot's posture and position changes through integration. (4) The global positioning system (GPS) provides accurate GPS data, effectively eliminating or reducing error sources such as atmospheric delay, helping the robot achieve centimeter-level positioning accuracy.
[0042] The collected SLAM dataset consists of two parts: the original file and the description file. The original file is stored in the bag file format and contains point cloud data acquired by the lidar, image data acquired by the optical imaging sensor, IMU data acquired by the inertial measurement unit, and GPS data acquired by the global positioning system. The description file uses the yaml file format and records the external parameter information from the lidar to the optical imaging sensor and from the lidar to the inertial measurement unit. Each mapping device usually corresponds to a corresponding description file.
[0043] S1-2. The mapping device uploads the SLAM dataset to the DDT front-end.
[0044] It should be noted that DDT (Dynamic Digital Twins), as a powerful digital twin technology editing and display platform, can collect, store and process massive amounts of data in real time, generate digital twin copies through in-depth data analysis and modeling, and provide strong support for application scenarios such as smart manufacturing, intelligent inspection, status monitoring and early warning analysis.
[0045] During the specific implementation, the mapping device uploads the organized SLAM dataset (including the bag package compressed in zip format and the external reference file in yaml format) to the DDT front end.
[0046] S1-3. The DDT front-end uploads the SLAM dataset to the DDT back-end.
[0047] Finally, the DDT front-end securely and efficiently transmits the received SLAM dataset to the DDT back-end, laying a solid foundation for subsequent data processing, 3D modeling, and robotics applications. This process ensures data integrity and accuracy, providing a strong foundation for building high-quality 3D maps of the environment.
[0048] S2, such as Figure 2As shown in the mapping stage in Figure 1, the DDT backend performs data preprocessing, radar-visual fusion, primitive segmentation, and 2D base map generation on the SLAM dataset to obtain primitive data. The primitive data is then uploaded to MinIO and the primitive database of the 3D map is constructed. The specific steps are as follows.
[0049] S2-1. Perform data preprocessing on the SLAM dataset.
[0050] On the DDT backend, data preprocessing of SLAM datasets is a crucial step. This step aims to eliminate errors and redundancies in the data and ensure the accuracy of subsequent processing and analysis. In practice, data preprocessing includes the following key steps:
[0051] S2-11. Routine Cleaning: First, perform a comprehensive inspection of the raw data in the SLAM dataset to identify and remove data points that are obviously erroneous or abnormal. These erroneous data may be caused by sensor failure, environmental interference, or data transmission errors.
[0052] S2-12, Denoising: After data cleaning, further denoising is performed on the data. Noise data typically manifests as small fluctuations or random deviations between data points, which can interfere with subsequent data analysis and processing. Denoising can smooth data curves and improve data accuracy and reliability.
[0053] S2-13. Deduplication: SLAM datasets sometimes contain duplicate data points or frames. This duplicate data not only increases the complexity of data processing but can also mislead subsequent analysis results. Therefore, during the data preprocessing stage, careful inspection and removal of duplicate data is necessary.
[0054] S2-14. Dedistortion: Point cloud data in SLAM datasets may be distorted due to sensor motion or environmental factors. Distortion can cause deviations in the shape, position, or orientation of the point cloud data, affecting 3D map construction and robot positioning accuracy. Therefore, dedistortion processing is required to restore the true shape and position of the point cloud data.
[0055] As can be seen from the above, when the DDT backend performs data preprocessing on the SLAM dataset, it will sequentially perform routine processing steps such as cleaning, denoising, deduplication, and dedistortion to ensure the accuracy and reliability of the data, laying a solid foundation for subsequent 3D modeling and robotic applications.
[0056] S2-2. Perform radar-visual fusion on the SLAM dataset after data preprocessing.
[0057] The DDT backend adopts a SLAM algorithm based on radar and vision fusion, and performs three-dimensional modeling based on the SLAM dataset after data preprocessing, aiming to build a three-dimensional map of the environment.
[0058] It should be noted that the laser-visual fusion SLAM algorithm, as a cutting-edge technology, is centered around the deep fusion of LiDAR and optical imaging sensor data to achieve simultaneous positioning and map construction. The essence of this technology lies in leveraging the complementarity of multimodal data to significantly improve the system's positioning accuracy and robustness.
[0059] In practice, the coordinates of the point cloud data and image data are precisely fused using the extrinsic parameters between the lidar and optical imaging sensors—the displacement and rotation angle relationships between them. This step ensures that the coordinates of the two data types are uniformly converted in 3D space, laying a solid foundation for subsequent processing.
[0060] Then, using the acceleration a and angular velocity ω contained in the IMU data, through calculus processing, the point cloud data and image data that are not unified in time and space are converted and mapped, solving the inconsistency between the point cloud data and image data caused by the time difference, so that the point cloud and image data can be unified in time and space, providing a strong guarantee for building an accurate three-dimensional map.
[0061] In the above steps, the efficient processing of the radar-visual fusion SLAM algorithm can construct a detailed and accurate 3D map of the environment, providing strong support for the robot's navigation, positioning, and environmental perception. Furthermore, the fusion of optical imaging sensors and lidar can fully utilize their complementary nature, improving the robustness and accuracy of 3D reconstruction. The data after radar-visual fusion can achieve more stable positioning and more accurate mapping in complex environments.
[0062] S2-3. Segment the three-dimensional map into primitives.
[0063] It's important to note that, based on a unified coordinate system, the DDT backend precisely divides the 3D map into a series of 50m x 50m squares, called primitives, and constructs a primitive database for the 3D map. This database contains all primitive information, with each primitive consisting of several point clouds. Each point cloud includes detailed location information (such as x, y, and z coordinates), height information (such as the point cloud's height relative to the ground), and color information (such as RGB values), providing rich data support for subsequent robot navigation, positioning, and environmental understanding.
[0064] In practice, to more effectively manage and utilize the resulting 3D map, the 3D map is segmented into primitives, including the following steps: First, each 3D point (x, y, z) in the 3D map is segmented according to its 2D coordinates (x, y). Second, the x and y coordinates are divided by a preset primitive unit size (in this embodiment, the primitive unit size is defined as 50 meters), resulting in regions of size equal to the primitive unit size × the primitive unit size. These regions are referred to as primitives.
[0065] In this way, all 3D points in the complete 3D map will be classified into corresponding primitives after dividing the x and y coordinates by the primitive unit size. Along the x and y directions of the 3D map, at intervals of 50 meters, the entire 3D map is cut into 50m x 50m local point cloud areas, and these local point cloud areas are constructed into a primitive database for subsequent management and utilization. In particular, the primitive database can be reused during robot inspections. The robot can dynamically load the surrounding primitives for positioning based on its current position, without having to load the entire 3D map as used in traditional technology. This method significantly reduces the amount of data during robot operation and improves inspection efficiency and accuracy.
[0066] S2-4. Generate a 2D base map.
[0067] In the DDT backend, based on the bird's-eye view of the three-dimensional map after segmentation, the three-dimensional map after segmentation is generated into a 2D base map to obtain the metadata of the 2D base map.
[0068] During implementation, the generation process must ensure that the 2D basemap maintains the same primitive composition and spatial structure as the 3D map to facilitate subsequent analysis and processing. Furthermore, the generated 2D basemap is also composed of primitives. These primitives should accurately reflect the corresponding elements in the 3D map, including geometric shapes such as points, lines, and surfaces. Each primitive should contain unique attribute information, such as position, size, and color, to facilitate identification and differentiation in subsequent applications.
[0069] Next, the generated 2D basemap and its metadata are saved in the metadata database. This step is a key link in data management and storage, ensuring the security and accessibility of the 2D basemap and its metadata.
[0070] S2-5. The DDT backend stores the metadata database in the Potree format in MinIO.
[0071] Potree, an open-source point cloud visualization tool based on JavaScript and WebGL, can easily load and render large-scale point cloud data in real time within modern web browsers, significantly improving the efficiency and access speed of point cloud data sharing. Its unique hierarchical storage data structure significantly boosts the efficiency and performance of real-time rendering.
[0072] In practice, the DDT backend stores the primitive database in the advanced Potree format within the MinIO high-performance distributed object storage system, enabling rapid generation and management of cloud-based storage files. Furthermore, thanks to Potree's superior data structure, subsequent robots, after acquiring current GPS data, can rapidly load primitives from the surrounding environment. By layering data, this layered approach creates a rendered, high-precision 3D map, significantly improving the efficiency of loading 3D models using this method.
[0073] S3, such as Figure 2 As shown in the editing phase, the DDT front-end is used to edit the metadata, triggering the DDT back-end to generate new metadata and update the metadata database in MinIO.
[0074] It should be noted that the scene editing includes not only the editing of point cloud data, but also the editing of 3D maps, and the 3D maps support export and interaction.
[0075] First, editing point cloud data involves copying, deleting, and moving point cloud data within the same 3D map. For example, if tables or other objects are temporarily added, moved, or removed from an area, editing the point cloud data can be used to eliminate the need for repeated scanning and reconstruction of the 3D map model.
[0076] Secondly, the editing of different 3D maps includes: map splicing, map expansion, map alignment, map semantic definition, and map navigation definition, which will be introduced one by one below.
[0077] (A) Map stitching refers to combining multiple existing 3D maps to form a new 3D map with a larger scope or more complete content. During this process, the content of the original 3D map remains unchanged.
[0078] During implementation, multiple 3D maps to be stitched together are manually selected based on actual needs, and the spatial transformation relationships between each 3D map, including translation vectors and rotation matrices, are determined. This step relies primarily on human spatial perception and understanding of the scene. By visually observing features in the map, such as corners and stairways, the relative positions and orientations of the different maps are determined. Operations such as translation, rotation, and scaling are then performed on the multiple maps to ensure they can be accurately stitched together, forming a larger or more complete new map, providing strong support for the robot's navigation and path planning.
[0079] For example, if a legged robot's motion space needs to be expanded from floors 1-2 to floors 3-4, the existing 3D maps for floors 1, 2, 3, and 4 can be stitched together to form a new 3D map covering all floors. This stitching process requires aligning the coordinates of these 3D maps one by one. The resulting global 3D map is a fusion of all the local maps, while the 3D maps for floors 1-4 remain unchanged.
[0080] (B) Map expansion refers to the process of scanning new data in the expanded area not covered by the original 3D map and fusing it with the original 3D map to form an expanded 3D map. The expanded map will overwrite (i.e., merge and update) the original 3D map.
[0081] In specific implementation, mapping equipment is first used to fully scan the unscanned expansion area, acquiring point cloud and image data for the expansion area. IMU and GPS data are also recorded to generate a SLAM dataset for the expansion area. The SLAM dataset is then transferred to a digital twin system (such as the DDT backend) for data preprocessing, radar-visual fusion, and primitive segmentation to generate primitive metadata for the 3D map of the expansion area. The 3D map of the expansion area is then merged with the original 3D map, overlaying the expanded map with the existing one and updating the existing primitive database. During the merging process, the expanded 3D map undergoes overall optimization, including re-segmentation (if necessary) to accommodate the new map scale and structure; and global geometric accuracy checks and repairs are performed to ensure the accuracy of object shapes, sizes, and positional relationships within the 3D map.
[0082] (C) Map alignment refers to the process of accurately aligning the old and new 3D maps after spatial information changes to ensure that the robot can reuse the original deployed tasks.
[0083] During implementation, the same space may experience changes in spatial information, such as the temporary addition of a podium or renovations. When a robot attempts to reuse an existing deployment task after this change, errors may exist in the 3D maps scanned by different robots or by the same robot at different times. Aligning the old and new 3D maps allows the robot to reuse the same task, thus avoiding duplicate deployments.
[0084] Specifically, the process begins with a rough manual adjustment phase. Representative feature points or areas in the old and new maps are identified and marked. Coordinate transformation tools are used to perform preliminary translation, rotation, and scaling operations, frequently switching views, and using measurement tools to check feature point relationships for optimization. Next, the algorithmic fine-tuning phase begins. Precise feature matching algorithms (such as ICP and image registration) are used to calculate an optimal coordinate transformation matrix. Error metrics are calculated to assess alignment accuracy. Visual inspection ensures seamless alignment. For any unsatisfactory results, the causes are analyzed and parameters are optimized or coarse adjustments are returned. Finally, the field verification phase begins. Robots are deployed along the planned route to observe navigation accuracy and mission execution, record data, and analyze any issues. Once verified, the aligned maps are officially deployed for mission deployment, with ongoing monitoring and evaluation to ensure long-term stability and effectiveness. Furthermore, the alignment of the old and new 3D maps only applies to the current 3D map in need of adjustment; the reference map remains unchanged.
[0085] (D) Map semantic definition refers to assigning specific semantic information to three-dimensional space through annotation.
[0086] In practice, users use specific tools or interfaces to annotate areas in the 3D map, assigning them specific semantic labels. When clicking on these areas, the system can display the semantic information previously annotated by the user.
[0087] For example, users can label areas as ramps, ponds, and so on, and assign corresponding semantic tags to these areas. When users click or query a labeled area, the system instantly displays the area's semantic information, such as "Ramp here," "Pond there," and so on. Furthermore, this semantic information can be linked with robots or other terminal devices to achieve more intelligent interactions and applications.
[0088] (E) Map navigation refers to the marking of recommended driving routes, non-drivable areas, electronic fences, speed limit areas, and other information on a three-dimensional map to assist the robot in navigation and path planning.
[0089] (E-1) The recommended route is a drivable route automatically calculated by the robot using an algorithm. This route is the optimal path calculated based on a combination of factors, including the current environment, the robot's state, and the task requirements.
[0090] (E-2) No-drivable areas are areas that need to be manually marked during robot mission deployment, such as fixed no-drivable areas like pools. Furthermore, when the robot approaches the boundaries of a no-drivable area, safety logic, such as emergency stops or obstacle avoidance, is activated to avoid these areas and ensure safe operation.
[0091] (E-3) Geo-fences refer to obstacles or areas temporarily placed to restrict a robot's range, improving safety and flexibility, and adapting to different scenarios. For example, when a fence is added during road maintenance, the robot will navigate around the obstacle or reroute to avoid the area. Once the maintenance is complete, the geo-fence can be removed, allowing the robot to traverse the road normally.
[0092] (E-4) Speed Limit Zones are areas where users can set the robot's speed. Setting speed limits protects the robot from damage while improving stability and safety. For example, when encountering complex terrain such as potholes outdoors, the robot's speed can be reduced to maintain a stable center of gravity, allowing it to navigate the pothole smoothly.
[0093] As can be seen from the above, the map navigation definition and related terms cover the key information required for robot navigation and path planning in three-dimensional maps, providing strong support for the robot's autonomous driving.
[0094] Furthermore, if Figure 2 As shown, when the method of the present invention performs scene editing, the digital twin system includes the following working processes:
[0095] S3-1. Request point cloud data: The DDT front-end sends a request to the DDT back-end to obtain point cloud data in Potree format.
[0096] S3-2. Return data URL and editing operations: After receiving the request, the DDT backend processes and returns the URL of the corresponding Potree format point cloud data and the allowed editing operation information to the DDT frontend.
[0097] S3-3. Query point cloud data: The DDT front-end uses the received URL to initiate a query to the MinIO object storage system, requesting the specified Potree format point cloud data.
[0098] S3-4. Loading point cloud data: MinIO responds to the query request, and the DDT front-end loads the required Potree format point cloud data from MinIO.
[0099] S3-5. Edit point cloud data: The DDT front-end uses the provided editing tools or interfaces to perform necessary scene editing operations on the loaded point cloud data.
[0100] S3-6, Upload editing operation: After the scene editing is completed, the DDT front-end uploads the edited point cloud data change information to the DDT back-end.
[0101] S3-7, Download point cloud data: After receiving the editing operation, the DDT backend downloads the corresponding Potree format point cloud data from MinIO as needed for subsequent processing.
[0102] S3-8, Apply editing operations: In scheduled tasks or real-time processing, the DDT backend applies the editing operations uploaded by the frontend to the downloaded Potree format point cloud data to update and process the data.
[0103] S3-9. Upload updated point cloud data: After the editing application is completed, the DDT backend will re-upload the updated Potree format point cloud data to MinIO to ensure that the latest state of the data is saved and shared.
[0104] Furthermore, post-processing of the 3D map constructed by this invention supports scene interaction and the export of usable URLs. These interactions include viewpoint setting, primitive insertion, object insertion, point cloud editing, interactive annotation, navigation annotation, and user customization. The exported URLs can be embedded in connected systems for use.
[0105] When implementing it specifically, Figure 2 As shown, when the DDT, robot, or other terminal requests that the editing operation has been fully applied to the point cloud data, the DD backend queries the corresponding data in MinIO and returns the data. The DDT, robot, or other terminal obtains the primitive data and applies it. For example: when the robot needs to call the primitive map to issue a task, it needs to request data from the DDT, obtain the primitive at the current location and the primitives around it, and then form a primitive map to issue the task. The digital twin system of the present invention includes the following working process:
[0106] C1-1, the DDT front-end, robot or other terminal requests the edited point cloud data in Potree format from the DDT back-end;
[0107] C1-2, the DDT backend returns the URL of the corresponding point cloud data to the DDT frontend, robot or other terminal;
[0108] C1-3, the DDT front-end, robot or other terminal queries the corresponding Potree format point cloud data in MinIO through the URL;
[0109] C1-4, download the corresponding Potree format point cloud data from MinIO by the DDT front-end, robot or other terminal;
[0110] C1-5. Point cloud data in the corresponding Potree format is used by the DDT front end, robot or other terminal.
[0111] like Figure 2 As shown, when the DDT, robot or other terminal requests that the editing operation is not fully applied to the point cloud data, the DDT or robot terminal control system or other terminal can edit the metadata; the edited operation information will be transmitted to the backend, and the backend will call the corresponding data of Minio for processing; the processed data will be updated to Minio and returned to the DDT, robot or other terminal at the same time; after obtaining the processed data, the DDT, robot or other terminal can perform application operations. For example: when the robot needs to perform editing operations on the existing three-dimensional map, such as map splicing, map expansion, etc., after the operation is completed, the front end of the robot control system will send a request instruction to the DDT backend; after receiving the request, the DDT backend will call the corresponding data from MinIO and process it; after the processing is completed, the result will be returned to the robot; then the robot can apply the edited map. The digital twin system of the present invention includes the following working process:
[0112] C2-1, the DDT front-end, robot or other terminal requests the unedited point cloud data in Potree format from the DDT back-end;
[0113] C2-2, the DDT backend queries MinIO for the corresponding point cloud data in Potree format;
[0114] C2-3, the DDT backend downloads the corresponding point cloud data in Potree format from MinIO;
[0115] C2-4, the DDT backend applies editing to the corresponding Potree format point cloud data;
[0116] C2-5. The DDT backend uploads the updated Potree format point cloud data to MinIO.
[0117] C2-6. MinIO returns a successful file upload to the DDT backend.
[0118] C2-7, the DDT backend returns the URL of the corresponding point cloud data to the DDT frontend, robot or other terminal;
[0119] C2-8. The DDT front-end downloads and exports the corresponding point cloud data in Potree format from MinIO, and the robot or other terminal downloads the corresponding point cloud data in Potree format from MinIO.
[0120] C2-9. Point cloud data in the corresponding Potree format is used by the DDT front end, robot or other terminal.
[0121] S4. Group and manage the graphic metadata through the DDT backend to form data blocks of different resolutions; each data block includes: point cloud data of one or more graphic elements, location information, height information, color information related to the point cloud data, and metadata information used to identify the resolution level of the data block and its position in the three-dimensional map.
[0122] During the implementation, the DDT backend precisely segments the 3D map into a series of 50m x 50m blocks, called primitives, based on a unified coordinate system. Primitive data is grouped and managed at 50m x 50m intervals. The primitive's point cloud data, along with its associated location, height, and color information, as well as metadata identifying the block's resolution level and location within the 3D map, are grouped together within the data block, creating data blocks of varying resolutions.
[0123] Then, when viewing the digital twin map of the robot's environment through the DDT front-end, the corresponding data blocks are organized and loaded from the graphic database through the octree structure, and the LOD technology is used to dynamically switch data blocks of different resolutions according to the position and viewing angle of the virtual camera. The loaded data blocks are then rendered in three dimensions to obtain the current digital twin map.
[0124] It should be noted that in digital twin maps, to improve loading efficiency and user experience, point cloud data loading is optimized, using block loading and LOD (Level of Detail) technology to quickly render high-precision 3D maps. Block loading technology divides the digital twin map into multiple small blocks (data blocks) and dynamically loads these small blocks based on the user's perspective and needs.
[0125] In specific implementation, an octree structure is used to spatially partition point cloud data, dividing the point cloud data into multiple levels of data blocks. Each level of data blocks is further subdivided into smaller data blocks. Deeper levels indicate finer data block divisions, with the highest level corresponding to the original point cloud resolution. The data blocks to be loaded are dynamically determined based on the robot's current position and viewing direction. This structure ensures the efficient organization of large-scale point cloud data while enabling dynamic loading requirements related to viewpoints. Furthermore, thanks to POTREE's data structure, layered data loading increases the speed of loading 3D models.
[0126] LOD technology dynamically adjusts the model's level of detail based on the distance between the object and the observer or other relevant factors, in order to optimize performance and visual effects. In digital twin maps, the application of LOD technology enables the system to load data blocks of different resolutions based on whether the user's perspective is zoomed in or out. The user's perspective is zoomed in or out through the virtual camera in the system. A virtual camera is a virtual device that simulates the functions of a real camera and is used to define parameters such as the user's perspective, position, and direction of observing a three-dimensional map scene. In a digital twin environment, the virtual camera allows users to change the perspective through operations such as zooming and dragging, just like using a real camera, thereby dynamically loading data blocks of different resolutions.
[0127] In specific implementation, when the user views the digital twin map of the robot's environment, the corresponding data block is loaded from the primitive database, and the position and viewing angle of the virtual camera are changed by operations such as zooming in and out with the mouse wheel and dragging the map. In addition, LOD technology is used to switch data blocks of different resolutions based on the distance change between the user's perspective and the point cloud when the user operates the virtual camera.
[0128] (1) When the virtual camera's viewing angle is zoomed out: In a digital twin map, zooming out is equivalent to increasing the focal length of the virtual camera. When the user scrolls the mouse wheel upward or pinches outward on a touch-enabled device, the viewing angle is zoomed out, the map gradually becomes smaller, and the displayed range increases, just like zooming out when using a real camera. At this time, there is no need to view overly fine map details, and the system will load low-resolution data blocks based on LOD technology. At this time, the user does not need to view overly fine map details, and low-resolution data blocks are sufficient to meet the needs, while reducing system resource usage and increasing loading speed.
[0129] (2) When the virtual camera's perspective is zoomed in: Zooming in corresponds to zooming out, which is equivalent to reducing the focal length of the virtual camera. When the user scrolls the mouse wheel downward or pinches inward using a two-finger gesture on a touch-enabled device, the perspective is zoomed in, the digital twin map gradually becomes larger, and the area of interest to which the user is focusing is gradually magnified, just like focusing the lens on a closer object when using a real camera. At this time, in order to allow the user to clearly observe the specific objects and terrain features in the environment, the system will load high-resolution data blocks based on LOD technology. High-resolution data blocks can clearly display the specific objects and terrain features in the environment, allowing users to observe and analyze the map in more detail.
[0130] Furthermore, after switching the data block resolution, the loaded data block must be rendered in 3D. 3D rendering refers to rendering the corresponding area or object based on the image, making the rendered effect closer to the real world. In practice, rendering can be performed using image data captured by the optical imaging sensor, and it also supports imported material rendering. Imported material formats are supported: STL, GLB, OBJ, and PLY, with file sizes not exceeding 50MB.
[0131] As can be seen above, the 3D map constructed by this invention supports exporting via URL or file format, allowing for quick embedding and display of rendered 3D maps, as well as interaction with other devices within the 3D map, such as enabling a robot to navigate to a specified location within the 3D map. Furthermore, the 3D map supports multiple viewing angles, including movement, zooming in and out, rotation, roaming, scene switching, and more.
[0132] This allows the robot to precisely navigate to designated locations within a 3D map, and users can freely view the map from multiple perspectives, including movement, zooming in and out, rotation, roaming, and scene switching, for an unprecedented immersive experience. Furthermore, when users view the digital twin map of the robot's environment, the system dynamically loads data blocks of the corresponding resolution by adjusting the position and perspective of the virtual camera based on user actions (such as zooming with the mouse wheel or dragging the map). For example, when the user is far away from the point cloud, the loaded low-resolution data block may contain simplified point cloud data of multiple 50m×50m primitives, which have undergone some compression or sampling to reduce the data volume. However, when the user approaches the point cloud, the loaded high-resolution data block contains more detailed and denser point cloud data, which can present finer map details.
[0133] In summary, the 3D map constructed by the method of the present invention not only supports easy export via URL or file format, but can also be quickly embedded in the DDT front-end for display, achieving seamless integration and efficient presentation, and also enabling rich interaction with other terminals. Furthermore, the application of block loading and LOD technology in digital twin maps enables the system to dynamically load data blocks of different resolutions based on the user's perspective and needs, and perform 3D rendering on the loaded data blocks, thereby achieving efficient point cloud data loading and map rendering.
[0134] Example 2: The present invention is based on a digital twin device of radar-vision fusion and primitive segmentation.
[0135] like Figure 3 As shown, the present invention further provides a digital twin device based on radar-visual fusion and primitive segmentation, comprising: at least one processor, and a memory communicatively connected to the at least one processor; in addition, any other appropriate components may also be included depending on the specific application.
[0136] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the above-mentioned digital twin method based on radar-visual fusion and primitive segmentation.
[0137] The processor may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the device to perform desired functions.
[0138] The memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, or flash memory. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor may execute these instructions to implement the functions of the methods and / or systems described in the various embodiments above.
[0139] Example 3: Computer-readable storage medium of the present invention.
[0140] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the digital twin method based on radar-visual fusion and primitive segmentation as described in the above method embodiment are implemented.
[0141] The instructions may be written in any combination of one or more programming languages to form program code for performing the operations of the embodiments of the present application, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user computing device, as a standalone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0142] The computer-readable storage medium may be any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0143] It should also be pointed out that technical personnel in the relevant field can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0144] Furthermore, the above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present application. Therefore, the present application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0145] Furthermore, this description is not intended to limit the embodiments of the present application to the forms disclosed herein. Although a number of exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof. However, such variations, modifications, alterations, additions, and sub-combinations do not detract from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A digital twin method based on radar-visual fusion and primitive segmentation, characterized in that: The digital twin system, consisting of mapping equipment, DDT, MinIO, and robots, includes the following steps: Through the mapping equipment equipped with lidar, optical imaging sensor, inertial measurement unit and global satellite positioning system, point cloud data, image data, IMU data and GPS data of the surrounding environment are collected to obtain SLAM data set, which is uploaded to MinIO through DDT front-end and DDT back-end in sequence. The DDT backend performs data preprocessing, radar-visual fusion, primitive segmentation, and 2D base map generation on the SLAM dataset to obtain primitive data. The primitive data is then stored in MinIO in Potree format to build a primitive database for the 3D map. Edit the scene metadata through the DDT front-end, trigger the DDT back-end to generate new metadata, and update the metadata database in MinIO; The DDT backend groups and manages the metadata of the graphics elements to form data blocks of different resolutions. Each data block includes point cloud data of one or more graphics elements, location information, height information, color information related to the point cloud data, and metadata information used to identify the resolution level of the data block and its location in the 3D map. When viewing the digital twin map of the robot's environment through the DDT front-end, the corresponding data blocks are organized and loaded from the graphic primitive database through the octree structure. Based on the position and viewing angle of the virtual camera, the LOD technology is used to dynamically switch data blocks of different resolutions. The loaded data blocks are then rendered in three dimensions to obtain the current digital twin map. According to the position of the virtual camera and the distance of the viewing angle, LOD technology is used to dynamically switch data blocks of different resolutions, including: when the viewing angle is far away, low-resolution data blocks are loaded according to LOD technology; When the perspective is zoomed in, high-resolution data blocks are loaded according to the LOD technology.
2. The digital twin method based on radar-visual fusion and primitive segmentation according to claim 1 is characterized in that: The SLAM data set includes: original files and description files; The original files include point cloud data, image data, IMU data, and GPS data; The description file includes external parameters from the laser radar to the optical imaging sensor, and external parameters from the laser radar to the inertial measurement unit.
3. The digital twin method based on radar-visual fusion and primitive segmentation according to claim 2 is characterized in that: The data preprocessing includes: cleaning, denoising, deduplication and dedistortion processing of the SLAM data set.
4. The digital twin method based on radar-visual fusion and primitive segmentation according to claim 3 is characterized in that: The laser-visual fusion includes: uniformly converting the three-dimensional spatial coordinates of point cloud data and image data through the displacement relationship and rotation angle relationship between the laser radar and the optical imaging sensor, and converting and mapping the point cloud data that are not uniform in time and space through the acceleration a and angular velocity ω in the IMU data to obtain a three-dimensional map of the surrounding environment.
5. The digital twin method based on radar-visual fusion and primitive segmentation according to claim 4 is characterized in that: The primitive segmentation includes: segmenting the three-dimensional map into 50m×50m primitives to obtain primitive data of the three-dimensional map; and each primitive is composed of a number of point clouds, and each point cloud includes relevant position information, height information, and color information.
6. The digital twin method based on radar-visual fusion and primitive segmentation according to claim 5 is characterized in that: The generating of the 2D base map includes: generating the 2D base map based on the bird's-eye view of the three-dimensional map after the segmentation of the primitives, and obtaining primitive data of the 2D base map.
7. A digital twin method based on radar-visual fusion and primitive segmentation according to any one of claims 1 to 6, characterized in that: The scene editing includes: copying, deleting and moving point cloud data of the same three-dimensional map, and map splicing, map expansion, map alignment, map semantic definition and map navigation definition of different three-dimensional maps.
8. A digital twin device based on radar-visual fusion and primitive segmentation, characterized in that: include: a memory and at least one processor, wherein the memory and the at least one processor are interconnected via a line; Wherein, a computer program is stored in the memory; the at least one processor calls the computer program in the memory to execute the steps of a digital twin method based on radar-visual fusion and primitive segmentation as described in any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a digital twin method based on radar-visual fusion and primitive segmentation are implemented as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-machine cloud edge collaborative map creation and dynamic digital twinning method and system
CN116030213A
Ground-based synthetic aperture radar rapid three-dimensional terrain rendering method and device
CN116824017A