Digital twinning method and device based on Leiyu fusion and primitive segmentation, and medium
Through the digital twin method based on Leixi Vision fusion and element segmentation, the problems of three-dimensional map error and slow loading speed in the prior art are solved, and the effects of high precision, real-time and efficient loading are achieved.
Patent Information
- Application Number
- CN202510586178.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-08
AI Technical Summary
The existing digital twin technology has errors and distortions when building three-dimensional maps, slow loading speed, lag in operation, and difficult to meet the requirements of high precision and real-time.
Using a digital twin method based on lightning vision fusion and element segmentation, data is collected through lidar and optical imaging sensors, lightning vision fusion and element segmentation are performed, data blocks of different resolutions are loaded dynamically, and loading efficiency is optimized using LOD technology.
It improves the accuracy and real-timeness of three-dimensional maps, improves loading speed and operation fluency, and realizes flexible model updates and efficient data loading and rendering.
Smart Images

Figure CN120101776A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a digital twin method, device and medium based on radar-visual fusion and primitive segmentation. The radar-visual fusion refers to the effective fusion of data acquired by a laser radar and an optical imaging sensor to improve the system's perception of a target, and belongs to the field of digital twin technology. Background Art
[0002] Inspection robots are robots that integrate multiple sensors, can simulate manual operations, and replace traditional manual inspections. They have been gradually applied in workplaces such as substations, computer rooms, sewage treatment plants, pipe corridors, and tunnels, where daily inspections require high intensity and personnel risks are high. In order to enable users to use the inspection robot's control system more realistically, intuitively, and conveniently, digital twin technology is usually used to build a bridge between the real world and the digital virtual world, and to construct a three-dimensional map of the work scene, thereby realizing the positioning and monitoring of the inspection robot.
[0003] However, the existing digital twin technology has limitations. Specifically, the common method to implement digital twin technology is manual modeling and adding visual interaction. Although automatic synchronization of virtual space to physical space is achieved, the constructed three-dimensional map not only has certain errors and distortions, but also generally adopts overall loading, which can easily lead to slow loading speed, operation jams, and poor experience.
[0004] Moreover, as users' demands for perception data become higher and higher, coupled with the fact that scenario applications such as high-precision digital twins have also put forward higher requirements for the accuracy and real-time performance of perception, in order to break through the limitations of existing perception devices and detection equipment, the emerging radar and vision fusion technology can achieve continuous tracking of inspection robot information by integrating lidar and optical imaging sensors, forming a full-process holographic perception of target characteristics, position, speed, trajectory, behavior and other information.
[0005] Therefore, the present invention urgently needs to develop a digital twin technology based on radar-vision fusion and primitive segmentation to break through the current technical limitations. Summary of the invention
[0006] In response to the above-mentioned existing technical problems, the present invention provides a digital twin method, device and medium based on radar and vision fusion and primitive segmentation, which completes digital twin modeling through radar and vision fusion and primitive segmentation, and dynamically loads data blocks of different resolutions according to the user's perspective and needs to improve loading efficiency and user experience.
[0007] To achieve the above object, the present invention provides a digital twin method based on radar-visual fusion and primitive segmentation, using a digital twin system mainly composed of a mapping device, DDT, MinIO, and a robot, including the following steps: Through the mapping equipment equipped with laser radar, optical imaging sensor, inertial measurement unit and global satellite positioning system, point cloud data, image data, IMU data and GPS data of the surrounding environment are collected to obtain SLAM data set, which is uploaded to MinIO through the front-end and back-end of DDT in turn; The DDT backend performs data preprocessing, radar-vision fusion, primitive segmentation, and 2D base map generation on the SLAM data set to obtain primitive data. The primitive data is then stored in MinIO in Potree format to build a primitive database for the 3D map. Edit the scene metadata through the DDT front-end, trigger the DDT back-end to generate new metadata, and update the metadata database in MinIO; The metadata of the graphics elements are grouped and managed through the DDT backend to form data blocks of different resolutions; each data block includes point cloud data of one or more graphics elements, location information, height information, color information related to the point cloud data, and metadata information used to identify the resolution level of the data block and its location in the 3D map; When viewing the digital twin map of the robot's environment through the DDT front-end, the corresponding data blocks are organized and loaded from the graphic metadata database through the octree structure, and the LOD technology is used to dynamically switch data blocks of different resolutions according to the position and viewing angle of the virtual camera. The loaded data blocks are then rendered in three dimensions to obtain the current digital twin map.
[0008] The method of the present invention further comprises: the SLAM data set comprises: an original file and a description file; The original files include point cloud data, image data, IMU data, and GPS data; The description file includes external parameters from the laser radar to the optical imaging sensor and external parameters from the laser radar to the inertial measurement unit.
[0009] In the method of the present invention, the data preprocessing includes: cleaning, denoising, de-duplication and de-distortion processing of the SLAM data set.
[0010] The method of the present invention further comprises: uniformly converting the three-dimensional spatial coordinates of the point cloud data and the image data through the displacement relationship and rotation angle relationship between the laser radar and the optical imaging sensor, and converting and mapping the point cloud data that are not uniform in time and space through the acceleration a and angular velocity ω in the IMU data to obtain a three-dimensional map of the surrounding environment.
[0011] The method of the present invention further comprises: dividing the three-dimensional map into 50m×50m primitives to obtain primitive data of the three-dimensional map; and each primitive is composed of a number of point clouds, and each point cloud includes relevant position information, height information, and color information.
[0012] The method of the present invention further comprises: generating the 2D base map based on the bird's-eye view of the three-dimensional map after the segmentation of the primitives, and obtaining the primitive data of the 2D base map.
[0013] Furthermore, the scene editing of the present invention includes: copying, deleting and moving point cloud data of the same three-dimensional map, and performing map splicing, map expansion, map alignment, map semantic definition and map navigation definition on different three-dimensional maps.
[0014] The present invention further utilizes LOD technology to dynamically switch data blocks of different resolutions according to the position of the virtual camera and the viewing angle, including: when the viewing angle is far away, low-resolution data blocks are loaded according to the LOD technology; when the viewing angle is close, high-resolution data blocks are loaded according to the LOD technology.
[0015] Secondly, the present invention provides a digital twin device based on radar-visual fusion and primitive segmentation, comprising: a memory and at least one processor, wherein the memory and the at least one processor are interconnected via a line; Wherein, a computer program is stored in the memory; and the at least one processor calls the computer program in the memory to execute the steps of the digital twin method based on radar-vision fusion and primitive segmentation.
[0016] Thirdly, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the digital twin method based on radar-vision fusion and primitive segmentation are implemented.
[0017] In summary, the present invention proposes a digital twin technology based on radar-visual fusion and primitive segmentation, which has the following technical advantages: 1. The present invention uses a laser radar to scan the target area to collect point cloud data, and combines optical imaging sensors to enrich the image information of objects around the target area. First, because the laser radar has strong penetration and anti-interference capabilities, and can rotate 360° as the operator moves when collecting data, the mapping efficiency of the present invention is greatly improved. Secondly, because the optical imaging sensor can provide rich target image information and support 360° panoramic shooting, the target features obtained by the present invention are restored with high fidelity and efficiency. Furthermore, the combination of laser radar and optical imaging sensors can construct a complete and high-precision map, improve real-time performance and accuracy, and enhance robustness.
[0018] 2. The present invention divides point cloud data into primitives for storage, supports dynamic loading of primitives for three-dimensional models, and after dividing the large model, only small batches of primitives within the field of view need to be loaded, thereby improving loading speed and operation fluency.
[0019] 3. The three-dimensional map constructed by the present invention is easy to update, and only needs to replace some updated primitives, which realizes flexible modification of the three-dimensional model and reduces work intensity.
[0020] 4. The three-dimensional map constructed by the present invention not only supports easy export in URL or file format, but can also be quickly embedded in the DDT front end for display, achieving seamless integration and efficient presentation, and also has the ability to interact with other terminals in a rich manner.
[0021] 5. The present invention applies block loading and LOD technology to digital twin maps, enabling the system to dynamically load data blocks of different resolutions according to the user's perspective and needs, and perform three-dimensional rendering on the loaded data blocks, thereby achieving efficient point cloud data loading and map rendering, significantly reducing the computational burden of the GPU, and improving loading efficiency and user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0023] Figure 1 A schematic diagram of the steps provided for an embodiment of the method of the present invention; Figure 2 A data flow timing diagram provided for an embodiment of the method of the present invention; Figure 3 A block diagram of the hardware composition principles of an embodiment of the device of the present invention. DETAILED DESCRIPTION
[0024] In the following description, specific details such as specific system structures and technologies are provided for illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application can also be implemented in other embodiments without these specific details.
[0025] Embodiment 1: The digital twin method of the present invention is based on radar-vision fusion and primitive segmentation.
[0026] like Figure 1 As shown, this embodiment provides a digital twin method based on radar-vision fusion and primitive segmentation, using a digital twin system composed of a mapping device, a robot, DDT, and MinIO, which is specifically introduced as follows.
[0027] S1, such as Figure 2 As shown in the mapping stage in Figure 1, the point cloud data, image data, IMU data and GPS data of the surrounding environment are collected through the mapping equipment equipped with laser radar, optical imaging sensor, inertial measurement unit and global satellite positioning system to obtain the SLAM (Simultaneous Localization and Mapping) data set of the surrounding environment, and then uploaded to MinIO through the DDT front-end and DDT back-end. The specific steps are as follows.
[0028] S1-1. Use mapping equipment to perceive the surrounding environment and obtain the SLAM data set of the surrounding environment.
[0029] In the specific implementation, in order to build a detailed and accurate SLAM (Simultaneous Localization and Mapping) data set, a mapping device equipped with laser radar, optical imaging sensor, inertial measurement unit (IMU) and global satellite positioning system is used. The mapping device is used to fully perceive the surrounding environment and collect various raw data required to build the SLAM data set.
[0030] It should be noted that these sensors have their own functions: (1) LiDAR obtains accurate distance information of the environment by accurately measuring the reflection time of the laser beam, laying a solid foundation for the robot's positioning and map construction. (2) Optical imaging sensors capture 360° images in all directions, providing rich environmental visual information, helping robots to deeply understand and perceive the environment. (3) Inertial measurement units (IMUs) integrate accelerometers and gyroscopes to accurately measure linear acceleration and angular velocity, and calculate the robot's posture and position changes through integral calculation. (4) Global Positioning Systems (GPS) provide accurate GPS data, effectively eliminating or reducing error sources such as atmospheric delay, helping robots achieve centimeter-level positioning accuracy.
[0031] In addition, the collected SLAM dataset contains two parts: the original file and the description file. The original file is stored in the bag file format, covering the point cloud data obtained by the lidar, the image data obtained by the optical imaging sensor, the IMU data obtained by the inertial measurement unit, and the GPS data obtained by the global positioning system. The description file uses the yaml file format, which records in detail the external parameter information from the lidar to the optical imaging sensor, and the external parameter information from the lidar to the inertial measurement unit. In addition, each mapping device usually corresponds to a description file.
[0032] S1-2, the mapping device uploads the SLAM dataset to the DDT front end.
[0033] It should be noted that DDT (Dynamic Digital Twins), as a powerful digital twin technology editing and display platform, can collect, store and process massive data in real time, generate digital twin copies through in-depth data analysis and modeling, and provide strong support for application scenarios such as smart manufacturing, smart inspection, status monitoring and early warning analysis.
[0034] During the specific implementation, the mapping device uploads the sorted SLAM data set (including the bag package compressed in zip format and the external reference file in yaml format) to the DDT front end.
[0035] S1-3, the DDT front end uploads the SLAM dataset to the DDT back end.
[0036] Finally, the DDT front-end transmits the received SLAM data set to the DDT back-end securely and efficiently, laying a solid foundation for subsequent data processing, 3D modeling and robot applications. This process ensures the integrity and accuracy of the data, providing a strong guarantee for building high-quality 3D maps of the environment.
[0037] S2, such as Figure 2 As shown in the mapping stage in the figure, the SLAM dataset is preprocessed, radar-vision fused, and primitives are segmented to generate a 2D base map through the DDT backend to obtain primitive data. The primitive data is then uploaded to MinIO and a primitive database for the 3D map is constructed. The specific steps are as follows.
[0038] S2-1. Perform data preprocessing on the SLAM dataset.
[0039] In the DDT backend, data preprocessing of SLAM datasets is a crucial step. This step aims to eliminate errors and redundancies in the data and ensure the accuracy of subsequent processing and analysis. In specific implementation, data preprocessing includes the following key steps: S2-11, General cleaning: First, conduct a comprehensive check on the raw data in the SLAM dataset to identify and remove those data points that are obviously erroneous or abnormal. These erroneous data may be caused by sensor failure, environmental interference, or data transmission errors.
[0040] S2-12, Denoising: Based on data cleaning, further denoising is performed on the data. Noise data usually manifests as small fluctuations or random deviations between data points, which may interfere with subsequent data analysis and processing. Denoising can smooth the data curve and improve the accuracy and reliability of the data.
[0041] S2-13, Deduplication: In SLAM datasets, duplicate data points or data frames sometimes appear. These duplicate data not only increase the complexity of data processing, but may also mislead subsequent analysis results. Therefore, in the data preprocessing stage, these duplicate data need to be carefully checked and removed.
[0042] S2-14, Dedistortion: For the point cloud data in the SLAM dataset, distortion may occur due to the influence of sensor movement or environmental factors. Distortion will cause deviations in the shape, position or direction of the point cloud data, thereby affecting the construction of the 3D map and the positioning accuracy of the robot. Therefore, the point cloud data needs to be dedistorted to restore its true shape and position information.
[0043] From the above, we can see that when the DDT backend performs data preprocessing on the SLAM dataset, it will perform conventional cleaning, denoising, deduplication and dedistortion processing steps in sequence to ensure the accuracy and reliability of the data, laying a solid foundation for subsequent 3D modeling and robotics applications. S2-2. Perform radar-vision fusion on the SLAM dataset after data preprocessing.
[0044] The DDT backend adopts a SLAM algorithm based on radar and vision fusion, and performs three-dimensional modeling based on the SLAM data set after data preprocessing, aiming to build a three-dimensional map of the environment.
[0045] It should be noted that the laser-vision fusion SLAM algorithm, as a cutting-edge technology, is based on the deep fusion of laser radar and optical imaging sensor data to achieve simultaneous positioning and map construction. The essence of this technology is to use the complementarity of multimodal data to significantly improve the positioning accuracy and robustness of the system.
[0046] In the specific implementation, first, the coordinates of the point cloud data and image data are accurately fused using the external parameters between the lidar and the optical imaging sensor, that is, the displacement relationship and rotation angle relationship between them. This step ensures that the coordinates of the two data in three-dimensional space can be uniformly converted, laying a solid coordinate foundation for subsequent processing.
[0047] Then, using the acceleration a and angular velocity ω contained in the IMU data, through calculus processing, the point cloud data and image data that are not unified in time and space are converted and mapped, which solves the inconsistency between the point cloud data and the image data caused by the time difference, making the point cloud and image data unified in time and space, providing a strong guarantee for building an accurate three-dimensional map.
[0048] In the above steps, through the efficient processing of the radar-visual fusion SLAM algorithm, a detailed and accurate three-dimensional map of the environment can be constructed, providing strong support for the robot's navigation, positioning and environmental perception. In addition, the fusion of optical imaging sensors and lidar can make full use of the complementarity of the two, improve the robustness and accuracy of three-dimensional reconstruction, and the data after radar-visual fusion can achieve more stable positioning and more accurate map construction in complex environments.
[0049] S2-3, segmenting the three-dimensional map into primitives.
[0050] It should be noted that based on the unified coordinate system, the DDT backend accurately divides the 3D map into a series of 50m x 50m small blocks, called primitives, and builds a primitive database for the 3D map. This database contains all primitive information, and each primitive is composed of several point clouds. Each point cloud contains detailed location information (such as x, y, z coordinates), height information (such as the height of the point cloud relative to the ground), and color information (such as RGB values), providing rich data support for subsequent robot navigation, positioning, and environmental understanding.
[0051] In the specific implementation, in order to more effectively manage and utilize the obtained three-dimensional map, the three-dimensional map is divided according to the primitives, including the following steps: First, each three-dimensional point (x, y, z) in the three-dimensional map is cut according to its coordinates (x, y) on the two-dimensional plane. Secondly, the x and y coordinates are divided by the preset primitive unit size (in this embodiment, the primitive unit size is actually defined as 50 meters) to obtain a block of areas with a size of the primitive unit size × the primitive unit size, which are called primitives.
[0052] In this way, all 3D points in the complete 3D map will be classified into corresponding primitives after dividing the x and y coordinates by the primitive unit size. Along the x and y directions of the 3D map, at intervals of 50 meters, the entire 3D map is cut into 50m x 50m local point cloud areas, and these local point cloud areas are constructed into a primitive database for subsequent management and utilization. In particular, the primitive database can be reused during robot inspections. The robot can dynamically load the surrounding primitives for positioning based on the current position without loading the entire 3D map in traditional technology. This method significantly reduces the amount of data during robot operation and improves inspection efficiency and accuracy.
[0053] S2-4. Generate a 2D base map.
[0054] At the back end of DDT, based on the bird's-eye view of the three-dimensional map after segmentation of the primitives, the three-dimensional map after segmentation of the primitives is generated into a 2D base map to obtain the primitive data of the 2D base map.
[0055] In the specific implementation, during the generation process, it is necessary to ensure that the 2D base map maintains the same primitive composition and spatial structure as the 3D map for subsequent analysis and processing. In addition, the generated 2D base map is also composed of primitives. These primitives should accurately reflect the corresponding elements in the 3D map, including geometric shapes such as points, lines, and surfaces. Each primitive should contain its unique attribute information, such as position, size, color, etc., so that it can be identified and distinguished in subsequent applications.
[0056] Next, the generated 2D base map and its metadata are saved in the metadata database. This step is a key link in data management and storage, ensuring the security and accessibility of the 2D base map and its metadata.
[0057] S2-5. The DDT backend stores the metadata database in Potree format to MinIO.
[0058] It should be noted that Potree, as an open source point cloud visualization tool based on JavaScript and WebGL, can easily load and render large-scale point cloud data in real time in modern web browsers, significantly improving the sharing efficiency and access speed of point cloud data. Its unique hierarchical storage data structure has injected strong impetus into the efficiency and performance of real-time rendering.
[0059] In the specific implementation, the DDT backend stores the primitive database in the advanced format of Potree in the MinIO high-performance distributed object storage system, realizing the rapid generation and management of cloud storage files. In addition, thanks to the excellent data structure of Potree, the subsequent robot can quickly load the primitives of the surrounding environment after obtaining the current GPS data, and present a rendered high-precision three-dimensional map by loading data in layers, thereby greatly improving the efficiency of loading three-dimensional models in the method of the present invention.
[0060] S3, such as Figure 2 As shown in the editing phase in Figure 1, the DDT front-end is used to edit the scene metadata, triggering the DDT back-end to generate new metadata and update the metadata database in MinIO.
[0061] It should be noted that the scene editing includes not only the editing of point cloud data, but also the editing of three-dimensional maps, and the three-dimensional maps support export and interaction.
[0062] First, the editing of point cloud data includes copying, deleting, moving, etc. of the point cloud data of the same 3D map. For example, when tables or other objects are temporarily added, moved, or removed in a certain area, the editing function of point cloud data can be used, so that there is no need to scan the map repeatedly to rebuild the 3D map model.
[0063] Secondly, the editing of different three-dimensional maps includes: map stitching, map expansion, map alignment, map semantic definition, and map navigation definition, which will be introduced one by one below.
[0064] (A) Map stitching refers to combining multiple existing 3D maps to form a new 3D map with a larger scope or more completeness, and in this process, the original 3D map content remains unchanged.
[0065] In specific implementation, according to actual needs, multiple 3D maps to be spliced are manually selected, and the spatial transformation relationship between each 3D map is determined separately, including the translation vector and rotation matrix. This step mainly relies on manual spatial perception and understanding of the scene. By visually observing the features in the map, such as corners and stairways, the relative position and direction between different maps are determined, and multiple maps are translated, rotated, scaled, and other operations are performed to ensure that they can be accurately spliced together to form a larger or more complete new map, providing strong support for the robot's navigation and path planning.
[0066] For example, when the motion space of a legged robot needs to be expanded from 1-2 floors to 3-4 floors, the existing 3D maps of floors 1, 2, 3, and 4 can be spliced to form a new 3D map covering all floors. The splicing process requires the coordinates of these 3D maps to be aligned in sequence. After splicing, a global 3D map that integrates all local maps is generated, and the 3D maps of floors 1-4 remain unchanged.
[0067] (B) Map expansion refers to the formation of an expanded three-dimensional map by scanning new data and integrating it with the original three-dimensional map in the expanded area not covered by the original three-dimensional map, and the expanded map will cover (i.e. merge and update) the original three-dimensional map.
[0068] In the specific implementation, the mapping equipment is first used to perform a comprehensive scan of the unscanned expansion area to obtain the point cloud data and image data of the expansion area, and the IMU data and GPS data are recorded at the same time to obtain the SLAM data set of the expansion area. The SLAM data set is then transmitted to the digital twin system (such as the DDT backend) for data preprocessing, radar fusion, and primitive segmentation to form the primitive data of the three-dimensional map of the expansion area. The three-dimensional map of the expansion area is then merged with the original three-dimensional map, the original three-dimensional map is overwritten with the expanded map, and the original primitive database is updated. In the merging process, the expanded three-dimensional map is optimized as a whole, including: re-segmentation of primitives (if necessary) to adapt to the new map scale and structure; global geometric accuracy inspection and repair of the three-dimensional map to ensure that the shape, size and position relationship of objects in the three-dimensional map are accurate.
[0069] (C) Map alignment refers to the process of accurately aligning the old and new 3D maps after the spatial information changes to ensure that the robot can reuse the original deployed tasks.
[0070] In the specific implementation, the same space may have changes in spatial information, such as the temporary addition of a podium or decoration. When the spatial information is changed, if the robot wants to reuse the original deployment task, there will be errors in the 3D maps scanned by different robots or the same robot in different time periods. If the old and new 3D maps can be aligned, the robot's task reuse can be achieved, thereby avoiding repeated deployment tasks.
[0071] In detail, we first enter the manual rough adjustment stage, identify and mark the representative feature points or areas in the new and old maps, use the coordinate transformation tool to perform preliminary translation, rotation, and scaling operations, frequently switch views, and use the measurement tool to check the relationship between feature points for optimization and adjustment. Then enter the algorithm precision alignment stage, use the precise feature matching algorithm (such as ICP algorithm, image registration algorithm) to calculate a better coordinate transformation matrix, and calculate the error index, evaluate the alignment accuracy, and then use visual inspection to ensure seamless connection. For unsatisfactory situations, analyze the causes and optimize the parameters or return to rough adjustment correction. Finally, enter the field verification stage, send the robot to drive along the predetermined path, observe the navigation accuracy and task execution, record data, and analyze problems. After verification, the aligned map will be officially applied to task deployment, and continuous monitoring and evaluation will be carried out to ensure the long-term stability and effectiveness of the map. In addition, the alignment of the new and old 3D maps only acts on the 3D map that needs to be adjusted at present, and the reference map will not be changed.
[0072] (D) Map semantic definition refers to the process of assigning specific semantic information to three-dimensional space through annotation.
[0073] In specific implementation, users use specific tools or interfaces to mark various areas in the 3D map and give them specific semantic labels. When clicking on these areas, the system can display the semantic information previously marked by the user.
[0074] For example, users can mark an area as a ramp, a pond, etc., and set corresponding semantic labels for these areas. When users click or query a marked area, the system can instantly display the semantic information of the area, such as "here is a ramp", "there is a pond", etc. In addition, this semantic information can also be linked with robots or other terminal devices to achieve more intelligent interaction and application.
[0075] (E) Map navigation is defined as marking recommended driving routes, non-drivable areas, electronic fences, speed limit areas and other information on a three-dimensional map to assist the robot in navigation and path planning.
[0076] (E-1) The recommended route refers to a drivable route automatically calculated by the robot based on the algorithm. The route is the optimal path calculated based on multiple factors such as the current environment, robot status, and task requirements.
[0077] (E-2) The non-drivable areas are those that need to be manually marked when the robot is deployed, such as fixed non-drivable areas such as pools. In addition, when the robot approaches the boundary of the non-drivable area, safety logic will be activated, such as emergency stop or obstacle avoidance, to avoid these areas and ensure the safe operation of the robot.
[0078] (E-3) The electronic fences mentioned above refer to some additional obstacles or areas that are used to temporarily limit the robot's driving range to improve the robot's driving safety and flexibility and adapt to the needs of different scenarios. For example, when a road section is under maintenance, a fence is added, and the robot will bypass the obstacle or re-plan its route to avoid the area; after the maintenance is completed, the electronic fence can be deleted and the robot can pass through the road section normally.
[0079] (E-4) The speed limit area refers to the area where the user can set the robot's driving speed. By setting the speed limit, the robot is protected from damage and the driving stability and safety are improved. For example, when encountering complex terrain such as potholes outdoors, in order to ensure the smooth operation of the robot body, the robot's driving speed can be reduced to ensure the stability of the center of gravity, so that it can pass through the potholes smoothly.
[0080] From the above, we can see that the map navigation definition and related terms cover the key information required for the robot's navigation and path planning in a three-dimensional map, providing strong support for the robot's autonomous driving.
[0081] Furthermore, if Figure 2 As shown, when the method of the present invention performs scene editing, the digital twin system includes the following working processes: S3-1. Request point cloud data: The DDT front end sends a request to the DDT back end to obtain point cloud data in Potree format.
[0082] S3-2. Return data URL and editing operations: After receiving the request, the DDT backend processes and returns the URL of the corresponding Potree format point cloud data and the allowed editing operation information to the DDT frontend.
[0083] S3-3. Query point cloud data: The DDT front-end uses the received URL to initiate a query to the MinIO object storage system, requesting to obtain the specified Potree format point cloud data.
[0084] S3-4, Loading point cloud data: MinIO responds to the query request, and the DDT front-end loads the required Potree format point cloud data from MinIO.
[0085] S3-5. Edit point cloud data: The DDT front end uses the provided editing tools or interfaces to perform necessary scene editing operations on the loaded point cloud data.
[0086] S3-6, upload editing operation: After the scene editing is completed, the DDT front end uploads the edited point cloud data change information to the DDT back end.
[0087] S3-7, Download point cloud data: After receiving the editing operation, the DDT backend downloads the corresponding Potree format point cloud data from MinIO as needed for subsequent processing.
[0088] S3-8, Apply editing operations: In scheduled tasks or real-time processing, the DDT backend applies the editing operations uploaded by the frontend to the downloaded Potree format point cloud data to update and process the data.
[0089] S3-9, upload updated point cloud data: After the editing application is completed, the DDT backend will re-upload the updated Potree format point cloud data to MinIO to ensure that the latest state of the data is saved and shared.
[0090] In addition, the three-dimensional map constructed by the present invention supports scene addition interaction and export of available URLs during post-processing. The interaction includes perspective setting, primitive insertion, object insertion, point cloud editing, interactive annotation, navigation annotation, and supports user customization. The exported URL can be embedded in the docked system for use.
[0091] When implementing it, Figure 2 As shown, when DDT, robot or other terminal requests that the editing operation has been fully applied to the point cloud data, the DD backend queries the corresponding data in MinIO and returns the data. DDT, robot or other terminal obtains the primitive data and applies it. For example: when the robot needs to call the primitive map to issue a task, it needs to request data from DDT, obtain the primitive at the current location and the primitives around it, and form a primitive map, so that the task can be issued. The digital twin system of the present invention includes the following working process: C1-1. The DDT front-end, robot or other terminal requests the edited point cloud data in Potree format from the DDT back-end; C1-2, the DDT backend returns the URL of the corresponding point cloud data to the DDT frontend, robot or other terminal; C1-3, the DDT front-end, robot or other terminal queries the corresponding point cloud data in Potree format in MinIO through the URL; C1-4, download the corresponding point cloud data in Potree format from MinIO by DDT front-end, robot or other terminal; C1-5. Point cloud data in the corresponding Potree format is used by the DDT front end, robot or other terminal.
[0092] like Figure 2As shown, when the DDT, robot or other terminal requests that the editing operation is not fully applied to the point cloud data, the DDT or robot terminal control system or other terminal can edit the metadata; the edited operation information will be transmitted to the back end, and the back end will call the corresponding data of Minio for processing; the processed data will be updated to Minio and returned to the DDT, robot or other terminal at the same time; after obtaining the processed data, the DDT, robot or other terminal can perform application operations. For example: when the robot needs to perform editing operations on the existing three-dimensional map, such as map splicing, map expansion, etc., after the operation is completed, the front end of the robot control system will send a request instruction to the DDT back end; after receiving the request, the DDT back end will call the corresponding data from MinIO and process it; after the processing is completed, the result will be returned to the robot; after that, the robot can apply the edited map. The digital twin system of the present invention includes the following working process: C2-1, the DDT front-end, robot or other terminal requests the unedited point cloud data in Potree format from the DDT back-end; C2-2, the DDT backend queries MinIO for the corresponding point cloud data in Potree format; C2-3, the DDT backend downloads the corresponding point cloud data in Potree format from MinIO; C2-4, the DDT backend applies and edits the corresponding point cloud data in Potree format; C2-5, the DDT backend uploads the updated point cloud data in Potree format to MinIO; C2-6. MinIO returns the successful upload of the file to the DDT backend; C2-7, the DDT backend returns the URL of the corresponding point cloud data to the DDT frontend, robot or other terminal; C2-8. The DDT front-end downloads and exports the corresponding point cloud data in Potree format from MinIO, and the robot or other terminal downloads the corresponding point cloud data in Potree format from MinIO; C2-9. Point cloud data in the corresponding Potree format is used by the DDT front end, robot or other terminal.
[0093] S4. Group and manage the graphic metadata through the DDT backend to form data blocks of different resolutions; each data block includes: point cloud data of one or more graphic elements, location information, height information, color information related to the point cloud data, and metadata information used to identify the resolution level of the data block and its position in the three-dimensional map.
[0094] In the specific implementation, in the step of segmenting the three-dimensional map, based on the unified coordinate system, the DDT backend accurately segments the three-dimensional map into a series of 50m×50m small squares, called primitives. Now, one primitive is set for every 50m×50m to group and manage the primitive data, and the point cloud data of the primitive, the location information, height information, color information related to the point cloud data, and the metadata information used to identify the resolution level of the data block and the location in the three-dimensional map are collected in the data block, thereby forming data blocks of different resolutions.
[0095] Then, when viewing the digital twin map of the robot's environment through the DDT front-end, the corresponding data blocks are organized and loaded from the graphic database through the octree structure, and the LOD technology is used to dynamically switch data blocks of different resolutions according to the position and viewing angle of the virtual camera. The loaded data blocks are then rendered in three dimensions to obtain the current digital twin map.
[0096] It should be noted that in the digital twin map, in order to improve loading efficiency and user experience, optimize the point cloud data loading method, use block loading and LOD (Level of Detail) technology, and quickly render high-precision three-dimensional maps. Among them, block loading technology refers to dividing the digital twin map into multiple small blocks (data blocks) and dynamically loading these small blocks of data according to the user's perspective and needs.
[0097] In the specific implementation, the octree structure is used to spatially divide the point cloud data, so that the point cloud data is divided into multiple levels of data blocks, and each level of data blocks is further subdivided into smaller data blocks. The deeper the level, the finer the data block division, and the highest level corresponds to the original point cloud resolution. The data blocks to be loaded are dynamically determined according to the current position and viewing direction of the robot. This structure not only ensures the efficient organization of large-scale point cloud data, but also realizes the dynamic loading requirements related to the viewpoint. In addition, thanks to the data structure of POTREE, the rate of loading 3D models is improved by loading data in layers.
[0098] LOD technology is a technology that dynamically adjusts the level of detail of the model according to the distance between the object and the observer or other relevant factors, in order to optimize performance and visual effects. In the digital twin map, the application of LOD technology enables the system to load data blocks of different resolutions according to the user's perspective. The user's perspective is zoomed in or out through the virtual camera in the system. A virtual camera is a virtual device that simulates the function of a real camera and is used to define parameters such as the user's perspective, position, and direction of observing a three-dimensional map scene. In a digital twin environment, a virtual camera enables users to dynamically load data blocks of different resolutions by changing the perspective through operations such as zooming and dragging, just like using a real camera.
[0099] In specific implementation, when the user views the digital twin map of the robot's environment, the corresponding data block is loaded from the primitive database, and the position and viewing angle of the virtual camera are changed through operations such as zooming in and out with the mouse wheel and dragging the map. The LOD technology is used to switch data blocks of different resolutions based on the distance change between the user and the point cloud when operating the virtual camera, that is, based on the user's perspective.
[0100] (1) When the virtual camera's viewing angle is zoomed out: In the digital twin map, zooming out is equivalent to increasing the focal length of the virtual camera. When the user scrolls the mouse wheel upward or expands outward through a two-finger pinch gesture on a touch-enabled device, the viewing angle is zoomed out, the map gradually becomes smaller, and the displayed range increases, just like zooming out when using a real camera. At this time, there is no need to view overly fine map details, and the system will load low-resolution data blocks based on LOD technology. At this time, users do not need to view overly fine map details, and low-resolution data blocks are sufficient to meet their needs. At the same time, it can reduce system resource usage and increase loading speed.
[0101] (2) When the virtual camera zooms in: Zooming in corresponds to zooming out, which is equivalent to reducing the focal length of the virtual camera. When the user scrolls the mouse wheel down or pinches inwards with a two-finger pinch gesture on a touch-enabled device, the perspective zooms in, the digital twin map gradually becomes larger, and the area of interest to which the user is concerned is gradually enlarged, just like focusing the lens on a closer object when using a real camera. At this time, in order to allow users to clearly observe specific objects and terrain features in the environment, the system will load high-resolution data blocks based on LOD technology. High-resolution data blocks can clearly display specific objects and terrain features in the environment, allowing users to observe and analyze the map in more detail.
[0102] Furthermore, after switching the resolution of the data block, the loaded data block needs to be rendered in three dimensions. Three-dimensional rendering refers to rendering the corresponding area or object according to the image, so that the rendered effect is close to the real world. In specific implementation, rendering can render the image data collected by the optical imaging sensor, and also supports imported material rendering. The formats of imported materials are supported: STL, GLB, OBJ, PLY, and the file size does not exceed 50MB.
[0103] As can be seen from the above, the three-dimensional map constructed by the present invention supports URL or file format export, can quickly embed and display the rendered three-dimensional map, or interact with other terminals in the three-dimensional map, such as supporting robots to drive to the specified location in the three-dimensional map. In addition, the three-dimensional map supports moving, zooming in, zooming out, rotating, roaming, scenes, switching and other different perspectives for viewing.
[0104] In this way, the robot can accurately drive to the designated location in the three-dimensional map, and the user can freely view the map from multiple different perspectives such as moving, zooming in, zooming out, rotating, roaming, and scene switching, and enjoy an unprecedented immersive experience. In addition, when the user views the digital twin map of the robot's environment, the system will change the position and perspective of the virtual camera according to the user's operation (such as mouse wheel zooming, dragging the map, etc.), thereby dynamically loading the data block of the corresponding resolution. For example, when the user is far away from the point cloud, the loaded low-resolution data block may contain simplified point cloud data of multiple 50m×50m primitives, which have undergone certain compression or sampling processing to reduce the amount of data; when the user is close to the point cloud, the loaded high-resolution data block contains more detailed and denser point cloud data, which can present finer map details.
[0105] In summary, the three-dimensional map constructed by the method of the present invention not only supports easy export in URL or file format, but also can be quickly embedded in the DDT front end for display, achieving seamless integration and efficient presentation, and also has the ability to interact with other terminals. In addition, the application of block loading and LOD technology in digital twin maps enables the system to dynamically load data blocks of different resolutions according to the user's perspective and needs, and perform three-dimensional rendering on the loaded data blocks, thereby achieving efficient point cloud data loading and map rendering.
[0106] Embodiment 2: The present invention is a digital twin device based on radar-vision fusion and primitive segmentation.
[0107] like Figure 3 As shown, the present invention further provides a digital twin device based on radar-vision fusion and primitive segmentation, comprising: at least one processor, and a memory communicatively connected to the at least one processor; in addition, any other appropriate components may also be included depending on the specific application scenario.
[0108] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the above-mentioned digital twin method based on radar-vision fusion and primitive segmentation.
[0109] The processor may be a central processing unit (CPU) or other forms of processing units with data processing capability and / or instruction execution capability, and may control other components in the device to perform desired functions.
[0110] The memory may include one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor may execute the instructions to implement the methods of the various embodiments described above and / or the functions of the system thereof.
[0111] Embodiment 3: Computer readable storage medium of the present invention.
[0112] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the digital twin method based on radar-vision fusion and primitive segmentation as described in the above method embodiment are implemented.
[0113] The instructions may be written in any combination of one or more programming languages to form program codes for executing the operations of the embodiments of the present application, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user computing device, as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0114] The computer readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can include, for example, but is not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0115] It should also be pointed out that technicians in the relevant field can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0116] Furthermore, the above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present application. Therefore, the present application is not intended to be limited to the aspects shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
[0117] Furthermore, this description is not intended to limit the embodiments of the present application to the forms disclosed herein. Although a number of exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof. However, these variations, modifications, changes, additions, and sub-combinations do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A digital twin method based on radar-visual fusion and primitive segmentation, characterized in that: The digital twin system mainly consists of mapping equipment, DDT, MinIO, and robots, including the following steps: Through the mapping equipment equipped with laser radar, optical imaging sensor, inertial measurement unit and global satellite positioning system, the point cloud data, image data, IMU data and GPS data of the surrounding environment are collected to obtain the SLAM data set, which is then uploaded to MinIO via the DDT front-end and DDT back-end in turn; The DDT backend performs data preprocessing, radar-vision fusion, primitive segmentation, and 2D base map generation on the SLAM data set to obtain primitive data. The primitive data is then stored in MinIO in Potree format to build a primitive database for the 3D map. Edit the scene metadata through the DDT front-end, trigger the DDT back-end to generate new metadata, and update the metadata database in MinIO; The metadata of the graphics elements are grouped and managed through the DDT backend to form data blocks of different resolutions; each data block includes point cloud data of one or more graphics elements, location information, height information, color information related to the point cloud data, and metadata information used to identify the resolution level of the data block and its location in the 3D map; When viewing the digital twin map of the robot's environment through the DDT front-end, the corresponding data blocks are organized and loaded from the graphic metadata database through the octree structure, and the LOD technology is used to dynamically switch data blocks of different resolutions according to the position and viewing angle of the virtual camera. The loaded data blocks are then rendered in three dimensions to obtain the current digital twin map.
2. According to claim 1, a digital twin method based on radar-visual fusion and primitive segmentation is characterized in that: The SLAM data set includes: an original file and a description file; The original files include point cloud data, image data, IMU data, and GPS data; The description file includes external parameters from the laser radar to the optical imaging sensor and external parameters from the laser radar to the inertial measurement unit.
3. The digital twin method based on radar-visual fusion and primitive segmentation according to claim 2 is characterized in that: The data preprocessing includes: cleaning, denoising, de-duplication and de-distortion processing of the SLAM data set.
4. The digital twin method based on radar-visual fusion and primitive segmentation according to claim 3 is characterized in that: The laser vision fusion includes: uniformly converting the three-dimensional spatial coordinates of point cloud data and image data through the displacement relationship and rotation angle relationship between the laser radar and the optical imaging sensor, and converting and mapping the point cloud data that are not uniform in time and space through the acceleration a and angular velocity ω in the IMU data to obtain a three-dimensional map of the surrounding environment.
5. The digital twin method based on radar-visual fusion and primitive segmentation according to claim 4 is characterized in that: The primitive segmentation includes: segmenting the three-dimensional map into 50m×50m primitives to obtain primitive data of the three-dimensional map; and each primitive is composed of a number of point clouds, and each point cloud includes relevant position information, height information, and color information.
6. The digital twin method based on radar-visual fusion and primitive segmentation according to claim 5 is characterized in that: The generating of the 2D base map comprises: generating the 2D base map based on the bird's-eye view of the three-dimensional map after the segmentation of the primitives, and obtaining the primitive data of the 2D base map.
7. A digital twin method based on radar-visual fusion and primitive segmentation according to any one of claims 1 to 6, characterized in that: The scene editing includes: copying, deleting and moving point cloud data of the same three-dimensional map, and performing map splicing, map expansion, map alignment, map semantic definition and map navigation definition on different three-dimensional maps.
8. A digital twin method based on radar-visual fusion and primitive segmentation according to any one of claims 1 to 6, characterized in that: According to the position of the virtual camera and the distance of the viewing angle, LOD technology is used to dynamically switch data blocks of different resolutions, including: when the viewing angle is far away, low-resolution data blocks are loaded according to LOD technology; when the viewing angle is close, high-resolution data blocks are loaded according to LOD technology.
9. A digital twin device based on radar-visual fusion and primitive segmentation, characterized in that: include: A memory and at least one processor, wherein the memory and the at least one processor are interconnected via a line; Wherein, a computer program is stored in the memory; and the at least one processor calls the computer program in the memory to execute the steps of a digital twin method based on radar-visual fusion and primitive segmentation as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a digital twin method based on radar-vision fusion and primitive segmentation as described in any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Multi-machine cloud edge collaborative map creation and dynamic digital twinning method and system
CN116030213A
Ground-based synthetic aperture radar rapid three-dimensional terrain rendering method and device
CN116824017A
Global level vector tile data compression construction and dynamic LOD loading method
CN118484438A
Digital twin point cloud editing asynchronous updating method
CN119579840A
Rendering scenes using a combination of raytracing and rasterization
US20200372703A1