Multi-sensor fusion three-dimensional reconstruction method applied to unmanned forklift

Through multi-sensor fusion technology and 3DGS Gaussian splashing technology, the problems of poor three-dimensional reconstruction accuracy and inability to intuitively present RGB features in the existing technology are solved, and high-precision, real-time capture and reconstruction of objects are achieved, meeting the accuracy and intuitive needs of smart factories.

CN120070753APending Publication Date: 2025-05-30MULTIWAY ROBOTICS TECH (SHENZHEN) CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510137768.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has problems in three-dimensional reconstruction with high manual modeling costs, poor modeling accuracy or inability to visually present RGB features.

Method used

Using multi-sensor fusion method, data is collected through 3D lidar and global exposure plane array cameras, combined with 3D GS Gaussian splattering technology, point clouds and images are fusion and rasterized, and a three-dimensional model with RGB characteristics is output.

Benefits of technology

Real-time and high-precision capture and reconstruction of object shapes, textures, positions and colors is achieved, ensuring a high consistency between digital twin models and physical entities, meeting the strict requirements of smart factories for model accuracy, and providing an intuitive visual experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070753A_ABST
    Figure CN120070753A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of unmanned forklift sensors, in particular to a multi-sensor fusion three-dimensional reconstruction method applied to an unmanned forklift, which comprises the following steps: 1) radar hardware layout: arranging a 3D laser radar equipped with two global exposure area-array cameras at a head position, the radar and the camera are converted to a vehicle body coordinate system through calibration; 2) multi-frame point cloud and preprocessing: (1) continuously collecting multi-frame point cloud data by using a radar, and distributing a timestamp for each frame of point cloud data; (2) performing motion distortion removal processing on the multi-frame point cloud data by using odometer information, namely converting all the point cloud data at different moments into a vehicle body coordinate system at the current moment so as to eliminate point cloud distortion generated by vehicle motion; according to the invention, real-time and high-precision capture and reconstruction of the shape, texture, position and color of an object are realized by fusing a real-time data acquisition technology of a camera and a laser radar and applying a 3DGS Gaussian splashing technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned forklift sensors, and specifically to a multi-sensor fusion three-dimensional reconstruction method applied to unmanned forklifts. Background Art

[0002] With the rapid development of the logistics industry, the application of automated guided vehicles (AGVs) in the industry has become increasingly widespread, becoming an indispensable and important part of smart factories. Smart factories, through highly automated, informatized, and intelligent production methods, have greatly improved production efficiency, reduced costs, and enhanced product quality. Digital twin technology, as one of the core technologies of smart factories, has played a crucial role.

[0003] Digital twin technology realizes the comprehensive, real-time, and accurate simulation and monitoring of physical entities by creating digital models corresponding to physical entities in a virtual space. This technology is widely applied in multiple links of smart factory production line design, production monitoring, energy efficiency management, quality control, etc. Through means such as simulation optimization, real-time monitoring, and predictive maintenance, it effectively improves the production efficiency and stability of the factory. At the same time, digital twin technology also facilitates remote maintenance, training, and data service-ization, providing strong support for the intelligent transformation of smart factories.

[0004] However, the key technology of digital twin is how to perform three-dimensional reconstruction of objects. There are mainly the following four ways of three-dimensional reconstruction: BIM modeling, drawing modeling, oblique photography modeling, and laser point cloud modeling. The first two methods mainly rely on manual model establishment. For unknown models, it is necessary to manually measure their three-dimensional features and establish a 3D model. This method has high maintenance costs and poor flexibility. Oblique photography modeling mainly obtains a large amount of image data through multi-angle, multi-direction, and multi-viewpoint shooting. These image data are processed through image processing, registration, stitching and other technologies, and finally a three-dimensional model is generated. Although this method has low costs, its modeling accuracy is poor, and for some objects without surface textures, an accurate model cannot be established. Laser point cloud modeling is a process of obtaining the spatial three-dimensional data of real-world targets through devices such as three-dimensional laser scanning or lidar, and then using modeling software to establish a high-precision three-dimensional model with the point cloud data as a reference. Although the model established by this method can accurately establish the three-dimensional model of an object, it has no RGB features and cannot intuitively present the real model of the object. Summary of the Invention

[0005] (1) Objectives of the Invention

[0006] In view of this, the objective of the present invention is to propose a multi-sensor fusion three-dimensional reconstruction method applied to unmanned forklifts, which solves the problems raised in the above background.

[0007] (2) Technical Solution

[0008] A multi-sensor fusion three-dimensional reconstruction method applied to an unmanned forklift, the method comprising the following steps:

[0009] 1) Radar hardware layout: Install a 3D lidar equipped with two global exposure area array cameras at the front of the vehicle head, and convert the radar and the camera to the vehicle body coordinate system through calibration;

[0010] 2) Multi-frame point cloud and preprocessing: (1) Continuously collect multi-frame point cloud data using the radar, and assign a timestamp to each frame of point cloud data; (2) Use odometer information to perform motion de-distortion processing on the multi-frame point cloud data, that is, convert the point cloud data at all different times to the vehicle body coordinate system at the current time to eliminate the point cloud distortion caused by vehicle movement;

[0011] 3) Point cloud and image fusion: According to the timestamp, obtain the image with the closest time and the radar data after motion de-distortion processing respectively, and convert the image and the radar data to the vehicle body coordinate system through the RT matrix from the radar to the vehicle body and from the camera to the vehicle body;

[0012] 4) Three-dimensional reconstruction: Use the 3DGS Gaussian splashing technology to convert the fused point cloud into a Gaussian distribution, and perform rasterization processing, outputting the positions (xyz), covariance matrices, colors RGB, and transparencies Alpha of all points to complete the three-dimensional reconstruction.

[0013] Preferably, the calibration process in step 1) includes:

[0014] Make a calibration board and place it in front of the radar and the camera, keeping the calibration board perpendicular to the vehicle body and the ground;

[0015] Obtain the corner point information of the calibration board by acquiring the rgb image through the camera, and obtain the 3D coordinates of the corner points by extracting the intensity information of the point cloud through the radar;

[0016] Calculate the RT transformation matrix from the radar to the camera through singular value decomposition (SVD), and then determine the transformation matrices R1T and R2T2 from the radar and the camera to the vehicle body according to the position relationship between the calibration board and the vehicle body coordinate system obtained by manual measurement.

[0017] Preferably, the motion de-distortion processing in step 2) further includes:

[0018] Assume that the pose change of the head and tail frame radars is T, and obtain the time difference ΔT between the head and tail point clouds during this period through the odometer;

[0019] Through a vehicle uniform motion hypothesis model or other suitable motion models, based on the timestamp and odometer information, all point clouds at different times are transformed into the vehicle body coordinate system at the current time, that is, the motion distortion of the point cloud is completed.

[0020] Preferably, the fusion of the point cloud and the image in step 3) further includes:

[0021] The left camera and the right camera select the image closest to the current time according to the timestamp and fuse it with the radar data after motion distortion processing.

[0022] Preferably, the 3D reconstruction in step 4) further includes:

[0023] By training a 3D Gaussian splash model, each point in the point cloud is initialized as a 3D Gaussian, and its initial spherical harmonic coefficients, rotation matrix, scaling coefficient, and transparency are calculated;

[0024] Using backpropagation and gradient descent techniques to update the model parameters until an optimized 3D Gaussian splash model is obtained;

[0025] Deploy the trained model to the device side, and output the position (xyz), covariance matrix, color RGB, and transparency Alpha of all points by inputting the point cloud after motion distortion processing and RGB information.

[0026] Preferably, the training process of the 3D Gaussian splash model further includes:

[0027] Observe the 3D Gaussian in the scene from the position and angle of the camera to obtain a pseudo RGB image;

[0028] Use this pseudo RGB image to compare with the real RGB image and calculate the L1 loss to optimize the model parameters.

[0029] From the above technical solutions, it can be seen that this application has the following beneficial effects:

[0030] 1. The invention realizes the real-time, high-precision capture and reconstruction of the shape, texture, position, and color of objects by integrating the real-time data acquisition technology of cameras and lidar, and applying 3DGS Gaussian splash technology. This innovative solution not only ensures a high degree of consistency between the digital twin model and the physical entity in the smart factory, provides strong support for real-time monitoring and predictive maintenance, but also meets the strict requirements of the smart factory for the accuracy of the digital twin model. At the same time, the generated 3D model with RGB features intuitively presents the color and texture of the object, providing a more intuitive and vivid visual experience for remote maintenance and training. Brief Description of the Drawings

[0031] Figure 1Schematic diagram of the specific process of 3D reconstruction of the present invention;

[0032] Figure 2 Schematic diagram of the hardware layout of the present invention;

[0033] Figure 3 Schematic diagram of the calibration board of the present invention. Detailed implementation manners

[0034] The following description is merely exemplary in nature and is not intended to limit the present disclosure, its application, and uses. It should be understood that in all these drawings, the same or similar reference numerals indicate the same or similar parts and features. Each drawing only schematically shows the concept and principle of the embodiments of the present disclosure, and does not necessarily show the specific dimensions and their ratios of the embodiments of the present disclosure. In a specific part of a specific drawing, the relevant details or structures of the embodiments of the present disclosure may be illustrated in an exaggerated manner.

[0035] Please refer to Figures 1-3 , an embodiment provided by the present invention:

[0036] A multi-sensor fusion 3D reconstruction method applied to an automated forklift, the method comprising the following steps:

[0037] 1) Radar hardware layout: Install a 3D lidar equipped with two global exposure area cameras at the front of the vehicle head, and convert the lidar and the cameras to the vehicle body coordinate system through calibration;

[0038] 2) Multi-frame point cloud and preprocessing: (1) Continuously collect multi-frame point cloud data using the lidar, and assign a timestamp to each frame of point cloud data; (2) Use the odometer information to perform motion de-distortion processing on the multi-frame point cloud data, that is, convert the point cloud data at all different times to the vehicle body coordinate system at the current time to eliminate the point cloud distortion caused by vehicle movement;

[0039] 3) Point cloud and image fusion: According to the timestamp, respectively obtain the image with the closest time and the lidar data after motion de-distortion processing, and convert the image and the lidar data to the vehicle body coordinate system through the RT matrix from the lidar to the vehicle body and from the camera to the vehicle body;

[0040] 4) 3D reconstruction: Use the 3DGS Gaussian splashing technology to convert the fused point cloud into a Gaussian distribution, and perform rasterization processing, outputting the positions (xyz), covariance matrices, colors RGB, and transparencies Alpha of all points to complete the 3D reconstruction.

[0041] Further, the calibration process in step 1) includes:

[0042] Make a calibration board and place it in front of the lidar and the cameras, keeping the calibration board perpendicular to the vehicle body and the ground;

[0043] Obtain the RGB image through the camera to extract the corner point information of the calibration board, and obtain the point cloud data through the radar. Extract the 3D coordinates of the corner points according to the intensity information of the point cloud;

[0044] Calculate the RT transformation matrix from the radar to the camera through singular value decomposition (SVD). Then, based on the positional relationship between the calibration board and the vehicle body coordinate system obtained by manual measurement, finally determine the transformation matrices R1T and R2T2 from the radar and the camera to the vehicle body. This process not only considers the relative positional relationship between the devices but also their integration with the entire vehicle system, enabling the calibration results to be directly applied to actual vehicle navigation, environmental perception, and other scenarios, enhancing the practicality and adaptability of the system.

[0045] Furthermore, the motion de-distortion process in step 2 also includes:

[0046] Assume that the pose change of the radar between the first and last frames is T, and obtain the time difference ΔT between the first and last point clouds during this time period through the odometer;

[0047] Based on the vehicle uniform motion assumption model or other suitable motion models, and according to the time stamps and odometer information, transform all point clouds at different times to the vehicle body coordinate system at the current time, that is, complete the motion de-distortion of the point cloud.

[0048] Furthermore, the point cloud and image fusion in step 3) also includes:

[0049] The left camera and the right camera select the image closest to the current time according to the time stamp and fuse it with the radar data after motion de-distortion processing.

[0050] Furthermore, the 3D reconstruction in step 4) also includes:

[0051] By training a 3D Gaussian splash model, initialize each point in the point cloud as a 3D Gaussian, and calculate its initial spherical harmonic coefficients, rotation matrix, scaling coefficient, and transparency;

[0052] Use backpropagation and gradient descent techniques to update the model parameters until an optimized 3D Gaussian splash model is obtained;

[0053] Deploy the trained model to the device side. By inputting the point cloud and RGB information after motion de-distortion processing, output the positions (xyz), covariance matrices, colors RGB, and transparencies Alpha of all points, and then the 3D reconstruction can be completed.

[0054] Furthermore, the training process of the 3D Gaussian splash model also includes:

[0055] Observe the 3D Gaussian in the scene from the position and angle of the camera to obtain a pseudo RGB image;

[0056] The pseudo RGB image is compared with the real RGB image to calculate the L1 loss for optimizing the model parameters.

[0057] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced by the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

Claims

1. A multi-sensor fusion 3D reconstruction method applied to unmanned forklifts, characterized by: The method comprises the following steps: 1) Radar hardware layout: A 3D laser radar equipped with two global exposure array cameras is installed at the front of the vehicle, and the radar and camera are converted to the vehicle body coordinate system through calibration; 2) Multi-frame point cloud and preprocessing: (1) Use radar to continuously collect multi-frame point cloud data and assign a timestamp to each frame of point cloud data; (2) Use odometer information to perform motion de-distortion processing on multi-frame point cloud data, that is, convert all point cloud data at different times to the vehicle coordinate system at the current time to eliminate point cloud distortion caused by vehicle motion; 3) Fusion of point cloud and image: According to the timestamp, the image closest in time and the radar data after motion dedistortion are obtained respectively, and the image and radar data are converted to the vehicle coordinate system through the RT matrix from radar to vehicle body and from camera to vehicle body; 4) 3D reconstruction: Use 3DGS Gaussian splashing technology to convert the fused point cloud into a Gaussian distribution, perform rasterization, and output the position (xyz), covariance matrix, color RGB, and transparency Alpha of all points to complete the 3D reconstruction.

2. A multi-sensor fusion 3D reconstruction method for unmanned forklift according to claim 1, characterized in that: The calibration process in step 1) includes: Make a calibration plate and place it in front of the radar and camera, keeping it perpendicular to the vehicle body and the ground; The camera obtains RGB images to extract the corner point information of the calibration plate, and the radar obtains point cloud data, and extracts the 3D coordinates of the corner points based on the intensity information of the point cloud; The RT transformation matrix from radar to camera is calculated through singular value decomposition (SVD), and then the transformation matrices R1T and R2T2 from radar and camera to vehicle body are finally determined based on the position relationship between the calibration plate and the vehicle body coordinate system obtained by manual measurement.

3. The multi-sensor fusion 3D reconstruction method for unmanned forklift according to claim 1 is characterized in that: The motion de-distortion processing in step 2 further includes: Assuming that the position change of the first and last frame radar is T, the time difference ΔT between the first and last point clouds in this time period is obtained by the odometer; Through the vehicle uniform speed assumption model or other appropriate motion model, according to the timestamp and odometer information, all point clouds at different times are transformed to the vehicle body coordinate system at the current time, that is, point cloud motion dedistortion is completed.

4. The multi-sensor fusion 3D reconstruction method for unmanned forklift according to claim 1 is characterized in that: The point cloud and image fusion in step 3) further includes: The left camera and the right camera select the image closest to the current moment according to the timestamp and fuse it with the radar data after motion dedistortion processing.

5. The multi-sensor fusion 3D reconstruction method for unmanned forklift according to claim 1 is characterized in that: The three-dimensional reconstruction in step 4) further includes: By training a 3D Gaussian splash model, each point in the point cloud is initialized as a 3D Gaussian, and its initial spherical harmonic coefficients, rotation matrix, scaling factor, and transparency are calculated; Use back propagation and gradient descent techniques to update model parameters until an optimized 3D Gaussian splash model is obtained; The trained model is deployed to the device, and the position (xyz), covariance matrix, color RGB, and transparency Alpha of all points are output by inputting the point cloud and RGB information after motion dedistortion.

6. The multi-sensor fusion 3D reconstruction method for unmanned forklift according to claim 5 is characterized in that: The training process of the 3D Gaussian splash model also includes: Observe the 3D Gaussian in the scene from the camera's position and angle to get a pseudo RGB image; The pseudo RGB image is compared with the real RGB image to calculate the L1 loss to optimize the model parameters.

Citation Information

Patent Citations

  • Semantic live-action three-dimensional reconstruction method and system of laser fusion multi-view camera

    CN113362247A

  • Three-dimensional reconstruction method for vehicle based on multi-sensor fusion

    CN113421325A

  • Point cloud and image fusion labeling method and system suitable for different batches of data

    CN116051656A

  • Forest region positioning and three-dimensional reconstruction method and system based on multi-sensor fusion

    CN116228969A

  • Real-time multi-mode sensing high-precision map construction method

    CN116817891A