Three-dimensional model reconstruction method based on multi-source input and apparatus
By integrating point cloud information and image information from multiple sources, and adjusting the geometric and optical residuals of the 3D model, the problems of incomplete 3D reconstruction and low realism in existing technologies are solved, and efficient and accurate 3D model reconstruction is achieved.
Patent Information
- Application Number
- PCT/CN2025/095518
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-09
- Filing Date
- 2025-05-16
- Publication Date
- 2025-11-27
AI Technical Summary
Existing 3D reconstruction methods are limited by single-view or multi-view image input, resulting in incomplete reconstruction models, low realism, insufficient precision and accuracy, and cumbersome data acquisition processes with low efficiency, making it difficult to meet the needs of high speed, batch processing, and high frequency.
By acquiring the point cloud information and image information of the object to be reconstructed, and combining the geometric residuals and optical residuals, the 3D model is adjusted to improve the modeling accuracy and realism.
It achieves high-precision and realistic 3D model reconstruction, improves modeling efficiency and completeness, supports rapid mapping and real-time feedback, and adapts to high-speed, batch, and high-frequency requirements.
Smart Images

Figure CN2025095518_27112025_PF_FP_ABST
Abstract
Description
A three-dimensional model reconstruction method and device based on multi-source input
[0001] The present disclosure claims priority to Chinese Patent Application No. 202410626989.0, entitled "A Three-Dimensional Reconstruction Method and Device Based on Cloud Computing Technology", filed on May 20, 2024, and Chinese Patent Application No. 202411096682.0, entitled "A Three-Dimensional Model Reconstruction Method and Device Based on Multi-Source Input", filed on August 9, 2024, the contents of which are incorporated herein by reference in their entirety. TECHNICAL FIELD
[0002] The present application relates to the field of three-dimensional model reconstruction, and in particular to a three-dimensional model reconstruction method and device based on multi-source input. BACKGROUND
[0003] Three-dimensional reconstruction (3D Reconstruction) is a mathematical process and computer technology for restoring the original three-dimensional information of a scene. Early three-dimensional model reconstruction methods usually take single-view or multi-view image information as input, and are limited by the input data. The reconstructed three-dimensional model is usually incomplete and has low realism. High-precision three-dimensional reconstruction schemes based on pure visual input have a cumbersome data collection process and low work efficiency. Due to the need for high dependence on a Structure from Motion (SFM) module, the applicable scenarios are limited, the degree of automation is low, it is difficult to meet the needs of high speed, batch, and high frequency, and the precision and accuracy are low with serious distortion.
[0004] In recent years, with the iteration of sensor hardware technology, reconstruction technologies that fuse three-dimensional sensor information such as Light Detection and Ranging Sensor (LiDAR) and Inertial Measurement Unit (IMU) have developed rapidly, aiming to obtain more dense and higher-precision three-dimensional models. SUMMARY
[0005] The present application provides a three-dimensional model reconstruction method and device based on multi-source input. By obtaining mapping results including point cloud information and image information of an object to be reconstructed, a three-dimensional model is established. Further, geometric residuals and optical residuals are determined based on the point cloud information and the image information, and the three-dimensional model is further adjusted to improve the modeling efficiency, the modeling precision, and the modeling realism.
[0006] In a first aspect, the application provides a three-dimensional model reconstruction method based on multi-source input. The method is applied to a three-dimensional model reconstruction platform. Specifically, the method includes the following steps: obtaining a mapping result of an object to be reconstructed, wherein the mapping result includes image information of the object to be reconstructed and point cloud information of the object to be reconstructed. Further, a three-dimensional model of the object to be reconstructed is established according to the mapping result. On this basis, target image information and target point cloud information of the object to be reconstructed under a target pose are determined from the mapping result. The target pose includes position information and attitude information, and then the three-dimensional model is instructed to generate a target depth image and a target digital image according to the target pose. Thus, the three-dimensional model reconstruction platform determines a geometric residual according to the target point cloud information and the target depth image, determines an optical residual according to the target image information and the target digital image, and adjusts the three-dimensional model according to the geometric residual and the optical residual.
[0007] In the scheme provided in the application, the three-dimensional model is established by obtaining the mapping result of the multi-source input including the point cloud information and the image information of the object to be reconstructed. The established three-dimensional model can output corresponding depth images and digital images based on a specific pose. Thus, the optical residual and the geometric residual are constructed by combining the target point cloud information and the target image information in the mapping result, and the three-dimensional model is adjusted according to the optical residual and the geometric residual, which jointly supervise the training of the three-dimensional model. Thus, the three-dimensional model is adjusted by simultaneously using the geometric supervision and the optical supervision to simultaneously possess the real optical information and the real geometric information, so that the geometric precision and the richness of the detail features of the three-dimensional model are obviously improved, and the rendering product has high reality.
[0008] In combination with the first aspect, in a possible implementation manner of the first aspect, the specific steps of obtaining the mapping result of the object to be reconstructed are as follows: obtaining initial image information and initial point cloud information of the object to be reconstructed, and collecting image inertia information corresponding to the collection of the initial image information and point cloud inertia information corresponding to the collection of the initial point cloud information, and then generating the mapping result according to the image information, the point cloud information, the point cloud inertia information and the image inertia information, so as to obtain the mapping result of the modeling object.
[0009] In the scheme provided in the application, the information of the multi-source input such as the image information, the point cloud information and the corresponding inertia information is obtained, the frame-by-frame point cloud, the frame-by-frame image and the corresponding motion prior information are fused, the object to be reconstructed is mapped, and high-precision mapping is realized. By introducing the corresponding inertia information and the point cloud information, pose fitting can be faster, and mapping efficiency can be improved.
[0010] With reference to the first aspect, in a possible implementation manner of the first aspect, the three-dimensional model reconstruction platform determines description sub-information according to the initial image information, where the description sub-information includes texture features of at least one frame of image in the initial image information, and then determines the key frame image according to the initial image information and the description sub-information.
[0011] In the scheme provided in the present application, the determination of the key frame image helps to screen out images that meet the requirements of specific description sub-information, for example, images with relatively rich texture features or detail richness. When adjusting the three-dimensional model, by selecting the key frame image to construct the optical residual, the adjustment efficiency of the three-dimensional model and the rendering effect of the three-dimensional model on the detail features can be further improved.
[0012] With reference to the first aspect, in a possible implementation manner of the first aspect, the three-dimensional model reconstruction platform can determine a target pose for observing the key frame image according to the key frame image.
[0013] In the scheme provided in the present application, by determining the target pose according to the key frame image, the digital image output by the three-dimensional model can be used to construct the optical residual with the key frame image with relatively rich texture features or detail richness, and at the same time, the depth image output by the three-dimensional model and the corresponding point cloud information with high accuracy can be used to construct the geometric residual. Therefore, the corresponding pose obtained from the key frame image as the basis for residual construction can further improve the efficiency and effect of the three-dimensional model training iteration.
[0014] With reference to the first aspect, in a possible implementation manner of the first aspect, the three-dimensional model reconstruction platform can further determine image pose information according to the image information and image inertia information, and determine point cloud pose information according to the point cloud information and point cloud inertia information, and then generate a point cloud map according to the image information, the point cloud information, the point cloud inertia information and the image inertia information, specifically including the following steps: generating the point cloud map according to the image information, the point cloud information, the point cloud pose information and the image pose information. Further, the three-dimensional model reconstruction platform can determine global pose information according to the point cloud pose information and the image pose information.
[0015] In the scheme provided in the present application, by fusing the image information and the image inertia information, the point cloud information and the point cloud inertia information, the corresponding image pose and the point cloud pose are determined, so that when mapping, the image information and the point cloud information can determine the adjacent relationship and the matching relationship through the corresponding image pose and the point cloud pose, and accurate mapping is realized. By establishing the global pose information, the image pose and the point cloud pose can be quickly matched and converted, and the three-dimensional model can output complete and accurate rendering images according to a certain pose in the point cloud pose or the image pose.
[0016] With reference to the first aspect, in a possible implementation manner of the first aspect, the three-dimensional model reconstruction platform performs loop detection on the mapping result to obtain matching information, the matching information is used to indicate a matching condition of a current frame and a historical frame, and the mapping result is adjusted according to the matching information.
[0017] In the scheme provided in the application, since there is cumulative measurement error in the acquisition device during long-time operation, especially when the acquisition device returns to a region for which mapping has been completed after long-time work, misalignment will occur between current sensor information and the already completed mapping result. Therefore, loop detection needs to be performed during mapping to search whether the current frame can form a loop with historical information, and a pose graph is constructed to adjust the mapping result and improve mapping accuracy.
[0018] With reference to the first aspect, in a possible implementation manner of the first aspect, the three-dimensional model reconstruction platform updates the mapping result according to the supplementary image information, the supplementary point cloud information, the supplementary image inertia information and the supplementary laser inertia information.
[0019] In the scheme provided in the application, by using multi-source input for mapping of the object to be reconstructed, the mapping speed is improved, and current mapping progress can be fed back to the user in real time, so that the user can know whether there is local area missing, detail missing and the like in the object to be reconstructed in a timely manner, and the corresponding area is supplemented for scanning in a targeted manner, so as to complete high-quality original data construction for mapping.
[0020] With reference to the first aspect, in a possible implementation manner of the first aspect, the user can configure adjustment of the three-dimensional model, and the three-dimensional model reconstruction platform obtains the iteration number input by the user, and indicates that the three-dimensional model is adjusted according to the iteration number, wherein the iteration number is used to indicate the number of adjustments of the three-dimensional model, and the three-dimensional model is output after the three-dimensional model meets the iteration number.
[0021] In the scheme provided in the application, the user controls the iteration training process of the three-dimensional model in a configured manner, so that the three-dimensional model is output after meeting the corresponding adjustment requirement, to be used for subsequent service.
[0022] The second aspect or any one of the implementation manners of the second aspect is an apparatus implementation corresponding to the first aspect or any one of the implementation manners of the first aspect, and the description in the first aspect or any one of the implementation manners of the first aspect is applicable to the second aspect or any one of the implementation manners of the second aspect, and will not be repeated here.
[0023] In a third aspect, the present application provides a computing device cluster, comprising at least one computing device, each computing device comprising a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method of the first aspect above and any one of the implementation manners of the first aspect above.
[0024] In a fourth aspect, the present application provides a computer program product comprising instructions, which, when executed by a computing device cluster, cause the computing device cluster to perform the method of the first aspect above and any one of the implementation manners of the first aspect above.
[0025] In a fifth aspect, the present application provides a computer-readable storage medium comprising computer program instructions, which, when executed by a computing device cluster, cause the computing device cluster to perform the method of the first aspect above and any one of the implementation manners of the first aspect above. BRIEF DESCRIPTION OF DRAWINGS
[0026] FIG. 1 is a schematic diagram of an application scenario of the method for reconstructing a three-dimensional model based on multi-source input according to an embodiment of the present application;
[0027] FIG. 2 is a schematic diagram of another application scenario of the method for reconstructing a three-dimensional model based on multi-source input according to an embodiment of the present application;
[0028] FIG. 3 is a schematic diagram of a flow of the method for reconstructing a three-dimensional model based on multi-source input according to an embodiment of the present application;
[0029] FIG. 4 is a schematic diagram of a mapping flow of the method for reconstructing a three-dimensional model based on multi-source input according to an embodiment of the present application;
[0030] FIG. 5 is a schematic diagram of another mapping flow of the method for reconstructing a three-dimensional model based on multi-source input according to an embodiment of the present application;
[0031] FIG. 6 is a schematic diagram of a key frame determination flow of the method for reconstructing a three-dimensional model based on multi-source input according to an embodiment of the present application;
[0032] FIG. 7 is a schematic diagram of a three-dimensional model reconstruction flow of the method for reconstructing a three-dimensional model based on multi-source input according to an embodiment of the present application;
[0033] FIG. 8 is a schematic diagram of a structure of an apparatus according to an embodiment of the present application;
[0034] FIG. 9 is a schematic diagram of a structure of a computing device for reconstructing a three-dimensional model based on multi-source input according to an embodiment of the present application;
[0035] FIG. 10 is a schematic diagram of a structure of a computing device cluster for reconstructing a three-dimensional model based on multi-source input according to an embodiment of the present application;
[0036] FIG. 11 is another structural schematic diagram of a computing device cluster for the method of reconstructing a three-dimensional model based on multi-source input according to an embodiment of the present application;
[0037] FIG. 12 is another structural schematic diagram of a computing device cluster for the method of reconstructing a three-dimensional model based on multi-source input according to an embodiment of the present application. DETAILED DESCRIPTION
[0038] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0039] Reference to “an embodiment” or “some embodiments” in this text means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The appearance of the phrase in various places in the specification does not necessarily all refer to the same embodiment, nor does it necessarily refer to a separate or alternative embodiment in which the other embodiments are mutually exclusive. A person of ordinary skill in the art explicitly and implicitly understands that the embodiments described herein can be combined with other embodiments.
[0040] Reference to “an embodiment” or “some embodiments” in this text means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The appearance of the phrase in various places in the specification does not necessarily all refer to the same embodiment, nor does it necessarily refer to a separate or alternative embodiment in which the other embodiments are mutually exclusive. A person of ordinary skill in the art explicitly and implicitly understands that the embodiments described herein can be combined with other embodiments.
[0041] First, the terms appearing in this text are explained as follows:
[0042] Three-dimensional model reconstruction (3D Model Reconstruction): refers to a process of converting an object or scene in the real world into a computer-processable three-dimensional digital model by using computer vision and graphics processing technology.
[0043] Object to be reconstructed: refers to the object or object to be reconstructed by computer vision and graphics processing technology, which can be a three-dimensional space, a scene, an object-level model, a digital human, etc. Depth image (Depth Image): Depth image is also called distance image, which refers to the image taking the distance (depth) from the image collector to each point in the scene as the pixel value, which directly reflects the geometric shape of the visible surface of the scene. Depth image can be calculated as point cloud data after coordinate conversion, and point cloud data with regular and necessary information can also be calculated as depth image data.
[0044] Digital image (Digital image): is the representation of an image with finite digital pixel values, represented by an array or matrix, whose illumination position and intensity are discrete. Digital image is an image that can be stored and processed by digital computer or digital circuit, which is obtained by digitizing analog image and taking pixel as the basic element. In an embodiment of the present application, the digital image can be a binary image, a color image, etc., such as RGB image, full-color image, etc.
[0045] Point cloud (Point Clouds): is a set of points in space, which can represent three-dimensional shapes or objects, usually obtained by three-dimensional scanners. For example, the position of each point in the point cloud is described by a set of Cartesian coordinates (X, Y, Z) {\textstyle(X,Y,Z)}, and some may contain color information (R, G, B) or object reflectance intensity information. In an embodiment of the present application, point cloud information can be point cloud directly obtained by laser scanner (LiDAR) or other sensors; it can also be point cloud directly obtained by other sensors such as sound wave radar.
[0046] Three-dimensional model: refers to the mathematical representation of a real or imaginary object in three-dimensional space. This representation is usually composed of a series of three-dimensional coordinate points, which are connected into lines and surfaces to form a complete geometric structure, which can accurately reproduce the shape, structure and appearance of the object, and realize the reconstruction of the scene. For example, 3D Gaussian splatting (3D Gaussian Splatting, 3DGS): a scene reconstruction and rendering technology based on 3D Gaussian body representation, which combines the advantages of explicit representation and implicit representation, can realize scene reconstruction and efficient real-time rendering based on pure image input, and generate synthetic data under new view angle, which is a new reconstruction paradigm representing the development direction of current computer vision field; for example, neural radiance field (Neural Radiance Field, NeRF): an implicit three-dimensional model reconstruction method based on deep neural network, which learns the radiation and color information of each point in the scene, so as to synthesize realistic images at any view angle. It samples points in three-dimensional space and predicts radiation and color for each point to construct the implicit representation of the scene.
[0047] Rasterization: A core step of 3D Gaussian Splatting rendering, is a mathematical process that converts the mathematical description of an object and the color information associated with the object into pixels on the screen for the corresponding position and the color used to fill the pixels. In 3D Gaussian, it refers to the mathematical process of rendering a 2D image from the distribution of Gaussian kernels and color information.
[0048] Inertial Measurement Unit (IMU): A device that measures the three-axis attitude angle (or angular rate) and acceleration of an object.
[0049] Pose: Refers to the position and orientation of an object relative to a reference coordinate system. Specifically, pose includes spatial position information and rotation direction information of the object. Pose can be used to describe the position and orientation of a rigid body object in any three-dimensional space. Among them, the position usually refers to the coordinates of the object center or a specific reference point in three-dimensional space; the attitude usually refers to the rotation angle or rotation matrix of the object, indicating the rotation of the object relative to the reference coordinate system.
[0050] Explicit representation: A traditional representation form of three-dimensional models, which explicitly models scenes or objects, allowing users to edit and view, including meshes, point clouds, voxels, etc.
[0051] Implicit representation: A representation form that describes the three-dimensional information of a scene in a parameterized manner based on machine learning methods such as deep neural networks, and constructs a mapping relationship from three-dimensional space coordinates to corresponding geometric / texture information.
[0052] Simultaneous Localization and Mapping (SLAM): An explicit three-dimensional model reconstruction method that runs in real time, hoping that the robot starts from an unknown location in an unknown environment, and locates its position and attitude through repeatedly observed map features (such as corners, columns, etc.) during movement, and then constructs a map incrementally according to its position, so as to achieve the purpose of simultaneous localization and mapping.
[0053] Structure from Motion (SFM): A non-real-time explicit three-dimensional model reconstruction method, given a series of images with a certain degree of overlap, to simultaneously estimate the position and attitude of the camera when each image is taken, and the sparse point cloud of the object or scene being photographed.
[0054] In the field of three-dimensional reconstruction, it is a common method to reconstruct the three-dimensional model of the object to be reconstructed based on single-view or multi-view image information as input data. However, due to the limitation of input data, the point cloud information is usually estimated by view image, which is time-consuming, low in efficiency, poor in accuracy, and the reconstructed three-dimensional model is not complete, the depth information is distorted seriously, and lacks of real sense. In the process of optimizing the three-dimensional model, the input of accurate geometric checking information is lacking, it is difficult to optimize and adjust the geometric depth information, and at the same time, the three-dimensional model reconstruction efficiency is low.
[0055] Based on this, the application provides a three-dimensional model reconstruction method based on multi-source input, by acquiring the mapping result including the point cloud information and image information of the object to be modeled, establishing a three-dimensional model, further determining the geometric residual and optical residual by the point cloud information and image information, and further adjusting the three-dimensional model, improving the modeling efficiency, accuracy and reality.
[0056] Please refer to FIG. 1, as shown in FIG. 1, FIG. 1 is a kind of application scene schematic diagram of three-dimensional model reconstruction method based on multi-source input provided in the embodiment of the application, specifically:
[0057] As an embodiment of the application, the three-dimensional model reconstruction platform 12 can run on the infrastructure 15, wherein the infrastructure 15 includes at least one computing device for computing in the three-dimensional model reconstruction process. And, user A can directly access the three-dimensional model reconstruction platform 12, control the three-dimensional model reconstruction process through the three-dimensional model reconstruction platform 12, and input corresponding operation instruction in the process. The three-dimensional model reconstruction platform 12 reconstructs the three-dimensional model 17 by acquiring the mapping result 16 of the object to be reconstructed, and reconstructs the three-dimensional model 17 according to the mapping result 16, which is the output result, can be called by subsequent data asset three-dimensional model reconstruction platform, digital twin simulation platform. Among them, digital asset three-dimensional model reconstruction platform is a digital asset platform that can uniformly manage and schedule the reconstructed three-dimensional model and various types of multi-modal synthesized data derived from the three-dimensional model. According to the digital asset format, it can be divided into original three-dimensional model, color image, depth map, color point cloud, semantic map, single segmentation result, and mesh model, etc. It can accumulate synthesized data, and can be directly applied to embodied large model training. In addition, digital twin simulation platform can directly load dense three-dimensional model, and real-time display high-fidelity rendering result of three-dimensional space in simulator interface to user, support user real-time roaming in simulation scene, and can realize ranging and navigation downstream tasks by using accurate geometric information of three-dimensional model.
[0058] Please continue to refer to FIG. 2. As shown in FIG. 2, FIG. 2 is another application scenario of the method for reconstructing a three-dimensional model based on multi-source input according to an embodiment of the present application. Specifically,
[0059] The three-dimensional model reconstruction platform 12 can be deployed on the infrastructure 14 to realize fusion mapping according to multi-source input and output mapping results 16. Meanwhile, the three-dimensional model reconstruction platform 12 can also be deployed on the infrastructure 15 to reconstruct a three-dimensional model 17 according to the mapping results 16 and output the three-dimensional model 17. In addition, the user A can directly access the three-dimensional model reconstruction platform 12. In particular, when the three-dimensional model reconstruction platform 12 performs mapping on the object to be reconstructed, the user can monitor the mapping results in real time through the three-dimensional model reconstruction platform 12 and input corresponding instructions to control the mapping process. For example, the infrastructure 14 is a computing device deployed on the user side to be close to the object to be reconstructed 10, to realize mapping and calculation nearby, and to allow the user to view the mapping results 16 in real time. When there is local area missing sampling and detail missing, the user can quickly supplement sampling to realize high-precision sampling. Based on this, the infrastructure 15 is deployed on the cloud. The infrastructure 15 provided by the cloud provides the computing capability of the three-dimensional model reconstruction platform 12 in the process of reconstructing a three-dimensional model according to the mapping results 16, to meet the computing requirements of three-dimensional model reconstruction. Meanwhile, the digital asset three-dimensional model reconstruction platform and the digital twin simulation platform can directly call the three-dimensional model 17 on the cloud, which is more convenient and efficient, and improves the use efficiency. In another example, the infrastructure 14 and the infrastructure 15 can both be computing devices deployed on the user side, that is, the three-dimensional model reconstruction platform 12 directly runs on the user side device, thereby realizing the user's full-process control over the fusion mapping and three-dimensional model reconstruction process of the three-dimensional model reconstruction platform 12.
[0060] For the object to be reconstructed 10, in order to obtain its multi-source input data, the object to be reconstructed 10 is data-acquired by the acquisition device 11. Among them, the acquisition device 11 can be various mapping devices such as robots, professional data acquisition devices, portable handheld devices, etc. It can also be a laser radar, an IMU, an RGB (RGB Color Mode) camera and other sensors required for high-precision mapping. Through the methods of laser radar and camera external parameter calibration, laser radar and IMU establishment, and sensor hardware trigger time stamp synchronization, the data space alignment between all sensors is completed. At the same time, it can also include a computing unit (responsible for data acquisition and high-precision mapping module deployment and operation), a storage unit (responsible for storing and managing multi-modal data sets), and a network unit (responsible for high-speed data transmission). The above acquisition device can be one or a combination of multiple acquisition devices, and the present application does not limit this. Thus, when the acquisition device 11 collects the corresponding data, the initial image information, the initial point cloud information, and the inertial information are obtained. Raw data, among them, in order to facilitate data processing when the three-dimensional model reconstruction platform 12 is mapping, the data structure and the unified application programming interface (Application Programming Interface, API) interface can be unified in the acquisition process, so as to improve the mapping efficiency.
[0061] In order to more clearly illustrate the embodiment of the present application, the 3D Gaussian model will be used as an example to illustrate the three-dimensional model in the subsequent embodiments, wherein the 3D Gaussian model is an example of the three-dimensional model protected by the present application. As shown in the three-dimensional model definition, in the field of three-dimensional model reconstruction, there are still many other models that can realize the three-dimensional reconstruction of the scene, which should not be understood as specific limitation here.
[0062] Please continue to refer to FIG. 3, as shown in FIG. 3, FIG. 3 is a flowchart of a three-dimensional model reconstruction method based on multi-source input provided by an embodiment of the present application, specifically:
[0063] Step S201. Obtain the mapping result of the object to be reconstructed, and the mapping result includes image information of the object to be reconstructed and point cloud information of the object to be reconstructed.
[0064] The three-dimensional model reconstruction platform 12 obtains a mapping result of the object to be reconstructed as input of the three-dimensional model reconstruction. The mapping result can be provided by a user according to business needs and characteristics of the object to be reconstructed by using different methods, which is not limited in the present application. It is worth noting that the mapping result includes image information of the object to be reconstructed and point cloud information of the object to be reconstructed. The image information can be initial image information of the object to be reconstructed collected by a collection device, or specific image information obtained after fusion and fitting, which can reflect color and shape information of the object to be reconstructed. The point cloud information can be laser point cloud information, which can accurately determine depth information and geometric information of the object to be reconstructed. Compared with shape determined by a structured light or a pure vision scheme, the embodiment has higher precision and accuracy.
[0065] In an embodiment of the present application, the mapping result can be generated by the three-dimensional model reconstruction platform 12. In step S201, refer to FIG. 4 for specific steps. FIG. 4 is another flowchart of the three-dimensional model reconstruction method based on multi-source input provided by the embodiment of the present application. Specifically:
[0066] In step S301, initial image information and initial point cloud information of the object to be reconstructed are obtained, as well as image inertia information corresponding to the collection of the initial image information and point cloud inertia information corresponding to the collection of the initial point cloud information.
[0067] The three-dimensional model reconstruction platform 12 obtains initial image information and initial point cloud information of the object to be reconstructed, as well as image inertia information corresponding to the collection of the initial image information and point cloud inertia information corresponding to the collection of the initial point cloud information through the collection device 11.
[0068] Specifically, refer to FIG. 5. FIG. 5 is a mapping flowchart of the three-dimensional model reconstruction method based on multi-source input provided by the embodiment of the present application. Specifically:
[0069] In step S3011, initial point cloud information of the object to be reconstructed is obtained.
[0070] The three-dimensional model reconstruction platform 12 obtains initial point cloud information of the object to be reconstructed through data reported by the collection device 11. For example, the initial point cloud information can be a set collected by the collection device 11 such as a laser radar according to a certain collection path and corresponding pose. The initial point cloud information includes frame-by-frame point cloud information.
[0071] In step S3012, inertia information of the object to be reconstructed is obtained.
[0072] The three-dimensional model reconstruction platform 12 obtains the inertia information of the object to be reconstructed when collecting the initial point cloud information and the inertia information when collecting the initial image information through the data reported by the collection device 11. The collection device 11 can simultaneously integrate a laser radar, an RGB camera and an IMU to realize synchronous data collection and improve the collection accuracy. The inertia information of the object to be reconstructed is also the motion prior information of the initial point cloud information and the initial image information, and the motion prior information can include angular velocity, acceleration and the like.
[0073] Step S3013. Obtain the initial image information of the object to be reconstructed.
[0074] The three-dimensional model reconstruction platform 12 obtains the initial image information of the object to be reconstructed through the data reported by the collection device 11. The initial image information can be collected by the collection device 11 such as an RGB camera, and includes the frame-by-frame image of the object to be reconstructed during the collection process.
[0075] Step S302. Generate a mapping result according to the initial image information, the initial point cloud information, the point cloud inertia information and the image inertia information.
[0076] For example, the three-dimensional model reconstruction platform 12 fuses the frame-by-frame laser point cloud and the motion prior information provided by the IMU. Further, the image features and the motion prior information provided by the IMU are fused to jointly estimate the pose state of the entire initial image information and initial point cloud information, and generate a mapping result.
[0077] Step S3024. Generate point cloud pose information.
[0078] The three-dimensional model reconstruction platform 12 fuses the frame-by-frame laser point cloud and the motion prior information provided by the IMU, matches the features of adjacent point clouds, and estimates the point cloud pose information in real time.
[0079] Step S3025. Determine the descriptor information according to the initial image information.
[0080] The three-dimensional model reconstruction platform 12 determines the descriptor information of the frame-by-frame image according to the initial image information reported by the collection device 11. The descriptor information can be Perspective-n-Point (PnP) descriptor information.
[0081] Specifically, please refer to FIG. 6. As shown in FIG. 6, FIG. 6 is a key frame determination process schematic diagram of the three-dimensional model reconstruction method based on multiple source inputs provided by the embodiment of the application, and the steps are as follows:
[0082] Step S3025. Determine the descriptor information according to the initial image information.
[0083] By defining the description sub-information, the three-dimensional model reconstruction platform 12 can determine the description sub-information of the image frame by frame according to the initial image information. For example, the description sub-information is a PnP description sub, and specifically, the PnP description sub-information can include: PnP residual: image pose reliability index, high residual represents unreliable pose accuracy; PnP point pair number: image texture feature richness index, low point pair number represents serious visual degradation. According to the preset rule, the three-dimensional model reconstruction platform 12 performs frame-by-frame judgment on the initial image information. The preset rule can also be configured by the user to filter images that meet specific rules.
[0084] Through this step, the input data can be filtered, especially the relevant image data that meets certain requirements can be filtered, thereby providing specific images for subsequent mapping and three-dimensional model reconstruction.
[0085] Step S303. Determine the key frame image according to the initial image information and the description sub-information.
[0086] For example, according to the preset rule of the description sub-information, such as the user inputting PnP residual less than a first threshold value and / or PnP point pair number greater than a second threshold value, the three-dimensional model reconstruction platform 12 filters the images in the initial image information that meet the description sub-information frame by frame according to the description sub-information. The filtered images have pose accuracy and image texture feature richness that meet the PnP description sub requirements, and are considered as key frame images.
[0087] In this example, by filtering the initial image information, the key frame images with rich visual features, accurate pose and visual angle are retained, and the images with serious visual degradation and large pose error are excluded. Further, the three-dimensional model reconstruction platform 12 can exclude interference images during the mapping process, thereby improving the mapping accuracy. At the same time, the key frame images can also be configured to check the optical residual and geometric residual in the three-dimensional model reconstruction, thereby improving the accuracy and accuracy of the three-dimensional model reconstruction.
[0088] Step S3026. Generate image pose information.
[0089] The three-dimensional model reconstruction platform 12 fuses image features and IMU-provided motion prior information to estimate the space-time displacement relationship between adjacent two frames of images and output image pose.
[0090] Step S3027. The three-dimensional model reconstruction platform 12 performs loop detection.
[0091] Since the IMU has accumulated measurement errors over a long period of time, when the data acquisition device returns to the area where mapping has been completed after a long period of time, the current sensor information will be misaligned with the already built point cloud map. Therefore, loop detection needs to be performed during mapping to search for whether the current frame can match the historical information to form a loop, and a pose graph is constructed. In this way, the accuracy of the mapping data is improved, and the mapping precision is ensured.
[0092] For example, a robot moves along a corridor and enters a room, and then returns to the corridor again. The robot continuously collects images and records its pose (position and attitude). Each image is processed to extract key features, such as feature points (Oriented FAST and Rotated BRIEF, ORB), etc. Feature descriptors can be used to represent these feature points. When the robot moves, the newly collected images are matched with the previously collected images. A feature matching algorithm is used to find similar feature points. If a sufficient number of matching feature points are found (such as exceeding a certain threshold), it is considered that there is a potential loop. For each loop candidate, the corresponding image and robot pose are recorded. A pose graph is constructed, in which the nodes represent the poses of the robot, and the edges represent the relative transformations between adjacent poses. For each loop candidate, a constraint edge is added, indicating that the robot returns to the previous position. The feature matching results of the loop candidates and the robot pose information are used to verify the authenticity of the loop. Further verification can be performed by calculating the geometric consistency between loop candidates, such as using the Random Sample Consensus (RANSAC) algorithm to estimate the best relative pose. Once the loop is confirmed, the pose graph needs to be updated to reflect the latest loop information. A graph optimization algorithm is used to minimize the errors in the pose graph, maintaining the consistency of the map.
[0093] Step S3028. The three-dimensional model reconstruction platform 12 performs back-end optimization.
[0094] After loop detection is completed, the three-dimensional model reconstruction platform 12 performs closed-loop optimization on the mapping results based on the matching information of the current loop, eliminating the mapping misalignment phenomenon caused by sensor errors. The matching results are optimized through the constraints of point cloud information and image information.
[0095] Step S3029. The three-dimensional model reconstruction platform 12 generates the mapping results.
[0096] Exemplarily, the three-dimensional model reconstruction platform 12 can generate a mapping result corresponding to the initial point cloud information and the initial image information according to the image pose information and the point cloud pose information after the above steps are completed. The mapping result includes image information and point cloud information. The image information can be a set of images in the initial image information, and the point cloud information can be a set of point clouds in the initial point cloud information. The image information includes key frame images. As an example of the present application, the three-dimensional model reconstruction platform 12 does not determine the descriptor information and does not identify the key frame images, and the image information of the mapping result is only a subset of the initial image information.
[0097] It is worth noting that the above mapping process steps S3011-S3029 are an embodiment of the present application, and any step combination and any execution order can fall within the protection scope of the present application. In other embodiments, for example, steps S3027, S3025 and S3028 can be additional steps of the three-dimensional model reconstruction platform 12 for mapping according to the initial image information, the initial point cloud information and the corresponding initial inertial information.
[0098] By taking the multi-source input mode to map the object to be reconstructed, the mapping efficiency can be greatly improved. The point cloud information can provide more accurate geometric information and depth information of the object to be reconstructed. At the same time, the introduction of IMU inertial information can more quickly and accurately realize the input of the motion prior information of the collection device when collecting the initial point cloud information and the initial image information. In the subsequent mapping process, the accuracy is higher when performing pose matching and fitting, and the problem of identifying rough geometric information, inertial information, etc. through a large amount of calculation based on image information alone is avoided, which greatly improves the mapping efficiency and accuracy.
[0099] On this basis, in an embodiment of the present application, the three-dimensional model reconstruction platform 12 can also provide mapping feedback to the user in real time during the mapping process. When data missing, image missing, detail loss, etc. occur, the user can take remedial measures such as supplementary collection in time. Specifically:
[0100] Step S3023. The three-dimensional model reconstruction platform 12 acquires the supplementary collection instruction.
[0101] Exemplarily, the three-dimensional model reconstruction platform 12 waits for the supplementary collection data of the collection device 11 according to the supplementary collection instruction input by the user A, including the supplementary collection image information, the supplementary collection point cloud information, the supplementary collection image inertial information and the supplementary collection point cloud inertial information, and updates the mapping result according to the above supplementary collection data. The mapping efficiency is improved, and the integrity of the mapping result of the object to be reconstructed is ensured.
[0102] Exemplarily, the three-dimensional model reconstruction platform 12 provides part of the interface for the user to customize the query of related content.
[0103] GetCameraInfo() CameraInfo: Get the camera parameters and status at the current time.
[0104] Object definition: current image frame timestamp, image number, image intrinsic parameters, image pose information, whether it is a key frame, image data.
[0105] GetLidarInfo() LidarInfo: Get the laser radar parameters and status at the current time.
[0106] Object definition: laser radar type, including solid-state laser radar, multi-line laser radar; laser radar line number, laser radar current frame timestamp, laser point cloud frame number, laser radar pose information, whether the current position detects a loop, laser point cloud frame number associated with the loop, laser point cloud raw data.
[0107] GetIMUInfo() IMUInfo: Get the IMU parameters and status at the current time.
[0108] Object definition: IMU type, including six-axis and nine-axis; IMU timestamp, IMU data.
[0109] GetMapInfo() MapInfo: Get the global point cloud mapping result at the current time.
[0110] Object definition: point cloud coordinates, point cloud color values.
[0111] The interfaces that depend on user input include:
[0112] Set the PnP descriptor threshold, which is used for automatic key frame extraction and input parameters:
[0113] Residual: PnP residual threshold;
[0114] PairNum: PnP point pair number threshold.
[0115] It is worth noting that the above mapping process is an embodiment of the present application, and it should be understood that the above mapping steps do not limit the mapping result in the present application in the process of three-dimensional model reconstruction.
[0116] Step S202. Establish a three-dimensional model of the object to be reconstructed according to the mapping result.
[0117] In an embodiment of the present application, taking a three-dimensional model as a 3D Gaussian model as an example, the three-dimensional model reconstruction platform 12 establishes a 3D Gaussian model according to the mapping result, which generally includes the following five parameters:
[0118] Position: also known as mean, representing the coordinates of the center position of the 3D Gaussian kernel in three-dimensional space;
[0119] Covariance: represents the shape distribution of 3D Gaussian kernel, 3 column vectors in covariance matrix, representing the 3 principal axis directions of Gaussian ellipsoid;
[0120] Scaling factor: represents the size of each 3D Gaussian kernel;
[0121] Opacity: represents the transparency information of 3D Gaussian kernel, the higher the opacity, the closer the Gaussian kernel to the object surface;
[0122] Lighting parameter: encodes the lighting information under different viewing angles, which can reflect the lighting changes in the three-dimensional scene.
[0123] Exemplarily, feature points or landmarks are selected from the mapping result. These feature points can be fixed points in the environment, such as marks on the wall, corners of the room, etc. On this basis, the mean vector μ of each feature point, μ is usually the estimated position of the feature point in the world coordinate system. Further, the covariance matrix Σ is determined, which describes the uncertainty of the position of the feature point. A 3D Gaussian model is established, and for each selected feature point, a 3D Gaussian distribution model is constructed using the mean vector and covariance matrix determined above, and the 3D Gaussian model is completed.
[0124] Step S203. Determine target image information and target point cloud information.
[0125] Specifically, the target image information and the target point cloud information of the object to be reconstructed under the target pose observation are determined from the mapping result, wherein the target pose includes position information and attitude information. Wherein, the target pose can be the position information and attitude information used to observe the target image information and the target point cloud information in the mapping result, which can be input by the user A, or can be filtered according to the preset rule by the three-dimensional model reconstruction platform 12, for example, set a random rule, and the three-dimensional model reconstruction platform 12 determines the target pose according to the preset rule.
[0126] Step S2031. Determine the target pose according to the key frame image.
[0127] In an embodiment of the present application, the determination of the target pose can be determined according to the key frame image determined by the three-dimensional model reconstruction platform 12 in the mapping process. Specifically, when the key frame image is determined according to the descriptor information, it has corresponding pose information, and when the target pose is determined, the pose information of the key frame image can be directly determined as the target pose. As described above, the key frame image is usually determined according to the descriptor information, which has the characteristics of high image texture feature richness and accurate pose. Therefore, by determining the target pose according to the key frame image, the adjustment and training efficiency and effect can be improved in the subsequent three-dimensional model adjustment process.
[0128] Step S204. Indicating the three-dimensional model to generate a target depth image and a target digital image according to the target pose.
[0129] The three-dimensional model reconstruction platform 12 instructs the 3D Gaussian model to generate a target depth image and a target digital image according to the target pose. For example, the target digital image is taken as an example, the coordinates of the Gaussian model are transformed from the world coordinate system to the corresponding coordinate system according to the target pose, the position of the 3D Gaussian model in the target pose coordinate system is projected onto the 2D image plane, the part exceeding the boundary of the image is cropped, the points on the image plane are sampled, and the probabilities of these points corresponding to the 3D Gaussian model are calculated. The sampled probability values are converted into gray or color values, thereby obtaining the final image.
[0130] Step S205. Determining a geometric residual according to the target point cloud information and the target depth image.
[0131] The three-dimensional model reconstruction platform 12 determines a geometric residual according to the target point cloud information of the object to be reconstructed under the target pose in the mapping result and the target depth image generated by the 3D Gaussian model according to the target pose. For example, the three-dimensional model reconstruction platform 12 can determine the target point cloud information corresponding to the target pose in the target point cloud information according to the target pose, compare the target point cloud information and the target depth image, then calculate the distance between the points and the expected position of the Gaussian distribution, and determine the geometric residual.
[0132] Step S206. Determining an optical residual according to the target image information and the target digital image.
[0133] The three-dimensional model reconstruction platform 12 determines an optical residual according to the target image information of the object to be reconstructed under the target pose in the mapping result and the target image information generated by the 3D Gaussian model according to the target pose. For example, the three-dimensional model reconstruction platform 12 can determine the target image information corresponding to the target pose in the target image information according to the target pose, if the target pose is the pose information of the key frame image, the target image information is the key frame image, the residual can be defined as the difference between the actual pixel value and the rendered pixel value, the optical residual is determined by comparing the target image information and the target digital image.
[0134] Step S207. Adjusting the three-dimensional model according to the geometric residual and the optical residual.
[0135] Further, the three-dimensional model reconstruction platform 12 adjusts the 3D Gaussian model according to the optical residual and the geometric residual obtained in the above steps. For example, a nonlinear optimization algorithm such as gradient descent is used to minimize the geometric residual and / or the optical residual, and the parameters of each Gaussian distribution are updated until convergence or the maximum number of iterations is reached.
[0136] Thus, the image information and point cloud information included in the mapping result generated by the multi-source input, and the target depth image and target digital image generated according to the target pose are respectively used to construct optical residual error and geometric residual error, and the optical residual error and the geometric residual error are introduced to adjust the three-dimensional model at the same time. In particular, the geometric residual error is introduced to supervise the training of the three-dimensional model through the point cloud information, so that the three-dimensional model is adjusted according to the real scale information after adjustment, and the geometric accuracy and the richness of the detail features are higher.
[0137] Step S208. Anti-aliasing calculation is performed according to the target pose.
[0138] For example, when there is a serious inconsistency in the resolution between the rendering view angle input by the user and the training view angle in the rendering process of the 3D Gaussian model, abnormal rendering may occur. This is because the opacity parameter trained at the original resolution is only valid for the resolution. When the user performs a large-scale zoom operation, the opacity parameter error will cause abnormal occlusion in the image plane. Based on this, an embodiment of the present application specifically calculates the compensation coefficient of the opacity wherein Σ is a covariance matrix, I is a unit matrix, the compensation coefficient ρ of the opacity at the current pose is calculated, and the opacity of the target digital image is compensated in the rendering process of the 3D Gaussian model. By compensating the opacity parameter for the current resolution, the rasterization framework can correctly render the texture details at different resolutions.
[0139] Step S209. The three-dimensional model reconstruction platform 12 instructs the three-dimensional model to adjust the pose in the rendering process.
[0140] An embodiment of the present application can derive the image pose relative to the rasterization rendering function by configuring the rasterization rendering function in the rasterization rendering, and realize the supervision of the image pose through the gradient information of the geometric residual error and the optical residual error. When the image pose deviates by more than a certain threshold, the image pose is adjusted to reduce the error. Thus, the ghosting phenomenon on the 3D Gaussian model reconstruction result is reduced, and the rendering quality is further improved.
[0141] The three-dimensional model reconstruction platform 12 can provide a corresponding interface for the user to configure and control the three-dimensional model reconstruction process. For example:
[0142] GetGaussMapInfo() GaussMapInfo: Obtain 3D Gaussian model parameters.
[0143] Object definition: Gaussian kernel center position, covariance, scaling coefficient, opacity, spherical harmonic function parameter.
[0144] GetTrainingInfo() TrainingInfo: Get the current state and feedback of the 3D Gaussian model training task.
[0145] Object definition: iteration steps, optical residual error, geometric residual error, peak signal-to-noise ratio index (used to evaluate image rendering quality), structural similarity (used to evaluate the similarity between the rendered image and the true value)
[0146] GetCameraOptimizationInfo() CameraOptimizationInfo: Get the image pose optimization result.
[0147] Object definition: key frame number, key frame timestamp, original pose, optimized image pose, image pose correction amplitude
[0148] Interfaces that depend on user input include:
[0149] SetConfig(Params, UseGeoLoss, UseCameraOpt, UseAntiAliasing): Set the training parameters and input parameters of the 3D Gaussian reconstruction module.
[0150] Params: including basic parameter configurations such as training iteration steps and optimizer learning rate, the reconstruction framework provides recommended settings according to the size of the reconstructed scene for user reference:
[0151] UseGeoLoss: whether to use geometric residual error supervision;
[0152] UseCameraOpt: whether to enable camera extrinsic parameter joint optimization;
[0153] UseAntiAliasing: whether to enable anti-aliasing calculation.
[0154] Based on this, in an embodiment of the present application, user A can configure whether to enable anti-aliasing calculation, training iteration steps, etc. through the three-dimensional model reconstruction platform 12. For example, user A inputs the iteration number, and the three-dimensional model reconstruction platform 12 adjusts the number of iterations of the three-dimensional model according to the iteration number input by user A. When the three-dimensional model is iterated according to the geometric residual error and the optical residual error, the adjusted three-dimensional model is output.
[0155] The embodiment of the present application also provides a three-dimensional model reconstruction platform 12, which will be described below with reference to FIG. 8. FIG. 8 is a structural schematic diagram of a three-dimensional model reconstruction platform according to an embodiment of the present application. Specifically, the following modules are included:
[0156] The acquisition module 901 is configured to acquire mapping results of the object to be reconstructed, wherein the mapping results comprise image information of the object to be reconstructed and point cloud information of the object to be reconstructed.
[0157] The establishment module 903 is configured to establish a three-dimensional model of the object to be reconstructed according to the mapping results.
[0158] The determination module 902 is further configured to determine target image information and target point cloud information of the object to be reconstructed under a target pose from the mapping results, wherein the target pose comprises position information and attitude information.
[0159] The indication module 904 is configured to instruct the three-dimensional model to generate a target depth image and a target digital image according to the target pose.
[0160] The determination module 902 is configured to determine a geometric residual according to the target point cloud information and the target depth image, and further configured to determine an optical residual according to the target image information and the target digital image.
[0161] The adjustment module 905 is configured to adjust the three-dimensional model according to the geometric residual and the optical residual.
[0162] In another embodiment of the present application, the three-dimensional model reconstruction platform 12 can further comprise:
[0163] The generation module 906 is configured to generate the mapping results according to the initial image information, the initial point cloud information, the point cloud inertia information and the image inertia information.
[0164] The update module 907 is configured to update the mapping results according to the supplementary image information, the supplementary point cloud information, the supplementary laser inertia information and the supplementary image inertia information.
[0165] The output module 908 is configured to output the three-dimensional model when the number of iterations is satisfied.
[0166] It is worth noting that the user in the above embodiment can be replaced by other multiple users, and the above modules can all realize corresponding technical functions, and the embodiments of the present application do not limit this.
[0167] The embodiments of the present application take the acquisition module 901, the determination module 902, the establishment module 903, the indication module 904 and the adjustment module 905 as examples for illustration, and similarly, the implementation modes of the generation module 906, the update module 907 and the output module 908 can refer to the implementation modes of the foregoing modules.
[0168] Specifically, the obtaining module 901, the determining module 902, the establishing module 903, the instructing module 904, and the adjusting module 905 can be implemented by software or by hardware. For example, the implementation of the obtaining module 901 is described below. Similarly, the implementation of the determining module 902, the establishing module 903, the instructing module 904, and the adjusting module 905 can refer to the implementation of the obtaining module 901.
[0169] As an example of the software functional unit, the obtaining module 901 can include code running on a computing instance. The computing instance can include at least one of a physical host (computing device), a virtual machine, and a container. Further, the computing instance can be one or more. For example, the obtaining module 901 can include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers running the code can be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers running the code can be distributed in the same availability zone or in different availability zones, and each availability zone includes one data center or multiple data centers in a similar geographical location. Generally, one region can include multiple availability zones.
[0170] As an example of the hardware functional unit, the obtaining module 901 can include at least one computing device, such as a server. Alternatively, the obtaining module 901 can be a device implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), etc. The PLD can be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0171] It should be noted that in other embodiments, the acquisition module 901 can be configured to perform any step of the method for reconstructing a three-dimensional model based on multi-source input, the determination module 902 can be configured to perform any step of the method for reconstructing a three-dimensional model based on multi-source input, the establishment module 903 can be configured to perform any step of the method for reconstructing a three-dimensional model based on multi-source input, the indication module 904 can be configured to perform any step of the method for reconstructing a three-dimensional model based on multi-source input, the adjustment module 905 can be configured to perform any step of the method for reconstructing a three-dimensional model based on multi-source input, the steps implemented by the acquisition module 901, the determination module 902, the establishment module 903, the indication module 904, and the adjustment module 905 can be specified as needed, and the acquisition module 901, the determination module 902, the establishment module 903, the indication module 904, and the adjustment module 905 respectively implement different steps in the method for reconstructing a three-dimensional model based on multi-source input to realize all functions of the three-dimensional model reconstruction platform.
[0172] The above describes the method, the three-dimensional model reconstruction platform, and the system of the embodiments of the present application in detail. In order to better implement the above-mentioned schemes of the embodiments of the present application, the related devices for implementing the above-mentioned schemes are also provided as follows.
[0173] The present application provides a computing device, which is described below with reference to FIG. 9. FIG. 9 is a structural schematic diagram of a computing device of a method for reconstructing a three-dimensional model based on multi-source input according to an embodiment of the present application. The computing device 900 includes a bus 911, a processor 912, a memory 910, and a communication interface 909. The processor 912, the memory 910, and the communication interface 909 communicate through the bus 911. The computing device 900 can be a server or a terminal device. It should be understood that the number of processors and memories in the computing device 900 is not limited in the present application.
[0174] The bus 911 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only one line is shown in FIG. 9, but it does not mean that there is only one bus or only one type of bus. The bus 911 can include a path for transmitting information between various components (for example, the memory 910, the processor 912, and the communication interface 909) of the computing device 900.
[0175] The processor 912 can include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), among other processors.
[0176] The memory 910 can include volatile memory (such as random access memory (RAM)), and the processor 912 can further include non-volatile memory (such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid-state drive (SSD)).
[0177] The memory 910 stores executable program code, and the processor 912 executes the executable program code to respectively implement the functions of the obtaining module 901, the determining module 902, the establishing module 903, the instructing module 904, and the adjusting module 905, so as to implement the method for reconstructing a three-dimensional model based on multi-source input. That is, the memory 910 stores instructions for executing the method for reconstructing a three-dimensional model based on multi-source input by the three-dimensional model reconstruction platform.
[0178] The communication interface 909 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 900 and other devices or communication networks.
[0179] The embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a notebook computer, or a smart phone.
[0180] Please refer to FIG. 10, which is a structural schematic diagram of a computing device cluster for the method for reconstructing a three-dimensional model based on multi-source input according to an embodiment of the present application. As shown in FIG. 10, the computing device cluster includes at least one computing device 900, and the memory 910 in one or more computing devices 900 in the computing device cluster can store the same instructions for executing the method for reconstructing a three-dimensional model based on multi-source input by the three-dimensional model reconstruction platform.
[0181] In some possible implementation manners, one or more of the computing devices 900 in the computing device cluster can also be configured to execute the instructions of the three-dimensional model reconstruction platform for performing the method for reconstructing a three-dimensional model based on multi-source input. In other words, the combination of one or more of the computing devices 900 can collectively execute the instructions of the three-dimensional model reconstruction platform for performing the method for reconstructing a three-dimensional model based on multi-source input.
[0182] It should be noted that the memories 910 in different computing devices 900 in the computing device cluster can store different instructions for performing part of the functions of the three-dimensional model reconstruction platform. That is, the instructions stored in the memories 910 in different computing devices 900 can implement the functions of one or more of the obtaining module 901, the determining module 902, the establishing module 903, the indicating module 904, and the adjusting module 905.
[0183] In some possible implementation manners, the memories 910 in one or more of the computing devices 900 in the computing device cluster can also respectively store instructions for performing part of the method for reconstructing a three-dimensional model based on multi-source input. In other words, the combination of one or more of the computing devices 900 can collectively execute the instructions for performing the method for reconstructing a three-dimensional model based on multi-source input.
[0184] The following refers to FIG. 11, which is another structural schematic diagram of a computing device cluster for the method for reconstructing a three-dimensional model based on multi-source input according to an embodiment of the present application. As shown in FIG. 11, two computing devices 900A and 900B are connected through the communication interfaces 909. The memory in the computing device 900A stores instructions for executing the determining module 902, the establishing module 903, and the adjusting module 905. The memory in the computing device 900B stores instructions for executing the obtaining module 901 and the indicating module 904. In other words, the memories 910 in the computing devices 900A and 900B collectively store the instructions of the three-dimensional model reconstruction platform for performing the method for reconstructing a three-dimensional model based on multi-source input.
[0185] The connection manner between the computing device cluster shown in FIG. 11 can be that, considering that the method for reconstructing a three-dimensional model based on multi-source input provided in the present application needs to perform a large amount of data transmission on the obtaining module 901, the functions are implemented by the computing device 900B to avoid overloading of the computing device 900A.
[0186] It should be understood that the functions of the computing device 900A shown in FIG. 11 can also be completed by multiple computing devices 900. Similarly, the functions of the computing device 900B can also be completed by multiple computing devices 900.
[0187] Please refer to Figure 12, which is another structure diagram of the computing device cluster of the method for reconstructing a three-dimensional model based on multi-source input according to an embodiment of the present application. In some possible implementation manners, one or more computing devices in the computing device cluster can be connected through a network. The network can be a wide area network or a local area network, etc. Figure 12 shows a possible implementation manner. As shown in Figure 12, two computing devices 900C and 900D are connected through a network. Specifically, the computing devices are connected to the network through communication interfaces in the computing devices. In this kind of possible implementation manner, the memory 910 in the computing device 900C stores instructions for executing the determining module 902, the establishing module 903 and the adjusting module 905. Meanwhile, the memory 910 in the computing device 900D stores instructions for executing the obtaining module 901 and the indicating module 904.
[0188] The connection manner between the computing devices in Figure 12 can be that, considering that the method for reconstructing a three-dimensional model based on multi-source input provided by the present application needs to perform a large amount of data transmission and needs to be connected through a network, the functions are relatively independent, in order to make the storage and computing performance optimal, the data transmission function is considered to be executed by the computing device 900D.
[0189] It should be understood that the functions of the computing device 900C shown in Figure 12 can also be completed by a plurality of computing devices 900. Similarly, the functions of the computing device 900D can also be completed by a plurality of computing devices 900.
[0190] In some possible implementation manners, the memory 910 of one or more computing devices 900 in the computing device cluster can also respectively store instructions for executing the method for reconstructing a three-dimensional model based on multi-source input. In other words, the combination of one or more computing devices 900 can collectively execute instructions for executing the method for reconstructing a three-dimensional model based on multi-source input.
[0191] The embodiment of the present application further provides a computer program product containing instructions. The computer program product can be software or a program product containing instructions, which can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, the at least one computing device is caused to execute the method for reconstructing a three-dimensional model based on multi-source input applied to a three-dimensional model reconstruction platform.
[0192] The embodiments of the present application further provide a computer readable storage medium. The computer readable storage medium can be any available medium or data storage device that can store the instructions of the computer device or a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk), etc. The computer readable storage medium includes instructions instructing the computer device to execute the above-mentioned method for performing the three-dimensional model reconstruction based on the multi-source input.
[0193] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the above-mentioned embodiments of the present application have been described in detail, those skilled in the art should understand that: it can still modify the technical solutions recorded in the above-mentioned embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the protection scope of the technical solutions of the embodiments of the present application.
[0194] Those skilled in the art can clearly understand the specific working process of the above-mentioned system, three-dimensional model reconstruction platform or unit, which can refer to the corresponding process in the above-mentioned method embodiments, and will not be repeated here.
[0195] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed in the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0196] The embodiments of the present application further provide a computer program product containing instructions. The computer program product can be software or program product containing instructions, which can run on a computer device or be stored in any available medium. When the computer program product runs on at least one computer device, it makes at least one computer device execute the above-mentioned method for performing the three-dimensional model reconstruction based on the multi-source input.
[0197] The embodiments of the present application also provide a computer readable storage medium. The computer readable storage medium can be any available medium or data storage device that can store the instructions of the computer device, or a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk), etc. The computer readable storage medium includes instructions indicating the computer device to execute the above-mentioned method for performing the three-dimensional model reconstruction based on the multi-source input.
[0198] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the protection scope of the technical solutions of the embodiments of the present application.
[0199] Those skilled in the art can clearly understand the specific working process of the above-mentioned system, three-dimensional model reconstruction platform or unit, which can refer to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0200] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited to this, any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for 3D model reconstruction based on multi-source input, characterized in that, The method is applied to a three-dimensional model reconstruction platform, and the method comprises: acquiring a mapping result of an object to be reconstructed, the mapping result comprising image information of the object to be reconstructed and point cloud information of the object to be reconstructed; establishing a three-dimensional model of the object to be reconstructed according to the mapping result; determining target image information and target point cloud information of the object to be reconstructed under a target pose from the mapping result, the target pose comprising position information and attitude information; instructing the three-dimensional model to generate a target depth image and a target digital image according to the target pose; determining a geometric residual according to the target point cloud information and the target depth image; determining an optical residual according to the target image information and the target digital image; adjusting the three-dimensional model according to the geometric residual and the optical residual.
2. The method of claim 1, wherein, The acquiring of the mapping result of the object to be reconstructed comprises: acquiring initial image information and initial point cloud information of the object to be reconstructed, and image inertia information corresponding to the acquisition of the initial image information and point cloud inertia information corresponding to the acquisition of the initial point cloud information by the acquisition device; generating the mapping result according to the initial image information, the initial point cloud information, the point cloud inertia information and the image inertia information.
3. The method according to claim 1 or 2, characterized in that, The method further comprises: determining description sub-information according to the initial image information, the description sub-information comprising texture features of at least one frame of image in the initial image information; determining a key frame image according to the initial image information and the description sub-information.
4. The method of claim 3, wherein, The method further comprises: determining the target pose for observing the key frame image according to the key frame image.
5. The method according to any one of claims 1 to 4, characterized in that, The method further comprises: acquiring supplementary image information and supplementary point cloud information of the object to be reconstructed, and supplementary image inertia information and supplementary laser inertia information; updating the mapping result according to the supplementary image information, the supplementary point cloud information, the supplementary laser inertia information and the supplementary image inertia information.
6. The method according to any one of claims 1 to 5, characterized in that, The method further comprises: acquiring an iteration number input by a user, the iteration number being used to indicate the number of adjustments of the three-dimensional model; outputting the three-dimensional model when the iteration number is met.
7. A three-dimensional model reconstruction platform, characterized in that, The three-dimensional model reconstruction platform comprises: an acquisition module, which is used to acquire a mapping result of an object to be reconstructed, the mapping result comprising image information of the object to be reconstructed and point cloud information of the object to be reconstructed; an establishment module, which is used to establish a three-dimensional model of the object to be reconstructed according to the mapping result; a determination module, which is used to determine target image information and target point cloud information of the object to be reconstructed under a target pose from the mapping result, the target pose comprising position information and attitude information; an instruction module, which is used to instruct the three-dimensional model to generate a target depth image and a target digital image according to the target pose; the determination module is further used to determine a geometric residual according to the target point cloud information and the target depth image; the determination module is further used to determine an optical residual according to the target image information and the target digital image; an adjusting module configured to adjust the three-dimensional model according to the geometric residual error and the optical residual error.
8. The three-dimensional model reconstruction platform of claim 7, wherein, The acquisition module is configured to acquire mapping results of the object to be reconstructed. The acquisition module is specifically configured to acquire initial image information and initial point cloud information of the object to be reconstructed, and image inertia information corresponding to acquisition of the initial image information and point cloud inertia information corresponding to acquisition of the initial point cloud information by the acquisition device. The three-dimensional model reconstruction platform further comprises: a generation module configured to generate the mapping results according to the initial image information, the initial point cloud information, the point cloud inertia information, and the image inertia information.
9. The three-dimensional model reconstruction platform of claim 7 or 8, wherein The determination module is further configured to determine description sub-information according to the initial image information, the description sub-information including texture features of at least one frame of image in the initial image information. The determination module is further configured to determine a key frame image according to the initial image information and the description sub-information.
10. The three-dimensional model reconstruction platform of claim 9, wherein The acquisition module is further configured to determine the target pose for observing the key frame image according to the key frame image.
11. The three-dimensional model reconstruction platform of any one of claims 7 to 10, wherein The acquisition module is further configured to acquire supplementary image information, supplementary point cloud information, supplementary image inertia information, and supplementary laser inertia information of the object to be reconstructed; and the three-dimensional model reconstruction platform further comprises: an updating module configured to update the mapping results according to the supplementary image information, the supplementary point cloud information, the supplementary laser inertia information, and the supplementary image inertia information.
12. The three-dimensional model reconstruction platform of any one of claims 7 to 11, wherein, The acquisition module is further configured to acquire an iteration number input by a user, the iteration number being used to indicate the number of adjustments of the three-dimensional model; and the three-dimensional model reconstruction platform further comprises: an output module configured to output the three-dimensional model when the iteration number is met. at least one computing device, each computing device including a processor and a memory; 13. A cluster of computing devices, characterized in that, the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method of any one of claims 1 to 6. The instructions, when executed by the cluster of computer devices, cause the cluster of computer devices to perform the method of any one of claims 1 to 6.
14. A computer program product comprising instructions, characterized in that, computer program instructions, which, when executed by a cluster of computing devices, cause the cluster of computing devices to perform the method of any one of claims 1 to 6.
15. A computer-readable storage medium, characterized in that,
Citation Information
Patent Citations
Three-dimensional reconstruction model generation method and system, electronic equipment and storage medium
CN114022639A
Three-dimensional object pose recognition method and device based on visual image, and electronic equipment
CN116309836A
Three-dimensional reconstruction method and system
CN116958452A
Three-dimensional model construction and rendering method and device, equipment and medium
CN117689826A
Simultaneous Localization and Mapping Method, Device, System and Storage Medium
US20230260151A1
Cited By
Gaussian scene suspension point elimination method, system and device and storage medium
CN121330304A