Three-dimensional reconstruction method and system based on multi-view geometric constraint, medium and equipment
Through a deep learning method based on multi-view geometric constraints, the high cost and complexity of existing three-dimensional reconstruction technologies are solved, low-cost and flexible three-dimensional reconstruction are achieved, and the ability to generalize real-world data is improved.
Patent Information
- Application Number
- CN202510094182.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
The existing three-dimensional reconstruction technology has high cost and complexity in use, making it difficult to achieve low cost and flexibility.
Deep learning method based on multi-view geometric constraints is adopted, and training data preparation is prepared by obtaining actual three-dimensional digital models of multiple different objects to be measured, and deep learning models are constructed to generate fine three-dimensional point cloud data from rough three-dimensional models.
A three-dimensional reconstruction technology with lower cost and higher flexibility is realized, reducing reconstruction costs and improving the ability to generalize real-world data.
Smart Images

Figure CN120014167A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of three-dimensional reconstruction technology, and in particular to a three-dimensional reconstruction method, system, medium and equipment based on multi-view geometric constraints. Background Art
[0002] In recent years, 3D reconstruction pre-measurement technology has been widely used in many fields, such as customized orthopedic insoles or footwear, medical diagnosis, e-commerce display, and cultural relics protection. The main challenge of current 3D reconstruction technology is the high cost, as well as the time and training required for professionals to adapt to the use of 3D scanners. Therefore, achieving low-cost and flexible 3D reconstruction is an important and challenging task.
[0003] Existing 3D reconstruction technologies can be roughly divided into: (1) Optical measurement technology. Industrial optical 3D scanners based on structured light or laser scanning can achieve high-precision reconstruction, but due to their high cost and limited scope of use, they are only applied on a small scale. (2) Deep learning technology. Current 3D foot shape reconstruction technologies based on deep learning can be divided into two categories: recovering the complete shape from incomplete data and predicting the 3D object model from the image. The first technology only requires a depth map from a single viewpoint as the input of the network, and then through forward propagation, the network generates complete 3D object data. The most effective method is SnowflakeNet, but there are two limitations to the development of this technology: 1) the size of the dataset needs to be increased; 2) the current data is highly artificial. The second technology directly reconstructs the 3D object model from the image. Usually, principal component analysis is used to compress the shape of the object into specific parameters. Then, the designed CNN (Convolutional Neural Network) architecture regresses these shape parameters given multiple object contours and corresponding camera positions. However, this technique relies on parameterized models and has poor generalization capabilities to real-world data. Summary of the invention
[0004] Based on this, it is necessary to propose a three-dimensional reconstruction method, system, medium and equipment based on multi-view geometric constraints to address the above problems.
[0005] A three-dimensional reconstruction method based on multi-view geometric constraints, the method comprising:
[0006] Acquire actual three-dimensional digital models of multiple different objects to be measured, perform a first slicing process on the actual three-dimensional digital models to obtain a first cross-sectional view; obtain images of the object to be measured at different perspectives based on the actual three-dimensional digital models, and construct a rough three-dimensional model of the object to be measured based on the images at different perspectives; perform a second slicing process on the rough three-dimensional model to obtain a second cross-sectional view; wherein the slicing density of the first slicing process is greater than the slicing density of the second slicing process; use the first cross-sectional view and the second cross-sectional view as training data sets for a deep learning model; construct a deep learning model, use the second cross-sectional view as input data of the deep learning model, use the first cross-sectional view as output data of the deep learning model, train the deep learning model, and obtain a trained deep learning model.
[0007] Acquire photos of the object to be tested at different viewing angles, construct a rough three-dimensional model of the object to be tested based on the photos at different viewing angles, and perform the second slicing process on the rough three-dimensional model to obtain a rough cross-sectional image of the object to be tested.
[0008] The rough cross-sectional image is input into a trained deep learning model to obtain a predicted cross-sectional image output by the deep learning model, and three-dimensional point cloud data of the object to be measured is reconstructed based on the predicted cross-sectional image.
[0009] The step of obtaining images of the object under test at different viewing angles based on the actual three-dimensional digital model and constructing a rough three-dimensional model of the object under test based on the images at different viewing angles specifically includes:
[0010] The images of the actual three-dimensional digital model of the object under test at different viewing angles are obtained through a virtual camera.
[0011] The edge extraction algorithm is used to extract the contour data of the object under test from images of different viewing angles.
[0012] A rough three-dimensional model of the object to be measured is constructed based on the contour data of the object to be measured.
[0013] The step of constructing a deep learning model, using the second cross-sectional image as input data of the deep learning model, using the first cross-sectional image as output data of the deep learning model, training the deep learning model, and obtaining a trained deep learning model specifically includes:
[0014] Construct a deep learning model, use the second cross-sectional image as input data of the deep learning model, and use the first cross-sectional image as output data of the deep learning model.
[0015] The value of the loss function is determined according to the true label and the predicted label of the pixel in the first slice image.
[0016] Parameters of the deep learning model are optimized according to the value of the loss function to determine a trained deep learning model.
[0017] The step of determining the value of the loss function according to the true label and the predicted label of the pixel in the first slice image specifically includes:
[0018] The loss function is:
[0019]
[0020] Among them, BC is the value of the loss function, N is the number of pixels in the first slice image, s is the number of slice images output by the network, and y i is the true label of the i-th pixel in the first slice image, is the predicted label of the i-th pixel in the first slice image.
[0021] The acquiring of photos of the object to be tested at different viewing angles, constructing a rough three-dimensional model of the object to be tested based on the photos at different viewing angles, and performing the second slicing process on the rough three-dimensional model to obtain a rough cross-sectional view of the object to be tested specifically includes:
[0022] Place the object to be tested on the calibration target, and use the camera to shoot the object to be tested and the calibration target from different angles to obtain photos of the object to be tested from different perspectives;
[0023] The intrinsic and extrinsic parameters of the camera at each viewing angle are calibrated by the calibration target, wherein the extrinsic parameters are the camera pose parameters at different viewing angles;
[0024] Use edge extraction algorithm to extract the contour data of the object to be tested from the photos of the object to be tested at different viewing angles;
[0025] Based on a multi-view geometric constraint algorithm, a rough three-dimensional model of the object to be measured is constructed using the intrinsic and extrinsic parameters of the camera at each viewing angle and the contour data of the object to be measured;
[0026] The second slicing process is performed on the rough three-dimensional model to obtain a rough cross-sectional image of the object to be tested.
[0027] The step of inputting the rough cross-sectional image into the trained deep learning model to obtain the predicted cross-sectional image output by the deep learning model, and reconstructing the three-dimensional point cloud data of the object to be measured based on the predicted cross-sectional image, specifically includes:
[0028] The three-dimensional point cloud data of the object to be measured is reconstructed into a grid model according to the Poisson reconstruction method.
[0029] A three-dimensional reconstruction system based on multi-view geometric constraints, the system comprising:
[0030] A deep learning model acquisition module is used to acquire actual three-dimensional digital models of multiple different objects to be measured, perform a first slicing process on the actual three-dimensional digital models to obtain a first cross-sectional view; obtain images of the object to be measured from different perspectives based on the actual three-dimensional digital models, and construct a rough three-dimensional model of the object to be measured based on the images from different perspectives; perform a second slicing process on the rough three-dimensional model to obtain a second cross-sectional view; wherein the slicing density of the first slicing process is greater than the slicing density of the second slicing process; use the first cross-sectional view and the second cross-sectional view as training data sets for a deep learning model; construct a deep learning model, use the second cross-sectional view as input data of the deep learning model, use the first cross-sectional view as output data of the deep learning model, train the deep learning model, and obtain a trained deep learning model.
[0031] The rough cross-sectional image acquisition module is used to acquire photos of the object to be tested from different perspectives, construct a rough three-dimensional model of the object to be tested based on the photos from different perspectives, and perform the second slicing process on the rough three-dimensional model to obtain a rough cross-sectional image of the object to be tested.
[0032] The three-dimensional model reconstruction module is used to input the rough section image into the trained deep learning model to obtain the predicted section image output by the deep learning model, and reconstruct the three-dimensional point cloud data of the object to be measured based on the predicted section image.
[0033] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the steps of the method described above.
[0034] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.
[0035] The embodiments of the present invention have the following beneficial effects:
[0036] The present invention first prepares training data by acquiring actual three-dimensional digital models of multiple different objects to be measured, performs a first slicing process on the actual three-dimensional digital model to obtain a first cross-sectional view with high density, further, constructs a rough three-dimensional model of the object to be measured based on the actual three-dimensional digital model, performs a second slicing process on the rough three-dimensional model to obtain a second cross-sectional view with low density, and uses the first cross-sectional view and the second cross-sectional view to train a deep learning model to achieve a three-dimensional reconstruction technology with lower cost and higher flexibility. The input of the deep learning model is the second cross-sectional view of the rough three-dimensional model, and the output is a predicted cross-sectional view that is as fine as the first cross-sectional view. Subsequently, for the object to be measured, a rough three-dimensional model is constructed using photos of the object from different perspectives, and a second slicing process is performed on the object to be measured to obtain a rough cross-sectional view. Finally, this rough cross-sectional view is input into the trained deep learning model, and the model outputs a predicted cross-sectional view. According to the predicted cross-sectional view, the complete three-dimensional point cloud data of the object to be measured can be reconstructed, thereby achieving three-dimensional reconstruction of the object to be measured. In summary, the present invention uses a large amount of real data to train a deep learning model, which can automatically learn from photos and generate three-dimensional point cloud data, greatly reducing the reconstruction cost and improving the generalization ability of real-world data. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0038] in:
[0039] Figure 1 A schematic flow chart of an embodiment of a 3D reconstruction method based on multi-view geometric constraints provided by the present invention;
[0040] Figure 2 A schematic flow chart of another embodiment of a 3D reconstruction method based on multi-view geometric constraints provided by the present invention;
[0041] Figure 3 A second slice diagram provided by the present invention;
[0042] Figure 4 A schematic structural diagram of an embodiment of a 3D reconstruction system based on multi-view geometric constraints provided by the present invention;
[0043] Figure 5 A schematic structural diagram of an embodiment of the device provided by the present invention;
[0044] Figure 6 A schematic structural diagram of an embodiment of the medium provided by the present invention. DETAILED DESCRIPTION
[0045] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0046] like Figure 1 As shown, Figure 1 A schematic flow chart of an embodiment of a 3D reconstruction method based on multi-view geometric constraints provided by the present invention. A 3D reconstruction method based on multi-view geometric constraints, characterized in that the method comprises:
[0047] S101: Acquire actual three-dimensional digital models of multiple different objects to be measured, perform a first slicing process on the actual three-dimensional digital models to obtain a first cross-sectional view; obtain images of the object to be measured at different perspectives based on the actual three-dimensional digital models, and construct a rough three-dimensional model of the object to be measured based on the images at different perspectives; perform a second slicing process on the rough three-dimensional model to obtain a second cross-sectional view; wherein the slicing density of the first slicing process is greater than the slicing density of the second slicing process; use the first cross-sectional view and the second cross-sectional view as training data sets for a deep learning model; construct a deep learning model, use the second cross-sectional view as input data of the deep learning model, use the first cross-sectional view as output data of the deep learning model, train the deep learning model, and obtain a trained deep learning model.
[0048] Exemplarily, a large number of actual three-dimensional digital models of human feet are scanned and acquired by precise measurement means, and the actual three-dimensional digital models of human feet are subjected to a first slicing process with denser slices to obtain first sectional images, for example, 101 slices are cut to obtain 101 first sectional images.
[0049] Furthermore, the images of the actual three-dimensional digital model of the object under test at different viewing angles are obtained through a virtual camera, and the contour data of the object under test in the images of the object under test at different viewing angles are extracted using an edge extraction algorithm, and a rough three-dimensional model of the object under test is constructed based on the contour data of the object under test. The rough three-dimensional model is subjected to a second slicing process along the Z axis according to a number of slices of 6, and 6 second cross-sectional images are obtained; wherein the slicing density of the first slicing process is greater than the slicing density of the second slicing process.
[0050] The first and second slice images are used as training data sets for the deep learning model. A deep learning model is constructed, and the second slice image is used as input data for the deep learning model, and the first slice image is used as output data for the deep learning model. The value of the loss function is determined based on the true label and predicted label of the pixels in the first slice image. The loss function is:
[0051]
[0052] Among them, BC is the value of the loss function, N is the number of pixels in the first slice image, s is the number of slice images output by the network, and y i is the true label of the i-th pixel in the first slice image, is the predicted label of the i-th pixel in the first slice image.
[0053] Furthermore, the parameters of the deep learning model are optimized according to the value of the loss function to determine the trained deep learning model.
[0054] S102: Acquire photos of the object to be tested from different perspectives, construct a rough three-dimensional model of the object to be tested based on the photos from different perspectives, and perform a second slicing process on the rough three-dimensional model to obtain a rough cross-sectional image of the object to be tested.
[0055] Exemplarily, the object to be tested is placed on a calibration target, and the object to be tested and the calibration target are photographed from different angles by a camera to obtain photos of the object to be tested from different perspectives; the intrinsic and extrinsic parameters of the camera at each perspective are calibrated by the calibration target, and the extrinsic parameters are the camera posture parameters at different perspectives; the edge extraction algorithm is used to extract the contour data of the object to be tested from the photos of the object to be tested from different perspectives; based on a multi-view geometric constraint algorithm, a rough three-dimensional model of the object to be tested is constructed using the intrinsic and extrinsic parameters of the camera at each perspective, as well as the contour data of the object to be tested; the rough three-dimensional model is subjected to a second slicing process to obtain a rough cross-sectional image of the object to be tested.
[0056] S103: Input the rough cross-sectional image into the trained deep learning model to obtain the predicted cross-sectional image output by the deep learning model, and reconstruct the three-dimensional point cloud data of the object to be measured based on the predicted cross-sectional image.
[0057] Exemplarily, a rough cross-section image is input into a trained deep learning model to obtain a predicted cross-section image output by the deep learning model; the rough cross-section image and the predicted cross-section image of the object to be measured are used as image data for accurate three-dimensional reconstruction, and the intrinsic and extrinsic parameters of the camera at each viewing angle are used to reversely project rays from the image data, and the coordinates of a series of intersection points are calculated by the intersection of rays from each viewpoint to reconstruct the three-dimensional point cloud data of the object to be measured.
[0058] It can be seen from the above description that the present invention first prepares training data by obtaining actual three-dimensional digital models of multiple different objects to be measured, performs a first slicing process on the actual three-dimensional digital model to obtain a first cross-sectional view with high density, and further, constructs a rough three-dimensional model of the object to be measured based on the actual three-dimensional digital model, performs a second slicing process on the rough three-dimensional model to obtain a second cross-sectional view with low density, and uses the first cross-sectional view and the second cross-sectional view to train a deep learning model to achieve a three-dimensional reconstruction technology with lower cost and higher flexibility. The input of the deep learning model is the second cross-sectional view of the rough three-dimensional model, and the output is a predicted cross-sectional view that is as fine as the first cross-sectional view. Subsequently, for the object to be measured, a rough three-dimensional model is constructed using photos of the object from different perspectives, and a second slicing process is performed on it to obtain a rough cross-sectional view. Finally, this rough cross-sectional view is input into the trained deep learning model, and the model outputs a predicted cross-sectional view. According to the predicted cross-sectional view, the complete three-dimensional point cloud data of the object to be measured can be reconstructed, thereby realizing the three-dimensional reconstruction of the object to be measured. In summary, the present invention uses a large amount of real data to train a deep learning model, which can automatically learn from photos and generate three-dimensional point cloud data, greatly reducing the reconstruction cost and improving the generalization ability of real-world data.
[0059] like Figure 2 As shown, Figure 2 A schematic flow chart of another embodiment of a 3D reconstruction method based on multi-view geometric constraints provided by the present invention. A 3D reconstruction method based on multi-view geometric constraints, the method comprising:
[0060] S201: Acquire actual three-dimensional digital models of a plurality of different objects to be measured, and perform a first slicing process on the actual three-dimensional digital models to obtain a first sectional view.
[0061] It should be noted that step S201 Figure 1 This has been discussed in detail in the implementation scenario shown and will not be repeated here.
[0062] S202: Acquire images of different viewing angles of the actual three-dimensional digital model of the object to be measured through a virtual camera.
[0063] S203: Extracting contour data of the object to be measured from images of the object to be measured at different viewing angles using an edge extraction algorithm.
[0064] For example, the actual three-dimensional digital model is placed in the virtual space, and virtual cameras (assuming there are 5 cameras) are arranged in the virtual space. Different virtual cameras shoot the object to be measured from 5 different angles, and these 5 perspectives should cover the whole picture of the object to be measured. The number of virtual cameras (the number of perspectives for shooting) is consistent with the number of cameras or the number of shooting angles of the system during actual measurement (shooting from different angles when shooting with one camera). After shooting, the edge extraction algorithm is used to extract the contour data of the object to be measured in the images of different perspectives of the object to be measured.
[0065] S204: constructing a rough three-dimensional model of the object to be measured according to the contour data of the object to be measured.
[0066] For example, based on the multi-view geometric constraint algorithm (the core of the algorithm is to use the intersection of camera cone space), a rough three-dimensional model of the object to be measured is constructed according to the contour data of the object to be measured. The premise of the execution of the multi-view geometric constraint algorithm is that the posture relationship of the camera under each viewing angle is known, and the posture relationship between two viewing angles is expressed by a rigid body transformation RT, that is, a rotation matrix R and a translation matrix T are multiplied.
[0067] S205: Perform a second slicing process on the rough three-dimensional model to obtain a second cross-sectional view; wherein the slicing density of the first slicing process is greater than the slicing density of the second slicing process.
[0068] For example, the rough three-dimensional model is cut into six faces along the Z axis to obtain six second section images, such as Figure 3 As shown, Figure 3 A second cross-sectional view provided by the present invention, Figure 3 What is shown is the final cross-section image, 6 two-dimensional cross-section images, which serve as the second cross-section image as the input data of the deep learning model.
[0069] S206: Using the first cross-sectional image and the second cross-sectional image as training data sets for the deep learning model.
[0070] S207: Construct a deep learning model, use the second cross-section image as input data of the deep learning model, and use the first cross-section image as output data of the deep learning model.
[0071] S208: Determine the value of the loss function according to the true label and the predicted label of the pixel in the first slice image.
[0072] S209: Optimize the parameters of the deep learning model according to the value of the loss function to determine the trained deep learning model.
[0073] It should be noted that steps S206-S209 are Figure 1 This has been discussed in detail in the implementation scenario shown and will not be repeated here.
[0074] S210: placing the object to be tested on the calibration target, photographing the object to be tested and the calibration target from different angles through a camera, and obtaining photographs of the object to be tested from different viewing angles.
[0075] For example, five cameras are arranged in the target area, and different cameras take pictures of the object to be tested, such as a human foot, from five different angles. The object to be tested is placed on the calibration target, and the cameras take pictures of the object to be tested and the calibration target from different angles. Pictures of the object to be tested from different perspectives are obtained.
[0076] It should be noted that the number of simulation cameras (the number of shooting angles) is consistent with the number of cameras or shooting angles in the system during actual measurement.
[0077] S211: Calibrate the intrinsic and extrinsic parameters of the camera at each viewing angle by calibrating the target, where the extrinsic parameters are the camera pose parameters at different viewing angles.
[0078] Exemplarily, the camera intrinsic parameters and camera extrinsic parameters at each viewing angle are calibrated on the calibration target. The camera extrinsic parameters are the camera pose parameters at different viewing angles, that is, the rigid body transformation RT. Specifically, the camera intrinsic parameters include but are not limited to:
[0079]
[0080] Among them, KK is the camera internal parameter, fx is the focal length parameter of the lens in the horizontal axis direction of the imaging plane, fy is the focal length parameter of the lens in the vertical axis direction of the imaging plane, cx is the horizontal coordinate of the center coordinate of the imaging plane, and cy is the vertical coordinate of the center coordinate of the imaging plane.
[0081] S212: Extracting contour data of the object to be tested from the photos of the object to be tested at different viewing angles using an edge extraction algorithm.
[0082] Exemplarily, the object to be tested in the photographs of the object to be tested at different viewing angles is segmented from the background, for example, the object to be tested is a human foot, and the contour data of the object to be tested in the photographs of the object to be tested at different viewing angles is extracted using an edge extraction algorithm.
[0083] S213: Based on the multi-view geometric constraint algorithm, a rough three-dimensional model of the object to be measured is constructed using the intrinsic and extrinsic parameters of the camera at each viewing angle and the contour data of the object to be measured.
[0084] S214: performing a second slicing process on the rough three-dimensional model to obtain a rough cross-sectional image of the object to be measured.
[0085] Exemplarily, based on the multi-view geometric constraint algorithm, a rough three-dimensional model of the object to be tested is constructed using the intrinsic and extrinsic parameters of the camera at each viewing angle and the contour data of the object to be tested. The rough three-dimensional model of the object to be tested is sliced along the Z axis according to 6 slices to generate 6 rough cross-sectional images of the object to be tested.
[0086] S215: Input the rough cross-sectional image into the trained deep learning model to obtain the predicted cross-sectional image output by the deep learning model, and reconstruct the three-dimensional point cloud data of the object to be measured based on the predicted cross-sectional image.
[0087] Exemplarily, a rough cross-section image is input into a trained deep learning model to obtain a predicted cross-section image output by the deep learning model; the rough cross-section image and the predicted cross-section image of the object to be measured are used as image data for accurate three-dimensional reconstruction, and the intrinsic and extrinsic parameters of the camera at each viewing angle are used to reversely project rays from the image data, and the coordinates of a series of intersection points are calculated by the intersection of rays from each viewpoint to reconstruct the three-dimensional point cloud data of the object to be measured.
[0088] S216: Reconstruct the three-dimensional point cloud data of the object to be measured into a mesh model according to the Poisson reconstruction method.
[0089] For example, in the Poisson reconstruction algorithm, an indicator function (value inside the model is 1, value outside the model is 0) is used to represent the implicit function. The problem is transformed into solving the Poisson equation by using the property that the gradient of the indicator function after smoothing is equal to the vector field obtained by the normal vector of the smooth surface. Finally, the improved marching cube algorithm is used to extract the isosurface. In addition, the three-dimensional point cloud data of the object to be measured is sampled equidistantly in the Z-axis direction. Due to the data noise, discretization processing and point cloud density approximation, its zero isosurface may deviate from the model surface. Therefore, random perturbations within the accuracy range are added to the point cloud before Poisson fusion to break the approximation of the point cloud density, thereby estimating a reasonable isosurface and completing the mesh reconstruction of the three-dimensional point cloud data.
[0090] From the above description, it can be seen that the present invention only relies on the camera on the smartphone and a calibration target calibrated with the camera intrinsic parameters and camera extrinsic parameters at various viewing angles. The user only needs to use the mobile phone to effectively and accurately reconstruct the complete three-dimensional point cloud data of the object to be measured, and then reconstruct the three-dimensional point cloud data of the object to be measured into a mesh model according to the Poisson reconstruction method. Its significant advantages lie in convenience, low cost and simple operation.
[0091] like Figure 4 As shown, Figure 4 A schematic diagram of a structure of an embodiment of a 3D reconstruction system based on multi-view geometric constraints provided by the present invention. A 3D reconstruction system 10 based on multi-view geometric constraints, the system comprising:
[0092] The deep learning model acquisition module 11 is used to obtain actual three-dimensional digital models of multiple different objects to be measured, perform a first slicing process on the actual three-dimensional digital models to obtain a first cross-sectional view; obtain images of the object to be measured from different perspectives based on the actual three-dimensional digital models, and construct a rough three-dimensional model of the object to be measured based on the images from different perspectives; perform a second slicing process on the rough three-dimensional model to obtain a second cross-sectional view; wherein the slicing density of the first slicing process is greater than the slicing density of the second slicing process; use the first cross-sectional view and the second cross-sectional view as training data sets for the deep learning model; construct a deep learning model, use the second cross-sectional view as input data of the deep learning model, use the first cross-sectional view as output data of the deep learning model, train the deep learning model, and obtain a trained deep learning model.
[0093] The rough cross-sectional image acquisition module 12 is used to acquire photos of the object to be tested at different viewing angles, construct a rough three-dimensional model of the object to be tested based on the photos of different viewing angles, and perform a second slicing process on the rough three-dimensional model to obtain a rough cross-sectional image of the object to be tested.
[0094] The three-dimensional model reconstruction module 13 is used to input the rough cross-section image into the trained deep learning model to obtain the predicted cross-section image output by the deep learning model, and reconstruct the three-dimensional point cloud data of the object to be measured based on the predicted cross-section image.
[0095] Exemplarily, in the deep learning model acquisition module 11, actual three-dimensional digital models of multiple different objects to be measured are obtained, and the actual three-dimensional digital models are subjected to a first slicing process to obtain a first cross-sectional view. Images of the actual three-dimensional digital model of the object to be measured from different perspectives are obtained through a virtual camera; the contour data of the object to be measured in the images of the object to be measured from different perspectives are extracted using an edge extraction algorithm; and a rough three-dimensional model of the object to be measured is constructed based on the contour data of the object to be measured. The rough three-dimensional model is subjected to a second slicing process to obtain a second cross-sectional view; wherein the slice density of the first slicing process is greater than the slice density of the second slicing process. The first cross-sectional view and the second cross-sectional view are used as training data sets for the deep learning model; a deep learning model is constructed, the second cross-sectional view is used as input data of the deep learning model, and the first cross-sectional view is used as output data of the deep learning model; the value of the loss function is determined based on the true label and predicted label of the pixel in the first cross-sectional view; the parameters of the deep learning model are optimized based on the value of the loss function to determine the trained deep learning model.
[0096] Furthermore, in the rough cross-sectional image acquisition module 12, the object to be measured is placed on the calibration target, and the object to be measured and the calibration target are photographed from different angles by a camera to obtain photos of the object to be measured from different perspectives; the intrinsic and extrinsic parameters of the camera at each perspective are calibrated by the calibration target, and the extrinsic parameters are the camera posture parameters at different perspectives; the edge extraction algorithm is used to extract the contour data of the object to be measured in the photos of the object to be measured from different perspectives; based on the multi-view geometric constraint algorithm, the intrinsic and extrinsic parameters of the camera at each perspective, as well as the contour data of the object to be measured are used to construct a rough three-dimensional model of the object to be measured; the rough three-dimensional model is subjected to a second slicing process to obtain a rough cross-sectional image of the object to be measured.
[0097] Finally, in the three-dimensional model reconstruction module 13, the rough cross-section image is input into the trained deep learning model to obtain the predicted cross-section image output by the deep learning model; and the three-dimensional point cloud data of the object to be measured is reconstructed based on the predicted cross-section image.
[0098] like Figure 5 As shown, Figure 5 The device 20 includes a memory 21 and a processor 22. The memory 21 stores a computer program, and the processor 22 executes the computer program when working to implement the following. Figure 1 and Figure 2 The method shown.
[0099] The specific technical details of a three-dimensional reconstruction method based on multi-view geometric constraints implemented when the above-mentioned device 20 executes a computer program have been discussed in detail in the aforementioned method steps, so they will not be repeated here.
[0100] like Figure 6 As shown, Figure 6 The structure diagram of an embodiment of the medium provided by the present invention is shown in FIG. The medium 30 stores at least one computer program 31, and the computer program 31 is executed by the processor 22 to implement the following Figure 1 and Figure 2 In one embodiment, the medium 30 may be a storage chip, a hard disk, a mobile hard disk, a USB flash drive, an optical disk, or other readable and writable storage tools, or a server, etc.
[0101] The above describes specific embodiments of the present specification, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily have to be performed in the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0102] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment, and non-volatile computer-readable storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0103] The apparatus, device, non-volatile computer-readable storage medium and method provided in the embodiments of this specification correspond to each other, and therefore, the apparatus, device, and non-volatile computer storage medium also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the corresponding apparatus, device, and non-volatile computer storage medium will not be repeated here.
[0104] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0105] For the convenience of description, the above device is described by being divided into various units according to their functions and described separately. Of course, when implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware. It should be understood by those skilled in the art that this specification embodiment can be provided as a method, system, or computer program product. Therefore, this specification embodiment can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification embodiment can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0106] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0107] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0108] The above disclosure is only the preferred embodiment of the present invention, which certainly cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.
Claims
1. A 3D reconstruction method based on multi-view geometric constraints, characterized in that: The method comprises: Acquire actual three-dimensional digital models of multiple different objects to be measured, perform a first slicing process on the actual three-dimensional digital models to obtain a first cross-sectional view; obtain images of the object to be measured at different perspectives based on the actual three-dimensional digital models, and construct a rough three-dimensional model of the object to be measured based on the images at different perspectives; perform a second slicing process on the rough three-dimensional model to obtain a second cross-sectional view; wherein the slicing density of the first slicing process is greater than the slicing density of the second slicing process; use the first cross-sectional view and the second cross-sectional view as training data sets for a deep learning model; construct a deep learning model, use the second cross-sectional view as input data of the deep learning model, use the first cross-sectional view as output data of the deep learning model, train the deep learning model, and obtain a trained deep learning model; Acquire photos of the object to be tested at different viewing angles, construct a rough three-dimensional model of the object to be tested based on the photos at different viewing angles, and perform the second slicing process on the rough three-dimensional model to obtain a rough cross-sectional image of the object to be tested; The rough cross-sectional image is input into a trained deep learning model to obtain a predicted cross-sectional image output by the deep learning model, and three-dimensional point cloud data of the object to be measured is reconstructed based on the predicted cross-sectional image.
2. The 3D reconstruction method based on multi-view geometric constraints according to claim 1, characterized in that: The step of obtaining images of the object under test at different viewing angles based on the actual three-dimensional digital model and constructing a rough three-dimensional model of the object under test based on the images at different viewing angles specifically includes: Acquire images of different viewing angles of the actual three-dimensional digital model of the object being measured through a virtual camera; The edge extraction algorithm is used to extract the contour data of the object under test from the images of different viewing angles of the object under test; A rough three-dimensional model of the object to be measured is constructed based on the contour data of the object to be measured.
3. The 3D reconstruction method based on multi-view geometric constraints according to claim 2, characterized in that: The step of constructing a deep learning model, using the second cross-sectional image as input data of the deep learning model, using the first cross-sectional image as output data of the deep learning model, and training the deep learning model to obtain a trained deep learning model specifically includes: Constructing a deep learning model, using the second cross-sectional image as input data of the deep learning model, and using the first cross-sectional image as output data of the deep learning model; Determine a value of a loss function based on the true label and the predicted label of the pixel in the first slice image; Parameters of the deep learning model are optimized according to the value of the loss function to determine a trained deep learning model.
4. The 3D reconstruction method based on multi-view geometric constraints according to claim 3, characterized in that: The determining the value of the loss function according to the true label and the predicted label of the pixel in the first slice image specifically includes: The loss function is: Among them, BC is the value of the loss function, N is the number of pixels in the first slice image, s is the number of slice images output by the network, and y i is the true label of the i-th pixel in the first slice image, is the predicted label of the i-th pixel in the first slice image.
5. The 3D reconstruction method based on multi-view geometric constraints according to claim 3, characterized in that: The obtaining of photos of the object to be tested at different viewing angles, constructing a rough three-dimensional model of the object to be tested based on the photos at different viewing angles, and performing the second slicing process on the rough three-dimensional model to obtain a rough cross-sectional view of the object to be tested specifically includes: Place the object to be tested on the calibration target, and use the camera to shoot the object to be tested and the calibration target from different angles to obtain photos of the object to be tested from different perspectives; The intrinsic and extrinsic parameters of the camera at each viewing angle are calibrated by the calibration target, wherein the extrinsic parameters are the camera pose parameters at different viewing angles; Use edge extraction algorithm to extract the contour data of the object to be tested from the photos of the object to be tested at different viewing angles; Based on a multi-view geometric constraint algorithm, a rough three-dimensional model of the object to be measured is constructed using the intrinsic and extrinsic parameters of the camera at each viewing angle and the contour data of the object to be measured; The second slicing process is performed on the rough three-dimensional model to obtain a rough cross-sectional image of the object to be tested.
6. The 3D reconstruction method based on multi-view geometric constraints according to claim 5, characterized in that: After inputting the rough section image into the trained deep learning model to obtain the predicted section image output by the deep learning model, and reconstructing the three-dimensional point cloud data of the object to be measured based on the predicted section image, the method specifically includes: The three-dimensional point cloud data of the object to be measured is reconstructed into a grid model according to the Poisson reconstruction method.
7. A 3D reconstruction system based on multi-view geometric constraints, characterized in that: The system comprises: A deep learning model acquisition module is used to acquire actual three-dimensional digital models of multiple different objects to be measured, perform a first slicing process on the actual three-dimensional digital models to obtain a first cross-sectional view; obtain images of the object to be measured at different perspectives based on the actual three-dimensional digital models, and construct a rough three-dimensional model of the object to be measured based on the images at different perspectives; perform a second slicing process on the rough three-dimensional model to obtain a second cross-sectional view; wherein the slicing density of the first slicing process is greater than the slicing density of the second slicing process; use the first cross-sectional view and the second cross-sectional view as training data sets for a deep learning model; construct a deep learning model, use the second cross-sectional view as input data of the deep learning model, use the first cross-sectional view as output data of the deep learning model, train the deep learning model, and obtain a trained deep learning model; a rough cross-sectional image acquisition module, used to acquire photos of the object to be tested from different perspectives, construct a rough three-dimensional model of the object to be tested based on the photos from different perspectives, and perform the second slicing process on the rough three-dimensional model to obtain a rough cross-sectional image of the object to be tested; The three-dimensional model reconstruction module is used to input the rough section image into the trained deep learning model to obtain the predicted section image output by the deep learning model, and reconstruct the three-dimensional point cloud data of the object to be measured based on the predicted section image.
8. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 6.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 6.