A three-dimensional scanning method, device, apparatus and storage medium

By establishing a graphical structure and texture map points, the problems of sparse reconstruction errors, long time consumption, and redundant calculations in existing 3D mapping methods are solved, achieving fast and accurate 3D model mapping and improving user experience.

CN116188664BActive Publication Date: 2026-07-21SHINING 3D TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHINING 3D TECH CO LTD
Filing Date
2022-12-22
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing 3D mapping methods suffer from sparse reconstruction errors, long processing times, excessive redundant calculations, and poor adaptability, resulting in a poor user experience.

Method used

By acquiring high-quality, multi-view, multi-frame images, a graphic structure is established, and texture map points on the surface of the 3D model are constructed. Based on the graphic structure and dataset, the pose of each frame of the captured image is determined for texture processing, reducing redundant calculations and improving texture efficiency and accuracy.

Benefits of technology

It achieves fast and accurate 3D model texturing, reduces the waste of manpower and resources, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116188664B_ABST
    Figure CN116188664B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a three-dimensional scanning method, device, equipment and storage medium. The scanning method comprises: acquiring a first image set, a second image set and a three-dimensional model of a scanned object, the first image set comprising a plurality of frames of shooting images, and the second image set comprising a plurality of frames of scanning images; establishing a graph structure between target frames of the plurality of frames of shooting images; constructing a data set of a surface of the three-dimensional model based on three-dimensional points in a three-dimensional model coordinate system and the plurality of frames of scanning images; and generating texture data of the three-dimensional model according to the graph structure and the data set. The method provided by the present disclosure can quickly and accurately complete scanning model mapping processing, thereby improving user experience.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Technology Neighborhood

[0002] This disclosure relates to the field of computer technology, and more particularly to a three-dimensional scanning method, apparatus, device, and storage medium. Background Technology

[0003] With the development of computer hardware and industrial scanning technology, three-dimensional (3D) texture mapping is widely used in various fields of computer graphics to increase the realism of 3D objects. 3D texture mapping is the fusion of high-quality multi-view images captured by camera devices and scanned models to improve the realism of the scanned objects displayed.

[0004] However, existing methods that use SFM (Structure From Motion) algorithms for sparse reconstruction followed by texturing are prone to sparse reconstruction errors and are time-consuming. Methods that use feature matching followed by texturing involve a large amount of redundant computation, and the matching process heavily relies on calibration results, resulting in poor adaptability. In summary, current 3D texturing methods have relatively low texturing efficiency and a poor user experience. Summary of the Invention

[0005] To address the aforementioned technical issues, this disclosure provides a 3D scanning method, apparatus, device, and storage medium that can quickly and accurately complete the texturing process of the scanned model, thereby improving the user experience.

[0006] In a first aspect, embodiments of this disclosure provide a three-dimensional scanning method, including:

[0007] Acquire a first image set, a second image set, and a 3D model of the scanned object. The first image set includes multiple captured images, and the second image set includes multiple scanned images.

[0008] Establish the graphic structure between the target frame images in the multi-frame captured images;

[0009] Based on the three-dimensional points in the coordinate system of the three-dimensional model and the multi-frame scan images, a dataset of the surface of the three-dimensional model is constructed.

[0010] Based on the graphic structure and the dataset, the texture data of the 3D model is generated.

[0011] Optionally, establishing the graphic structure between the target frame images in the multi-frame captured images includes:

[0012] Feature matching is performed based on the extracted features of the multi-frame captured images to obtain the feature matching relationship between the multi-frame captured images;

[0013] Based on the feature matching relationship between the multi-frame captured images, a graphical structure is established between the target frame captured images in the multi-frame captured images.

[0014] Optionally, constructing the dataset of the three-dimensional model surface based on the three-dimensional points in the three-dimensional model coordinate system and the multi-frame scan images includes:

[0015] Features of the multi-frame scanned images are extracted, wherein the features of each frame scanned image include multiple two-dimensional points and a feature descriptor corresponding to each two-dimensional point;

[0016] Based on the features of the multi-frame scan images, determine the feature descriptor corresponding to each three-dimensional point of the three-dimensional model;

[0017] Using the three-dimensional model as a framework, a dataset of the surface of the three-dimensional model is constructed based on the three-dimensional points in the coordinate system of the three-dimensional model and the feature descriptor corresponding to each three-dimensional point.

[0018] Optionally, generating the texture data of the 3D model based on the graphic structure and the dataset includes:

[0019] Based on the aforementioned graphic structure, the first frame image is determined from the target frame captured image;

[0020] The first pose of the first frame image is calculated based on the dataset, and the depth map data of the first frame image in the three-dimensional model coordinate system is determined.

[0021] Based on the first pose and the depth map data, the first frame image is textured, and after all target frame images are textured, the texture data of the three-dimensional model is generated.

[0022] Optionally, the step of calculating the first pose of the first frame image based on the dataset and determining the depth map data of the first frame image in the three-dimensional model coordinate system includes:

[0023] The first frame image and the dataset are subjected to feature matching to obtain a first matching pair with a first matching relationship, wherein the first matching relationship reflects the matching relationship between two-dimensional points and three-dimensional points;

[0024] The first pose of the first frame image is calculated based on the first matching pair, and the camera parameters of the camera device that captured the first frame image are determined.

[0025] The depth map data of the first frame captured image in the three-dimensional model coordinate system is determined based on the camera parameters.

[0026] Optionally, determining the camera parameters of the imaging device for capturing the first frame image includes:

[0027] The first matching pair is expanded based on the first pose and the feature descriptor corresponding to each 3D point in the dataset;

[0028] The camera parameters of the imaging device for capturing the first frame image are determined based on the expanded first matching pair.

[0029] Optionally, after determining the depth map data of the first frame captured image in the three-dimensional model coordinate system, the method further includes:

[0030] Based on the graphic structure, a second frame image that matches the first frame image is determined in the target frame image;

[0031] The first frame image and the second frame image are subjected to feature matching to obtain a second matching pair with the first matching relationship;

[0032] The second pose of the second frame image is calculated based on the second matching pair, and the camera parameters are jointly optimized based on the second matching pair and the first matching pair.

[0033] The depth map data of the second frame image in the three-dimensional model coordinate system is determined based on the optimized camera parameters.

[0034] Optionally, the step of performing feature matching between the first frame image and the second frame image to obtain a second matching pair having the first matching relationship includes:

[0035] The first frame image and the second frame image are subjected to feature matching to obtain a third matching pair with a second matching relationship, wherein the second matching relationship reflects the matching relationship between two-dimensional points;

[0036] Based on the depth map data of the first frame image and the third matching pair, a second matching pair with the first matching relationship is obtained.

[0037] Secondly, embodiments of this disclosure provide a three-dimensional scanning device, comprising:

[0038] The acquisition module is used to acquire a first image set, a second image set, and a 3D model of the scanned object. The first image set includes multiple captured images, and the second image set includes multiple scanned images.

[0039] The module is used to establish the graphic structure between the target frame images in the multi-frame captured images;

[0040] A construction module is used to construct a dataset of the surface of the three-dimensional model based on the three-dimensional points in the coordinate system of the three-dimensional model and the multi-frame scan images;

[0041] The generation module is used to generate texture data of the three-dimensional model based on the graphic structure and the dataset.

[0042] Thirdly, embodiments of this disclosure provide an electronic device, including:

[0043] Memory;

[0044] Processor; and

[0045] Computer programs;

[0046] The computer program is stored in the memory and configured to be executed by the processor to implement the three-dimensional scanning method described above.

[0047] Fourthly, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the three-dimensional scanning method described above.

[0048] This disclosure provides a 3D scanning method, comprising: acquiring a first image set, a second image set, and a 3D model of a scanned object; the first image set including multiple captured images, and the second image set including multiple scanned images; establishing a graphical structure between the target captured images in the multiple captured images; constructing a dataset of the 3D model surface based on 3D points in the 3D model coordinate system and the multiple scanned images; and generating texture data of the 3D model according to the graphical structure and the dataset. The method provided by this disclosure can quickly and accurately complete the texturing processing of the scanned model, improving the user experience. Attached Figure Description

[0049] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0050] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 A flowchart illustrating a three-dimensional scanning method provided in an embodiment of this disclosure;

[0052] Figure 2 A flowchart illustrating another three-dimensional scanning method provided in this embodiment of the disclosure;

[0053] Figure 3 This is a schematic diagram of the structure of a three-dimensional scanning device provided in an embodiment of the present disclosure;

[0054] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0055] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0056] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0057] Currently, 3D texture mapping is widely used in various fields to enhance the realism of 3D models. One feasible existing technical solution involves inputting high-quality multi-view images of the scanned object, offline reconstructing the 3D data of the scanned object using the SFM algorithm, and then registering the offline reconstructed data with the model data using a scaled 3D point cloud registration algorithm. This yields the correspondence between the high-quality multi-view images and the scanned model, completing the texture mapping algorithm process. However, this sparse reconstruction method based on the SFM algorithm is time-consuming and prone to sparse reconstruction errors on scanned models with limited features or repetitive features, leading to reconstruction failure. After a failure, high-quality multi-view images must be reacquired, introducing uncertainty and wasting manpower and resources through repeated acquisition. Another feasible existing technical solution involves pre-calibrating the acquisition device for high-quality multi-view images to obtain the camera intrinsic parameters of the acquisition device. Subsequently, the image features of the high-quality multi-view images and the scanned images are calculated. The correspondence between 2D-2D points (two-dimensional points and two-dimensional points) between frames is constructed through feature matching. Since the scanned images obtained by the scanning device are RGBD data, the correspondence between 2D-3D points (two-dimensional points and three-dimensional points) can be constructed. Then, the camera extrinsic parameters of each high-quality multi-view image are calculated frame by frame using the Ransac (Random Sample Consensus) algorithm combined with the pose solving algorithm (Perspective N Point, PNP). However, this feature matching strategy based on the Ransac and PNP algorithms processes each frame as a unit, performing feature matching between all frames of high-quality multi-view images and all frames of scanned images one by one. However, a significant proportion of features overlap between the two types of images, leading to a large amount of redundant computation during feature matching, resulting in low robustness and a high computational load. Furthermore, because the computation relies heavily on a calibration framework—that is, it depends heavily on pre-calibrated results—the algorithm has poor adaptability to scanned models of different sizes, and the repeated calibration process easily leads to a waste of human and material resources. In summary, existing 3D texturing methods involve a large number of repetitive operations, easily resulting in a waste of human and material resources, and the texturing efficiency is relatively low, leading to a poor user experience.

[0058] To address the aforementioned technical problems, this disclosure provides a scanning method. It establishes a graphical structure reflecting effective feature matching relationships using high-quality, multi-view, multi-frame captured images. A texture map (dataset) is constructed from the acquired multi-frame scanned images, including 3D points on the surface of the scanned model and corresponding feature descriptors. Subsequently, the pose of each captured image is determined based on the graphical structure and the dataset, and texturing is performed according to the pose to obtain the texture data of the 3D model. This facilitates the subsequent display of a realistic 3D model to the user. The method provided by this disclosure, without requiring the user to repeatedly calibrate camera intrinsic parameters, achieves a relatively accurate algorithm logic of scanning first and then texturing through simple operation logic. It outperforms existing texturing methods in terms of accuracy, computational efficiency, and robustness to texture features of different captured images, reducing the waste of manpower and resources and improving the user experience. A detailed description is provided through at least one of the following embodiments.

[0059] Figure 1 This is a flowchart illustrating a three-dimensional scanning method provided in an embodiment of this disclosure. Applied to a terminal or server, the following embodiment uses a server executing the three-dimensional scanning method as an example. The server acquires a first image set, a second image set, and a three-dimensional model (scanning model) of the object to be scanned. The second image set and the three-dimensional model can be sent by the scanning device, and the three-dimensional model can also be obtained by the server performing three-dimensional reconstruction based on the second image set. The first image set can be a sequence of images acquired by a high-quality image acquisition device. In one feasible implementation, after the scanning device captures the second image set and constructs the three-dimensional model, it directly transmits the second image set and the three-dimensional model to the server. The acquisition device also transmits the acquired first image set to the server. Subsequently, the server executes the three-dimensional scanning method based on the first image set, the second image set, and the three-dimensional model to obtain the texture data of the three-dimensional model. In another feasible implementation, the scanning device transmits the second image set to the server, the server performs three-dimensional reconstruction based on the second image set to obtain the three-dimensional model, and then the server executes the three-dimensional scanning method based on the first image set, the second image set, and the three-dimensional model. Other possible embodiments of this disclosure are not limited here.

[0060] The three-dimensional scanning method specifically includes, for example: Figure 1 The following steps S110 to S140 are shown:

[0061] S110. Obtain the first image set, the second image set, and the 3D model of the scanned object.

[0062] The first image set includes multiple captured images, and the second image set includes multiple scanned images.

[0063] Understandably, the process involves acquiring a first image set, a second image set, and a 3D model of the scanned object. The first image set can be a set of multiple frames captured by a high-quality image acquisition device, denoted as Img. H High-quality image acquisition equipment can be a DSLR camera. The second image set is a set of multiple scanned images acquired by the scanning device, denoted as Img. L The scanning device can possess 3D reconstruction capabilities, performing real-time 3D reconstruction during image acquisition to generate a 3D model of the scanned object. Alternatively, it can perform 3D reconstruction after all scanned images have been acquired. It can also directly transmit a second image set, which the server then uses to perform 3D reconstruction to generate the 3D model of the scanned object. The specific 3D reconstruction method is not limited. It is understood that the acquisition order and timing of the first and second image sets are not limited and can be determined according to user needs.

[0064] S120. Establish the graphic structure between the target frame images in the multi-frame captured images.

[0065] Understandably, based on the above S110, after obtaining the first image set, a graph structure is established between the target frame images in the multi-frame captured images. The number of target frame images is less than the number of multi-frame captured images, that is, there will be multiple target frame captured images. The established graph structure can reflect the feature matching relationship between multiple target frame captured images.

[0066] Optionally, the graphical structure established in S120 above is implemented through the following steps:

[0067] Feature matching is performed based on the extracted features of the multi-frame captured images to obtain the feature matching relationship between the multi-frame captured images.

[0068] Based on the feature matching relationship between the multi-frame captured images, a graphical structure is established between the target frame captured images in the multi-frame captured images.

[0069] Understandably, a scale-invariant feature transform (SIFT) is performed on each frame of the multi-frame image to extract texture features, resulting in the features of each frame. Specifically, these features can be local features of each frame. After feature extraction, feature matching is performed pairwise on each of the multi-frame images to obtain the feature matching relationships between them. Subsequently, a graphical structure is established between the target frame images based on the feature matching relationships between the multi-frame images. This graphical structure reflects the feature information of the target frame images with effective feature matching relationships.

[0070] S130. Based on the three-dimensional points in the coordinate system of the three-dimensional model and the multi-frame scan images, construct a dataset of the surface of the three-dimensional model.

[0071] Understandably, based on the above S110, the three-dimensional points (3D data) in the three-dimensional model coordinate system are determined, and a dataset of the three-dimensional model surface is constructed based on the three-dimensional points on the three-dimensional model surface and multiple frame scan images. This dataset can be understood as texture map points (MapPoint). The texture map points contain the three-dimensional points in the three-dimensional model coordinate system and the corresponding image SIFT feature descriptors of the three-dimensional points.

[0072] Optionally, the dataset for constructing the 3D model surface in S130 above can be implemented through the following steps:

[0073] Features are extracted from the multi-frame scan images, wherein the features of each frame scan image include multiple two-dimensional points and a feature descriptor corresponding to each two-dimensional point.

[0074] Based on the features of the multi-frame scanned images, the feature descriptor corresponding to each three-dimensional point of the three-dimensional model is determined.

[0075] Using the three-dimensional model as a framework, a dataset of the surface of the three-dimensional model is constructed based on the three-dimensional points in the coordinate system of the three-dimensional model and the feature descriptor corresponding to each three-dimensional point.

[0076] Understandably, SIFT feature extraction is performed on multiple frames of scanned images to obtain local features for each frame. Each frame's local features include multiple two-dimensional points (2D data) and a corresponding SIFT feature descriptor for each point. Subsequently, based on the local features of each frame, the feature descriptor for each 3D point on the 3D model surface is determined. Each 3D point has a fixed and unique identifier within the 3D model. Using these 3D points on the 3D model surface as a framework, texture map points are defined for the surface. These texture map points include 3D points in the 3D model coordinate system and a corresponding feature descriptor for each point.

[0077] S140. Generate texture data for the three-dimensional model based on the graphic structure and the dataset.

[0078] Understandably, based on the above S120 and S130, the pose and corresponding camera parameters of each frame of the multiple target frame images are calculated according to the graph structure and dataset. The 3D model is textured based on the pose of each frame of the captured images. After the textures of all target frame images are textured, global optimization (Bundle Adjustment, BA) and texture fusion operations are performed to obtain the texture data of the 3D model.

[0079] Understandably, after obtaining the texture data of the 3D model, this texture data is displayed to the user. The texture data of the 3D model refers to fusing high-quality, multi-view captured images onto the 3D model to present the user with a realistic 3D model of the scanned object. Subsequently, after all target frame images have been textured, the user can adjust the pose of each frame image based on the displayed texture results.

[0080] Understandably, after determining the pose of each captured image frame and completing the texturing, the texturing result of that captured image frame (current frame) can be displayed. That is, the texturing of each captured image frame is performed after calculating the pose of each frame, and the texturing result of that frame is displayed in real time so that the user can make adjustments.

[0081] This disclosure provides a 3D scanning method. Based on target frame images with effective feature matching relationships acquired in a first image set, a graphic structure is established. This structure reduces redundant data, further reducing the computational load and time for subsequent feature matching. After SIFT feature extraction from multiple frames, a texture map is constructed based on 3D points on the surface of a 3D model and multiple scanned images with different poses. This map includes 3D points and their corresponding feature descriptors. By establishing a matching relationship between the 3D model and the scanned images, the method facilitates rapid and accurate determination of the correspondence between the captured images and the 3D points on the surface of the 3D model when matching the captured images with the texture map points, thus accelerating the image capture process. The system improves image mapping efficiency. Finally, it automatically determines the camera parameters corresponding to each target frame image based on the graphic structure and texture map points, and continuously optimizes the camera parameters. This eliminates the need for users to repeatedly calibrate the camera intrinsic parameters, reducing repetitive operations. It can also perform single-frame mapping based on the pose of each target frame image in the 3D model coordinate system. After all target frames have completed mapping, it performs global BA optimization and texture fusion to obtain the texture data of the 3D model. This facilitates the subsequent presentation of a realistic 3D model to users. The entire process of obtaining texture data is highly feasible, has a relatively high accuracy, and is robust to images with different texture features, further improving the user experience.

[0082] Based on the above embodiments, Figure 2 This is a flowchart illustrating a three-dimensional scanning method provided in an embodiment of the present disclosure. Optionally, generating texture data of the three-dimensional model based on the graphic structure and the dataset specifically includes, for example: Figure 2 The following steps S210 to S230 are shown:

[0083] S210. Based on the graphic structure, determine the first frame image in the target frame image.

[0084] Understandably, the graph structure reflects the texture feature dataset among target frame images with effective feature matching relationships. The SIFT feature extraction algorithm extracts the texture features of the captured images. At least one target frame image with a texture feature count greater than a first threshold in the graph structure is selected as the first frame image. For example, the target frame image with the most texture features in the graph structure can be selected as the first frame image. In this case, only one first frame image needs to be determined, and the number of first frame images is not limited and can be determined according to user needs. "Most texture features" refers to having the most effective matching points. The first frame image can be considered the initial frame.

[0085] S220. Calculate the first pose of the first frame captured image based on the dataset, and determine the depth map data of the first frame captured image in the three-dimensional model coordinate system.

[0086] Understandably, based on the above S210, after determining the initial frame, the first pose of the first frame image is calculated according to the dataset. The first pose can be understood as the initial value. After calculating the first pose, the depth map data of the first frame image in the three-dimensional model coordinate system is determined. Understandably, the original first frame image is an RGB image, and the depth map data is the depth map corresponding to the first frame image. The first frame image with the obtained depth map can be understood as an RGBD image. That is to say, the two-dimensional points included in the first frame image can be regarded as three-dimensional points by adding the depth map data, and the correspondence between the shooting and the three-dimensional model is established based on this.

[0087] Optionally, the determination of the first pose and depth map data of the first frame of the captured image is achieved through the following steps:

[0088] The first frame image and the dataset are subjected to feature matching to obtain a first matching pair with a first matching relationship, wherein the first matching relationship reflects the matching relationship between two-dimensional points and three-dimensional points.

[0089] The first pose of the first frame image is calculated based on the first matching pair, and the camera parameters of the camera device that captured the first frame image are determined.

[0090] The depth map data of the first frame captured image in the three-dimensional model coordinate system is determined based on the camera parameters.

[0091] Understandably, there are several methods for determining the first pose of the first frame image: The first method involves registering the SIFT features of the first frame image with image features in the dataset to obtain first matching pairs with a first matching relationship. This first matching relationship refers to the matching relationship between two-dimensional and three-dimensional points. Each first matching pair consists of multiple valid 2D-3D in-line point pairs. Since the image feature points in the first frame image are two-dimensional, and the image feature points in the dataset are three-dimensional, multiple first matching pairs with a first matching relationship can be obtained. Subsequently, the Ransac and PNP pose determination algorithms are used to calculate the first pose of the first frame image based on these first matching pairs. The second method involves manually selecting multiple first matching pairs with a 2D-3D matching relationship, and then using the Ransac and PNP pose determination algorithms to calculate the first pose of the first frame image based on these first matching pairs. After determining the first pose, the camera intrinsics are initialized using the Exchangeable Image File Format (Exif) of any frame from the multi-frame image capture. The Exif information includes the camera's (high-quality image acquisition device) attribute information and capture data. Then, the camera parameters of the imaging device capturing the first frame are updated according to the first matching pair. These camera parameters include intrinsic and extrinsic parameters. Finally, the depth map data of the first frame in the 3D model coordinate system is updated using the updated camera parameters.

[0092] Optionally, the camera parameters of the imaging device for capturing the first frame of the image are determined, specifically through the following steps:

[0093] The first matching pair is expanded based on the first pose and the feature descriptor corresponding to each 3D point in the dataset.

[0094] The camera parameters of the imaging device for capturing the first frame image are determined based on the expanded first matching pair.

[0095] Understandably, after obtaining the first pose, each 3D point in the dataset is back-projected based on the camera intrinsics and extrinsic parameters from the first pose, i.e., a 3D-to-2D back-projection is performed to obtain the back-projection result. Then, the first matching pair is expanded based on the feature descriptors corresponding to the 3D points and the back-projection result, i.e., more 2D-to-3D point pairs are added. The camera intrinsics and extrinsic parameters of the first frame image are optimized using a local BA optimization algorithm based on the expanded first matching pairs. Simultaneously, invalid matching pairs are filtered out from the expanded first matching pairs; invalid matching pairs are those whose back-projection residuals are greater than a second threshold. The camera intrinsics and extrinsic parameters of the first frame image are then updated based on the filtered first matching pairs. Finally, the depth map data of the first frame image in the 3D model coordinate system is calculated based on the updated camera intrinsics and extrinsic parameters of the first frame image. For example, the dataset contains 1000 data points. By registering the SIFT features of the first frame image with the 1000 data points in the dataset, 100 first matching pairs are obtained. Then, an augmentation operation is performed to obtain 300 first matching pairs. Filtering the 300 first matching pairs yields 200 valid first matching pairs, thereby improving the accuracy of calculating the camera parameters of the first frame image.

[0096] S230. Based on the first pose and the depth map data, the first frame image is textured, and after all target frame images are textured, the texture data of the three-dimensional model is generated.

[0097] Understandably, based on the above S220, the first frame of the captured image is textured according to the first pose and depth map data. This process continues until all target frames of the captured image are textured, generating texture data for the 3D model. This facilitates the subsequent display of a realistic 3D model to the user based on the texture data. Understandably, texture processing can be performed on each frame of the captured image after calculating the pose and depth map data, or texture processing can be performed uniformly after determining the pose and depth map data of all target frames of the captured image. The specific texture method is not limited here and can be determined according to the user's needs.

[0098] Optionally, after determining the depth map data of the first frame captured image in the three-dimensional model coordinate system, the method further includes:

[0099] Based on the graphic structure, a second frame image that matches the first frame image is determined in the target frame image.

[0100] The first frame image and the second frame image are subjected to feature matching to obtain a second matching pair that has the first matching relationship.

[0101] The second pose of the second frame image is calculated based on the second matching pair, and the camera parameters are jointly optimized based on the second matching pair and the first matching pair.

[0102] The depth map data of the second frame image in the three-dimensional model coordinate system is determined based on the optimized camera parameters.

[0103] Understandably, based on the graph structure, the second frame image, which has the most feature matching points with the first frame image among multiple target frame images, is determined. The second frame image can also be understood as the secondary matching frame. The feature points in the second frame image are two-dimensional points, while the feature points in the first frame image can be considered three-dimensional points after being calculated into the depth map data. Therefore, feature matching between the first and second frames images yields second matching pairs with a first matching relationship, resulting in multiple second matching pairs with a 3D-2D matching relationship. Then, the second pose of the second frame image is calculated based on the second matching pairs using the Ransac and PNP pose solving algorithms. Subsequently, the camera intrinsic and extrinsic parameters are jointly optimized using an incremental BA optimization algorithm based on the first and second matching pairs, i.e., the camera parameters are jointly optimized using the first and second frames images. Finally, the depth map data of the second frame image in the three-dimensional model coordinate system is determined based on the jointly optimized camera parameters, i.e., the two-dimensional feature points in the second frame image are mapped to three-dimensional feature points.

[0104] Understandably, if the initial value calculation of the target frame image fails, the user can be prompted to manually determine the first matching pair and recalculate the pose of the target frame image to improve the robustness of the algorithm to 3D models of different sizes.

[0105] Optionally, the above process yields a second matching pair with a first matching relationship, specifically achieved through the following steps:

[0106] The first frame image and the second frame image are subjected to feature matching to obtain a third matching pair with a second matching relationship, wherein the second matching relationship reflects the matching relationship between two-dimensional points.

[0107] Based on the depth map data of the first frame image and the third matching pair, a second matching pair with the first matching relationship is obtained.

[0108] Understandably, matching the two-dimensional features in the first and second captured images yields a third matching pair with a second matching relationship. This second matching relationship refers to a 2D-to-2D feature matching relationship. Based on the depth map data of the first captured image and the third matching pair, a second matching pair with a first matching relationship is obtained. This first matching relationship reflects a 3D-to-2D feature matching relationship. In other words, the two-dimensional matching relationship between the first and second captured images is first determined, and then mapped to a three-dimensional matching relationship. Alternatively, the third matching pair with a second matching relationship can be determined without specifying it. Instead, the 2D-to-2D feature matching relationship between the initial frame and the next matched frame can be directly mapped to a 3D-to-2D feature matching relationship to obtain a second matching pair with a first matching relationship.

[0109] Understandably, after obtaining the second matching pair with the first matching relationship, the expansion and filtering methods described above for the first matching pair can be used to expand and filter the second matching pair. Subsequently, the camera parameters are jointly optimized based on the expanded and filtered first matching pair and the similarly expanded and filtered second matching pair.

[0110] Understandably, the third frame image with the most feature matching points of the second frame image can also be determined based on the graph structure. The method for calculating the pose and depth map data of the third frame image is the same as the method for determining the second frame image, which will not be elaborated here, until the pose and depth map data of all target frame images are determined.

[0111] The three-dimensional scanning method provided in this disclosure solves the pose by combining the texture map points constructed from multiple scanned images and three-dimensional models with the graphic structure of high-quality multi-view images. This reduces redundant computation and speeds up the overall computation time. Furthermore, it eliminates the need for additional sparse reconstruction. In conjunction with the algorithm logic for manually correcting the pose, the mapping accuracy and robustness of the target frame captured image are higher.

[0112] Figure 3 This is a schematic diagram of a three-dimensional scanning device provided in an embodiment of the present disclosure. The three-dimensional scanning device provided in this embodiment can execute the processing flow provided in the above-described three-dimensional scanning method embodiments, such as... Figure 3 As shown, the scanning device 300 includes an acquisition module 310, an establishment module 320, a construction module 330, and a generation module 340, wherein:

[0113] The acquisition module 310 is used to acquire a first image set, a second image set, and a three-dimensional model of the scanned object. The first image set includes multiple captured images, and the second image set includes multiple scanned images.

[0114] The module 320 is used to establish the graphic structure between the target frame images in the multi-frame captured images;

[0115] The construction module 330 is used to construct a dataset of the surface of the three-dimensional model based on the three-dimensional points in the coordinate system of the three-dimensional model and the multi-frame scan images;

[0116] The generation module 340 is used to generate texture data of the three-dimensional model based on the graphic structure and the dataset.

[0117] Optionally, module 320 is used for:

[0118] Feature matching is performed based on the extracted features of the multi-frame captured images to obtain the feature matching relationship between the multi-frame captured images;

[0119] Based on the feature matching relationship between the multi-frame captured images, a graphical structure is established between the target frame captured images in the multi-frame captured images.

[0120] Optionally, building module 330 is used for:

[0121] Features of the multi-frame scanned images are extracted, wherein the features of each frame scanned image include multiple two-dimensional points and a feature descriptor corresponding to each two-dimensional point;

[0122] Based on the features of the multi-frame scan images, determine the feature descriptor corresponding to each three-dimensional point of the three-dimensional model;

[0123] Using the three-dimensional model as a framework, a dataset of the surface of the three-dimensional model is constructed based on the three-dimensional points in the coordinate system of the three-dimensional model and the feature descriptor corresponding to each three-dimensional point.

[0124] Optionally, the generation module 340 is used for:

[0125] Based on the aforementioned graphic structure, the first frame image is determined from the target frame captured image;

[0126] The first pose of the first frame image is calculated based on the dataset, and the depth map data of the first frame image in the three-dimensional model coordinate system is determined.

[0127] Based on the first pose and the depth map data, the first frame image is textured, and after all target frame images are textured, the texture data of the three-dimensional model is generated.

[0128] Optionally, the generation module 340 is used for:

[0129] The first frame image and the dataset are subjected to feature matching to obtain a first matching pair with a first matching relationship, wherein the first matching relationship reflects the matching relationship between two-dimensional points and three-dimensional points;

[0130] The first pose of the first frame image is calculated based on the first matching pair, and the camera parameters of the camera device that captured the first frame image are determined.

[0131] The depth map data of the first frame captured image in the three-dimensional model coordinate system is determined based on the camera parameters.

[0132] Optionally, the generation module 340 is used for:

[0133] The first matching pair is expanded based on the first pose and the feature descriptor corresponding to each 3D point in the dataset;

[0134] The camera parameters of the imaging device for capturing the first frame image are determined based on the expanded first matching pair.

[0135] Optionally, the generation module 340 is used for:

[0136] Based on the graphic structure, a second frame image that matches the first frame image is determined in the target frame image;

[0137] The first frame image and the second frame image are subjected to feature matching to obtain a second matching pair with the first matching relationship;

[0138] The second pose of the second frame image is calculated based on the second matching pair, and the camera parameters are jointly optimized based on the second matching pair and the first matching pair.

[0139] The depth map data of the second frame image in the three-dimensional model coordinate system is determined based on the optimized camera parameters.

[0140] Optionally, the generation module 340 is used for:

[0141] The first frame image and the second frame image are subjected to feature matching to obtain a third matching pair with a second matching relationship, wherein the second matching relationship reflects the matching relationship between two-dimensional points;

[0142] Based on the depth map data of the first frame image and the third matching pair, a second matching pair with the first matching relationship is obtained.

[0143] Figure 3 The scanning device shown in the embodiment can be used to execute the technical solution of the above method embodiment. Its implementation principle and technical effect are similar, and will not be repeated here.

[0144] Figure 4 This is a schematic diagram of an electronic device provided in an embodiment of the present disclosure. The electronic device provided in this embodiment can execute the processing flow provided in the above embodiments, such as... Figure 4 As shown, the electronic device 400 includes a processor 410, a communication interface 420, and a memory 430; wherein, the computer program is stored in the memory 430 and configured to be executed by the processor 410 as described above in the three-dimensional scanning method.

[0145] In addition, this disclosure also provides a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the three-dimensional scanning method described in the above embodiments.

[0146] Furthermore, this disclosure also provides a computer program product, which includes a computer program or instructions that, when executed by a processor, implement the three-dimensional scanning method described above.

[0147] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0148] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A three-dimensional scanning method, characterized in that, The method includes: Acquire a first image set, a second image set, and a 3D model of the scanned object. The first image set includes multiple captured images, and the second image set includes multiple scanned images. Establish the graphic structure between the target frame images in the multi-frame captured images; Based on the 3D points in the 3D model coordinate system and the multi-frame scan images, a dataset of the 3D model surface is constructed, including: extracting features from the multi-frame scan images, wherein the features of each scan image include multiple 2D points and a feature descriptor corresponding to each 2D point; determining the feature descriptor corresponding to each 3D point of the 3D model based on the features of the multi-frame scan images; and constructing the dataset of the 3D model surface based on the 3D points in the 3D model coordinate system and the feature descriptor corresponding to each 3D point, using the 3D model as a framework. Generating texture data for the 3D model based on the graphical structure and the dataset includes: determining a first frame image in the target frame captured images based on the graphical structure; calculating the first pose of the first frame captured image based on the dataset, and determining the depth map data of the first frame captured image in the coordinate system of the 3D model; applying texture mapping to the first frame captured image based on the first pose and the depth map data, until all target frame captured images have been textured, thereby generating texture data for the 3D model.

2. The method according to claim 1, characterized in that, The step of establishing the graphic structure between the target frame images in the multi-frame captured images includes: Feature matching is performed based on the extracted features of the multi-frame captured images to obtain the feature matching relationship between the multi-frame captured images; Based on the feature matching relationship between the multi-frame captured images, a graphical structure is established between the target frame captured images in the multi-frame captured images.

3. The method according to claim 1, characterized in that, The step of calculating the first pose of the first frame image based on the dataset and determining the depth map data of the first frame image in the three-dimensional model coordinate system includes: The first frame image and the dataset are subjected to feature matching to obtain a first matching pair with a first matching relationship, wherein the first matching relationship reflects the matching relationship between two-dimensional points and three-dimensional points; The first pose of the first frame image is calculated based on the first matching pair, and the camera parameters of the camera device that captured the first frame image are determined. The depth map data of the first frame captured image in the three-dimensional model coordinate system is determined based on the camera parameters.

4. The method according to claim 3, characterized in that, The camera parameters of the imaging device used to capture the first frame image include: The first matching pair is expanded based on the first pose and the feature descriptor corresponding to each 3D point in the dataset; The camera parameters of the imaging device for capturing the first frame image are determined based on the expanded first matching pair.

5. The method according to claim 3, characterized in that, After determining the depth map data of the first frame captured image in the three-dimensional model coordinate system, the method further includes: Based on the graphic structure, a second frame image that matches the first frame image is determined in the target frame image; The first frame image and the second frame image are subjected to feature matching to obtain a second matching pair with the first matching relationship; The second pose of the second frame image is calculated based on the second matching pair, and the camera parameters are jointly optimized based on the second matching pair and the first matching pair. The depth map data of the second frame image in the three-dimensional model coordinate system is determined based on the optimized camera parameters.

6. The method according to claim 5, characterized in that, The step of performing feature matching between the first frame image and the second frame image to obtain a second matching pair having the first matching relationship includes: The first frame image and the second frame image are subjected to feature matching to obtain a third matching pair with a second matching relationship, wherein the second matching relationship reflects the matching relationship between two-dimensional points; Based on the depth map data of the first frame image and the third matching pair, a second matching pair with the first matching relationship is obtained.

7. A three-dimensional scanning device, characterized in that, The device includes: The acquisition module is used to acquire a first image set, a second image set, and a 3D model of the scanned object. The first image set includes multiple captured images, and the second image set includes multiple scanned images. The module is used to establish the graphic structure between the target frame images in the multi-frame captured images; A construction module is used to construct a dataset of the surface of the three-dimensional model based on the three-dimensional points in the coordinate system of the three-dimensional model and the multi-frame scan images. This includes: extracting features from the multi-frame scan images, wherein the features of each scan image include multiple two-dimensional points and a feature descriptor corresponding to each two-dimensional point; determining the feature descriptor corresponding to each three-dimensional point of the three-dimensional model based on the features of the multi-frame scan images; and constructing the dataset of the surface of the three-dimensional model based on the three-dimensional points in the coordinate system of the three-dimensional model and the feature descriptor corresponding to each three-dimensional point, using the three-dimensional model as a framework. The generation module is used to generate texture data of the three-dimensional model based on the graphic structure and the dataset, including: determining a first frame image in the target frame captured images based on the graphic structure; calculating the first pose of the first frame captured image based on the dataset, and determining the depth map data of the first frame captured image in the coordinate system of the three-dimensional model; applying texture mapping to the first frame captured image based on the first pose and the depth map data, until all target frame captured images are textured, and then generating texture data of the three-dimensional model.

8. An electronic device, characterized in that, include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the three-dimensional scanning method as described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the three-dimensional scanning method as described in any one of claims 1 to 6.