Fast object reconstruction method, system and device based on 3D Gaussian sputtering segmentation

Through the rapid reconstruction method of objects based on 3D Gaussian sputtering segmentation, the problem of lack of geometric constraints at the Gaussian points at the boundary during rapid reconstruction of objects is solved, and efficient and accurate object boundary segmentation and reconstruction are achieved, which significantly improves the reconstruction quality and speed.

CN119784959BActive Publication Date: 2025-05-27SICHUAN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510283127.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-05-27
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

In the prior art, the Gaussian points at the boundary when objects are rapidly reconstructed in rapid development, and they are prone to span multiple objects, thus creating ambiguity during segmentation, and the target object boundary cannot be accurately obtained, which has the problem of low reconstruction quality.

Method used

By providing a rapid reconstruction method for objects based on 3D Gaussian sputtering segmentation, it includes acquiring the appearance image of the target object, acquiring image acquisition quality data and extracting three-dimensional reconstruction information, optimizing the image data set, reconstructing voxels, cropping voxels, generating depth maps, segmenting the target object mask, and obtaining a standard 3D Gaussian sputtering model of the target object through iterative optimization.

Benefits of technology

Consistent segmentation across perspective angles is achieved, which significantly improves the clarity of the target object boundary, improves the segmentation speed, improves the reconstruction quality, and ensures the accuracy of the target object boundary.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119784959B_ABST
    Figure CN119784959B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system and device for rapid reconstruction of an object based on 3D Gaussian sputtering segmentation, belonging to the technical field of image data processing. The method includes the following steps: collecting an appearance image of a target object, obtaining image quality and three-dimensional reconstruction information, and screening and optimizing an image set according to the image quality; converting sparse point clouds into voxels, cropping the optimized image set to obtain 2D regions from each perspective, and using the camera pose and the 2D regions to crop the voxels to obtain target voxels; rendering the target voxels as a depth map, generating sampling points according to the depth map, using the sampling points to segment the optimized image set to obtain a target object mask, and cropping to obtain segmentation results from each perspective; projecting the segmentation results back to the initial perspective, generating a depth map, comparing it with the target perspective mask to obtain a consistency evaluation value, and iteratively segmenting the results according to the evaluation value to obtain a set of target object masks; reconstructing a 3D Gaussian sputtering model according to the cropped voxels and Gaussian points.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image data processing, and in particular to a method, system and device for rapid object reconstruction based on 3D Gaussian sputtering segmentation. Background Art

[0002] As a breakthrough technology in the field of new perspective synthesis, 3D Gaussian sputtering has wide applications in 3D reconstruction and virtual reality. Existing segmentation methods first use 3D Gaussian sputtering to reconstruct the entire scene, and then expand the 2D segmentation result generated by the 2D basic segmentation model SM to 3D to segment the 3D Gaussian sputtering model of the target object.

[0003] Existing methods for rapid object reconstruction are mainly implemented through a variety of technical means, including the use of AI technology to perform three-dimensional modeling by taking photos of real objects, processing depth data obtained by two-dimensional or three-dimensional laser scanners based on laser scanning data, three-dimensional reconstruction of mirror object surfaces based on the principle of light reflection, and a new method of using 2D diffusion models to complete incomplete 3D objects.

[0004] For example, the invention patent with announcement number: CN109360267B announces a method for fast three-dimensional reconstruction of thin objects, including: fixing the thin object to be reconstructed on a base with marking points, collecting depth images and color images through a depth camera, identifying the marking points, and obtaining a transformation matrix, performing coordinate transformation on the point cloud of the object to be reconstructed according to the transformation matrix, and finally splicing the point clouds of the two perspectives of the transformed object to complete the three-dimensional reconstruction of the sample. The present invention splices the point clouds of two perspectives based on the marking points to achieve three-dimensional reconstruction of thin objects, so for thin objects with a thickness of about 2mm-30mm, the reconstruction speed is fast, the reconstruction accuracy is high, and the operation is simple.

[0005] For example, the invention patent with announcement number: CN110288642B discloses a method for rapid reconstruction of three-dimensional objects based on a camera array, including: in structured light projection and camera array acquisition, firstly calibrating the digital projector and the camera array, using the camera array to acquire the three-dimensional scene to be measured, obtaining the projection result light stripes and the light stripe images captured by the camera array, and establishing the correspondence between the same point in space on different images; then calculating the offset of each deformed stripe according to the determined shooting center depth plane; then reconstructing the modulation of the three-dimensional point cloud according to the adjustment of the deformed stripe center depth plane, calculating the corresponding depth distance of the deformed stripe image and the modulation of the deformed stripe image at different focusing depths; registering and reconstructing the three-dimensional point cloud, iterating the nearest point algorithm to solve the coordinate transformation, and obtaining a complete reconstruction model of the three-dimensional scene to be measured.

[0006] However, in the process of implementing the technical solution of the invention in the embodiments of the present application, the present application found that the above technology has at least the following technical problems:

[0007] In the prior art, when objects are rapidly reconstructed, Gaussian points at the boundaries lack geometric constraints and easily cross multiple objects, which causes ambiguity during segmentation and makes it impossible to accurately obtain the boundaries of the target objects, resulting in low reconstruction quality. Summary of the invention

[0008] The present invention provides a method, system and device for rapid object reconstruction based on 3D Gaussian sputtering segmentation, thereby solving the problem in the prior art that Gaussian points at the boundaries of objects lack geometric constraints during rapid object reconstruction, easily cross multiple objects, thus causing ambiguity during segmentation, and failing to accurately obtain the boundaries of the target objects, resulting in low reconstruction quality. The present invention realizes a complete process from image acquisition to three-dimensional reconstruction, and then to image processing and 3D model generation.

[0009] The present invention provides a method for rapid object reconstruction based on 3D Gaussian sputtering segmentation, comprising the following steps: collecting an appearance image of a target object, obtaining image collection quality data and extracting 3D reconstruction information, wherein the 3D reconstruction information includes an initial image set, camera posture information, a sparse point cloud and an accurate mask of a target view; performing quality screening on an optimized image data set according to the image collection quality data to obtain an optimized image data set; reconstructing the sparse point cloud into voxels, preliminarily cropping the optimized image data set to obtain a 2D image area under projection of each view angle, performing voxel cropping on the optimized image data set according to the camera posture information and the 2D image area under projection of each view angle to obtain cropped voxels of each view angle, cropping according to view angle continuity to obtain cropped target voxels of each view angle, and rendering the cropped target voxels of each view angle as a depth map; generating sampling points according to the depth map, segmenting the optimized image data set according to the sampling points, generating a target object mask, and performing voxel cropping on the target object mask to obtain segmentation results of each view angle; projecting the current voxel of the segmentation result back to the initial prompt view angle to generate a depth map, using the accurate mask of the target view angle to obtain the depth map, and performing the depth map according to the D 1′ and D 1 A consistency evaluation value is obtained through comprehensive analysis, and the segmentation results are iterated according to the consistency evaluation value to obtain a target object mask set; the optimized image data set is masked according to the target object mask set to obtain an RGB image of the target object, and the target object is reconstructed based on the voxels cropped from each perspective to obtain an initial 3D Gaussian sputtering model; the initial 3D Gaussian sputtering model is initially restored according to the depth map of the original scene voxels and the depth map of the target object voxels to obtain a restored 3D Gaussian sputtering model; and the Gaussian points of the restored 3D Gaussian sputtering model are filtered according to the voxels of the target object to obtain a standard 3D Gaussian sputtering model.

[0010] Furthermore, the optimized image data set is quality screened according to the image acquisition quality data, and the step of obtaining the optimized image data set includes: the image acquisition quality data includes scanning device performance data and shooting environment data; a scanning device performance evaluation value is obtained according to a comprehensive analysis of the scanning device performance data, and evaluation feedback is performed according to the scanning device performance evaluation value; a shooting environment evaluation value of each initial image is obtained according to a comprehensive analysis of the shooting environment data; an image acquisition quality evaluation value of each initial image is obtained according to a comprehensive analysis of the scanning device performance evaluation value and the shooting environment evaluation value of each initial image, and the initial image set is quality screened according to the image acquisition quality evaluation value of each initial image to obtain the optimized image data set.

[0011] Furthermore, the step of obtaining a scanning device performance evaluation value based on a comprehensive analysis of the scanning device performance data includes: the scanning device performance data includes scanning time, scanning resolution and scanning frequency; obtaining a reference scanning time, an allowable deviation scanning time, a critical scanning resolution, a reference scanning frequency and an allowable deviation scanning frequency from an object reconstruction database; and obtaining a scanning device performance evaluation value through comprehensive analysis.

[0012] Furthermore, the step of providing evaluation feedback based on the scanning device performance evaluation value includes: obtaining a scanning device performance evaluation threshold from an object reconstruction database; comparing the scanning device performance evaluation value with the scanning device performance evaluation threshold; if the scanning device performance evaluation value is greater than or equal to the scanning device performance evaluation threshold, no additional processing is performed; if the scanning device performance evaluation value is less than the scanning device performance evaluation threshold, adjustment feedback is provided to the scanning device.

[0013] Furthermore, the step of obtaining a shooting environment evaluation value of each initial image based on a comprehensive analysis of the shooting environment data includes: the shooting environment data includes light intensity, light angle and light change rate; obtaining a reference light intensity, an allowable deviation light intensity, a reference light angle, an allowable deviation light angle and a critical light change rate from an object reconstruction database; and obtaining a shooting environment evaluation value of each initial image through a comprehensive analysis.

[0014] Furthermore, the step of obtaining an image acquisition quality evaluation value of each initial image by comprehensively analyzing the scanning device performance evaluation value and the shooting environment evaluation value of each initial image includes: obtaining a critical scanning device performance evaluation value and a critical shooting environment evaluation value from an object reconstruction database; and obtaining an image acquisition quality evaluation value of each initial image by comprehensively analyzing the scanning device performance evaluation value, the shooting environment evaluation value of each initial image, the critical scanning device performance evaluation value and the critical shooting environment evaluation value.

[0015] Furthermore, the image acquisition quality evaluation value of each initial image is obtained as follows:

[0016] ;

[0017] In the formula, ζQ k represents the image acquisition quality assessment value of the kth initial image, α 7 represents the image acquisition quality assessment impact factor corresponding to the scanning device performance evaluation value, α 8 represents the image acquisition quality assessment impact factor corresponding to the shooting environment assessment value, ζP 1 represents the performance evaluation value of the scanning device, ζP 0 represents the critical scanning equipment performance evaluation value, ζL 1k represents the shooting environment evaluation value of the kth initial image, ζL 0 Indicates the critical shooting environment evaluation value.

[0018] Furthermore, the initial image set is quality screened according to the image acquisition quality assessment value of each initial image to obtain the step of optimizing the image data set, including: obtaining the image acquisition quality assessment threshold from the object reconstruction database; comparing the image acquisition quality assessment value of each initial image with the image acquisition quality assessment threshold; if the image acquisition quality assessment value of an initial image is greater than or equal to the image acquisition quality assessment threshold, marking the initial image as a qualified image; if the image acquisition quality assessment value of an initial image is less than the image acquisition quality assessment threshold, marking the initial image as an unqualified image, and reacquiring the initial image of the angle; and counting the qualified images to obtain the optimized image data set.

[0019] The embodiment of the present application provides a rapid object reconstruction system based on 3D Gaussian sputtering segmentation, comprising a data acquisition module, an image screening module, a voxel processing module, and an object reconstruction module; wherein the data acquisition module is used to acquire an appearance image of a target object, acquire image acquisition quality data, and extract three-dimensional reconstruction information, wherein the three-dimensional reconstruction information includes an initial image set, camera pose information, a sparse point cloud, and an accurate mask of a target viewing angle; the image screening module is used to perform quality screening on an optimized image data set according to the image acquisition quality data to obtain an optimized image data set; the voxel processing module is used to reconstruct the sparse point cloud into voxels, and perform preliminary The object reconstruction module is used to project the current voxel of the segmentation result back to the initial prompt view to generate a depth map, and the target voxel after cropping from each view is obtained by cropping according to the continuity of the view. The target voxel after cropping from each view is rendered as a depth map; each sampling point is generated according to the depth map, and the optimized image data set is segmented according to each sampling point to generate a target object mask, and the target object mask is subjected to voxel cropping to obtain the segmentation result of each view; the object reconstruction module is used to project the current voxel of the segmentation result back to the initial prompt view to generate a depth map, and the depth map is obtained by using the accurate mask of the target view. 1′ and D 1 A consistency evaluation value is obtained through comprehensive analysis, and the segmentation results are iterated according to the consistency evaluation value to obtain a target object mask set; the optimized image data set is masked according to the target object mask set to obtain an RGB image of the target object, and the target object is reconstructed based on the voxels cropped from each perspective to obtain an initial 3D Gaussian sputtering model; the initial 3D Gaussian sputtering model is initially restored according to the depth map of the original scene voxels and the depth map of the target object voxels to obtain a restored 3D Gaussian sputtering model; and the Gaussian points of the restored 3D Gaussian sputtering model are filtered according to the voxels of the target object to obtain a standard 3D Gaussian sputtering model.

[0020] An embodiment of the present application provides a device for rapid reconstruction of objects based on 3D Gaussian sputtering segmentation, comprising: a processor and a memory for storing instructions executable by the processor; when the processor is configured to execute the instructions, the device for rapid reconstruction of objects based on 3D Gaussian sputtering segmentation implements a method for rapid reconstruction of objects based on 3D Gaussian sputtering segmentation.

[0021] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0022] 1. The present invention provides a method, system and device for rapid reconstruction of objects based on 3D Gaussian sputtering segmentation, thereby achieving cross-viewing consistent segmentation in 2D images, and then locally reconstructing the target object through 3D Gaussian sputtering, thereby achieving not only clear segmentation of the target object boundary, but also a 2 to 6 times increase in segmentation speed.

[0023] 2. The present invention introduces the 2D mask of the target object to supervise the 3D Gaussian sputtering training and adds geometric constraints, thereby enhancing the model's ability to finely capture the boundaries of the target object, thereby achieving a significant improvement in boundary clarity.

[0024] 3. The present invention iteratively optimizes the 3D geometric representation of the target object through perspective continuity, which helps the model gradually eliminate errors caused by observation errors or data noise, thereby continuously approaching the true shape and structure of the target object, thereby significantly improving the accuracy of the segmentation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 A flow chart of a method for rapid object reconstruction based on 3D Gaussian sputtering segmentation provided in an embodiment of the present application.

[0026] Figure 2 This is an example of SAM segmentation failure provided in an embodiment of the present application.

[0027] Figure 3 This is an example of sampling failure caused by an obstruction provided in an embodiment of the present application.

[0028] Figure 4 A graph showing changes in the performance evaluation value of a scanning device based on a rapid object reconstruction method provided in an embodiment of the present application.

[0029] Figure 5 A schematic diagram of the structure of a system for rapid object reconstruction based on 3D Gaussian sputtering segmentation provided in an embodiment of the present application. DETAILED DESCRIPTION

[0030] The embodiments of the present application provide a method, system and device for rapid reconstruction of objects based on 3D Gaussian sputtering segmentation, which solves the problem that the Gaussian points at the boundaries of objects in the prior art lack geometric constraints during rapid reconstruction, easily cross multiple objects, and thus cause ambiguity during segmentation, and cannot accurately obtain the boundaries of the target objects, resulting in low reconstruction quality. By acquiring an appearance image of the target object, image acquisition quality data is obtained and three-dimensional reconstruction information is extracted, wherein the three-dimensional reconstruction information includes an initial image set, camera pose information, a sparse point cloud and an accurate mask of the target perspective; the optimized image data set is quality screened according to the image acquisition quality data to obtain an optimized image data set; the sparse point cloud is reconstructed into voxels , perform preliminary cropping on the optimized image data set to obtain the 2D image area under each viewing angle projection, perform voxel cropping according to the camera posture information and the 2D image area under each viewing angle projection, obtain the cropped voxels of each viewing angle, and obtain the target voxels cropped from each viewing angle by cropping according to the viewing angle continuity, and render the cropped target voxels from each viewing angle into a depth map; generate each sampling point according to the depth map, segment the optimized image data set according to each sampling point, generate a target object mask, and perform voxel cropping on the target object mask to obtain the segmentation results of each viewing angle; project the current voxel of the segmentation result back to the initial prompt viewing angle to generate a depth map, and use the precise mask of the target viewing angle to obtain the depth map, and according to D 1′ and D 1 A consistency evaluation value is obtained through comprehensive analysis, and the segmentation results are iterated according to the consistency evaluation value to obtain a target object mask set; the optimized image data set is masked according to the target object mask set to obtain an RGB image of the target object, and the target object is reconstructed based on the voxels cropped from each perspective to obtain an initial 3D Gaussian sputtering model; the initial 3D Gaussian sputtering model is initially restored according to the depth map of the original scene voxels and the depth map of the target object voxels to obtain a restored 3D Gaussian sputtering model; the Gaussian points of the restored 3D Gaussian sputtering model are filtered according to the voxels of the target object to obtain a standard 3D Gaussian sputtering model, realizing a complete process from image acquisition to three-dimensional reconstruction, and then to image processing and 3D model generation.

[0031] The technical solution in the embodiment of the present application is to solve the problem that the Gaussian points at the boundary lack geometric constraints during the rapid reconstruction of the above-mentioned objects, and are easy to cross multiple objects, thereby causing ambiguity during segmentation, and the boundary of the target object cannot be accurately obtained, resulting in low reconstruction quality. The overall idea is as follows:

[0032] By acquiring the appearance image of the target object, the image acquisition quality data is obtained and the 3D reconstruction information is extracted, wherein the 3D reconstruction information includes the initial image set, the camera posture information, the sparse point cloud and the precise mask of the target view; the optimized image data set is quality screened according to the image acquisition quality data to obtain the optimized image data set; the sparse point cloud is reconstructed into voxels, and the optimized image data set is preliminarily cropped to obtain the 2D image area under the projection of each view angle, and voxel cropping is performed according to the camera posture information and the 2D image area under the projection of each view angle to obtain the cropped voxels of each view angle, and the cropped target voxels of each view angle are obtained by cropping according to the continuity of the view angle, and the cropped target voxels of each view angle are rendered as a depth map; each sampling point is generated according to the depth map, and the optimized image data set is segmented according to each sampling point to generate a target object mask, and the target object mask is voxel cropped to obtain the segmentation results of each view angle; the current voxel of the segmentation result is projected back to the initial prompt view angle to generate a depth map, and the depth map is obtained by using the precise mask of the target view angle, and the depth map is obtained according to the D 1′ and D 1 A consistency evaluation value is obtained through comprehensive analysis, and the segmentation results are iterated according to the consistency evaluation value to obtain a target object mask set; the optimized image data set is masked according to the target object mask set to obtain an RGB image of the target object, and the target object is reconstructed based on the voxels cropped from each perspective to obtain an initial 3D Gaussian sputtering model; the initial 3D Gaussian sputtering model is initially restored according to the depth map of the original scene voxels and the depth map of the target object voxels to obtain a restored 3D Gaussian sputtering model; the Gaussian points of the restored 3D Gaussian sputtering model are filtered according to the voxels of the target object to obtain a standard 3D Gaussian sputtering model, thereby achieving 3D model generation.

[0033] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0034] like Figure 1 As shown, it is a flow chart of a method for rapid reconstruction of an object based on 3D Gaussian sputtering segmentation provided in an embodiment of the present application, the method comprising the following steps: acquiring an appearance image of the target object, obtaining image acquisition quality data and extracting 3D reconstruction information, wherein the 3D reconstruction information comprises an initial image set, camera posture information, a sparse point cloud and an accurate mask of the target perspective; performing quality screening on the optimized image data set according to the image acquisition quality data to obtain an optimized image data set; reconstructing the sparse point cloud into voxels, performing preliminary cropping on the optimized image data set to obtain a 2D image area under each perspective projection, performing voxel cropping according to the camera posture information and the 2D image area under each perspective projection to obtain cropped voxels for each perspective, cropping according to perspective continuity to obtain cropped target voxels for each perspective, and rendering the cropped target voxels for each perspective as a depth map.

[0035] The formula for performing preliminary cropping on the optimized image dataset is as follows:

[0036] ;

[0037] Where M i Represents G target In the perspective P i The 2D image area under projection, G target Represents the 3D geometric expression of the target object, Π is the projection function defined by the internal and external parameters of the camera, where i is the number of each view, i=1,2,3,...,n, and n is the total number of views.

[0038] The present invention reconstructs sparse point clouds into voxels as scene geometry representation, and combines the idea of ​​visual shell to scene (scene initial geometry) for cropping. As the number of viewing angles used for cropping increases, the cropping result gradually approaches G target .

[0039] The visual shell is the largest enclosing geometric body formed by inversely rendering the multi-view 2D outline into 3D space. In order to facilitate the clipping operation based on the visual shell, the present invention converts the scene point cloud into voxels V. Voxels are cube grids regularly distributed in space, which are not only more efficient in calculation and storage, but also provide a consistent geometric representation that can more completely describe the surface information of objects under different viewing angles. The method for voxel clipping is as follows:

[0040] ;

[0041] Where V n represents the voxel after cropping at n viewpoints, П -1 is to change the viewing angle P i The following outline is the inverse rendering operation to 3D space.

[0042] Use user initial prompt M u After the first clipping of the voxel, we get the voxel V 1 However, since it is difficult to completely exclude the redundant voxels behind the target object by cropping with a single viewpoint, V 1 Usually rough, it requires further iterative cropping through multiple perspectives to get close to G target .

[0043] The cropping process is based on the principle of perspective continuity. The self-hinting strategy generates hint points to guide SAM to achieve accurate segmentation, thereby completing cross-view consistency segmentation and finally obtaining the accurate mask of the target object under all perspectives. The goal of cross-view consistency segmentation is to obtain the 3D geometric expression G of the target object. targetTo this end, the present invention utilizes the perspective continuity of the picture set to iteratively crop redundant voxels. Perspective continuity means that the perspective change amplitude of two adjacent frames in the picture set is small, and each change in perspective will only cause a small number of non-target voxels to appear at the edge of the object.

[0044] Specifically, in each iteration, the formula for rendering the target voxels cropped from each view into a depth map is as follows:

[0045] ;

[0046] Where D(x,y) represents the depth value of the pixel with coordinates (x,y) on the depth map, z i represents the depth value of the target voxel, P i Represents the camera pose information, ∏ is the projection function, pos i is the coordinate of the target voxel at view angle i, V is the voxel set, and the indicator function δ is in voxel V i Output 1 when projected onto pixel (x,y), otherwise output 0.

[0047] The sampling points are generated according to the depth map, the optimized image data set is segmented according to the sampling points, a target object mask is generated, and the target object mask is voxel-cut to obtain the segmentation results of each viewing angle.

[0048] The depth map D has the following two problems: (1) Non-target pixel interference: changes in viewing angle may cause incompletely cropped voxels to generate a small number of non-target pixels. (2) Rough expression: Voxels are only approximate geometric expressions of the target object, resulting in the generated mask only being able to roughly represent the approximate area of ​​the target object. In order to solve these problems, the present invention introduces a self-hinting strategy to generate sampling points on the depth map, thereby guiding SAM to accurately segment the target object. The core of the self-hinting strategy is to generate representative sampling points from the depth map to guide SAM to obtain a more accurate target object mask.

[0049] First, the depth map D is converted into a binary rough mask M c , as the basis for sampling. The selection of sampling points is crucial. K-means clustering is used to c The pixels in the cluster are grouped to capture the multiple subclasses that the target object may contain. ; As sampling points, the calculation formula of the cluster center set is as follows:

[0050] ;

[0051] Among them, C is the cluster center set, which represents the selected sampling points, is the coordinate set of all target pixels in the mask, M c represents a binary rough mask, is the coordinate of the jth cluster center, and the argmin() function represents the independent variable that minimizes the objective function, where j is the number of each cluster center, j=1,2,...,n, and n is the total number of cluster centers. The cluster center effectively captures the structural features of the target object and can provide a stable sampling point for segmentation. Through the sampling point C, SAM performs a segmentation on the original image I at the current viewing angle. t Segmentation is performed to generate a more accurate target object mask M f , for M f Using the formula for voxel clipping of the optimized image dataset, the voxels at the edge that are not sufficiently excluded under the current viewing angle can be clipped.

[0052] Through the above method, all perspectives are gradually cropped, and the voxel set gradually approaches G target , and store the target object mask generated in each iteration as the result of cross-view consistent segmentation.

[0053] The current voxel of the segmentation result is projected back to the initial prompt view to generate a depth map, and the depth map is obtained using the accurate mask of the target view. 1′ and D 1 A consistency evaluation value is obtained through comprehensive analysis, and the segmentation results are iterated according to the consistency evaluation value to obtain a set of target object masks.

[0054] Even if the cue points are reasonably distributed, the segmentation results of SAM may still be unsatisfactory. A common problem is that only a local area of ​​the target object is segmented. Figure 2 As shown, if such an error mask is directly used for clipping, voxels belonging to the target object will be incorrectly clipped.

[0055] In order to detect segmentation errors, the present invention designs an evaluation method based on the depth consistency ratio (DCR). The segmentation result M of the current perspective is obtained. f i After that, the current voxel is first clipped, and the clipped voxel is V i′ Then, V is rendered as the depth map D according to the formula of the target voxel at the current view. i′ Project back to the initial prompt view to generate the depth map D 1′ Since the initial hint mask M u is absolutely correct, the initial clipping voxel V 1 The depth map D rendered at the initial hint perspective 1 This is also completely correct. Therefore, by comparing D 1′ and D 1 The consistency of the segmentation results is evaluated, and the calculation formula for consistency evaluation is as follows:

[0056] ;

[0057] In the formula, DCR represents the consistency assessment value, D 1′ (x) represents the generated image depth map of the x-th pixel, D 1 (x) represents the correct depth map of the x-th pixel, o ′ It represents the depth difference tolerance threshold, and the H function is an indicator function that returns 1 when the condition is met. When DCR is less than 0.95, the SAM segmentation result is considered to have failed. Where x is the pixel number, x=1,2,3,...,X, and X is the total number of pixels in the depth map.

[0058] When segmentation fails, the resampling process is entered to adjust the position and distribution of the cue points and redirect the SAM segmentation. Resampling adjusts the initial cluster center of the k-means clustering to generate a new set of cue points. The new cue points are used to guide SAM to generate a new mask, which is then re-evaluated. To ensure the efficiency of the algorithm and prevent an infinite loop, a maximum number of resampling times is set. If the requirements are still not met after multiple resamplings, the segmentation result of the current view is rejected and the next view iteration is directly entered.

[0059] The optimized image data set is masked according to the target object mask set to obtain an RGB image of the target object, and the target object is reconstructed according to the cropped voxels of each viewpoint to obtain an initial 3D Gaussian sputtering model. The initial 3D Gaussian sputtering model is initially restored according to the depth map of the original scene voxels and the depth map of the target object voxels to obtain a restored 3D Gaussian sputtering model, and the Gaussian points of the restored 3D Gaussian sputtering model are filtered according to the voxels of the target object to obtain a standard 3D Gaussian sputtering model.

[0060] After completing the iteration of all viewpoints, a set of finely segmented target object masks is obtained. . Combining these masks, we can optimize the image dataset Perform mask processing to obtain the RGB image of the target object. n Downsampling is performed to generate the point cloud required for 3D Gaussian sputtering initialization. These RGB images are used as supervision and input into the 3D Gaussian sputtering training pipeline to guide the accurate reconstruction of the target object. Since the target object in the supervision image has a clear boundary, under this natural geometric constraint, the reconstructed 3D Gaussian sputtering model of the target object also has a clear boundary.

[0061] In some scenes, the target object may be partially occluded by other objects at a certain viewing angle. In order to ensure that the 3D Gaussian sputtering model of the target object can still fully render the target object image under the occlusion viewing angle, it is necessary to introduce additional strategies into the original segmentation process. First, the sampling area problem under the occlusion viewing angle needs to be solved. If the sampling is still performed directly on the pixel point set X, some of the generated sampling points will be located at the position of the occluder in the original image, which will cause the segmentation result to contain part of the occluder, such as Figure 3 To avoid this problem, the depth information of scene voxels and target object voxels can be used to identify the unoccluded pixel set X u The specific steps are as follows:

[0062] (1) Render the depth map D of the original scene voxels scene , and use it as the depth template S D ;

[0063] (2) Render the depth map D of the target object voxels target , using the depth template S D Conduct testing;

[0064] (3) The pixels that pass the depth template test constitute the unoccluded pixel set X u The formula for identifying the set of pixels that are not blocked is as follows:

[0065] ;

[0066] Then replace X with X u , and the calculation formula of the cluster center set is used to guide the SAM model segmentation to obtain the mask M under the occlusion perspective u Since some areas are blocked, the mask M u It is usually incomplete, and the obscured parts are marked as the background color.

[0067] The 3D Gaussian sputtering model in the present invention uses a set of three-dimensional Gaussian distributions G to preserve the properties required for rendering. The mean of each Gaussian distribution represents its position in three-dimensional space, and the covariance represents its scale. Given a specific camera pose, the 3D Gaussian sputtering projected the three-dimensional Gaussian into two-dimensional space, and then mixed a set of ordered Gaussian distributions G that overlapped with the light r. r To calculate the color C of the light. Let g i Represents G r The i-th Gaussian distribution in , this process can be expressed as:

[0068] ;

[0069] ;

[0070] In the formula, C(r) represents the color of light r, c gi represents the Gaussian point g i The color, α gi represents the Gaussian point g j The opacity, w gi Indicates the contribution of the current Gaussian point to the color value.

[0071] If the incomplete mask M is used u As a supervisory image, the 3D Gaussian sputtering model will converge to an incomplete target object under occlusion. This is because the 3D Gaussian sputtering model strictly follows the supervisory image for fitting, and a large number of meaningless background-colored Gaussian points will be generated in the occluded area to fit M. u The occluded area of ​​the background color.

[0072] To solve this problem, the present invention uses the voxels of the target object to filter the Gaussian points. It is not reliable to directly determine whether the center of the Gaussian point is located inside the voxel, because the voxel is only a rough representation of the target object, and there may be holes on the surface, which can easily lead to false elimination. Therefore, the present invention adopts a more relaxed clipping strategy, that is, to determine whether the center of the Gaussian point is located in the directed bounding box of the voxel. The directed bounding box can more accurately reflect the 3D spatial geometric distribution of the target object.

[0073] The goal of the directed bounding box is to find the best bounding box V n The core of the minimum rotation box is to calculate the main direction, which is calculated by calculating V n The centroid of the voxel is redefined to calculate the covariance matrix Σ, which is obtained by performing eigenvalue decomposition on Σ. The calculation formula of the covariance matrix is:

[0074] ;

[0075] Where Σ represents the covariance matrix, V represents the orthogonal eigenvector matrix, which is used to represent the main direction of the point cloud, and Λ represents the diagonal eigenvalue matrix, which is used to represent the variance in the main direction.

[0076] Using V n The directed bounding box of the 3D Gaussian sputtering model is used to filter the Gaussian points in the 3D Gaussian sputtering model, which can quickly and effectively remove meaningless Gaussian points, and finally obtain a clean 3D Gaussian sputtering model of the target object. Through this processing, the algorithm proposed in the present invention can effectively deal with the segmentation and 3D Gaussian sputtering model convergence problems in occluded scenes, thereby ensuring that the target object can be fully rendered even under occluded viewing angles.

[0077] In this embodiment, the input of the present invention is basically consistent with the input of 3D Gaussian sputtering, including an optimized image data set, camera pose information, a sparse point cloud, and an accurate mask of the target perspective, wherein the represents the optimized image dataset, Represents the camera pose information, M u It represents the precise mask of the target view. It is a hint to the user about the target object under a single view. Because various hint forms (such as points, strokes, bounding boxes, etc.) can be converted into the target object mask through the basic 2D segmentation model. The relevant formula of the basic 2D segmentation model is: SAM uses an encoder-decoder architecture. Encoder S e Receives image F as input and outputs the corresponding feature f F :f F =S e (F); Decoder S d Receive feature f I and a set of prompts P as input, the prompts p∈P can be points, boxes, texts and masks, and output the corresponding 2D segmentation mask M SAM , expressed in the form of a bitmap, namely: M SAM =S d (f F ,P).

[0078] In addition, the object reconstruction database is used to store relevant data of the object rapid reconstruction method based on 3D Gaussian sputtering segmentation, including reference scanning time, allowable deviation scanning time, critical scanning resolution, critical scanning frequency, scanning equipment performance evaluation influencing factors corresponding to the scanning time, and scanning equipment performance evaluation threshold, etc. The data in the object reconstruction database can be obtained by consulting public databases such as ShapeNet or relevant academic papers, and can also be obtained through cooperation and sharing with 3D scanning companies.

[0079] Furthermore, the optimized image data set is quality screened according to the image acquisition quality data, and the step of obtaining the optimized image data set includes: the image acquisition quality data includes scanning device performance data and shooting environment data; a scanning device performance evaluation value is obtained according to a comprehensive analysis of the scanning device performance data, and evaluation feedback is performed according to the scanning device performance evaluation value; a shooting environment evaluation value of each initial image is obtained according to a comprehensive analysis of the shooting environment data; an image acquisition quality evaluation value of each initial image is obtained according to a comprehensive analysis of the scanning device performance evaluation value and the shooting environment evaluation value of each initial image, and the initial image set is quality screened according to the image acquisition quality evaluation value of each initial image to obtain the optimized image data set.

[0080] In this embodiment, in the image-based 3D reconstruction process, the image acquisition quality is crucial. In order to ensure the accuracy and efficiency of subsequent reconstruction, the acquired images need to be firstly screened for quality. Through this step, we can ensure that the image data used for 3D reconstruction has a high quality level. This can not only improve the accuracy and efficiency of reconstruction, but also reduce errors and unnecessary calculations in subsequent processing.

[0081] Furthermore, the step of obtaining a scanning device performance evaluation value based on a comprehensive analysis of the scanning device performance data includes: the scanning device performance data includes scanning time, scanning resolution and scanning frequency; obtaining a reference scanning time, an allowable deviation scanning time, a critical scanning resolution, a reference scanning frequency and an allowable deviation scanning frequency from an object reconstruction database; and obtaining a scanning device performance evaluation value through comprehensive analysis.

[0082] The scanning device performance evaluation value is obtained as follows:

[0083] ;

[0084] In the formula, ζP 1 represents the performance evaluation value of the scanning device, α 1 represents the scanning device performance evaluation impact factor corresponding to the scanning time, α 2 represents the scanning device performance evaluation impact factor corresponding to the scanning resolution, α 3 represents the scanning device performance evaluation impact factor corresponding to the scanning frequency, T 1 Indicates the scan time, T 0 Indicates the reference scan time, T 2 Indicates the allowable deviation scanning time, O 1 Indicates the scanning resolution, O 0 represents the critical scanning resolution, N 1 Indicates the scanning frequency, N 0 Indicates the reference scanning frequency, N 2 It represents the allowable deviation scanning frequency, and e is a natural constant.

[0085] α 1 , α 2 and α 3The scanning device performance evaluation influencing factors corresponding to the scanning time, scanning resolution and scanning frequency preset in the object reconstruction database respectively represent the numerical values ​​of the influence of the scanning time, scanning resolution and scanning frequency on the scanning device performance evaluation value, and can be directly obtained from the object reconstruction database when used. These relationships are organized into a "mapping set", that is, a lookup table or corresponding rules. When the value of a certain scanning device performance data is actually monitored, this value can be searched according to the corresponding mapping set, and then the corresponding scanning device performance evaluation influencing factor can be obtained. This impact factor is a number between 0 and 1, which represents the evaluation result of the influence of the current scanning device performance data on the scanning device performance. For example, for scanning time, there is a mapping set of scanning time and scanning device performance evaluation influencing factors corresponding to scanning time. By inputting the value of scanning time, the scanning device performance evaluation influencing factors corresponding to scanning time can be obtained; for scanning resolution, there is a mapping set of scanning resolution and scanning device performance evaluation influencing factors corresponding to scanning resolution. By inputting the value of scanning resolution, the scanning device performance evaluation influencing factors corresponding to scanning resolution can be obtained; for scanning frequency, there is a mapping set of scanning frequency and scanning device performance evaluation influencing factors corresponding to scanning frequency. By inputting the value of scanning frequency, the scanning device performance evaluation influencing factors corresponding to scanning frequency can be obtained. These mapping relationships can be many-to-one or one-to-one.

[0086] In this embodiment, the scanning time can use a timer to record the time required for the scanning device to complete a complete object scan from startup; the scanning resolution and scanning frequency can be obtained by referring to the technical specifications or instructions of the device. The scanning resolution refers to the spatial interval between each sampling point in the point cloud data, and the scanning frequency refers to the number of times the scanning device collects point cloud data per second, that is, the number of scans completed per second. There is a correlation between the three. For example, increasing the scanning resolution or increasing the scanning frequency may lead to longer scanning time; increasing the scanning frequency can increase the scanning speed, but if the scanning resolution is too low, the accuracy and quality of the reconstruction may be reduced. The performance evaluation value of the scanning device obtained by comprehensive analysis can reflect the comprehensive performance of the device under specific tasks, including speed, accuracy and data processing requirements.

[0087] The scanning device performance evaluation factor corresponding to the scanning time is set to 0.3, the scanning device performance evaluation factor corresponding to the scanning resolution is set to 0.4, the scanning device performance evaluation factor corresponding to the scanning frequency is set to 0.3, the scanning time is set to 5 minutes, the reference scanning time is set to 5 minutes, the allowable deviation scanning time is set to 2 minutes, the scanning resolution is set to 1000 μm, the critical scanning resolution is set to 1200 μm, the reference scanning frequency is set to 1000 Hz, and the allowable deviation scanning frequency is set to 100 Hz. When the scanning frequency is continuously increased, the scanning device performance evaluation value is calculated. As shown in Table 1, the scanning device performance evaluation value data table based on the object fast reconstruction method.

[0088] Table 1 Data table of performance evaluation values ​​of scanning equipment based on object rapid reconstruction method

[0089]

[0090] like Figure 4 As shown in Table 1 and Figure 4 It can be seen that the scanning device performance evaluation influencing factor corresponding to the scanning time, the scanning device performance evaluation influencing factor corresponding to the scanning resolution, the scanning device performance evaluation influencing factor corresponding to the scanning frequency, the scanning time, the reference scanning time, the allowable deviation scanning time, the scanning resolution, the critical scanning resolution, the reference scanning frequency and the allowable deviation scanning frequency remain unchanged. When the scanning frequency continues to increase, the closer the scanning frequency is to the reference scanning frequency, the greater the scanning device performance evaluation value will be.

[0091] Furthermore, the step of providing evaluation feedback based on the scanning device performance evaluation value includes: obtaining a scanning device performance evaluation threshold from an object reconstruction database; comparing the scanning device performance evaluation value with the scanning device performance evaluation threshold; if the scanning device performance evaluation value is greater than or equal to the scanning device performance evaluation threshold, no additional processing is performed; if the scanning device performance evaluation value is less than the scanning device performance evaluation threshold, adjustment feedback is provided to the scanning device.

[0092] In this embodiment, the scanning device performance evaluation threshold is a critical value used to determine whether the performance of the current scanning device meets the reconstruction requirements. When adjusting the scanning device for feedback, the performance of the scanner can be improved (for example, the scanning resolution can be improved, the data acquisition density can be increased by increasing the scanning frequency, etc.) or the number of scanners can be increased to improve the overall scanning efficiency and quality, and the equipment can be calibrated regularly to ensure the stability of the performance of the scanning equipment and reduce the low performance evaluation value caused by equipment instability; at the same time, an automatic adjustment mechanism is established, and when the performance evaluation value is detected to be lower than the threshold, the system can automatically adjust the scanning parameters, for example, appropriately reduce the resolution to increase the scanning speed without affecting the reconstruction quality; if the scanning time is too long, for example, the time required to scan a complete object normally is 5 minutes, but the scanning is still not completed after 10 minutes, the scanning frequency can be increased to speed up data acquisition, and real-time feedback is given to the device to ensure that the scanning quality is maintained at an appropriate level each time it is used.

[0093] Furthermore, the step of obtaining a shooting environment evaluation value of each initial image based on a comprehensive analysis of the shooting environment data includes: the shooting environment data includes light intensity, light angle and light change rate; obtaining a reference light intensity, an allowable deviation light intensity, a reference light angle, an allowable deviation light angle and a critical light change rate from an object reconstruction database; and obtaining a shooting environment evaluation value of each initial image through a comprehensive analysis.

[0094] The method for obtaining the shooting environment evaluation value of each initial image is as follows:

[0095] ;

[0096] In the formula, ζL 1k represents the shooting environment evaluation value of the kth initial image, α 4 Indicates the shooting environment assessment impact factor corresponding to the light intensity, α 5 Indicates the shooting environment assessment impact factor corresponding to the illumination angle, α 6 Indicates the shooting environment assessment impact factor corresponding to the illumination change rate, I 1k represents the illumination intensity of the kth initial image, I 0 Indicates the reference light intensity, I 2 Indicates the allowable deviation of light intensity, A 1k represents the illumination angle of the kth initial image, A 0 Indicates the reference illumination angle, A 2 Indicates the allowable deviation illumination angle, V 1k represents the rate of illumination change of the kth initial image, V 0 represents the critical illumination change rate, e is a natural constant, k is the number of each initial image, k=1,2,3,...,K, K is the total number of initial images.

[0097] α 4 , α 5 and α 6 The shooting environment assessment impact factors corresponding to the light intensity, light angle and light change rate preset in the object reconstruction database respectively represent the numerical values ​​of the influence of light intensity, light angle and light change rate on the shooting environment assessment value, and can be directly obtained from the object reconstruction database when used. These relationships are organized into a "mapping set", that is, a lookup table or corresponding rules. When the value of the shooting environment data of a certain initial image is actually monitored, this value can be searched according to the corresponding mapping set, and then the corresponding shooting environment assessment impact factor can be obtained. This impact factor is a number between 0 and 1, which represents the evaluation result of the influence of the current shooting environment data on the shooting environment assessment value. For example, for light intensity, there is a mapping set of light intensity and shooting environment assessment impact factor corresponding to light intensity. By inputting the value of light intensity, the shooting environment assessment impact factor corresponding to light intensity can be obtained; for light angle, there is a mapping set of light angle and shooting environment assessment impact factor corresponding to light angle. By inputting the value of light angle, the shooting environment assessment impact factor corresponding to light angle can be obtained; for light change rate, there is a mapping set of light change rate and shooting environment assessment impact factor corresponding to light change rate. By inputting the value of light change rate, the shooting environment assessment impact factor corresponding to light change rate can be obtained. These mapping relationships can be many-to-one or one-to-one.

[0098] In this embodiment, the light intensity can be measured using a light intensity meter or a photometer. Too strong or too weak light may cause the image quality to deteriorate, thereby affecting the quality of point cloud generation. At the same time, too strong light may cause overexposure, resulting in loss of image details, while too weak light may make the image too dim and difficult to identify the outline of the object. The light angle can be marked as perpendicular to the camera direction as the reference light angle when taking pictures, and measured using a protractor or a device with an angle measurement function (such as a laser rangefinder). An inappropriate light angle may cause shadows and reflections in the image to interfere with object recognition, thereby affecting the quality of point cloud generation. The rate of light change can be detected using a dedicated light change monitoring device such as a light intensity meter. When quickly reconstructing an object, light changes help to better capture and separate the details and texture features of the object, but the stability of the shooting environment needs to be ensured to avoid the impact of too fast light changes on image and point cloud generation. The three factors work together. In an ideal shooting environment, these three factors are coordinated: moderate light intensity, reasonable light angle, and gentle light changes. When one factor is unbalanced (for example, too high light intensity or inappropriate light angle), it may affect the image quality, resulting in image distortion or loss of details. The shooting environment evaluation value obtained by comprehensive analysis is used to reflect the overall quality of the image in terms of lighting. At the same time, it can be judged whether the image meets specific lighting requirements based on the size of the evaluation value. When the shooting environment evaluation value of an initial image is less than the shooting environment evaluation threshold preset in the object reconstruction database, the lighting conditions can be adjusted in real time by introducing a dynamic lighting compensation mechanism. For example, the automatic exposure function can be used to adjust the image exposure time, or the automatic white balance can be used to correct color and lighting changes; the image brightness can also be increased or decreased by adjusting the exposure compensation, so that the image is clearer and more visible in a low-light environment; when there is too much noise, the ISO setting of the camera can be increased to balance the image brightness and noise; when the light is strong, the shutter speed can be increased to reduce the light exposure and avoid overexposure.

[0099] Furthermore, the step of obtaining an image acquisition quality evaluation value of each initial image by comprehensively analyzing the scanning device performance evaluation value and the shooting environment evaluation value of each initial image includes: obtaining a critical scanning device performance evaluation value and a critical shooting environment evaluation value from an object reconstruction database; and obtaining an image acquisition quality evaluation value of each initial image by comprehensively analyzing the scanning device performance evaluation value, the shooting environment evaluation value of each initial image, the critical scanning device performance evaluation value and the critical shooting environment evaluation value.

[0100] In this embodiment, the scanning device performance evaluation value provides a hardware upper limit for image acquisition, while the shooting environment evaluation value describes the lighting conditions during actual image acquisition. The relationship between the two is reflected in that the performance of the scanning device determines the potential quality of the image, while the shooting environment evaluation value reflects the actual influencing factors during the image acquisition process. The image acquisition quality evaluation value obtained through comprehensive analysis can objectively reflect the overall quality level of the image and provide data support for device tuning and optimization of the acquisition environment. This comprehensive analysis helps to improve the stability and quality of image acquisition and ensure that the final image meets the requirements.

[0101] Furthermore, the image acquisition quality evaluation value of each initial image is obtained as follows:

[0102] ;

[0103] In the formula, ζQ K represents the image acquisition quality assessment value of the kth initial image, α 7 represents the image acquisition quality assessment impact factor corresponding to the scanning device performance evaluation value, α 8 represents the image acquisition quality assessment impact factor corresponding to the shooting environment assessment value, ζP 1 represents the performance evaluation value of the scanning device, ζP 0 represents the critical scanning equipment performance evaluation value, ζL 1K represents the shooting environment evaluation value of the kth initial image, ζL 0 Indicates the critical shooting environment evaluation value.

[0104] In this embodiment, α 7 and α 8The image acquisition quality assessment impact factors corresponding to the scanning device performance assessment value and the shooting environment assessment value preset in the object reconstruction database respectively represent the numerical values ​​of the influence degree of the scanning device performance assessment value and the shooting environment assessment value on the image acquisition quality assessment value, and can be directly obtained from the object reconstruction database when used. These relationships are organized into a "mapping set", that is, a lookup table or a corresponding rule. When the value of a certain data is actually monitored, the value can be searched according to the corresponding mapping set, and then the corresponding image acquisition quality assessment impact factor can be obtained. This impact factor is a number between 0 and 1, which represents the evaluation result of the influence degree of the current data on the image acquisition quality assessment value. For example, for the scanning device performance assessment value, there is a scanning device performance assessment value and an image acquisition quality assessment impact factor corresponding to the scanning device performance assessment value as a mapping set. By inputting the value of the scanning device performance assessment value, the image acquisition quality assessment impact factor corresponding to the scanning device performance assessment value can be obtained; for the shooting environment assessment value, there is a shooting environment assessment value and an image acquisition quality assessment impact factor corresponding to the shooting environment assessment value as a mapping set. By inputting the value of the shooting environment assessment value, the image acquisition quality assessment impact factor corresponding to the shooting environment assessment value can be obtained. These mapping relationships can be many-to-one or one-to-one.

[0105] Furthermore, the initial image set is quality screened according to the image acquisition quality assessment value of each initial image to obtain the step of optimizing the image data set, including: obtaining the image acquisition quality assessment threshold from the object reconstruction database; comparing the image acquisition quality assessment value of each initial image with the image acquisition quality assessment threshold; if the image acquisition quality assessment value of an initial image is greater than or equal to the image acquisition quality assessment threshold, marking the initial image as a qualified image; if the image acquisition quality assessment value of an initial image is less than the image acquisition quality assessment threshold, marking the initial image as an unqualified image, and reacquiring the initial image of the angle; and counting the qualified images to obtain the optimized image data set.

[0106] In this embodiment, the image acquisition quality assessment threshold is a standard value used to distinguish qualified and unqualified images. By comparing the image acquisition quality assessment value of each initial image with the image acquisition quality assessment threshold, the image quality in the obtained optimized image data set is higher, which helps to improve the quality of the final reconstructed model. Low-quality images may cause errors or omissions in the reconstruction process, affecting the accuracy of object reconstruction. In object reconstruction or computer vision applications, processing low-quality images often requires more computing resources and may lead to increased errors. By screening out qualified images, redundant data can be reduced and computing efficiency can be improved.

[0107] like Figure 5As shown, it is a structural schematic diagram of a system for rapidly reconstructing an object based on 3D Gaussian sputtering segmentation provided in an embodiment of the present invention. The system for rapidly reconstructing an object based on 3D Gaussian sputtering segmentation provided in an embodiment of the present application comprises: a data acquisition module, an image screening module, a voxel processing module, and an object reconstruction module; wherein the data acquisition module is used to acquire an appearance image of a target object, acquire image acquisition quality data, and extract 3D reconstruction information, wherein the 3D reconstruction information comprises an initial image set, camera pose information, a sparse point cloud, and an accurate mask of the target viewing angle; the image screening module is used to perform quality screening on an optimized image data set according to the image acquisition quality data to obtain an optimized image data set; the voxel processing module is used to convert the sparse point cloud into a pixel mask; The cloud is reconstructed into voxels, the optimized image data set is preliminarily cropped to obtain the 2D image area under each viewing angle projection, voxel cropping is performed according to the camera posture information and the 2D image area under each viewing angle projection to obtain the cropped voxels of each viewing angle, and the target voxels cropped from each viewing angle are obtained by cropping according to the continuity of the viewing angle, and the target voxels cropped from each viewing angle are rendered as a depth map; each sampling point is generated according to the depth map, the optimized image data set is segmented according to each sampling point, a target object mask is generated, and the target object mask is voxel cropped to obtain the segmentation result of each viewing angle; the object reconstruction module is used to project the current voxel of the segmentation result back to the initial prompt viewing angle to generate a depth map, and the depth map is obtained by using the precise mask of the target viewing angle, and the depth map is obtained according to the D 1′ and D 1 A consistency evaluation value is obtained through comprehensive analysis, and the segmentation results are iterated according to the consistency evaluation value to obtain a target object mask set; the optimized image data set is masked according to the target object mask set to obtain an RGB image of the target object, and the target object is reconstructed based on the voxels cropped from each perspective to obtain an initial 3D Gaussian sputtering model; the initial 3D Gaussian sputtering model is initially restored according to the depth map of the original scene voxels and the depth map of the target object voxels to obtain a restored 3D Gaussian sputtering model; and the Gaussian points of the restored 3D Gaussian sputtering model are filtered according to the voxels of the target object to obtain a standard 3D Gaussian sputtering model.

[0108] An embodiment of the present application also provides a device for rapid reconstruction of objects based on 3D Gaussian sputtering segmentation, comprising: a processor and a memory for storing instructions executable by the processor; when the processor is configured to execute the instructions, the device for rapid reconstruction of objects based on 3D Gaussian sputtering segmentation implements a method for rapid reconstruction of objects based on 3D Gaussian sputtering segmentation.

[0109] In summary, the embodiment of the present application acquires image acquisition quality data and extracts 3D reconstruction information by acquiring an appearance image of a target object, wherein the 3D reconstruction information includes an initial image set, camera pose information, a sparse point cloud, and an accurate mask of a target view; performs quality screening on an optimized image data set according to the image acquisition quality data to obtain an optimized image data set; reconstructs the sparse point cloud into voxels, performs preliminary cropping on the optimized image data set to obtain a 2D image area under projection of each view angle, performs voxel cropping on the optimized image data set according to the camera pose information and the 2D image area under projection of each view angle to obtain cropped voxels of each view angle, obtains cropped target voxels of each view angle through cropping according to view angle continuity, and renders cropped target voxels of each view angle into a depth map; generates sampling points according to the depth map, segments the optimized image data set according to the sampling points, generates a target object mask, and performs voxel cropping on the target object mask to obtain segmentation results of each view angle; projects the current voxel of the segmentation result back to the initial prompt view angle to generate a depth map, obtains the depth map using the accurate mask of the target view angle, and renders the cropped target voxels of each view angle into a depth map according to the D 1′ and D 1 A consistency evaluation value is obtained through comprehensive analysis, and the segmentation results are iterated according to the consistency evaluation value to obtain a target object mask set; the optimized image data set is masked according to the target object mask set to obtain an RGB image of the target object, and the target object is reconstructed based on the voxels cropped from each perspective to obtain an initial 3D Gaussian sputtering model; the initial 3D Gaussian sputtering model is initially restored according to the depth map of the original scene voxels and the depth map of the target object voxels to obtain a restored 3D Gaussian sputtering model; the Gaussian points of the restored 3D Gaussian sputtering model are filtered according to the voxels of the target object to obtain a standard 3D Gaussian sputtering model, realizing a complete process from image acquisition to three-dimensional reconstruction, and then to image processing and 3D model generation.

[0110] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0111] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0112] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0113] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0114] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0115] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A method for fast object reconstruction based on 3D Gaussian sputtering segmentation, characterized in that: The following steps are involved: Capture the appearance image of the target object, obtain the image acquisition quality data and extract the 3D reconstruction information, where the 3D reconstruction information includes the initial image set, camera pose information, sparse point cloud and accurate mask of the target view; Performing quality screening on the initial image set according to the image acquisition quality data to obtain an optimized image data set; Reconstruct the sparse point cloud into voxels, perform preliminary cropping on the optimized image data set to obtain the 2D image area under each viewing angle projection, perform voxel cropping based on the camera pose information and the 2D image area under each viewing angle projection to obtain the cropped voxels at each viewing angle, crop the cropped target voxels at each viewing angle based on the viewing angle continuity, and render the cropped target voxels at each viewing angle into a depth map; Generate sampling points according to the depth map, segment the optimized image data set according to the sampling points, generate a target object mask, and perform voxel cropping on the target object mask to obtain the segmentation results of each perspective; Project the current voxel of the segmentation result back to the initial prompt view to generate a depth map D 1′ , using the accurate mask of the target perspective to obtain the depth map D1, according to D 1′ The consistency evaluation value is obtained by comprehensive analysis with D1, and the segmentation result is iterated according to the consistency evaluation value to obtain the target object mask set; The optimized image data set is masked according to the target object mask set to obtain an RGB image of the target object, and the target object is reconstructed according to the cropped voxels of each viewpoint to obtain an initial 3D Gaussian sputtering model. The initial 3D Gaussian sputtering model is initially restored according to the depth map of the original scene voxels and the depth map of the target object voxels to obtain a restored 3D Gaussian sputtering model, and the Gaussian points of the restored 3D Gaussian sputtering model are filtered according to the voxels of the target object to obtain a standard 3D Gaussian sputtering model.

2. The method for rapid object reconstruction based on 3D Gaussian sputtering segmentation according to claim 1, characterized in that: The step of performing quality screening on the initial image set according to the image acquisition quality data to obtain an optimized image data set comprises: The image acquisition quality data includes scanning equipment performance data and shooting environment data; Obtain a scanning device performance evaluation value based on comprehensive analysis of scanning device performance data, and provide evaluation feedback based on the scanning device performance evaluation value; Obtaining a shooting environment evaluation value of each initial image based on a comprehensive analysis of the shooting environment data; An image acquisition quality evaluation value of each initial image is obtained by comprehensive analysis based on the scanning equipment performance evaluation value and the shooting environment evaluation value of each initial image, and the initial image set is quality screened based on the image acquisition quality evaluation value of each initial image to obtain an optimized image data set.

3. The method for rapid object reconstruction based on 3D Gaussian sputtering segmentation according to claim 2, characterized in that: The step of obtaining a scanning device performance evaluation value based on comprehensive analysis of the scanning device performance data comprises: The scanning device performance data, including scanning time, scanning resolution and scanning frequency; Obtaining a reference scanning time, an allowable deviation scanning time, a critical scanning resolution, a reference scanning frequency, and an allowable deviation scanning frequency from an object reconstruction database; The performance evaluation value of the scanning equipment is obtained through comprehensive analysis.

4. The method for rapid object reconstruction based on 3D Gaussian sputtering segmentation according to claim 2, characterized in that: The step of providing evaluation feedback according to the performance evaluation value of the scanning device comprises: Obtaining a scanning device performance evaluation threshold from an object reconstruction database; The scanning device performance evaluation value is compared with the scanning device performance evaluation threshold. If the scanning device performance evaluation value is greater than or equal to the scanning device performance evaluation threshold, no additional processing is performed. If the scanning device performance evaluation value is less than the scanning device performance evaluation threshold, adjustment feedback is provided to the scanning device.

5. The method for rapid object reconstruction based on 3D Gaussian sputtering segmentation according to claim 2, characterized in that: The step of obtaining the shooting environment evaluation value of each initial image according to the comprehensive analysis of the shooting environment data comprises: The shooting environment data includes light intensity, light angle and light change rate; Obtain reference illumination intensity, allowable deviation illumination intensity, reference illumination angle, allowable deviation illumination angle and critical illumination change rate from an object reconstruction database; The shooting environment evaluation value of each initial image is obtained through comprehensive analysis.

6. The method for rapid object reconstruction based on 3D Gaussian sputtering segmentation according to claim 2, characterized in that: The step of obtaining the image acquisition quality evaluation value of each initial image by comprehensive analysis based on the scanning device performance evaluation value and the shooting environment evaluation value of each initial image comprises: Obtaining a critical scanning device performance evaluation value and a critical shooting environment evaluation value from an object reconstruction database; The image acquisition quality evaluation value of each initial image is obtained by comprehensive analysis based on the scanning device performance evaluation value, the shooting environment evaluation value of each initial image, the critical scanning device performance evaluation value and the critical shooting environment evaluation value.

7. The method for rapid object reconstruction based on 3D Gaussian sputtering segmentation according to claim 6, characterized in that: The image acquisition quality evaluation value of each initial image is obtained as follows: ; In the formula, ζQ k represents the image acquisition quality evaluation value of the kth initial image, e is a natural constant, α7 represents the image acquisition quality evaluation influencing factor corresponding to the scanning device performance evaluation value, α8 represents the image acquisition quality evaluation influencing factor corresponding to the shooting environment evaluation value, ζP1 represents the scanning device performance evaluation value, ζP0 represents the critical scanning device performance evaluation value, ζL 1k represents the shooting environment evaluation value of the kth initial image, and ζL0 represents the critical shooting environment evaluation value.

8. The method for rapid object reconstruction based on 3D Gaussian sputtering segmentation according to claim 2, characterized in that: The step of performing quality screening on the initial image set according to the image acquisition quality evaluation value of each initial image to obtain an optimized image data set comprises: Obtaining an image acquisition quality assessment threshold from an object reconstruction database; The image acquisition quality assessment value of each initial image is compared with the image acquisition quality assessment threshold. If the image acquisition quality assessment value of an initial image is greater than or equal to the image acquisition quality assessment threshold, the initial image is marked as a qualified image. If the image acquisition quality assessment value of an initial image is less than the image acquisition quality assessment threshold, the initial image is marked as an unqualified image, and the initial image of the angle is reacquired. The qualified images are counted to obtain the optimized image data set.

9. A system for rapid object reconstruction based on 3D Gaussian sputtering segmentation, applying the method for rapid object reconstruction based on 3D Gaussian sputtering segmentation according to any one of claims 1 to 8, characterized in that: It includes a data acquisition module, an image screening module, a voxel processing module, an object reconstruction module and an object reconstruction database; The data acquisition module is used to acquire the appearance image of the target object, obtain the image acquisition quality data and extract the three-dimensional reconstruction information, wherein the three-dimensional reconstruction information includes the initial image set, the camera pose information, the sparse point cloud and the accurate mask of the target perspective; The image screening module is used to perform quality screening on the initial image set according to the image acquisition quality data to obtain an optimized image data set; The voxel processing module is used to reconstruct the sparse point cloud into voxels, perform preliminary cropping on the optimized image data set to obtain the 2D image area under each viewing angle projection, perform voxel cropping according to the camera posture information and the 2D image area under each viewing angle projection to obtain the cropped voxels of each viewing angle, crop the target voxels after cropping according to the viewing angle continuity, and render the cropped target voxels of each viewing angle into a depth map; generate each sampling point according to the depth map, segment the optimized image data set according to each sampling point, generate a target object mask, and perform voxel cropping on the target object mask to obtain the segmentation result of each viewing angle; The object reconstruction module is used to project the current voxel of the segmentation result back to the initial prompt view to generate a depth map, and obtain the depth map using the accurate mask of the target view. 1′ A consistency evaluation value is obtained by comprehensive analysis with D1, and the segmentation result is iterated according to the consistency evaluation value to obtain a target object mask set; the optimized image dataset is masked according to the target object mask set to obtain an RGB image of the target object, and the target object is reconstructed according to the voxels cropped from each perspective to obtain an initial 3D Gaussian sputtering model; the initial 3D Gaussian sputtering model is initially restored according to the depth map of the original scene voxels and the depth map of the target object voxels to obtain a restored 3D Gaussian sputtering model; the Gaussian points of the restored 3D Gaussian sputtering model are filtered according to the voxels of the target object to obtain a standard 3D Gaussian sputtering model.

10. A device for rapid object reconstruction based on 3D Gaussian sputtering segmentation, characterized in that: include: A processor and a memory, wherein the memory is used to store instructions executable by the processor; when the processor is configured to execute the instructions, the object rapid reconstruction device based on 3D Gaussian sputtering segmentation implements the object rapid reconstruction method based on 3D Gaussian sputtering segmentation as described in any one of claims 1-8.

Citation Information

Patent Citations

  • A method for rapid 3D reconstruction of thin objects

    CN109360267B

  • A fast 3D object reconstruction method based on camera array

    CN110288642B

  • On-line dense reconstruction method and system based on RGB-D sensor

    CN117710469A

  • Method for recovering three-dimensional human body appearance from single image in real time

    CN118521711A