Method, system, device and medium for classification and reconstruction based on out-of-field data
Through multi-resolution three-plane network architecture and posture refinement and compensation code optimization technology, the problems of inaccurate posture labeling and shape distortion in the real data set are solved, and high-quality category reconstruction is achieved.
Patent Information
- Application Number
- CN202311048665.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-18
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2043-08-18
AI Technical Summary
The prior art has difficulty handling real data sets without accurate posture and mask annotations, resulting in low reconstruction quality and inability to achieve high-quality category reconstruction.
A multi-resolution three-plane network architecture is adopted, combining posture refinement and compensation code based on symbol distance function, roughly labeled data sets are optimized to achieve high-quality category reconstruction of real data sets.
High-quality category reconstruction of real datasets without accurate poses and mask annotations is achieved, which is superior to other category reconstruction schemes, and is able to robustly handle noise and distortion in wild datasets.
Smart Images

Figure CN117078854B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of three-dimensional reconstruction technology, and particularly to a method, system, device and medium for classification reconstruction based on in-the-wild data. Background Art
[0002] Multi-view reconstruction of real-world objects is crucial for various applications, such as autonomous driving and virtual reality. However, inaccurate annotations, camera parameter deviations, and lens distortions in real-world data often pose significant challenges to the reconstruction quality. Synthetic datasets provide accurate annotations and pose more manageable reconstruction problems, but they lack the realistic style of real-world objects. In addition, most existing classification reconstruction methods can only handle synthetic data, where the training network reconstructs all instances of an object class. Summary of the Invention
[0003] The embodiments of the present application provide a method, system, device and medium for classification reconstruction based on in-the-wild data, which can process real datasets without accurate pose and mask annotations and achieve high-quality class reconstruction.
[0004] To solve the above technical problems, in a first aspect, the embodiments of the present application provide a method for classification reconstruction based on in-the-wild data, including the following steps: First, obtain an in-the-wild dataset; the in-the-wild dataset includes multiple instances, and each instance includes multiple images describing the instance; then, perform annotation processing on the in-the-wild dataset to obtain a roughly annotated dataset; next, construct a classification reconstruction network, and based on the classification reconstruction network, perform optimization processing on the roughly annotated dataset; the classification reconstruction network is a multi-resolution three-plane network architecture; finally, train the classification reconstruction network to reconstruct all instances of the object class in the in-the-wild dataset.
[0005] In some exemplary embodiments, performing optimization processing on the roughly annotated dataset includes: performing pose optimization processing on the roughly annotated dataset; performing compensation code optimization processing on the roughly annotated dataset.
[0006] In some exemplary embodiments, performing pose optimization processing on the roughly annotated dataset includes: respectively using a translator, a rotator, and a scaler to transform the sampling coordinates of the roughly annotated dataset, and optimizing the rough annotation to obtain correct coordinates.
[0007] In some exemplary embodiments, performing compensation code optimization processing on the roughly annotated dataset includes: splicing the feature vector with a compensation code based on a signed distance function to obtain spliced data; using the spliced data as the input of a signed distance function decoder to hide the distortion information of the current view and complete classification reconstruction; using denotes the compensation code; where i and j denote the instance and the ground truth image respectively.
[0008] In some exemplary embodiments, the classification and reconstruction network includes a feature modulation module and a spanning module, and the feature modulation module is integrated into the spanning module; the spanning module is expressed as in formula (1):
[0009] G = {G (f) , G (m) , G (b)} (1)
[0010] where z i denotes the latent code of the i-th instance style, and denotes the learnable content at the k-th resolution level. It is initialized by uniformly sampling within the coordinate system range. Θ k denotes the explicit feature at the k-th level. Θ k is obtained by inputting into the spanning module and modulating it with z i . It is aligned with the orthogonal feature planes aligned along three axes; where r k is the spatial resolution and η k is the number of feature channels.
[0011] In some exemplary embodiments, the process of obtaining Θ k is shown in formula (2):
[0012]
[0013] where is the upsampled feature; and W k-1 is the learnable coefficient.
[0014] In some exemplary embodiments, the classification and reconstruction network is trained by minimizing the training formula; the training formula is shown in formula (3):
[0015] l all = l color + 0.005l eikonal + 0.003l sparisty (3)
[0016] where l color is the color loss function, C is the ground truth of the corresponding instance; l eikonal is the regularization function, l sparisty is the sparse regularization term function, and l sparisty= ∑ x f nld (s(x)); where and γ reg are constant scaling parameters.
[0017] In a second aspect, an embodiment of the present application further provides a classification and reconstruction system based on in-the-wild data, including: a dataset acquisition module, a rough annotation processing module, a classification and reconstruction network construction module, an optimization processing module, and a classification and reconstruction network training module connected in sequence; the dataset acquisition module is used to acquire an in-the-wild dataset; the in-the-wild dataset includes multiple instances, and each instance includes multiple images describing the instance; the rough annotation processing module is used to perform annotation processing on the in-the-wild dataset to obtain a roughly annotated dataset; the classification and reconstruction network construction module is used to construct a classification and reconstruction network; the optimization processing module is used to perform optimization processing on the roughly annotated dataset according to the classification and reconstruction network; the classification and reconstruction network is a multi-resolution three-plane network architecture; the classification and reconstruction network training module is used to train the classification and reconstruction network to reconstruct all instances of the object classes in the in-the-wild dataset.
[0018] In addition, the present application also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above method for classification and reconstruction based on in-the-wild data.
[0019] In addition, the present application also provides a computer-readable storage medium storing a computer program, and the computer program realizes the above method for classification and reconstruction based on in-the-wild data when executed by a processor.
[0020] The technical solution provided by the embodiment of the present application has at least the following advantages:
[0021] The embodiment of the present application provides a method, a system, a device, and a medium for classification and reconstruction based on in-the-wild data. The method includes the following steps: First, acquire an in-the-wild dataset; the in-the-wild dataset includes multiple instances, and each instance includes multiple images describing the instance; then, perform annotation processing on the in-the-wild dataset to obtain a roughly annotated dataset; next, construct a classification and reconstruction network, and perform optimization processing on the roughly annotated dataset based on the classification and reconstruction network; the classification and reconstruction network is a multi-resolution three-plane network architecture; finally, train the classification and reconstruction network to reconstruct all instances of the object classes in the in-the-wild dataset.
[0022] The present application provides a method for classification and reconstruction based on in-the-wild data, which can process real datasets without accurate pose and mask annotations. The present application adopts a multi-resolution triplane network structure, thus achieving high-quality class reconstruction for real datasets, which is superior to other class reconstruction schemes. In class reconstruction, the present application introduces pose refinement and a compensation code based on the signed distance function as effective methods to solve inaccurate pose annotation and shape distortion. For in-the-wild datasets (i.e., without annotations or with incorrect annotations), the present application adopts a pose refinement method to optimize rough annotations and can robustly perform reconstruction. In class reconstruction, for instances with image distortion, the present application adopts a compensation code optimization method to hide the distortion information of the current view and complete compatible reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not constitute limitations on the embodiments. Unless otherwise stated, the figures in the drawings do not constitute a scale limitation.
[0024] Figure 1 It is a schematic flowchart of a method for classification and reconstruction based on in-the-wild data provided by an embodiment of the present application;
[0025] Figure 2 It is a schematic flowchart of a method for classification and reconstruction based on in-the-wild data provided by another embodiment of the present application;
[0026] Figure 3 It is a schematic diagram of the reconstruction result of a method for classification and reconstruction based on in-the-wild data provided by an embodiment of the present application;
[0027] Figure 4 It is a schematic structural diagram of a system for classification and reconstruction based on in-the-wild data provided by an embodiment of the present application;
[0028] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] As can be seen from the background art, the existing classification and reconstruction methods have the technical problem that they can only process synthetic data and cannot perform classification and reconstruction on in-the-wild data.
[0030] In recent years, great attention has been paid to three-dimensional reconstruction of a single scene or object using neural network methods. A prominent example is the Neural Radiance Field (NeRF) method, which reconstructs 3D scenes from a set of 2D images by training a deep neural network. NeRF can generate high-fidelity, realistic images and can render scenes from any angle. NeuS is another surface rendering technique that projects 3D objects onto 2D images using a differentiable surface rendering method and then compares it with the input images for supervision to obtain an accurate object surface. One of its cores is the use of the signed distance function, that is, calculating the distance of points in space from the surface, and finally extracting the zero-value surface to represent the surface. There is also a related technology Instance-ngp, which divides the space into different resolutions and uses different networks to fit each resolution to achieve higher quality and details in single-scene reconstruction. However, these methods are only applicable to the representation of a single scene or object and cannot be generalized to the category reconstruction of multiple instances.
[0031] In addition, EG3D proposes to use the tri-planar method to represent the entire space. This is mainly because when the resolution in space is high, it will consume huge resources to represent these space points. The tri-planar means that the network only directly represents three two-dimensional planes: xoy, yoz, and xoz. Each plane interpolates with each other to indirectly represent the points of the entire three-dimensional space. This will greatly reduce the number of network parameters and improve the resolution number of network representation.
[0032] In the prior art, there has been little research on methods for category reconstruction. SRN uses a multi-layer fully connected network to complete simple category reconstruction of synthetic data. CodeNeRF decouples the appearance color and object shape and represents them with different hidden layer codes, completing the category reconstruction of synthetic data and improving controllability and editability.
[0033] The biggest drawback of the current prior art is that it cannot simultaneously meet the following capabilities: (1) Category reconstruction, that is, a network reconstructs multiple instances of a class of objects. (2) High-quality reconstruction of real-world datasets. Real-world data has a wider color distribution, greater shape diversity, and more noise than synthetic data, so the fitting difficulty is much greater. But it is also more important. (3) Handling annotation biases, image noise, and lens distortion in real data.
[0034] Therefore, to solve the above technical problems, the embodiments of the present application provide a method for classification and reconstruction based on in-the-wild data, including the following steps: First, obtain an in-the-wild data set; the in-the-wild data set includes multiple instances, and each instance includes multiple images describing the instance; then, perform annotation processing on the in-the-wild data set to obtain a roughly annotated data set; next, construct a classification and reconstruction network, and based on the classification and reconstruction network, perform optimization processing on the roughly annotated data set; the classification and reconstruction network is a multi-resolution tri-plane network architecture; finally, train the classification and reconstruction network to reconstruct all instances of the object classes in the in-the-wild data set. The embodiments of the present application provide a method for classification and reconstruction based on in-the-wild data, which can simultaneously meet the above three requirements, that is, it can process real data sets without accurate pose and mask annotations and achieve high-quality class reconstruction.
[0035] The following will elaborate on each embodiment of the present application with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in each embodiment of the present application, many technical details are presented for the reader to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented.
[0036] See Figure 1 , the embodiments of the present application provide a method for classification and reconstruction based on in-the-wild data, including the following steps:
[0037] Step S1, obtain an in-the-wild data set; the in-the-wild data set includes multiple instances, and each instance includes multiple images describing the instance.
[0038] Step S2, perform annotation processing on the in-the-wild data set to obtain a roughly annotated data set.
[0039] Step S3, construct a classification and reconstruction network, and based on the classification and reconstruction network, perform optimization processing on the roughly annotated data set; the classification and reconstruction network is a multi-resolution tri-plane network architecture.
[0040] Step S4, train the classification and reconstruction network to reconstruct all instances of the object classes in the in-the-wild data set.
[0041] It should be noted that the in-the-wild data in the in-the-wild data set in step S1 includes real data with uncontrollable factors such as annotation deviation, image noise, and lens distortion. These data are called in-the-wild data.
[0042] This application proposes a novel method to address the challenge of classifying and reconstructing real-world objects under field conditions. This application utilizes camera pose optimization techniques and compensation codes based on the Signed Distance Function (SDF) to handle inaccurate annotations and uncontrolled camera settings. Additionally, this application introduces a multi-resolution tri-plane network architecture to achieve accurate classification and reconstruction.
[0043] To evaluate the method for classification and reconstruction based on in-the-wild data provided by this application, the method of this application is evaluated on the Multi-View Major Car (MVMC) dataset, which contains a limited number of real-world car instances with inaccurate annotations. The experimental results demonstrate the effectiveness of the method provided by this application in reconstructing objects under difficult conditions. The method of this application can process real-world datasets without accurate pose and mask annotations and achieve accurate classification and reconstruction with state-of-the-art performance.
[0044] From the perspective of input and output, the input of the method of this application is an image dataset of a category, which can be a car, a chair, etc. For example, for cars, the dataset contains different types of cars, that is, multiple instances, and there are multiple images describing each instance. Note that the images in this dataset may or may not have camera parameter annotations. The output of the method of this application is the 3D reconstruction result of each instance of the dataset.
[0045] In step S1, an in-the-wild dataset is obtained. By given an in-the-wild classification multi-instance multi-view dataset, which contains N ins instances, and each instance contains N img images, this application denotes the dataset as where I i,j represents the j-th image of the i-th instance . The method proposed by this application includes two stages: data processing and classification and reconstruction network. The former processes problems such as noisy data, inaccurate annotations, and distortion. This application adopts a multi-resolution strategy in the network architecture, where six levels of resolution increase gradually as follows: r = [16, 32, 64, 128, 256, 512]. For any 3D position input of the i-th instance this application sets the size of the output feature vector v k to be η k at each level, k ∈ {1, 2, 3, 4, 5, 6}.
[0046] In some embodiments, in step S2, the roughly annotated dataset is optimized, including:
[0047] Step S201: Optimize the pose of the roughly annotated dataset.
[0048] Step S202: Perform compensation code optimization processing on the roughly annotated dataset.
[0049] In some embodiments, step S201 performs pose optimization processing on the roughly annotated dataset, including: separately using a translator, a rotator, and a scaler to transform the sampling coordinates of the roughly annotated dataset, and optimizing the rough annotation to obtain correct coordinates.
[0050] To process the raw data containing inaccurate annotations and pose estimations, the present application utilizes a shape pose method that estimates the initial camera pose of an object and then refines it as needed. The present application further utilizes the advanced segmentation model Mask2Former to segment the objects in the image and obtain a foreground mask. At this time, the camera pose still belongs to the rough annotation. Next, the present application will introduce how to further solve the inaccuracy problem: camera pose refinement.
[0051] Given that the pose annotations obtained at this stage may be biased, the present application introduces a learning-based pose refinement method to optimize the pose estimation during the reconstruction process. The present application formulates this process as follows. Since the camera pose in image I i,j may be biased, the sampling coordinates do not correspond to the correct image coordinates. To solve this problem, the present application uses a translator, a rotator, and a scaler to perform a transformation to obtain the correct coordinates Specifically, the present application applies a general SE(3) transformation to translation and rotation, and scales each coordinate by the scaling factor of the scaler. In addition, the present application parameterizes the transformation coefficients. Since the pose deviations of each image are inconsistent, the degree of transformation needs to be controlled separately. Note that the coordinate transformation coefficients vary with different images. During training, all query point coordinates are processed after transformation, while in the inference phase they remain untransformed.
[0052] In some embodiments, step S202 performs compensation code optimization processing on the roughly annotated dataset, including: concatenating the feature vector with the compensation code based on the signed distance function to obtain concatenated data; using the concatenated data as the input of the signed distance function decoder to hide the distortion information of the current view and complete classification reconstruction; using to represent the compensation code; where i and j respectively represent the instance and the ground truth image.
[0053] To solve the problem of inconsistent shapes of different views in the dataset, the present application introduces a learnable compensation code corresponding to the j-th ground truth image of the i-th instance to explain this variation. The calculation formula for the signed distance function value of x is s = D(v, β i,j) Among them, D is the signed distance function decoder network. Specifically, in this application, the code β i,j and the feature vector v are concatenated as the input of the signed distance function decoder network D.
[0054] In some embodiments, the classification and reconstruction network includes a feature modulation module and a spanning module, and the feature modulation module is integrated into the spanning module; the spanning module is represented by formula (1):
[0055] G = {G (f) , G (m) , G (b)} (1)
[0056] Among them, z i is used to represent the latent code of the i-th instance style, and is used to represent the learnable content at the k-th resolution level. It is initialized by uniformly sampling from within the coordinate system range. Θ k is used to represent the explicit feature at the k-th level. Θ k is obtained by inputting into the spanning module and modulating it with z i . It is aligned with the orthogonal feature planes aligned along three axes; among them, r k is the spatial resolution, and η k is the number of feature channels.
[0057] This application adopts a tri-planar representation structure at each resolution level. This application integrates the feature modulation module into the spanning module G = {G (f) , G (m) , G (b)}, and uses to represent the latent code of the i-th instance style. And is the learnable content at the k-th resolution level, which is initialized by uniformly sampling from within the coordinate system range.
[0058] In the tri-planar formula, the explicit feature at the k-th level is aligned with the orthogonal feature planes aligned along three axes.
[0059] To obtain Θ k , this application inputs into the module G k and modulates it with the code z i .
[0060] This application applies skip connections to the middle layer of the network G k starting from the second layer. For example, at the second layer (k = 2), this application extracts the intermediate features Upsample them starting from the first level. Then, the present application multiplies the upsampled features by the learnable coefficient w 1 and adds them to the features at the current level.
[0061] In some embodiments, the process of obtaining Θk is shown in Equation (2):
[0062]
[0063] where is the upsampled feature; w k-1 is the learnable coefficient.
[0064] Given a 3D position the present application first projects it onto each feature plane in Θ k and retrieves the corresponding feature vectors through bilinear interpolation. The present application calculates the feature vector v k by adding these three vectors. Finally, the features v k from each level are concatenated together to create the final feature vector v.
[0065] To address the problem of inconsistent view shapes in the dataset, the present application introduces a learnable compensation code corresponding to the j-th ground truth image of the i-th instance to account for this variation. The present application concatenates the code β i,j and the feature vector v as the input to the signed distance function decoder network D.
[0066] For in-the-wild data, the number of available views in the ground truth is usually limited, which may lead to incomplete shape information when targeting new views. To address this issue, the present application proposes calculating the average of the shape compensation codes as follows:
[0067]
[0068] The present application utilizes all available views of the current instance by calculating the average of the shape compensation codes. This provides a comprehensive representation of the instance and allows for accurate reconstruction of novel views. The present application obtains the color value of x using the radiance network R. Specifically, the present application concatenates the feature vector v, the ray direction and the position x and inputs them into the radiance network R to obtain the RGB color value c = R(r d , x, v).
[0069] The above part introduced the data processing and class reconstruction network structure of the present application. Now, how to train will be introduced.
[0070] This application uses the volume rendering program in NeuS. Sample k points along the ray to calculate the approximate pixel color of the current ray as where is the cumulative transmittance, and the opacity value α i is calculated according to the signed distance function value s, that is, defined in the same way as NeuS.
[0071] This application uses a total of three loss functions: 1. Color loss function, C is the ground truth of the corresponding instance. To extract the foreground object, this application uses a segmentation map to mask the regions in the ground truth image that are irrelevant to the foreground. 2. Eikonal regularization, This application regularizes the signed distance function network D at the sampled points by adding an Eikonal term, where s is the calculated value of the signed distance function and x is the corresponding input coordinate point. 3. Sparse regularization term, l sparisty = ∑ x f nld (s(x)), where and γ reg is a constant scaling parameter. Due to the unconstrained height and usually inaccurate geometry in the occluded regions, this application applies a regularization term to penalize the learned density values. Specifically, this application calculates the density values based on the uniformly sampled signed distance function values for the purpose of penalty.
[0072] In some embodiments, the classification and reconstruction network is trained by minimizing the training formula; the training formula is shown in Formula (3):
[0073] l all = l color + 0.005l eikonal + 0.003l sparisty (3)
[0074] where, l color is the color loss function, C is the ground truth of the corresponding instance; l eikonal is the regularization function, l sparisty is the sparse regularization term function, l sparisty = ∑ x f nld (s(x)); where, and γ reg are constant scaling parameters.
[0075] Through experiments and simulations on the method provided by this application, the verification results meet the expectations. Figure 3 Shows the new view rendering effect reconstructed. As Figure 3As shown, through the method for classifying and reconstructing objects in the wild real world of the present application, high-quality class reconstruction for real datasets is achieved.
[0076] In summary, the present application provides a method for classifying and reconstructing objects in the wild real world, which can process real datasets without accurate pose and mask annotations. From wild data to completing the reconstruction of all instances of this class, this entire processing flow. The method provided by the present application uses a pose refinement method and a compensation code based on the signed distance function as effective methods to solve inaccurate pose annotation and shape distortion. In addition, the present application also provides a novel multi-resolution tri-plane network architecture, which can achieve high-quality class reconstruction.
[0077] Refer to Figure 4 , an embodiment of the present application further provides a classification and reconstruction system based on wild data, including: a dataset acquisition module 101, a rough annotation processing module 102, a classification and reconstruction network construction module 103, an optimization processing module 104, and a classification and reconstruction network training module 105 connected in sequence; the dataset acquisition module 101 is used to acquire a wild dataset; the wild dataset includes multiple instances, and each instance includes multiple images describing the instance; the rough annotation processing module 102 is used to perform annotation processing on the wild dataset to obtain a roughly annotated dataset; the classification and reconstruction network construction module 103 is used to construct a classification and reconstruction network; the optimization processing module 104 is used to optimize the roughly annotated dataset according to the classification and reconstruction network; the classification and reconstruction network is a multi-resolution tri-plane network architecture; the classification and reconstruction network training module 105 is used to train the classification and reconstruction network to reconstruct all instances of the object class in the wild dataset.
[0078] Specifically, the optimization processing module 104 includes a pose optimization processing unit 1041 and a compensation code optimization processing unit 1042. The pose optimization processing unit 1041 is used to perform pose optimization processing on the roughly annotated dataset. The compensation code optimization processing unit 1042 is used to perform compensation code optimization processing on the roughly annotated dataset.
[0079] Refer to Figure 5 , another embodiment of the present application provides an electronic device, including: at least one processor 110; and a memory 111 communicatively connected to the at least one processor; wherein, the memory 111 stores instructions executable by the at least one processor 110, and the instructions are executed by the at least one processor 110 to enable the at least one processor 110 to execute any of the above method embodiments.
[0080] Among them, the memory 111 and the processor 110 are connected in a bus manner. The bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors 110 and the memory 111 together. The bus may also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, etc., which are well known in the art, and thus will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver may be a single component or multiple components, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices over a transmission medium. The data processed by the processor 110 is transmitted over a wireless medium via an antenna. Further, the antenna also receives data and transmits the data to the processor 110.
[0081] The processor 110 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory 111 can be used to store data used by the processor 110 when executing operations.
[0082] Another embodiment of the present application relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the above method embodiments are implemented.
[0083] That is, those skilled in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by instructing relevant hardware through a program. This program is stored in a storage medium and includes several instructions to enable a device (which may be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the above methods of various embodiments of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical disks, etc., which can store program codes.
[0084] With the above technical solutions, the embodiments of the present application provide a method, system, device, and medium for classifying and reconstructing in-the-wild data. The method includes the following steps: First, obtain an in-the-wild data set; the in-the-wild data set includes multiple instances, and each instance includes multiple images describing the instance; then, perform annotation processing on the in-the-wild data set to obtain a roughly annotated data set; next, construct a classification and reconstruction network, and based on the classification and reconstruction network, perform optimization processing on the roughly annotated data set; the classification and reconstruction network is a multi-resolution three-plane network architecture; finally, train the classification and reconstruction network to reconstruct all instances of the object classes in the in-the-wild data set.
[0085] The present application provides a method for classification and reconstruction based on in-the-wild data, which can process real datasets without accurate pose and mask annotations. The present application adopts a multi-resolution tri-plane network structure, thus achieving high-quality class reconstruction for real datasets, which is superior to other class reconstruction schemes. In class reconstruction, the present application introduces pose refinement and a compensation code based on the signed distance function as effective methods to solve inaccurate pose annotation and shape distortion. For in-the-wild datasets (i.e., without annotations or with incorrect annotations), the present application adopts a pose refinement method to optimize the rough annotations and can perform reconstruction robustly. In class reconstruction, for instances with image distortion, the present application adopts a compensation code optimization method to hide the distortion information of the current perspective and complete compatible reconstruction.
[0086] Those of ordinary skill in the art can understand that the above embodiments are specific examples for implementing the present application. In actual applications, various changes can be made to them in form and details without departing from the spirit and scope of the present application. Any person skilled in the art can make their own changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be subject to the scope defined by the claims.
Claims
1. A method for classification and reconstruction based on in-the-wild data, characterized in that, Including: Obtain an in-the-wild dataset; the in-the-wild dataset includes multiple instances, and each instance includes multiple images describing the instance; Perform annotation processing on the in-the-wild dataset to obtain a roughly annotated dataset; Construct a classification and reconstruction network, and based on the classification and reconstruction network, perform optimization processing on the roughly annotated dataset; the classification and reconstruction network is a multi-resolution tri-plane network architecture; Train the classification and reconstruction network to reconstruct all instances of the object classes in the in-the-wild dataset; The classification and reconstruction network includes a feature modulation module and a spanning module, and the feature modulation module is integrated into the spanning module; The spanning module is represented by formula (1): (1) Among them, using z i represents the latent code of the i-th instance style, ; using represents the learnable content of the k-th resolution level, initialized by uniformly sampling within the coordinate system range, ; Adopt represents the k-th level explicit feature, is obtained by inputting into the said spanning module and modulating it with z i ; ; Orthogonal feature planes aligned along three axes; wherein, is the spatial resolution, is the number of feature channels.
2. The method for classification and reconstruction based on in-the-wild data according to claim 1, characterized in that, Performing optimization processing on the roughly annotated dataset includes: Performing pose optimization processing on the roughly annotated dataset; Performing compensation code optimization processing on the roughly annotated dataset.
3. The method for classification and reconstruction based on in-the-wild data according to claim 2, characterized in that, Performing pose optimization processing on the roughly annotated dataset includes: Respectively use a translator, a rotator, and a scaler to transform the sampling coordinates of the roughly annotated dataset, and optimize the rough annotation to obtain correct coordinates.
4. The method for classification and reconstruction based on in-the-wild data according to claim 2, characterized in that, Performing compensation code optimization processing on the roughly annotated dataset includes: Concatenate the feature vector with the compensation code based on the signed distance function to obtain concatenated data; Use the concatenated data as the input of the signed distance function decoder to hide the distortion information of the current view and complete classification and reconstruction; Adopt represent the compensation code; where i and j respectively represent the instance and the true value image.
5. The method for classification and reconstruction based on in-the-wild data according to claim 1, characterized in that, Obtain The process is shown in Equation (2) as follows: (2) Among them, is the upsampled feature; is the learnable coefficient.
6. The method for classification and reconstruction based on in-the-wild data according to claim 1, characterized in that, Train the classification and reconstruction network by minimizing the training formula; The training formula is shown in formula (3): (3) Among them, is the color loss function, , where C is the true value of the corresponding instance; is the approximate pixel color value of the current ray; is a regularization function, ; is a sparse regularization term function, ; where, and are constant scaling parameters; s is the calculated value of the signed distance function; x is the coordinate point corresponding to the input.
7. A system for classification and reconstruction based on in-the-wild data, characterized in that, Including: A dataset acquisition module, a rough annotation processing module, a classification and reconstruction network construction module, an optimization processing module, and a classification and reconstruction network training module connected in sequence; The dataset acquisition module is used to obtain an in-the-wild dataset; the in-the-wild dataset includes multiple instances, and each instance includes multiple images describing the instance; The rough annotation processing module is used to perform annotation processing on the in-the-wild dataset to obtain a roughly annotated dataset; The classification and reconstruction network construction module is used to construct a classification and reconstruction network; The optimization processing module is used to perform optimization processing on the roughly annotated dataset according to the classification and reconstruction network; the classification and reconstruction network is a multi-resolution tri-plane network architecture; The classification and reconstruction network training module is used to train the classification and reconstruction network to reconstruct all instances of the object classes in the in-the-wild dataset; The classification and reconstruction network includes a feature modulation module and a spanning module, and the feature modulation module is integrated into the spanning module; The spanning module is represented by formula (1): (1) Among them, z is used i to represent the latent code of the i-th instance style, ; is used to represent the learnable content at the k-th resolution level, which is initialized by uniformly sampling within the coordinate system range, ; Adopt represents the k-th level explicit feature, is obtained by inputting into the said spanning module and modulating it with z i ; ; Alignment of orthogonal feature planes aligned along three axes; wherein, is the spatial resolution, is the number of feature channels.
8. An electronic device, characterized in that, Including: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for classification and reconstruction based on in-the-wild data according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for classification and reconstruction based on in-the-wild data according to any one of claims 1 to 6.
Citation Information
Patent Citations
Weak light multi-view geometric reconstruction method based on deep learning
CN114332355A
Training method of three-dimensional scene reconstruction device for multi-camera system
CN115619928A