A layer-focused illustration single-view 3D reconstruction method and system
By using a layer-focused single-view 3D reconstruction method for illustrations, the single-view illustration is divided into feature and effect areas. High-quality 3D representations are generated using the layer-CLIP model and latent diffusion model. This solves the quality and efficiency problems of 2D illustration 3D reconstruction in existing technologies and achieves high-precision 3D generation.
Patent Information
- Application Number
- CN202411580766.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-06
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2044-11-06
AI Technical Summary
Existing technologies for 3D reconstruction of 2D illustrations suffer from problems such as color block stacking, large computational load, lack of understanding of lighting layers, and insufficient retrieval of specific accessory information, resulting in poor quality of generated 3D assets.
A layer-focused illustration single-view 3D reconstruction method is adopted. By segmenting the illustration single view into feature region and effect region layers, semantic information is extracted using the layer-CLIP model, and high-quality 3D representation is generated by combining the latent diffusion model, including non-natural line extraction, initial point cloud generation, point cloud segmentation, and latent-three-plane module processing.
It improves the positional accuracy of 2D illustration details in 3D mapping, reduces feature offset, enhances the quality and computational efficiency of 3D generation, expands the generation range, and reduces the risk of hallucinations.
Smart Images

Figure CN119444995B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to image processing methods and systems, specifically to a layer-focused illustration single-view method. Figure 3 D reconstruction methods and electronic devices, storage media, and products. Background Technology
[0002] In recent years, AIGC's work has become a hot topic in the entertainment, remote sensing, and commercial fields. This trendy and complex technology has attracted increasing attention and investment as related communities have developed and hardware products have evolved. As a major direction in the 3D field, 3D generation and reconstruction has become simpler with the development of large-scale models and 3D representation methods; even personal computers can deploy models to infer impressive 3D assets. The following are some existing popular technologies:
[0003] 1. Train a new diffusion model using 3D data (i.e., a 3D diffusion model) to directly generate 3D assets based on conditions and maintain strong 3D consistency. A corresponding technical example is Get3D.
[0004] 2. Directly applying 2D diffusion models to 3D generation can handle various text prompts and produce highly detailed and complex geometry and appearance. A corresponding technology example is Dreamface.
[0005] 3. Obtain a sufficient number of multi-view images according to the new view generation approach, apply sparse view reconstruction methods or fractional distillation sampling (SDS) optimization, and fuse these multi-view images into a 3D shape to produce high-quality 3D shape creation. Corresponding technology examples include Instant3d / Dreamfusion.
[0006] Although these technologies have all produced impressive models and generated truly excellent 3D assets, they still have some shortcomings that need to be addressed in extensive experimentation and application.
[0007] Firstly, the first technical approach faces difficulties when scaling up to a large generative domain because 3D data is often hard to obtain and expensive. The current 3D datasets are much smaller than 2D datasets, which results in the generated 3D assets being inadequate in handling complex text prompts and generating complex / detailed geometry and appearance.
[0008] Then, for the second technical method, since the 2D diffusion model cannot understand the camera view, the generated 3D assets are difficult to achieve geometric consistency, and the generated 3D assets are often misaligned, especially for complex instances.
[0009] Finally, the third technical method causes significant inefficiency in the process of indirect generation of multi-view images. In addition, the quality of the generated shape is highly dependent on the fidelity and continuity of the multi-view images, often leading to loss of detail or reconstruction failure.
[0010] Moreover, the aforementioned technologies are mostly applied and optimized for problems involving natural images, and there are still shortcomings in optimizing the 3D generation of 2D illustrations. Among them, the 3D reconstruction problem of 2D illustrations mainly includes complex and extreme color schemes and black outlines based on unnatural lines, but current research suffers from problems such as color block stacking, large computational load, lack of understanding of lighting layers, and insufficient retrieval of specific accessory information. Summary of the Invention
[0011] To address the shortcomings of the existing technology, the present invention aims to provide a single-view illustration with layer focus. Figure 3 The 3D reconstruction method described in this application is applicable to generating 3D models from single-view illustrations and can improve the quality of 3D reconstruction.
[0012] The technical solution provided by this invention is as follows:
[0013] Firstly, this application provides a layer-focused illustration single-view... Figure 3 The D reconstruction method includes the following steps:
[0014] Obtain an illustration single-view dataset T, which includes multiple original illustration single views;
[0015] The original illustration single view in dataset T is segmented for the first time to obtain each feature region layer of the original illustration single view; each feature region layer is segmented for the second time to obtain the effect region layer in each feature region layer; region text pairs containing the original illustration single view and its layers are generated, along with the weights of the corresponding layers when inputting to the image encoder, the total number of layers corresponding to each feature region layer of the original illustration single view, and the two-dimensional position information of the effect region layer obtained in the second segmentation within the feature region layer obtained in the first segmentation; wherein, the layers of the original illustration single view include feature region layers and effect region layers; wherein, each region text pair includes a layer obtained from the segmentation of the original illustration single view and a corresponding text description, the region being a layer in the original illustration single view, and the text content being a concise description of these layers;
[0016] A layer-CLIP model is constructed and trained, comprising an image encoder and a text encoder. The image encoder includes an RGB image convolutional layer and a layer convolutional layer, used to input the original illustration single view and its layers, respectively. The text encoder is used to input the text description of the corresponding layer. The text encoder has an auxiliary channel for inputting the two-dimensional position information of the effect region layer obtained from the second segmentation within the feature region layer obtained from the first segmentation. The trained layer-CLIP model has the function of outputting the semantic information and layer information label c1 corresponding to the original illustration single view, its layers, and the corresponding text description.
[0017] For the original illustration single view, the layer-CLIP model is used to obtain tokens c1 containing corresponding semantic information and layer information;
[0018] The original illustration single view is input into a 3D diffusion model based on the single view. Combined with the marker c1, the corresponding point cloud is diffused using the Latent Diffusion Model (LDM) according to the layer number corresponding to each feature region, to obtain the final 3D representation of the original illustration single view.
[0019] In one possible implementation, the single-view-based 3D diffusion model includes an unnatural line extraction module, an initial point cloud generation module, a point cloud segmentation module, a point cloud-layer diffusion module, and a latent-triplane module.
[0020] The process involves inputting the original single-view illustration into a 3D diffusion model based on that single-view, combining it with marker c1, and applying a latent diffusion model to the corresponding point cloud according to the layer number corresponding to each feature region to perform a corresponding number of diffusion operations, thereby obtaining the final 3D representation of the original single-view illustration. This includes:
[0021] The unnatural line extraction module extracts a first illustration single view from the original illustration single view; the first illustration single view is the illustration single view after removing the unnatural outlines from the original illustration single view.
[0022] The initial point cloud generation module selects the corresponding basic model based on the character's posture in the first illustration single view to generate the initial point cloud P;
[0023] The point cloud segmentation module divides the initial point cloud P into different point cloud parts, and performs basic coloring of each point cloud part based on the color of the region with the smallest weight value in the corresponding marker c1, resulting in a point cloud with basic coloring. Each point cloud part divided from the initial point cloud P corresponds one-to-one with each feature region layer obtained after the first segmentation. For example, if the first segmentation yields four parts: head, body, hand, and foot, subsequent point cloud segmentation will also be based on these four parts.
[0024] The point cloud-layer diffusion module maps the base-colored point cloud to a latent representation z. Based on the latent representation z, the label c1, and the total number of layers corresponding to each feature region layer, a latent diffusion model is used to perform diffusion a corresponding number of times over time t, outputting the final latent representation z. m ;
[0025] Latent-three-plane module, based on the final latent representation z m The final 3D (Triplane) representation is obtained.
[0026] In one possible implementation, extracting the first illustration single view from the original illustration single view includes:
[0027] S101. Extract the unnatural contour line x of the original illustration single view through differential Gaussian sketch extraction; during extraction, a convex hull around the key structure is created using an existing face marker detector to prevent the lines of the key structure from being extracted into the unnatural contour line x, thus avoiding the lines of the key structure being filled when the pixel color near the lines in the drawing is generated by subsequent linear interpolation.
[0028] S102. Use a shallow convolutional neural network to extract preliminary features of the original illustration single view;
[0029] S103. Based on the preliminary features, perform multiple linear interpolations (Lerp) on the region where the non-natural outline x is located to regenerate the pixel color of the non-natural outline x in the original illustration single view, and obtain the first illustration single view.
[0030] In one possible implementation, the step of selecting a corresponding basic model based on the character's posture in the first illustration single view to generate an initial point cloud P includes:
[0031] S201. Select a basic model based on the character's posture in the first illustration single view as the initial 3D model;
[0032] S202. Based on the initial 3D model, construct a triangular mesh M. Use multilayer perceptrons (MLPs) to predict the signed distance function (SDF) value and texture color of each vertex in the triangular mesh M. Then, convert the vertex SDF values and colors of M into point clouds, denoted as pt. M (p M c M ), where p M ∈R 3 This refers to the position of the point cloud, equal to the vertex coordinates of M, c M ∈R 3 This refers to the color of the point cloud, which is the same as the color of the vertices of M;
[0033] S203, in pt M Perform noisy point cloud growth and color perturbation on the surrounding point cloud, including:
[0034] First, calculate pt. M A bounding box is drawn on the surface, and then a noisy point cloud pt is uniformly grown within the bounding box of the non-natural contour line x on the point cloud. r (p r c r ), where p r and c r These represent the location and color of the noise point cloud, respectively.
[0035] S204. Filter the noisy point cloud based on location p. M Construct a K-dimensional tree, based on the nearest point found in the K-dimensional tree and p. r The distance between points is retained within a set distance threshold;
[0036] S205, pt M and pt r The positions and colors are merged to obtain the final initial point cloud P.
[0037] The quality of initialization can be improved by growing noisy point clouds, perturbing colors, and filtering point clouds; during the point cloud filtering process, fast searching can be achieved by constructing a K-dimensional tree.
[0038] In one possible implementation, based on the latent representation z, the label c1, and the total number of layers corresponding to each feature region layer, a latent diffusion model is used, and diffusion is performed a corresponding number of times over time t to output the final latent representation z. m ,include:
[0039] Based on the latent representation z, the label c1, and the total number of layers corresponding to each feature region layer, noise diffusion is performed on the latent representation z. During each noise diffusion step, a cross-attention layer is used to enhance the information of the corresponding layer in the first label c1 and the final latent representation z. m The interaction between them, wherein the noise diffusion order is carried out in ascending order of the weight values in the first label c1, and the diffusion time t is appropriately increased when the weight has only one decimal place; wherein, for the part of the feature region layer with a total of m layers in the latent representation z, m noise diffusions are performed.
[0040] In one possible implementation, based on the final latent representation z m The final 3D (Triplane) representation is obtained, including:
[0041] In obtaining the final potential representation z m Then, it is reshaped into a three-plane representation of z. reshape And by perpendicularly connecting the three planes in the height dimension, we obtain z concat Upsampling is performed to obtain a high-resolution three-plane feature map; the upsampling process is as follows: a convolutional decoder is used to progressively upsample the explicit final latent representation to obtain the final 3D representation.
[0042] In a second aspect, this application provides an electronic device, including: a memory and a processor;
[0043] The memory is used to store computer programs;
[0044] The processor is used to invoke the computer program to execute the method described above.
[0045] Thirdly, this application provides a computer-readable storage medium storing a computer program that, when run on an electronic device, causes the electronic device to perform the method described above.
[0046] Fourthly, this application provides a computer program product, including a computer program that, when run on an electronic device, causes the electronic device to perform the method described above.
[0047] The specific implementation methods of the second to fourth aspects of this application can refer to the implementation methods of the first aspect, and will not be elaborated here.
[0048] The technical solution provided by this invention has the following beneficial technical effects:
[0049] 1. The layer-focused illustration single-view method used in this invention Figure 3The 3D reconstruction method fully simulates the drawing process of 2D illustration, that is, to understand the detailed information on the illustration by dividing it into layers. It greatly solves the problems of extreme color block stacking, interference of highlight and shadow layers, and loss of complex accessory information. At the same time, it uses a mature 2D diffusion model to optimize 3D generation, reducing the need for high-quality 3D assets, and accurately positions the 2D details when mapping to 3D, reducing the problem of feature offset.
[0050] 2. The present invention extracts unnatural contour lines x, which not only reduces artifact interference caused by illustration lines, but also makes it convenient to use the area where the unnatural contour lines x are located as the main target for the growth of noisy point clouds and color perturbation in the initial 3D point cloud, thereby further improving the quality of the initial point cloud.
[0051] 3. By segmenting the point cloud, the complex human point cloud is divided into multiple parts. Each part is operated on separately in a low-dimensional latent space, rather than directly in a high-dimensional 3D space. This greatly improves the scalability and computational efficiency of the model, while maintaining the quality of the generated 3D content and avoiding the discontinuity problem in the synthesis of new perspectives.
[0052] 4. The construction of the weighted dataset not only expands the scope of generation but also reduces the risk of illusions when generating the required 3D assets. Attached Figure Description
[0053] Figure 1 This is a flowchart illustrating the method of an embodiment of this application;
[0054] Figure 2 This is a schematic diagram of the non-natural line extraction module in an embodiment of this application;
[0055] Figure 3 This is a schematic diagram of the initial point cloud generation module in an embodiment of this application;
[0056] Figure 4 This is a schematic diagram of the point cloud segmentation module in an embodiment of this application;
[0057] Figure 5 This is a schematic diagram of the point cloud-layer diffusion module in an embodiment of this application;
[0058] Figure 6 This is a schematic diagram of a potential three-plane module in an embodiment of this application;
[0059] Figure 7 This is a diagram showing the correspondence between single-view segmentation, point cloud segmentation, and point cloud-layer diffusion in the embodiments of this application. Detailed Implementation
[0060] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0061] To better understand the embodiments of this application, the relevant models or technologies involved in the embodiments of this application are described below:
[0062] 1. SAM Model as a Foundational Model: SAM, proposed by Meta, is a general-purpose model for image segmentation tasks. SAM consists of three parts: a powerful image encoder computes the image embedding, a cue encoder embeds the cue, and then the two information sources are combined in a lightweight mask decoder to predict the segmentation mask. It can be used plug-and-play with cue to solve various tasks involving object and image distribution, including edge detection, object proposal generation, instance segmentation, and text-to-mask prediction. It can output the segmentation mask in real-time upon cueing for interactive use.
[0063] 2. CLIP stands for Contrastive Language-Image Pre-training, which is a pre-trained model based on contrastive text-image pairs. CLIP is a multimodal model based on contrastive learning. CLIP's training data consists of text-image pairs: an image and its corresponding text description. Through contrastive learning, the model can learn the matching relationship between text-image pairs.
[0064] 2. kd-tree: A tree-like data structure that stores instance points in k-dimensional space for fast retrieval. It is mainly used for searching key data in multi-dimensional space (such as range search and nearest neighbor search).
[0065] 3. PointNet: A deep learning network architecture specifically designed for processing point cloud data. It can learn feature representations directly from point clouds without complex voxelization or shape assumptions, segmenting point clouds into different objects or parts of objects.
[0066] This application discloses a layer-focused illustration single-view method. Figure 3The D reconstruction method includes: collecting a suitable single-view illustration dataset T for reconstruction; processing the original single-view illustrations in T using SAM to generate region text pairs containing the original single-view illustrations and their layer channels to train a fine-tuned layer-CLIP model; S3: extracting and removing unnatural contour lines from each original single-view illustration in T using an unnatural line extraction module; selecting a basic model (referring to a basic, universal, free, and easily accessible human body model) based on the character's posture in each original single-view illustration in T as the initial 3D model to generate an initial point cloud P; performing noise growth and color perturbation at the corresponding positions of the unnatural contour lines in the initial point cloud; obtaining a basic colored point cloud N from the initial point cloud P through point cloud segmentation and coloring, and inputting it into the point cloud-layer diffusion module to map it into the latent space and perform diffusion; obtaining a label c1 containing corresponding semantic and layer information from each original single-view illustration in T through the layer-CLIP model, combining it with the total number of corresponding layers, inputting it into the point cloud-layer diffusion module, and outputting the final latent representation z after multiple diffusions over time t. m Based on the final latent representation z m The final 3D representation is obtained, namely the three-plane representation.
[0067] The following is in conjunction with the appendix Figures 1-7 This application will be described in detail.
[0068] like Figure 1 As shown, this application discloses a layer-focused illustration single-view method. Figure 3 D reconstruction method, which includes the following steps:
[0069] S1. Obtain the illustration single-view dataset T, which includes multiple original illustration single views;
[0070] The dataset T images are required to be high resolution, front view, neutral expression, and undone. In some embodiments, the dataset T may be the Hololive character dataset, the Honkai Impact 3rd character dataset, or the Honkai: StarRail character dataset. The Hololive character dataset is from the virtual idol company Hololive or VirtualYoutuberFandomWiki; the Honkai Impact 3rd character dataset is from the official website database of miHoYo.
[0071] Preferably, the dataset T is the Vroid3D dataset, which consists of 11.2k 3D anime character datasets from VroidHub created by the University of Maryland-College Park. High-resolution, neutral-expression, uncropped images from the Vroid3D dataset can be used as the primary image dataset for 3D generation. For the selected dataset, frontal view images can be filtered out using a frontal face detection method.
[0072] S2. Perform a first segmentation on the original illustration single view in dataset T to obtain each feature region layer of the original illustration single view; perform a second segmentation on each feature region layer to obtain the effect region layer in each feature region layer; generate region text pairs containing the original illustration single view and its layers, the weights of the corresponding layers when inputting to the image encoder, and the total number of layers corresponding to each feature region layer of the original illustration single view.
[0073] The original illustration single view includes a feature area layer and an effect area layer;
[0074] The feature area layer corresponds to various features of a person, such as head, hands, feet, body, facial features, hair, and accessories;
[0075] The effect area layer corresponds to the effect areas on various features of the character, such as highlights and shadows;
[0076] After segmentation, the weights set when inputting the image encoder to the corresponding layers are set according to the segmentation results. The weight values are set between [0, 1]. In the first segmentation, the weight of the foreground region is set to 1 based on the occlusion relationship, and the weights of the subsequent regions are decreased by a gradient of 0.1. In the second segmentation, the weights of the feature region layers obtained in the first segmentation are decreased by a gradient of 0.01.
[0077] Record the total number of layers corresponding to each feature region layer;
[0078] For example, if a certain feature region layer obtained from the first segmentation is segmented a second time, resulting in 4 effect region layers, then the total number of layers corresponding to that feature region layer is 5.
[0079] Obtain the two-dimensional position information of the effect area layer obtained from the second segmentation in the feature area layer obtained from the first segmentation according to the proportion;
[0080] In the steps described above, the segmentation of the original single-view illustrations in dataset T differs from that of natural images. By performing secondary segmentation on the original single-view illustrations in the dataset and recording illustration layer information such as highlights and shadows, the interaction information is recorded in a multi-layered manner before being passed to downstream tasks. This significantly reduces problems such as color block mixing, artifact interference, and loss of accessory information. It is particularly suitable for generating 3D models from single-view illustration datasets, such as generating character head models from game illustrations.
[0081] In some embodiments, the Segment All Model (SAM) can be used to segment the original illustration single view in dataset T.
[0082] Each region text pair includes a layer obtained by segmenting the original illustration single view and its corresponding text description. The region is a layer in the original illustration single view, and the text content is a concise description of these layers (such as in English).
[0083] S3. Construct and train a layer-CLIP model, wherein the layer-CLIP model includes an image encoder and a text encoder; the image encoder includes an RGB image convolutional layer and a layer convolutional layer, which are used to input the original illustration single view and its layers respectively; the text encoder is used to input the text description of the corresponding layer; the trained layer-CLIP model has the function of outputting the semantic information and layer information tokens c1 corresponding to the original illustration single view, its layers and the corresponding text description.
[0084] Compared to the existing CLIP model, the above-mentioned layer-CLIP model adds an adaptive layer convolutional layer to the VisionTransformer (ViT) structure of the image encoder. This is equivalent to adding a layer channel to the image encoder. This convolutional layer is used to process the input data of the layer channel and runs in parallel with the RGB convolutional layer. It allows the CLIP image encoder to accept additional layer channels as input. The weight values are set between [0, 1], where 1 represents the foreground and 0 represents the background. When the input precision is 0.1, it is the image obtained from the first segmentation. When the precision is 0.01, it is the image obtained from the second segmentation. A 3*3 convolutional kernel is used for the first segmented image, and a 1*1 convolutional kernel is used for the second segmented image.
[0085] In the CLIP model text encoder settings, an auxiliary channel is added to input the two-dimensional position information of the effect region layer obtained from the second segmentation in the feature region layer obtained from the first segmentation. This channel is open to accept input only when the layer image after the second segmentation is input into the image encoder. In other cases, the input of this channel can be set to (-2, -2) and is regarded as invalid input.
[0086] For example, an auxiliary channel can be added to the CLIP model text encoder settings to input the two-dimensional position information of the layer after the second segmentation in the first segmentation layer image, with a range of x ~ (-1, 1) and y ~ (-1, 1). For instance, for a hand, the position information is (0, 0) with the center of the back of the hand as the origin, while for the shadow on the middle finger of the hand, the position information may be (0.1, -0.9). The specific position information can be selected by scaling it down proportionally.
[0087] When using it in practice, it simultaneously accepts data from the original channels (RGB channels) and layer channels to train and fine-tune the layer-CLIP model. That is, the original illustration single view and its segmented layers are simultaneously input into the image encoder.
[0088] Image data from GRIT-20m can be used to train fine-tuning layers - CLIP on RGBL region-text pairs.
[0089] S4. For the original illustration single view, obtain tokens c1 containing corresponding semantic information and layer information through the layer-CLIP model;
[0090] S5. Input the original illustration single view into the 3D diffusion model based on the single view, and combine it with the marker c1. Diffuse the corresponding point cloud using the Latent Diffusion Model (LDM) according to the layer number corresponding to each feature region to obtain the final 3D representation corresponding to the original illustration single view.
[0091] In some embodiments, the single-view-based 3D diffusion model includes an unnatural line extraction module, an initial point cloud generation module, a point cloud segmentation module, a point cloud-layer diffusion module, and a latent-triplane module.
[0092] The process involves inputting the original single-view illustration into a 3D diffusion model based on that single-view, combining it with marker c1, and applying a latent diffusion model to the corresponding point cloud according to the layer number corresponding to each feature region to perform a corresponding number of diffusion operations, thereby obtaining the final 3D representation of the original single-view illustration. This includes:
[0093] The unnatural line extraction module extracts a first illustration single view from the original illustration single view; the first illustration single view is the illustration single view after removing the unnatural outlines from the original illustration single view.
[0094] The initial point cloud generation module selects the corresponding basic model based on the character's posture in the first illustration single view to generate the initial point cloud P;
[0095] like Figure 4As shown, the point cloud segmentation module divides the initial point cloud P into different point cloud parts, and performs basic coloring of each point cloud part based on the color of the region with the smallest weight value in the corresponding marker c1, resulting in a basic-colored point cloud. Each point cloud part divided from the initial point cloud P corresponds one-to-one with each feature region layer obtained after the first segmentation. For example, if the first segmentation yields four parts—head, body, hands, and feet—subsequent point cloud segmentation will also be based on these four parts. Figure 7 As shown.
[0096] In some embodiments, the machine learning model PointNet can be used to segment the initial point cloud P into different point cloud parts N.
[0097] The point cloud-layer diffusion module maps the base-colored point cloud to a latent representation z. Based on the latent representation z, the label c1, and the total number of layers corresponding to each feature region layer, a latent diffusion model is used to perform diffusion a corresponding number of times over time t, outputting the final latent representation z. m ;
[0098] Latent-three-plane module, based on the final latent representation z m The final 3D (Triplane) representation is obtained.
[0099] In some embodiments, such as Figure 2 As shown, the non-natural line extraction module extracts the first illustration single view from the original illustration single view, including:
[0100] S101. Extract the unnatural contour line x of the original illustration single view by using the difference Gaussian sketch extraction. During extraction, a convex hull is created around the key structure (such as facial features, including eyes, nose, mouth, eyebrows, and ears) using an existing facial marker detector to prevent the lines of the key structure from being extracted into the unnatural contour line x, thus avoiding the lines of the key structure being filled when the pixel color near the lines in the drawing is generated by subsequent linear interpolation.
[0101] S102. Use a shallow convolutional neural network to extract preliminary features of the original illustration single view;
[0102] S103. Based on the preliminary features, perform multiple linear interpolations (Lerp) on the region where the non-natural outline x is located to regenerate the pixel color of the non-natural outline x in the original illustration single view, and obtain the first illustration single view.
[0103] In some embodiments, such as Figure 3 As shown, the initial point cloud generation module selects the corresponding basic model based on the character's posture in the first illustration single view to generate the initial point cloud P, including:
[0104] S201. Select a basic model based on the character's posture in the first illustration single view as the initial 3D model;
[0105] S202. Based on the initial 3D model (asset), construct a triangular mesh M. Use multilayer perceptrons (MLPs) to predict the signed distance function (SDF) value and texture color of each vertex in the triangular mesh M. Then, convert the vertex SDF values and colors of M into point clouds, denoted as pt. M (p M c M ), where p M ∈R 3 This refers to the position of the point cloud, equal to the vertex coordinates of M, c M ∈R 3 This refers to the color of the point cloud, which is the same as the color of the vertices of M;
[0106] For example, the grid size is 128. 3 This means there are 128 points in each coordinate axis direction, for a total of 128×128×128 vertices. The signed distance function SDF(x, y, z) represents the distance of each point to the nearest surface of the initial 3D model. It is signed, where a negative value indicates that the point is inside the surface and a positive value indicates that the point is outside the surface. In MLPs, an SDF value is calculated for each vertex (x, y, z) to represent its position relative to the surface of the initial 3D model. Each vertex (x, y, z) not only has position information, but also associated texture color information, which is usually represented as (r, g, b), representing the red, green and blue color components.
[0107] S203, in pt M Perform noisy point cloud growth and color perturbation on the surrounding point cloud, including:
[0108] First, calculate pt. M A bounding box (BBox) is drawn on the surface, and then a noisy point cloud pt is uniformly grown within the BBox covering the non-natural contour line x on the point cloud. r (p r c r ), where p r and c r These represent the location and color of the noise point cloud, respectively.
[0109] The BBox is used to find the smallest bounding rectangle of the point cloud in 3D space. It iterates through all the points in the point cloud and records the minimum and maximum values of each point on the X, Y, and Z coordinate axes. The bounding box can be represented by any two opposite corner points of its eight corner points, usually the minimum and maximum points.
[0110] Since the selected illustration images are all front views, the 2D image coordinates are mapped to 3D space, ignoring the changes in the Z-axis. The 3D point cloud coordinates corresponding to each pixel (u, v) in the non-natural contour line x of the 2D image are: Where D is the distance of the 3D point in the depth direction (Z-axis), and f is the focal length of the camera;
[0111] S204. Filter the noisy point cloud based on location p. M Construct a KD-tree based on the nearest point found in the KD-tree and p. r The distance between points is retained to keep the points within a set distance threshold; that is, iterate through each generated point, find the nearest point and check whether the distance condition is met.
[0112] By constructing a K-dimensional tree, fast searching can be achieved.
[0113] For example, the distance threshold mentioned above can be a normalized distance of 0.01.
[0114] Among them, for the color of the noisy point cloud, make c r With c M Similar, but with some perturbations added:
[0115] c r =c M +a
[0116] The value of 'a' is randomly sampled between 0 and 0.2.
[0117] S205, pt M and pt r The positions and colors are merged to obtain the final initial point cloud P.
[0118] The quality of initialization can be improved by growing noisy point clouds, perturbing colors, and filtering point clouds; during the point cloud filtering process, fast searching can be achieved by constructing a K-dimensional tree.
[0119] In some embodiments, such as Figure 5 As shown, the point cloud-layer diffusion module maps the base-colored point cloud to a latent representation z, including:
[0120] Fourier features are used to represent the positional structure of the point cloud after basic coloring, and a series of learnable tokens are introduced. Cross-attention layers are used to query point cloud features, allowing 3D information from the point cloud to be injected into latent labels. Subsequently, multiple self-attention layers are used to enhance the representation of these labels, resulting in a latent representation of the point cloud after basic coloring. Where r*r represents the resolution of the latent representation, d e Denotes the channel dimension of e, d z This represents the z-channel dimension.
[0121] Fourier features are used to represent the positional information of a point cloud. For the position P of each point... i It can be mapped to a high-dimensional space through Fourier feature mapping:
[0122] γ(p i )=[cos(2πB P i ), sin(2πB P) i )]
[0123] Where B∈R r×3 It is a Fourier basis matrix, usually randomly sampled.
[0124] In some embodiments, based on the latent representation z, the label c1, and the total number of layers corresponding to each feature region layer, a latent diffusion model is used to perform diffusion a corresponding number of times over time t, outputting the final latent representation z. m ,include:
[0125] A cross-attention layer is used during each noise diffusion step to enhance the corresponding layer information in the first label c1 and the final latent representation z. m The interaction between them, wherein the noise diffusion order is carried out in ascending order of the weight values in the first label c1, and the diffusion time t can be appropriately increased when the weight has only one decimal place; wherein, for the part of the feature region layer with a total of m layers in the latent representation z, m noise diffusions are performed.
[0126] In some embodiments, LoRA is used to apply the model weights ε of the potential diffusion model during each potential diffusion process. i Fine-tuning training was performed; and the semantic information in c1 was also added as interpolation to the data ε. i In the model, all LoRA matrices are concatenated to obtain weight data representing a portion of a layer. After training on N different instances, the final model weight dataset D = {ε1, ε2, ..., ε} is obtained. N};
[0127] In this context, the model weights of the latent diffusion model are its parameters, which are optimized during training to minimize the loss function. The model weights encode the knowledge the model learns from the input data and determine how the model maps the input to the output.
[0128] Because certain layer attributes often appear together during layer training, they may become entangled during subsequent retrieval. For example, reddish hair and warm highlights may limit the accuracy of the model when editing or generating specific attributes. Therefore, based on the information of label c1, when encountering attribute entanglement, the result of fine-tuning the training of the smaller weight in the model of label c1 layer is used to cover the result of the other side.
[0129] Principal component analysis is performed on the model weight dataset. The model weight dataset is then dimensionality-reduced while retaining the principal components. The data points are compressed into fewer parameters, retaining only the most important information. After dimensionality reduction of the dataset, the model weight dataset is used as data to train a semantic retrieval tool. A hypergraph is constructed based on the semantic information in c1 as labels.
[0130] A new model can be generated by sampling the model weight dataset D based on a semantic information S, thus creating a new layer. A hypergraph is defined for each layer, and the semantic information of the marker c1 embedded in the model weights is used for retrieval. Then, the relevant model weights in the model are edited. For example, to obtain the right hand of a woman wearing a red bracelet and black tattoo in sunlight, a hypergraph that conforms to this semantics is constructed. A hyperedge connects the model weights of warm highlights, red bracelet, woman's right hand, and black tattoo in the model weight dataset. Based on the combination relationship of this hypergraph, the model weights that conform to this semantics are obtained, and then diffusion is continued according to these model weights.
[0131] The above steps, by injecting semantic information and layer information from the image into the latent space, accurately map image features onto the point cloud. This ensures that the high-frequency details of the 3D assets generated by the diffusion model are aligned with the conditional image (original illustration single view). At the same time, the model weights from each diffusion are used as new data to construct a model weight dataset that generates corresponding 3D asset features, expanding the range of generated objects and reducing the risk of creating illusions.
[0132] In some embodiments, such as Figure 6 As shown, the latent-three-plane module is based on the final latent representation z. m The final 3D (Triplane) representation is obtained, including:
[0133] In obtaining the final potential representation z m Then, it is reshaped into a three-plane representation of z. reshape And by perpendicularly connecting the three planes in the height dimension, we obtain To prevent erroneous mixing of planes along the channel dimension, the latent-triplane module then... concat Upsampled to a high-resolution three-plane feature map, with an upsampling factor of f.
[0134] The upsampling process involves using a convolutional decoder (convolutional network) to progressively upsample the explicit final latent representation and obtain the final 3D representation.
[0135] The reshaped three planes are represented as follows:
[0136] z reshape =reshape(z, (3, r, r, d) z ))
[0137] Connect the three planes perpendicularly along the height dimension to obtain
[0138]
[0139] By using convolutional networks to progressively upsample the explicit final latent representation, compared to using a Transformer decoder, which can effectively upsample through transposed convolutions while providing a different feature extraction method than the encoder, thus achieving feature complementarity;
[0140] This method is used to generate head models of game illustration characters, virtual anchors, or 3D figurines, with a preference for generating head models of game illustration characters.
[0141] Below are anime illustrations featuring characters wearing animal ear headbands. Figure 3 Let's take D reconstruction as an example to illustrate.
[0142] 1. Input an anime illustration
[0143] 2. Layer segmentation and weight setting
[0144] Feature region layer segmentation: The SAM model is used to perform the first segmentation of the 2D illustration to obtain the basic feature layer, that is, to divide the head, face, ears, hair ornaments, and other areas of the 2D illustration into separate layers. Lighting and shadow effect layers are further segmented within the ear and hair ornament feature area.
[0145] Layer description and text generation: Based on the segmentation, each layer and its layer description are obtained, and region text pairs are generated (such as "head region", "ear hair ornament", "highlight effect"). The label containing the corresponding semantic information and layer information is obtained through the layer-CLIP model, which is used for layer control during generation.
[0146] Weight setting: After segmentation, set the weights of the corresponding layers when inputting the image encoder according to the segmentation results.
[0147] 3. Initial point cloud generation
[0148] Select a base model: Based on the character's facial and body features, select a base model with head and ear outlines as the initial 3D model.
[0149] Triangular mesh and point cloud generation: A triangular mesh is constructed based on the selected primitive model. The signed distance function (SDF) of the mesh vertices and the texture color are predicted using a multilayer perceptron (MLP). The vertex SDF values and colors of the mesh are converted into an initial point cloud to ensure that the contour and color information of features such as ears and hair ornaments are fully preserved.
[0150] 4. Point cloud noise growth and color perturbation
[0151] Unnatural line extraction: Unnatural outlines of the illustration (such as the black edges of the ear hair ornaments) are extracted using the differential Gaussian filtering method, and the position of the outlines in the illustration layer is recorded.
[0152] Noise point cloud growth: Generate a bounding box (BBox) within the mapped area of the ear hair ornament, and grow a noise point cloud uniformly within the bounding box to match the noise point cloud density with the ear hair ornament area.
[0153] KD-Tree filtering and color perturbation: The KD-Tree structure is used to filter the noisy point cloud, retaining noisy points within the distance threshold by nearest neighbor distance, and a slight color perturbation is added to the ear area to make the color coordinate with the surrounding hair accessories, generating a natural hair accessory color transition effect.
[0154] 5. Point cloud - layer diffusion processing
[0155] Application of the latent diffusion model: The point cloud is mapped to the latent representation space z, and the region is diffused a corresponding number of times according to the layer weights (e.g., the weight of the ear hair ornament is 0.9). A cross-attention layer is used during the diffusion process to allow the layer information of the ear hair ornament and the point cloud features to interact.
[0156] Diffusion order: Diffusion is performed in ascending order of layer weight values to ensure that the structure of the background layer and foreground layers such as hair accessories is improved layer by layer. If the weight precision of the ear hair accessories is 0.1, the diffusion time t is appropriately increased to improve the effect.
[0157] 6. The latent three-plane module generates a three-plane representation.
[0158] Three-plane construction: The final potential representation after diffusion is reshaped into a three-plane representation, and the three planes are stacked vertically along the height dimension to ensure that the hair accessory layer is not confused when generated.
[0159] Upsampling generates high-resolution 3D feature maps: The three-plane feature maps are progressively upsampled through a convolutional decoder to generate a high-resolution three-dimensional representation, making details such as ears and hair ornaments clearly visible in the 3D model.
[0160] 7. Final 3D model output
[0161] Model Output: The final generated 3D model includes the complete structure and texture of the head and ear ornaments. The highlights, shadows, and edge line details of the ear ornaments are naturally rendered in the 3D model, achieving a visual effect consistent with the original illustration.
[0162] This process makes full use of layer information, noise growth, and multi-layer diffusion methods to ensure that ear ornaments and other details are accurately rendered during the generation process and conform to the characteristics of two-dimensional illustrations.
[0163] This application also provides an electronic device, including: a memory and a processor;
[0164] The memory is used to store computer programs;
[0165] The processor is used to invoke the computer program to execute the method described above.
[0166] This application also provides a computer-readable storage medium storing a computer program that, when run on an electronic device, causes the electronic device to perform the method described above.
[0167] This application also provides a computer program product, including a computer program that, when run on an electronic device, causes the electronic device to perform the method described above.
[0168] This application also provides specific implementations of a system, electronic device, computer-readable storage medium, and computer program product. These specific implementations can be referred to in the above-described methods and will not be repeated here.
[0169] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for 3D reconstruction of a single-view illustration using layer-focusing, characterized in that, Includes the following steps: Obtain an illustration single-view dataset T, which includes multiple original illustration single views; The original illustration single view in dataset T is segmented for the first time to obtain each feature region layer of the original illustration single view; each feature region layer is segmented for the second time to obtain the effect region layer in each feature region layer; region text pairs containing the original illustration single view and its layers, the weights of the corresponding layers when inputting to the image encoder, the total number of layers corresponding to each feature region layer of the original illustration single view, and the two-dimensional position information of the effect region layer obtained in the second segmentation in the feature region layer obtained in the first segmentation are generated; wherein, the layers of the original illustration single view include feature region layers and effect region layers; wherein, each region text pair includes a layer obtained from the segmentation of the original illustration single view and the corresponding text description; A layer-CLIP model is constructed and trained, comprising an image encoder and a text encoder. The image encoder includes an RGB image convolutional layer and a layer convolutional layer, used to input the original illustration single view and its layers, respectively. The text encoder is used to input the text description of the corresponding layer. The text encoder has an auxiliary channel for inputting the two-dimensional position information of the effect region layer obtained from the second segmentation within the feature region layer obtained from the first segmentation. The trained layer-CLIP model has the function of outputting the semantic information and layer information label c1 corresponding to the original illustration single view, its layers, and the corresponding text description. For the original illustration single view, the label c1 containing the corresponding semantic information and layer information is obtained through the layer-CLIP model; The original illustration single view is input into a 3D diffusion model based on the single view. Combined with the marker c1, the corresponding point cloud is diffused a certain number of times using the latent diffusion model according to the layer number corresponding to each feature region, so as to obtain the final 3D representation of the original illustration single view.
2. The method according to claim 1, characterized in that, The single-view-based 3D diffusion model includes a non-natural line extraction module, an initial point cloud generation module, a point cloud segmentation module, a point cloud-layer diffusion module, and a potential-three-plane module. The process involves inputting the original single-view illustration into a 3D diffusion model based on that single-view, combining it with marker c1, and applying a latent diffusion model to the corresponding point cloud according to the layer number corresponding to each feature region to perform a corresponding number of diffusion operations, thereby obtaining the final 3D representation of the original single-view illustration. This includes: The unnatural line extraction module extracts a first illustration single view from the original illustration single view; the first illustration single view is the illustration single view after removing the unnatural outlines from the original illustration single view. The initial point cloud generation module selects the corresponding basic model based on the character's posture in the first illustration single view to generate the initial point cloud P; The point cloud segmentation module divides the initial point cloud P into different point cloud parts, and performs basic coloring of each point cloud part according to the color of the region with the smallest weight value in the marker c1 corresponding to each point cloud part, to obtain the point cloud after basic coloring; wherein, each point cloud part divided into the initial point cloud P corresponds one-to-one with each feature region layer obtained after the first segmentation. The point cloud-layer diffusion module maps the base-colored point cloud into a latent representation z. Based on the latent representation z, the marker c1, and the total number of layers corresponding to each feature region layer, a latent diffusion model is used to perform a corresponding number of diffusions over time t, outputting the final latent representation z. m ; The latent-three-plane module is based on the final latent representation z. m The final 3D representation is obtained, namely the three-plane representation.
3. The method according to claim 2, characterized in that, The step of extracting the first illustration single view from the original illustration single view includes: S101. Extract the unnatural contour lines x of the original illustration single view using a difference Gaussian sketch; during extraction, a convex hull is created around the key structure using an existing face marker detector; S102. Use a shallow convolutional neural network to extract preliminary features of the original illustration single view; S103. Based on the preliminary features, perform multiple linear interpolations on the region where the non-natural outline x is located to regenerate the pixel color of the non-natural outline x in the original illustration single view, and obtain the first illustration single view.
4. The method according to claim 3, characterized in that, The step of selecting a corresponding basic model based on the character's posture in the first illustration single view to generate an initial point cloud P includes: S201. Select a basic model based on the character's posture in the first illustration single view as the initial 3D model; S202. Based on the initial 3D model, construct a triangular mesh M. Use multilayer perceptrons (MLPs) to predict the signed distance function (SDF) value and texture color of each vertex in the triangular mesh M. Then, convert the vertex SDF values and colors of M into point clouds, denoted as pt. M (p M c M ), where p M ∈R 3 This refers to the position of the point cloud, equal to the vertex coordinates of M, c M ∈R 3 This refers to the color of the point cloud, which is the same as the color of the vertices of M; S203, in pt M Perform noisy point cloud growth and color perturbation on the surrounding point cloud, including: First, calculate pt. M A bounding box is drawn on the surface, and then a noisy point cloud pt is uniformly grown within the bounding box of the non-natural contour line x on the point cloud. r (p r c r ), where p r and c r These represent the location and color of the noise point cloud, respectively. S204. Filter the noisy point cloud based on location p. M Construct a K-dimensional tree, based on the nearest point found in the K-dimensional tree and p. r The distance between points is retained within a set distance threshold; S205, pt M and pt r The positions and colors are merged to obtain the final initial point cloud P.
5. The method according to claim 4, characterized in that, Based on the latent representation z, the label c1, and the total number of layers corresponding to each feature region layer, a latent diffusion model is adopted. Diffusion is performed a corresponding number of times over time t to output the final latent representation z. m ,include: Based on the latent representation z, the label c1, and the total number of layers corresponding to each feature region layer, noise diffusion is performed on the latent representation z. During each noise diffusion step, a cross-attention layer is used to enhance the information of the corresponding layer in label c1 and the final latent representation z. m The interaction between them, wherein the noise diffusion order is carried out in ascending order of the weight values in label c1, and the diffusion time t is appropriately increased when the weight has only one decimal place; wherein, for the part of the feature region layer with a total of m layers in the latent representation z, the noise diffusion is carried out m times.
6. The method according to claim 5, characterized in that, Based on the final latent representation z m To obtain the final 3D representation, including: In obtaining the final potential representation z m Then, it is reshaped into a three-plane representation of z. reshape And by perpendicularly connecting the three planes in the height dimension, we obtain z concat Upsampling yields a high-resolution three-plane feature map; the upsampling process involves progressively upsampling the displayed latent representation using a convolutional decoder to obtain the final 3D representation.
7. An electronic device, characterized in that, include: Memory and processor; The memory is used to store computer programs; The processor is configured to invoke the computer program to perform the method as described in any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed on an electronic device, causes the electronic device to perform the method as described in any one of claims 1 to 6.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is run on an electronic device, it causes the electronic device to perform the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
3D modeling method
CN115496861A
Three-dimensional target reconstruction method based on non-calibration single view
CN118470221A