Generation Method, Device, Equipment and Storage Medium of Assembled Building Block Model
Through deep learning diffusion model and image processing technology, the building block model map is automatically generated, which solves the problem of inefficiency of traditional design methods, and realizes an efficient and automated mapping design process, improving the efficiency and quality of the design.
Patent Information
- Application Number
- CN202510055153.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-01-14
AI Technical Summary
In traditional building block model design, the texture design is inefficient and time-consuming, especially when dealing with complex textures or fine patterns, it is difficult to achieve the expected results, which limits the innovation and diversity of the design.
The deep learning diffusion model is used to generate initial map renderings, and automatic map is achieved through cutout processing and projection mapping. Combining deep learning technology and image processing technology, the generation process from user input to the final map model is automated.
It significantly improves the efficiency and quality of building block map design, solves the time-consuming and labor-consuming problem of traditional manual design methods, and realizes full automatic generation of the final map model from user input to the.
Smart Images

Figure CN119478318B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a method, device, equipment and storage medium for generating an assembled building block model. Background Art
[0002] In the field of building block model design, texture mapping design is a key link, which can significantly improve the visual effect and expressiveness of building block models. Traditional building block texture mapping design methods mainly rely on manual operations by designers. Designers usually need to use professional graphic design software to draw the texture maps of each building block one by one according to user requirements and internal specifications. This process involves multiple steps, including pattern design, color adjustment, texture addition, and matching with the shape of the building blocks.
[0003] However, this traditional manual design method has a significant defect: low efficiency and time-consuming. For a set of complex building block products, designers often need to spend a lot of time to complete the texture mapping design. Especially when dealing with complex textures or delicate patterns, the design process is even more time-consuming, and sometimes it is even difficult to achieve the expected effect. This low efficiency not only increases the product development cycle but also limits the innovation and diversity of the design. Summary of the Invention
[0004] The main purpose of the present invention is to solve the technical problem of low efficiency and time-consuming in the texture mapping design method in the existing building block model design process;
[0005] The first aspect of the present invention provides a method for generating an assembled building block model, and the method for generating the assembled building block model includes:
[0006] Obtain a source image and a model rendering image to be texture mapped, and generate an initial texture rendering image according to the source image and the model rendering image to be texture mapped through a deep learning diffusion model;
[0007] Use the initial texture rendering image as the target texture rendering image, and perform matte extraction processing on the target texture rendering image to obtain a matte bitmap;
[0008] Use the matte bitmap as a projection texture map, and perform projection mapping on the projection texture map according to the preset block structure of the assembled building block model to automatically texture map the assembled building block model and obtain an assembled building block model with texture mapping completed.
[0009] Optionally, in the first implementation manner of the first aspect of the present invention, the obtaining a source image and a model rendering image to be texture mapped, and generating an initial texture rendering image according to the source image and the model rendering image to be texture mapped through a deep learning diffusion model includes:
[0010] Obtain the source image and the model rendering image to be pasted, and perform multi-scale feature extraction on the source image to obtain a set of feature maps containing visual features at multiple scales;
[0011] According to the geometric structure of the model rendering image to be pasted, perform adaptive spatial transformation on the visual features in the set of feature maps to obtain a set of transformed feature maps aligned with the model rendering image;
[0012] Combine the set of transformed feature maps and the model rendering image as conditional inputs and input them into a pre-trained deep learning diffusion model to obtain an initial noise map;
[0013] Perform an iterative denoising process on the initial noise map, generating or updating a denoised image in each iteration until a preset number of iterations is reached, and use the denoised image corresponding to the number of iterations as the initial pasted image rendering.
[0014] Optionally, in the second implementation manner of the first aspect of the present invention, the step of combining the set of transformed feature maps and the model rendering image as conditional inputs and inputting them into a pre-trained deep learning diffusion model to obtain an initial noise map includes:
[0015] Combine the set of transformed feature maps and the model rendering image as conditional inputs and input them into a pre-trained deep learning diffusion model, where the deep learning diffusion model includes a conditional embedding module and a U-Net denoising network;
[0016] Use the conditional embedding module to encode the set of transformed feature maps and the model rendering image to obtain a conditional embedding vector, and generate a random noise map with the same size as the model rendering image according to a preset initial noise distribution;
[0017] Input the conditional embedding vector and the random noise map into the U-Net denoising network to obtain a predicted noise, and sample and update the random noise map according to the predicted noise and a preset noise schedule to obtain an initial noise map.
[0018] Optionally, in the third implementation manner of the first aspect of the present invention, after using the initial pasted image rendering as the target pasted image rendering, it further includes:
[0019] Judge whether the initial pasted image rendering as the target pasted image rendering is a high-resolution rendering;
[0020] If not, perform multi-scale decomposition on the initial pasted image rendering to obtain a set of sub-images with different frequency components, and perform upsampling processing on each sub-image in the set of sub-images according to a preset super-resolution factor to obtain an enlarged set of sub-images;
[0021] Input the enlarged sub - image set into a pre - trained super - resolution diffusion model, and perform iterative denoising processing on each enlarged sub - image to obtain a sub - image set with enhanced high - frequency details;
[0022] Perform adaptive fusion on the sub - image set with enhanced high - frequency details to obtain a preliminary super - resolution image, and use an edge - aware detail enhancement network to process the preliminary super - resolution image to obtain a high - resolution texture image;
[0023] Update the target texture rendering map according to the high - resolution texture image.
[0024] Optionally, in the fourth implementation manner of the first aspect of the present invention, the process of performing matte extraction on the target texture rendering map to obtain a matte bitmap includes:
[0025] Generate a rendering map of the texture model to be pasted with the same shape and perspective according to the shape and perspective of the high - resolution texture image;
[0026] Perform per - pixel difference calculation on the high - resolution texture image and the rendering map of the texture model to be pasted, take the absolute value of the difference result, and sum multiple color channels to obtain a difference matrix;
[0027] Perform binarization processing on the difference matrix according to a preset threshold, retain pixels greater than the threshold, and set pixels less than the threshold to transparent to obtain a preliminary matte bitmap;
[0028] Apply the Sobel operator to the preliminary matte bitmap for edge detection to remove abnormal edges and obtain a matte bitmap.
[0029] Optionally, in the fifth implementation manner of the first aspect of the present invention, the process of using the matte bitmap as a projection texture and performing projection mapping on the projection texture according to the block structure of a preset assembled building block model to automatically texture the assembled building block model to obtain a textured assembled building block model includes:
[0030] Use the matte bitmap as a projection texture, and detect whether there is an adjustment instruction for the matte bitmap used as the projection texture. If so, convert the matte bitmap into a corresponding texture vector map, and update the projection texture according to the texture vector map;
[0031] According to the preset block structure of the assembled building block model, perform segmentation processing on the projection texture to obtain a set of sub - projection textures corresponding to each building block of the block structure;
[0032] Perform geometric transformation and projection calculation on each sub - projection texture in the set of sub - projection textures to obtain a projection texture that matches the surface of the corresponding building block;
[0033] According to the three-dimensional geometric information of the building blocks, perform UV coordinate mapping on the projection texture map to obtain texture coordinate information, and associate the texture coordinate information with the three-dimensional models of the corresponding building blocks to obtain a set of building block models with texture information;
[0034] Assemble and render the set of building block models with texture information to obtain an assembled building block model with the texture completed.
[0035] Optionally, in the sixth implementation manner of the first aspect of the present invention, the performing geometric transformation and projection calculation on each sub-projection texture map in the set of sub-projection texture maps to obtain a projection texture map matching the surface of the corresponding building block includes:
[0036] Perform edge detection and contour analysis on each sub-projection texture map in the set of sub-projection texture maps to obtain a set of control points representing the main contour of the sub-projection texture map;
[0037] According to the three-dimensional geometric information of the corresponding building block, calculate the normal vector and the main direction of the surface of the building block to obtain a surface parameterization description;
[0038] Use the set of control points and the surface parameterization description to construct a projection mapping function and calculate the mapping relationship from the sub-projection texture map to the surface of the building block;
[0039] According to the mapping relationship, perform geometric transformation and projection calculation on the sub-projection texture map to obtain a projection texture map matching the surface of the corresponding building block.
[0040] The second aspect of the present invention provides a device for generating an assembled building block model, and the device for generating an assembled building block model includes:
[0041] An image diffusion module, configured to obtain a source picture and a model rendering picture to be textured, and generate an initial texture rendering picture according to the source picture and the model rendering picture to be textured through a deep learning diffusion model;
[0042] A matte extraction module, configured to use the initial texture rendering picture as a target texture rendering picture and perform matte extraction processing on the target texture rendering picture to obtain a matte bitmap;
[0043] A texture projection module, configured to use the matte bitmap as a projection texture map and perform projection mapping on the projection texture map according to the preset block structure of the assembled building block model to realize automatic texturing of the assembled building block model and obtain an assembled building block model with the texture completed.
[0044] In a third aspect of the present invention, there is provided a generating device for an assembled building block model, including: a memory and at least one processor, wherein instructions are stored in the memory, and the memory and the at least one processor are interconnected by a line; the at least one processor invokes the instructions in the memory to cause the generating device of the assembled building block model to execute the steps of the above-mentioned generating method of the assembled building block model.
[0045] In a fourth aspect of the present invention, there is provided a computer-readable storage medium, in which instructions are stored, and when it runs on a computer, it causes the computer to execute the steps of the above-mentioned generating method of the assembled building block model.
[0046] The above-mentioned generating method, device, equipment and storage medium of the assembled building block model obtain a source picture and a model rendering picture to be pasted, and generate an initial pasted picture rendering through a deep learning diffusion model according to the source picture and the model rendering picture to be pasted; use the initial pasted picture rendering as the target pasted picture rendering, and perform matte extraction processing on the target pasted picture rendering to obtain a matte bitmap; use the matte bitmap as a projection texture map, and perform projection mapping on the projection texture map according to the preset block structure of the assembled building block model to realize automatic texturing of the assembled building block model and obtain an assembled building block model with texturing completed. This method combines deep learning technology and image processing technology to realize the full-automatic generation process from user input to the final textured model, significantly improving the efficiency and quality of building block texturing design and solving the problem of time-consuming and laborious traditional manual design methods.
[0047] Other features and advantages of the present invention will be described in the following description, and in part, will be obvious from the description, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the description, claims and drawings.
[0048] To make the above-mentioned objectives, features and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given in conjunction with the accompanying drawings and are described in detail as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a schematic diagram of the first embodiment of the generating method of the assembled building block model in the embodiment of the present invention;
[0050] Figure 2 It is a schematic diagram of an embodiment of the generating device of the assembled building block model in the embodiment of the present invention;
[0051] Figure 3 It is a schematic diagram of an embodiment of the generating equipment of the assembled building block model in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0053] The terms "including" and "having" and any variations thereof mentioned in the embodiments of the present invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device including a series of steps or units is not limited to the listed steps or units, but optionally further includes other unlisted steps or units, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0054] To facilitate the understanding of this embodiment, a method for generating an assembled building block model disclosed in the embodiments of the present invention will be introduced in detail first. As Figure 1 shown, this method includes the following steps:
[0055] 101. Obtain a source picture and a model rendering picture to be textured, and generate an initial texture rendering picture based on the source picture and the model rendering picture to be textured through a deep learning diffusion model;
[0056] In one embodiment of the present invention, the obtaining of the source picture and the model rendering picture to be textured, and generating the initial texture rendering picture based on the source picture and the model rendering picture to be textured through the deep learning diffusion model includes: obtaining the source picture and the model rendering picture to be textured, and performing multi-scale feature extraction on the source picture to obtain a set of feature maps containing visual features of multiple scales; performing adaptive spatial transformation on the visual features in the set of feature maps according to the geometric structure of the model rendering picture to be textured to obtain a set of transformed feature maps aligned with the model rendering picture; combining the set of transformed feature maps and the model rendering picture as conditional inputs and inputting them into a pre-trained deep learning diffusion model to obtain an initial noise map; performing an iterative denoising process on the initial noise map, generating or updating a denoised image in each iteration until a preset number of iterations is reached, and using the denoised image corresponding to the number of iterations as the initial texture rendering picture.
[0057] Specifically, first, obtain the source image and the model rendering image to which the texture is to be applied. The source image is usually a high-resolution real-person photo or other reference image uploaded by the user, and the format may include common image formats such as JPEG and PNG. The system will preprocess these images, including unifying the size, adjusting the color space (such as converting from RGB to the Lab color space), and performing basic image enhancement (such as contrast adjustment). The model rendering image to which the texture is to be applied is the three-dimensional rendering result of the building block model pre-generated by 3D modeling software (such as Blender or Maya). This rendering image contains the geometric information, material properties, and lighting information of the building block model, and is usually saved in a high-resolution PNG format to retain transparency information. After obtaining these two inputs, the system will perform a preliminary analysis on them, including extracting the basic attributes of the images (such as size, color range) and performing a simple quality check (such as detecting whether the image is complete, whether there are serious noise or blurring).
[0058] Specifically, perform multi-scale feature extraction on the source image to obtain a set of feature maps containing visual features at multiple scales. This step uses a pre-trained deep convolutional neural network, such as VGG19 or ResNet50. Each layer of the network will extract features at different scales and abstraction levels. Specifically, the shallow network (such as the first few convolutional layers) mainly captures low-level visual features, such as edge, texture, and color information; the middle-level network captures more complex structures, such as shapes and local patterns; the deep network extracts high-level semantic features, such as parts or the overall structure of an object. In actual operation, multiple key layers of the network will be selected as feature extraction points, usually including the last layer of each convolutional block. The feature maps extracted from these layers will be saved to form a set of multi-scale feature maps. Each feature map is a three-dimensional tensor, whose spatial dimensions gradually decrease as the network depth increases, while the number of channels gradually increases. For ease of subsequent processing, these feature maps will be normalized to the same numerical range (such as between 0 and 1). In addition, the feature maps will be spatially upsampled so that their spatial dimensions are consistent with the original image, which can make it easier to perform feature alignment and fusion in subsequent steps.
[0059] Specifically, according to the geometric structure of the model rendering diagram of the texture to be pasted, the visual features in the feature map set are adaptively spatially transformed to obtain a set of transformed feature maps aligned with the model rendering diagram. This step uses a Spatial Transformer Network (STN) or a Deformable Convolutional Network (DCN). First, the system analyzes the geometric structure of the model rendering diagram and extracts key point and contour information. This is usually achieved through edge detection algorithms (such as the Canny edge detector) and contour extraction algorithms (such as the findContours function in OpenCV). The extracted geometric information is used to guide the parameter prediction of the transformation network. The Spatial Transformer Network consists of three main parts: a localization network, a grid generator, and a sampler. The localization network receives the model rendering diagram and the current feature map as inputs and outputs transformation parameters (such as an affine transformation matrix). The grid generator creates a sampling grid based on these parameters, and the sampler uses this grid to resample the feature map. This process is performed for each feature map in the feature map set, and the result is a set of transformed feature maps that are spatially aligned with the model rendering diagram. In the implementation, residual connections are also added to preserve the original feature information and prevent information loss caused by excessive deformation. In addition, to handle features at different scales, a multi-scale Spatial Transformer Network may be used, with each scale responsible for processing feature maps at the corresponding resolution.
[0060] Specifically, the set of transformed feature maps and the model rendering diagram are combined as conditional inputs and fed into a pre-trained deep learning diffusion model to obtain an initial noise map. This step first fuses the set of transformed feature maps and the model rendering diagram, usually using channel concatenation or an attention mechanism. The fused data is used as a conditional input to guide the generation process of the diffusion model. The diffusion model is a generative model based on gradually removing noise, and its core idea is to regard the image generation process as a process of gradually reducing noise. In the training phase, the model learns how to gradually add noise to a clear image until it becomes pure noise; in the generation phase, this process is performed in reverse. Specifically in this application, the diffusion model receives the fused conditional input and generates an initial noise map with the same size as the target texture. This initial noise map is a random noise distribution, usually following a Gaussian distribution, and its variance is determined according to a predefined noise schedule. The generation of the initial noise map marks the start of the denoising process and provides a starting point for subsequent iterative denoising.
[0061] Specifically, an iterative denoising process is performed on the initial noise map. In each iteration, a denoised image is generated or updated until a preset number of iterations is reached, and the final denoised image is used as the initial texture rendering map. This process is the core of the diffusion model for generating images. In each iteration, the model performs the following steps: First, according to the current noise level (determined by the preset noise schedule) and the conditional input, a neural network with a U-Net structure is used to predict the noise at the current time step. This U-Net network is the main component of the diffusion model. It processes the input at different spatial scales and can capture the multi-scale features of the image. The predicted noise is then subtracted from the current image to obtain an updated, less noisy image. This process is repeated multiple times, usually dozens to hundreds of iterations. The choice of the number of iterations needs to balance the generation quality and computational efficiency. Each iteration slightly reduces the noise in the image while gradually introducing more details and structures. The conditional input plays a crucial role throughout the process, ensuring that the generated image matches the visual features of the source picture and the geometric structure of the model. The image obtained from the final iteration is used as the initial texture rendering map, which combines the visual features of the source picture and the geometric structure of the model, laying the foundation for subsequent super-resolution processing and matte extraction.
[0062] Further, the step of combining the set of transformed feature maps and the model rendering map as conditional input and inputting them into a pre-trained deep learning diffusion model to obtain the initial noise map includes: combining the set of transformed feature maps and the model rendering map as conditional input and inputting them into a pre-trained deep learning diffusion model, where the deep learning diffusion model includes a conditional embedding module and a U-Net denoising network; using the conditional embedding module to encode the set of transformed feature maps and the model rendering map to obtain a conditional embedding vector, and generating a random noise map with the same size as the model rendering map according to a preset initial noise distribution; inputting the conditional embedding vector and the random noise map into the U-Net denoising network to obtain the predicted noise, and sampling and updating the random noise map according to the predicted noise and a preset noise schedule to obtain the initial noise map.
[0063] Specifically, first, the transformed feature map set and the model rendering map are combined as conditional inputs. This step first preprocesses the feature map set. Each feature map passes through a 1x1 convolutional layer to adjust the number of channels, making the number of channels of all feature maps consistent, usually set to 256 or 512. Then, the bilinear interpolation method is used to adjust the spatial resolution of all feature maps to be the same as that of the model rendering map. For the model rendering map, it first undergoes feature extraction through a small convolutional network (e.g., containing 3 3x3 convolutional layers, each followed by BatchNorm and ReLU activation functions) to obtain a feature representation with the same number of channels as the adjusted feature maps. Next, the channel concatenation method is adopted to concatenate the processed feature map set and the model rendering map features in the channel dimension. Specifically, if there are N feature maps and the number of channels of each feature map and the model rendering map feature is C, then the number of channels of the concatenated tensor will be (N + 1)*C. This concatenation operation preserves all the spatial information of the inputs, enabling the subsequent network to make full use of multi-scale features and model geometric information. Finally, a global average pooling operation is applied to the concatenated tensor to obtain a global feature vector, which will be used as the input to the conditional embedding module.
[0064] Specifically, next, the conditional embedding module is used to encode the transformed feature map set and the model rendering map to obtain a conditional embedding vector. The conditional embedding module is a multi-layer perceptron (MLP) network, usually containing 3 to 4 fully connected layers, each followed by a LayerNorm and a GELU activation function. The first fully connected layer maps the input global feature vector to a higher dimension (such as 1024 or 2048), the middle layers keep this dimension unchanged, and the last layer compresses the features to the desired embedding dimension (such as 256 or 512). Between each fully connected layer, a Dropout layer (with a dropout rate usually set to 0.1) is also inserted to prevent overfitting. In addition, to enhance the model's sensitivity to the input, a self-attention mechanism is introduced in the conditional embedding module. Specifically, before the last fully connected layer, a multi-head self-attention layer (usually with 8 attention heads) is used, allowing the model to establish long-range dependencies between different feature dimensions. Finally, the conditional embedding vector is normalized by a LayerNorm layer to stabilize the subsequent training process. At the same time, a random noise map is generated according to a preset initial noise distribution. Here, the standard normal distribution N(0, 1) is used as the initial noise distribution. The torch.randn function in PyTorch is used to generate a random tensor with the same spatial dimensions (height H and width W) and the same number of channels (usually 3, corresponding to the three RGB channels) as the model rendering map. Each element of this tensor is independently sampled from the standard normal distribution. To ensure that the generated noise is compatible with other parts of the model in terms of the numerical range, the noise is scaled, usually multiplied by a small coefficient (such as 0.1).
[0065] Finally, the conditional embedding vector and the random noise map are input into the U-Net denoising network. The encoder part of the U-Net network contains 4 downsampling blocks, each block containing two 3x3 convolutional layers, each convolutional layer followed by a BatchNorm and a LeakyReLU activation function, and then a 2x2 max pooling layer. The decoder part correspondingly has 4 upsampling blocks, each block first using a transposed convolution for upsampling, then concatenating with the feature map of the corresponding layer in the encoder, and then followed by two 3x3 convolutional layers, also followed by a BatchNorm and a LeakyReLU. The conditional embedding vector is injected into each convolutional block by addition or multiplication. Specifically, the embedding vector is mapped to the same number of channels as the feature map through a linear layer, and then element-wise addition or multiplication is performed. The output of the U-Net is the predicted noise, which has the same dimension as the input noise map. Then, the random noise map is sampled and updated according to the predicted noise and a preset noise schedule. The noise schedule uses a cosine decay function:
[0066] ;
[0067] where \(t\) is the current step number, \(T\) is the total number of steps, and \(s\) is a small constant (such as 0.008). The update formula is:
[0068] ;
[0069] , where is the predicted noise, is additional random noise. This update process is repeated multiple times (usually 50 to 100 times) to finally obtain the initial noise map.
[0070] 102. Use the initial texture rendering map as the target texture rendering map, and perform matte extraction on the target texture rendering map to obtain a matte bitmap;
[0071] In an embodiment of the present invention, after using the initial texture rendering map as the target texture rendering map, it further includes: determining whether the initial texture rendering map used as the target texture rendering map is a high-resolution rendering map; if not, performing multi-scale decomposition on the initial texture rendering map to obtain a set of sub-images with different frequency components, and performing upsampling processing on each sub-image in the set of sub-images according to a preset super-resolution factor to obtain an enlarged set of sub-images; inputting the enlarged set of sub-images into a pre-trained super-resolution diffusion model, performing iterative denoising processing on each enlarged sub-image to obtain a set of sub-images with enhanced high-frequency details; performing adaptive fusion on the set of sub-images with enhanced high-frequency details to obtain a preliminary super-resolution image, and using an edge-aware detail enhancement network to process the preliminary super-resolution image to obtain a high-resolution texture image; updating the target texture rendering map according to the high-resolution texture image.
[0072] Specifically, the initial texture rendering map generated by the deep learning diffusion model according to the source image and the model rendering map of the texture to be pasted can directly be a high-resolution texture. However, in the case of directly outputting a high-resolution texture, the cost of texture generation is relatively high. Therefore, considering the cost, the initial texture rendering map may be output as a high-resolution texture or a low-resolution texture. When it is output as a low-resolution texture, it is necessary to perform ultra-high-resolution processing on the initial texture rendering map, and use the new texture rendering map obtained by the processing as the target texture rendering map.
[0073] First, perform multi-scale decomposition on the initial texture rendering image to obtain a set of sub-images with different frequency components. This step uses the Laplacian Pyramid algorithm for multi-scale decomposition. Specifically, when implementing, first blur the original image with a 5x5 Gaussian kernel, and then perform 2x downsampling to obtain the low-frequency sub-image. Upsample this low-frequency sub-image back to the original size and subtract it from the original image to obtain the high-frequency sub-image. This process is repeated recursively, usually obtaining 4 to 5 scales of sub-images. Each sub-image represents the detailed information of the original image within a specific frequency range. Next, upsample each sub-image according to a preset super-resolution factor (usually 2x or 4x). Here, the Bicubic Interpolation method is used for upsampling, which can better preserve the edge information while maintaining the smoothness of the image. For each sub-image, first calculate the interpolation kernel, and then perform a convolution operation on the image to obtain the enlarged sub-image. This process will preserve the unique frequency characteristics of each sub-image, providing a basis for subsequent detail enhancement.
[0074] Next, input the set of enlarged sub-images into a pre-trained super-resolution diffusion model. This model adopts the structure of a Conditional Diffusion Probabilistic Model, which includes a U-Net backbone network and a time embedding module. The U-Net network contains 4 downsampling blocks and 4 upsampling blocks, each block contains 3 residual blocks, and each residual block contains two 3x3 convolutional layers, with the SiLU activation function used in the middle. The time embedding module uses the sine position encoding method to convert the time step into a 512-dimensional embedding vector. For each enlarged sub-image, the model performs a 200-step iterative denoising process. In each step, the model predicts the noise residual, and then updates the image state according to a predefined noise schedule (using a cosine decay function). Through this iterative process, the model gradually enhances the high-frequency details of each sub-image, generating a set of sub-images with enhanced high-frequency details.
[0075] Finally, perform adaptive fusion on the set of sub-images with enhanced high-frequency details to obtain a preliminary super-resolution image. This step uses a Multi-scale Attention Fusion Network. This network first applies a 3x3 convolutional layer to each sub-image to extract features, and then uses a channel attention mechanism and a spatial attention mechanism to calculate channel weights and spatial weights respectively. The channel attention is achieved through global average pooling and max pooling followed by two fully connected layers, and the spatial attention is achieved through a convolutional layer and a Sigmoid activation function. Multiply these two attention weights with the feature map to obtain a weighted feature map. Then, use a deconvolution layer to upsample all the weighted feature maps to the same spatial resolution and fuse them by element-wise addition. Finally, apply a 1x1 convolutional layer to adjust the number of channels to obtain a preliminary super-resolution image. Next, process it using an edge-aware detail enhancement network. This network adopts a residual learning structure and contains 20 residual blocks, each of which contains two 3x3 convolutional layers and a LeakyReLU activation function. An edge detection branch is introduced in the middle layer of the network, and the Sobel operator is used to extract edge information and fuse it with the backbone feature map. The last layer uses a Tanh activation function to ensure that the output values are in the range of [-1, 1]. This network can further enhance the edge and texture details of the image and finally output a high-quality high-resolution texture map image.
[0076] In one embodiment of the present invention, the process of matting the target texture rendering map to obtain a matted bitmap includes: generating a texture model rendering map with the same shape and perspective according to the shape and perspective of the high-resolution texture map image; performing per-pixel difference calculation on the high-resolution texture map image and the texture model rendering map, and taking the absolute value of the difference result and summing it over multiple color channels to obtain a difference matrix; performing binarization processing on the difference matrix according to a preset threshold, retaining the pixels greater than the threshold and setting the pixels less than the threshold to be transparent to obtain a preliminary matted bitmap; applying the Sobel operator to the preliminary matted bitmap for edge detection and removing abnormal edges to obtain a matted bitmap.
[0077] Specifically, first, according to the shape and perspective of the high-resolution texture image, a rendering of the model to be textured with the same shape and perspective is generated. This step is implemented using the OpenGL rendering engine. In the specific process, first, the metadata of the high-resolution texture image is read to extract its resolution and camera parameters. Then, the pre-prepared 3D building block model is loaded into the OpenGL environment. The virtual camera parameters are set to match the perspective of the texture image, including the camera position, orientation, and field of view. To ensure the rendering quality, multisample anti-aliasing and anisotropic filtering are enabled. A parallel light source is used to simulate natural lighting, and ambient occlusion is applied to enhance the sense of depth. No textures are applied during the rendering process, only retaining the geometric information of the model. The purpose of doing this is to generate a rendering that exactly corresponds to the original texture image in terms of spatial structure but only contains geometric information. This rendering will be used as a reference for the subsequent matte extraction process to help identify which areas belong to the model itself and which areas are the added texture content. In this way, the texture content can be precisely separated from the geometric shape of the model, laying a foundation for the subsequent matte extraction process.
[0078] Next, a per-pixel difference calculation is performed on the high-resolution texture image and the rendering of the model to be textured. First, the two images are converted to the Lab color space because the Lab space is more in line with human visual perception. Then, the pixel differences between the two images are calculated, and the absolute value of the difference result is taken to eliminate the influence of positive and negative differences. Subsequently, a weighted sum is performed on the three channels in the Lab space to merge the three channels into a scalar value representing the overall difference degree at each pixel position. The use of a weighted sum instead of a simple sum is because the human eye has different sensitivities to brightness and different colors. Usually, a higher weight is given to the L channel because the human eye is more sensitive to brightness changes. The final resulting difference matrix has the same spatial dimensions as the input images but only has one channel. This difference matrix clearly shows the differences between the original texture image and the rendering of the pure geometric model, effectively highlighting the added texture content. Through this method, the modifications and additions made by the texture to the original model can be accurately identified, providing a key information basis for the subsequent matte extraction process.
[0079] Then, the difference matrix is binarized according to a preset threshold. The threshold is adaptively determined by Otsu's method. The core idea of Otsu's method is to find the optimal threshold by maximizing the between-class variance. The specific implementation process includes calculating the histogram of the difference matrix to obtain the pixel number distribution of gray levels. Then, calculate the total number of pixels, the cumulative sum array, and the cumulative mean array. Next, traverse all possible thresholds. For each threshold, divide the image into foreground and background parts and calculate the between-class variance. Select the threshold that maximizes the between-class variance as the final binarization threshold. The advantage of this method is that it can automatically select the optimal threshold according to the actual gray distribution of the image without manual intervention and is applicable to images with different lighting conditions and contents. The difference matrix is binarized using the determined threshold. Pixels greater than the threshold are retained (set to 1), and pixels less than the threshold are set to transparent (set to 0) to obtain a preliminary matte bitmap. This preliminary matte bitmap clearly distinguishes the texture area to be retained and the background area to be removed.
[0080] Finally, apply the Sobel operator to the preliminary matte bitmap for edge detection to remove abnormal edges and obtain the matte bitmap. The Sobel operator consists of two 3x3 convolution kernels, which are used to detect edges in the horizontal and vertical directions respectively. By performing a convolution operation on the preliminary matte bitmap and calculating the gradient magnitude, significant edges can be identified. Set a gradient threshold, and set the edge pixels with a gradient magnitude less than the threshold to transparent in the original mask, thereby removing abnormal edges and obtaining the final matte bitmap. This process can effectively eliminate some noise and discontinuous edges generated by thresholding, making the matte result smoother and more accurate.
[0081] 103. Use the matte bitmap as a projected texture, and perform projection mapping on the projected texture according to the block structure of a preset assembled building block model to achieve automatic texturing of the assembled building block model and obtain the assembled building block model with texturing completed.
[0082] In an embodiment of the present invention, the process of using the matte bitmap as a projection texture map and performing projection mapping on the projection texture map according to the block structure of a preset assembled building block model to automatically texture the assembled building block model and obtain a textured assembled building block model includes: using the matte bitmap as a projection texture map and detecting whether there is an adjustment instruction for the matte bitmap used as the projection texture map. If there is, perform vector graph conversion and adjustment on the matte bitmap according to the adjustment instruction to obtain a corresponding textured vector graph, and update the projection texture map according to the textured vector graph; according to the preset block structure of the assembled building block model, perform segmentation processing on the projection texture map to obtain a set of sub-projection texture maps corresponding to each building block of the block structure; perform geometric transformation and projection calculation on each sub-projection texture map in the set of sub-projection texture maps to obtain a projection texture map that matches the surface of the corresponding building block; according to the three-dimensional geometric information of the building block, perform UV coordinate mapping on the projection texture map to obtain texture coordinate information, and associate the texture coordinate information with the three-dimensional model of the corresponding building block to obtain a set of building block models with texture information; perform assembly and rendering processing on the set of building block models with texture information to obtain a textured assembled building block model.
[0083] Specifically, first, use the matte bitmap as a projection texture map and detect whether there is an adjustment instruction for the matte bitmap used as the projection texture map. The purpose of this step is to provide an opportunity for manual intervention and adjustment, increasing the flexibility of the generation process. The process of detecting the adjustment instruction involves the design of the user interface and the implementation of the interaction logic. The system continuously listens for user input, including interaction behaviors such as mouse clicks, keyboard operations, or touch screen touches. When a relevant adjustment instruction is detected, the system triggers the corresponding processing flow. If there is an adjustment instruction, perform vector graph conversion and adjustment on the matte bitmap according to the instruction. Vector graph conversion usually uses contour tracing algorithms, such as the Moore-Neighbor algorithm, to extract the edge contour of the bitmap, and then uses the Douglas-Peucker algorithm to simplify the contour. The simplified contour points are fitted with Bezier curves to generate a smooth vector path. The adjustment process may include operations such as path editing, color modification, adding or deleting elements, etc., which all require the support of specialized vector graph editing tools. After the adjustment is completed, the system regenerates the bitmap according to the modified vector graph and updates the projection texture map. This process not only allows designers to make fine adjustments but also ensures the continuity and consistency of subsequent processing.
[0084] Specifically, according to the segmented structure of the preset assembled building block model, the projection map is segmented to obtain a set of sub-projection maps corresponding to each building block of the segmented structure. The implementation of this step involves complex geometric calculations and spatial partitioning algorithms. Specifically, it is necessary to first analyze the segmented structure information of the assembled building block model, which usually includes the position, size, shape, and relative relationship of each building block. Then, project this three-dimensional spatial information onto a two-dimensional plane to form a plane segmentation scheme corresponding to the projection map. Next, use computational geometry algorithms, such as plane segmentation trees or region growing algorithms, to segment the projection map. During the segmentation process, it is necessary to consider the connection relationships and overlapping situations between the building blocks to ensure that the segmentation results can accurately reflect the structure of the three-dimensional model. For complex curved surfaces or irregularly shaped building blocks, more advanced surface segmentation algorithms may be required. After segmentation, each sub-projection map corresponds to a specific building block and retains part of the information of the original map. The key to this step is to maintain the accuracy and integrity of the segmentation to ensure that the subsequent projection mapping can accurately correspond to each building block.
[0085] Specifically, perform geometric transformation and projection calculation on each sub-projection map in the set of sub-projection maps to obtain a projection map that matches the surface of the corresponding building block. This process involves complex three-dimensional to two-dimensional projection transformation. First, it is necessary to obtain the three-dimensional geometric information of each building block, including its position, orientation, and the normal vector of the surface in the entire model. Then, apply a series of geometric transformations to each sub-projection map, including rotation, scaling, and translation, to make it consistent with the surface shape and orientation of the corresponding building block. This step is usually implemented using an affine transformation matrix. Next, perform projection calculation to project the transformed two-dimensional projection map onto the surface of the building block in three-dimensional space. This projection process needs to consider factors such as perspective and lighting to ensure that the projected map visually matches the surface of the building block perfectly. For curved surfaces or irregularly shaped building blocks, more complex non-linear projection methods, such as spherical projection or cylindrical projection, may be required. During the entire process, special attention needs to be paid to maintaining the continuity of the map, especially at the edges and corners of the building blocks, to avoid obvious breaks or distortions. The result of this step is a set of projection maps that perfectly match the surfaces of each building block, preparing for the subsequent UV coordinate mapping.
[0086] Specifically, based on the three-dimensional geometric information of the building blocks, perform UV coordinate mapping on the projection texture map to obtain texture coordinate information, and associate the texture coordinate information with the three-dimensional models of the corresponding building blocks to obtain a set of building block models with texture information. UV coordinate mapping is the process of unfolding the surface of a three-dimensional model onto a two-dimensional plane, and this step is crucial for ensuring that the texture is correctly applied to the three-dimensional model. First, a UV unwrapping map needs to be created for each building block. This process usually involves complex surface parameterization algorithms, such as the minimum stretch energy method or the spectral method. For building blocks with regular shapes, simple planar projection can be used, while for complex shapes, more advanced methods, such as spherical projection or cylindrical projection, may be required. Once the UV unwrapping map is created, the pixels of the projection texture map need to be mapped into this UV space. This process needs to take into account the resolution of the texture map and the surface area of the building block to ensure that the texture does not stretch or compress. During the mapping process, the edges of the texture map also need to be processed to ensure that the texture smoothly transitions at the seams of the building blocks. After completing the UV coordinate mapping, the obtained texture coordinate information is associated with the three-dimensional models of the building blocks. This is usually achieved by adding UV coordinates to the vertex data of the three-dimensional model. Finally, each building block becomes a three-dimensional model with accurate texture information, forming a complete set of building block models with texture information.
[0087] Specifically, assemble and render the set of building block models with texture information to obtain the assembled building block model with the texture completed. This step involves the construction of a three-dimensional scene and high-quality rendering. First, according to the preset assembly structure, place each building block model with texture information in the correct position and orientation. This process requires precise spatial transformation calculations to ensure that each building block can be perfectly joined together. During the assembly process, the contact surfaces between the building blocks also need to be processed to ensure that there are no visual gaps or overlaps. Next, set the lighting environment of the scene, including ambient light, direct light, and indirect light, to create a realistic lighting effect. It may also be necessary to add Ambient Occlusion to enhance the three-dimensional sense of the model. Then, apply advanced rendering techniques, such as Physically Based Rendering (PBR), considering properties such as the reflectivity and roughness of the material, to make the texture and the building block material look more realistic. For transparent or semi-transparent building blocks, the refraction and scattering of light also need to be correctly processed. Finally, use high-quality anti-aliasing techniques, such as MSAA or FXAA, to render the entire scene to obtain the final high-quality image or an interactive 3D model. This final assembled building block model not only contains accurate geometric information but also incorporates highly customized texture designs, presenting a unique visual effect.
[0088] Further, performing geometric transformation and projection calculation on each sub-projection texture map in the sub-projection texture map set to obtain a projection texture map matching the surface of the corresponding building block includes: performing edge detection and contour analysis on each sub-projection texture map in the sub-projection texture map set to obtain a control point set representing the main contour of the sub-projection texture map; calculating the normal vector and main direction of the surface of the building block according to the three-dimensional geometric information of the corresponding building block to obtain a surface parametric description; using the control point set and the surface parametric description to construct a projection mapping function and calculate the mapping relationship from the sub-projection texture map to the surface of the building block; and performing geometric transformation and projection calculation on the sub-projection texture map according to the mapping relationship to obtain a projection texture map matching the surface of the corresponding building block.
[0089] Specifically, first, perform edge detection and contour analysis on each sub-projection texture map in the sub-projection texture map set to obtain a control point set representing the main contour of the sub-projection texture map. The implementation of this step involves complex image processing and geometric analysis techniques. Specifically, first, the projection texture map needs to be converted into a high-resolution bitmap format to apply traditional image processing algorithms. Then, the Canny edge detection operator is used to perform edge detection on the image. The Canny operator first performs Gaussian blurring on the image to reduce noise, then calculates the gradient magnitude and direction of the image, then performs non-maximum suppression to thin the edges, and finally obtains the final edge map through double thresholding and edge connection. After obtaining the edge map, a contour tracking algorithm, such as the Moore-Neighbor algorithm, is used to extract continuous contour lines. To reduce the data volume and retain key features, polygon approximation is performed on the extracted contours, which is usually implemented using the Douglas-Peucker algorithm. This algorithm simplifies the contour by recursively dividing the curve into line segments and removing points with small deviations. Finally, curve fitting is performed on the simplified contour, usually using cubic spline interpolation, to obtain a set of control point sets that can accurately represent the main contour of the sub-projection texture map. These control points not only retain the key features of the image but also greatly reduce the data volume, providing an efficient input for subsequent projection mapping.
[0090] Specifically, based on the three-dimensional geometric information of the corresponding building block, the normal vector and the main direction of the building block surface are calculated to obtain a surface parameterization description, which involves complex three-dimensional geometric calculations and surface analysis techniques. First, the geometric information of a single building block needs to be extracted from the three-dimensional model data, including vertex coordinates, patch information, etc. Then, the normal vector of each patch is calculated, usually by taking the cross product of vectors of three non-collinear vertices on the patch. For curved surface building blocks, the average normal vector needs to be calculated at each vertex, which can be achieved by weighted averaging the normal vectors of adjacent patches. Next, the principal component analysis (PCA) method is used to determine the main direction of the building block surface. PCA calculates the covariance matrix of vertex coordinates and performs eigenvalue decomposition on it to obtain the main direction vectors of the surface. These main direction vectors are crucial for the subsequent parameterization process. Surface parameterization is the process of mapping a three-dimensional surface to a two-dimensional plane, usually using isometric parameterization or minimum distortion parameterization methods. For complex shapes, the surface may need to be divided into multiple segments, and each segment is parameterized separately and then spliced. The parameterization process needs to solve an optimization problem, with the goal of minimizing the deformation energy while keeping the area and angle distortion minimized. The finally obtained surface parameterization description contains the mapping relationship from three-dimensional space to two-dimensional parameter space, providing a basis for subsequent texture mapping projection.
[0091] Specifically, using the control point set and the surface parameterization description, a projection mapping function is constructed, and calculating the mapping relationship between the sub-projection texture and the building block surface involves complex mathematical modeling and optimization calculations. First, the control points of the sub-projection texture need to be located in the parameterized two-dimensional space. This is usually achieved by minimizing the correspondence error between the control points in the two-dimensional space and the three-dimensional surface, and non-linear least squares methods such as the Levenberg-Marquardt algorithm can be used. Once the positions of the control points in the parameter space are determined, a mapping function from the two-dimensional parameter space to the three-dimensional surface can be constructed. This mapping function is usually represented by spline interpolation or radial basis function (RBF). For spline interpolation, bicubic splines or NURBS surfaces can be used; for RBF, common basis functions include multiquadric functions or Gaussian functions. The process of constructing the mapping function is actually to solve an interpolation problem, which needs to ensure that the mapping function exactly matches at the control points and smoothly transitions in other regions. In addition, the bijectivity of the mapping needs to be considered to ensure that there are no folds or overlaps. To improve the quality of the mapping, additional constraint conditions may need to be introduced, such as maintaining the area ratio or minimizing the distortion. The finally obtained projection mapping function not only needs to accurately reflect the correspondence between the sub-projection texture and the building block surface, but also needs to ensure the continuity and smoothness of the mapping, laying a foundation for subsequent texture mapping projection.
[0092] Specifically, according to the mapping relationship, performing geometric transformation and projection calculation on the sub-projection texture map to obtain a projection texture map that matches the surface of the corresponding building block involves complex graphics transformation and rendering techniques. First, the sub-projection texture map needs to be discretized into a high-resolution bitmap for pixel-level transformation. Then, for each pixel in the bitmap, use the previously constructed projection mapping function to calculate its corresponding position on the surface of the building block. This process is actually an inverse mapping, and numerical methods such as Newton's method are needed to solve the non-linear equation. For the mapped position, interpolation calculation needs to be performed to obtain accurate color values, usually using bilinear interpolation or bicubic interpolation. When performing projection, the curvature of the building block surface and the viewing angle also need to be considered, and the projection result needs to be appropriately deformed and distorted to ensure that it looks correct in three-dimensional space. In addition, occlusion and edge problems also need to be handled to ensure that the texture map can smoothly transition at the edges and corners of the building block. To improve the rendering quality, anti-aliasing techniques such as super-sampling anti-aliasing (SSAA) or multi-sampling anti-aliasing (MSAA) may also need to be applied. Finally, combine the processed texture map information with the geometric information of the building block to generate a complete three-dimensional building block model with accurate texture mapping. This final projection texture map not only needs to visually match the surface of the building block perfectly, but also needs to ensure that the correct appearance can be maintained under different angles and lighting conditions.
[0093] In this embodiment, by obtaining the source picture and the model rendering picture to be textured, and generating an initial texture rendering picture according to the source picture and the model rendering picture to be textured through a deep learning diffusion model; taking the initial texture rendering picture as the target texture rendering picture, and performing matte extraction processing on the target texture rendering picture to obtain a matte bitmap; taking the matte bitmap as the projection texture map, and performing projection mapping on the projection texture map according to the preset block structure of the assembled building block model to realize automatic texturing of the assembled building block model, and obtaining an assembled building block model with texturing completed. This method combines deep learning technology and image processing technology to realize the full-automatic generation process from user input to the final textured model, significantly improving the efficiency and quality of building block texture design, and solving the problem of time-consuming and laborious traditional manual design methods.
[0094] The method for generating an assembled building block model in the embodiments of the present invention has been described above. Next, the device for generating an assembled building block model in the embodiments of the present invention will be described. Please refer to Figure 2 , an embodiment of the device for generating an assembled building block model in the embodiments of the present invention includes:
[0095] An image diffusion module 201, configured to obtain a source picture and a model rendering picture to be textured, and generate an initial texture rendering picture according to the source picture and the model rendering picture to be textured through a deep learning diffusion model;
[0096] The matte extraction module 202 is used to render the initial texture map rendering as the target texture map rendering, and perform matte extraction processing on the target texture map rendering to obtain a matte bitmap;
[0097] The texture map projection module 203 is used to use the matte bitmap as the projection texture map, and perform projection mapping on the projection texture map according to the preset block structure of the assembled building block model, so as to realize automatic texture mapping of the assembled building block model and obtain the assembled building block model with texture mapping completed.
[0098] In the embodiment of the present invention, the generating device of the assembled building block model runs the above-mentioned generating method of the assembled building block model. The generating device of the assembled building block model obtains a source picture and a model rendering to be textured, and generates an initial texture map rendering according to the source picture and the model rendering to be textured through a deep learning diffusion model; renders the initial texture map rendering as the target texture map rendering, and performs matte extraction processing on the target texture map rendering to obtain a matte bitmap; uses the matte bitmap as the projection texture map, and performs projection mapping on the projection texture map according to the preset block structure of the assembled building block model, so as to realize automatic texture mapping of the assembled building block model and obtain the assembled building block model with texture mapping completed. This method combines deep learning technology and image processing technology to realize the full-automatic generation process from user input to the final textured model, significantly improves the efficiency and quality of building block texture design, and solves the problem of time-consuming and laborious traditional manual design methods.
[0099] Above Figure 2 The generating device of the assembled building block model in the embodiment of the present invention is described in detail from the perspective of modular functional entities. Next, the generating device of the assembled building block model in the embodiment of the present invention will be described in detail from the perspective of hardware processing.
[0100] Figure 3FIG. 0 is a schematic structural diagram of a generating device for an assembled building block model provided by an embodiment of the present invention. The generating device 300 for the assembled building block model may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPU) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 for storing application programs 333 or data 332 (for example, one or more mass storage device terminals). Among them, the memory 320 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the generating device 300 for the assembled building block model. Further, the processor 310 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the generating device 300 for the assembled building block model to implement the steps of the above-mentioned method for generating an assembled building block model.
[0101] The generating device 300 for the assembled building block model may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 331, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art can understand that Figure 3 The shown structural diagram of the generating device for the assembled building block model does not limit the generating device for the assembled building block model provided by the present invention, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.
[0102] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium, and when the instructions run on a computer, the computer is caused to execute the steps of the method for generating an assembled building block model.
[0103] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system or device and unit can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.
[0104] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0105] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of the present invention.
Claims
1. A method for generating an assembled building block model, characterized in that The method for generating the assembled building block model includes: Obtain a source image and a model rendering to be textured, and generate an initial texture rendering through a deep learning diffusion model based on the source image and the model rendering to be textured; Use the initial texture rendering as the target texture rendering, and perform matte extraction on the target texture rendering to obtain a matte bitmap; Use the matte bitmap as a projection texture, and perform projection mapping on the projection texture according to the preset block structure of the assembled building block model to achieve automatic texturing of the assembled building block model, and obtain an assembled building block model with texturing completed; The step of using the matte bitmap as a projection texture and performing projection mapping on the projection texture according to the preset block structure of the assembled building block model to achieve automatic texturing of the assembled building block model and obtain an assembled building block model with texturing completed includes: using the matte bitmap as a projection texture, and detecting whether there is an adjustment instruction for the matte bitmap used as the projection texture. If so, perform vector graph conversion and adjustment on the matte bitmap according to the adjustment instruction to obtain a corresponding texture vector graph, and update the projection texture according to the texture vector graph; According to the preset block structure of the assembled building block model, perform segmentation processing on the projection texture to obtain a set of sub-projection textures corresponding to each building block of the block structure; Perform geometric transformation and projection calculation on each sub-projection texture in the set of sub-projection textures to obtain a projection texture matching the surface of the corresponding building block; According to the three-dimensional geometric information of the building block, perform UV coordinate mapping on the projection texture to obtain texture coordinate information, and associate the texture coordinate information with the three-dimensional model of the corresponding building block to obtain a set of building block models with texture information; Perform assembly and rendering processing on the set of building block models with texture information to obtain an assembled building block model with texturing completed.
2. The generation method of the assembled building block model according to claim 1, characterized in that, The step of obtaining a source image and a model rendering to be textured, and generating an initial texture rendering through a deep learning diffusion model based on the source image and the model rendering to be textured includes: Obtain a source image and a model rendering to be textured, and perform multi-scale feature extraction on the source image to obtain a set of feature maps containing multi-scale visual features; According to the geometric structure of the model rendering to be textured, perform adaptive spatial transformation on the visual features in the set of feature maps to obtain a set of transformed feature maps aligned with the model rendering; Combine the set of transformed feature maps and the model rendering as conditional inputs, and input them into a pre-trained deep learning diffusion model to obtain an initial noise map; Perform an iterative denoising process on the initial noise map, generating or updating a denoised image in each iteration until a preset number of iterations is reached, and use the denoised image corresponding to the number of iterations as the initial texture rendering.
3. The generation method of the assembled building block model according to claim 2, characterized in that, The step of combining the set of transformed feature maps and the model rendering as conditional inputs, and inputting them into a pre-trained deep learning diffusion model to obtain an initial noise map includes: Combine the transformed feature map set and the model rendering map as conditional inputs and input them into a pre-trained deep learning diffusion model, where the deep learning diffusion model includes a conditional embedding module and a U-Net denoising network; Use the conditional embedding module to encode the transformed feature map set and the model rendering map to obtain a conditional embedding vector, and generate a random noise map with the same size as the model rendering map according to a preset initial noise distribution; Input the conditional embedding vector and the random noise map into the U-Net denoising network to obtain a predicted noise, and sample and update the random noise map according to the predicted noise and a preset noise schedule to obtain an initial noise map.
4. The method for generating an assembled building block model according to claim 1, wherein After using the initial texture map rendering as the target texture map rendering, it further includes: Determine whether the initial texture map rendering as the target texture map rendering is a high-resolution rendering; If not, perform multi-scale decomposition on the initial texture map rendering to obtain a set of sub-images with different frequency components, and perform upsampling processing on each sub-image in the set of sub-images according to a preset super-resolution factor to obtain an enlarged set of sub-images; Input the enlarged set of sub-images into a pre-trained super-resolution diffusion model, and perform iterative denoising processing on each enlarged sub-image to obtain a set of sub-images with enhanced high-frequency details; Perform adaptive fusion on the set of sub-images with enhanced high-frequency details to obtain a preliminary super-resolution image, and use an edge-aware detail enhancement network to process the preliminary super-resolution image to obtain a high-resolution texture image; Update the target texture map rendering according to the high-resolution texture image.
5. The method for generating an assembled building block model according to claim 4, wherein The matte processing of the target texture map rendering to obtain a matte bitmap includes: Generate a to-be-textured model rendering with the same shape and perspective according to the shape and perspective of the high-resolution texture image; Perform per-pixel difference calculation on the high-resolution texture image and the to-be-textured model rendering, and sum the multiple color channels after taking the absolute value of the difference result to obtain a difference matrix; Perform binarization processing on the difference matrix according to a preset threshold, retain the pixels greater than the threshold, and set the pixels less than the threshold to be transparent to obtain a preliminary matte bitmap; Apply the Sobel operator to the preliminary matte bitmap for edge detection to remove abnormal edges and obtain a matte bitmap.
6. The method for generating an assembled building block model according to claim 1, wherein, The geometric transformation and projection calculation of each sub-projection texture in the sub-projection texture set to obtain a projection texture matching the surface of the corresponding building block includes: Perform edge detection and contour analysis on each sub-projection texture in the sub-projection texture set to obtain a control point set representing the main contour of the sub-projection texture; Calculate the normal vector and the main direction of the surface of the building block according to the three-dimensional geometric information of the corresponding building block to obtain a surface parameterization description; Use the control point set and the surface parameterization description to construct a projection mapping function and calculate the mapping relationship from the sub-projection texture to the surface of the building block; According to the mapping relationship, perform geometric transformation and projection calculation on the sub-projection texture to obtain a projection texture matching the surface of the corresponding building block.
7. A generating device for an assembled building block model, characterized in that, The generating device of the assembled building block model includes: An image diffusion module, configured to obtain a source picture and a model rendering picture to be textured, and generate an initial textured rendering picture according to the source picture and the model rendering picture to be textured through a deep learning diffusion model; A matte extraction module, configured to use the initial textured rendering picture as a target textured rendering picture, and perform matte extraction processing on the target textured rendering picture to obtain a matte bitmap; A texture projection module, configured to use the matte bitmap as a projected texture, and perform projection mapping on the projected texture according to the preset block structure of the assembled building block model to achieve automatic texturing of the assembled building block model, and obtain an assembled building block model with texturing completed; the step of using the matte bitmap as a projected texture and performing projection mapping on the projected texture according to the preset block structure of the assembled building block model to achieve automatic texturing of the assembled building block model and obtain an assembled building block model with texturing completed includes: using the matte bitmap as a projected texture, and detecting whether there is an adjustment instruction for the matte bitmap used as the projected texture, if so, performing vector graph conversion and adjustment on the matte bitmap according to the adjustment instruction to obtain a corresponding textured vector graph, and updating the projected texture according to the textured vector graph; according to the preset block structure of the assembled building block model, performing segmentation processing on the projected texture to obtain a set of sub-projected textures corresponding to each building block of the block structure; performing geometric transformation and projection calculation on each sub-projected texture in the set of sub-projected textures to obtain a projected texture matching the surface of the corresponding building block; according to the three-dimensional geometric information of the building block, performing UV coordinate mapping on the projected texture to obtain texture coordinate information, and associating the texture coordinate information with the three-dimensional model of the corresponding building block to obtain a set of building block models with texture information; performing assembly and rendering processing on the set of building block models with texture information to obtain an assembled building block model with texturing completed.
8. A generating device for an assembled building block model, characterized in that, The generating device of the assembled building block model includes: a memory and at least one processor, and instructions are stored in the memory; The at least one processor calls the instructions in the memory to enable the generating device of the assembled building block model to execute the steps of the generating method of the assembled building block model according to any one of claims 1-6.
9. A computer-readable storage medium having instructions stored thereon, characterized in that, When the instructions are executed by the processor, the steps of the generating method of the assembled building block model according to any one of claims 1-6 are implemented.
Citation Information
Patent Citations
Face image processing method, user equipment, storage medium and device
CN110688962A
Image super-resolution method based on reference image
CN117372259A
Map generation method and device, equipment, computer program product and storage medium
CN118135114A