Three-dimensional model reconstruction method based on neural radiation field

By combining conditional diffusion model, CNN-MLP optimization and attention style transfer technology, end-to-end design from sketch generation to high-quality rendering is realized, solving the problems of inefficiency and style structure dislocation in the traditional design process, and providing efficient and personalized design tools.

CN120298570APending Publication Date: 2025-07-11THE INST OF AUTOMATION HEILONGJIANG ACADEMY OF SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510202026.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Traditional design processes rely on manual experience, inefficient design efficiency, limited generation results, insufficient diversity and personalized response capabilities, and style transfer technology ignores local details, resulting in misalignment of artistic style and design structure.

Method used

Combining the conditional diffusion model, CNN-MLP optimization module and attention-enhanced style transfer technology, a multi-stage collaborative framework is used to achieve dynamic conditional control and local consistency optimization from sketch generation to high-quality rendering, ensuring that the generated results meet user needs and style consistency.

Benefits of technology

It realizes end-to-end design from concept to finished product, improves the diversity, efficiency and quality of design generation, solves the conflict between style and structure in traditional methods, and provides efficient personalized design tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298570A_ABST
    Figure CN120298570A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional model reconstruction method based on a neural radiation field, and relates to the technical field of model reconstruction. The reconstruction method comprises the following steps of: combining a conditional diffusion model with an attention-enhanced style migration technology to realize a full-process controllable design from sketch generation to high-quality rendering: generating a diversified sketch by the diffusion model through noise iteration and conditional embedding, and breaking through the mode collapse limitation of a traditional model; the CNN-MLP optimization module improves the line and structure rationality through a joint loss function; the attention mechanism driven style migration technology accurately balances the consistency of global styles and local details; according to the method, a multi-stage collaborative framework is provided, a diffusion model, CNN-MLP optimization, attention style migration and GAN super-resolution rendering are integrated into a unified system for the first time, and end-to-end design from a concept to a finished product is achieved; a dynamic condition control mechanism is developed, and user keywords and style labels are embedded in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of model reconstruction, and specifically relates to a three-dimensional model reconstruction method based on neural radiance fields. Background Art

[0002] The current creative design field faces multiple technical bottlenecks: The traditional design process highly depends on manual experience, resulting in low design efficiency, and the generated results are restricted by the subjective cognition of designers, making it difficult to break through the inherent style paradigm. Although existing generation models can achieve automated design, their generated content lacks diversity, and their ability to respond to user personalized needs is weak. In addition, style transfer technology often ignores local details due to global optimization, causing a dislocation between the artistic style and the design structure, further restricting the innovation and practicality of design results. Summary of the Invention

[0003] To solve the problems mentioned in the above background art, the purpose of the present invention is to provide a three-dimensional model reconstruction method based on neural radiance fields.

[0004] A three-dimensional model reconstruction method based on neural radiance fields of the present invention has a reconstruction method as follows: Combining a conditional diffusion model with an attention-enhanced style transfer technology realizes a full-process controllable design from sketch generation to high-quality rendering: The diffusion model generates diverse sketches through noise iteration and conditional embedding, breaking through the mode collapse limitation of traditional models; The CNN-MLP optimization module improves the rationality of lines and structures through a joint loss function; The style transfer technology driven by the attention mechanism precisely balances the consistency between the global style and local details.

[0005] Preferably, the sketch generation generates diverse initial sketches that meet user requirements through a diffusion model; the specific process is as follows: The system initializes a Gaussian noise image as the starting point and gradually denoises it through multiple steps of iteration of the diffusion model to generate a sketch; The diffusion model converts random noise into a structured image through the reverse diffusion process, and each step of iteration is optimized based on the result of the previous step.

[0006] Preferably, the style transfer is to transfer the user-specified artistic style (such as oil painting, watercolor, etc.) to the optimized sketch while ensuring the consistency of local details and the natural fusion of styles. First, the system adopts a neural network-based style transfer algorithm, uses the optimized sketch as the content image, and the user-specified style image as the style reference, and extracts multi-level features of the content and style images through the pre-trained convolutional neural network VGG-19.

[0007] Preferably, the content features are mainly extracted from the deep convolutional layers of the network to capture the global structure and main contours of the sketch; the style features are extracted from multiple convolutional layers of the network, and the Gram matrix of the feature maps is calculated to characterize the style texture and color distribution; to enhance the local consistency of style transfer, the system introduces an attention mechanism, which dynamically adjusts the transfer intensity of the style features by calculating the attention weights between the content image and the style image features.

[0008] Preferably, the goal of the high-quality rendering stage is to convert the sketch after style transfer into a high-resolution and detail-rich final design image through generative adversarial networks and super-resolution techniques. The system adopts an architecture based on conditional GAN, where the generator is responsible for mapping the low-resolution sketch into a high-resolution image, and the discriminator distinguishes the generated image from the real high-resolution art image through adversarial training, thereby driving the generator to optimize the details.

[0009] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0010] First, a multi-stage collaborative framework is proposed, which integrates the diffusion model, CNN-MLP optimization, attention style transfer, and GAN super-resolution rendering into a unified system for the first time to achieve end-to-end design from concept to finished product;

[0011] Second, a dynamic conditional control mechanism is developed to endow the diffusion model with fine-grained generation direction regulation ability through the real-time embedding of user keywords and style tags;

[0012] Third, a local consistency optimization strategy is designed to use the attention map to constrain the local feature alignment of style transfer and solve the style-structure conflict problem in traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] For ease of description, the present invention is described in detail by the following specific embodiments and the accompanying drawings.

[0014] Figure 1 It is a schematic diagram of the GAN architecture;

[0015] Figure 2 It is a schematic diagram of the feature space entropy value;

[0016] Figure 3 It is a schematic diagram of innovation;

[0017] Figure 4 It is a schematic diagram of the generation time. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be described below through specific embodiments shown in the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. The structures, ratios, sizes, etc. depicted in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those skilled in this technology to understand and read, and are not used to limit the conditions under which the present invention can be implemented. Therefore, they do not have any technical substance. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the efficacy that the present invention can produce and the objectives that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention. In addition, in the following description, the description of well-known structures and technologies is omitted to avoid unnecessarily confusing the concepts of the present invention.

[0019] Here, it should also be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution according to the present invention are shown in the drawings, while other details less relevant to the present invention are omitted.

[0020] The following technical solutions are adopted in this specific embodiment:

[0021] I. Sketch generation:

[0022] Sketch generation is the core step of the intelligent creative design system, and its goal is to generate diverse initial sketches that meet user requirements through a diffusion model. The specific process is as follows: First, the system initializes a Gaussian noise image as the starting point and gradually denoises it through multiple steps of iteration of the diffusion model to generate a sketch. The diffusion model transforms random noise into a structured image through a reverse diffusion process, and each step of iteration is optimized based on the result of the previous step. To control the generation direction and meet user requirements, the system introduces a conditional diffusion model and uses the keywords entered by the user ("modern architecture" or "natural scenery") or style tags ("abstract" or "realistic") as conditional inputs. These conditional information are transformed into high-dimensional vectors through an embedding layer and fused with the features of the noise image to ensure that the generated sketch is consistent with user requirements in terms of content and style. During the iteration process, the system adopts an adaptive step size strategy, dynamically adjusting the denoising step size according to the complexity of the current image to balance the generation efficiency and image quality. In addition, to enhance the diversity of the sketches, the system introduces a random seed during each generation to ensure that different sketch variants can be generated even with the same user input. The finally generated sketches are preliminarily screened to remove the results that obviously do not meet user requirements or have low quality, and high-quality and diverse sketches are retained for subsequent optimization and rendering. Table 1 shows the performance comparison under different parameter settings during the sketch generation process, including generation time, diversity index (entropy value based on the feature space distribution), innovation index (score based on visual uniqueness), and user matching degree (the degree of compliance between the generated sketch and user requirements, with a score range of 0-10):

[0023] Table 1: Performance under Different Parameters

[0024]

[0025]

[0026] II. Sketch Optimization:

[0027] The goal of the sketch optimization stage is to improve the line quality and structural rationality of the initial sketch through the synergy of CNN and MLP. The specific process is as follows: First, the initial sketch generated by the diffusion model is input into a pre-trained CNN for feature extraction. This CNN uses a modified version of the multi-scale convolutional layer VGG-16 to capture the global contour, local details, and spatial relationships of the sketch, generating a high-dimensional feature map. Each channel of the feature map corresponds to visual information at different levels of abstraction (edges, textures, shapes, etc.). Through the feature fusion module, the multi-scale features are weighted and integrated to form a comprehensive feature vector. Subsequently, this feature vector is input into the MLP network, which consists of three fully connected layers. The activation function uses LeakyReLU to alleviate the gradient vanishing problem. Its task is to optimize the line smoothness, structural coherence, and proportional coordination of the sketch through non-linear mapping. Specifically, the first layer of the MLP reduces the dimensionality of the feature vector and extracts key structural parameters; the middle layer analyzes the topological relationships between lines through the self-attention mechanism, identifies and repairs broken or redundant line segments; the last layer outputs the optimized sketch parameter matrix, which is reconstructed into the optimized sketch image through deconvolution operations. To constrain the optimization process, the system defines a joint loss function, including line smoothness loss (measuring the consistency of adjacent pixel gradients), and its formula is:

[0028]

[0029] where S i,j represents the grayscale value of the sketch at pixel position (i, j), are the gradient operators in the horizontal and vertical directions respectively. The structure alignment loss (similarity measurement based on a predefined template), and its formula is:

[0030]

[0031] where is the optimized sketch,[[]]ID=19]] is the real hand-drawn sketch, φCNN represents the feature vector extracted by the CNN, and N is the number of samples; and the adversarial loss (distinguishing the optimized sketch from the real hand-drawn sketch through a lightweight discriminator). Finally, the weights of the CNN and MLP are iteratively updated through backpropagation. During the training process, the Adam optimizer is used to dynamically adjust the learning rate, and transfer learning is performed on the public sketch dataset (QuickDraw) to enhance the generalization ability.

[0032] III. Style Transfer:

[0033] ​The goal of the style transfer stage is to transfer the user-specified artistic style (such as oil painting, watercolor, etc.) to the optimized sketch while ensuring the consistency of local details and the natural integration of the style. First, the system adopts a neural network-based style transfer algorithm, using the optimized sketch as the content image and the user-specified style image as the style reference. It extracts multi-level features of the content and style images through the pre-trained convolutional neural network VGG-19. The content features are mainly extracted from the deep convolutional layers of the network to capture the global structure and main contours of the sketch; the style features are extracted from multiple convolutional layers of the network, and the style texture and color distribution are characterized by calculating the Gram matrix of the feature maps (representing the correlation between feature channels). To enhance the local consistency of style transfer, the system introduces an attention mechanism. By calculating the attention weights between the content image and the style image features, it dynamically adjusts the transfer intensity of the style features.

[0034] In the self-attention module adopted by the attention mechanism, the similarity scores between each position of the content feature map and all positions of the style feature map are calculated to generate an attention map, which is used to guide the weighted fusion of the style features. This process ensures that the style transfer can be consistent with the structure of the content image in the local area (specific lines or textures in the sketch), avoiding distortion or blurring caused by style transfer. Next, the system generates the target image through iterative optimization. The objective function consists of content loss, style loss, and attention consistency loss. The content loss is defined as the Euclidean distance between the generated image and the content image in the deep feature space to ensure that the generated image retains the main structure of the sketch; the style loss is defined as the difference between the generated image and the style image on the multi-layer Gram matrix to ensure that the generated image matches the target style; the attention consistency loss ensures the accuracy of local style transfer by comparing the attention maps of the generated image and the content image. During the training process, the L-BFGS optimization algorithm is used to efficiently solve the objective function, and the image details are gradually optimized through a multi-scale generation strategy.

[0035] IV. High-quality rendering:

[0036] The goal of the high-quality rendering stage is to convert the style-transferred sketch into a high-resolution and detail-rich final design image through generative adversarial networks (GANs) and super-resolution techniques. The system adopts an architecture based on conditional GANs, where the Generator is responsible for mapping the low-resolution sketch into a high-resolution image, and the Discriminator distinguishes the generated image from the real high-resolution art image through adversarial training, thereby driving the Generator to optimize the details. The core structure of the Generator contains an encoder-decoder framework: the encoder consists of multiple convolutional layers and downsampling layers to extract the deep features of the input sketch; residual blocks are introduced in the middle part to enhance the feature expression ability and alleviate the vanishing gradient problem; the decoder gradually upsamples through transposed convolutional layers and combines skip connections to fuse the low-level features (such as edges and textures) of the encoder with the high-level features of the decoder to ensure detail restoration. The Discriminator adopts a PatchGAN structure to distinguish the authenticity of local regions of the image through local receptive fields and outputs a spatial resolution matrix to guide the Generator to optimize local details. During the adversarial training process, the goal of the Generator is to minimize the adversarial loss and the content loss. To further improve the image resolution, the system integrates super-resolution techniques, adds a subpixel convolution layer at the end of the Generator to rearrange the low-resolution feature map into high-resolution pixels, and combines perceptual loss to optimize the visual quality. During training, a progressive strategy is adopted. First, the basic model is trained with low-resolution images, the resolution is gradually increased, and the network parameters are fine-tuned. Finally, a detail-rich image with a 4-fold resolution increase (from 256×256 to 1024×1024) is generated. As Figure 1 shown.

[0037] Example:

[0038] I. Experimental Setup:

[0039] The experiment aims to verify the performance of the intelligent creative design system based on the diffusion model. The experimental design focuses on design generation efficiency, diversity, innovation, and rendering quality. The experiment uses the publicly available datasets QuickDraw (containing 500,000 hand-drawn sketches) and WikiArt (covering 10 art styles such as oil paintings and watercolors) as data sources, with 80% used for training and 20% for testing. The comparison methods include traditional GAN generation models, neural style transfer algorithms (Gatys et al.), and Transformer-based design generation systems. The system parameter configurations are as follows: the diffusion model iterates for 1000 steps, and the noise schedule adopts a cosine decay strategy; the CNN feature extraction network is improved based on VGG-16, retaining the first 10 convolutional layers and adding adaptive pooling; the number of attention mechanism heads in the style transfer module is 8, and the cosine similarity is used for calculating attention weights; the GAN renderer sets the number of generator residual blocks to 9, the discriminator uses the PatchGAN structure, the adversarial training learning rate is 0.0002, and the upsampling factor of the super-resolution module is 4. The evaluation metrics include diversity (entropy value based on the feature space distribution), innovation (visual uniqueness score, averaged from the independent scores of 10 professional designers), generation time (total time from generation to rendering of a single sketch), and rendering quality (peak signal-to-noise ratio PSNR and structural similarity SSIM). The training environment uses 4 NVIDIA A100 GPUs for parallel acceleration, is implemented using the PyTorch framework, and all experiments are repeated 5 times and averaged to reduce the impact of randomness. The ablation experiment further validates the contributions of the attention mechanism, super-resolution module, and conditional diffusion model to the system performance.

[0040] II. Experimental Results:

[0041] In the baseline model selection and diversity metric evaluation phase, GLIDE and DALL-E2 are first selected as the comparison baselines, and parallel generation tests are conducted based on a unified test dataset (containing 2,000 sets of user input instructions and style constraints). The calculation of the feature space entropy value extracts the 2,048-dimensional feature vectors of the generated samples through the pre-trained Inception-v3 network. After reducing the dimension to 64-dimensional latent space by t-SNE, the kernel density estimation method is used to construct the probability density function of the feature distribution, and the distribution dispersion of the generated samples in batch mode is calculated. The grid division accuracy is set to 0.01 standard deviation units to ensure the stability of density estimation. In the comparison experiments, each model loads the official pre-trained weights and fixes the random seed. Under the condition of completely aligned input (including the same text prompt encoder and style label embedding layer), the comparability of the generation process is ensured by controlling the generation temperature parameter (temperature = 0.7) and the sampling steps (DDIM 50 steps). Finally, each model generates 500 sets of output samples for batch feature analysis, and the feature extraction process enables bilinear interpolation to uniformly scale to the input specification of 299×299 resolution. The calculation results of the feature space entropy value are as Figure 2 shown.

[0042] Experimental data shows that the feature space entropy value of this system shows a significant upward trend (5.84 → 6.15) with the increase of sampling steps, reaches a stable state after 34 steps, and its final entropy value is increased by 12.4% and 15.4% compared with GLIDE (5.47) and DALL-E 2 (5.33) respectively. This phenomenon stems from the multi-stage generation framework adopted by this system: in the early iteration stage (2 - 16 steps), through the progressive injection mechanism of Gaussian noise in the conditional diffusion model, the exploration scope of the solution space is expanded while preserving the user's intention; in the middle stage (16 - 30 steps), with the help of the attention-guided style transfer module, the diversity expression of style elements is maintained during the feature decoupling process; in the later stage (after 30 steps), through the adversarial training strategy of the GAN rendering engine, while improving the generation quality, the detail innovation is maintained. Compared with the baseline model, the single-stage generation mode of GLIDE leads to the stagnation of the entropy value growth after 30 steps (5.35 → 5.47), exposing the inherent problem of the diversity decay of traditional diffusion models in long-sequence generation; while DALL-E 2 is limited by the sequence generation characteristics of the autoregressive architecture, and its entropy value curve shows a significant low-slope growth (5.08 → 5.33), and the discrete token prediction mechanism leads to low efficiency in exploring the latent space. Further analysis shows that the dynamic noise scheduling algorithm of this system realizes the precise matching of the noise intensity and the generation stage. At the key turning point (step 20), through the intervention of the super-resolution module, the local diversity of the feature space is transformed into global innovation, while the baseline model lacks such cross-module cooperation mechanism, resulting in the generated samples falling into local optimality at the texture detail (GLIDE) or composition logic (DALL-E 2) level. Technical dissection shows that the phenomenon of the entropy value saturation of this system at step 40 coincides exactly with the regularization intensity threshold of the feature space of the MLP optimizer, verifying the effectiveness of the quality-diversity balance mechanism in the system design.

[0043] Innovative results are as Figure 3 shown.

[0044] In the visual distinctiveness scoring, this system significantly leads GLIDE (7.19 points) and DALL-E 2 (8.48 points) with an average score of 9.02 points. This is due to the technical characteristics of the system's integration of conditional diffusion models and attention mechanisms: In Designer 1, the system's surreal compositions generated through dynamic noise injection (appearance probability <5%) received extremely high evaluations. However, DALL-E 2 has obvious shortcomings in terms of texture innovation and element association logic in its generated samples due to being restricted by the discrete coding space of VQ-VAE. It is worth noting that GLIDE performs moderately well in structured innovation (scored 8.1 by Designer 8), but is severely lacking in style breakthrough, revealing the deficiencies of traditional diffusion models in cross-style fusion capabilities. Comparing the scoring distribution curves, this system exhibits a right-skewed distribution (skewness -1.2) feature, verifying its ability to generate breakthrough innovation solutions; while the normal distribution of DALL-E 2 (skewness 0.3) reflects that its innovation is strongly restricted by the distribution of training data.

[0045] Finally, the generation time was tested, and the results are as Figure 4 shown.

[0046] A systematic evaluation of the generation efficiency was conducted, and the results show that this system demonstrates a significant time advantage in the sketch generation task. Specifically, the generation time of this system is controlled within the range of 7.3 - 9.7 seconds. In contrast, the generation times of GLIDE and DALL-E 2 reach 9.9 - 12.6 seconds and 9.9 - 12.9 seconds respectively. This efficiency difference is mainly due to three aspects of architecture optimization: First, this system adopts a hierarchical feature extraction mechanism, shortening the sketch parsing time by 32% through a pre-computation module; Second, the parallel computing architecture based on dynamic pruning effectively reduces the redundant computing amount by 83%; Finally, the adaptive resource allocation algorithm improves the GPU video memory utilization rate to 91%, which has a significant advantage compared to the average of 75% of the comparison models.

[0047] III. Rendering Quality:

[0048] The rendering quality was tested through structural similarity SSIM. Table 2 shows the test results:

[0049] Table 2: Rendering Quality

[0050]

[0051]

[0052] According to the Structural Similarity (SSIM) test results in Table 2, this system performs superiorly in terms of rendering quality. An SSIM value close to 1 indicates higher image quality and better structural similarity. In the test, the SSIM value of this system ranges from 0.91 to 0.96, showing stable overall performance, indicating that the generated images maintain good structural similarity with real images. In contrast, the SSIM values of GLIDE and DALL-E 2 are both between 8.3 and 9.3, significantly lower than this system. This shows that even under similar generation conditions, the rendering quality of GLIDE and DALL-E 2 is relatively low, possibly limited by their generation algorithms or model architectures. In addition, the SSIM value of GLIDE remains almost unchanged, while the SSIM value of DALL-E 2 fluctuates slightly, which further emphasizes the advantage of this system in maintaining high rendering quality. In summary, this system performs outstandingly in terms of image rendering quality and is suitable for application scenarios with high requirements for image quality.

[0053] In summary, the intelligent creative design system based on the diffusion model demonstrates significant advantages in enhancing the diversity, efficiency, and quality of image generation. Through experiments, it is observed that the system shows higher diversity in terms of the feature space entropy value, indicating that it can explore the generation space more comprehensively, thus creating more innovative and unique design works. In the test of sketch generation time, the response speed of the system is significantly better than that of GLIDE and DALL-E 2, showing its efficiency in processing user input. Through the evaluation of Structural Similarity (SSIM), this system also achieves excellent results in terms of rendering quality, with the SSIM value significantly higher than that of other comparison models, reflecting the high consistency between the generated images and real images in terms of structure. These results indicate that the design system based on the diffusion model can not only generate high-quality images quickly but also provide rich creative options to meet the needs of users for personalized and diverse designs. This system is expected to play an important role in fields such as art creation, product design, and advertising marketing, providing powerful tools for designers and creators and promoting the development of the creative industry.

[0054] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to embrace all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention.

[0055] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only includes an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A three-dimensional model reconstruction method based on neural radiance fields, characterized in that: Its reconstruction method is as follows: By combining the conditional diffusion model with the attention-enhanced style transfer technology, it realizes a fully controllable design from sketch generation to high-quality rendering. The diffusion model generates diverse sketches through noise iteration and conditional embedding, breaking through the mode collapse limitation of traditional models. The CNN-MLP optimization module improves the rationality of lines and structures through a joint loss function. The style transfer technology driven by the attention mechanism precisely balances the consistency between the global style and local details.

2. The three-dimensional model reconstruction method based on neural radiance fields according to claim 1, characterized in that: The sketch generation generates diverse initial sketches that meet user requirements through the diffusion model. The specific process is as follows: The system initializes a Gaussian noise image as the starting point and gradually denoises it through multiple steps of iteration of the diffusion model to generate a sketch. The diffusion model converts random noise into a structured image through the reverse diffusion process, and each step of iteration is optimized based on the result of the previous step.

3. A three-dimensional model reconstruction method based on a neural radiance field according to claim 1, characterized in that: The style transfer is to transfer the user-specified artistic style to the optimized sketch while ensuring the consistency of local details and the natural integration of styles. First, the system adopts a neural network-based style transfer algorithm, using the optimized sketch as the content image and the user-specified style image as the style reference. It extracts multi-level features of the content and style images through the pre-trained convolutional neural network VGG-19. The content features are extracted from the deep convolutional layers of the network to capture the global structure and main contours of the sketch. The style features are extracted from multiple convolutional layers of the network, and the Gram matrix of the feature maps is calculated to characterize the style texture and color distribution. To enhance the local consistency of the style transfer, the system introduces an attention mechanism, and dynamically adjusts the transfer intensity of the style features by calculating the attention weights between the content image and the style image features.

4. A three-dimensional model reconstruction method based on neural radiance fields according to claim 1, characterized in that: The goal of the high-quality rendering stage is to convert the style-transferred sketch into a high-resolution and detail-rich final design image through generative adversarial networks and super-resolution techniques. The system adopts a conditional GAN-based architecture, where the generator is responsible for mapping the low-resolution sketch to a high-resolution image, and the discriminator distinguishes the generated image from the real high-resolution artistic image through adversarial training, thereby driving the generator to optimize the details.

Citation Information

Cited By

  • VR real-time rendering method and system for low-bit-rate cloud rendering plug flow

    CN121074333A