An image style transfer method and system based on a national traditional building data set
By constructing a dataset of traditional ethnic architecture to train a LoRA model and combining depth map and edge map constraints, the problems of low architectural design efficiency and lack of regional cultural characteristics in existing technologies are solved. This achieves efficient and accurate transfer of traditional ethnic architectural styles and generates high-quality images consistent with the original design.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HONG KONG UNIV OF SCI & TECH (GUANGZHOU)
- Filing Date
- 2026-04-03
- Publication Date
- 2026-07-10
AI Technical Summary
Existing technologies in architectural design suffer from low efficiency, structural distortion, and lack of regional cultural characteristics. In particular, style transfer algorithms based on GANs lead to distortion of architectural geometry, and general large models lack fine-grained knowledge of specific ethnic architecture, resulting in images that lack regional cultural characteristics.
A dataset of traditional ethnic architecture is constructed. The LoRA model is trained to obtain global style guidance capabilities. The external structure of the buildings is constrained by depth maps and edge maps. Features are extracted using ControlNet and IPAdapter models to generate images of traditional ethnic architectural styles.
It accurately learns and restores the fine-grained visual features of specific ethnic architecture, solves the problems of distorted architectural lines and perspective errors, ensures the consistency between the generated image and the original design, and improves the regional cultural recognition and generation efficiency.
Smart Images

Figure CN122368249A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image transfer technology, and in particular to an image style transfer method and system based on a dataset of traditional ethnic architecture. Background Technology
[0002] Currently, in the fields of architectural design, urban and rural planning, and cultural heritage protection, design schemes are typically presented in the form of preliminary models (white models). To showcase the design effect, the traditional process requires designers to use rendering software (such as Lumion, 3ds Max) for tedious material mapping, lighting settings, and rendering, which is time-consuming and labor-intensive. With the development of AIGC technology, using generative AI for style transfer is becoming a trend.
[0003] There are three existing technical solutions. The first is traditional manual rendering, which relies on manual adjustment of material parameters and lighting. This method has a long production cycle (usually several days) and requires extremely high professional skills from designers, making it difficult to meet the rapid iteration needs of large-scale design projects. The second is style transfer based on GAN (Generative Adversarial Network): While early algorithms (such as CycleGAN) can transfer color and texture, they often lead to severe distortion of architectural geometry (such as curved lines and perspective errors), failing to meet the stringent requirements of architectural design for structural accuracy. The third is general Stable Diffusion generation: Although the generated image quality is high, the general large model lacks fine-grained knowledge of specific ethnic architecture (such as the appearance of specific mortise and tenon structures and decorative patterns). The generated images are often "seemingly similar" but actually generic ancient architectural styles, lacking regional cultural characteristics. Moreover, without strong constraints, it is prone to "AI illusions," resulting in inconsistencies between the generated image and the original design. Summary of the Invention
[0004] To address the problems of low efficiency, structural distortion, and lack of regional cultural characteristics in existing technologies, this invention provides an image style transfer method and system based on a dataset of ethnic traditional architecture.
[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: An image style transfer method based on a dataset of ethnic traditional architecture includes the following steps: S1. Obtain a dataset of traditional ethnic buildings, which contains multiple images of traditional ethnic buildings; train a LoRA model on the diffusion model based on the dataset of traditional ethnic buildings to obtain a LoRA model with global style guidance capability; S2. Obtain the original design white film drawing to be migrated, which is used to preserve the overall structural outline of the building's exterior; based on the original design white film drawing, obtain a depth map and an edge map, whereby the depth map reflects the three-dimensional structure of the building's exterior and the edge map outlines the building's exterior contour lines; S3. Obtain a first control feature based on the depth map, and a second control feature based on the edge map; construct prompt words, and obtain text features based on the prompt words; acquire a reference image, and obtain reference style features based on the reference image; S4. Load the LoRA model into the diffusion model to obtain the loaded diffusion model; S5. Based on the original design white film diagram, the first control feature, the second control feature, the text feature, and the reference style feature, a latent feature map is generated through the loaded diffusion model; an image of ethnic traditional architectural style is generated based on the latent feature map.
[0006] A style transfer system based on a dataset of ethnic traditional architecture includes the following modules: The LoRA style model training module is used to acquire a dataset of ethnic traditional architecture, which contains multiple images of ethnic traditional architecture; the diffusion model is trained using LoRA based on the dataset of ethnic traditional architecture to obtain a LoRA model with global style guidance capabilities; The structural constraint preprocessing module is used to obtain the original design white film drawing to be transferred, which is used to preserve the overall structural outline of the building's exterior; and to obtain a depth map and an edge map based on the original design white film drawing, whereby the depth map reflects the three-dimensional structure of the building's exterior and the edge map outlines the building's exterior contour lines. A multimodal feature extraction module is used to obtain a first control feature based on the depth map and a second control feature based on the edge map; construct prompt words and obtain text features based on the prompt words; acquire a reference image and obtain reference style features based on the reference image; The model loading module is used to load the LoRA model into the diffusion model to obtain the loaded diffusion model; The fusion generation module is used to generate a latent feature map based on the original design white film image, the first control feature, the second control feature, the text feature, and the reference style feature through the loaded diffusion model; and to generate an image of ethnic traditional architectural style based on the latent feature map.
[0007] Compared with existing technologies, its advantages are as follows: By training the LoRA model using a dataset of traditional ethnic architecture, this invention can accurately learn and reproduce the fine-grained visual features of specific ethnic architecture, solving the problem of general models lacking regional cultural recognition. This invention obtains depth and edge maps based on the original design white film map, and uses the depth and edge maps to constrain the external three-dimensional structure of the building and outline the external contour lines of the building, effectively overcoming the defects of architectural line distortion and perspective errors in traditional style transfer, and ensuring the structural consistency between the generated map and the original white film map. Attached Figure Description
[0008] Figure 1 This is a flowchart illustrating the steps of a style transfer method based on a dataset of traditional ethnic architecture provided by this invention. Figure 2 This is a flowchart of the process for outputting images of traditional ethnic architectural styles provided in Embodiment 1 of the present invention; Figure 3 This is a flowchart of the process of outputting images of traditional Guangxi ethnic architectural styles in Embodiment 2 of the present invention. Detailed Implementation
[0009] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.
[0010] Example 1 like Figure 1 As shown, this embodiment provides an image style transfer method based on a dataset of traditional ethnic architecture, including the following steps: S1. Obtain a dataset of traditional ethnic buildings, which contains multiple images of traditional ethnic buildings; train the LoRA model on the diffusion model based on the traditional ethnic building dataset to obtain a LoRA model with global style guidance capabilities; S2. Obtain the original design white sheet drawing to be migrated. The original design white sheet drawing is used to preserve the overall structural outline of the building's exterior. Based on the original design white sheet drawing, obtain the depth map and edge map. The depth map is used to reflect the three-dimensional structure of the building's exterior, and the edge map is used to outline the building's exterior contour lines. S3. Obtain the first control feature based on the depth map and the second control feature based on the edge map; construct prompt words and obtain text features based on the prompt words; obtain a reference image and obtain reference style features based on the reference image; S4. Load the LoRA model into the diffusion model to obtain the loaded diffusion model; S5. Based on the original design white film map, first control feature, second control feature, text feature and reference style feature, a latent feature map is generated through the loaded diffusion model; an image of ethnic traditional architectural style is generated based on the latent feature map.
[0011] Step S1: Obtain the ethnic traditional architecture dataset, which contains multiple images of ethnic traditional architecture; train the diffusion model using LoRA based on the ethnic traditional architecture dataset to obtain a LoRA model with global style guidance capabilities.
[0012] In step S1, the ethnic traditional architecture dataset contains multiple images of ethnic traditional architecture. The specific steps for obtaining these images are as follows: S1.1. Collect data on traditional ethnic architecture in the field and integrate authoritative and publicly available resources to obtain an original dataset of traditional ethnic architecture; S1.2. Perform multi-level quality verification on the original ethnic traditional architecture dataset, retain samples with complete external structural information and style features, and clean and preprocess them to obtain the ethnic traditional architecture dataset to be processed. S1.3. The dataset of ethnic traditional buildings to be processed is initially labeled using a large language model to generate preliminary labeling results with descriptions of the external scene of the buildings and structured knowledge. Domain experts verify and correct the preliminary labeling results and supplement ethnic cultural attribute information. Finally, through double-blind cross-validation, the association between text description and image is established based on the core features of the building's exterior to obtain multiple images of ethnic traditional buildings.
[0013] Furthermore, the diffusion model is either a StableDiffusion model or a Flux model, the loss function used for LoRA training is the L2 loss function, and a labeler is used to generate descriptive text centered on the external features of the building as training labels.
[0014] By building its own dataset of traditional ethnic architecture and fine-tuning with LoRA, the model can accurately generate unique ethnic architectural elements (such as double-eaved pyramidal roofs, bronze drum patterns, and stilted wooden structures), solving the problem of general models being "unsuitable for local conditions". The generated images have a very high degree of regional cultural recognition.
[0015] The flowchart for outputting images of traditional ethnic architectural styles in this embodiment is as follows: Figure 2 As shown, Figure 2 This demonstrates the process of generating an image in the traditional ethnic architectural style from the original design white model. After inputting the original white model, the preprocessing module generates a depth map and an edge map. Figure 2 The preprocessing module in this example is the structural constraint preprocessing module of the style transfer system based on a dataset of ethnic traditional architecture. The depth map is used to capture the three-dimensional spatial information of the building. The core principle of the method in this embodiment when generating the depth map is to generate a structural constraint vector based on the depth value, forcing the generated map to match the three-dimensional layout of the original space. The edge map is used to extract the external contour line information of the building. The core principle of the method in this embodiment when generating the edge map is to generate a contour constraint vector based on the edge lines, limiting the attachment boundaries of ethnic elements. Acquire the dataset. Figure 2The dataset serves as the reference image in step 3. The IPAdapter model uses the reference dataset detail map to obtain reference style features, providing local detail enhancement and improving the accuracy of pattern texture and color reproduction. First control features are extracted through the depth map, which locks the spatial structure (ensuring functional areas do not shift or proportions are not distorted). Second control features are extracted through the edge map, which constrains the detail contours (preventing ethnic elements from covering key functional areas). An architectural LoRA model is loaded into the diffusion model, providing global style guidance and injecting ethnic style elements as a foundation. In the style-structure fusion module, the original white film image, the first control features, the second control features, and the reference style are processed through the loaded diffusion model, sampled by KSampler, and decoded by VAE to generate an initial stylized image. Figure 2 The initial stylized image in this embodiment is the ethnic traditional architectural style image output by the method of this embodiment. The method of this embodiment also includes ImageScale super-resolution processing, which enlarges the initial stylized image to a resolution of 2048×2048, and outputs an interior design drawing that combines ethnic style and practical function.
[0016] Step S2: Obtain the original design white film drawing to be migrated. The original design white film drawing is used to preserve the overall structural outline of the building's exterior. Based on the original design white film drawing, obtain the depth map and edge map. The depth map is used to reflect the three-dimensional structure of the building's exterior, and the edge map is used to outline the building's exterior contour lines.
[0017] By utilizing the dual ControlNet collaborative constraint mechanism of "Depth Map and Canny Map," this embodiment solves the pain points of distorted architectural lines and perspective errors in traditional style transfer. The depth map locks in the 3D spatial layout, while the canny map locks in the detailed outlines of doors and windows, ensuring that the generated image is completely consistent with the original white film image, thus meeting the implementation requirements of professional designs.
[0018] Step S3: Obtain the first control feature based on the depth map and the second control feature based on the edge map; construct prompt words and obtain text features based on the prompt words; acquire a reference image and obtain reference style features based on the reference image.
[0019] Furthermore, the acquisition of the first control feature, the second control feature, the text feature, and the reference style feature specifically involves: The ControlNet model is used to process the depth map and edge map to obtain the first control feature and the second control feature. External style instructions are generated using a cue word generation tool. These instructions include both positive and negative cue words. The external style instructions are then input into the text encoder of the diffusion model to obtain text features. A reference image is obtained and input into the IPAdapter model. The features of the reference image are extracted through the IPAdapter model to obtain the reference style features.
[0020] Step S4 loads the LoRA model into the diffusion model to obtain the loaded diffusion model.
[0021] Step S5 generates a latent feature map based on the original design white film map, the first control feature, the second control feature, the text feature, and the reference style feature through the loaded diffusion model; and generates an image of ethnic traditional architectural style based on the latent feature map.
[0022] In step S5, the original design white film image is encoded into the latent space by the encoder of the variational autoencoder to obtain the initial latent space features, which serve as the starting point for iterative denoising. The iterative denoising process is controlled by the KSampler sampler. Within a preset number of sampling steps, the following operations are performed in a loop: input the latent space features of the current step, the current time step, the first control feature, the second control feature, the text feature, and the reference style feature into the loaded diffusion model, and output the noise corresponding to the current step; constrain the first control feature and the second control feature with the first weight and the second weight respectively; the first weight is greater than the second weight.
[0023] The KSampler sampler updates the latent space features of the current step based on the preset sampling algorithm and the noise corresponding to the current step, and obtains the latent space features of the next step; until the iteration is completed, the latent feature map is obtained. The latent feature map is decoded by a variational autoencoder to generate an image of traditional ethnic architectural style.
[0024] Example 2 Based on the image style transfer method based on a dataset of traditional ethnic architecture described in Example 1, this example provides an image style transfer method based on a dataset of traditional ethnic architecture in Guangxi, including the following steps: S1. Obtain a dataset of traditional ethnic buildings, which contains multiple images of traditional ethnic buildings; train the LoRA model on the diffusion model based on the traditional ethnic building dataset to obtain a LoRA model with global style guidance capabilities; S2. Obtain the original design white sheet drawing to be migrated. The original design white sheet drawing is used to preserve the overall structural outline of the building's exterior. Based on the original design white sheet drawing, obtain the depth map and edge map. The depth map is used to reflect the three-dimensional structure of the building's exterior, and the edge map is used to outline the building's exterior contour lines. S3. Obtain the first control feature based on the depth map and the second control feature based on the edge map; construct prompt words and obtain text features based on the prompt words; obtain a reference image and obtain reference style features based on the reference image; S4. Load the LoRA model into the diffusion model to obtain the loaded diffusion model; S5. Based on the original design white film map, first control feature, second control feature, text feature and reference style feature, a latent feature map is generated through the loaded diffusion model; an image of ethnic traditional architectural style is generated based on the latent feature map.
[0025] Step S1: Obtain the ethnic traditional architecture dataset, which contains multiple images of ethnic traditional architecture; train the diffusion model using LoRA based on the ethnic traditional architecture dataset to obtain a LoRA model with global style guidance capabilities.
[0026] In this embodiment, the Guangxi ethnic traditional architecture dataset contains multiple images of Guangxi ethnic traditional architecture. The specific steps for obtaining these images are as follows: S1.1. Collect data on traditional ethnic architecture in Guangxi on-site and integrate authoritative and publicly available resources to obtain an original dataset of traditional ethnic architecture; In this embodiment, a professional team was organized to conduct systematic field collection in key areas of multiple cities in Guangxi, based on the sampling principles of "regional representativeness," "style typicality," and "external structural integrity." At the same time, authoritative and publicly available resources were integrated to ensure the comprehensiveness and reliability of the data on the external features of traditional ethnic buildings in Guangxi. Key areas for on-site data collection: In areas inhabited by ethnic minorities, priority will be given to architectural samples with intact external structures, clear mortise and tenon joints / bracket sets, and distinctive ethnic features (such as Dong drum towers and Zhuang stilted houses); In areas inhabited by immigrants from Guangdong and Guangxi in eastern Guangxi, architectural remains reflecting the integration of immigrant and local cultures will be collected (such as the facades of arcade buildings and the exterior wall styles of Hakka walled villages); In border areas, modern architectural samples that blend Chinese and Western architectural styles will be added (such as church-style facades and exotic roofs).
[0027] Public and authoritative resource integration: Data comes from official channels such as the Guangxi Housing and Urban-Rural Development Department, cultural protection units at all levels, and digital cultural heritage platforms. Strictly adhering to the principles of "clear copyright, qualified quality, and semantic complementarity," data includes high-definition panoramic images of building exteriors and multi-angle elevation views, supplementing external scenes not covered by on-site collection, and removing resources with unclear external structures, missing key features, or weak relevance.
[0028] S1.2. Perform multi-level quality verification on the original ethnic traditional architecture dataset, retain samples with complete external structural information and style features, and clean and preprocess them to obtain the ethnic traditional architecture dataset to be processed. In this embodiment, with "external structural integrity", "style recognition" and "model compatibility" as the core objectives, standardized tools and parameter optimization strategies are used for processing, and multi-level quality verification is implemented. The focus is on retaining samples that can fully reflect the external main form of the building, the overall structural layout, the external environmental background and cultural scene information.
[0029] After retaining samples with complete external structural information and stylistic characteristics, low-quality data with incomplete external details and distorted structural proportions are removed. Professional tools are used for precise segmentation to capture the dynamic interaction between the building exterior and light and shadow, focusing on retaining segments of the overall external structure of the building from different angles. External structural descriptions are added to each video segment, and specific information on building exterior forms from the cultural heritage knowledge base is linked. Annotated text and structured cultural knowledge are organized, and the terminology of building exterior structures is standardized (such as "raised beam roof structure" and "stilt house exterior wall layout"). A spatiotemporal sequence association with the video data is established, supporting cross-modal retrieval, resulting in a dataset of ethnic traditional buildings to be processed.
[0030] S1.3. The dataset of ethnic traditional buildings to be processed is initially labeled using a large language model to generate preliminary labeling results with descriptions of the external scene of the buildings and structured knowledge. Domain experts verify and correct the preliminary labeling results and supplement ethnic cultural attribute information. Finally, through double-blind cross-validation, the association between text description and image is established based on the core features of the building's exterior to obtain multiple images of ethnic traditional buildings.
[0031] In this embodiment, the video data is processed using large language models such as Qianwen Vision to initially annotate the dataset of ethnic traditional buildings to be processed, generating a summary of the building's external scene description and structured knowledge (including building external type, external style features, overall structural attributes, core external structures, etc.). Scholars in the field of Guangxi ethnic architecture are invited to verify and correct the initial annotations, supplementing information on the external cultural attributes of ethnic buildings, conducting an assessment of the cultural value of external structures, and standardizing regional and ethnic architectural terminology. Double-blind cross-validation: Based on the core features of the building's exterior (such as roof shape, exterior wall structure, and column layout), a precise association is established between the text and video clips, resulting in multiple images of Guangxi ethnic traditional buildings. This step ensures semantic consistency and annotation accuracy, adapting to cross-modal fine-tuning requirements.
[0032] Furthermore, the diffusion model can be a StableDiffusion model, an SDXL model, an SD1.5 model, or a Flux model. The loss function used for LoRA training is an L2 loss function. A labeler is used to generate descriptive text with the building's external features as the core as training labels.
[0033] In this embodiment, an NVIDIA RTX 4060 graphics card (16GB VRAM) is configured, and deep learning frameworks such as Python, PyTorch, and CUDA are installed. The SDXL model is used as the base model to ensure the quality of the generated building exterior and the accuracy of the overall structural reproduction.
[0034] Training parameters were set as follows: AdamW8bit optimizer, U-Net learning rate 1e-4, text encoder learning rate 1e-5, and cosine annealing learning rate scheduler and early stopping mechanism to prevent overfitting. Training strategy: batch size of 4, image resolution of 1024×512, L2 loss function, and gradient clipping and weight decay to optimize training stability. A tagger was used to generate descriptive text centered on the external features of buildings (e.g., "external structure of Dong drum tower with pointed roof" and "external wall of Zhuang stilted house"), and ethnic style keywords were manually added. The tag editor was used to adjust the image tag information to strengthen the training guidance of external structure and style features.
[0035] Furthermore, it is preferable to set Rank=128, Alpha=128, and enable Gradient Checkpointing to save video memory and ensure the best feature capture capability under limited hardware resources.
[0036] Step S2: Obtain the original design white film drawing to be migrated. The original design white film drawing is used to preserve the overall structural outline of the building's exterior. Based on the original design white film drawing, obtain the depth map and edge map. The depth map is used to reflect the three-dimensional structure of the building's exterior, and the edge map is used to outline the building's exterior contour lines.
[0037] In this embodiment, the style transfer process deeply integrates with the core components of the ComfyUI workflow, focusing on the external features of traditional ethnic architecture in Guangxi. Based on ComfyUI, an end-to-end automated workflow is encapsulated, supporting dynamic prompt generation and batch processing, achieving "one-click generation." Compared to traditional rendering, efficiency is improved several times (reduced from days to minutes), significantly lowering rendering costs for architectural design institutes and cultural tourism planning departments during the scheme presentation stage.
[0038] Furthermore, the first control tool in the two-stage ControlNet structure obtains the original design white film image to be transferred (focusing on preserving the overall external structural outline), and the second control tool generates the corresponding depth map (reflecting the external three-dimensional structure) and edge map (outlining the external contour lines), providing data support for subsequent external structure control.
[0039] Step S3: Obtain the first control feature based on the depth map and the second control feature based on the edge map; construct prompt words and obtain text features based on the prompt words; acquire a reference image and obtain reference style features based on the reference image.
[0040] The ControlNet model is used to process the depth map and edge map to obtain the first and second control features. The ControlNet preprocessor can be replaced with MLSD (Line Detection, suitable for modern geometric architecture) or Segment (Semantic Segmentation, suitable for complex scenes) to adapt to different types of architectural line features.
[0041] External style instructions, including positive and negative prompts, are generated using a prompt word generation tool. These instructions are then input into the text encoder of a diffusion model to obtain text features. A reference image is acquired and input into an IPAdapter model. The IPAdapter model extracts features from the reference image to obtain the reference style features.
[0042] In this embodiment, the prompt word construction stage integrates a prompt word generation tool to generate targeted external style instructions. Positive prompt words include "Zhuang ethnic bronze drum pattern exterior wall, Dong ethnic wind and rain bridge exterior wood carving window lattice, blue tile roof exterior shape, stilted exterior wall structure"; negative prompt words supplement "external structure deformation, lack of ethnic external elements, modern minimalist exterior style" to avoid style deviation and conflict with external structure. In the model preparation stage, the basic design pre-trained model and the dedicated LoRA model for Guangxi ethnic architecture dataset are loaded simultaneously, and the LoRA weight is set to 0.6-0.8 to ensure the stability of the basic external space style and give full play to the guiding role of ethnic external style characteristics; then, the style features of Guangxi traditional residential exterior detail images (such as the external shape of drum tower brackets, Zhuang brocade pattern exterior wall, and wind and rain bridge exterior wood carving) are injected through the IPAdapter (image prompt adapter) tool, and the IPAdapter index is configured to 0.5 to strengthen the accurate fit between external local details and ethnic elements.
[0043] Step S4 loads the LoRA model into the diffusion model to obtain the loaded diffusion model.
[0044] Step S5 generates a latent feature map based on the original design white film map, the first control feature, the second control feature, the text feature, and the reference style feature through the loaded diffusion model; and generates an image of ethnic traditional architectural style based on the latent feature map.
[0045] In step S5, the original design white film image is encoded into the latent space by the encoder of the variational autoencoder to obtain the initial latent space features, which serve as the starting point for iterative denoising. The iterative denoising process is controlled by the KSampler sampler. Within a preset number of sampling steps, the following operations are performed in a loop: input the latent space features of the current step, the current time step, the first control feature, the second control feature, the text feature, and the reference style feature into the loaded diffusion model, and output the noise corresponding to the current step; constrain the first control feature and the second control feature with the first weight and the second weight respectively; the first weight is greater than the second weight.
[0046] In this embodiment, the relevant models of the control depth map and edge map are loaded by the ControlNet loader respectively, and applied step by step in the ControlNet control node: In the first stage, the depth map control is enabled with a first weight of 0.8 to accurately lock the three-dimensional structure, overall proportion and spatial layout of the original building exterior, ensuring that the core external structures such as exterior walls, roof, and columns do not shift; In the second stage, the edge map control is enabled with a second weight of 0.65 to accurately constrain the line integrity and attachment position of the ethnic external patterns (such as avoiding patterns from covering the exterior door and window openings and destroying the overall outline).
[0047] The KSampler sampler updates the latent space features of the current step based on the preset sampling algorithm and the noise corresponding to the current step, and obtains the latent space features of the next step; until the iteration is completed, the latent feature map is obtained. The latent feature map is decoded by a variational autoencoder to generate the image of the ethnic traditional architectural style.
[0048] In this embodiment, the KSampler node is configured with 30 sampling steps, a sampler, and a CFGScale (without classifier guidance coefficient) of 7.5 to achieve a balance between the intensity of external style expression and the fidelity of the overall structure. After image decoding by the VAE (Variational Autoencoder) decoder, the generated image is further enlarged to a high resolution of 2048x2048 using the ImageScale tool to improve the presentation of external structural details and style elements. The image comparison component focuses on comparing the integrity of the overall external structure of the building before and after migration, the naturalness of the integration of external style elements, and the coordination between external structure and style elements. If there are problems such as external structural deformation, obscuring of the core external structure by ethnic elements, or a disconnect between style and external form, the LoRA weights, IPAdapter strength, or ControlNet parameters can be adjusted in reverse. Finally, the optimized style migration result is saved using the image saving component, forming a design scheme that combines ethnic external cultural characteristics with the practicality of the external structure.
[0049] The flowchart for outputting images of traditional Guangxi ethnic architectural styles in this embodiment is as follows: Figure 3 As shown. First, obtain the style image. Figure 3The style images in this embodiment represent the original ethnic traditional architecture dataset. Semantic annotation of these images is performed using the Qianwen model to obtain preliminary annotation results. Experts in Guangxi ethnic architecture are invited to correct and supplement these preliminary annotation results. Double-blind cross-validation is used to verify and quality control the corrected and supplemented results. This text annotation process yields the ethnic traditional architecture dataset. Next, prompt words are constructed from the style images, describing ideal features and style details. Then, reference images are obtained, which are the original design white-film images in this embodiment. Depth maps and line drawings are extracted from the reference images, with the line drawings representing the edge maps. Finally, model training and inference are performed. A LoRA model is trained based on the ethnic traditional architecture dataset. The trained LoRA model is then loaded into a diffusion model. This embodiment uses the SDXL model. Based on the depth map, line drawings, and prompt word control information, the loaded SDXL diffusion model generates a latent feature map. Based on the latent feature map, the Guangxi ethnic traditional architecture style image is output.
[0050] Example 3 This embodiment provides a style transfer system based on a dataset of ethnic traditional architecture, including the following modules: The LoRA style model training module is used to obtain a dataset of ethnic traditional architecture, which contains multiple images of ethnic traditional architecture. Based on the ethnic traditional architecture dataset, the diffusion model is trained using LoRA to obtain a LoRA model with global style guidance capabilities. The structural constraint preprocessing module is used to obtain the original design white model to be transferred. The original design white model is used to preserve the overall structural outline of the building's exterior. Based on the original design white model, a depth map and an edge map are obtained. The depth map is used to reflect the three-dimensional structure of the building's exterior, and the edge map is used to outline the building's exterior contour lines. The multimodal feature extraction module is used to obtain the first control feature based on the depth map and the second control feature based on the edge map; construct prompt words and obtain text features based on the prompt words; acquire reference images and obtain reference style features based on the reference images; The model loading module is used to load the LoRA model into the diffusion model to obtain the loaded diffusion model. The fusion generation module is used to generate a latent feature map based on the original design white film map, the first control feature, the second control feature, the text feature, and the reference style feature through a loaded diffusion model; and to generate an image of ethnic traditional architectural style based on the latent feature map.
[0051] In summary, this invention provides a style transfer method and system based on a dataset of ethnic traditional architecture. This invention constructs a dataset of ethnic traditional architecture containing natural language descriptions and feature labels, trains a LoRA model, and establishes the overall style tone of ethnic traditional architecture. Dual ControlNets respectively achieve macro-level control of the building's external spatial structure and micro-level constraints on external pattern details, avoiding the failure of external functional construction caused by style transfer. The trained LoRA model works in conjunction with the IPAdapter to accurately extract the detailed features of the building's exterior within the dataset, improving the accuracy of external style restoration. The prompt word generation mechanism uses a prompt word splicing component to achieve random combination of ethnic external elements, effectively avoiding the uniformity of transfer results and ensuring that each generation presents a rich and diverse expression of the external style of ethnic traditional architecture while maintaining a high degree of adaptability to the original external structure. This invention effectively overcomes the defects of traditional style transfer, such as distorted building lines and perspective errors, generating high-quality images of ethnic traditional architectural styles.
[0052] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and substitutions can be made without departing from the technical principles of the present invention, and these improvements and substitutions should also be considered within the scope of protection of the present invention.
Claims
1. A method for image style transfer based on a dataset of traditional ethnic architecture, characterized in that, Includes the following steps: S1. Obtain a dataset of traditional ethnic buildings, which contains multiple images of traditional ethnic buildings; train a LoRA model on the diffusion model based on the dataset of traditional ethnic buildings to obtain a LoRA model with global style guidance capability; S2. Obtain the original design white film drawing to be migrated, which is used to preserve the overall structural outline of the building's exterior; based on the original design white film drawing, obtain a depth map and an edge map, whereby the depth map reflects the three-dimensional structure of the building's exterior and the edge map outlines the building's exterior contour lines; S3. Obtain a first control feature based on the depth map, and a second control feature based on the edge map; construct prompt words, and obtain text features based on the prompt words; Obtain a reference image, and obtain reference style features based on the reference image; S4. Load the LoRA model into the diffusion model to obtain the loaded diffusion model; S5. Based on the original design white film diagram, the first control feature, the second control feature, the text feature, and the reference style feature, a latent feature map is generated through the loaded diffusion model; an image of ethnic traditional architectural style is generated based on the latent feature map.
2. The method according to claim 1, characterized in that, In step S1, the acquisition of the ethnic traditional architecture dataset, which contains multiple images of ethnic traditional buildings, includes: S1.
1. Collect data on traditional ethnic architecture in the field and integrate authoritative and publicly available resources to obtain an original dataset of traditional ethnic architecture; S1.
2. Perform multi-level quality verification on the original ethnic traditional architecture dataset, retain samples with complete external structural information and style characteristics, and perform cleaning and preprocessing to obtain the ethnic traditional architecture dataset to be processed. S1.
3. The dataset of ethnic traditional buildings to be processed is initially labeled using a large language model to generate preliminary labeling results with descriptions of the external scene of the buildings and structured knowledge. Domain experts verify and correct the preliminary labeling results and supplement ethnic cultural attribute information. Finally, through double-blind cross-validation, the association between text description and image is established based on the core features of the external architecture to obtain the multiple images of ethnic traditional buildings.
3. The method according to claim 1, characterized in that, In step S1, the diffusion model is either the StableDiffusion model or the Flux model.
4. The method according to claim 1, characterized in that, In step S1, the diffusion model is trained using LoRA based on the ethnic traditional architecture dataset to obtain a LoRA model with global style guidance capabilities. This includes: the loss function used in the LoRA training is the L2 loss function, and a labeler is used to generate descriptive text with the external features of the buildings as the core as training labels.
5. The method according to claim 1, characterized in that, In step S3, obtaining the first control feature based on the depth map and the second control feature based on the edge map includes: processing the depth map and the edge map using the ControlNet model to obtain the first control feature and the second control feature.
6. The method according to claim 1, characterized in that, In step S3, the construction of cue words and the acquisition of text features based on the cue words include: generating external style instructions through a cue word generation tool, the external style instructions including positive cue words and negative cue words; inputting the external style instructions into the text encoder of the diffusion model to obtain the text features.
7. The method according to claim 1, characterized in that, In step S3, obtaining a reference image and obtaining reference style features based on the reference image includes: obtaining a reference image, inputting the reference image into an IPAdapter model, extracting features from the reference image through the IPAdapter model, and obtaining the reference style features.
8. The method according to claim 1, characterized in that, In step S5, the original design white film image is encoded into the latent space by the encoder of the variational autoencoder to obtain the initial latent space features, which serve as the starting point for iterative denoising. The iterative denoising process is controlled by a KSampler sampler, which performs the following operations in a loop within a preset number of sampling steps: inputting the latent space features of the current step, the current time step, the first control feature, the second control feature, the text feature, and the reference style feature into the loaded diffusion model; the loaded diffusion model outputs the noise corresponding to the current step; the KSampler sampler updates the latent space features of the current step according to the preset sampling algorithm and the noise corresponding to the current step, and obtains the latent space features of the next step. The latent feature map is obtained after all iterations are completed. The latent feature map is decoded by a variational autoencoder to generate the image of the ethnic traditional architectural style.
9. The method according to claim 8, characterized in that, In step S5, the first control feature and the second control feature are constrained with a first weight and a second weight, respectively; The first weight is greater than the second weight.
10. A style transfer system based on a dataset of traditional ethnic architecture, characterized in that, Includes the following modules: The LoRA style model training module is used to acquire a dataset of ethnic traditional architecture, which contains multiple images of ethnic traditional architecture; the diffusion model is trained using LoRA based on the dataset of ethnic traditional architecture to obtain a LoRA model with global style guidance capabilities; The structural constraint preprocessing module is used to obtain the original design white film drawing to be transferred, which is used to preserve the overall structural outline of the building's exterior; and to obtain a depth map and an edge map based on the original design white film drawing, whereby the depth map reflects the three-dimensional structure of the building's exterior and the edge map outlines the building's exterior contour lines. A multimodal feature extraction module is used to obtain a first control feature based on the depth map and a second control feature based on the edge map; construct prompt words and obtain text features based on the prompt words; Obtain a reference image, and obtain reference style features based on the reference image; The model loading module is used to load the LoRA model into the diffusion model to obtain the loaded diffusion model; The fusion generation module is used to generate a latent feature map based on the original design white film image, the first control feature, the second control feature, the text feature, and the reference style feature through the loaded diffusion model; and to generate an image of ethnic traditional architectural style based on the latent feature map.