Artificial intelligence image generation system and method suitable for historical block facade updating
By integrating the ComfyUI workflow of the Flux model loader, image input module, DeepSeek language interaction module, Controlnet processing module, LoRA module and image generation module, the problem of insufficient adaptability to local culture in existing technologies is solved, and convenient localized street renovation image generation is achieved, thereby improving design efficiency and interactivity.
Patent Information
- Application Number
- CN202510818593.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-30
AI Technical Summary
Existing image generation technology cannot effectively reflect the unique features and characteristics of local culture. Traditional methods have limited support for ordinary citizens and non-professional designers, lack flexibility and convenience, and it is difficult to generate professional design solutions that conform to actual scenarios.
The ComfyUI workflow, consisting of the Flux model loader, image input module, DeepSeek language interaction module, Controlnet processing module, LoRA module, sampler module, and image generation module, is used in combination with a mobile app to implement an image generation system that generates street renovation plans that conform to local characteristics through natural language input.
It improves design efficiency and interactivity, can easily generate street renovation plans that conform to local characteristics, supports mobile operation, ensures that the generated images are highly consistent with the actual scene, and is suitable for ordinary citizens and non-professional designers.
Smart Images

Figure CN120726162A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image generation technology, and in particular to an artificial intelligence image generation system and method suitable for updating facades of historical blocks. Background Art
[0002] With the advancement of urbanization, the renovation of historical streets faces the challenges of diverse needs and complex designs. While existing image generation technologies can generate images of street renovations, most models lack adaptability to local culture and characteristics, and the generated images often fail to capture the unique character of a specific area. Furthermore, traditional large models often require complex operations and high-end hardware support, and can only be used via a PC, lacking flexibility and convenience. In particular, traditional methods lack effective constraints on spatial structures such as building outlines and street layouts, resulting in an inadequate match between the generated images and the actual scene. In particular, during the image generation process, traditional methods have limited support for ordinary citizens and non-professional designers, making it difficult to generate professional design solutions through plain language. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide an artificial intelligence image generation system and method suitable for the renewal of facades in historical blocks, which can easily and quickly obtain street renovation plans that conform to local characteristics, greatly improving the efficiency and interactivity of the design.
[0004] In order to solve the above technical problems, the present invention provides an artificial intelligence image generation system suitable for the renovation of historical block facades, comprising:
[0005] The Flux model loader module is used to load basic models and components, provide basic parameters for image generation, environment configuration, and model loading;
[0006] Image input module, used to collect existing street images;
[0007] The DeepSeek language interaction module is used to receive personalized street renovation requirements input by users through natural language, parse the content and generate structured control instructions;
[0008] Controlnet processing module, used to generate structural constraints;
[0009] a sampler module for collecting image input data, structured control instructions, and structural constraints of street images and generating graph generation data;
[0010] LoRA module, used to perform localized style adjustment on image generation data;
[0011] The graph generation module is used to decode the graph generation data to generate a street transformation image and output the street transformation result;
[0012] The mobile app module is used for user input, image acquisition and upload, and displays the final street transformation results;
[0013] The Flux model loader module, image input module, DeepSeek language interaction module, LoRA module, sampler module, Controlnet processing module, image generation module and mobile mini-program module are combined to form the ComfyUI workflow, and the street transformation image generation is completed through the ComfyUI workflow.
[0014] Furthermore, the Flux model loader module includes a UNET loader, which is used to load the flux large model. The flux large model is specifically the flux1-schnel lfp8.safetensors large model. The flux large model is used to provide the environment and parameters for the image generation process and dynamically update the configuration in the image processing process. A FluxForwardOverrider node is added to the flux large model. The FluxForwardOverrider node is used to merge functions or modify existing behaviors to ensure that the original model remains unchanged while supporting advanced modifications.
[0015] Furthermore, the image input module is connected to street cameras, drones and / or mobile devices to collect and transmit street images in real time.
[0016] Furthermore, the LoRA module is obtained by training the manually annotated LoRA training module, which first extracts and calculates the input data of the collected street image samples, and forms a historical block dataset based on manual annotation for training;
[0017] The manually labeled LoRA training module incorporates historical street architectural landscape elements for training, classifies and professionally identifies the characteristics of historical street landscape elements, and makes the generated images have highly localized features.
[0018] Furthermore, the sampler module dynamically adjusts the sampling steps and sampling rate according to different detail areas of the street image to ensure that the image details are clear and distortion-free during the generation process; the sampler module uses the model sampling algorithm flux node to perform model sampling.
[0019] Furthermore, the Controlnet module performs multi-level structural control on the image generation process by loading the Controlnet constraint model, including building edge reinforcement, street element positioning and scale adjustment.
[0020] Furthermore, the mobile mini-program module includes an image input module, a text input module, and an image preview module, and users can provide feedback and make modifications based on the generated image.
[0021] An artificial intelligence image generation method suitable for updating facades in historic blocks, using any of the above-described systems, comprises the following steps:
[0022] Step S1: Train the Lora model to obtain the core model;
[0023] Step S2: Use a camera or sensor to capture images of existing streets or upload street view images that need to be modified and input them into the ComfyUI workflow;
[0024] Step S3: parse the user input requirements through natural language processing technology and convert them into structured control instructions;
[0025] Step S4: During the image generation process, the sampler module samples the street view image and generates graph generation data based on structured control instructions and ControlNet structural constraints. The graph generation data is also adjusted for local style by the core model.
[0026] Step S5: the image generation module decodes the adjusted image generation data to generate a street reconstruction image and outputs the street reconstruction result;
[0027] Step S6: The user views the generated street reconstruction result map through the mobile mini program. If it meets the requirements, the generation is completed. If it does not meet the requirements, jump to step S3 to continue execution.
[0028] Furthermore, the training process of the Lora model is as follows: first install the Lora model trainer Fluxgym and the operating environment, collect street view images obtained from field surveys, crop the street view images in batches to 512*512 pixels, identify the building roof shape, analyze the wall color spectrum, extract the decorative component pattern, and calculate the street space ratio of the collected street view images, use Fluxgym's annotation tool to annotate the data, and manually annotate the historical block elements in the street view images. Use N manually annotated street view images as training samples, adjust the training parameters and start the training process; after the training is completed, evaluate and optimize based on the loss value of the Lora model, and select the best performing model as the core model for image generation.
[0029] Furthermore, the ComfyUI_Bxb component is installed in the ComfyUI node manager, and the image input, text input, and image generation components are encapsulated into a mobile mini-program, so that users can view and adjust the generated street reconstruction design through the mobile terminal.
[0030] Beneficial effects of the present invention:
[0031] By integrating the DeepSeek language model, the present invention can convert the natural language input of ordinary users into professional-level street design drawings, realizing more convenient and low-threshold image generation. At the same time, by introducing the ControlNet module, edge detection and semantic segmentation technology are used to accurately control the building structure and street proportions to ensure the spatial rationality of the generated results. At the same time, the Lora model generates street renovation images that conform to local culture and architectural style through special training on local characteristics, solving the problem that existing large models cannot achieve localized design. In addition, the present invention also enables users to experience image generation on their mobile phones anytime and anywhere through mobile applets. Whether they are citizens, designers or government workers, they can quickly and easily obtain street renovation plans that conform to local characteristics, greatly improving the efficiency and interactivity of the design. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a system structure diagram of the present invention. DETAILED DESCRIPTION
[0033] The present invention will be further described below with reference to the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it. However, the embodiments are not intended to limit the present invention.
[0034] Reference Figure 1 As shown, an embodiment of the artificial intelligence image generation system applicable to the renovation of historical block facades of the present invention specifically includes the following modules:
[0035] The Flux model loader module relies on cloud computing power to deploy the basic model framework and build a multimodal generation environment compatible with historical block renovation tasks; the image input module collects street status images through the mobile terminal matrix, and integrates denoising and distortion correction technologies to improve input quality; the DeepSeek language interaction module uses intent recognition algorithm to analyze user renovation needs and generate quantifiable style control parameters; the LoRA module innovatively adopts a multi-scale feature partitioning recognition mechanism to construct a three-dimensional annotation including building facade partitioning, component positioning layer, and material semantic layer for core style elements such as the roof type of historical block buildings (hip roof / hard roof / hanging roof), wall color spectrum, decorative component pattern, and street height-width ratio. The system couples regional features through a dynamic weight allocation algorithm. The sampler module incorporates dynamic control technology to balance style innovation with the need to preserve traditional features during the generation process. The Controlnet processing module employs a three-dimensional composite constraint mechanism. Based on the feature distribution heat map output by LoRA model training, it simultaneously applies edge detection to lock building outlines, depth measurement to ensure spatial scale, and semantic segmentation to protect historical components. The image generation module, based on the improved ComfyUI workflow, introduces a feature partitioning attention mechanism to achieve adaptive weighted fusion of multiple control signals. The mobile mini-program module encapsulates the complexity of the underlying technology and provides an interactive interface for the entire process, from image acquisition to solution generation. The Flux model loader module, image input module, DeepSeek language interaction module, LoRA module, sampler module, Controlnet processing module, image generation module, and mobile mini-program module form the ComfyUI workflow, which completes the generation of street renovation images.
[0036] Specifically, the Flux model loader module is used to load the required basic models and components, provide basic parameters for image generation, environment configuration, and deep learning model loading, to ensure the smooth progress of subsequent image processing; the Flux model loader module includes: the UNET loader is used to load the flux large model, and by loading the flux1-schnellfp8.safetensors large model, it provides the necessary environment and parameters for the subsequent image generation process, and dynamically updates various configurations in the image processing process. The innovative addition of the FluxForwardOverrider node in the workflow can merge other functions into the model or modify existing behaviors, ensuring that the original model remains unchanged, thereby maintaining its integrity while still supporting advanced modifications, maximizing the integration of parameter modifications for large models, and improving the problem that traditional parameter adjustments require direct changes to the model structure and destroy the original architecture, so as to facilitate the injection of domain knowledge parameters through the overlay layer and retain the general capabilities of the basic model.
[0037] The image input module is used to collect existing street images and generate street image input data. It supports image input in multiple formats and can collect images in real time through sensors or cameras. The image input module supports multiple sensor interfaces and can be connected to street cameras, drones or other mobile devices to collect and transmit street images in real time.
[0038] The DeepSeek language interaction module is used to receive personalized street renovation requirements input by users through natural language, parse the text content and generate control instructions, thereby adjusting the style, layout and other features of the street image; the DeepSeek language interaction module uses natural language processing technology to support complex user input requirements, generate highly targeted control parameters, and assist in generating specific street renovation images. The Ollama Options V2 node and its affiliated Ollama Generate V2 node and Ollama Connectivity V2 node are added to the workflow construction to call local deepseek and complete the generation of prompt words, and the String Replace node is used to identify key information in the text generated by deepseek as prompt words in the subsequent map generation work. In workflow design, the Ollama Options V2 node addresses the complex and error-prone parameter configuration of large local models. Through a parameter sandbox mechanism, standardized call templates are pre-set, enabling one-click parameter management. The Ollama Connectivity V2 node addresses the challenge of local model communication stability by employing connection pooling and dynamic load balancing technology to ensure reliable interaction with models like DeepSeek, mitigating network fluctuations and resource preemption risks. The Ollama Generate V2 node transforms ambiguous user commands into structured prompts through contextual analysis and logical reorganization, improving model understanding accuracy. The String Replace node utilizes regular expressions and semantic analysis for dual filtering, accurately extracting entity relationships and core descriptors from generated text, eliminating redundant information that interferes with image generation. These four nodes are linked together to form a complete chain of "parameter configuration → model call → command optimization → information extraction," achieving an automated closed loop from language generation to cross-modal creation.
[0039] The LoRA module performs localized style adjustments on images based on the cultural background and architectural style of a specific region. The LoRA module is trained through the manually labeled LoRA training module. Specifically, the collected street images are used to identify the roof form (hip roof / hip roof / hanging roof), analyze the wall color spectrum, extract the pattern of decorative components, and calculate the street space ratio. Based on the manual annotation of historical block elements, a historical block data set is formed and trained to generate street images with the characteristics of historical blocks, and the style can be adjusted according to user needs to meet local culture and design standards. Specifically, independent training is conducted for each specific region, and historical street architectural landscape elements are integrated into the training. The characteristics of historical street landscape elements are classified and professionally identified, so that the generated street images have highly localized characteristics and can adjust the design style according to specific needs, such as adding green belts and adjusting building layouts.
[0040] The sampler module performs diversified sampling adjustments on various parts of the image during the image generation process to optimize image stability, detail expression, and output quality, preventing noise or distortion during the generation process. Specifically, the sampler module dynamically adjusts the sampling steps and sampling rate based on the different detailed areas of the street image, ensuring clear details during the image generation process and overall image stability and distortion. The model sampling algorithm flux node is used in the workflow design. The model sampling algorithm flux node can improve the efficiency of model sampling. It can process data quickly and accurately, allowing the model to obtain the required data more quickly and improve the accuracy of model sampling. Through its careful organization and classification of data, it can enable the model to better understand the data and obtain more accurate sampling results.
[0041] The Controlnet processing module is used to introduce structural constraints during the image generation process. It uses edge detection, semantic segmentation, or depth estimation technology to precisely control building outlines, street layouts, and detailed features to ensure that the generated image conforms to the actual spatial structure and requirements. Specifically, the Controlnet processing module loads the Controlnet constraint model to perform multi-level structural control on the image generation process, including building edge enhancement, street element positioning, and scale adjustment, to improve the spatial rationality and detail accuracy of the generated image. The AIO Aux Preprocessor node is used in the workflow construction to simplify image preprocessing using various auxiliary processors in the ControlNet framework. This node is linked to the ApplyControlNet node to implement depth control and edge control of the image, thereby improving the control of the details of the overall image and reducing the chance of blind opening of the box during image generation.
[0042] The image generation module, based on the image generation technology of the ComfyUI workflow, combines the structural constraints of ControlNet to process the input image generation data, optimizes and transforms street images through techniques such as style transfer, and generates street transformation images that meet the expected requirements. Specifically, the image generation module uses multiple image processing and style optimizations, combined with the structural constraints of ControlNet, to ultimately output street transformation images that meet design requirements and support the user's diverse design needs. The ApplyTeaCachePatch node and UltimatesDupscale node are used in the workflow design. The ApplyTeaCachePatch node focuses on solving the collaborative problem of real-time optimization of AI models and dynamic cache updates. To address the bottleneck of traditional model performance tuning requiring service interruption or full parameter reloading, this node uses incremental cache patch injection technology to seamlessly apply performance enhancement patches while keeping the model running. Its core function is to achieve zero-downtime model upgrades. In principle, it receives a model object (model input parameter) and parses the rel_l 1_thresh threshold to control patch sensitivity. Combined with the device resource allocation strategy specified by cache_device, it dynamically compares the differences between model parameters and cached patches. Wan_coefficients is then used to adjust the stability and performance balance of the initial step. The final output is a patched model that retains the original functionality but improves execution efficiency. Secondly, after image generation, the UltimatesDupscale node is inserted to address the conflict between detail loss and computational resources in high-resolution image generation. This node addresses the pain points of traditional super-resolution algorithms, which are prone to artifacts, blurred textures, or memory overload when the magnification is increased. This node uses multi-scale adaptive enhancement technology to simultaneously optimize resolution and visual fidelity in a single inference. Its core function is to achieve intelligent lossless upscaling of low-resolution input. In principle, it receives the input image (source_image) and parses the target_scale parameter to dynamically build a multi-stage upscaling flow. It combines the adversarial denoising module controlled by denoise_intensity to suppress high-frequency noise, and uses the hardware acceleration strategy specified by the device_priority parameter to implement resource-sensitive computing. Finally, it outputs an upscaled_image that retains the original details, thereby generating a 4K high-definition image and significantly improving the quality of the original image.
[0043] The mobile mini-program module is used for image acquisition and upload, providing a convenient user interface for designers and citizens to review and evaluate renovation plans. It also supports multi-platform interaction and displays the final street renovation results. Specifically, to make it easier for users to use the mini-program, the encapsulated mobile mini-program design only displays the image input module, text input module, and image preview module. These modules receive input images through the user interface and display the corresponding renovation plan. Users can provide feedback and make modifications based on the system-generated images, greatly enhancing user engagement.
[0044] Based on the aforementioned system, the application also discloses an AI-powered image generation method for updating facades in historic districts, aiming to address the localization and personalization needs of existing street designs. By integrating the DeepSeek language model and the LoRA model, the invention can generate street design images that are consistent with local characteristics and cultural backgrounds based on user natural language input. Specifically, it includes the following steps:
[0045] First, train the LoRA model through the manual annotation LoRA training module:
[0046] To achieve local street design characteristics, the present invention first needs to train the LoRA model so that it can adapt to the culture and architectural style of a specific region. In this step, the specific operations are as follows:
[0047] Installing the trainer: First, you need to install Fluxgym, the LoRA model trainer, and configure the relevant runtime environment. Preparing training data: Collect and prepare street view imagery data for a specific area, batch-crop the collected street view imagery, and annotate these images. Annotation: Manually annotate street elements such as buildings, green belts, and roads in detail. Detailed annotation of historical and cultural block elements is performed according to the classification of historical and cultural block elements, allowing the model to recognize and understand various design elements.
[0048] Start training: Select the flux model to be trained. This time, we use flux_dev8. Adjust the training parameters (Repeat trains per image, 10, max train epochs, 16) to suit your computer's performance. Select a VRAM allocation suitable for your computer's performance. Then, start the LoRA model training process. The goal of training is to enable the model to learn local characteristics from the dataset by adjusting the design style.
[0049] Screening the best model: By analyzing the loss value during the training process, evaluating the performance of different models, and finally selecting the LoRA model with the best effect, and applying it to the image processing in the subsequent steps, that is, obtaining the core model.
[0050] Users can then use cameras or other sensors to capture images of existing streets. These images are fed into the system and processed through the following steps:
[0051] Image acquisition: Users capture street images through mobile devices or other devices to ensure that the image quality meets design requirements. ComfyUI is equipped with an image input module to receive and transmit image data from cameras or sensors. The image input module inputs this image data into the ComfyUI workflow for use.
[0052] Users can input personalized street design requirements through natural language. For example, they might enter requirements such as "increase green belts" or "adjust architectural style." The DeepSeek language model converts these requirements into structured control instructions that the system can understand:
[0053] Deploy the DeepSeekR1 model: Deploy the DeepSeekR1 model locally to ensure that the system can efficiently process language input. Configure the OLLAMA runner: Install and configure the OLLAMA runner to support the operation of the DeepSeek model.
[0054] Text input and component integration: Configure the components required for DeepSeek and OLLAMA to run in ComfyUI, and integrate text input and related components into the ComfyUI workflow to ensure seamless integration of language processing and image generation.
[0055] In order to improve the quality of the generated image, the system will use the sampler module to optimize the details of the processed image generation data. The specific steps are as follows:
[0056] Configure the sampler component: Configure the sampler component in ComfyUI and adjust the sampling parameters during the image generation process to ensure that the generated image details are clear and stable with reduced distortion.
[0057] Image optimization: The generated image is fine-tuned through a sampler to ensure that the image quality meets the design standards. The Lora model processes the generated image data based on user needs and the culture and architectural style of a specific region to ensure that the generated street design is consistent with local characteristics. This process includes:
[0058] Configure the Lora model loader: Configure the Lora model loader in ComfyUI to ensure that the system can load and use it. The trained Lora model is used to adjust the model weight parameters (model strength 0.6).
[0059] Localized style adjustment: The Lora model will adjust the architectural style, road layout, green elements, etc. of the street design based on the loaded local characteristic data to make it conform to the local culture and architectural style.
[0060] During the image optimization process, ControlNet structural constraints were also added. The ControlNet module was used to perform spatial structural control on the building outlines and street layouts in the generated images, ensuring the physical consistency between the design and the actual scene:
[0061] Configure the ControlNet component: Load a pre-trained ControlNet model (edge detection model or semantic segmentation model) in ComfyUI and set structural constraints for key parameters such as building height and road width.
[0062] Apply constraint techniques: For example, edge detection can be used to enforce building boundaries, and depth estimation can be used to adjust street perspective scale to avoid structural distortion or unreasonable layout in the generated image.
[0063] After image optimization, we can obtain image generation data. The image generation data is used to generate the final street transformation image through the image generation module. In the image generation step, the system will perform style optimization on the processed image based on the ComfyUI workflow to generate the final street transformation image. The specific steps include:
[0064] Configure image output components: Configure image output components in ComfyUI to ensure that the generated images can be saved and transmitted to users.
[0065] Style Optimization: Based on user needs and local characteristics, the image style is further optimized to ensure that the final street design meets the user's personalized requirements. The generated image node is linked to the magnified model (4x-UltraSharp model) for image magnification and generates 4K high-definition image output to ensure output image quality.
[0066] The mobile mini-program module then displays the results and provides feedback. The generated street renovation images are then displayed to users via the mobile mini-program. Users can review the images and provide feedback to ensure that the design meets their expectations. This step includes:
[0067] Mini program packaging: Install the ComfyUI_Bxb component in the ComfyUI node manager, and package the image input, text input, and image generation components into mini programs, allowing users to view and modify design plans at any time through their mobile phones.
[0068] User feedback and modification: Users provide feedback on the generated images through the mini-program. If modification is required, the system will make adjustments based on the feedback to ensure that the final design meets the user's needs.
[0069] After user feedback and modification, the system will generate a street renovation plan that meets the user's personalized needs and regional characteristics. The final results include:
[0070] Generate final image: Based on all processing and feedback, the system generates the final street improvement design image.
[0071] Provide reference images: The final street design images will be provided as reference for designers or relevant departments for actual street reconstruction implementation.
[0072] The above embodiments are only preferred embodiments for fully illustrating the present invention, and the protection scope of the present invention is not limited thereto. Any equivalent substitution or modification made by those skilled in the art based on the present invention is within the protection scope of the present invention.
Claims
1. An artificial intelligence image generation system suitable for updating the facades of historical blocks, characterized by: include: The Flux model loader module is used to load basic models and components, provide basic parameters for image generation, environment configuration, and model loading; Image input module, used to collect existing street images; The DeepSeek language interaction module is used to receive personalized street renovation requirements input by users through natural language, parse the content and generate structured control instructions; Controlnet processing module, used to generate structural constraints; a sampler module for collecting image input data, structured control instructions, and structural constraints of street images and generating graph generation data; LoRA module, used to perform localized style adjustment on image generation data; The graph generation module is used to decode the graph generation data to generate a street transformation image and output the street transformation result; The mobile app module is used for user input, image acquisition and upload, and displays the final street transformation results; The Flux model loader module, image input module, DeepSeek language interaction module, LoRA module, sampler module, Controlnet processing module, image generation module and mobile mini-program module are combined to form the ComfyUI workflow, and the street transformation image generation is completed through the ComfyUI workflow.
2. The artificial intelligence image generation system for updating facades in historic blocks according to claim 1, characterized in that: The Flux model loader module includes a UNET loader, which is used to load the flux large model. The flux large model is specifically the flux1-schnel lfp8.safetensors large model. The flux large model is used to provide the environment and parameters for the image generation process and dynamically update the configuration during the image processing process. A FluxForwardOverrider node is added to the flux large model. The FluxForwardOverrider node is used to merge functions or modify existing behaviors to ensure that the original model remains unchanged while supporting advanced modifications.
3. The artificial intelligence image generation system for updating facades in historic blocks according to claim 1, characterized in that: The image input module is connected to street cameras, drones and / or mobile devices to collect and transmit street images in real time.
4. The artificial intelligence image generation system for updating facades in historic blocks according to claim 1, characterized in that: The LoRA module is trained through the manually annotated LoRA training module. First, the input data of the collected street image samples is extracted and calculated, and the historical block dataset is formed based on manual annotation for training. The manually labeled LoRA training module incorporates historical street architectural landscape elements for training, classifies and professionally identifies the characteristics of historical street landscape elements, and makes the generated images have highly localized features.
5. The artificial intelligence image generation system for updating facades in historic blocks according to claim 1, characterized in that: The sampler module dynamically adjusts the sampling steps and sampling rate according to the different detail areas of the street image to ensure that the image details are clear and distortion-free during the generation process; the sampler module uses the model sampling algorithm flux node to perform model sampling.
6. The artificial intelligence image generation system for updating facades in historic blocks according to claim 1, characterized in that: The Controlnet processing module performs multi-level structural control on the image generation process by loading the Controlnet constraint model, including building edge reinforcement, street element positioning and scale adjustment.
7. The artificial intelligence image generation system for updating facades in historic blocks according to claim 1, characterized in that: The mobile mini-program module includes an image input module, a text input module, and an image preview module. Users can provide feedback and make modifications based on the generated image.
8. An artificial intelligence image generation method suitable for updating the facades of historical blocks, characterized by: The system according to any one of claims 1 to 7 comprises the following steps: Step S1: Train the Lora model to obtain the core model; Step S2: Use a camera or sensor to capture an existing street view image or upload a street view image that needs to be modified and input it into the ComfyUI workflow; Step S3: parse the user input requirements through natural language processing technology and convert them into structured control instructions; Step S4: During the image generation process, the sampler module samples the street view image and generates graph generation data based on structured control instructions and ControlNet structural constraints. The graph generation data is also adjusted for local style by the core model. Step S5: the image generation module decodes the adjusted image generation data to generate a street reconstruction image and outputs the street reconstruction result; Step S6: The user views the generated street reconstruction result map through the mobile mini program. If it meets the requirements, the generation is completed. If it does not meet the requirements, jump to step S3 to continue execution.
9. The artificial intelligence image generation method for updating the facades of a historic block according to claim 8, characterized in that: The training process of the Lora model is as follows: First, install the Lora model trainer Fluxgym and its operating environment, collect street view images obtained from field research, batch-crop these street view images to 512*512 pixels, perform building roof shape recognition, wall color spectrum analysis, decorative component pattern extraction, and street space ratio calculation on the collected street view images, use Fluxgym's annotation tools for data annotation, and manually annotate historical block elements in the street view images. N manually annotated street view images are used as training samples, and the training parameters are adjusted and the training process is started. After training is completed, the loss value of the Lora model is evaluated and optimized, and the best performing model is selected as the core model for image generation.
10. The artificial intelligence image generation method for updating the facades of a historic block according to claim 8, characterized in that: Install the ComfyUI_Bxb component in the ComfyUI Node Manager and encapsulate the image input, text input, and image generation components into a mobile applet to facilitate users to view and adjust the generated street renovation design through mobile devices.
Citation Information
Cited By
A street surface update identification method fusing time-series street view data and remote sensing data
CN122368611A
A street surface update identification method fusing time-series street view data and remote sensing data
CN122368611B