Country style and appearance design image generation method and system based on knowledge graph and large model

By constructing a knowledge graph and multi-level labeling system covering the core elements and hierarchical classification of rural landscape design, combined with the basic diffusion model architecture, the problem of unstable quality of rural landscape design image generation in existing technologies is solved, and efficient and accurate image generation is achieved.

CN120747285AActive Publication Date: 2025-10-03GUANGDONG URBAN & RURAL PLANNING & DESIGN INST

Patent Information

Application Number
CN202511180385.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-10-03
Estimated Expiration
2045-08-22

AI Technical Summary

Technical Problem

The existing large-scale image generation models lack an accurate understanding of the professional terminology and regional cultural symbols in the field of rural landscape design, resulting in unstable quality of generated images and difficulty in accurately matching rural scenes.

Method used

Construct a knowledge graph covering the core elements and hierarchical classification of rural landscape design, annotate sample images based on a multi-level label system to form a dedicated training set, and train through the basic diffusion model architecture to generate rural landscape design images.

Benefits of technology

The efficiency and quality of rural landscape design image generation have been improved, ensuring that the generated images achieve optimal results in terms of regional adaptability and style consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747285A_ABST
    Figure CN120747285A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image generation, and discloses a country style and appearance design image generation method and system based on a knowledge graph and a large model, and the method comprises the steps: constructing a core element covering country style and appearance design and a knowledge graph; constructing a multi-level label system based on the knowledge graph; pre-processed rural style and appearance sample images are obtained, the rural style and appearance sample images are labeled based on a multi-level label system, and a special training set in the field of rural style and appearance design is obtained; constructing a basic diffusion large model architecture, and training the basic diffusion large model architecture by using the special training set for the country style and appearance design field to obtain a trained country style and appearance design image generation model; and generating a rural style and appearance design image by using the trained rural style and appearance design image generation model. According to the method, the problems that a general large model is insufficient in domain knowledge support and insufficient in data labeling specialty in rural style and appearance design image generation are solved, and the rural style and appearance design image generation efficiency and quality are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image generation technology, and more specifically, to a method and system for generating rural landscape design images based on a knowledge graph and a large model. Background Art

[0002] Rural landscape design, a key means of improving rural environmental quality and stimulating endogenous rural vitality, is experiencing increasing demand. Traditional design methods, relying on manual modeling and rendering, are inefficient and limited in accuracy, making them inadequate for current rural landscape design needs. Therefore, the introduction of intelligent tools is inevitable. AI-based large-scale image generation models, with their powerful learning capabilities and efficient rendering, offer a new approach to rural landscape design.

[0003] The current mainstream image generation large models (such as StableDiffusion and Midjourney) mainly rely on general corpora or material libraries for training, and lack an accurate understanding of professional terminology and regional cultural symbols in the field of rural landscape design. As a result, the generated images have problems such as unstable quality, incompatibility with rural regional characteristics, and difficulty in accurately fitting rural scenes. At the same time, the existing professional large models in the field of urban and rural planning mostly focus on urban planning design, urban landscape design or general architectural design, and lack the systematic construction of a full-factor knowledge graph in the field of rural landscape design and high-quality annotation and training of image data, resulting in the existing technology having defects in quality and efficiency in rural landscape design image generation. Summary of the Invention

[0004] In order to overcome the defects of low image generation quality and efficiency in existing rural landscape design image generation technologies, the present invention proposes the following technical solutions: In a first aspect, the present invention proposes a method for generating rural landscape design images based on a knowledge graph and a large model, comprising: Construct a knowledge map covering the core elements and hierarchical classification of rural landscape design; Constructing a multi-level labeling system based on the knowledge graph; Obtaining preprocessed rural landscape sample images, and labeling the rural landscape sample images based on the multi-level labeling system to obtain a special training set for the rural landscape design field; Constructing a basic diffusion model framework, and training the basic diffusion model framework using the rural landscape design field-specific training set to obtain a trained rural landscape design image generation model; Use the trained rural landscape design image generation model to generate rural landscape design images.

[0005] As the preferred technical solution, a knowledge graph covering the core elements and hierarchical classification of rural landscape design is constructed, including: Perform hierarchical structural division on the pre-screened core entities of rural landscape to obtain hierarchical classification results; According to the hierarchical classification results, a structured triple containing the head entity, relationship and tail entity is constructed; Map the head entity, relationship, and tail entity in the structured triple to a vector space, and optimize the vector representation of the head entity vector and the relationship vector so that the spatial distance between the sum of the head entity vector and the relationship vector and the tail entity vector meets a preset distance threshold condition; Based on the vectorized triples that meet the preset distance threshold condition, a rural landscape design knowledge graph is constructed.

[0006] As a preferred technical solution, the basic diffusion model architecture includes a variational autoencoder module, a noise prediction module, a text encoding module and a multi-channel control module; the variational autoencoder module includes an encoding unit and a decoding unit; The encoding unit is used to map the input sample image to a low-dimensional latent space to obtain a latent variable; The text encoding module is used to encode the input design requirement text to obtain a semantic embedding vector; The multi-channel control module is used to extract multimodal features of the input sample image and the design requirement text, and perform feature fusion processing on the multimodal features to generate multimodal conditional fusion features; the multimodal features include edge and morphological control features, spatial perception and semantic control features, style and content editing features, and special function features; The noise prediction module is used to fuse the latent variable, the semantic embedding vector and the multimodal conditional fusion feature, and perform noise prediction and denoising operations on the fusion result to obtain a denoised latent variable; The decoding unit is used to reconstruct the denoised latent variables into a rural landscape design image.

[0007] As a preferred technical solution, the noise prediction module fuses the latent variable, the semantic embedding vector and the multimodal conditional fusion feature, and performs noise prediction and denoising operations on the fusion result to obtain the denoised latent variable, including: Perform cross attention fusion on the latent variable, the semantic embedding vector and the multimodal conditional fusion feature to generate an attention fusion feature , whose expression is as follows:

[0008] Where, is the current time step t The latent variables, is the semantic embedding vector, is the multimodal conditional fusion feature, is the attention query matrix, is the bond matrix K The transpose of is the dimension of a single attention head in the attention mechanism, is the value vector of the attention mechanism, is a learnable parameter matrix used to map latent variables and semantic embedding vectors to the attention query matrix Q, is used to map latent variables and semantic embedding vectors into key matrices K The learnable parameter matrix of Attention-based fusion features , noise prediction is performed through the frozen main branch and conditional branch of the multi-channel control module, and its expression is as follows:

[0009] Where, represents the total prediction noise, represents the prediction noise of the frozen main branch, is a scaling factor that controls the strength of the condition, represents the prediction noise of conditional branches; Based on the total prediction noise, construct the latent variable The ordinary differential equation for the inverse evolution of is expressed as follows:

[0010] Where, is the rate of change function of inverse denoising, is the current time step t The noise scheduling parameters, is the current time step t Diffusion parameters; According to the latent variables The ordinary differential equation of the inverse evolution of the forward diffusion of the final noisy latent variable At the beginning, T is the maximum time step, and the sampler is used to iteratively solve the differential relationship of the ordinary differential equation. When the time step regresses to the initial stage, the denoised latent variable is generated. .

[0011] As a preferred technical solution, the edge and morphology control features include Canny edge features, probabilistic edge features, straight line morphology features and line draft morphology features; The multi-channel control module extracts edge and morphological control features of the input sample image, including: The Canny operator is used to perform Gaussian filtering, Sobel operator gradient calculation, non-maximum suppression and double threshold processing on the input sample image in sequence to generate Canny edge features. ;in, and are the height and width of the feature, respectively; Extract the probabilistic edge features of the input sample image based on PIDI-Net, and generate probabilistic edge features based on the probabilistic edge features ; Perform straight line structure detection on the input sample image, parametrically extract the endpoint coordinates of the straight line, and generate straight line morphological features ;in, is the number of detected lines; Extract multi-scale line draft features of the input sample image based on the HED network, and generate line draft morphological features based on the multi-scale line draft features .

[0012] As a preferred technical solution, the spatial perception and semantic control features include depth information features and semantic segmentation features; The multi-channel control module extracts spatial perception and semantic control features of the input sample image, including: The MiDaS network is used to predict the monocular depth of the input sample image and generate depth information features based on the monocular depth. ;in, and are the height and width of the feature, respectively; Based on the Deeplabv3+ model, perform semantic segmentation on the input sample image according to the ADE20K protocol to generate semantic segmentation features ;in, is the number of object categories, and each channel corresponds to the segmentation mark of a type of rural landscape elements; As a preferred technical solution, the style and content editing features include style vector features, low-rank adaptation matrix features and editing mask features; The multi-channel control module extracts the style and content editing features of the input sample image and the design requirement text, including: Use the IP-Adapter image prompt adapter to encode the input sample image and obtain the style vector feature ;in, d is the dimension of the feature; Based on the design requirement text, a low-rank adaptation matrix feature is constructed through the T2I-Adapter text-to-image adapter. ; Generate edit mask features based on design requirement text ; The elements in the edit mask feature use binary values ​​to identify the areas in the input sample image that need to be redrawn.

[0013] Construct local redrawing mask features based on the rural landscape restoration requirements in the design requirements text ; Among them, the local redrawing mask feature Used to constrain the diffusion model to generate rural landscape details in a specified area.

[0014] As a preferred technical solution, the basic diffusion model architecture is trained using the dedicated training set in the field of rural landscape design, including: In all cross-attention layers of the noise prediction module of the basic diffusion model, a low-rank adapter based on LoRA is injected. The first stage of training is performed by constructing the first objective function including CLIP semantic alignment loss and LPIPS perceptual loss to obtain a preliminary training model. Obtain the binary evaluation results and the number of edits for the generated rural landscape design images; constructing a reward function based on the binary evaluation result and the number of edits; Based on the reward function, using a policy optimization algorithm, constructing a second objective function including an importance sampling ratio and an advantage function of the generated action; The preliminary training model is trained in the second stage based on the second objective function to obtain a trained rural landscape design image generation model.

[0015] As a preferred technical solution, the first objective function The expression is as follows:

[0016] Where, and are weight parameters, Design images of rural landscapes generated by the model, For the corresponding design text, Sample images from a dedicated training set for rural landscape design. Used to measure the semantic matching between the generated graph and the text, Image features extracted by the VGG network; Reward Function The expression is as follows:

[0017] Where, To generate binary evaluation results for the rural landscape design images, =1 means passed, =0 means reject, Indicates the number of times the rural landscape design images need to be manually edited; Second objective function The expression is as follows:

[0018] Where, To find the mathematical expectation of the data at time step t during the training process, is the importance sampling ratio, Indicates the current policy Generating status in the countryside Take action The probability of Indicates the old policy before the update Generating status in the countryside Take action The probability of is the advantage function used to measure the action of generating rural landscape design images a Compared to the value gain of the average action, Indicates that Clip to interval [ ].

[0019] In a second aspect, the present invention further proposes a rural landscape design image generation system based on a knowledge graph and a large model, which is applied to the rural landscape design image generation method based on a knowledge graph and a large model as described in any solution of the first aspect, comprising: The first building block is used to construct a knowledge graph covering the core elements and hierarchical classification of rural landscape design; A second building module is used to build a multi-level labeling system based on the knowledge graph; a labeling module, configured to obtain pre-processed rural landscape sample images, label the rural landscape sample images based on the multi-level labeling system, and obtain a special training set for the rural landscape design field; The third construction module is used to construct a basic diffusion model framework, and train the basic diffusion model framework using the dedicated training set in the field of rural landscape design to obtain a trained rural landscape design image generation model; The generation module uses the trained rural landscape design image generation model to generate rural landscape design images.

[0020] The beneficial effects of the present invention include at least: The present invention addresses the problems of insufficient domain knowledge support and lack of professional data annotation in the generation of rural landscape design images by general large models. By constructing a knowledge graph covering the core elements and hierarchical classification of rural landscape design, the model's semantic understanding of rural domain terms and features is strengthened; a multi-level labeling system is built based on the knowledge graph to accurately annotate the pre-processed rural landscape sample images to form a dedicated training set; the basic diffusion large model architecture is trained with the training set to enable the model to deeply learn the professional characteristics of rural landscape, and finally generate images with the help of the trained model, thereby improving the efficiency and quality of rural landscape design image generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 A schematic flow chart of a method for generating rural landscape design images based on a knowledge graph and a large model provided in an embodiment of the present invention.

[0022] Figure 2 A schematic diagram of the technical architecture of the rural landscape design image generation method provided by an embodiment of the present invention.

[0023] Figure 3 This is an example of the rural landscape design knowledge graph provided by an embodiment of the present invention.

[0024] Figure 4 Among them, (a) is the original sample image I of the rural landscape provided by an embodiment of the present invention, (b) is the Canny edge feature map provided by an embodiment of the present invention, (c) is the straight line morphological feature map provided by an embodiment of the present invention, (d) is the line draft morphological feature map provided by an embodiment of the present invention, (e) is the probabilistic edge feature map provided by an embodiment of the present invention, (f) is the depth information feature map provided by an embodiment of the present invention, (g) is the semantic segmentation feature map provided by an embodiment of the present invention, and (h) is the graffiti constraint feature map provided by an embodiment of the present invention.

[0025] Figure 5 This is a training loss curve provided by an embodiment of the present invention.

[0026] Figure 6 A schematic diagram of an X / Y visualization chart provided by an embodiment of the present invention.

[0027] Figure 7 The embodiment of the present invention provides a rural landscape design image generation system based on knowledge graph and large model.

[0028] Figure 8 A schematic diagram of the flow of the multi-dimensional control optimization strategy provided by an embodiment of the present invention.

[0029] Figure 9Among them, (a) is the original sample image II of the rural landscape provided by an embodiment of the present invention, (b) is the generated image of the StableDiffusion native basic model provided by an embodiment of the present invention, (c) is the generated image of the rural landscape design image generation model trained by the present invention provided by an embodiment of the present invention, and (d) is the generated image of the general large model provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0030] The following will describe embodiments of the present invention with reference to the accompanying drawings and preferred technical solutions. Those skilled in the art will readily understand other advantages and benefits of the present invention from the contents disclosed in this specification. The present invention may also be implemented or applied through different specific embodiments, and the details in this specification may be modified or altered based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred technical solutions are intended only to illustrate the present invention and are not intended to limit the scope of protection of the present invention.

[0031] It should be noted that the illustrations provided in the following embodiments are merely schematic illustrations of the basic concept of the present invention. Therefore, the illustrations only show components related to the present invention and are not drawn according to the number, shape, and size of components in actual implementation. In actual implementation, the type, quantity, and proportion of each component may be changed arbitrarily, and the component layout may also be more complex.

[0032] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the embodiments of the present invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the embodiments of the present invention.

[0033] Example 1 This embodiment proposes a method for generating rural landscape design images based on knowledge graph and large model. Figure 1 As shown, Figure 1 This is a flow chart of a method for generating a rural landscape design image based on a knowledge graph and a large model provided in this embodiment. The method includes the following steps: S1: Construct a knowledge graph covering the core elements and hierarchical classification of rural landscape design; S2: Construct a multi-level labeling system based on the knowledge graph; S3: Obtain preprocessed rural landscape sample images, and label the rural landscape sample images based on the multi-level labeling system to obtain a special training set for the rural landscape design field.

[0034] In one example, renderings and photos of rural landscapes with regional characteristics and local cultural features were collected through multiple channels, including design websites, on-site photography, and rural planning and production projects. The collected images were then subjected to denoising, contrast enhancement, and size normalization using the OpenCV library. The image resolution was uniformly adjusted to no less than 512×512 pixels, and the format was uniformly converted to PNG. Images with unclear subjects, disorganized expressions, or ambiguous content were discarded, resulting in a high-quality dataset of sample rural landscape images. Subsequently, based on the rural landscape knowledge graph, a self-developed annotation tool embedded in the rural landscape knowledge graph was used to manually annotate the sample images. The labeling system was set to a four-level system. The first-level labels divided the rural landscape into six major scenes: natural environment, architecture, landscape, facilities, plants, and general scenes. The second and third-level labels further refined the rural landscape categories and specific elements, such as topography, building types, waterfront landscapes, road facilities, trees and plants, and seasonal weather. The fourth-level labels were further refined to include the characteristic attributes, style, and material of elements such as stone arch bridges, sloping roofs, and pebble tree ponds. Compared with the WD1.4 labeler, the labeling professional level of this self-developed labeling tool has increased by more than 75%; compared with BooruDatasetTagManager, the labeling efficiency has increased by more than 50%. Finally, based on the labeling results of this multi-level labeling system, a special training set for the field of rural landscape design was constructed.

[0035] S4: constructing a basic diffusion model architecture, and training the basic diffusion model architecture using the rural landscape design field-specific training set to obtain a trained rural landscape design image generation model; S5: Generate rural landscape design images using the trained rural landscape design image generation model.

[0036] During the specific implementation process, we first sorted out the core elements of rural landscape design, such as rural architecture, landscape, and natural environment, divided the multi-level classification structure of "scene-category-component", established the correlation between elements, and constructed a knowledge graph for rural landscape design; then, based on the hierarchical classification of the knowledge graph, we designed multi-level labeling rules covering scene types and element attributes; then, we collected images from channels such as design renderings and rural real-life photos, and after pre-processing such as denoising and size normalization, we screened high-quality samples and annotated them according to multi-level labeling rules to form a special training set for the field of rural landscape design; then, we selected the basic diffusion model of the StableDiffusion class, imported the special training set, set the training parameters to carry out model training, and obtained the rural landscape design image generation model; finally, we input the text requirements for rural landscape design, and generated the corresponding rural landscape design images through the trained model.

[0037] It is understandable that in order to address the problems of insufficient domain knowledge support and lack of professional data annotation in the generation of rural landscape design images by general large models, we can strengthen the model's semantic understanding of rural domain terms and features by constructing a knowledge graph covering the core elements and hierarchical classification of rural landscape design; build a multi-level labeling system based on the knowledge graph, and accurately annotate the pre-processed rural landscape sample images to form a dedicated training set; rely on the training set to train the basic diffusion large model architecture, so that the model can deeply learn the professional characteristics of rural landscape, and finally generate images with the help of the trained model, so as to improve the efficiency and quality of rural landscape design image generation.

[0038] Example 2 This embodiment makes improvements based on the rural landscape design image generation method based on knowledge graph and large model proposed in embodiment 1. Figure 2 As shown, Figure 2 This is a schematic diagram of the technical architecture of the rural landscape design image generation method provided by an embodiment of the present invention. It should be noted that in this architecture, the base model and its core modules form the platform's technical support layer. This module provides the fundamental capabilities for image generation. This layer includes key components such as the VAE (Variational Autoencoder) module, the U-Net module, and the CLIP module. The VAE module extracts image features in the latent space through efficient encoding and decoding processes, ensuring high-quality and detailed image generation. The U-Net module utilizes its symmetrical structure and skip connections to effectively preserve details during image generation while improving image resolution and accuracy. The CLIP module uses comparative learning between text and images to ensure the platform can accurately generate design images related to rural landscapes based on natural language descriptions, enhancing the system's semantic understanding. The base model is built on StableDiffusionXL, which provides powerful image generation capabilities and produces high-quality visual content.

[0039] The platform's specialized engine is a large, specialized model for generating rural landscape design images. Building on this foundational model, this implementation utilizes LoRA (Low-Rank Adaptation) technology to conduct targeted training on this specialized model. This module, deployed on the AIGC Creative Design Platform, enhances the model's adaptability to specific rural landscape characteristics. This module enables the platform to tailor generated images to the culture, style, and needs of specific rural areas, ensuring optimal regional adaptability and stylistic consistency.

[0040] The application platform serves as the user interface for the entire platform, offering a series of functional modules to support the entire rural landscape design process. The platform includes four functional modules, including the Basic Creation Module, which encompasses four key functions: "Cultural Image," "Image-Based Image," "Partial Redrawing," and "High-Definition Restoration." The Design Scenario Module encompasses rural design scenarios such as agricultural landscapes, portal nodes, waterfront landscapes, and road landscapes. Users can select the appropriate scenario for their design needs. The Inspiration Market and My Gallery modules provide designers with a space for creative exchange and storage, allowing them to share inspiration and manage their work.

[0041] Optionally, a knowledge graph covering the core elements and hierarchical classification of rural landscape design is constructed, including: Perform hierarchical structural division on the pre-screened core entities of rural landscape to obtain hierarchical classification results; According to the hierarchical classification results, a structured triple containing the head entity, relationship and tail entity is constructed; Map the head entity, relationship, and tail entity in the structured triple to a vector space, and optimize the vector representation of the head entity vector and the relationship vector so that the spatial distance between the sum of the head entity vector and the relationship vector and the tail entity vector meets a preset distance threshold condition; Constructing a rural landscape design knowledge graph based on the vectorized triples that meet the preset distance threshold condition; ;

[0042] In one example, as shown in Table 1 and Figure 3 As shown in Table 1, the core entity system of rural landscape design is defined in a four-level classification architecture system, covering the six core areas of "natural environment, architecture, landscape, facilities, plants, and general scenes". First, the core entities of architecture and landscape are pre-screened and hierarchically divided: the "farmhouse building" under the architecture category is subdivided into the four-level entity "sloping roof farmhouse" according to the "roof form", and the "public space landscape" under the landscape category is subdivided into the four-level entity "Lingnan style public space" according to the "regional style"; secondly, a structured triple is constructed: with "sloping roof farmhouse" as the head entity h 、"Style Type" is the relationship r 、"Lingnan style public space" is the last entity t , forming a triplet ( h , r , t); Then, the head entity, relationship, and tail entity are mapped to a low-dimensional vector space through the TransE algorithm, and the vector representation is optimized so that the spatial distance (such as the Euclidean distance) between the sum of the head entity vector and the relationship vector and the tail entity vector meets the preset distance threshold; finally, the iteration is extended to the remaining four core entities, such as natural environment, facilities, plants, and general scenes (for example, the "mountain environment" of the "natural environment" category is subdivided into "multi-slope mountain", and the "stone wall" of the "facility" category is subdivided into "Lingnan stone wall", and triples (({Lingnan stone wall},{spatial association},{Lingnan banyan tree})) are constructed). Based on the vectorized triples that meet the distance threshold, a rural landscape design knowledge graph is systematically constructed.

[0043] Optionally, the basic diffusion model architecture includes a variational autoencoder module, a noise prediction module, a text encoding module and a multi-channel control module; the variational autoencoder module includes an encoding unit and a decoding unit; The encoding unit is used to map the input sample image to a low-dimensional latent space to obtain a latent variable.

[0044] In this embodiment, the variational autoencoder module adopts a variational autoencoder (VAE).

[0045] It should be noted that the variational autoencoder is a probability-based generative model whose core goal is to achieve efficient data generation and reconstruction by learning the potential distribution of data. VAE assumes that the input data By latent variables Generated by the decoder, i.e. , which follows the prior distribution , usually a standard Gaussian distribution , encoder Parameterize the posterior distribution through the neural network and output the mean and variance , which maps the input data into the latent space. Decoder The data is reconstructed from the latent variables, usually using Gaussian distribution or Bernoulli distribution. Since the true posterior distribution is difficult to solve directly, VAE introduces an approximate posterior through variational inference. , and maximizes the evidence lower bound (ELBO) on the log-likelihood:

[0046] The loss function consists of two parts: (1) reconstruction loss, which measures the difference between the input data and the decoded output, and is usually calculated using mean square error (MSE) or cross entropy; (2) KL divergence, which constrains the similarity between the latent distribution and the prior distribution to avoid overfitting.

[0047] Input sample image , after VAE encoder Mapping to a low-dimensional latent space:

[0048] In this embodiment, the potential distribution is constrained by variational inference , reducing the computational complexity of the diffusion model. Forward noise injection, the latent variable Perform a Markov chain noise superposition and define a forward diffusion process:

[0049] in Controlled by the cosine scheduling strategy, the noise level is kept from to Monotonically increasing, eventually .

[0050] The text encoding module is used to encode the input design requirement text to obtain a semantic embedding vector.

[0051] In this embodiment, the text encoding module adopts CLIP (Contrastive Language-Image Pre-training) TextEncoder.

[0052] It should be noted that CLIP is a multimodal model whose text encoder is based on the Transformer architecture. It achieves semantic alignment between text and images through contrastive learning. Its structure uses a decoder-only Transformer, including multi-head self-attention (MHA) and feed-forward network (FFN) layers. After the input text is word-embedded and position-encoded, the self-attention mechanism captures contextual relationships and ultimately extracts a token vector as the overall representation of the text. CLIP is trained on a large-scale image-text pair dataset, using the infoNCE loss as the objective function:

[0053] in and are the normalized embeddings of images and texts, respectively, is the temperature parameter. This loss maximizes the similarity of positive sample pairs and suppresses negative sample pairs. CLIP maps text and images into the same semantic space, so that text descriptions can be directly used as guiding conditions for the generation model.

[0054] During the iterative optimization of the large model, the design requirement text y is input and the semantic embedding vector is generated through the pre-trained CLIP text encoder , where L is the sequence length and d=768 is the embedding dimension, and the visualization is:

[0055] The multi-channel control module is used to extract multimodal features of the input sample image and the design requirement text, and perform feature fusion processing on the multimodal features to generate multimodal conditional fusion features; the multimodal features include edge and morphological control features, spatial perception and semantic control features, style and content editing features, and special function features; The noise prediction module is used to fuse the latent variable, the semantic embedding vector and the multimodal conditional fusion feature, and perform noise prediction and denoising operations on the fusion result to obtain a denoised latent variable; In this implementation, the noise prediction module adopts the U-Net+Scheduler structure.

[0056] It is important to note that the U-Net's symmetrical encoder-decoder structure and skip connections make it a core architecture for denoising tasks in diffusion models. Its design consists of an encoder (contracting path), a decoder (expanding path), and skip connections. This involves progressive downsampling through convolution and pooling to extract multi-scale features without sacrificing global context. Then, upsampling through deconvolution or interpolation gradually restores spatial resolution. Skip connections are used to fuse local details corresponding to the encoder, concatenating feature maps from each encoder layer with those from the decoder layer. This allows for the fusion of low-level details (such as edges) with high-level semantics (such as object shape), enhancing segmentation and image generation accuracy. Between the encoder and decoder, the UNetModel's skip connection mechanism connects features from corresponding encoder layers to those from corresponding decoder layers, helping to preserve more spatial information and detailed features. During image generation, the UNetModel employs downsampling and upsampling, and integrates multiple modules, including the ResBlock residual module, the timestep_embedding timestep embedding module, and the SpatialTransformer spatial transformer.

[0057] The decoding unit is used to reconstruct the denoised latent variables into a rural landscape design image.

[0058] Optionally, this embodiment constructs a conditional reverse process , using U-Net noise prediction network Estimate the noise residual and use Controlnet as a multi-channel control module to introduce multimodal conditional input to enhance the control ability of text-to-image diffusion. The core idea is to freeze and train the weights of the diffusion model and build a trainable conditional encoding branch to achieve precise alignment between the generated content and the input constraints. Controlnet consists of two parallel networks: the frozen main branch retains all the parameters of the trained diffusion model (StableDiffusionU-Net) to ensure that the original generation ability is not degraded; the trainable conditional branch copies the encoding structure of the main branch and is connected to the main branch through a zero-initialized convolution layer (Zero-Conv). The weights are close to zero in the initial stage to avoid interfering with the original model. The network integrates three parts of information: the noise latent variable , time step embedding (injecting temporal information through sinusoidal positional encoding), text conditional embedding and multimodal conditional fusion features Perform cross attention fusion to generate attention fusion features , whose expression is as follows:

[0059] Where, is the current time step t The latent variables, is the semantic embedding vector, is the multimodal conditional fusion feature, is the attention query matrix, is the bond matrix K The transpose of is the dimension of a single attention head in the attention mechanism, is the value vector of the attention mechanism, is a learnable parameter matrix used to map latent variables and semantic embedding vectors to the attention query matrix Q, is used to map latent variables and semantic embedding vectors into key matrices K The learnable parameter matrix of Attention-based fusion features , noise prediction is performed through the frozen main branch and conditional branch of the multi-channel control module, and its expression is as follows:

[0060] Where, represents the total prediction noise, represents the prediction noise of the frozen main branch, is a scaling factor that controls the strength of the condition, represents the prediction noise of conditional branches; The mean prediction of the inverse process is performed by the noise predictor drive:

[0061] Based on the total prediction noise, construct the latent variable The ordinary differential equation for the inverse evolution of is expressed as follows:

[0062] Where, is the rate of change function of inverse denoising, is the current time step t The noise scheduling parameters, is the current time step t Diffusion parameters; According to the latent variables The ordinary differential equation of the inverse evolution of the forward diffusion of the final noisy latent variable At the beginning, T is the maximum time step, and the sampler is used to iteratively solve the differential relationship of the ordinary differential equation. When the time step regresses to the initial stage, the denoised latent variable is generated. .

[0063] Optionally, the edge and morphology control features include Canny edge features, probabilistic edge features, straight line morphology features, and line draft morphology features; The multi-channel control module extracts edge and morphological control features of the input sample image, including: The Canny operator is used to perform Gaussian filtering, Sobel operator gradient calculation, non-maximum suppression and double threshold processing on the input sample image in sequence to generate Canny edge features. ;in, and are the height and width of the feature respectively. The U-Net conditional branch learns the edge-texture mapping relationship through the convolution layer, forcing the generated image to be Structural alignment.

[0064] Extract the probabilistic edge features of the input sample image based on PIDI-Net, and generate probabilistic edge features based on the probabilistic edge features , preserve fuzzy boundary information and enhance texture continuity through channel weighting.

[0065] Perform straight line structure detection on the input sample image, parametrically extract the endpoint coordinates of the straight line, and generate straight line morphological features through spatial attention guidance. ;in, is the number of detected lines; Extract multi-scale line features of the input sample image based on the HED (Hierarchical Edge Detection) network, and generate line morphological features based on the multi-scale line features , retaining hair-level details and injecting style information through AdalN.

[0066] Optionally, the spatial perception and semantic control features include depth information features and semantic segmentation features; The multi-channel control module extracts spatial perception and semantic control features of the input sample image, including: The MiDaS network is used to predict the monocular depth of the input sample image and generate depth information features based on the monocular depth. ;in, and are the height and width of the feature, respectively. The values ​​encode the spatial hierarchy of the scene, and the depth map generates resolution through spatial attention adjustment.

[0067] Based on the Deeplabv3+ model, perform semantic segmentation on the input sample image according to the ADE20K protocol to generate semantic segmentation features ;in, is the number of object categories, each channel corresponds to the segmentation mark of a type of rural landscape elements, and the segmentation map is locally generated through the weighted control of the category channel.

[0068] Optionally, the style and content editing features include style vector features, low-rank adaptation matrix features, and editing mask features; The multi-channel control module extracts the style and content editing features of the input sample image and the design requirement text, including: Use the IP-Adapter image prompt adapter to encode the input sample image and obtain the style vector feature ;in, d The dimension of the feature is injected into the generation process through a lightweight adapter, and the broad modal attention fuses style and text Based on the design requirement text, a low-rank adaptation matrix feature is constructed through the T2I-Adapter text-to-image adapter. ; Construct a low-rank adaptation matrix Fine-tuning text-image alignment, and text-conditional dynamic enhancement.

[0069] Generate edit mask features based on design requirement text ; The elements in the edit mask feature use binary values ​​to identify the areas in the input sample image that need to be redrawn.

[0070] Optionally, the special function features include graffiti constraint features and local redraw mask features; The multi-channel control module extracts special functional features of the input sample image and design requirement text, including: Extract the user's hand-drawn rural landscape graffiti input from the design requirement text and generate graffiti features As a loose constraint, allowing AI to freely fill in details, graffiti enhances the creative space through sparse attention, where and are the height and width of the feature respectively, and 3 is the number of channels; Construct local redrawing mask features based on the rural landscape restoration requirements in the design requirements text ; Among them, the local redrawing mask feature Used to constrain the diffusion model to generate rural landscape details in a specified area. Based on the mask Specify the repair area, constrain the diffusion model to keep the original image in a certain area, and use latent space mixing and gradient guidance.

[0071] In one example, if Figure 4 As shown, Figure 4 Among them, (a) is the original sample image I of rural landscape provided by the embodiment of the present invention, which presents the original image of rural landscape and serves as the basic reference for the generation task; (b) is the Canny edge feature map provided by the embodiment of the present invention, which extracts the outline of the building and the environment through Canny edge detection, providing a basis for constraining the structural boundary of the model; (c) is the straight line morphological feature map provided by the embodiment of the present invention, which uses MLSD multi-level straight line detection to identify linear features such as building beams and columns, roads, etc., and strengthen geometric structure control; (d) is the line draft morphological feature map provided by the embodiment of the present invention, which generates Linear fine line drafts, refines the detailed lines of building doors, windows, decorations, etc., and improves the learning accuracy of the model for the structure; (e) is the probability provided by the embodiment of the present invention (f) The depth information feature map provided by the embodiment of the present invention constructs a depth perception map to distinguish the spatial layers of buildings, vegetation, and distant views, and assists the model in shaping the three-dimensional sense of depth; (g) The semantic segmentation feature map provided by the embodiment of the present invention completes semantic segmentation, annotates functional areas such as buildings, vegetation, and roads with colors, and clarifies the logic of element distribution; (h) The graffiti constraint feature map provided by the embodiment of the present invention displays the Scribble graffiti control results, converts the user's hand-drawn constraints into guidance information that can be parsed by the model, and realizes the customization of creative details. The above sub-graphs jointly construct a multimodal control input system to support the ControlNet algorithm in accurately regulating the structure, layer, region, and creativity in the process of generating rural landscapes.

[0072] Optionally, the basic diffusion model architecture is trained using the rural landscape design field-specific training set, including: In all cross-attention layers of the noise prediction module of the basic diffusion model, a low-rank adapter based on LoRA is injected. The first stage of training is performed by constructing the first objective function including CLIP semantic alignment loss and LPIPS perceptual loss to obtain a preliminary training model. Obtain the binary evaluation results and the number of edits for the generated rural landscape design images; constructing a reward function based on the binary evaluation result and the number of edits; Based on the reward function, using a policy optimization algorithm, constructing a second objective function including an importance sampling ratio and an advantage function of the generated action; The preliminary training model is trained in the second stage based on the second objective function to obtain a trained rural landscape design image generation model.

[0073] It should be noted that in all cross-attention layers of the noise prediction module of the basic diffusion model, a low-rank adapter based on LoRA is injected, and the first objective function containing CLIP semantic alignment loss and LPIPS perceptual loss is constructed. Carry out the first stage of training to obtain the preliminary training model; the first objective function The expression is as follows:

[0074] Where, and are weight parameters, Design images of rural landscapes generated by the model, For the corresponding design text, Sample images from a dedicated training set for rural landscape design. Used to measure the semantic matching between the generated graph and the text, Image features extracted by the VGG network; Table 2. First stage training parameters

[0075] In this embodiment, during the model training phase, a gradient descent optimization framework is used to drive iterative parameter updates. Relying on the backpropagation algorithm, the gradient of the loss function (such as cross-entropy loss, mean squared error, etc.) with respect to the network weights is accurately derived, enabling dynamic parameter adjustment. As shown in Table 2, to balance training efficiency and convergence stability, 50 training epochs are set, and gradient updates are accelerated through mini-batch data input. Cosine annealing, dynamic decay, and learning rate scheduling mechanisms are also introduced. A high learning rate is used to rapidly explore the parameter space in the early stages, and the learning rate is gradually reduced in the later stages to achieve fine convergence. During training, the loss function (Loss) value continues to decrease and stabilize to below 0.1, and the validation set loss curve (validation loss) does not increase significantly, effectively avoiding the risk of overfitting and ensuring the model's ability to fit the training data.

[0076] like Figure 5 As shown, Figure 5 The training loss curve provided by the embodiment of the present invention, wherein the horizontal axis is the training round epoch, and the vertical axis is the loss value loss. Figure 5 As shown in the figure, in the initial stage (epoch0-10), the loss value oscillates and decreases rapidly due to the randomness of the small-batch gradient update, reflecting that the model quickly explores the parameter space and captures the basic pattern of the data with a high learning rate in the early stage of learning rate scheduling; in the mid-term stage (epoch10-35), the loss fluctuation amplitude increases but the mean continues to decline, which is due to the local perturbation of the gradient direction / amplitude caused by the difference in the distribution of small-batch samples. At the same time, the dynamic adjustment of the learning rate prompts the model to jump out of the local optimum and continuously optimize the feature fitting ability; in the later stage (epoch35-50), the fluctuation amplitude gradually narrows, and the loss stably converges to the range below 0.1, indicating that the model parameter update tends to be stable and the fit with the training set data reaches a high level.

[0077] In order to further evaluate the generalization performance of the model and the system, this embodiment uses X / Y / visualization charts to carry out multi-dimensional testing, such as Figure 6 As shown, Figure 6 This is a schematic diagram of an X / Y visualization chart provided by an embodiment of the present invention, where the X-axis corresponds to the model at different training rounds, and the Y-axis corresponds to the model's different weight parameters. Based on the test results, the study used GridSearch to determine the optimal weight combination, and combined it with an early stopping mechanism to select the model round with the best performance on the validation set (e.g., round 35) as the final training result.

[0078] From the perspective of engineering deployment, the feasibility of implementation is evaluated by quantifying inference latency and memory footprint, and model pruning is introduced to compress redundant parameters, and quantization is introduced to reduce the computational accuracy requirements. While ensuring the accuracy of pose estimation, computing efficiency and hardware adaptability are optimized, ultimately achieving a synergy between high precision and engineering practicality.

[0079] In the second phase, we emphasize fine-tuning and alignment optimization. Based on the preliminary training model, we build a reward function that integrates manual scoring, semantic matching, and editing costs. , and the second objective function constructed by using the proximal strategy optimization algorithm The preliminary training model is trained in the second stage to obtain a trained rural landscape design image generation model; Among them, the reward function The expression is as follows:

[0080] Where, To generate binary evaluation results for the rural landscape design images, =1 means passed, =0 means reject, Indicates the number of times the rural landscape design image needs to be manually edited, which follows an exponential decay distribution; Second objective function The expression is as follows:

[0081] Where, To find the mathematical expectation of the data at time step t during the training process, is the importance sampling ratio, Indicates the current policy Generating status in the countryside Take action The probability of Indicates the old policy before the update Generating status in the countryside Take action probability; is the advantage function, through GAE ( ) calculation, used to measure the action of generating rural landscape design images a The value gain compared to the average action; Indicates that Clip to interval [ ], clipping threshold .

[0082] The second phase significantly increased the adoption rate of designs and significantly improved the accuracy of generating relevant professional terms such as "sloping roof" and "Guangfu style." After evaluation, unsuccessful models were merged or retrained. By optimizing the accuracy of the labeling system and iterating hyperparameter combination strategies, the team ultimately achieved precise expression of professional design language.

[0083] Example 3 like Figure 7 As shown, this embodiment proposes a rural landscape design image generation system based on knowledge graph and large model, which is applied to the rural landscape design image generation method based on knowledge graph and large model as described in the above embodiment, including: a first construction module 100, a second construction module 200, a labeling module 300, a third construction module 400 and a generation module 500.

[0084] Among them, the first construction module 100 is used to construct a knowledge graph covering the core elements and hierarchical classification of rural landscape design; the second construction module 200 is used to construct a multi-level labeling system based on the knowledge graph; the labeling module 300 is used to obtain pre-processed rural landscape sample images, label the rural landscape sample images based on the multi-level labeling system, and obtain a special training set for the field of rural landscape design; the third construction module 400 is used to construct a basic diffusion large model architecture, and use the special training set for the field of rural landscape design to train the basic diffusion large model architecture to obtain a trained rural landscape design image generation model; the generation module 500 uses the trained rural landscape design image generation model to generate a rural landscape design image.

[0085] It should be noted that the above explanation of the embodiment of the rural landscape design image generation method based on knowledge graph and large model is also applicable to the rural landscape design image generation system based on knowledge graph and large model of this embodiment, and will not be repeated here.

[0086] This embodiment focuses on the high-frequency and typical business needs in rural planning and design. Using the rural landscape design image generation system based on knowledge graphs and large models described in this embodiment as the core technology foundation, we independently developed the AIGC creative design platform, and simultaneously carried out the construction of a professional application scenario system and the exploration of AI-assisted design workflows: On the one hand, innovative modules for specialized application scenarios such as "entrances, building facade renovations, beautiful main streets, road environments, small parks and squares, and bird's-eye views" are constructed to form a multi-faceted design image generation system covering rural landscapes from node details to overall layouts. The rural landscape knowledge graph constructed by the first building block provides domain knowledge support for the refined semantic definition of each scene; the multi-level labeling system established by the second building block ensures the accuracy of structured modeling of key elements of scenes such as "road environments, small parks and squares"; combined with the ControlNet multimodal control mechanism, the model is infused with domain adaptation capabilities through the collaborative creation of a rural landscape training set by the annotation module and the third building block, achieving semantic alignment, morphological control, and landscape adaptation during the image generation process. Typical applications include collaborative design of the entrance signage system and landscape, style normalization and intelligent material replacement of building facades, and optimization of the pedestrian environment and creation of a commercial atmosphere on beautiful main streets, precisely matching the needs of rural design businesses.

[0087] On the other hand, we are exploring iterative human-machine collaborative AI-assisted design workflows, with the core process being implemented through the system's functional modules: After designers submit their design requirements through natural language descriptions or sketch input, the system invokes a multi-task learning framework (based on the semantic parsing capabilities of knowledge graphs) to simultaneously analyze functional requirements, style preferences, and technical constraints. Designers use the platform's interactive annotation tools to perform feedback operations such as geometric corrections, material replacements, and style adjustments on the initial output of the generated module. After 3-5 rounds of human-machine iterative optimization, the final output is a solution that meets professional design standards and possesses regional characteristics. Field data confirms that this workflow, supported by the knowledge graph and large-scale model system, has increased the efficiency of rural design by over 30%, achieving a coordinated advancement in the depth of professional scenario coverage and the upgrading of design effectiveness.

[0088] To effectively evaluate the quality of rural landscape image generation models, this implementation uses objective evaluation metrics to conduct mathematical and statistical analysis of generated images, quantifying their technical quality characteristics. Simultaneously, the model leverages the subjective perceptions of professional designers to conduct an industry standard compatibility assessment, accurately measuring the degree to which the generated images align with actual production requirements. This collaborative evaluation approach encompasses both technical performance verification and business requirement compatibility verification, forming a comprehensive and reliable closed-loop quality verification process.

[0089] In the objective evaluation dimension, Fréchet Inception Distance (FID) was selected as the core quantitative indicator. FID is a classic indicator widely used in the performance evaluation of generative models. Its technical principle is as follows: the feature distributions of the generated image set and the real image set are mapped to a high-dimensional space through a pre-trained Inception network, and then fitted into Gaussian distributions respectively; by calculating the Fréchet distance between the two Gaussian distributions, the similarity of the semantic features of the two image sets is quantified. Specifically, the smaller the FID value, the closer the high-level semantic features of the generated image are to the real image, and the better the diversity and realism of the generated results; when the FID value is 0, it can be determined that the feature distributions of the two image sets are completely consistent. Compared with traditional evaluation indicators, FID has two core advantages: first, it is based on the high-level semantic features extracted by the Inception network, which is more closely related to human visual perception and more in line with actual design needs; second, it quantifies the differences by statistical distribution distance, and the calculation results are less affected by random noise and have better stability. The mathematical expression for calculating FID is:

[0090] in, : The mean of the features extracted from the real image; : the mean of the features extracted from the generated image; : Covariance matrix of features extracted from real images; : Covariance matrix of features extracted from real images; : trace of the matrix.

[0091] In this embodiment, a multi-dimensional control optimization strategy is proposed, such as Figure 8 As shown, Figure 8 The flow chart of the multi-dimensional control optimization strategy provided by the embodiment of the present invention is as follows in the rural image generation optimization framework based on the ControlNet module: Figure 8The process logic shown uses the "current state map" as input and relies on a multimodal conditional control mechanism to perform high-dimensional feature encoding and decoupled representation learning on the input image. The pre-trained StableDiffusion model is selected as the basic generation architecture, integrating four parallel ControlNet branches: line morphology (LineArt), scale perspective (DepthMap), linear contour (MLSD), and material semantic segmentation (Segmentation). These branches process the structured control signals presented by the control module. The line morphology branch matches the "line morphology" subgraph and extracts sketch features using a lightweight residual convolutional network to maintain the original scene composition logic. The scale perspective branch matches the "scale perspective" subgraph and uses a monocular depth estimation network to establish 3D spatial constraints to ensure geometric consistency between the generated result and the original image. The linear contour branch matches the "linear contour" subgraph and uses a Hough transform baseline segment detector to enhance the rigid features of the building structure to anchor key linear elements. The material semantic segmentation branch matches the "material segmentation" subgraph and relies on a CLIP-driven semantic parsing module to accurately map material properties such as vegetation, buildings, and sky.

[0092] Entering the parameter optimization stage, an adversarial training strategy is adopted in conjunction with a progressive learning rate decay mechanism to implement end-to-end fine-tuning of the ControlNet weight matrix; dynamic gradient clipping and mixed precision training are simultaneously introduced to stabilize the training process, and a multi-scale loss function is constructed through KL divergence and perceptual loss to balance the fidelity and creativity of the generated images.

[0093] During the inference phase, a parameterized control intensity adjustment mechanism is designed to dynamically adjust the conditional injection weights based on a cosine annealing strategy. This ensures that line control maintains a high intensity (0.8-1.0) in the early stages of generation to ensure structural stability, and decays to 0.4-0.6 in the later stages of generation to release the potential for detailed creation of the diffusion model. For depth maps and material segmentation, a hierarchical attention mechanism is used to perform spatially adaptive control intensity allocation, imposing strong constraints (≥0.6) on building outlines and loosening constraints (≤0.4) on natural elements such as vegetation and sky to enhance generation diversity.

[0094] Experimental results demonstrate that this solution significantly improves the FID metric compared to the baseline model. User studies have also verified that the generated images demonstrate outstanding advantages in structural rationality and aesthetic expression. Through the full-link mapping process from "existing map → control module → generated drawings," the generated results closely align with the existing map in terms of form, scale, and elements, achieving a dialectical unity of controllable generation and creative divergence. This provides a smart generation solution for intelligent rural landscape design that combines engineering robustness with artistic expression.

[0095] like Figure 9 As shown, Figure 9Among them, (a) is the original sample image II of the rural landscape provided by the embodiment of the present invention, which presents the real status of rural buildings and environment, clearly displays the original status of building materials, spatial layout and surrounding landscape, and serves as a reference for model training; (b) is the generated image of the StableDiffusion native basic model provided by the embodiment of the present invention, which is the generation result of the StableDiffusion native basic model. Compared with the current status map, it is not much different and the image quality is poor; (c) is the generated image of the rural landscape design image generation model trained by the present invention provided by the embodiment of the present invention, which is the output of the model after training by the present invention. Relying on the rural landscape knowledge graph and the special training set, while retaining the local texture of the building, it systematically optimizes the landscape hierarchy, strengthens the style unity, and conforms to the rural landscape. It realizes the modification of the building facade and roads in line with professional needs, and realizes the integration of local characteristics and design aesthetics; (d) is the generated image of the general large model provided by the embodiment of the present invention. Although it enhances the sense of picture atmosphere through rendering, it has problems such as distortion of rural building details and vague expression of regional characteristics. Its style does not conform to the needs of rural planning and design.

[0096] Based on this modular and structured application scenario system, the generated rural landscape design elements are more complete and comprehensive, encompassing the rural natural environment, architecture, landscape, facilities, native plants, and more. The matching degree of local rural characteristics reaches 90% (compared to 60% for the general large-scale model), the proportional error of landscape elements is ≤5%, and the FID (Frechet Inception Distance) score in the "style generation" task is 12.7, significantly better than the general model (FID=29.4). At the same time, in rural planning and design projects, it can assist in solving the image expression needs of early communication, mid-term concept generation, and later project implementation. It can shorten design time by over 30% and reduce the per capita mapping and design workload by approximately 20%, effectively contributing to the improvement of rural living environments.

[0097] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples without contradiction.

[0098] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "N" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0099] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing a custom logical function or step of a process, and the scope of the preferred embodiments of the invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of the invention pertain.

[0100] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiment, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logical functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array, a field programmable gate array, etc.

[0101] Those skilled in the art will understand that all or part of the steps in the method for implementing the above-mentioned embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0102] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. A method for generating rural landscape design images based on knowledge graph and large model, characterized in that: include: Construct a knowledge map covering the core elements and hierarchical classification of rural landscape design; Constructing a multi-level labeling system based on the knowledge graph; Obtaining preprocessed rural landscape sample images, and labeling the rural landscape sample images based on the multi-level labeling system to obtain a special training set for the rural landscape design field; Constructing a basic diffusion model framework, and training the basic diffusion model framework using the rural landscape design field-specific training set to obtain a trained rural landscape design image generation model; Use the trained rural landscape design image generation model to generate rural landscape design images.

2. The rural landscape design image generation method based on knowledge graph and large model according to claim 1 is characterized in that: Construct a knowledge graph covering the core elements and hierarchical classification of rural landscape design, including: Perform hierarchical structural division on the pre-screened core entities of rural landscape to obtain hierarchical classification results; According to the hierarchical classification results, a structured triple containing the head entity, relationship and tail entity is constructed; Map the head entity, relationship, and tail entity in the structured triple to a vector space, and optimize the vector representation of the head entity vector and the relationship vector so that the spatial distance between the sum of the head entity vector and the relationship vector and the tail entity vector meets a preset distance threshold condition; Based on the vectorized triples that meet the preset distance threshold condition, a rural landscape design knowledge graph is constructed.

3. The rural landscape design image generation method based on knowledge graph and large model according to claim 1 is characterized in that: The basic diffusion model architecture includes a variational autoencoder module, a noise prediction module, a text encoding module and a multi-channel control module; the variational autoencoder module includes an encoding unit and a decoding unit; The encoding unit is used to map the input sample image to a low-dimensional latent space to obtain a latent variable; The text encoding module is used to encode the input design requirement text to obtain a semantic embedding vector; The multi-channel control module is used to extract multimodal features of the input sample image and the design requirement text, and perform feature fusion processing on the multimodal features to generate multimodal conditional fusion features; the multimodal features include edge and morphological control features, spatial perception and semantic control features, and style and content editing features; The noise prediction module is used to fuse the latent variable, the semantic embedding vector and the multimodal conditional fusion feature, and perform noise prediction and denoising operations on the fusion result to obtain a denoised latent variable; The decoding unit is used to reconstruct the denoised latent variables into a rural landscape design image.

4. The method for generating rural landscape design images based on knowledge graph and large model according to claim 3 is characterized in that: The noise prediction module fuses the latent variable, the semantic embedding vector and the multimodal conditional fusion feature, and performs noise prediction and denoising operations on the fusion result to obtain the denoised latent variable, including: Perform cross attention fusion on the latent variable, the semantic embedding vector and the multimodal conditional fusion feature to generate an attention fusion feature , whose expression is as follows: Where, is the current time step t The latent variables, is the semantic embedding vector, is the multimodal conditional fusion feature, is the attention query matrix, is the bond matrix K The transpose of is the dimension of a single attention head in the attention mechanism, is the value vector of the attention mechanism, is a learnable parameter matrix used to map latent variables and semantic embedding vectors to the attention query matrix Q, is used to map latent variables and semantic embedding vectors into key matrices K The learnable parameter matrix of Attention-based fusion features , noise prediction is performed through the frozen main branch and conditional branch of the multi-channel control module, and its expression is as follows: Where, represents the total prediction noise, represents the prediction noise of the frozen main branch, is a scaling factor that controls the strength of the condition, represents the prediction noise of conditional branches; Based on the total prediction noise, construct the latent variable The ordinary differential equation for the inverse evolution of is expressed as follows: Where, is the rate of change function of inverse denoising, is the current time step t The noise scheduling parameters, is the current time step t Diffusion parameters; According to the latent variables The ordinary differential equation of the inverse evolution of the forward diffusion of the final noisy latent variable At the beginning, T is the maximum time step, and the sampler is used to iteratively solve the differential relationship of the ordinary differential equation. When the time step regresses to the initial stage, the denoised latent variable is generated. .

5. The method for generating rural landscape design images based on knowledge graph and large model according to claim 3 or 4, characterized in that: The edge and morphological control features include Canny edge features, probabilistic edge features, straight line morphological features and line draft morphological features; The multi-channel control module extracts edge and morphological control features of the input sample image, including: The Canny operator is used to perform Gaussian filtering, Sobel operator gradient calculation, non-maximum suppression and double threshold processing on the input sample image in sequence to generate Canny edge features. ;in, and are the height and width of the feature, respectively; Extract the probabilistic edge features of the input sample image based on PIDI-Net, and generate probabilistic edge features based on the probabilistic edge features ; Perform straight line structure detection on the input sample image, parametrically extract the endpoint coordinates of the straight line, and generate straight line morphological features ;in, is the number of detected lines; Extract multi-scale line draft features of the input sample image based on the HED network, and generate line draft morphological features based on the multi-scale line draft features .

6. The method for generating rural landscape design images based on knowledge graph and large model according to claim 3 or 4, characterized in that: The spatial perception and semantic control features include depth information features and semantic segmentation features; The multi-channel control module extracts spatial perception and semantic control features of the input sample image, including: The MiDaS network is used to predict the monocular depth of the input sample image and generate depth information features based on the monocular depth. ;in, and are the height and width of the feature, respectively; Based on the Deeplabv3+ model, perform semantic segmentation on the input sample image according to the ADE20K protocol to generate semantic segmentation features ;in, is the number of object categories, and each channel corresponds to the segmentation mark of a type of rural landscape elements.

7. The method for generating rural landscape design images based on knowledge graph and large model according to claim 3 or 4, characterized in that: The style and content editing features include style vector features, low-rank adaptation matrix features and editing mask features; The multi-channel control module extracts the style and content editing features of the input sample image and the design requirement text, including: Use the IP-Adapter image prompt adapter to encode the input sample image and obtain the style vector feature ;in, d is the dimension of the feature; Based on the design requirement text, a low-rank adaptation matrix feature is constructed through the T2I-Adapter text-to-image adapter. ; Generate edit mask features based on design requirement text ; The elements in the edit mask feature use binary values ​​to identify the areas in the input sample image that need to be redrawn.

8. The rural landscape design image generation method based on knowledge graph and large model according to claim 1 is characterized in that: The basic diffusion model architecture is trained using the rural landscape design field-specific training set, including: In all cross-attention layers of the noise prediction module of the basic diffusion model, a low-rank adapter based on LoRA is injected. The first stage of training is performed by constructing the first objective function including CLIP semantic alignment loss and LPIPS perceptual loss to obtain a preliminary training model. Obtain the binary evaluation results and the number of edits for the generated rural landscape design images; constructing a reward function based on the binary evaluation result and the number of edits; Based on the reward function, using a policy optimization algorithm, constructing a second objective function including an importance sampling ratio and an advantage function of the generated action; The preliminary training model is trained in the second stage based on the second objective function to obtain a trained rural landscape design image generation model.

9. The method for generating rural landscape design images based on knowledge graph and large model according to claim 8, characterized in that: The first objective function The expression is as follows: Where, and are weight parameters, Design images of rural landscapes generated by the model, For the corresponding design text, Sample images from a dedicated training set for rural landscape design. Used to measure the semantic matching between the generated graph and the text, Image features extracted by the VGG network; Reward Function The expression is as follows: Where, To generate binary evaluation results for the rural landscape design images, =1 means passed, =0 means reject, Indicates the number of times the rural landscape design images need to be manually edited; Second objective function The expression is as follows: in, Where, To find the mathematical expectation of the data at time step t during the training process, is the importance sampling ratio, Indicates the current policy Generating status in the countryside Take action The probability of Indicates the old policy before the update Generating status in the countryside Take action The probability of is the advantage function used to measure the action of generating rural landscape design images a Compared to the value gain of the average action, Indicates that Clip to interval [ ].

10. A rural landscape design image generation system based on knowledge graph and large model, characterized by: include: The first building block is used to construct a knowledge graph covering the core elements and hierarchical classification of rural landscape design; A second building module is used to build a multi-level labeling system based on the knowledge graph; a labeling module, configured to obtain pre-processed rural landscape sample images, label the rural landscape sample images based on the multi-level labeling system, and obtain a special training set for the rural landscape design field; The third construction module is used to construct a basic diffusion model framework, and train the basic diffusion model framework using the dedicated training set in the field of rural landscape design to obtain a trained rural landscape design image generation model; The generation module uses the trained rural landscape design image generation model to generate rural landscape design images.

Citation Information

Patent Citations

  • Knowledge graph-based automatic identification and migration method for style and appearance style of Mongolian building

    CN118780171A

  • Multimodal semantic analysis and image retrieval

    US20240354336A1

Cited By

  • Public art design scheme generation method and system based on knowledge graph

    CN121435364A

  • Craniofacial generation system based on edge guidance and semantic regulation and control

    CN121527251A

  • Building image in-place continuation design generation method and building image in-place continuation design generation system

    CN122310621A