Diffusion model-based child-friendly guiding community public space updating intention graph model training method

By sorting out the three-level tag set and training the LoRA model, the problem that the general big model cannot understand planning terms and generate design elements in community updates is solved, and high-quality community update intention map generation is achieved.

CN120032020APending Publication Date: 2025-05-23TONGJI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510156182.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In the application scenario of community updates, the general model cannot accurately understand the planning professional terms and the design elements of the community public space generated, resulting in poor results in the generated images.

Method used

By sorting out the three-level label set of child-friendly guided spaces, and collecting sample pictures in a directional way for labeling, generating a sample library of child-friendly guided spaces, and training the LoRA model, ensuring that the model can understand the planning terminology and generate the correct design elements.

Benefits of technology

The model's understanding of important terms in child-friendly oriented community updates is realized, and the common design elements and their spatial relationships are accurately generated, which improves the quality of the community update intention map.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032020A_ABST
    Figure CN120032020A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence content generation, and provides a child-friendly guiding community public space updating intention graph model training method based on a diffusion model, and the method comprises the steps: 1, combing a child-friendly guiding space label set, the child-friendly guidance space label set is a three-level label set of child-friendly guidance-multiple function types-multiple element details; 2, directionally collecting various sample pictures on the basis of the child-friendly guide space label set, performing three-level label labeling on the preprocessed pictures, and performing label frequency inspection and directional supplement on the labeled picture set until a final child-friendly guide space sample library is obtained; and step 3, training the LoRA model by using the child-friendly guide space sample library to obtain a reference LoRA model. The model trained based on the method can simply and quickly generate a community updating intention graph meeting requirements, facilitates communication between designers and residents, and facilitates public participated design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to AIGC (Artificial Intelligence Generated Content) technology, and specifically relates to a child-friendly oriented community public space renewal intention map model training method based on a diffusion model, which is applied to the generation of design intention maps for community renewal. Background Art

[0002] Community renewal is a key strategy to stimulate urban vitality. Under the concept of "people's city", public participatory design of community renewal needs to be further strengthened.

[0003] AIGC is a way for creators to guide artificial intelligence models to generate various contents through instructions (prompt words). The advancement of AIGC technology and the emergence of a general image generation model have enabled rapid image generation, providing the possibility for rapid communication between the public and planners in assisting community renewal. The image generation model has learned a large amount of image-text data from the Internet and has general image generation knowledge. In addition, ControlNet can control the neural network structure of the diffusion model by adding additional conditions. In the process of text-to-image generation, conditional inputs such as graffiti, edge mapping, segmentation mapping, pose key points, etc. are used to make the generated image closer to the input image, and a controllable image generation effect is obtained. This provides a prerequisite for AIGC to generate controllable images, and the use of diffusion models to generate images has been applied in many fields. However, in the field of urban planning, the general large model lacks an in-depth understanding of specific tasks and planning industry knowledge. When using prompt words for image generation in the application scenario of community renewal, the following problems will arise:

[0004] (1) The model lacks understanding of planning terminology (conceptual vocabulary). For example, when the prompt word "child-friendly" is input, some scenes related to children will appear in the generated pictures, but they do not match the community public space. At the same time, the generated scenes will also be distorted, mixed, and unable to integrate into the original picture.

[0005] (2) The model is not very effective in generating design elements (objective things) of community public spaces. For example, when the prompt word “children’s climbing slope” is input, the model cannot generate climbing slopes in children’s play facilities, but instead generates elements such as stairs and climbing frames. When the prompt word “shallow pool” is input, the model cannot accurately generate shallow pools in children’s activity areas. When the prompt word “popular science garden” is input, only green lawns can be generated in most cases.

[0006] (3) Other image corruption issues may occur, such as the scale and perspective relationship of the generated elements are incorrect, the style is not realistic, etc. And the generated elements cannot be guaranteed to be completely correct.

[0007] In summary, the general large model cannot be directly used in participatory design of community renewal.

[0008] In order to enable the image model to master the professional knowledge of community renewal planning, fine-tuning of the large model is required. Among them, LoRA, as a low-parameter fine-tuning method, is fast and efficient, so it is a widely used fine-tuning method. At present, a variety of LoRA models have appeared in various fields, and existing patents also have a variety of LoRA model training methods for multiple fields, but the LoRA model training methods in other fields are not completely consistent with the LoRA model training methods in the field of community renewal. In the field of child-friendly community renewal, there is still a lack of planning-specialized LoRA models and training methods. When using existing image models and LoRA models, there will be problems such as misunderstanding of planning professional terms and inaccurate generation of design elements, making it difficult to generate high-quality community renewal intention maps.

[0009] The Chinese invention patent application with publication number CN 118570567 A discloses a method and system for generating a planning intention map based on an image generation model. It includes: obtaining feature labels and sample images, preprocessing the sample images to form a sample library; adjusting and acceptance testing the preset basic model through the diffusion model; adjusting the standard CheckPoint model through the LoRA method to obtain the LoRA model; generating images through the image generation interface. This method summarizes the entire process of generating a planning intention map, but in the labeling stage, it lacks the combing of professional knowledge and the construction of label sets; in terms of sample image screening, it lacks accurate description of screening conditions. Therefore, the use effect in specific planning fields is limited, such as the field of child-friendly community public space renewal, and cannot meet actual needs. Summary of the invention

[0010] In response to the above problems, the present invention provides a method for training a child-friendly oriented community public space renewal intention map model based on a diffusion model. The model trained by the method of the present invention can understand some important terms in child-friendly oriented community renewal, facilitate the generation of intention maps with corresponding effects, and can accurately generate commonly used design elements in child-friendly oriented community renewal; it can accurately generate the correct spatial relationship between design elements. Therefore, the model trained by the method of the present invention can simply and quickly generate a community renewal intention map that meets the requirements, and facilitate communication between designers and residents, which is convenient for public participatory design.

[0011] The technical solution of this application is as follows:

[0012] A method for training a child-friendly community public space renewal intention graph model based on a diffusion model includes the following steps:

[0013] Step 1: sorting out a child-friendly guidance space label set, wherein the child-friendly guidance space label set is a three-level label set of "child-friendly guidance - multiple functional types - multiple element details";

[0014] Step 2: Based on the child-friendly oriented space label set, various sample images are collected in a targeted manner, the pre-processed images are labeled with three-level labels, and the labeled image set is tested for label frequency and supplemented in a targeted manner until the final child-friendly oriented space sample library is obtained;

[0015] Step 3: Use the child-friendly oriented space sample library to train the LoRA model and obtain the baseline LoRA model.

[0016] Further, the three-level tag set of "update guide - function type - element details", wherein the first level is the update guide tag, the second level is the function type tag, and the third level is the element detail tag;

[0017] The specific steps for label set sorting are:

[0018] 1.1 Based on planning expertise, the first-level update orientation label "child-friendly orientation" is refined and split into multiple functional types according to the activity characteristics of the venue, and the second-level functional type label is constructed;

[0019] 1.2 Based on the experience of space design, sort out the key elements needed for the training of the child-friendly community public space renewal intention map model, and construct the third-level element detail labels;

[0020] 1.3 Establish the correspondence between the function type label and the element detail label;

[0021] 1.4 Obtain a three-level label set of “child-friendly orientation – multiple feature types – multiple element details”.

[0022] Furthermore, the step 2 comprises:

[0023] 2.1 Based on the child-friendly oriented space label set constructed in step 1, extensively collect sample images according to multiple functional types and element detail requirements;

[0024] 2.2 Label the preprocessed sample images using the child-friendly oriented spatial label set, where:

[0025] The third-level element detail labels are selected based on the actual situation of the drawing;

[0026] Select at least one of the second-level functional type labels;

[0027] The first-level label "child-friendly orientation" is mandatory;

[0028] 2.3 Perform sample library index test: set the label frequency test index, if the frequency of each label is higher than the label frequency test index, the test is passed;

[0029] 2.4 If the test fails, collect additional sample images for the relevant labels and repeat steps 2.2, 2.3, and 2.4 until the test passes;

[0030] 2.5 Generate a sample library of child-friendly guided spaces.

[0031] Furthermore, the step 3 comprises:

[0032] 3.1 Select a suitable CheckPoint model, use the child-friendly guided space sample library obtained in step 2 as the training set, adjust the standard CheckPoint model through the LoRA method, and obtain several alternative LoRA models;

[0033] 3.2 For each candidate LoRA model, acceptance test and weight screen test are performed on the LoRA model through preset chart scripts;

[0034] 3.3 All LoRA models that pass the test can be used as benchmark LoRA models for the generation of child-friendly community public space renewal intention maps; if all alternative LoRA models fail the test, the parameters are adjusted and the model with the best test performance among the alternative LoRA models is further fine-tuned until it passes the test and the benchmark LoRA model is obtained.

[0035] Compared with the prior art, this application has the following advantages and beneficial effects:

[0036] (1) A three-level knowledge graph is used to sort out the child-friendly community update label set and correctly collect and label the corresponding images, so that the trained LoRA model can generate correct images of planning terms and detailed design elements in the label set.

[0037] (2) Based on the child-friendly oriented spatial label set, various sample images are collected in a targeted manner, and three-level labeling and indicator testing are performed to generate a sample library, so that the LoRA model trained based on the sample library can understand the correct element details, element spatial relationships, and image style.

[0038] (2) Using the trained child-friendly community public space renewal model, designers can quickly generate intention maps, which can be used to provide inspiration references, quickly exchange ideas, generate usable renderings, and discuss and feedback design plans with the public. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1The present invention is a schematic diagram of a training method and application flow of a child-friendly oriented community public space updating intention graph model based on a diffusion model.

[0040] Figure 2 The present invention is a flowchart of sorting out a child-friendly guide space tag set according to an embodiment of the present invention.

[0041] Figure 3 The figure is a schematic diagram of the process of obtaining a sample library of child-friendly guided spaces according to an embodiment of the present invention.

[0042] Figure 4 Schematic diagram of function type label in the embodiment.

[0043] Figure 5 It is a schematic diagram of labeling sample images in an embodiment. DETAILED DESCRIPTION

[0044] The technical solution provided by the present application will be further described below in conjunction with specific embodiments and accompanying drawings. The advantages and features of the present application will become more apparent with the following description.

[0045] A method for training a child-friendly community public space renewal intention graph model based on a diffusion model includes the following steps:

[0046] Step 1. Sort out the set of labels for child-friendly spaces

[0047] Based on planning expertise and spatial design experience, the conceptual vocabulary "child-friendly orientation" that is difficult for machines to understand and learn in model training is refined and split into a three-level label set of "update orientation-function type-element details". Among them, the first level is the update orientation label, which has 1; the second level is the function type label, which has 4; the third level is the element detail label, which has 31. The specific steps include:

[0048] 1.1 Based on planning expertise, interpret the first-level conceptual vocabulary (vocabulary that describes a comprehensive function rather than lists specific elements) in the community public space renewal scene, and refine the first-level renewal orientation label "child-friendly orientation" into multiple functional types according to the activity characteristics of the venue, and construct the second-level functional type label. There are 4 functional type labels in the embodiment, specifically sports type, leisure and recreation type, natural science type, and water-friendly activity type. 1.2 Based on spatial design experience, sort out and list the details of various elements involved in the child-friendly space scene; based on the understanding of the prompt words and machine learning characteristics of the existing general large model, screen out non-essential elements that can be automatically generated by the general large model, retain the key elements required for the training of the child-friendly orientation community public space renewal intention map model, and construct the third-level element detail label, including "special venue type", "special facility type", and "general facility type". Each category includes several element details. Among them, "special venues" refer to terrain or pavement related to children's activities, which are the main space for venue activities, including: children's colors, children's pavement, slow-moving pavement, plank roads, lawns, micro-topography, science gardens, sand pits, shallow pools, and runways; "special facilities" refer to devices or facilities related to children's activities, which are functional components located in the venue space, including: amusement facilities, slides, graffiti, art installations, safety protection, seats, tables and chairs, children's climbing slopes, swings, and dry fountains; "general facilities" refer to devices or facilities that are common in community public space renewal scenes and are not related to children's activities. They are important basic components in the venue space, including people, trees, grass, flowers, flower beds, paths, roofs, walls, steps, cars, and buildings.

[0049] There are 31 element detail tags in total in the embodiment.

[0050]

[0051]

[0052] 1.3 Establish the correspondence between the function type label and the element detail label, such as "exclusive", "optional", "not included", etc. "Exclusive" means that the element may only appear in the corresponding function type and is its characteristic identification element. If the element appears, it must belong to the corresponding function type; "optional" means that the element may appear in multiple corresponding function types; "not included" means that the element will not appear in the corresponding function type.

[0053]

[0054] 1.4 Obtain a three-level label set of “child-friendly orientation – multiple feature types – multiple element details”.

[0055]

[0056] Step 2. Obtain a sample library of child-friendly spaces

[0057] Based on the child-friendly oriented space label set constructed in step 1, various sample images are collected in a targeted manner, the pre-processed images are labeled with three-level labels, and the labeled image set is tested for label frequency and supplemented in a targeted manner until the final child-friendly oriented space sample library is obtained. The specific steps include:

[0058] 2.1 Based on the child-friendly oriented space label set, sample images are widely collected according to multiple functional types and element detail requirements.

[0059] 2.2 Label the sample images after preprocessing (such as removing text, adjusting angles, and cropping) using the child-friendly oriented spatial label set. Among them:

[0060] The third-level element detail labels are selected based on the actual situation of the drawing. Site elements account for 10%-50% of the screen area, and facility elements account for 3%-10% of the screen area. If the screen ratio is too low, the element cannot be effectively identified, resulting in invalid element labels; if the screen ratio is too high, the combination relationship between the element and the environment cannot be learned, resulting in distortion of the element scale in the generated image. The second-level functional type label is selected based on the corresponding relationship of "element details-functional type" in the label set. At least one is selected, and different functional type labels can be used in combination at the same time. If a "proprietary" relationship appears, the functional type corresponding to the element must be selected; if no "proprietary" relationship appears, the functional type with the most "optional" relationships is selected.

[0061] The first-level label "Child-friendly" is required for all sample images.

[0062] 2.3 Perform sample library index test. Set the label frequency test index to 10%, and summarize the number of times each label appears. If the frequency of a label is lower than 10%, the number of samples of the label is too small, and the LoRA model cannot effectively understand its content, which is considered to have failed the test. If the frequency of each label is higher than 10%, it passes the test.

[0063] 2.4 If the test fails, collect additional sample images for the relevant labels and repeat steps 2.2, 2.3, and 2.4 until the test passes.

[0064] 2.5 Generate a sample library of child-friendly guided spaces.

[0065] Step 3. LoRA model training, the specific steps include:

[0066] 3.1 Select a suitable CheckPoint model, such as SD1.5, SDXL, the urban design large model, etc. Use the sample library of child-friendly oriented spaces obtained in step 2 as the training set, and adjust the standard CheckPoint model through the LoRA method to obtain several alternative LoRA models.

[0067] 3.2 For each alternative LoRA model, conduct acceptance tests and weight screen tests through a preset chart script. If the following conditions are met simultaneously, the test is passed; otherwise, the test fails:

[0068] Condition (1): When the prompt contains the third-level element detail label, the graphic elements generated by loading this LoRA model conform to the third-level element detail label.

[0069] Condition (2): When the prompt contains the second-level function type label, the generated graphic by loading this LoRA model contains elements with the third-level element detail label that has a "proprietary" or "optional" relationship with it.

[0070] Condition (3): The detail attributes of the third-level element detail label elements generated by loading this LoRA model (such as the color and surface flatness of the sandpit, the marking direction of the runway, the ripple texture of the water surface, etc.) are consistent with the style of the sample library.

[0071] Condition (4): The spatial relevance between the graphic elements generated by loading this LoRA model conforms to the basic knowledge of planning and the principles of spatial perspective (such as children's play facilities are located on the ground with children's pavement, rather than in the road space).

[0072] 3.3 For all LoRA models that pass the test in 3.2, they can be used as the benchmark LoRA model for generating the updated intention map of the child-friendly oriented community public space. If all alternative LoRA models fail the test in 3.2, adjust parameters such as the learning rate, number of epochs, number of parallel processes, optimizer method, learning precision, etc., and further fine-tune the model with the best performance (the model that meets the most conditions) among the alternative LoRA models until the test is passed to obtain the benchmark LoRA model.

[0073] Furthermore, in the image generator loaded with the benchmark LoRA model, according to the child-friendly oriented space label set and the update requirements of the specific scenario, input the prompt and the base map of the scene to be updated with the redrawn area delimited, and perform spatial constraints through various Controlnet methods to generate the updated intention map.

[0074] The above description is only a description of the preferred embodiments of the present application, and is not intended to limit the scope of the present application. Any changes or modifications made by any person skilled in the art based on the above disclosed technical contents shall be deemed as equivalent effective embodiments and shall fall within the scope of protection of the technical solution of the present application.

Claims

1. A method for training a child-friendly community public space renewal intention graph model based on a diffusion model, characterized in that: The steps include: Step 1: sorting out a child-friendly guidance space label set, wherein the child-friendly guidance space label set is a three-level label set of "child-friendly guidance - multiple function types - multiple element details"; Step 2: Based on the child-friendly oriented space label set, various sample images are collected in a targeted manner, the pre-processed images are labeled with three-level labels, and the labeled image set is tested for label frequency and supplemented in a targeted manner until the final child-friendly oriented space sample library is obtained; Step 3: Use the child-friendly oriented space sample library to train the LoRA model and obtain the baseline LoRA model.

2. The method according to claim 1, characterized in that Step 1: The three-level tag set of "update guide - function type - element details", wherein the first level is the update guide tag, the second level is the function type tag, and the third level is the element detail tag; The specific steps of label set combing are: 1.1 Based on planning expertise, the first-level update orientation label "child-friendly orientation" is refined and split into multiple functional types according to the activity characteristics of the site, and the second-level functional type label is constructed; 1.2 Based on the experience of space design, the key elements needed for the training of the intention map model for the renewal of child-friendly community public spaces were sorted out, and the third-level element detail labels were constructed; 1.3 Establish the correspondence between the functional type label and the element detail label; 1.4 Obtain a three-level label set of "child-friendly orientation - multiple function types - multiple element details".

3. The method according to claim 2, characterized in that There is one first-level update guide label, which is "child-friendly guide"; There are four second-level functional type labels, namely sports type, leisure and recreation type, natural science type, and water-related activity type; There are 31 third-level element detail labels, including three categories: "Specialized venues", "Specialized facilities", and "General facilities". Each category contains several element details, including: "Specialized venues" refer to terrain or pavement related to children's activities, which are the main spaces for venue activities, including: children's colors, children's pavement, slow-moving pavement, plank roads, lawns, micro-topography, science gardens, sand pits, wading pools, and running tracks; "Specialized facilities" refer to devices or facilities related to children's activities, which are functional components located in the venue space, including: amusement facilities, slides, graffiti, art installations, safety protection, seats, tables and chairs, children's climbing slopes, swings, and dry fountains; "General facilities" refer to common devices or facilities in community public space renewal scenarios that are not related to children's activities. They are important basic components in the site space, including people, trees, grass, flowers, flower beds, paths, roofs, walls, steps, cars, and buildings.

4. The method according to claim 1, characterized in that The step 2 comprises: 2.1 Based on the child-friendly oriented space label set constructed in step 1, extensively collect sample images according to multiple functional types and element detail requirements; 2.2 Label the preprocessed sample images using the child-friendly oriented spatial label set: 2.3 Perform sample library index test: set the label frequency test index, if the frequency of each label is higher than the label frequency test index, the test is passed; 2.4 If the test fails, collect additional sample images for the relevant labels and repeat steps 2.2, 2.3, and 2.4 until the test passes; 2.5 Generate a sample library of child-friendly guided spaces.

5. The method according to claim 4, characterized in that In step 2.2, the preprocessed sample images are labeled using a child-friendly oriented spatial label set, wherein: The third-level element detail labels are selected based on the actual situation of the drawing; Select at least one of the second-level functional type labels; The first-level label "Child-friendly" is required.

6. The method according to claim 1, characterized in that The step 3 comprises: 3.1 Select a suitable CheckPoint model, use the child-friendly guided space sample library obtained in step 2 as the training set, adjust the standard CheckPoint model through the LoRA method, and obtain several alternative LoRA models; 3.2 For each candidate LoRA model, acceptance test and weighted screen test are performed through preset chart scripts; 3.3 All LoRA models that pass the test can be used as benchmark LoRA models for the generation of child-friendly community public space renewal intention maps; if all alternative LoRA models fail the test, the parameters are adjusted and the model with the best test performance among the alternative LoRA models is further fine-tuned until it passes the test and the benchmark LoRA model is obtained.

7. The method according to claim 6, characterized in that In step 3.2, for each candidate LoRA model, acceptance test and weighted screen test are performed through a preset chart script. If the following conditions are met at the same time, the test is passed, otherwise it is failed: Condition (1): When the prompt word contains the third-level element detail label, the screen elements generated by loading the LoRA model conform to the third-level element detail label; Condition (2): When the prompt word contains a second-level functional type label, the screen generated by loading the LoRA model contains an element with a third-level feature detail label that has a "proprietary" or "optional" relationship with it; Condition (3): The detail attributes of the third-level feature detail label element generated by the LoRA model are consistent with the style of the sample library; Condition (4): The spatial correlation between the image elements generated by the LoRA model complies with the basic knowledge of planning and the principle of spatial perspective.

Citation Information

Patent Citations

  • Planning intention graph generation method and system based on image generation model

    CN118570567A