Interactive generation method of urban functional composition ratio based on dedicated large language model

Through the interactive generation method of urban functional composition ratio based on a dedicated large language model, the problems of traditional urban functional composition ratio being time-consuming, labor-intensive and highly subjective are solved, the scientificity, rationality and efficiency of urban functional composition ratio are improved, and intelligent and accurate urban planning suggestions are provided.

CN119721045BActive Publication Date: 2025-10-03SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411768364.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-10-03
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

The traditional method of matching urban functional composition is difficult to balance multiple influencing factors such as population, economic development, resources and environment. It is time-consuming and labor-intensive, lacks comprehensive and integrated evaluation and optimization methods, is highly subjective, and affects the long-term development of urban functional composition.

Method used

An interactive generation method for urban functional composition ratio based on a dedicated large language model is adopted. Through data collection, preprocessing, semantic analysis, model training and feedback reinforcement learning, an urban design text case database is constructed. Natural language processing technology and machine learning algorithms are used for automatic classification and training to generate urban functional composition ratio suggestions, which are then modified and output through an interactive operation platform.

Benefits of technology

It has achieved the improvement of the scientificity, rationality and efficiency of the urban function composition ratio, provided targeted and comprehensive urban information, improved the intelligence and accuracy of the generation process, and enhanced the flexibility and scalability of the urban function composition ratio plan.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119721045B_ABST
    Figure CN119721045B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for interactively generating the urban function composition ratio based on a dedicated large language model, comprising the following steps: collecting urban design text case data; screening out corpus data corresponding to the planning area and zoning type of the design site; using natural language processing technology to perform semantic analysis and processing on the associated background information of the design site, converting it into prompt words related to the function ratio decision as a sample set; using the LLama‑7B pre-trained large language model, and sending the preliminary generation results of the pre-trained model to the user; retraining the large language model generated for the urban function composition ratio to form a dedicated large language model; and outputting the urban function composition ratio to an interactive operation platform, visualizing the urban functions and enabling interactive modification. The present invention solves the problem that the traditional urban function composition ratio is difficult to balance multiple influencing factors such as population, economic development, resource and environmental conditions, and is time-consuming and labor-intensive.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of urban planning, and in particular to a method for interactively generating urban function composition ratios based on a dedicated large language model. Background Art

[0002] The proportion of urban functional composition is a key task in urban planning and design, and it is an important foundation for the balanced and efficient distribution of urban functions. A reasonable proportion of urban functional composition is a scientific utilization of urban resources, which can improve the overall efficiency of the city, meet the diverse needs of residents, optimize urban environmental issues, and promote sustainable urban development. Cities are complex systems. Traditional proportions of urban functional composition require comprehensive consideration of planning objectives, population structure, economic development level, resource and environmental conditions, and other aspects. Collecting and analyzing this information is time-consuming and relies on the experience and judgment of urban planning experts and professionals. It lacks a comprehensive evaluation and optimization method, is subject to strong subjectivity, and suffers from an imperfect analytical perspective, which hinders the long-term development of urban functional composition. Summary of the Invention

[0003] Purpose of the Invention: To address the shortcomings of existing technologies and matching processes, the present invention provides an interactive method for generating urban functional composition matching based on a dedicated large-scale language model. This method addresses the time-consuming and labor-intensive nature of traditional urban functional composition matching, which struggles to balance multiple influencing factors such as population, economic development, and resource and environmental conditions. The present invention offers comprehensive, automated, intelligent, and multi-dimensional analysis capabilities, fully considering the actual conditions and needs of urban development and providing a scientific, reasonable, and sustainable matching solution for urban functional composition.

[0004] The technical solution adopted by the present invention is: a method for interactively generating urban functional composition ratios based on a dedicated large language model, comprising:

[0005] (1) Obtain the public standard text data of urban planning related departments and urban design text case data through the Python webpage data automatic collector, clean and preprocess the collected text data, and split it into complete corpus. Each corpus data is given attribute information describing the corpus characteristics, including corpus type, planning area, and zoning type, to form an urban design text case database;

[0006] (2) According to the planning area and zoning type of the design site input by the user, the urban design text case database is screened to select the corpus data corresponding to the planning area and zoning type of the design site;

[0007] (3) Using natural language processing technology to perform semantic analysis and processing on the contextual information of the design site, converting it into prompt words related to functional ratio decision making as a sample set;

[0008] (4) Use LLama-7B to pre-train the large language model and send the preliminary generated results of the pre-trained model to the user;

[0009] (5) Based on user feedback on the initial generated results, the large language model generated by the urban functional composition ratio is retrained using human feedback reinforcement learning (RLHF) to form a dedicated large language model;

[0010] (6) The city function composition ratio is output to the interactive operation platform, and the city functions are visualized and can be modified interactively.

[0011] Furthermore, the step (1) collects urban design text case data to form an urban design text case database, including:

[0012] We collected text data related to work reports, statistical yearbooks, and economic information, social, and demographic standards from official government websites at all levels, websites of planning-related functional departments, and the national standard full-text disclosure system. We also used a Python web data automatic collector to obtain text data from publicly available winning entries in the Shanghai Urban Design Challenge, hosted by the Shanghai Municipal Planning and Natural Resources Bureau.

[0013] The types of normative text data include urban planning-related planning documents, planning design documents, planning scheme evaluation reports, planning implementation evaluation reports, and various current norms, regulations, standards, indicators, guidelines, guidelines, and guidance opinions related to urban planning;

[0014] The collected text data was cleaned and preprocessed using Python 3.0-based computer automatic recognition and extraction to remove duplicate data, correct erroneous information, and split into complete corpora. Each corpus data was assigned attribute information describing the corpus characteristics, including corpus type, planning area, and zoning type.

[0015] The corpus types are divided into three categories: general knowledge, counting, and association. Among them, general knowledge corpus data includes generalized and basic urban design specifications and models, counting corpus data includes urban design-related indicators and quantitative data, and association corpus data includes descriptive sentences and textual expressions with urban planning expertise.

[0016] The planning area includes administrative and geographical coordinate information;

[0017] The zoning types are residential areas, comprehensive service areas, commercial and business areas, industrial development areas, logistics and warehousing areas, green space and leisure areas, transportation hub areas, strategic reserve areas, urban flexible development areas, and special-purpose areas. The classification is based on the "Guidelines for the Preparation of Municipal Land and Space Master Plans" compiled by the Ministry of Natural Resources. An urban design text case database is constructed, which includes a general database, a counting database, and an association database. Corpus data is imported into the three types of databases according to the corpus type.

[0018] The urban design text case database relies on the MySQL database management system and a storage medium for storing corpus data.

[0019] Furthermore, the step (2) acquires associated case data and design site-related background information according to the user's selection, including:

[0020] Based on the planning area and zoning type of the design site input by the user, an SQL query statement is written to filter the urban design text case database, and the corpus data corresponding to the planning area and zoning type of the design site are filtered out from the general database, count database, and association database respectively;

[0021] According to the site selected by the user, the background information associated with the design site is obtained.

[0022] Furthermore, the step (3) constructs city function composition ratio prompt words and instruction sentences as a training sample set, including:

[0023] A classification system for prompt words and instruction statements is established based on the six land use functions of commercial services, industrial and mining storage, residential, public administration and public services, transportation, and special land. The classification is based on GB / T 21010-2017 "Land Use Status Classification";

[0024] Use natural language processing technology and machine learning algorithms to automatically classify text data into prompt words and instruction sentences;

[0025] The semantic similarity calculation is used to analyze the association between the prompt word A and the instruction sentence B. The degree of association between them is evaluated by comparing their distance or similarity in the semantic space. The specific formula is:

[0026]

[0027] Where A·B is the dot product of vectors A and B, ||A||×||B|| is the modulus of vectors A and B respectively, and the calculation result range is [-1, 1]. Select the strongly associated prompt words and instruction sentences with calculation results greater than 0.5 to form a sample set of prompt words and instruction sentences, and identify the sample set of prompt words and instruction sentences in the selected corpus data;

[0028] Clean and annotate the constructed sample set to remove duplicates and incomplete data;

[0029] The sample set is divided into a training set and a test set, where 70% of the data is used as a training set and 30% of the data is used as a test set;

[0030] Natural language processing technology is used to perform semantic analysis and processing on the relevant background information of the design site, and convert it into prompt words related to functional ratio decisions.

[0031] Furthermore, the step (4) uses the llama-7B pre-training large language model method to include:

[0032] Build a large language model based on llama-7B;

[0033] Import prompt words and instruction sentences into the large language model module for pre-training;

[0034] Prompt-based Word fine-tuning The method performs pre-training, freezes all parameters of the main model llama-7B, and continuously updates the sentence embedding vector corresponding to the filled slot in the prompt word structure template through multiple iterations in the embedding layer; the number of iterations is set to 20; after every four iterations, an initial result of the pre-trained language large model is generated and returned to the user; the initial training result I = (I1 ~ I5);

[0035] The user asks questions to the initial training result I=(I1~I5) based on the test set data in S5 according to the prompt template, and generates a candidate answer set N=(N1~N5) containing multiple candidate answers. i )_; Calculate the text similarity T between the candidate answers and the test set data based on the TF-IDF algorithm, and select the initial training model with the highest similarity;

[0036] The text similarity T is calculated by using the Jieba word segmentation algorithm in pycharm to segment the candidate answer N i Split into multiple words, call the TF-IDF algorithm to calculate the similar word area M between the candidate answer and the test set data, and finally calculate

[0037] Furthermore, the step (5) is retrained using Reinforcement Learning with Human Feedback (RLHF) based on the user's feedback on the preliminary generated results. The specific method is as follows:

[0038] Based on the instruction statement in step (3), the initial training model is asked questions, the model samples and generates multiple candidate answers, and the user sorts the candidate answers; the preference sorting data set A = {sorting number, candidate answer number} is obtained and imported into the database for storage;

[0039] The pre-trained language model is trained using the gradient descent algorithm based on the preference ranking data. The loss values ​​of the prediction results, the selected answers, and the wrong answers are calculated. The training result with the minimum learning loss value is regarded as the reward model.

[0040] The learning loss value H is calculated as follows

[0041] H(p,q)=-∑ x p(x)logq(x)

[0042] Where p(x) is the true label label of the current input, and q(x) is the model's predicted value for each label label;

[0043] The enhanced training method is as follows: based on the instruction statement in step (3), the initial training model is asked multiple questions in an ultra-deep computer, and the generated results are optimized by the calculation results of the reward model; the training is stopped after the number of iterations preset by the user is reached, and a dedicated large language model is obtained; the model file is exported in py format and locally deployed on an ultra-deep computer. The ultra-deep computer configuration requirements are a 2*8-core processor, 80G video memory, and an NVIDIA DGX-2 deep learning system.

[0044] Furthermore, the step (6) outputs the city function composition ratio to the interactive operation platform, and the specific method is as follows:

[0045] Obtaining information on the recommended ratio of urban functional land use; the user inputs a site prompt word and a question template in Chinese into the deployed dedicated large language model; the dedicated large language model index construction tool calls the large language model API to access the large language model module, and obtains a text of a certain urban functional ratio under certain site conditions, wherein the recommended text includes tags [district functional positioning], [district main functional composition], [building carrier of a certain function], [functional proportion], [reference specifications and case sources] and corresponding content; outputting all texts to an interactive operation platform, which is a preset program that can be used for dialogue and includes a conversation unit, an intention recognition unit and a historical record unit; after receiving the generated results, the user continues to initiate a conversation modification instruction, and the intention recognition unit recognizes and generates intention parameters and feeds them back to the dedicated large language model to complete a new round of modified content output, which is stored by the historical record unit until the output content meets the user's expectations; the modified content is integrated by using data integration and translation equipment, and the modified content is printed into an urban functional ratio planning report by a printing device.

[0046] Beneficial effects:

[0047] (1) The present invention uses a dedicated large language model-based interactive generation system for urban functional composition ratios. This system uses a Python webpage data collector to acquire publicly available standard text data from urban planning departments and urban design case studies. The collected text data is cleaned and preprocessed, split into complete corpora, and each piece of corpus data is assigned attribute information describing the corpus characteristics, including corpus type, planning area, and zoning type, to form an urban design case database. This ensures the standardization and accuracy of the data, significantly improving the rationality, scientificity, and efficiency of generating urban functional composition ratios.

[0048] (2) The present invention is based on a dedicated large language model for interactively generating the proportion of urban functional composition. According to the planning area and zoning type of the design site input by the user, the urban design text case database is screened, and the corpus data corresponding to the planning area and zoning type of the design site are screened out from the general database, the counting database and the association database respectively. At the same time, according to the site selected by the user, the background information related to the design site is obtained, so that the urban information and theoretical basis for generating the urban functional composition are more targeted and comprehensive.

[0049] (3) The present invention is based on a city function composition ratio interactive generation system based on a dedicated large language model. It establishes a classification system for prompt words and instruction sentences according to the six land use functions of commercial services, industrial and mining storage, residential, public management and public services, transportation, and special land. It uses natural language processing technology and machine learning algorithms to automatically classify prompt words and instruction sentences for text data, and uses semantic similarity calculation to analyze the association relationship between prompt words and instruction sentences to form a prompt word and instruction sentence sample set; it uses natural language processing technology to perform semantic analysis and processing on the associated background information of the design site, and converts it into prompt words related to function ratio decision-making, which can fully improve the scientificity and accuracy of the sample set used to train and generate the large language model for urban function composition ratio.

[0050] (4) The present invention is based on a city function composition ratio interactive generation system based on a dedicated large language model, constructs a large language model based on llama-7B, imports prompt words and instruction sentences into the large language model module for pre-training, performs pre-training based on the prompt word fine-tuning method, freezes all parameters of the main model llama-7B, and continuously updates the sentence embedding vector corresponding to the filling slot in the prompt word structure template through multiple iterations in the embedding layer. After every four iterations, an initial result of the pre-trained language large model is generated and returned to the user; the user asks questions to the initial training results based on the prompt template based on the test set data, generates a candidate answer set containing multiple candidate answers, and screens the initial training model with the highest similarity in the generated results, which greatly improves the organicity and accuracy of the large language model used to generate the city function composition ratio.

[0051] (5) The present invention is an interactive generation system for city function composition ratio based on a dedicated large language model. Questions are asked to the initial training model, and the model samples and generates multiple candidate answers. The user sorts the candidate answers and imports them into the database for storage. The pre-trained language model is trained based on the preference sorting data using a gradient descent algorithm. The loss values ​​of the prediction results and the selected answers and wrong answers are calculated. The training result with the minimum learning loss value is regarded as a reward model, making the large language model used to generate the city function composition ratio more intelligent, efficient and accurate.

[0052] (6) The present invention is an interactive generation system for city function composition and matching based on a dedicated large language model. Users can input Chinese prompt words and question templates for the site into the deployed dedicated large language model to obtain a text of a city function matching suggestion under certain site conditions, and output all texts to a dialog-enabled interactive operation platform containing a conversation unit, an intention recognition unit, and a history record unit. After receiving the generated results, the user continues to initiate a conversation modification instruction, and the large language model completes a new round of modification content output until the output content meets the user's expectations. The modified content is integrated with data and printed into a city function matching planning report, which greatly improves the flexibility and scalability of the city function composition matching plan. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 Flow chart of the method of the present invention;

[0054] Figure 2 In the embodiment of the present invention, the designer asks questions to the dedicated large language model and outputs the technical logic of the content. DETAILED DESCRIPTION

[0055] The following takes the urban design of Chuanhuai New District in Zhoukou City as an example to illustrate the specific implementation process of the present invention. Figure 1 , showing the entire process of the present invention.

[0056] Specifically, step 1 uses web crawler technology to acquire information, using a Python web data collector to obtain public standard text data from urban planning-related departments and urban design text case data. This includes collecting and acquiring text data related to work reports, statistical yearbooks, and economic information, social and demographic standards from official websites of governments at all levels, websites of planning-related functional departments, and the national standard full-text disclosure system. Furthermore, the Python web data collector is used to obtain text data from publicly available winning entries in the Shanghai Urban Design Challenge, hosted by the Shanghai Municipal Planning and Natural Resources Bureau.

[0057] The types of normative text data include urban planning-related planning documents, planning design documents, planning scheme evaluation reports, planning implementation evaluation reports, and various current norms, regulations, standards, indicators, guidelines, guidelines, and guidance opinions related to urban planning;

[0058] In step 2, the collected text data is cleaned and preprocessed through computer automatic recognition and extraction using Python 3.0. This removes duplicate data, corrects erroneous information, and splits the data into complete corpora. Each piece of corpus data is assigned attribute information that describes the characteristics of the corpus, including corpus type, planning area, and zoning type.

[0059] The corpus types are divided into three categories: general knowledge, counting, and association. Among them, general knowledge corpus data includes generalized and basic urban design specifications and models, counting corpus data includes urban design-related indicators and quantitative data, and association corpus data includes descriptive sentences and textual expressions with urban planning expertise.

[0060] All three types of corpora are stored in txt format. The content examples are as follows:

[0061] General type:

[0062] [Weave the urban green space system, sort out the gray street corner spaces in the city, build pocket parks, and provide public spaces for the public to relax and socialize.]

[0063] [Relying on high-speed rail to form a comprehensive service sector]

[0064] Counting type:

[0065] [Opening windows over a large area allows for natural ventilation to remove heat and moisture from the room while preventing excessive solar radiation from entering. Therefore, the window ratio in public buildings should not exceed 0.85, and glass with a high shading coefficient should be used.]

[0066] [The urban area has approximately 2.69 square kilometers of parks and green spaces, or 2.99 square meters per person]

[0067] [The ratio of blue and green land exceeds 25%, of which green land accounts for 18% of the total land area and water area accounts for 4% of the total area.]

[0068] Association Class

[0069] The Super Central Axis includes the high-strength core of the business portal, the core of Sanshui Huiwo Park, the core of the government office, and the core of the performing arts park. These four cores drive the development of the urban axis of the Chuanhuai New Area.

[0070] [The renewal sector serves as a new engine for the development of the central part of the city, connecting with important urban functional sectors such as the central business district and the innovation and entrepreneurship cluster.]

[0071] The planning area includes administrative and geographical coordinate information;

[0072] The zoning types are residential areas, comprehensive service areas, commercial and business areas, industrial development areas, logistics and warehousing areas, green space and leisure areas, transportation hub areas, strategic reserved areas, urban flexible development areas, and special-purpose areas. The classification is based on the "Guidelines for the Preparation of Municipal Land and Space Master Plans" compiled by the Ministry of Natural Resources;

[0073] Construct an urban design text case database, which includes a general knowledge database, a counting database, and an association database. Import corpus data into the three types of databases according to corpus type;

[0074] The urban design text case database relies on the MySQL database management system and a storage medium for storing corpus data.

[0075] Step 3: Input the planning area and zoning type of the design site, write an SQL query statement, filter the urban design text case database, and filter the corpus data corresponding to the planning area and zoning type of the design site in the general database, count database, and association database respectively;

[0076] For example, [SELECT latitude between(33°25N,33°49N),logitude between(115°00W,115°38'W),partition type = 'commercial business district' FROM 'general knowledge corpus' WHERE 'file location']

[0077] According to the site selected by the user, the background information associated with the design site is obtained.

[0078] Step 4: Establish a classification system for prompt words and instruction statements based on the six land use functions of commercial services, industrial and mining storage, residential, public administration and public services, transportation, and special land. The classification is based on GB / T21010-2017 "Land Use Status Classification";

[0079] Natural language processing technology and machine learning algorithms are used to automatically classify prompt words and instruction sentences for text data; this embodiment does not limit the algorithm, as long as it has classification capabilities.

[0080] The prompt words and template examples are as follows: The bracket information containing the mask is the fill slot, and the internal format of the fill slot is a text vector.

[0081] "You are an experienced urban designer. Please analyze the key information here based on the main functions of the site and the related information of the upper-level planning, {information mask}. Your analysis should reflect that this information is obtained from the site description provided by the user. The site location hierarchy and site zoning type must be known information provided by the user. Analyze and provide the land use function positioning {area function positioning mask}, the main functional composition ratio {function ratio mask}, and the main building types {building carrier mask}. The analysis process should be detailed and the reference document title {reference specifications and case sources} should be given. Speak like a senior urban and rural planner and have a gentle and friendly tone. The person you are talking to can be

[0082] {Other urban designers, students, and scientific researchers};”

[0083] The semantic similarity calculation is used to analyze the association between the prompt word A and the instruction sentence B. The degree of association between them is evaluated by comparing their distance or similarity in the semantic space. The specific formula is:

[0084]

[0085] Where A·B is the dot product of vectors A and B, ||A||×||B|| is the modulus of vectors A and B respectively, and the calculation result range is [-1, 1]. Select the strongly associated prompt words and instruction sentences with calculation results greater than 0.5 to form a sample set of prompt words and instruction sentences, and identify the sample set of prompt words and instruction sentences in the selected corpus data;

[0086] Clean and annotate the constructed sample set to remove duplicates and incomplete data;

[0087] The sample set is divided into a training set and a test set, where 70% of the data is used as a training set and 30% of the data is used as a test set;

[0088] Natural language processing technology is used to perform semantic analysis and processing on the relevant background information of the design site, and convert it into prompt words related to functional ratio decisions.

[0089] Step 5

[0090] The base large language model used in this invention is the llama-7B model. This open-source large language model is based on the General Language Model (GLM) framework and has up to 1 trillion training parameters. It is commercially available. Its technology is similar to ChatGPT and is currently one of the best-performing open-source large language models. Furthermore, this invention uses a cue word fine-tuning method to fine-tune the large language model.

[0091] First, deploy a large language model based on llama-7B.

[0092] Import the prompt word and instruction sentence files into the large language model for pre-training.

[0093] The present invention is based on the prompt Word fine-tuning The method performs pre-training, freezing all parameters of the main llama-7B model. The embedding layer then iterates repeatedly to update the sentence embedding vectors corresponding to the filled slots in the prompt word structure template. The number of iterations is set to 20. After every four iterations, an initial result of the pre-trained language model is generated and returned to the user. The initial training result, I = (I1-I5), is obtained.

[0094] The user asks questions to the initial training result I=(I1~I5) based on the test set data in S5 according to the prompt template, and generates a candidate answer set N=(N1~N5) containing multiple candidate answers. i ). Calculate the text similarity T between the candidate answers and the test set data based on the TF-IDF algorithm, and select the initial training model with the highest similarity in the generated results.

[0095] The text similarity T is calculated by using the Jieba word segmentation algorithm in pycharm to segment the candidate answer N i Split into multiple words, call the TF-IDF algorithm to calculate the similar word area M between the candidate answer and the test set data, and finally calculate

[0096] Step 6 asks questions to the initial training model based on the instruction statement in step 4, for example

[0097] {

[0098] “Question”: Please write 5 common characteristic scenes in urban design.

[0099] “input”:””,

[0100] “output”: OK, now we give you the following solution:

[0101] Candidate answer 1: Waterfront park, cultural square, historical district, commercial pedestrian street, ecological green corridor;

[0102] Candidate answer 2: Old street style, ancient bridge charm, deep courtyard, city green lung, folk custom pavilion;

[0103] Candidate answer 3: Waterfront oasis, ancient alleys, window to the future, green lung heart, cultural square;

[0104] }

[0105] The model samples and generates multiple candidate answers, which are then ranked by the user from satisfactory to unsatisfactory. The resulting preference ranking dataset A = {ranking number, candidate answer number} is imported into the database for storage.

[0106] The gradient descent algorithm is used to train the pre-trained language model based on the preference ranking data. The loss values ​​of the prediction results, the selected answers, and the wrong answers are calculated. The training result with the minimum learning loss value is regarded as the reward model.

[0107] The learning loss value H is calculated as follows

[0108] H(p,q)=-∑ x p(x)logq(x)

[0109] Where p(x) is the true label of the current input, and q(x) is the model's predicted value for each label.

[0110] The enhanced training method is as follows: Based on the instruction statement in step (3), the initial training model is asked multiple questions in the ultra-deep computer, and the generated results are optimized by the calculation results of the reward model. The training is stopped after the number of iterations preset by the user is reached, and a dedicated large language model is obtained. The model file is exported in py format and deployed locally on the ultra-deep computer. The ultra-deep computer configuration requirements are 2*8 core processors, 80G video memory, and equipped with NVIDIA DGX-2 deep learning system.

[0111] Step 7: Obtain recommended information on the urban functional land use ratio. The user enters a Chinese site prompt and a question template into the deployed dedicated large language model. The dedicated large language model index construction tool calls the large language model API to access the large language model module and obtains a text recommendation for a specific urban functional land use ratio under specific site conditions. The recommended text includes the tags [Area Function Positioning], [Area Main Function Composition], [Architectural Carrier for a Certain Function], [Function Ratio], [Reference Standards and Case Sources], and corresponding content.

[0112] In this example, the output result text is as follows:

[0113] {

[0114] "Instruction": We are participating in a district-scale urban design competition. The location is the Chuanhuai New District in Zhoukou City, Henan Province. The area within the red line is 2 square kilometers. It is a new district to be developed. Our initial idea is to create a commercial and business district. Based on this information, can you give some suggestions on functional positioning and proportions?

[0115] “input”:””,

[0116] “output”: OK, now we recommend the following solutions for you:

[0117] 1. Area Function Positioning: Financial Business Center Innovation Demonstration Zone

[0118] 2. The main functions of the area are: business office land, cultural and leisure land, commercial retail land, public green land, transportation and infrastructure land,

[0119] 3. Building carrier:

[0120] Business office space: corporate headquarters, financial trading centers, international hotels, high-end business hotels, high-end maker spaces

[0121] Cultural and leisure land: music exhibitions, gymnasiums, art galleries

[0122] Commercial retail land: high-end shopping, charming water street

[0123] Public green space: street parks, rooftop greenways

[0124] Transportation and infrastructure land: including land for roads and public facilities

[0125] 4. Functional ratio: Based on the blue-green network, industry leadership, and mixed functions, it is recommended that the land use for business offices account for 30%, retail commercial land for 20%, R&D land for 10%, administrative offices for 10%, life service facilities for 10%, residential land for 10%, and other land for 10%;

[0126] 5. Reference standards and case studies: Luoyang Luolong Science and Technology Park Financial Service Center, Tangshan Caofeidian New District Financial Business Center; Reference standard: GB / T 21010-2017 "Land Use Status Classification";

[0127] }

[0128] All texts are output to the interactive operation platform, which is a pre-set program that can be used for dialogue, including a conversation unit, an intention recognition unit, and a history record unit. After receiving the generated results, the user can continue to initiate conversation modification instructions.

[0129] For example,

[0130] [{mask}, can you give me more details?]

[0131] [Write a few more {mask}], where the mask information is the prompt word

[0132] In this example, the question is

[0133] Regarding the blue-green network, industry leadership, and hybrid functions, please elaborate further and output the complete text in conjunction with historical records.

[0134] The intent recognition unit identifies and generates intent parameters and feeds them back to the dedicated large language model to complete a new round of modified content output, which is then stored by the history recording unit until the output content meets user expectations.

[0135] The modified output is as follows:

[0136] {

[0137] "output": OK, after modification, we now recommend the following solution for you:

[0138] 1. Area Function Positioning: Financial Business Center Innovation Demonstration Zone

[0139] 2. The main functions of the area are: business office land, cultural and leisure land, commercial retail land, public green land, transportation and infrastructure land,

[0140] 3. Building carrier:

[0141] Business office space: corporate headquarters, financial trading centers, international hotels, high-end business hotels, high-end maker spaces

[0142] Cultural and leisure land: music exhibitions, gymnasiums, art galleries

[0143] Commercial retail land: high-end shopping, charming water street

[0144] Public green space: street parks, rooftop greenways

[0145] Transportation and infrastructure land: including land for roads and public facilities

[0146] 4. Functional Proportion: We recommend starting with blue-green networks, industry leadership, and functional mixing:

[0147] Blue-green network: Protect and restore the city’s water systems and green spaces, establish a comprehensive blue-green network system, organically connect various areas in the area, and form a complete ecological network

[0148] Industry leadership: Facing the demands of new economic forms in future urban development, we will take financial centers, headquarters economy, new internet economy and other future-oriented modern service industries as industrial support.

[0149] Functional mixing: Urban construction land emphasizes mixed land use and complex functions. Among them, new industrial mixed and residential mixed land include residential, office, R&D, service and other functions, forming block units that integrate work and residence.

[0150] 30% for business offices, 20% for retail and commercial use, 10% for research and development, 10% for administrative offices, 10% for living service facilities, 10% for residential use, and 10% for other uses;

[0151] 5. Reference standards and case studies: Luoyang Luolong Science and Technology Park Financial Service Center, Tangshan Caofeidian New District Financial Business Center; Reference standard: GB / T 21010-2017 "Land Use Status Classification";

[0152] }

[0153] The modified content is integrated by using data integration and translation equipment, and the modified content is printed into an urban function allocation planning report through a printing device.

[0154] For detailed process, please refer to Figure 2 .

[0155] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention, and such changes and modifications fall within the scope of the invention as claimed.

Claims

1. An interactive generation method for urban functional composition ratio based on a dedicated large language model, characterized by: include: (1) Obtain the public standard text data of urban planning related departments and urban design text case data through the Python webpage data automatic collector, clean and preprocess the collected text data, and split it into complete corpus. Each corpus data is given attribute information describing the corpus characteristics, including corpus type, planning area, and zoning type, to form an urban design text case database; (2) According to the planning area and zoning type of the design site input by the user, the urban design text case database is screened to select the corpus data corresponding to the planning area and zoning type of the design site; (3) Using natural language processing technology to perform semantic analysis and processing on the contextual information of the design site, converting it into prompt words related to functional ratio decision making as a sample set; (4) Use LLama-7B to pre-train the large language model and send the preliminary generated results of the pre-trained model to the user; (5) Based on user feedback on the initial generated results, human feedback reinforcement learning is used to retrain the large language model generated by the urban functional composition ratio to form a dedicated large language model; (6) Output the city function composition ratio to the interactive operation platform, visualize the city functions and enable interactive modification; The step (3) constructs the city function composition ratio prompt words and instruction sentences as a training sample set, including: A classification system for prompt words and instruction statements is established based on the six land use functions of commercial services, industrial and mining storage, residential, public administration and public services, transportation, and special land. The classification is based on GB / T 21010-2017 "Land Use Status Classification"; Use natural language processing technology and machine learning algorithms to automatically classify text data into prompt words and instruction sentences; The semantic similarity calculation is used to analyze the association between the prompt word A and the instruction sentence B. The degree of association between them is evaluated by comparing their distance or similarity in the semantic space. The specific formula is: Where A·B is the dot product of vectors A and B, ||A||×||B|| is the modulus of vectors A and B respectively, and the calculation result range is [-1, 1]. Select the strongly associated prompt words and instruction sentences with calculation results greater than 0.5 to form a sample set of prompt words and instruction sentences, and identify the sample set of prompt words and instruction sentences in the selected corpus data; Clean and annotate the constructed sample set to remove duplicates and incomplete data; The sample set is divided into a training set and a test set, where 70% of the data is used as a training set and 30% of the data is used as a test set; Use natural language processing technology to perform semantic analysis and processing on the relevant background information of the design site, and convert it into prompt words related to functional ratio decision-making; The step (4) uses the llama-7B pre-training large language model method to include: Build a large language model based on llama-7B; Import prompt words and instruction sentences into the large language model module for pre-training; Pre-training is performed based on the prompt word fine-tuning method. All parameters of the main model l lama-7B are frozen. The sentence embedding vector corresponding to the filled slot in the prompt word structure template is continuously updated through multiple iterations in the embedding layer. The number of iterations is set to 20. After every four iterations, an initial result of the pre-trained language model is generated and returned to the user. The initial training result I = (I1-I5) is obtained. The user asks questions to the initial training result I = (I1 ~ I5) based on the test set data in step (3) according to the prompt template, and generates a candidate answer set N = (N1 ~ N5) containing multiple candidate answers. i )_; Calculate the text similarity T between the candidate answers and the test set data based on the TF-IDF algorithm, and select the initial training model with the highest similarity; The text similarity T is calculated by using the Jieba word segmentation algorithm in pycharm to segment the candidate answer N i Split into multiple words, call the TF-IDF algorithm to calculate the similar word area M between the candidate answer and the test set data, and finally calculate 2. The interactive generation method of urban function composition ratio based on a dedicated large language model according to claim 1 is characterized in that: The step (1) collects urban design text case data to form an urban design text case database, including: We collected text data related to work reports, statistical yearbooks, and economic information, social, and demographic standards from official government websites at all levels, websites of planning-related functional departments, and the national standard full-text disclosure system. We also used a Python web data automatic collector to obtain text data from publicly available winning entries in the Shanghai Urban Design Challenge, hosted by the Shanghai Municipal Planning and Natural Resources Bureau. The types of normative text data include urban planning-related planning documents, planning design documents, planning scheme evaluation reports, planning implementation evaluation reports, and various current norms, regulations, standards, indicators, guidelines, guidelines, and guidance opinions related to urban planning; The collected text data was cleaned and preprocessed using Python 3.0-based computer automatic recognition and extraction to remove duplicate data, correct erroneous information, and split into complete corpora. Each corpus data was assigned attribute information describing the corpus characteristics, including corpus type, planning area, and zoning type. The corpus types are divided into three categories: general knowledge, counting, and association. Among them, general knowledge corpus data includes generalized and basic urban design specifications and models, counting corpus data includes urban design-related indicators and quantitative data, and association corpus data includes descriptive sentences and textual expressions with urban planning expertise. The planning area includes administrative and geographical coordinate information; The zoning types are residential areas, comprehensive service areas, commercial and business areas, industrial development areas, logistics and warehousing areas, green space and leisure areas, transportation hub areas, strategic reserve areas, urban flexible development areas, and special-purpose areas. The classification is based on the "Guidelines for the Preparation of Municipal Land and Space Master Plans" compiled by the Ministry of Natural Resources. An urban design text case database is constructed, which includes a general database, a counting database, and an association database. Corpus data is imported into the three types of databases according to the corpus type. The urban design text case database relies on the MySQL database management system and a storage medium for storing corpus data.

3. The interactive generation method of urban function composition ratio based on a dedicated large language model according to claim 2 is characterized in that: The step (2) acquires relevant case data and design site-related background information according to the user's selection, including: Based on the planning area and zoning type of the design site input by the user, an SQL query statement is written to filter the urban design text case database, and the corpus data corresponding to the planning area and zoning type of the design site are filtered out from the general database, count database and association database respectively; according to the site selected by the user, the background information related to the design site is obtained.

4. The interactive generation method of urban function composition ratio based on a dedicated large language model according to claim 3 is characterized in that: The step (5) is to use human feedback reinforcement learning to perform retraining based on the user's feedback on the preliminary generated results. The specific method is as follows: Based on the instruction statement in step (3), the initial training model is asked questions, the model samples and generates multiple candidate answers, and the user sorts the candidate answers; the preference sorting data set A = {sorting number, candidate answer number} is obtained and imported into the database for storage; The pre-trained language model is trained using the gradient descent algorithm based on the preference ranking data. The loss values ​​of the prediction results, the selected answers, and the wrong answers are calculated. The training result with the minimum learning loss value is regarded as the reward model. The learning loss value H is calculated as follows H(p,q)=-∑ x p(x)logq(x) In the formula, p(x) is the true label label of the current input, and q(x) is the model's predicted value for each label label; The enhanced training method is as follows: based on the instruction statement in step (3), the initial training model is asked multiple questions in an ultra-deep computer, and the generated results are optimized by the calculation results of the reward model; the training is stopped after the number of iterations preset by the user is reached, and a dedicated large language model is obtained; the model file is exported in py format and locally deployed on an ultra-deep computer. The ultra-deep computer configuration requirements are a 2*8-core processor, 80G video memory, and an NVIDIA DGX-2 deep learning system.

5. The interactive generation method of urban function composition ratio based on a dedicated large language model according to claim 4 is characterized in that: The step (6) outputs the city function composition ratio to the interactive operation platform. The specific method is: Obtaining information on the recommended ratio of urban functional land use; the user inputs a Chinese prompt word and a question template for the site into the deployed dedicated large language model; the dedicated large language model index construction tool calls the large language model API to access the large language model module to obtain a text of a certain urban functional ratio under certain site conditions, the text of which includes the tags [area functional positioning], [area main functional composition], [building carrier of a certain function], [functional ratio], [reference standards and case sources] and corresponding content; all texts are output to the interactive operation platform, which is a pre-set program with dialogue capabilities that includes a conversation unit, an intention recognition unit, and a history recording unit; After receiving the generated results, the user continues to initiate a conversation modification instruction. The intention recognition unit identifies the generated intention parameters and feeds them back to the dedicated large language model to complete a new round of modified content output. The historical record unit stores the modified content until the output content meets the user's expectations. The modified content is integrated by using data integration and translation equipment, and the modified content is printed into an urban function ratio planning report through a printing device.

Citation Information

Patent Citations

  • Bi-level energy management system with grid reinforcement for solar energy

    AU2021105856A4

  • Zero-sample large model generation code detection method and system

    CN117608648A