Urban planning knowledge graph construction method based on large language model

Through a method based on the large language model, naming entity recognition, relationship extraction, triple evaluation and text classification technologies are used to solve the problem of insufficient application of unstructured text data in the existing technology, and the rapid construction of urban planning knowledge maps without labeled corpus is achieved, improving efficiency and accuracy.

CN120031118APending Publication Date: 2025-05-23SUN YAT SEN UNIV

Patent Information

Application Number
CN202510186545.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing urban geographical knowledge graph research lacks direct application of unstructured text data, and the traditional KG construction method relies on labeled corpus to train deep learning models, which is inefficient, and faces greater challenges especially in the Chinese vertical NLP tasks.

Method used

Using a method based on a large language model, urban planning knowledge graphs are constructed through named entity recognition, relationship extraction, triple evaluation and text classification, which can quickly build knowledge graphs without labeled corpus and reduce manual participation.

Benefits of technology

It realizes the rapid construction of knowledge graphs without labeled corpus, improves task efficiency, reduces manual participation, and enhances the generation of plug-in knowledge bases through syntactic analysis and retrieval, improving the accuracy of Chinese text understanding and text classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031118A_ABST
    Figure CN120031118A_ABST
Patent Text Reader

Abstract

The invention provides an urban planning knowledge graph construction method based on a large language model. The method comprises the following steps: preprocessing an urban planning text; constructing a cue word template; using the cue word template to guide the large language model to extract planning knowledge entities in the city planning text; instructing the large language model to organize the entities into a structured triple meeting text semantics; guiding the large language model to perform mutual evaluation on the structured triples, and screening the triples based on an evaluation result; and respectively classifying head and tail entities and relationships of the triple set stored in the knowledge graph by using a classifier. According to the method, the information extraction task and the knowledge graph construction are completed through the large language model, and the task efficiency is improved; syntactic analysis information is added into a triple extraction task, so that the understanding ability of a large language model on a Chinese text is effectively enhanced, and the model is helped to automatically complement omission information in clauses; and the text classification accuracy of the large language model is improved by using a manner of generating a plug-in knowledge base through retrieval enhancement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge graph construction, and in particular to a method for constructing an urban planning knowledge graph based on a large language model. Background Art

[0002] Organizing the planning knowledge about urban development and spatial layout in planning documents into a structured form that is easy for the public and researchers to understand, learn and use can provide important data supplements for existing land use monitoring and urban computational simulation research. Knowledge Graph (KG), as a way to organize entities, attributes and their relationships in the real world into a semantic network composed of nodes and edges, provides a feasible path for structured storage of planning knowledge.

[0003] With the rise of Large Language Model (LLM), it has demonstrated excellent capabilities in text understanding and generation, providing a new paradigm for the natural language processing (NLP) work necessary for building KG. General LLM can achieve high-quality information extraction (IE) with few or even zero samples, which provides a basis for the automated construction of KG.

[0004] However, existing research on urban geographic knowledge graphs mainly focuses on top-level design based on ontology models and the construction of spatiotemporal knowledge graphs for urban computing based on (semi-)structured data, lacking direct application to unstructured text data. This may be because traditional KG construction methods rely on the use of annotated corpora to train deep learning models, which requires a lot of manpower and time costs and is inefficient. In addition, research on applying LLM to NLP tasks in the Chinese vertical field is relatively limited. Since the grammatical semantics and syntactic structure of Chinese are more complex than those of English, it brings greater challenges to the understanding and processing of LLM. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention provides a method for constructing an urban planning knowledge graph based on a large language model. The named entity recognition, relationship extraction, triple evaluation and text classification of the present invention can effectively and quickly construct a knowledge graph without labeled corpus, improve task efficiency and reduce manual participation.

[0006] The technical solution of the present invention is: a method for constructing an urban planning knowledge graph based on a large language model, comprising the following steps:

[0007] S1), preprocessing the urban planning text to obtain a preprocessed text;

[0008] S2), constructing a prompt word template, through which the identity of the large language model is specified, the task steps are introduced, the thinking of the large language model is guided, and the large language model is stimulated to complete various NLP tasks;

[0009] S3), using prompt word templates to guide the large language model to extract planning knowledge entities in urban planning texts;

[0010] S4), based on the preprocessed text, the prompt word template and the named entity recognition result, instructing the large language model to organize the entities into structured triples that satisfy the text semantics;

[0011] S5) Instruct the large language model to evaluate the structured triples, and screen the triples based on the evaluation results to obtain:

[0012] G = {E, R, F};

[0013] Among them, G represents the urban planning knowledge graph (knowledge base); E represents the entity set in the knowledge base; R represents the relationship set in the knowledge base; F represents the triple set stored in the knowledge graph;

[0014] The triple set F stored in the knowledge graph is expressed as:

[0015] F={(h,r,t)|h,t∈E,r∈R};

[0016] Among them, (h, r, t) represents a triple; h and t represent the head and tail entities respectively; r represents the relationship from the head entity to the tail entity.

[0017] S6) Construct a classifier based on the large language model, and use the classifier to classify the head and tail entities and relations of the triple set F stored in the knowledge graph, and use the classification results as the attributes of each component of the triple (h, r, t).

[0018] Preferably, in step S1), the preprocessing includes sentence segmentation and segmentation of the urban planning text; and component syntactic analysis CON and dependency syntactic analysis DEP processing of the sentences.

[0019] Preferably, in step S1), the urban planning text is segmented into sentences and paragraphs, wherein sentences are used as basic processing units and paragraphs are used to provide contextual background information to help the large language model understand semantics.

[0020] Preferably, in step S1), the natural language processing library HanLP is used to perform syntactic analysis, component syntactic analysis CON is used to express the compositional structural relationship between adjacent words in a sentence, and dependency syntactic analysis DEP is used to express the hierarchical relationship between words in a sentence.

[0021] Preferably, in step S2), the prompt word template includes five parts, namely:

[0022] Task description: used to define the role identity of the large language model through system parameters, inform the large language model of the corresponding role through system parameters, and describe the task requirements through user parameters;

[0023] Candidate target list: for named entity recognition tasks, triple evaluation tasks, and attribute classification tasks;

[0024] Task example: used to guide a large language model to complete the mapping from input to output;

[0025] Task emphasis: repeatedly emphasize matters needing attention and use a strong commanding tone to require the large language model to pay attention to and implement implementation details;

[0026] Second round of dialogue: The initial results are verified and optimized through the second round of dialogue.

[0027] Preferably, in step S3), using the prompt word template to guide the large language model to extract planning knowledge entities in the urban planning text specifically includes the following steps:

[0028] S31), taking the sentence as the processing unit, construct five groups of task examples in the form of sentence and corresponding entity recognition results:

[0029] S32), under the guidance of the prompt word template, the sentences to be processed are input in sequence, so that the large language model recognizes the named entities of the urban planning text;

[0030] S33) re-inputting the entity results obtained in step S32) into the large language model, requiring the large language model to verify and re-extract the results to ensure recognition quality.

[0031] Preferably, in step S4), based on the preprocessed text, the prompt word template and the named entity recognition result, the large language model is instructed to organize the entities into structured triples that satisfy the text semantics, which specifically includes the following steps:

[0032] S41), using the preprocessed text of step S1) and the named entity recognition results of each sentence of step S3) as inputs of the large language model, and using the recognized entities as candidate targets of the organizational planning knowledge triples;

[0033] S42), using the prompt word template to guide the input data of the large language model step S41) to learn contextual semantics, and organize the entities into SPO triples that conform to the semantic relationship of the subject, predicate, and object of the text data.

[0034] Preferably, in step S5), the large language model is guided to perform mutual evaluation on the structured triples, and the triples are screened based on the evaluation results, which specifically includes the following steps:

[0035] S51), adding a thinking chain CoT example to the prompt word template to guide reasoning, guiding the large language model to refer to the natural language inference model NLI, expanding each SPO triple into 3 complete short sentences as a hypothesis, taking the corresponding paragraph of the triple as a premise, and judging whether the hypothesis is implied in the premise;

[0036] Take three complete sentences as hypotheses and instruct the model to judge whether the hypothesis can be derived from the premise, that is, whether the premise implies the hypothesis, and give the answer: yes / no.

[0037] Finally, the command model combines the above reasoning results and comprehensively evaluates the SPO triples from five dimensions: semantics, consistency, authenticity, accuracy, and comprehensibility.

[0038] S52), instructing the large language models to evaluate the quality of the generated triples in a pairwise comparison and ranking manner, and collectively forming a set of answers;

[0039] S53), instruct the large language model to evaluate the quality of the generated triples in a multi-dimensional scoring manner. Based on a set of results obtained in step S52), the large language model is used to score the generated triples in the five dimensions mentioned in step S51) with a score of 1 to 5. After the mutual evaluation is completed, the triples with too low scores are filtered out to obtain the final results.

[0040] Preferably, in step S6), a classifier is constructed based on a large language model, and the head and tail entities and relations of the triple set F stored in the knowledge graph are classified by the classifier, and the classification results are used as attributes of each component of the triple (h, r, t), which specifically includes the following steps:

[0041] S61), select some entities generated in step S3) as training data, obtain vector representation of entity words, store entity words, corresponding vectors and corresponding classifications in a vector database as an enhanced retrieval generation plug-in knowledge base;

[0042] S62), using the large language model as a classifier, inputting the entities to be classified in sequence, obtaining the vector representation of the entities to be classified, calculating the cosine similarity between the entities and the existing records in the vector database, and returning the top five records with the highest similarity; instructing the large language model to use the labels of the returned records as a reference to assign labels to the entities to be classified;

[0043] S63) instructing the large language model to refine the classification results in step S62) based on general knowledge.

[0044] The beneficial effects of the present invention are:

[0045] 1. The present invention completes the information extraction task based on a general large language model that has not been fine-tuned, and can effectively and quickly construct a knowledge graph without annotated corpus, thereby improving task efficiency and reducing manual participation;

[0046] 2. The present invention adds syntactic analysis information to the triple extraction task, which effectively enhances the large language model's ability to understand Chinese text and helps the model automatically complete omitted information in sentences;

[0047] 3. The present invention uses retrieval enhancement to generate a plug-in knowledge base to improve the accuracy of text classification by a large language model. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a schematic diagram of the process of the present invention;

[0049] Figure 2 It is a schematic diagram of the evaluation process of the triplet of the present invention;

[0050] Figure 3 A schematic diagram of the process of adding attributes for text classification of the present invention. DETAILED DESCRIPTION

[0051] The specific implementation of the present invention will be further described below in conjunction with the accompanying drawings:

[0052] like Figure 1 As shown, this embodiment provides a method for constructing an urban planning knowledge graph based on a large language model, comprising the following steps:

[0053] S1), preprocessing the urban planning text to obtain a preprocessed text;

[0054] In this embodiment, the preprocessing includes sentence segmentation and segmentation processing of the urban planning text; and component syntactic analysis CON and dependency syntactic analysis DEP processing of the sentences.

[0055] By segmenting the urban planning text into sentences and paragraphs, with sentences as the basic processing unit, the large language model can focus more on text details. Paragraphs are used to provide contextual information to help the large language model understand semantics.

[0056] In this embodiment, the natural language processing library HanLP is used for syntactic analysis, including word segmentation, part-of-speech tagging, syntactic analysis, etc. Component syntactic analysis CON is used to express the compositional structure relationship between adjacent words in a sentence, and dependency syntactic analysis DEP is used to express the hierarchical relationship between words in a sentence. At the same time, word embedding is obtained through text-embedding-3-large.

[0057] S2), constructing a prompt word template, through which the identity of the large language model is specified, the task steps are introduced, the thinking of the large language model is guided, and the large language model is stimulated to complete various NLP tasks;

[0058] In this embodiment, the large language model adopts GPT-4o, or Qwen-Max, or Doubao-pro, or GLM-4.

[0059] In this embodiment, the prompt word template includes five parts, namely:

[0060] Task description: used to define the role identity of the large language model through system parameters, inform the large language model of the professional capabilities of the corresponding role through system parameters, and describe the task requirements through user parameters;

[0061] Candidate target list: used for named entity recognition tasks, triple evaluation tasks, and attribute classification tasks; it is organized in the form of "Candidate target: concise definition". This list can help the large language model to clarify the entities to be extracted, the dimensions to be evaluated, and the groups to be classified, while achieving understanding and distinction of professional terms.

[0062] Task examples: used to guide the large language model to complete the mapping from input to output; in the named entity recognition task, five sets of task examples are given, including input of the sentence to be processed and output of the entity recognition result. The large language model will extract the results based on the task example reference, think about the reasoning process and learn the output format.

[0063] Task emphasis: repeatedly emphasize matters needing attention and use a strong commanding tone to require the large language model to pay attention to and implement implementation details;

[0064] Second round of dialogue: The second round of dialogue verifies and optimizes the initial results, improves the accuracy of entity recognition and relationship extraction, and improves the quality of knowledge graph construction.

[0065] S3), using the prompt word template to guide the large language model to extract planning knowledge entities in the urban planning text; this embodiment uses the large language model Qwen-Max to extract five types of planning knowledge entities with strong representativeness in the planning text, including: location, land use function, planning concept, orientation and planning action, specifically including the following steps:

[0066] S31), firstly, based on the prompt word template, the identity of the large language model Qwen-Max is specified as: a planner who can easily identify various geographical planning knowledge entities in the text;

[0067]

[0068]

[0069]

[0070] Add the above five groups of examples to the task description;

[0071] Then, the task is briefly described as: identifying various types of place named entities in the text given by the user and outputting them strictly in a structured format;

[0072] S32) Provide a list of candidate targets as follows:

[0073] Location: a real geographical entity that can be located on a map, including provinces, cities, districts, towns, streets, and more specific locations such as various POIs including schools, shopping malls, and scenic spots;

[0074] Land use function: the definition of land use type and functional layout in urban planning, such as "high-tech industrial cluster", "industrial land", "residential land", "ecological protection area", etc.;

[0075] Planning concepts: planning goals and concepts proposed in urban planning, such as "expanding to the south, optimizing the north, advancing to the east, linking to the west", "three verticals and five horizontals, one ring and multiple corridors" for Guangzhou planning;

[0076] Direction: a synonym for indicating direction, including east, south, west, north, etc. In this embodiment, it is usually a phrase containing a directional pronoun, such as "central part", "northern section", etc.;

[0077] Planning action: In urban planning, a phrase indicating a planned action or measure, such as "advance", "renewal", "protection", etc.

[0078] S33) provides five sets of task examples, with structural references:

[0079] "input":"'Prioritize the development space of advanced manufacturing, strategic emerging industries and urban industries, and vigorously promote the construction of value innovation parks."';

[0080] "output":"'{"Location":[],"Land use function":["Advanced manufacturing","Strategic emerging industries","Urban industries","Value innovation park"],"Direction":[],"Concept":[],"Planning behavior":["Priority protection","Promote"]}"'};

[0081] S34), based on the guidance of the information in steps S31) to S32), while emphasizing the input-output mapping and output format of the large language model Qwen-Max referenced in the example provided in step S33), the first round of entity recognition results are obtained, and the results are stored in the large language model Qwen-Max for retaining the assistant parameter information of the historical conversation.

[0082] S35) Set up a second round of dialogue, take the sentence and corresponding entity results as input, command the large language model Qwen-Max to re-extract entities based on the extracted content and experience of the previous stage, and obtain the final result of this embodiment.

[0083] S4), based on the preprocessed text, the prompt word template and the named entity recognition result, instruct the large language model to organize the entities into structured triples that satisfy the text semantics; specifically comprising the following steps:

[0084] S41), using the preprocessed text in step S1) and the named entity recognition results of each sentence in step S3) as inputs to the large language models GPT-4o, Doubao-pro-32k and GLM-4-AirX, and using the recognized entities as candidate targets for organizational planning knowledge triples;

[0085] S42), use the prompt word template to guide the large language models GPT-4o, Doubao-pro-32k and GLM-4-AirX to learn contextual semantics according to the input data of step S41), and organize the entities into SPO triples (head entity, relationship, tail entity) that conform to the semantic relationship of subject, predicate, and object of the text data.

[0086] S5) Instruct the large language models GPT-4o, Doubao-pro-128k and GLM-4-Air to evaluate each other on the structured triples, and screen the triples based on the evaluation results to obtain:

[0087] G = {E, R, F};

[0088] Among them, G represents the urban planning knowledge graph (knowledge base); E represents the entity set in the knowledge base; R represents the relationship set in the knowledge base; F represents the triple set stored in the knowledge graph;

[0089] The triple set F stored in the knowledge graph is expressed as:

[0090] F={(h,r,t)|h,t∈E,r∈R};

[0091] Among them, (h, r, t) represents a triple; h and t represent the head and tail entities respectively; r represents the relationship from the head entity to the tail entity.

[0092] like Figure 2 As shown, the specific steps include:

[0093] S51) Add a thinking chain CoT example to the prompt word template to guide reasoning, guide the large language model to refer to the natural language inference model NLI, expand each SPO triple into 3 complete short sentences as a hypothesis, and use the corresponding paragraph of the triple as the premise to determine whether the hypothesis is contained in the premise, for example:

[0094] Input text paragraph: Baiyun East Area: ... (previous text) maintain Baiyun Mountain Scenic Area and other ecological environment resources. (latter text) ....

[0095] The input text paragraph corresponds to the triple: <Baiyun East District, Weiyu, Baiyun Mountain Scenic Area>

[0096] The prompt model, as an objective and fair evaluation expert, carefully reads the context paragraph (i.e., the input text paragraph) and uses it as a premise. The corresponding triples are expanded, such as: 1. The Baiyun East District maintains the Baiyun Mountain Scenic Area; 2. The Baiyun Mountain Scenic Area is protected by the Baiyun East District; 3. The Baiyun East District of Guangzhou maintains the ecological environment resources of Baiyun Mountain.

[0097] Take three complete sentences as hypotheses and instruct the model to judge whether the hypothesis can be derived from the premise, that is, whether the premise implies the hypothesis, and give the answer: yes / no.

[0098] Finally, the command model combines the above reasoning results and comprehensively evaluates the SPO triples from five dimensions: semantics, consistency, authenticity, accuracy, and comprehensibility.

[0099] S52) instructing the large language models to evaluate the quality of the generated triples in a pairwise comparison and ranking manner, and collectively form a set of answers; that is, the large language model A determines which answer between the large language model B and the large language model C is better.

[0100] S53), instruct the large language model to evaluate the quality of the generated triples in a multi-dimensional scoring manner. Based on a set of results obtained in step S52), the large language model is used to score the generated triples in the five dimensions mentioned in step S51) with a score of 1 to 5. After the mutual evaluation is completed, the triples with too low scores are filtered out to obtain the final results.

[0101] S6) Construct a classifier based on the large language model, and use the classifier to classify the head and tail entities and relations of the triple set F stored in the knowledge graph, and use the classification results as the attributes of each component of the triple (h, r, t). Figure 3 As shown, the specific steps include:

[0102] S61), filter the entity types, locations, planning concepts, and land and energy use generated in step S3), select 20% as training data, obtain the vector representation of entity words through the text-embedding-3-large tool provided by openai, and store the entity words, corresponding vectors, and corresponding classifications in the vector database as an enhanced retrieval generation plug-in knowledge base;

[0103] S62), using the large language model GPT-4o as a classifier, inputting the entities to be classified in sequence, and at the same time, the text-embedding-3-large tool obtains the vector representation of the entity to be classified, calculates its cosine similarity with the existing records in the vector database, and returns the top five records with the highest similarity; instructing the large language model GPT-4o to use the labels of the returned records as a reference to assign labels to the entities to be classified;

[0104] S63) instructs the large language model to refine the classification results in step S62) based on general knowledge (categories such as commercial office, transportation, ecological agriculture and forestry, etc.).

[0105] After the above steps S1) to S6), the triples containing attributes are organized into JSON format files that are easy to parse and are commonly used for knowledge graph storage, and the planning knowledge graph for planning files based on the large language model is automatically constructed. The files can be imported into a graph database such as Neo4j for visual display and management. Organizing unstructured planning texts into a knowledge graph in the form of a directed graph consisting of nodes and edges can not only intuitively express planning information, but also be more conducive to its application to downstream tasks such as land use simulation, question-answering systems, and recommendation systems.

[0106] The above embodiments and descriptions are only for illustrating the principles and best embodiments of the present invention. Without departing from the spirit and scope of the present invention, the present invention may be subject to various changes and improvements, all of which fall within the scope of the present invention to be protected.

Claims

1. A method for constructing an urban planning knowledge graph based on a large language model, characterized in that: The following steps are involved: S1), preprocessing the urban planning text to obtain a preprocessed text; S2), constructing a prompt word template, through which the identity of the large language model is specified, the task steps are introduced, the thinking of the large language model is guided, and the large language model is stimulated to complete various NLP tasks; S3), using prompt word templates to guide the large language model to extract planning knowledge entities in urban planning texts; S4), based on the preprocessed text, the prompt word template and the named entity recognition result, instructing the large language model to organize the entities into structured triples that satisfy the text semantics; S5), guiding the large language model to evaluate the structured triples, and screening the triples based on the evaluation results; S6) Construct a classifier based on the large language model, and use the classifier to classify the head and tail entities and relations of the triple set stored in the knowledge graph, and use the classification results as the attributes of each component of the triple.

2. The method for constructing an urban planning knowledge graph based on a large language model according to claim 1, characterized in that: In step S1), the preprocessing includes sentence segmentation and segmentation processing of the urban planning text; and component syntactic analysis CON and dependency syntactic analysis DEP processing of the sentences.

3. The method for constructing an urban planning knowledge graph based on a large language model according to claim 2, characterized in that: In step S1), the urban planning text is divided into sentences and paragraphs, wherein sentences are used as basic processing units and paragraphs are used to provide contextual background information to help the large language model understand semantics.

4. The method for constructing an urban planning knowledge graph based on a large language model according to claim 3 is characterized in that: In step S1), the natural language processing library HanLP is used to perform syntactic analysis, the component syntactic analysis CON is used to express the compositional structural relationship between adjacent words in a sentence, and the dependency syntactic analysis DEP is used to express the hierarchical relationship between words in a sentence.

5. The method for constructing an urban planning knowledge graph based on a large language model according to claim 1, characterized in that: In step S2), the prompt word template includes five parts, namely: Task description: used to define the role identity of the large language model through system parameters, inform the large language model of the corresponding role through system parameters, and describe the task requirements through user parameters; Candidate target list: for named entity recognition tasks, triple evaluation tasks, and attribute classification tasks; Task example: used to guide a large language model to complete the mapping from input to output; Task emphasis: repeatedly emphasize matters needing attention and use a strong commanding tone to require the large language model to pay attention to and implement implementation details; Second round of dialogue: The initial results are verified and optimized through the second round of dialogue.

6. The method for constructing an urban planning knowledge graph based on a large language model according to claim 1, characterized in that: In step S3), the prompt word template is used to guide the large language model to extract planning knowledge entities in the urban planning text, which specifically includes the following steps: S31), taking the sentence as the processing unit, constructing five groups of task examples in the form of sentence and corresponding entity recognition results, and adding the five groups of examples to the task description; S32), under the guidance of the prompt word template, the sentences to be processed are input in sequence, so that the large language model recognizes the named entities of the urban planning text; S33) re-inputting the entity results obtained in step S32) into the large language model, and verifying and re-extracting the results through the large language model.

7. The method for constructing an urban planning knowledge graph based on a large language model according to claim 1, characterized in that: In step S4), based on the preprocessed text, the prompt word template and the named entity recognition result, the large language model is instructed to organize the entities into structured triples that satisfy the text semantics, which specifically includes the following steps: S41), using the preprocessed text of step S1) and the named entity recognition results of each sentence of step S3) as inputs of the large language model, and using the recognized entities as candidate targets of the organizational planning knowledge triples; S42), using the prompt word template to guide the input data of the large language model step S41) to learn contextual semantics, and organize the entities into SPO triples that conform to the semantic relationship of subject, predicate, and object of the text data.

8. The method for constructing an urban planning knowledge graph based on a large language model according to claim 1, characterized in that: In step S5), the large language model is guided to evaluate the structured triples, and the triples are screened based on the evaluation results, which specifically includes the following steps: S51), adding a thinking chain CoT example to the prompt word template to guide reasoning, guiding the large language model to refer to the natural language inference model NLI, expanding each SPO triple into 3 complete short sentences as a hypothesis, taking the corresponding paragraph of the triple as a premise, and judging whether the hypothesis is implied in the premise; Take three complete sentences as hypotheses and instruct the model to judge whether the hypothesis can be derived from the premise, that is, whether the premise implies the hypothesis, and give the answer: yes / no; Finally, the command model combines the above reasoning results to comprehensively evaluate the SPO triples from five dimensions: semantics, consistency, authenticity, accuracy, and comprehensibility; S52), instructing the large language models to evaluate the quality of the generated triples in a pairwise comparison and ranking manner, and collectively forming a set of answers; S53), instruct the large language model to evaluate the quality of the generated triples in a multi-dimensional scoring manner. Based on a set of results obtained in step S52), the large language model is used to score the generated triples in the five dimensions in step S51) with a score of 1 to 5. After the mutual evaluation is completed, the triples with low scores are filtered out to obtain the final result.

9. The method for constructing an urban planning knowledge graph based on a large language model according to claim 8, characterized in that: In step S5), the screening result is expressed as: G = {E, R, F}; Among them, G represents the urban planning knowledge graph (knowledge base); E represents the entity set in the knowledge base; R represents the relationship set in the knowledge base; F represents the triple set stored in the knowledge graph; The triple set F stored in the knowledge graph is expressed as: F={(h,r,t)|h,t∈E,r∈R}; Among them, (h, r, t) represents a triple; h and t represent the head and tail entities respectively; r represents the relationship from the head entity to the tail entity.

10. The method for constructing an urban planning knowledge graph based on a large language model according to claim 9, characterized in that: In step S6), a classifier is constructed based on the large language model, and the head and tail entities and relations of the triple set F stored in the knowledge graph are classified by the classifier, and the classification results are used as the attributes of each component of the triple (h, r, t), which specifically includes the following steps: S61), select some entities generated in step S3), and obtain vector representations of entity words, and store the entity words, corresponding vectors, and corresponding classifications in a vector database as an enhanced retrieval generation plug-in knowledge base; S62), using the large language model as a classifier, inputting the entities to be classified in sequence, obtaining the vector representation of the entities to be classified, calculating the cosine similarity between the entities and the existing records in the vector database, and returning the top five records with the highest similarity; Instruct the large language model to use the labels of the returned records as a reference to assign labels to the entities to be classified; S63) instructing the large language model to refine the classification results in step S62) based on general knowledge.

Citation Information

Patent Citations

  • Knowledge graph construction method and device based on large model and medium

    CN117725995A

  • Power industry knowledge graph construction method fused with large-scale language model

    CN118627604A

  • Knowledge graph and large model fusion intelligent question answering method oriented to network security

    CN119474294A

  • Intelligent matching system with ontology-aided relation extraction

    US20180232443A1

Cited By

  • Intelligent distribution and high-efficiency execution method and system for regional transportation tasks

    CN120612029A

  • Urban update planning scheme optimization method and system driven by large language model, terminal and storage medium

    CN121303470A

  • A method, system, terminal, and storage medium for optimizing urban renewal planning schemes driven by a large language model.

    CN121303470B