An urban brain-based industry classification driven graph retrieval enhancement method

By introducing the industry classification-driven method of the City Brain into the graph retrieval, performing character fragment decomposition and first character combination, and combining the overlapping features of tags and fields, a stable path set is constructed. This solves the problems of coarse tag recognition granularity and unstable path structure in the existing technology, and achieves more efficient graph data filtering and path accuracy.

CN122489694APending Publication Date: 2026-07-31数字金华技术运营有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
数字金华技术运营有限公司
Filing Date
2026-04-20
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies fail to perform character structure decomposition and combination standardization operations when processing industry classification labels, resulting in coarse label recognition granularity and misclassification. The lack of an overlap comparison mechanism between labels and fields leads to inaccurate matching results. Path recognition does not combine connection direction and node continuity for structural stability judgment, easily generating semantically interrupted or skipped path chains. The processing methods at the end-result expression level are simplistic, resulting in repetitive link expressions or logical fragmentation, affecting the efficiency of graph data filtering and path accuracy.

Method used

By obtaining the industry classification table and business tag fields from the City Brain resource list catalog, character fragment segmentation and initial character combination are performed to generate industry affiliation description mapping groups; graph nodes are extracted and overlap rate comparison is performed to generate a tag integration node list; a path clustering structure sequence is constructed based on the connection direction field, and path combinations with complete structures are selected; query terms and path descriptions are matched to generate tag-guided query path clusters, and deduplication and node chain concatenation are combined with description fields to enhance the logical closure of the path structure.

Benefits of technology

It improves the accuracy of label grouping and discrimination, the stability of node selection, and the precision of path structure expression, thereby improving the overall efficiency of map data selection and path accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122489694A_ABST
    Figure CN122489694A_ABST
Patent Text Reader

Abstract

This invention relates to the field of graph retrieval enhancement technology, specifically a graph retrieval enhancement method driven by industry classification based on a city brain. The method includes acquiring classification and label fields, performing character segmentation and overlap detection, generating attribution mapping groups, extracting graph nodes and constructing a comparison relationship to generate an integrated list, aggregating path directions to generate a structural sequence, matching labels and keywords to generate query path clusters, and splicing end information to generate an enhanced path set. This invention introduces character fragment decomposition and initial character grouping mechanisms to improve the standardization of label expression, combines overlapping features to achieve fine mapping of field attribution, enhances node recognition consistency based on field comparison, integrates path direction and node quantity to construct a stable structure, and utilizes keyword position matching to strengthen semantic association. It also demonstrates how deduplication and link splicing improve structural closure and expression accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of graph retrieval enhancement technology, and in particular to a graph retrieval enhancement method driven by industry classification based on the city brain. Background Technology

[0002] The field of knowledge graph retrieval enhancement technology encompasses methods that combine knowledge graph representation with information retrieval techniques. The core concept is to use knowledge graphs to structurally represent entities and their relationships, enabling information resources to be organized as nodes and edges. This helps computers better understand and reason about semantic relationships across different domains. Knowledge graphs, as an important semantic knowledge base, are widely used in big data analysis, intelligent question-answering systems, recommendation systems, and other scenarios. They can improve the accuracy and reliability of retrieval systems by enhancing semantic understanding and reasoning capabilities. Retrieval enhancement technology combines knowledge graphs with large language models, effectively reducing the "illusion" problem that occurs during language model generation, improving the accuracy of generated content, enhancing its ability to grasp new knowledge, and increasing the transparency and credibility of generated results by displaying the sources of retrieved information.

[0003] Among them, the industry-classification-driven graph retrieval enhancement method based on the City Brain refers to optimizing traditional retrieval enhancement techniques by combining the City Brain's industry classification system with knowledge graph technology to address complex cross-industry and multi-dimensional query problems. It encompasses technical aspects such as industry-classification-driven graph index construction, binding industry classification tags to graph entities, and knowledge graph-based semantic matching and reasoning extension. Specifically, it uses the industry classification system as a guide for retrieval, based on entity relationships and semantic networks within the graph, to retrieve highly relevant contextual information from industry classifications and integrate it with the generative model. This enhances the retrieval and generation capabilities of the large language model in complex query scenarios, ensuring more accurate answer generation.

[0004] Existing technologies fail to perform character structure decomposition and combination standardization operations when processing industry classification labels, resulting in coarse label recognition granularity and easy classification errors. In the process of graph node recognition, no overlap comparison mechanism between labels and fields is introduced, and the matching results lack accuracy when there are naming deviations. In the path recognition stage, structural stability judgment is not combined with connection direction and node continuity, which easily generates path chains with semantic interruptions or jumps. At the end result expression level, the processing method for explanatory information is simple, and structural deduplication and field concatenation operations are not performed, resulting in problems such as repeated link expressions or logical fragmentation. Overall, it affects the filtering efficiency and path accuracy of graph data under complex queries. Summary of the Invention

[0005] To address the technical problems existing in the prior art, this invention provides an industry-classification-driven graph retrieval enhancement method based on a city brain. The technical solution is as follows:

[0006] An industry-classification-driven graph retrieval enhancement method based on the city brain includes the following steps:

[0007] S1: Obtain the industry classification table, business tag field and belonging scenario field from the City Brain resource list directory. Match the first character group of the classification table with the tag field. Calculate the frequency of occurrence of the matching group in combination with the scenario field. Filter the most frequent groups and map them to the belonging content to generate industry belonging description mapping groups.

[0008] S2: Call the tag field in the industry affiliation description mapping group, extract the corresponding node in the graph, obtain the name field and the affiliated tag field and perform an overlap rate comparison, filter and re-identify nodes with consistent source field affiliation scenarios, and generate a list of tag integration nodes;

[0009] S3: Based on the tag-integrated node list, construct node pair combinations, extract the connection direction field of adjacent paths and perform continuity judgment, count the number of intermediate nodes and sort them in ascending order, filter the path combinations with complete structures, and generate a path clustering structure sequence.

[0010] S4: Based on the starting field of the path clustering structure sequence, match the tag content of the industry affiliation description mapping group, extract the descriptive terms in the path and compare their order with the query terms, filter the path set with reasonable structure, and generate a tag-guided query path cluster.

[0011] As a further embodiment of the present invention, the industry affiliation description mapping group includes a combination of classification tag letters, a set of similar tag fields, and a group of high-frequency scene tags; the tag integration node list includes a name mapping field, a tag aggregation field, and a node affiliation identifier; the path aggregation structure sequence includes a directional continuous segment, a path node quantity value, and a structurally stable path group; and the tag-guided query path cluster includes a tag starting field group, a keyword matching sequence, and a path stability sequence.

[0012] As a further aspect of the present invention, the step of obtaining the industry affiliation description mapping group is as follows:

[0013] S101: Based on the industry management classification table fields registered in the City Brain Resource List Catalog, perform character segmentation operation for each field content, extract continuous Chinese characters or letter combinations using delimiter recognition, further call the extracted fragment content of each group, obtain the first character of each fragment and form a character combination, construct a character first character group that corresponds one-to-one with the field content, and generate a character first group set.

[0014] S102: Based on the business tag field and each character combination in the character first group set, extract all substring combinations respectively, call the character combination and business tag field value to perform character fragment comparison operation, compare the character overlap length with the total length of the character combination, calculate the character overlap rate value, and determine whether it is a similar tag matching item based on the condition that the character overlap rate value is greater than the set character overlap judgment threshold, and generate tag similarity matching group;

[0015] S103: Call the field group in the tag similarity matching group, perform tag affiliation mapping and marking operation according to its corresponding affiliation scene field, perform counting statistics on the scene field corresponding to each field group, filter the top five field groups with the highest frequency of occurrence as high-frequency tag groups, and establish industry affiliation description mapping group.

[0016] As a further aspect of the present invention, the step of obtaining the tag integration node list is as follows:

[0017] S201: Based on all the tag fields in the industry affiliation description mapping group, obtain the node set in the graph data. For each node in the node set, extract its name field and auxiliary label field. Call the name field and auxiliary label field to perform character overlap calculation, calculate the character overlap length of the two fields, and compare the overlap length with the total number of characters to obtain the field overlap data.

[0018] S202: Based on the field overlap data, perform a screening operation on the overlap rate of the source field and the affiliated tag field for the belonging scene value. Calculate the overlap rate of the belonging scene value of the source field and the tag field of each node, and filter out inconsistent nodes according to the preset threshold of the belonging scene value overlap rate, and retain the nodes that meet the conditions to obtain a set of consistent nodes.

[0019] S203: Call the field data in the consistent node set, construct a field correspondence mapping table between nodes, mark all nodes in the node set that meet the conditions, and after marking the correspondence between nodes and fields in the mapping table, generate a list of tagged integrated nodes.

[0020] As a further aspect of the present invention, the step of obtaining the path convergence structure sequence is as follows:

[0021] S301: Based on the node combination in the tag integration node list, obtain the path field between adjacent nodes in the graph, extract the direction field of each connection in the path, perform character value aggregation operation for each continuous connection direction field in the path, merge the characters of the continuous direction fields, generate the connection sequence of each path, and obtain the connection sequence data.

[0022] S302: Based on the connection sequence data, perform a counting operation on the number of intermediate nodes between the start and end nodes of each path, calculate the number of intermediate nodes for each path, and sort all paths in ascending order by the number of intermediate nodes to obtain node number sorting data.

[0023] S303: Based on the number of nodes, sort the data, filter the path combinations with stable structural continuity in the path sorting, remove the unstable paths, and retain the path combinations with stable structural continuity to construct the final path set and generate a path aggregation structure sequence.

[0024] As a further aspect of the present invention, the step of obtaining the tag-guided query path cluster set is as follows:

[0025] S401: Based on the starting point field in the path aggregation structure sequence, extract the starting point field in the path and the tag description field in the industry affiliation description mapping group, call the tag description field and the path starting point field, perform a character value equality judgment operation, retain the path node set with consistent field content, and generate a path set with consistent field content.

[0026] S402: Based on the path description terms in the consistent path set, call the query input keywords, perform the term sequence position arrangement operation, compare the order and continuity of the keyword characters in the path, eliminate non-continuous path term combinations, obtain paths that meet the continuous matching conditions, and obtain a continuous matching path group.

[0027] S403: Call the intermediate node count field in the continuous matching path group, perform a size comparison operation based on the node count value and construct an ascending sort sequence, filter path groups whose node count fluctuation range does not exceed two nodes based on the position of the sort sequence, extract the path data in which the node and label combination are consistent, and generate a label-guided query path cluster.

[0028] As a further aspect of the present invention, the method further includes:

[0029] S5: Extract the end node name and description fields based on the tag-guided query path cluster, detect and remove duplicate character combinations, reorganize the remaining fields into a description chain, classify the paths according to connection integrity, and generate a graph retrieval enhanced path set;

[0030] The enhanced path set for graph retrieval includes a set of end node names, a group of concatenated description fields, and a bidirectional connection path chain.

[0031] As a further aspect of the present invention, the step of obtaining the enhanced path set for map retrieval is as follows:

[0032] S501: Based on all paths in the tag-guided query path cluster, extract the name information field and the auxiliary description field of the terminal node in each path, call the continuous character group in the auxiliary description field, perform the duplicate content detection operation, identify the number of times the character group appears repeatedly in the path set, and remove the character group that appears more than once to obtain the non-duplicate character group data.

[0033] S502: Call the character content in the non-repeating character group data, perform field value concatenation operation according to the original character order, construct the character chain between nodes in each path, and at the same time retain the bidirectional connection field content of the first and last nodes in the path node combination chain, and generate a bidirectional connection structure chain according to the structural order of the concatenated path.

[0034] S503: Based on the path combinations in the bidirectional connection structure chain, group them according to the node connection integrity field, calculate the number of complete nodes in the node combination chain within each path group, sort them in descending order according to the integrity field value, filter out path combinations with continuous and uninterrupted structures, and establish a graph retrieval enhanced path set.

[0035] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:

[0036] In this invention, the standardization of tag structure expression is improved by introducing a character fragment decomposition and first character combination mechanism for the classification field. The fine mapping of field affiliation is achieved by combining the overlap features between the first character group of the tag and the classification. The consistency of node recognition is improved by relying on the character overlap and word count comparison between the name field and the subordinate tags. A stable path set is constructed by integrating the continuity of the path connection direction and the sorting analysis of the number of intermediate nodes. The consistent matching between the path and the query semantics is strengthened by using the arrangement relationship between keywords and descriptive terms. The logical closure of the terminal structure is enhanced by combining the deduplication of the description field and the node chain splicing method. Overall, the stability of tag grouping and node filtering and the accuracy of path structure expression are improved. Attached Figure Description

[0037] Figure 1 This is a flowchart of the method of the present invention;

[0038] Figure 2 This is a flowchart illustrating the process of obtaining the industry affiliation description mapping group for this invention.

[0039] Figure 3 This is a flowchart illustrating the process of obtaining the list of integrated tags in this invention.

[0040] Figure 4 This is a flowchart illustrating the process of obtaining the path convergence structure sequence of the present invention.

[0041] Figure 5 This is a flowchart illustrating the process of obtaining the tag-guided query path cluster set in this invention.

[0042] Figure 6 This is a flowchart illustrating the process of obtaining the enhanced path set for map retrieval in this invention. Detailed Implementation

[0043] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0044] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0045] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0046] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0047] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0048] Please see Figure 1 This invention provides a technical solution: an industry-classification-driven graph retrieval enhancement method based on the city brain, comprising the following steps:

[0049] S1: Obtain the industry management classification table, business tag field and belonging scenario field registered in the city brain resource list directory. Perform character segmentation on each content in the classification table field and retain the first character group. Perform character overlap detection and similarity judgment on the business tag field and the segmented character group. Perform tag belonging mapping mark on the matching field group. Perform matching group frequency statistics on the belonging scenario field and filter the group with the highest occurrence frequency to generate industry belonging description mapping group.

[0050] S2: Call all the marked fields in the industry attribution description mapping group, extract the set of nodes that match the fields in the graph data and extract the name field and the affiliated label field, calculate the character overlap of the two fields and compare the number of words, perform the attribution scenario value overlap rate screening on the source field and the label field and filter out inconsistent nodes, retain the field comparison relationship to build a mapping table and mark the node set, and generate a list of integrated label nodes;

[0051] S3: Based on the node combination in the tag integration node list, retrieve the path field between adjacent nodes in the graph and extract the connection direction field, aggregate the character values ​​of the continuous connection direction field in the path and construct the connection sequence, count the number of intermediate nodes in the path of the start and end nodes and sort them in ascending order, filter the path combination with stable structural continuity in the sorting to construct the path set, and generate the path aggregation structure sequence.

[0052] S4: Based on the starting point field in the path clustering structure sequence, extract the field that is consistent with the industry affiliation description mapping group label description, perform position matching on the description terms in the path and the query input keywords and filter out discontinuous combinations, perform a comparison operation on the number of intermediate nodes in the retained path set and rearrange the sequence, filter the stable paths corresponding to the tags and nodes according to the rearrangement results, and generate a tag-guided query path cluster set;

[0053] S5: Based on all paths in the tag-guided query path cluster set, extract the end node name information field and the auxiliary description field, perform duplicate detection on consecutive character groups in the description field and filter out duplicate structures, concatenate the characters in the remaining fields in order and retain the bidirectional connection information of the node combination chain, and group and rearrange the path chain set according to the node connection integrity to generate the graph retrieval enhanced path set.

[0054] The industry attribution description mapping group includes classification tag letter combinations, similar tag field sets, and scenario high-frequency tag groups. The tag integration node list includes name mapping fields, tag aggregation fields, and node attribution identifiers. The path aggregation structure sequence includes directional continuous segments, path node quantity values, and structurally stable path groups. The tag-guided query path cluster set includes tag starting field groups, keyword matching sequences, and path stability sequences. The graph retrieval enhanced path set includes end node name sets, description field splicing groups, and bidirectional connection path chains.

[0055] Please see Figure 2 The steps to obtain the industry affiliation description mapping group are as follows:

[0056] S101: Based on the field of the industry management classification table registered in the urban brain resource inventory directory, perform character segment division operations on the content of each field. Adopt the delimiter recognition method to extract continuous Chinese characters or letter combinations. Further call each group of segment content after extraction, obtain the first character of each segment, and form a character combination, construct a character first-character group corresponding to the field content one by one, and generate a character first-group set;

[0057] Before executing the solution, a database needs to be constructed first. That is, execute the docker pull neo4j:5.11 command to obtain the official image of Neo4j version 5.11 from Docker Hub. In the startup command for creating a container, map the plugin directory containing the APOC core library apoc-5.11.0-core.jar to the / plugins path of the container by means of volume mounting. At the same time, use the -e NEO4J_AUTH=neo4j / your_strong_password parameter to initialize the authentication information of the database, complete the creation and security configuration of the database instance. After the database instance is ready, start constructing the knowledge graph. Import the industry classification data into the graph database to generate the knowledge graph. That is, process the industry classification data into the format of the neo4j graph database. For the names of each level of industry categories, use the vector model to extract vector features, and then import the processed industry classification data into the neo4j graph database to obtain the knowledge graph;

[0058] Refine and generate the core nodes and labels from the original industry classification data, integrate and verify these nodes, and form a standardized data list (urban brain resource inventory) that can be imported into the graph database. Based on the fields of the industry management classification table registered in the inventory directory, this table includes field ID, field content, and attribution scenarios, as shown in Table 1 specifically. For the content of each field, for example, "Intelligent Transportation - Bus Route Planning" with field ID 001, perform character segment division operations. This operation is based on a preset delimiter set (including "-", "_", ""), and divides the field content into multiple character segments, that is, "Intelligent Transportation" and "Bus Route Planning". From each divided character segment, that is, "Intelligent Transportation" and "Bus Route Planning", extract its first Chinese character, which are "智" and "公" respectively. Then combine the multiple extracted first characters in the order of their appearance in the original field content to form a character combination, that is, "智公". Associate this character combination with the original field ID "001". Finally, gather the character combinations generated after the above processing of all field contents, such as "智公", "公视", "环监", "医保", etc., to construct a character first-character group that forms a one-to-one correspondence with the original field content, and generate a character first-group set.

[0059] Table 1: Industry Management Classification Table

[0060]

[0061] As shown in Table 1, this table lists some industry management classification fields in the urban brain resource inventory, providing basic data for the subsequent generation of the first characters of characters.

[0062] S102: According to each character combination in the business label field and the character first group set, extract all its substring combinations respectively, call the character combination and the business label field value to perform a character segment comparison operation, compare the character overlap length with the total length of the character combination, calculate the character overlap rate value, and take the condition that the character overlap rate value is greater than the set character overlap determination threshold to determine whether it is a similar label matching item, and generate a label similarity matching group;

[0063] According to the business label field, such as "intelligent bus", and the character combination in the character first group set, such as "zhi gong" from field ID "001", extract all the substring combinations of both. For "zhi gong", its substring combinations are {"zhi", "gong", "zhi gong"}, and for "intelligent bus", its substring combinations are {"zhi", "hui", "gong", "jiao", "zhi hui", "hui gong", "gong jiao", "zhi hui gong", "hui gong jiao", "intelligent bus"}. Call the character combination "zhi gong" and the business label field value "intelligent bus" to perform a character segment comparison operation, specifically to find the longest common substring between the two sets of substring combinations of the two strings. Here, the longest common substrings are "zhi" and "gong", and their total length is 2. Then, calculate the character overlap rate value, and the calculation method of this value is the ratio of the character overlap length to the total length of the character combination. , substitute the values for calculation. , compare the calculated character overlap rate value of 1.0 with the set character overlap determination threshold. This threshold is determined by testing a data set containing 5000 labeled samples. In the test, the threshold increases from 0.1 to 0.9 in steps of 0.1, and the precision and recall rate under each threshold are evaluated. Finally, 0.7 corresponding to the highest F1 score is selected as the character overlap determination threshold. Since the calculated overlap rate value of 1.0 is greater than the character overlap determination threshold of 0.7, it is determined that the business label "intelligent bus" and the character combination "zhi gong" are similar label matching items. Finally, all matching pairs that meet this condition, such as ("intelligent bus", "zhi gong"), ("public security video", "gong shi"), etc., are gathered to generate a label similarity matching group.

[0064] S103: Call the field groups in the label similarity matching group, perform label attribution mapping marking operations according to their corresponding attribution scenario fields, perform count statistics on the scenario fields corresponding to each field group, screen the top five field groups with the highest occurrence frequency as the high-frequency label groups, and establish an industry attribution description mapping group;

[0065] The system calls up tag similarity matching groups, such as a set containing matching items ("Smart Bus", "Smart Public Transport"), ("Bus Planning", "Smart Public Transport"), and ("Traffic Video", "Public Security Vision"). Based on the original field ID corresponding to each character combination in the matching group, it backtracks to the industry management classification table shown in Table 1 to find its corresponding belonging scenario field. For example, "Smart Public Transport" corresponds to field ID "001", and its belonging scenario is "Transportation", while "Public Security Vision" corresponds to field ID "002", and its belonging scenario is "Public Safety". For each matching item, a tag belonging mapping operation is performed, associating the business tag with the belonging scenario of its matching character combination to form a (business tag, belonging scenario) mapping pair, such as ("Smart Bus", "Transportation"), ("Bus Planning", "Transportation"), and ("Traffic Video", "Public Safety"). Then, all mapping pairs are processed by... The scenarios are grouped, and the number of labeled field groups under each scenario is counted. For example, in the "Transportation" scenario, the number of field groups counted is 2 (from "Smart Bus" and "Bus Planning"), and in the "Public Safety" scenario, the number of field groups counted is 1 (from "Traffic Video"). Finally, the count values ​​of all scenarios are sorted in descending order, and the top five field groups with the highest frequency are selected as high-frequency label groups. Assuming the statistical results are "Transportation" frequency 58, "Public Safety" frequency 45, "Environmental Protection" frequency 32, "Healthcare" frequency 28, "Urban Management" frequency 21, and "Education and Tourism" frequency 15, then the five scenarios of "Transportation", "Public Safety", "Environmental Protection", "Healthcare" and "Urban Management" and their corresponding field groups are selected to establish industry-specific description mapping groups.

[0066] Please see Figure 3 The steps to obtain the list of tag integration nodes are as follows:

[0067] S201: Based on all the tag fields in the industry affiliation description mapping group, obtain the node set in the graph data. For each node in the node set, extract its name field and auxiliary label field. Call the name field and auxiliary label field to perform character overlap calculation, calculate the character overlap length of the two fields, and compare the overlap length with the total number of characters to obtain the field overlap data.

[0068] Based on all the labeled fields in the industry-specific description mapping group, such as "Intelligent Transportation - Bus Route Planning", the set of nodes related to these fields is obtained from the graph database. Assuming node A is obtained, its name field is "Bus Route Planning" and its attached label field is "Intelligent Transportation". For the name field "Bus Route Planning" and the attached label field "Intelligent Transportation" of node A, the character overlap is calculated. This calculation process is as follows: the strings of the two fields are converted into character sets, that is, the character set of the name field is {"public", "transport", "line", "road", "plan", "plan"}, and the character set of the attached label field is {"intelligent", "smart", "transport", "connection"}. Then the intersection of the two sets is calculated, that is, {"intersection"}. The number of elements in the intersection is the character overlap length. In this example, the overlap length is 1. Finally, the overlap length 1 is compared with the total number of characters in the two fields (the name field length is 6, and the attached label field length is 4) to obtain the field overlap data, that is, (node ​​A, overlap length 1, total number of characters 10). This process is repeated for each node in the node set to obtain the field overlap data of all nodes.

[0069] S202: Based on the field overlap data, perform a screening operation on the overlap rate of the source field and the affiliated label field for the belonging scenario value. Calculate the overlap rate of the belonging scenario value of the source field and the label field of each node, and filter out inconsistent nodes according to the preset belonging scenario value overlap rate threshold, and retain the nodes that meet the conditions to obtain a set of consistent nodes.

[0070] Based on the field overlap data, such as (Node A, overlap length 1, total number of characters 10), and the source field (i.e., the field content in the industry management classification table, such as "Smart Transportation - Bus Route Planning") and the attached tag field (such as "Smart Transportation") corresponding to each node, query the corresponding belonging scenario value of these two fields in the industry belonging description mapping group. Assuming that the belonging scenario of "Smart Transportation - Bus Route Planning" is "Transportation," and the tag "Smart Transportation" also maps to the "Transportation" scenario, perform a belonging scenario value overlap rate filtering operation for each node. The overlap rate is defined as follows: if the belonging scenario values ​​of the source field and the tag field are completely identical, the overlap rate is 1.0; otherwise, it is 0. In this example, the belonging scenario of both fields is "Transportation." The element A has the keyword "transportation," so its scenario overlap rate is 1.0. The calculated overlap rate is then compared to a preset scenario overlap rate threshold of 0.9. This threshold is based on a test of 1000 graph nodes. The results show that when the threshold is 0.9 (i.e., requiring complete scenario consistency), the selected nodes have the highest accuracy in subsequent applications, reaching 95%. Since node A's overlap rate of 1.0 is greater than the threshold of 0.9, this node is retained. Conversely, if another node B's source field belongs to the scenario "transportation," while its tag field belongs to the scenario "urban management," its overlap rate is 0, less than 0.9, and this node will be filtered out. Finally, all nodes that pass the filtering, such as node A, are retained to obtain a consistent node set.

[0071] S203: Call the field data in the consistent node set, construct a field mapping table between nodes, mark all nodes in the node set that meet the conditions, and after marking the correspondence between nodes and fields in the mapping table, generate a list of tagged integrated nodes;

[0072] The consistent node set obtained in step S202 is called, for example, a set containing nodes A, B, and C. The field data of each node in the set is extracted, including the node ID, node name field, associated tag field, and source field. For example, the data of node A is (ID:'node_A', name:'bus route planning', tag:'intelligent transportation', source:'intelligent transportation-bus route planning'). Based on these field data, a mapping table is constructed for the field correspondence between nodes, as shown in Table 2. This table clearly lists the correspondence between each retained node and its related field content. Then, all nodes in the node set that meet the conditions are marked in this mapping table. The marking operation is to add a "verified" column to the mapping table and set the value of this column of all nodes in the consistent node set to "yes". Finally, after the marking operation of all nodes is completed and the field correspondence of all consistent nodes is fully recorded in the mapping table, a tag-integrated node list is generated.

[0073] Table 2: List of Tag Integration Nodes

[0074]

[0075] See Table 2, which shows the list of tag integration nodes generated after filtering and verification. It contains the explicit mapping relationship between the nodes and their core fields. The data in this list is the basis for the final import into the neo4j database to form the entities and relationships of the knowledge graph.

[0076] Please see Figure 4 The steps for obtaining the path clustering structure sequence are as follows:

[0077] S301: Based on the node combination in the tag integration node list, obtain the path field between adjacent nodes in the graph, extract the direction field of each connection in the path, perform character value aggregation operation for each continuous connection direction field in the path, merge the characters of the continuous direction fields, generate the connection sequence of each path, and obtain the connection sequence data.

[0078] After the knowledge graph is constructed, index building is required to achieve efficient retrieval. In the graph database, a query index is built for specific fields of the knowledge graph. Specifically, for the "Node Name" field in Table 2's "Label Integration Node List," the Cypher command `CREATE FULLTEXT INDEX nodeNameIndex FOR (n:Category) ON EACH[n.name]` is executed to create a full-text index. For the vector features generated for each node name in step 2-2 (assuming they are stored in the embedding attribute), the command `CREATE VECTOR INDEX nodeEmbeddingIndex FOR (n:Category) ON n.embedding OPTIONS` is executed. The vector index is created using `{indexConfig:{vector.dimensions:768,vector.similarity_function:'cosine'}}`. An API service is built using the Flask framework and neo4j-driver library in Python. This service provides a ` / query` endpoint that receives text and vectors, and internally calls the corresponding index for querying. After the index is built, the knowledge graph can be retrieved. That is, before using the large language model to answer user queries, the user's query for the large language model is obtained, and it is checked whether there is a locally running vector model or an online vector model service interface. If not, or if the interface call fails, a vector model instance is created locally, and the built index query service is called to query the knowledge graph in combination with the vector model.

[0079] By integrating node combinations in the generated tag list, such as node A (bus route planning) and its adjacent node D (transportation hub), the path field connecting these two nodes is obtained from the graph data. Assuming the path contains a connection from A to D with the character value "contains in", the character value aggregation operation is performed on each consecutive connection direction field in the path. If after A to D, D connects to E (city center area) with the connection direction "located in", then the two consecutive direction fields "contains in" and "located in" are merged to generate the connection sequence of the path, i.e., "contains in, located in". This operation is repeated for all paths between node combinations in the list to obtain a set of connection sequence data for each path, such as {"contains in, located in", "associated with, impacted", "monitored, reported"}, generating connection sequence data.

[0080] S302: Based on the connection sequence data, perform a counting operation on the number of intermediate nodes between the start and end nodes of each path, calculate the number of intermediate nodes for each path, and sort all paths in ascending order by the number of intermediate nodes to obtain the node count sorted data.

[0081] Based on the node combinations in the tag-integrated node list, such as node A (bus route planning) and its adjacent node D (transportation hub), the path field connecting these two nodes is obtained from the graph data. Assuming that the path contains a connection from A to D, and the character value of its direction field is "contained in", a character value aggregation operation is performed on each consecutive connection direction field in the path. If after A to D, D connects to E (city center area), and the connection direction is "located in", then the two consecutive direction fields "contained in" and "located in" are merged to generate the connection sequence of the path, i.e., "contained in, located in". This operation is repeated for all paths between node combinations in the list to obtain a set of connection sequence data for each path, such as {"contained in, located in", "associated with, impacted", "monitored, reported"}, generating connection sequence data.

[0082] S303: Sort the data based on the number of nodes, filter the path combinations with stable structural continuity in the path sorting, remove the unstable paths, and retain the path combinations with stable structural continuity to construct the final path set and generate a path clustering structure sequence.

[0083] Based on the obtained node count sorting data, i.e., [(P1,1),(P3,1),(P4,2),(P2,3)], we filter path combinations with stable structural continuity in the path sorting. The criterion for structural continuity is: in the sorted data, the absolute value of the difference in the number of intermediate nodes of continuous paths is not greater than 1. For example, P1 and P3 both have 1 intermediate node, with a difference of 0, satisfying the condition; P3 and P4 have 1 and 2 intermediate nodes respectively, with a difference of 1, satisfying the condition; and P4 and P2 have 2 and 3 intermediate nodes respectively, with a difference of 1, also satisfying the condition. Therefore, the initial judgment for (P1,P3,P4,P2) is... If a structurally continuous combination exists with another set of sorted data [(P5,2),(P6,5)] whose difference in the number of nodes is 3, which is greater than 1, then path P6 in this combination is considered structurally unstable and is removed. Among the path combinations that meet the above conditions, structurally unstable paths are removed. In this example, since all paths meet the continuity condition, no path is removed. Finally, all path combinations that are determined to be structurally continuous and stable after filtering, namely [(P1,1),(P3,1),(P4,2),(P2,3)], are retained to construct the final path set and generate the path clustering structure sequence.

[0084] Please see Figure 5 The steps to obtain the tag-guided query path cluster are as follows:

[0085] S401: Based on the starting field in the path aggregation structure sequence, extract the starting field in the path and the tag description field in the industry affiliation description mapping group, call the tag description field and the path starting field, perform a character value equality judgment operation, retain the path node set with consistent field content, and generate a path set with consistent field content.

[0086] For each path in the path aggregation structure sequence, such as path P1, its starting point field, namely the name of node A, "bus route planning," is extracted. Simultaneously, the label description field, such as "smart bus," is extracted from the industry affiliation description mapping group established in step S103. A character value equality check is performed between the label description field "smart bus" and the path starting point field "bus route planning." This operation compares the characters in the two strings one by one; the check is true only when the two strings are completely identical. In this example, "smart bus" and "bus route planning" are not equal, so path P1 is not retained. If the starting point field of another path P5 is "smart bus," then it is completely consistent with the label description field "smart bus," and the check is true; path P5 is retained. Finally, all path nodes whose starting point field is completely consistent with a certain label description field are aggregated to generate a set of paths with consistent fields.

[0087] S402: Based on the path description entries in the field-consistent path set, call the query input keywords, perform an operation to arrange the sequence positions of the entries, compare the order positions and continuity status of the keyword characters appearing in the path, eliminate the path entry combinations with non-consecutive arrangements, obtain the paths that meet the continuous matching conditions, and get the continuous matching path group;

[0088] Based on the path P5 included in the field-consistent path set, its path description entries are ("Smart Bus", "Connect", "Transportation Hub", "Impact", "Road Congestion"). At the same time, call the user's query input keywords, assumed to be "Traffic Congestion". For the description entry sequence of path P5, perform an operation to arrange the sequence positions of the entries. This operation aims to check the order and continuity of the characters in the query keyword appearing in the path entries. Specifically: Split the query keyword "Traffic Congestion" into characters "Jiao", "Tong", "Yong", "Du", and then search for these characters in the path description entries ("Smart Bus", "Connect", "Transportation Hub", "Impact", "Road Congestion"). It is found that "Jiao" and "Tong" appear in "Transportation Hub", and "Yong" and "Du" appear in "Road Congestion", but not in the continuous form of "Traffic Congestion". Therefore, this path does not meet the continuous matching conditions. If the description entries of another path P6 are ("Traffic Analysis", "Prediction", "Traffic Congestion Situation"), then the keyword "Traffic Congestion" appears continuously as a complete entry, meeting the conditions. Then, eliminate the path entry combinations with non-consecutive arrangements, such as P5, and retain the paths that meet the continuous matching conditions, such as P6, obtain the paths that meet the continuous matching conditions, and get the continuous matching path group.

[0089] S403: Call the intermediate node quantity field in the continuous matching path group, perform a size comparison operation according to the node quantity value and construct an ascending sorted list, filter the path group with the node quantity fluctuation range not exceeding two nodes based on the sorted list position, extract the path data with consistent node and label combinations, and generate a label-oriented query path cluster set;

[0090] Call paths P6 and P7 in the continuous matching path group, and extract the number of intermediate nodes of these two paths. Assuming that P6 has 2 intermediate nodes and P7 has 3 intermediate nodes, perform a size comparison operation on all paths in the path group according to their "number of intermediate nodes" value, and construct an ascending sorted sequence to obtain [P6(2),P7(3)]. Based on the position of this sorted sequence, filter out path groups whose node number fluctuation range does not exceed two nodes. The fluctuation range is calculated as the difference between the maximum number of nodes and the minimum number of nodes in the group. In this example, the fluctuation is 3-2=1, which does not exceed the set fluctuation range threshold. The value is 2, so the path group [P6,P7] is retained. If there is another path P8 with 5 nodes, the fluctuation after adding it to the group becomes 5-2=3, which is greater than 2, so P8 will be excluded. Finally, in the selected path group [P6,P7], the path data with consistent node and label combinations are extracted. That is, it is checked whether the starting label and internal node type of these paths belong to the same preset category. If the starting label of P6 and P7 are both "traffic analysis" and the intermediate nodes are both of the "data analysis model" type, then it is determined that the combination is consistent. Finally, these path data are aggregated to generate a label-guided query path cluster.

[0091] Please see Figure 6 The steps for obtaining the enhanced path set for graph retrieval are as follows:

[0092] S501: Based on the tag-guided query path cluster set, extract the name information field and the auxiliary description field of the terminal node in each path, call the continuous character group in the auxiliary description field, perform the duplicate content detection operation, identify the number of times the character group appears repeatedly in the path set, and remove the character group that appears more than once to obtain the non-duplicate character group data.

[0093] Based on the tag-guided query path cluster set, all paths are included. For example, path P6 has the terminal node "Traffic Congestion Status". The name information field "Traffic Congestion Status" and the auxiliary description field of the terminal node are extracted. Assuming that the auxiliary description field is "This status is derived from the analysis of traffic flow data of urban main roads, and duplicate data has been removed", the continuous character groups in the auxiliary description field, such as "urban main roads", "traffic flow data", and "data analysis", are called to perform a duplicate content detection operation. This operation compares the content of the auxiliary description field of each character group with the content of all other paths in the path cluster set and counts the total number of occurrences. Assuming that "urban main roads" appears 3 times in the entire cluster set, while "traffic flow data" appears 1 time, then the character groups with more than 1 occurrence are removed. Here, "urban main roads" is removed because it appears 3 times, while "traffic flow data" is retained because it appears 1 time. All the character groups that are not removed, i.e., the non-duplicate character groups, are gathered together to obtain the non-duplicate character group data.

[0094] S502: Call the character content in the non-repeating character group data, perform field value concatenation operation according to the original character order, construct the character chain between nodes in each path, and at the same time retain the bidirectional connection field content of the first and last nodes in the path node combination chain, and generate a bidirectional connection structure chain according to the structural order of the concatenated path.

[0095] The process involves calling non-repeating character set data, such as the non-repeating character set {“traffic flow data”} corresponding to path P6, and non-repeating character sets from other nodes in the path. Assuming the node combination chain of path P6 is (Traffic Analysis -> Data Model -> Traffic Congestion Status), its corresponding non-repeating character sets are {“Traffic Flow Prediction”}, {“Regression Analysis Model”}, and {“Traffic Flow Data”}. Following the original order of the nodes in the path, the process performs field value concatenation on the contents of these character sets to construct a character chain between nodes, namely “Traffic Flow Prediction Regression Analysis Model Traffic Flow Data”. Simultaneously, it retains the bidirectional connection field content of the first and last nodes in the path node combination chain. Assuming the starting point “Traffic Analysis” has an input connection field “Data Source”, and the ending point “Traffic Congestion Status” has an output connection field “Warning Signal”, these two fields are retained. Based on the structural order of the concatenated path, i.e., (Input field: “Data Source”, Concatenated character chain: “Traffic Flow Prediction Regression Analysis Model Traffic Flow Data”, Output field: “Warning Signal”), a bidirectional connection structure chain is generated.

[0096] S503: Based on the path combinations in the bidirectional connection structure chain, group them according to the node connection integrity field, calculate the number of complete nodes in the node combination chain within each path group, sort them in descending order according to the integrity field value, filter out path combinations with continuous and uninterrupted structures, and establish a graph retrieval enhanced path set.

[0097] The bidirectional connection structure chain set generated by step S502, for example, includes the structure chains corresponding to paths P6 and P7. It is grouped according to the node connection integrity field of each path. This field is a Boolean value, which is determined by checking whether there is a broken connection in the path. If all nodes in the path are directly or through intermediate nodes, the integrity is true. Assuming that the integrity of P6 is true and that of P7 is false, P6 is assigned to the "complete group" and P7 is assigned to the "incomplete group". In each group, the number of complete nodes in the node combination chain in each group is calculated. For P6 in the "complete group", the number of nodes is 3. This step is not applicable to the paths in the "incomplete group". Then, only the paths in the "complete group" are sorted in descending order according to the number of complete nodes. Assuming that there is also path P8 in the "complete group", the number of nodes is 4, then the sorting is [P8(4),P6(3)]. Finally, from this sorting list, all structurally continuous and uninterrupted path combinations are selected, that is, all paths with the integrity field being true, and a graph retrieval enhanced path set is established.

[0098] Table 3: Enhanced Path Sets for Graph Retrieval

[0099]

[0100] As shown in Table 3, this table lists the final set of graph retrieval enhancement paths, which is the final retrieval result and provides structured and highly relevant contextual information for subsequent large language model generation.

[0101] The information in the "Graph Retrieval Enhanced Path Set" shown in Table 3, such as the bidirectional connection structure chain content of path P6, is formatted into a descriptive text: "Based on the knowledge graph, data source-driven traffic analysis can predict traffic congestion through regression analysis models and issue early warning signals. Its core analysis chain is 'traffic flow prediction regression analysis model vehicle flow data'." This text is used as a prefix (part of the Prompt) and concatenated with the user's original question, "How to predict and alleviate traffic congestion?" The complete Prompt is then sent to a large language model (such as GPT-4). When generating the answer, the model will provide a more professional answer containing specific technical paths and logical relationships based on this precise context.

[0102] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for enhancing graph search based on industry classification driven by urban brain, characterized in that, Includes the following steps: S1: Obtain the industry classification table, business tag field and belonging scenario field from the City Brain resource list directory. Match the first character group of the classification table with the tag field. Calculate the frequency of occurrence of the matching group in combination with the scenario field. Filter the most frequent groups and map them to the belonging content to generate industry belonging description mapping groups. S2: Call the tag field in the industry affiliation description mapping group, extract the corresponding node in the graph, obtain the name field and the affiliated tag field and perform an overlap rate comparison, filter and re-identify nodes with consistent source field affiliation scenarios, and generate a list of tag integration nodes; S3: Based on the tag-integrated node list, construct node pair combinations, extract the connection direction field of adjacent paths and perform continuity judgment, count the number of intermediate nodes and sort them in ascending order, filter the path combinations with complete structures, and generate a path clustering structure sequence. S4: Based on the starting field of the path clustering structure sequence, match the tag content of the industry affiliation description mapping group, extract the descriptive terms in the path and compare their order with the query terms, filter the path set with reasonable structure, and generate a tag-guided query path cluster. 2.The Urban Brain based industry classification driven graph search enhancement method of claim 1, wherein: The industry affiliation description mapping group includes a combination of classification tag letters, a set of similar tag fields, and a group of high-frequency scenario tags. The tag integration node list includes a name mapping field, a tag aggregation field, and a node affiliation identifier. The path aggregation structure sequence includes a continuous directional segment, a number of path nodes, and a structurally stable path group. The tag-guided query path cluster includes a tag starting field group, a keyword matching sequence, and a path stability sequence. 3.The Urban Brain based industry classification driven graph search enhancement method of claim 1, wherein, The steps for obtaining the industry affiliation description mapping group are as follows: S101: Based on the industry management classification table fields registered in the City Brain Resource List Catalog, perform character segmentation operation for each field content, extract continuous Chinese characters or letter combinations using delimiter recognition, further call the extracted fragment content of each group, obtain the first character of each fragment and form a character combination, construct a character first character group that corresponds one-to-one with the field content, and generate a character first group set. S102: Based on the business tag field and each character combination in the character first group set, extract all substring combinations respectively, call the character combination and business tag field value to perform character fragment comparison operation, compare the character overlap length with the total length of the character combination, calculate the character overlap rate value, and determine whether it is a similar tag matching item based on the condition that the character overlap rate value is greater than the set character overlap judgment threshold, and generate tag similarity matching group; S103: Call the field group in the tag similarity matching group, perform tag affiliation mapping and marking operation according to its corresponding affiliation scene field, perform counting statistics on the scene field corresponding to each field group, filter the top five field groups with the highest frequency of occurrence as high-frequency tag groups, and establish industry affiliation description mapping group. 4.The Urban Brain based industry classification driven graph search enhancement method of claim 1, wherein, The steps for obtaining the list of tag integration nodes are as follows: S201: Based on all the tag fields in the industry affiliation description mapping group, obtain the node set in the graph data. For each node in the node set, extract its name field and auxiliary label field. Call the name field and auxiliary label field to perform character overlap calculation, calculate the character overlap length of the two fields, and compare the overlap length with the total number of characters to obtain the field overlap data. S202: Based on the field overlap data, perform a screening operation on the overlap rate of the source field and the affiliated tag field for the belonging scene value. Calculate the overlap rate of the belonging scene value of the source field and the tag field of each node, and filter out inconsistent nodes according to the preset threshold of the belonging scene value overlap rate, and retain the nodes that meet the conditions to obtain a set of consistent nodes. S203: Call the field data in the consistent node set, construct a field correspondence mapping table between nodes, mark all nodes in the node set that meet the conditions, and after marking the correspondence between nodes and fields in the mapping table, generate a list of tagged integrated nodes. 5.The Urban Brain based industry classification driven graph search enhancement method of claim 1, wherein, The steps for obtaining the path convergence structure sequence are as follows: S301: Based on the node combination in the tag integration node list, obtain the path field between adjacent nodes in the graph, extract the direction field of each connection in the path, perform character value aggregation operation for each continuous connection direction field in the path, merge the characters of the continuous direction fields, generate the connection sequence of each path, and obtain the connection sequence data. S302: Based on the connection sequence data, perform a counting operation on the number of intermediate nodes between the start and end nodes of each path, calculate the number of intermediate nodes for each path, and sort all paths in ascending order by the number of intermediate nodes to obtain node number sorting data. S303: Based on the number of nodes, sort the data, filter the path combinations with stable structural continuity in the path sorting, remove the unstable paths, and retain the path combinations with stable structural continuity to construct the final path set and generate a path aggregation structure sequence. 6.The Urban Brain based industry classification driven graph search enhancement method of claim 1, wherein, The steps for obtaining the tag-guided query path cluster are as follows: S401: Based on the starting point field in the path aggregation structure sequence, extract the starting point field in the path and the tag description field in the industry affiliation description mapping group, call the tag description field and the path starting point field, perform a character value equality judgment operation, retain the path node set with consistent field content, and generate a path set with consistent field content. S402: Based on the path description terms in the consistent path set, call the query input keywords, perform the term sequence position arrangement operation, compare the order and continuity of the keyword characters in the path, eliminate non-continuous path term combinations, obtain paths that meet the continuous matching conditions, and obtain a continuous matching path group. S403: Call the intermediate node count field in the continuous matching path group, perform a size comparison operation based on the node count value and construct an ascending sort sequence, filter path groups whose node count fluctuation range does not exceed two nodes based on the position of the sort sequence, extract the path data in which the node and label combination are consistent, and generate a label-guided query path cluster.

7. The industry-classification-driven graph retrieval enhancement method based on the city brain according to claim 1, characterized in that, The method further includes: S5: Extract the end node name and description fields based on the tag-guided query path cluster, detect and remove duplicate character combinations, reorganize the remaining fields into a description chain, classify the paths according to connection integrity, and generate a graph retrieval enhanced path set; The enhanced path set for graph retrieval includes a set of end node names, a group of concatenated description fields, and a bidirectional connection path chain.

8. The industry-classification-driven graph retrieval enhancement method based on the city brain according to claim 7, characterized in that, The steps for obtaining the enhanced path set for the graph retrieval are as follows: S501: Based on all paths in the tag-guided query path cluster, extract the name information field and the auxiliary description field of the terminal node in each path, call the continuous character group in the auxiliary description field, perform the duplicate content detection operation, identify the number of times the character group appears repeatedly in the path set, and remove the character group that appears more than once to obtain the non-duplicate character group data. S502: Call the character content in the non-repeating character group data, perform field value concatenation operation according to the original character order, construct the character chain between nodes in each path, and at the same time retain the bidirectional connection field content of the first and last nodes in the path node combination chain, and generate a bidirectional connection structure chain according to the structural order of the concatenated path. S503: Based on the path combinations in the bidirectional connection structure chain, group them according to the node connection integrity field, calculate the number of complete nodes in the node combination chain within each path group, sort them in descending order according to the integrity field value, filter out path combinations with continuous and uninterrupted structures, and establish a graph retrieval enhanced path set.