Construction method and device of traffic field knowledge graph and medium thereof
By combining large language models and other deep learning models, the text data in the transportation field is deeply analyzed and processed, and the error problem of LLM when generating knowledge graphs in the transportation field is solved, achieving high accuracy and efficiency knowledge graph construction.
Patent Information
- Application Number
- CN202411871801.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, LLM may generate incorrect entity recognition and relationship extraction when generating knowledge graphs in the traffic field, resulting in inaccurate graphs.
By combining large language models and other deep learning models, text data in the transportation field is deeply analyzed and processed, entity annotation, relationship extraction and knowledge organization are realized, and knowledge graphs with rich semantic relationships are constructed.
It effectively solves the problems of inaccurate entity recognition and unclear relationships between entities, improves the accuracy and efficiency of the knowledge graph, and ensures the clarity and efficiency of the graph.
Smart Images

Figure CN120045720A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of knowledge graph construction, and more particularly to a method, device, and medium for constructing a knowledge graph in the transportation field. Background Art
[0002] The technology of using a large language model (LLM) to generate a knowledge graph is an important advancement in the field of artificial intelligence in recent years. This technology realizes the goal of automatically constructing a knowledge graph through the semantic understanding and generation capabilities of the LLM, significantly improving the construction efficiency and accuracy of the graph. When constructing a knowledge graph, the LLM usually collects information from domain corpora and open knowledge graphs, extracts relevant triples using the LLM, and forms a preliminary knowledge graph. Then, a validator checks the authenticity and reliability of the generated triples to ensure the accuracy of the graph. Finally, a pruning tool further controls the generation direction to reduce redundant information and make the knowledge graph clearer and more efficient. However, there are also some drawbacks in the existing technology: 1) The LLM may generate content that is inconsistent with facts or fictional when generating text, that is, the so-called "hallucination phenomenon", because the model may learn inaccurate or misleading information during the training process, resulting in errors when generating the knowledge graph; 2) The LLM may not accurately understand the context in complex or ambiguous sentences, which will affect the extraction of entities and relationships.
[0003] Therefore, it is a technical problem to be solved to provide a method that can solve the problems of inaccurate entity recognition and unclear relationships between entities. Summary of the Invention
[0004] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a method, device, and medium for constructing a knowledge graph in the transportation field. Through the combination of a large language model and other deep learning models, the text data in the transportation field is deeply analyzed and processed to achieve entity annotation, relationship extraction, and knowledge organization, and a knowledge graph with rich semantic relationships can be constructed.
[0005] The purpose of the present invention can be achieved by the following technical solutions:
[0006] According to a first aspect of the present invention, there is provided a method for constructing a knowledge graph in the transportation field, the method comprising:
[0007] Obtain a set of supporting text materials and perform data preprocessing to obtain a data set, wherein the data preprocessing includes data cleaning and large language model classification;
[0008] Construct a material knowledge graph based on the set of supporting text materials after data preprocessing, wherein the nodes of the material knowledge graph are texts and classification topics, and the edges are the association relationships between the nodes;
[0009] Extract each character or word of each text node based on the described material knowledge graph, and perform entity pre-annotation based on the character or word.
[0010] Use a large language model to extract entities and relationships between entities in the text in combination with the entity pre-annotation results, and perform entity alignment and relationship disambiguation.
[0011] Construct a knowledge graph based on the aligned entities and disambiguated relationships.
[0012] As a preferred technical solution, the method of preprocessing is as follows:
[0013] Perform data cleaning on the support text material set, that is, remove all texts containing garbled characters, incorrect codes, and incomplete content in the support material text set.
[0014] Use a large language model to extract the publication time, title, and main content of each text in the support text material set after data cleaning.
[0015] Determine the sequence relationship of text publication based on the comparison of timestamps based on the publication time, evaluate similarity based on the title, and perform semantic analysis based on the main content.
[0016] Combine the semantic analysis results, similarity, and sequence relationship to eliminate duplicate texts.
[0017] Classify the support text material set after removing duplicate texts based on the theme and main content to obtain a data set.
[0018] As a preferred technical solution, the method of constructing a material knowledge graph is as follows:
[0019] Perform unstructured processing on each text in the data set in units of paragraphs to form a semantic graph.
[0020] Use a large language model to generate a theme summary for each text based on the semantic graph, and perform text classification based on the theme summary in combination with a preset classification standard.
[0021] Construct a material knowledge graph based on the text classification results.
[0022] As a preferred technical solution, the method of performing entity pre-annotation is as follows:
[0023] Extract each character or word of the text node in the material knowledge graph.
[0024] Perform context-aware encoding on each character or word to generate a semantic vector representation, and its expression is:
[0025] {h 1 ,h2 ,..., h n}\ =\ BERT(w 1 , w 2 ,..., w n ),
[0026] where h n represents the semantic representation of the nth word, w n represents the nth character or word in the text node, and BERT represents the BERT operation;
[0027] Combine the previous and subsequent semantic representations, capture the hidden state of each semantic representation, and generate a hidden state sequence;
[0028] Define multiple prediction label sequences based on the preset entity classification types, and the elements of the prediction label sequences represent the entity classification types of the characters or words corresponding to the hidden states in the hidden state sequence;
[0029] Evaluate each prediction label sequence based on the hidden state sequence using a scoring function, and select the prediction label sequence with the highest score as the entity pre-annotation result, and its expression is:
[0030]
[0031] where X represents the hidden state sequence; y represents the prediction label sequence; y i represents the ith prediction label in the prediction label sequence, represents the ith hidden state, and n represents the number of elements in the hidden state sequence and the prediction label sequence; represents the weight matrix; represents the transition weight.
[0032] As a preferred technical solution, the preset entity classification types include: means of transportation, traffic facilities, traffic routes, organizations, traffic events, traffic signals, traffic regulations, geographical locations, time, mileage / distance, speed, cost, and traffic participants.
[0033] As a preferred technical solution, the method for entity alignment is: obtaining different descriptions of the same entity based on the entity pre-annotation result, and mapping the different descriptions to a unified standard.
[0034] As a preferred technical solution, the method for relationship disambiguation is: using a large language model to extract the descriptive words or characters of the relationships between entities, and combining the context where the descriptive words or characters are located. If it is a polysemous description of the same relationship, it is normalized to the same relationship descriptive word or character; if it is the same descriptive word or character for different relationships, it is disambiguated into different relationship descriptive words or characters.
[0035] As a preferred technical solution, the method for constructing a knowledge graph based on the aligned entities and disambiguated relationships is as follows:
[0036] Extract the aligned entities and create a node for each entity;
[0037] Construct edges for each node based on the disambiguated relationships to generate an initial knowledge graph;
[0038] Use the large language model to extract the attribute information of each entity based on the dataset and optimize the initial knowledge graph based on the attribute information;
[0039] Verify the consistency of the optimized initial knowledge graph. If it is consistent, visualize it to obtain the final knowledge graph; if it is not consistent, correct the optimized initial knowledge graph based on the verification result and verify again until it is consistent.
[0040] According to the second aspect of the present invention, there is provided an electronic device for constructing a knowledge graph in the transportation field, including a memory and a processor. A computer program is stored on the memory, and when the processor executes the program, the above method is implemented.
[0041] According to the third aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above method is implemented.
[0042] Compared with the prior art, the present invention has the following advantages:
[0043] 1) First, the present invention constructs a material knowledge graph based on the relationships between texts, pre-labels entities for each text node on the basis of the material knowledge graph, captures the hidden state of each word or phrase by combining the context, and performs entity pre-labeling based on the hidden state. Finally, on the basis of entity pre-labeling, a large language model is used for further entity recognition, which maximally solves the problems of inaccurate multi-entity recognition and unclear relationships between entities, and can also show high accuracy and efficiency when processing large-scale unstructured text data;
[0044] 2) In order to improve the accuracy of recognizing the relationships between entities, the present invention also performs entity standardization and alignment as well as relationship disambiguation, which can effectively reduce problems such as entity redundancy and relationship confusion, and ensure the accuracy of the constructed knowledge graph;
[0045] 3) For the constructed knowledge graph, the present invention also uses a large language model to supplement semantic hierarchical attributes, so that the knowledge graph can not only reflect the relationships between entities, but also express the detailed features of each entity and relationship, providing more comprehensive knowledge. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is the flowchart of the method of the present invention;
[0047] Figure 2 is the technical roadmap of the present invention. Specific embodiments
[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0049] This embodiment provides a method for constructing a knowledge graph in the transportation field based on a large language model and text data. First, the large language model is used to perform topic analysis and classification on the literature, then the entities in the text are labeled and extracted, and finally a knowledge graph with rich semantic relationships is constructed to provide support for decision-making in the transportation field.
[0050] Specifically, the process of this method is as follows Figure 1 shown, and technical filling is carried out on the basis of the technology of Figure 1 to form a technical roadmap as shown in Figure 2 shown, which specifically includes the following steps:
[0051] S1. Obtain a set of supporting text materials and perform data preprocessing to obtain a data set. Among them, the data preprocessing includes data cleaning and large language model classification.
[0052] S11. Determine a series of specific keywords related to the transportation field, such as "traffic flow", "public transportation efficiency", "traffic safety management", etc., and formulate a systematic retrieval strategy to conduct a systematic literature search in multiple academic databases, aiming to cover the latest research results in this field; by accessing the government public platform, obtain relevant policy documents and statistical data to ensure the authority and accuracy of the information; use a search engine to conduct a wide literature search to collect industry reports and other relevant materials, and construct a set of supporting material texts based on the above collected materials.
[0053] S12. Clean the text in the set of supporting material texts according to its format, and remove all texts containing garbled characters, incorrect codes, and incomplete content in the set of supporting material texts.
[0054] S13. Use the large language model to extract the publication time, title, and main content of each text in the set of supporting text materials after data cleaning.
[0055] S14. Determine the chronological relationship of text publication by comparing timestamps based on the publication time, evaluate similarity based on the title, and perform semantic analysis on the main content.
[0056] S15. Combine the semantic analysis results, similarity, and chronological relationship to eliminate duplicate texts. This process effectively eliminates redundant data and improves the uniqueness of the dataset.
[0057] S16. Classify the set of supporting text materials after eliminating duplicate texts based on the theme and main content in combination with their specific application background in the transportation field to obtain the dataset.
[0058] Step S1 provides systematic content organization for constructing the knowledge graph. Through the cleaning and classification processes of this step, the quality of the data is significantly improved, ensuring the scientificity and rigor of subsequent analysis.
[0059] S2. Construct a material knowledge graph based on the set of supporting text materials after data preprocessing.
[0060] S21. Unstructurally process each text in the dataset in paragraphs to form a semantic graph. Specifically, define the dataset as S = {m 1 , m 2 , …, m n}, which contains multiple texts m i , and for each text, m i = {p 1 , p 2 , …, p n} contains multiple paragraphs p j .
[0061] The unstructured processed data contains a large number of complex fact and relationship networks and can be abstracted into a semantic graph. In this semantic graph, entities and relationships represent the interrelated facts and information in the document.
[0062] S22. Use a large language model to generate a theme summary for each text based on the semantic graph and perform text classification based on the theme summary in combination with a preset classification standard.
[0063] This step also uses a large language model for semantic analysis to ensure accurate discrimination of the theme categories of the documents. Specifically, the model will divide the texts into appropriate theme categories according to the semantic similarity of the text content and the classification criteria. This classification process not only simplifies the structure of the material set but also lays a foundation for subsequent knowledge organization.
[0064] S23. Construct a material knowledge graph based on the text classification results. In the material knowledge graph, use edges to connect text m iWith its affiliated subject classification structure, a directed graph is formed. The nodes in this graph represent specific documents and classification topics, and the edges express the association relationship between documents and classifications. In this way, the material knowledge graph provides a systematic semantic navigation structure for the entire material set.
[0065] S3. Based on the material knowledge graph, extract each word or character of each text node, and use the combined model BERT-BiLSTM-crf to perform entity pre-annotation based on the word or character.
[0066] S31. Extract each word or character of the text nodes in the material knowledge graph, denoted as X = {w 1 , w 2 ,..., w n}, where w i represents the i-th word or character in the text.
[0067] S32. Use the BERT model to generate semantic vector representations through the multi-layer Transformer structure for each word or character. Its expression is:
[0068] {h 1 , h 2 ,..., h n} = BERT(w 1 , w 2 ,..., w n ),
[0069] where h n represents the semantic representation of the n-th word, w n represents the n-th word or character in the text node, and BERT represents the BERT operation.
[0070] S33. Use the BiLSTM model to learn the temporal information in the semantic vectors by combining the front and back semantic representations through the forward and backward LSTM modules to capture the hidden state of each semantic representation, generate a hidden state sequence, and output the hidden layer state by combining the forward and backward hidden states of each word or character.
[0071] S34. Define multiple prediction label sequences based on the preset entity classification types, and the elements of the prediction label sequences represent the entity classification types of the words or characters corresponding to the hidden states in the hidden state sequence. The prediction label sequence is denoted as: y = {y 1 , y 2 ,..., y n}, where y n represents the n-th prediction label.
[0072] Among them, the preset entity classification types include:
[0073] 1) Means of transportation: Refers to various transportation vehicles, such as "sedan", "subway", "airplane".
[0074] 2) Transportation facilities: Include roads, bridges, airports, stations, etc., such as "expressway", "flyover".
[0075] 3) Transportation routes: Involve specific road names, flight routes, railway lines, etc., such as "Beijing-Tibet Expressway", "Beijing-Shanghai High-Speed Railway".
[0076] 4) Organizations: Refers to traffic management departments, transportation companies, etc., such as "China National Railway Group Co., Ltd.", "XX Municipal Transportation Commission".
[0077] 5) Traffic events: Include traffic accidents, traffic control, traffic jams and other events, such as "rear-end accident", "traffic control measures".
[0078] 6) Traffic signals: Involve traffic signs, signal lights, etc., such as "red light", "speed limit sign".
[0079] 7) Traffic regulations: Include traffic rules, laws and regulations, etc., such as "Regulations on Punishment for Traffic Violations".
[0080] 8) Geographical locations: Geographical location descriptions specific to the transportation field, such as "city center", "port".
[0081] 9) Time: Time points or time periods related to transportation, such as "peak hours", "departure time".
[0082] 10) Mileage / distance: Involve the measurement of distance, such as "10 kilometers", "500 meters".
[0083] 11) Speed: Involve the description of speed, such as "speed limit 60 km / h".
[0084] 12) Fees: Transportation-related fees, such as "toll fee", "ticket price".
[0085] 13) Traffic participants: Refers to pedestrians, drivers, passengers, etc., such as "pedestrian", "driver".
[0086] S35. Input the hidden state sequence into the CRF layer. The CRF evaluates each predicted label sequence through a defined scoring function and selects the predicted label sequence with the highest score as the entity pre-annotation result. Its expression is:
[0087]
[0088] Among them, X represents the hidden state sequence; y represents the predicted label sequence; y i represents the i-th predicted label in the predicted label sequence. represents the i-th hidden state, n represents the number of elements in the hidden state sequence and the predicted label sequence; represents the weight matrix; represents the transfer weight.
[0089] S4. Use a large language model combined with entity pre-annotation results to extract entities in the text and the relationships between entities, and perform entity alignment and relationship disambiguation.
[0090] Directly using a large language model to extract entities in the field of transportation may bring certain limitations. First, the field of transportation has a wealth of professional terms and specific entity types, which are often not fully understood and accurately distinguished by general large language models. Second, the text in the field of transportation usually contains complex contexts and multi-level relationships, such as traffic events, management organizations and facilities. Failure to directly classify through annotations may lead to semantic deviations in the model's understanding, thereby affecting the accuracy of entity classification. In addition, since the construction of a transportation knowledge graph requires high accuracy and consistency, the results directly extracted by a large language model alone are often difficult to meet the high-precision requirements of this field. Using the pre-annotated results as an auxiliary input together with the text to be extracted provides accurate recognition clues for the large language model, significantly improving the effect of entity extraction. This annotation assistance first improves the accuracy of recognition, ensuring that the model avoids deviations in the classification of professional terms and entities in the field of transportation, especially when dealing with complex multi-level semantic relationships, such as the relationship between traffic events and traffic facilities. The guidance of the annotation results not only helps the model accurately understand the context structure of the text, but also enables it to more efficiently process specific entities and relationships in the field of transportation, ensuring that the model achieves higher accuracy and consistency in classification and relationship recognition. In addition, the annotation results are pre-screened to remove redundant information, making the extraction results of the large language model more refined, ensuring that the knowledge graph has high information purity and low noise during the construction process, and ultimately providing accurate and targeted support for the construction of knowledge graphs in the transportation field.
[0091] In detail, the relationships between entities include:
[0092] 1) Relationship between transportation and other entities
[0093] Traveling on: transportation tool - transportation route (e.g. "car" traveling on "Beijing-Tibet Expressway").
[0094] Stop at: transportation tool - transportation facility (e.g. "subway" stops at "subway station").
[0095] Use: Transportation - Traffic signals / traffic regulations (e.g. "Cars" obey "speed limit signs").
[0096] Charge: Means of transportation - Expenses (e.g., a "sedan" needs to pay "toll fees").
[0097] Managed by: Means of transportation - Organization (e.g., "buses" are managed by the "Transportation Committee").
[0098] 2) Relationships between transportation facilities and other entities
[0099] Located in: Transportation facilities - Geographical location (e.g., the "highway" is located in the "city center").
[0100] Connected to: Transportation facilities - Transportation routes (e.g., the "flyover" connects the "Jingzang Expressway" and the "Beijing-Shanghai High-Speed Railway").
[0101] Managed by: Transportation facilities - Organization (e.g., the "airport" is managed by the "XX City Transportation Committee").
[0102] 3) Relationships between transportation routes and other entities
[0103] Passes through: Transportation routes - Geographical location (e.g., the "Beijing-Shanghai High-Speed Railway" passes through the "city center").
[0104] Contains: Transportation routes - Transportation facilities (e.g., the "Jingzang Expressway" contains "toll stations").
[0105] Affected by: Transportation routes - Traffic incidents (e.g., the "Beijing-Shanghai High-Speed Railway" is closed due to "traffic control").
[0106] 4) Relationships between organizations and other entities
[0107] Issued by: Organization - Traffic signals / Traffic regulations / Expenses (e.g., the "Transportation Committee" issues "speed limit signs").
[0108] Managed by: Organization - Means of transportation / Transportation facilities / Transportation routes (e.g., the "XX City Transportation Committee" manages "public transportation").
[0109] Responds to: Organization - Traffic incidents (e.g., the "Transportation Bureau" responds to "rear-end accidents").
[0110] 5) Relationships between traffic incidents and other entities
[0111] Occurs at: Traffic incidents - Transportation facilities / Transportation routes / Geographical location (e.g., a "rear-end accident" occurs on the "highway").
[0112] Occurs during: Traffic incidents - Time (e.g., "traffic jams" occur during "peak hours").
[0113] Involves: Traffic incidents - Traffic participants / Means of transportation (e.g., a "rear-end accident" involves a "sedan" and a "pedestrian").
[0114] 6) Relationships between traffic signals and other entities
[0115] Applicable to: Traffic signal - Vehicle / Traffic route (e.g., "Speed limit sign" is applicable to "Sedans" on the "Beijing-Tibet Expressway").
[0116] Issued at: Traffic signal - Traffic facility / Location (e.g., "Red light" is set up at the "Intersection" in the "City center").
[0117] 7) Relationships between traffic regulations and other entities
[0118] Stipulates: Traffic regulation - Speed / Fee / Traffic participant / Vehicle (e.g., "Traffic violation punishment regulations" stipulate that "Pedestrians" must abide by "Red lights").
[0119] 8) Relationships between locations and other entities
[0120] Includes: Location - Traffic facility / Traffic route (e.g., "City center" includes "High-speed sections").
[0121] Distance: Location - Mileage / Distance (e.g., The "City center" is 10 kilometers away from the "Port").
[0122] 9) Relationships between time and other entities
[0123] Restricts: Time - Vehicle / Traffic event (e.g., "Peak hours" restrict the passage of "Sedans").
[0124] Applicable to: Time - Fee (e.g., Ticket prices are discounted during "Night hours").
[0125] 10) Relationships between mileage / distance and other entities
[0126] Restricts: Mileage / Distance - Vehicle (e.g., "Heavy vehicles" are prohibited from driving within "500 meters").
[0127] 11) Relationships between speed and other entities
[0128] Restricts: Speed - Vehicle / Traffic route (e.g., "Speed limit 60 km / h" restricts "Sedans" on the "Highway").
[0129] 12) Relationships between fee and other entities
[0130] Applicable to: Fee - Vehicle / Traffic route / Traffic facility (e.g., "Toll fee" is applicable to "Private cars" on the "Highway").
[0131] 13) Relationships between traffic participants and other entities
[0132] Ride: Traffic participants - means of transportation (e.g., "passengers" ride on "subways").
[0133] Occurs in: Traffic participants - traffic incidents (e.g., "pedestrians" encounter "rear-end accidents")
[0134] Specifically, entity alignment and relation disambiguation include:
[0135] S41. Obtain different descriptions of the same entity based on the entity pre-annotation results and map the different descriptions to a unified standard. This process helps to clearly define the category of each entity and the information represented by the entity. For example, "XX City Transportation Commission" and "Beijing Transportation Commission" are mapped to a unified standard. The semantic understanding ability of the large language model in dealing with synonyms, abbreviations, name variants, etc. helps to identify and merge entities with the same semantics. Further, through the semantic similarity analysis ability of the model, entity alignment is performed based on features such as the context information and geographical location of the entity to ensure the accurate mapping of polysemous words in a specific context, thus effectively reducing redundancy.
[0136] S42. Use the large language model to extract descriptive words or characters of the relationships between entities, and combine the context where the descriptive words or characters are located. If it is a polysemous description of the same relationship, it is normalized to the same relationship descriptive word or character; if it is the same descriptive word or character for different relationships, disambiguation is performed to different relationship descriptive words or characters. For example, similar relationship expressions such as "travel on" and "pass by" are unified as the "pass through" relationship through the semantic understanding of the model. In addition, for the processing of polysemous relationships, the large language model can combine context information to distinguish the meanings of the same relationship in different contexts, so as to accurately locate the specific reference of the relationship. For example, "pass by" can refer to the passage of a means of transportation or the location where an event occurs in different scenarios, and the model ensures the uniqueness of the relationship expression through semantic disambiguation.
[0137] S5. Construct a knowledge graph based on the aligned entities and disambiguated relationships.
[0138] S51. Extract the aligned entities and create a node for each entity.
[0139] S52. Construct edges for each node based on the disambiguated relationships to generate an initial knowledge graph.
[0140] S53. Use a large language model to extract the attribute information of each entity based on the dataset, and optimize the initial knowledge graph based on the attribute information, where the attribute information includes: the type of transportation vehicle, speed limit, etc., the occurrence time and location of traffic events, and the background information between entities. For example, the attribute information about "highway" can not only include its name and location, but also introduce detailed data in other dimensions such as "number of lanes" and "construction year" to improve the richness and accuracy of the graph. It can not only reflect the relationships between entities, but also express the detailed characteristics of each entity and relationship, providing more comprehensive knowledge.
[0141] S54. Verify the consistency of the initial knowledge graph after multiple iterations of optimization to ensure the structural rationality and logical consistency of all nodes and edges. The role of the large language model in this stage is to identify potential logical conflicts or redundant information through model reasoning, so as to further adjust the layout and expression of entities and relationships.
[0142] If it is consistent, visualize it to obtain the final knowledge graph; if it is not consistent, correct the optimized initial knowledge graph based on the verification result and verify it again until it is consistent.
[0143] In addition, for the constructed knowledge graph, this embodiment also uses Neo4j for visualization. As a graph database, Neo4j has powerful graph data storage and query capabilities, and can intuitively display each entity and relationship in the form of nodes and edges. Through the graphical interface of Neo4j, it is convenient to browse, query and analyze the graph to obtain key information and insights in the traffic field. Through Neo4j, the knowledge graph is not only efficient in storage and query, but also can display the complex relationship network between entities in a graphical way. In addition, Neo4j supports the dynamic update and interactive query of the graph, ensuring the real-time and accuracy of the knowledge graph in the traffic field. After being integrated into the question answering system, the knowledge graph based on Neo4j can provide decision support for traffic management, policy making, academic research, etc., and can quickly obtain relevant information in the traffic field.
[0144] This embodiment also provides an electronic device for constructing a knowledge graph in the traffic field. The electronic device includes a central processing unit (CPU), which can execute various appropriate actions and processes according to the computer program instructions stored in the read-only memory (ROM) or the computer program instructions loaded from the storage unit into the random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other through a bus. The input / output (I / O) interface is also connected to the bus.
[0145] Multiple components in the device are connected to the I / O interface, including: an input unit, such as a keyboard, a mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a disk, an optical disc, etc.; and a communication unit, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0146] The processing unit executes the various methods and processes described above, such as methods S1 to S5. For example, in some embodiments, methods S1 to S5 may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of methods S1 to S5 described above may be executed. Alternatively, in other embodiments, the CPU may be configured to execute methods S1 to S5 by any other suitable means (e.g., by means of firmware).
[0147] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, by way of non-limitation, exemplary types of hardware logic components that may be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0148] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0149] In the context of the present invention, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0150] As described above, the foregoing are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily conceive of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for constructing a knowledge graph in the field of transportation, characterized in that: The method includes: Acquire a set of supporting text materials and perform data preprocessing to obtain a data set, wherein the data preprocessing includes data cleaning and large language model classification; Construct a material knowledge graph based on the supporting text material set after data preprocessing, wherein the nodes of the material knowledge graph are texts and classification topics, and the edges are the associations between the nodes; Based on the material knowledge graph, extract each word or phrase from each text node, and perform entity pre-labeling based on the word or phrase; Use a large language model combined with entity pre-labeling results to extract entities in the text and the relationships between entities, and perform entity alignment and relationship disambiguation; Build a knowledge graph based on the aligned entities and disambiguated relations.
2. The method for constructing a knowledge graph in the field of transportation according to claim 1, characterized in that: The pretreatment method is: Performing data cleaning on the supporting text material set, that is, removing all texts containing garbled characters, wrong characters, and incomplete content in the supporting material text set; Use the big language model to extract the release time, title and main content of each text in the supporting text material set after data cleaning; Performing a timestamp comparison based on the release time to determine the order of text release, evaluating similarity based on the title, and performing a semantic analysis based on the main content; Combine semantic analysis results, similarity and order to eliminate duplicate texts; Based on the themes and main contents, the supporting text material set after removing duplicate texts is classified to obtain a data set.
3. The method for constructing a knowledge graph in the field of transportation according to claim 1, characterized in that: The method for constructing a material knowledge graph is: Each text in the data set is subjected to unstructured processing in units of paragraphs to form a semantic graph; Generate a topic summary of each text based on the semantic graph using a large language model, and classify the text based on the topic summary combined with a preset classification standard; Build a material knowledge graph based on text classification results.
4. The method for constructing a knowledge graph in the field of transportation according to claim 1, characterized in that: The method for pre-marking entities is as follows: Extract each word or phrase of the text node in the material knowledge graph; Context-aware encoding is performed on each character or word to generate a semantic vector representation, which is expressed as follows: {h1,h2,...,h n }=BERT(w1,w2,...,w n ), Among them, h n represents the semantic representation of the nth word, w n Represents the nth word or term in a text node, and BERT represents the BERT operation; Combine the previous and next semantic representations, capture the hidden state of each semantic representation, and generate a hidden state sequence; Defining a plurality of prediction label sequences based on a preset entity classification type, wherein the elements of the prediction label sequences represent entity classification types of characters or words corresponding to hidden states in the hidden state sequence; Based on the hidden state sequence, each predicted label sequence is evaluated using a score function, and the predicted label sequence with the highest score is selected as the entity pre-labeling result, and its expression is: Among them, X represents the hidden state sequence; y represents the predicted label sequence; y i represents the i-th predicted label in the predicted label sequence, represents the i-th hidden state, n represents the number of elements in the hidden state sequence and the predicted label sequence; represents the weight matrix; represents the transfer weight.
5. The method for constructing a knowledge graph in the field of transportation according to claim 4, characterized in that: The preset entity classification types include: transportation tools, transportation facilities, transportation routes, institutional organizations, traffic events, traffic signals, traffic regulations, geographic locations, time, mileage / distance, speed, fees and traffic participants.
6. The method for constructing a knowledge graph in the field of transportation according to claim 4, characterized in that: The entity alignment method is: obtaining different descriptions of the same entity based on the entity pre-labeling results, and mapping the different descriptions to a unified standard.
7. The method for constructing a knowledge graph in the field of transportation according to claim 1, characterized in that: The method for relationship disambiguation is: using a large language model to extract descriptive words or descriptive characters of the relationship between entities, combining the context in which the descriptive words or descriptive characters are located, if it is a polysemous description of the same relationship, normalizing it into the same relationship descriptive word or descriptive character; if it is the same descriptive word or descriptive character of different relationships, disambiguation processing is performed into different relationship descriptive words or descriptive characters.
8. The method for constructing a knowledge graph in the field of transportation according to claim 1, characterized in that: The method for constructing a knowledge graph based on aligned entities and disambiguated relationships is: Extract the aligned entities and create a node for each entity; Build edges for each node based on the disambiguated relationships to generate an initial knowledge graph; Extracting attribute information of each entity based on the data set using the large language model, and optimizing the initial knowledge graph based on the attribute information; Verify the consistency of the optimized initial knowledge graph. If it is consistent, visualize it to obtain the final knowledge graph. If it is inconsistent, modify the optimized initial knowledge graph based on the verification results and verify it again until it is consistent.
9. An electronic device for constructing a knowledge graph in the field of transportation, comprising a memory and a processor, wherein a computer program is stored in the memory, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Traffic knowledge graph construction method and device, electronic equipment and storage medium
CN114676203A
Urban traffic knowledge graph construction method
CN115344707A
Energy industry knowledge graph construction method and device based on multi-source heterogeneous data fusion technology
CN117313849A
Knowledge graph construction method and device for electric power operation text, medium and chip
CN118469006A
Cited By
Electric power information query method and device based on knowledge graph, equipment, medium and product
CN121255735A