Enhanced road disease knowledge extraction method and system based on large model self-learning
The road disease knowledge map is constructed through the large language model BERT and self-learning mechanism, which solves the problem of complex semantic understanding in road disease detection, achieves high accuracy and efficient knowledge extraction, reduces the risk of error transmission, and improves the scientificity and efficiency of decision-making.
Patent Information
- Application Number
- CN202510353791.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art is difficult to effectively deal with complex professional terms and semantic understanding in road disease detection, resulting in recognition errors or omissions, and the knowledge graph construction process relies on cumbersome manual rules and insufficient generalization ability of machine learning algorithms.
The pre-trained large language model BERT is used for entity and relationship extraction, combined with hierarchical relationship coding rules and expert review mechanism, dynamically update the relationship intensity value through the self-learning mechanism, build a road disease knowledge graph, and use incremental learning and dynamic adjustment of weights to reduce the risk of error propagation.
It significantly improves the accuracy and decision-making efficiency of road disease knowledge extraction, successfully identify and correct 83.2% of semantic conflicts, and the false positive rate is less than 5%, supporting multidimensional relationship reasoning and rapid absorption of the latest research results.
Smart Images

Figure CN120338081A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of road disease detection and maintenance, and particularly to a method and system for enhancing road disease knowledge extraction based on large model self-learning. Background Art
[0002] With the rapid development of the road disease detection and maintenance field, a large amount of text data has been generated, such as standards in the road disease detection and maintenance field, scientific research papers, detection and maintenance reports, etc. These data contain rich knowledge about traffic facilities, diseases, treatment processes, etc. A knowledge graph is a knowledge representation and management tool under the background of big data, which describes semantic information and relationships in the form of nodes and edges, and is an intuitive and highly understandable knowledge representation and reasoning framework. Constructing a knowledge graph in the road disease field can integrate a large amount of scattered knowledge related to road diseases, and provide strong support for road maintenance decision-making, disease prediction, and prevention and control technology research and development.
[0003] The process of constructing a domain knowledge graph usually includes three steps: knowledge extraction, knowledge fusion, and knowledge processing, and the construction process usually relies on deep learning algorithms. Knowledge extraction methods mainly rely on rule-based pattern matching and simple machine learning algorithms. Rule-based methods require a large number of complex rules to be manually written. For different types of road disease entities and relationships, numerous and detailed rules need to be formulated, such as for disease entities like "cracks" and "potholes" and relationships like "disease - cause" and "disease - treatment". Although machine learning algorithms can handle some complex situations, they perform unsatisfactorily when dealing with complex professional terms and semantic understanding in the transportation field. For example, when identifying professional disease terms such as "pumping" and "bleeding", errors or omissions are likely to occur.
[0004] In recent years, large language models (LLMs) have been proposed as a type of language model with powerful language understanding, analysis, and generation capabilities. Their extremely strong language capabilities have brought an opportunity for the simplification and upgrade of knowledge graph construction methods. Large models can effectively solve the problem that the generalization ability of extraction methods based on rules or statistical learning is poor and they cannot adapt to new terms and complex semantics. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for enhancing road disease knowledge extraction based on large model self-learning to solve the problems raised in the above background art.
[0006] To achieve the above purpose, the present invention provides the following technical solution: A method for enhancing road disease knowledge extraction based on large model self-learning, including the following steps:
[0007] Performing entity and relationship extraction on standard documents in the road disease field through pre-training the large model BERT to generate initial knowledge triples;
[0008] Construct a decision table for road disease knowledge extraction, define a hierarchical relationship coding rule, and quantify the relationship strength value between entities. The relationship strength value is dynamically fused and calculated from the initial expert score, co-occurrence frequency driven by data, and attention weight.
[0009] Train a self-learning model based on the decision table to extract knowledge from papers and unstructured texts, and generate a weighted road disease knowledge graph.
[0010] Dynamically update the relationship strength value through an incremental learning mechanism. When the detected relationship strength deviation exceeds the threshold, trigger an expert review, update the decision table, and optimize the knowledge graph.
[0011] Preferably, the relationship strength quantification method includes: experts assign an initial strength value to typical entity relationship pairs and score based on specification clarity and historical cases; calculate the co-occurrence frequency score and context semantic attention weight score of entity pairs and dynamically fuse them.
[0012] Preferably, the hierarchical relationship coding rule includes: the parent class code represents the macro relationship type, such as "disease-treatment process", and the subclass code refines the specific relationship, such as "transverse crack - grooving and grouting", in the format of "parent class code.subclass serial number"; the coding table stores fields such as the head and tail nodes of entities, relationship strength, update time, and confidence.
[0013] Preferably, conflict detection and dynamic update include: when new data is input, calculate the deviation between the current relationship strength and the historical value. If the deviation is higher than the weight parameter, trigger an expert review.
[0014] Preferably, the knowledge graph is applied to the recommendation of road disease treatment plans, specifically including: sorting and recommending priorities according to the relationship strength value, and the plan with a strength value higher than a specific weight parameter is used as the primary recommendation; supporting chain relationship reasoning, and generating a complete treatment chain by combining the strengths of "disease - process" and "process - material".
[0015] An extraction system for a method of enhancing road disease knowledge extraction based on large model self-learning, including a road disease knowledge extraction model 1 BERT, a road disease knowledge extraction decision table, a self-learning road disease knowledge extraction model 2, a primary knowledge graph for road disease detection and maintenance, and a decision effect detection result table.
[0016] Preferably, the road disease knowledge extraction model 1 BERT is composed of road disease detection and maintenance prompt words and the BERT model. Domain experts input the road disease knowledge extraction objectives, requirements, and methods into the BERT model, and require BERT to interpret the task. The experts check and correct to supplement and correct the model's thinking. After guiding the model to be consistent with human thinking, the road disease knowledge extraction model 1 BERT is formed. The road disease knowledge extraction model 1 BERT realizes the extraction of entities and relationships in the standard documents in the field of road disease detection and maintenance.
[0017] Preferably, the road disease knowledge extraction decision table includes graph relationship encoding, relationship strength quantification, and decision table weight self-update. The graph relationship encoding assigns a unique code to different types of road relationships. The "road-disease" relationship code is 1, the "disease-treatment process" relationship code is 2, and the "bridge-structural component" relationship code is 3, which facilitates subsequent processing and recognition of relationships. Quantify the strength of the relationship, represented by a decimal between 0 and 1. The initial quantified relationship strength is generated by the road disease knowledge extraction model 1 BERT extracting standard specification documents, and the subsequent relationship strength values are checked and calibrated by experts based on the decision effect detection result table.
[0018] Preferably, the self-learning road disease knowledge extraction model 2 is trained with the road disease knowledge extraction decision table as the knowledge base to obtain the self-learning road disease knowledge extraction model 2. The road disease knowledge extraction decision table is updated after each knowledge extraction task, so the self-learning road disease knowledge extraction model 2 continuously self-learns to improve the extraction accuracy.
[0019] Preferably, for the primary knowledge graph of road disease detection and maintenance, relevant entity and relationship construction is carried out by inputting the text knowledge sources of relevant papers and policy documents in the road field into the self-learning road disease knowledge extraction model 2, and the road disease knowledge extraction model 2 extracts according to the decision-making situation of the road disease knowledge extraction decision table.
[0020] The decision effect detection result table is generated based on the strength calculation results driven by data. After the self-learning road disease knowledge extraction model 2 extracts the primary knowledge graph of road disease detection and maintenance from the paper file knowledge extraction, the impact on the strength relationship after the paper extraction is calculated based on the primary knowledge graph, forming the decision effect detection result table, and the result value is dynamically updated into the road disease knowledge extraction decision table.
[0021] Compared with the prior art, the beneficial effects of the present invention are:
[0022] The method and system for extracting knowledge of road diseases based on self - learning of large models proposed by the present invention effectively reduces the risk of the spread of incorrect relationships through the constraint of the relationship weight table and the expert review mechanism. Among 1000 test data, the conflict detection module successfully identifies and corrects 83.2% of semantic conflicts (such as the incorrect association of "seam tape" with "subgrade settlement"), and the false alarm rate is less than 5%.
[0023] In the scenario of recommending road disease treatment plans, the sorting mechanism based on relationship strength values significantly improves the decision - making efficiency and scientificity, and at the same time supports the joint reasoning of multi - dimensional relationships.
[0024] Through the incremental learning and weight dynamic adjustment mechanism, the knowledge graph can quickly absorb the latest research results. Brief Description of the Drawings
[0025] Figure 1 It is the flowchart of the method of the present invention. Detailed Embodiments
[0026] In order to clearly and completely describe the objectives, technical solutions of the present invention and make the advantages more clearly understood, the following further details the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are part of the embodiments of the present invention, rather than all of the embodiments, and are only used to explain the embodiments of the present invention, not to limit the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0027] Embodiment 1, please refer to Figure 1 The present invention provides a technical solution: a method for extracting knowledge of road diseases based on self - learning of large models, including the following steps:
[0028] Step 1: An expert constructs a task prompt Prompt to train the road disease knowledge extraction model 1 (BERT), guiding the model to be consistent with human thinking in terms of task objectives, requirements, etc.;
[0029] Step 2: Input the relevant standard documents for road disease detection and maintenance into the road disease knowledge extraction model 1 (BERT), and the road disease knowledge extraction model 1 (BERT) extracts the entities and relationships in the standard documents;
[0030] Step 3: Input the relevant papers on road disease detection and maintenance into the self - learning road disease knowledge extraction model 2. The model 2 determines the weights according to the road disease knowledge extraction decision table, extracts the entities and relationships in the paper documents, and constructs a primary knowledge graph for road disease detection and maintenance;
[0031] Step 4: The self-learning road disease knowledge extraction model 2 affects the strength relationship after extracting the calculation paper based on the data-driven strength calculation result, forming a decision effect detection result table. The expert examines the decision effect detection result table. If the instance strength of a certain relationship is lower than the threshold, the road disease knowledge extraction decision table is updated.
[0032] Step 5: The self-learning road disease knowledge extraction model 2 performs knowledge extraction on the papers and newly input papers based on the updated decision table, continuously improving the domain knowledge graph.
[0033] Example 2: On the basis of Example 1, an extraction system for a method of enhancing road disease knowledge extraction based on large model self-learning is proposed, including a road disease knowledge extraction model 1 (BERT), a road disease knowledge extraction decision table, a self-learning road disease knowledge extraction model 2, a primary road disease detection and maintenance knowledge graph, and a decision effect detection result table.
[0034] Road disease knowledge extraction model 1 (BERT): It is composed of road disease detection and maintenance prompt words and a BERT model. The domain expert inputs the road disease knowledge extraction objectives, requirements, and methods into the BERT model, and requires the BERT to explain the task. The expert checks and corrects to supplement and correct the model's thinking, and forms the road disease knowledge extraction model 1 (BERT) after guiding the model to be consistent with human thinking. The road disease knowledge extraction model 1 (BERT) realizes the extraction of standard document entities and relationships in the field of road disease detection and maintenance.
[0035] Road disease knowledge extraction decision table: It includes graph relationship encoding, relationship strength quantization, and decision table weight self-update. The graph relationship encoding assigns a unique code to different types of road relationships. For example, the "road-disease" relationship is encoded as 1, the "disease-treatment process" relationship is encoded as 2, the "bridge-structural component" relationship is encoded as 3, etc., which is convenient for subsequent processing and identification of relationships; the strength of the relationship is quantified, represented by a decimal between 0 and 1. The initial quantified relationship strength is generated by the road disease knowledge extraction model 1 (BERT) extracting standard specification documents, and the subsequent relationship strength values are checked and calibrated by the expert based on the decision effect detection result table.
[0036]
[0037] Road disease knowledge extraction decision schematic table
[0038] Self-learning Road Disease Knowledge Extraction Model 2: Using the road disease knowledge extraction decision table as the knowledge base, the road disease knowledge extraction model 1 (BERT) is trained to obtain the self-learning road disease knowledge extraction model 2. The road disease knowledge extraction decision table is updated after each knowledge extraction task. Therefore, the self-learning road disease knowledge extraction model 2 continuously self-learns to improve the extraction accuracy.
[0039] Primary Knowledge Graph for Road Disease Detection and Maintenance: Inputting text knowledge sources such as relevant papers and policy documents in the road field into the self-learning road disease knowledge extraction model 2, the model 2 extracts relevant entities and relationships according to the decision-making situation of the road disease knowledge extraction decision table to construct the primary knowledge graph for road disease detection and maintenance.
[0040] Decision Effect Detection Result Table: Generated based on the data-driven strength calculation results. After the self-learning road disease knowledge extraction model 2 extracts the primary knowledge graph for road disease detection and maintenance from the knowledge extraction of papers and other documents, the impact on the strength relationship after the paper extraction is calculated based on the primary knowledge graph to form the decision effect detection result table, and the result value is dynamically updated to the road disease knowledge extraction decision table.
[0041] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for extracting knowledge of reinforced road diseases based on self-learning of large models, characterized in that: It includes the following steps: Extract entities and relationships from the standard documents in the field of road diseases through the pre-trained large model BERT to generate initial knowledge triples; Construct a decision table for road disease knowledge extraction, define hierarchical relationship encoding rules, and quantify the relationship strength values between entities. The relationship strength values are dynamically fused and calculated from the initial scores of experts and the co-occurrence frequency and attention weights driven by data; Train a self-learning model based on the decision table to extract knowledge from papers and unstructured texts to generate a weighted road disease knowledge graph; Dynamically update the relationship strength values through an incremental learning mechanism. When it is detected that the relationship strength deviation exceeds the threshold, trigger expert review, update the decision table, and optimize the knowledge graph.
2. The method for extracting and strengthening road disease knowledge based on large model self-learning according to claim 1, wherein: The relationship strength quantification method includes: experts assign initial strength values to typical entity relationship pairs and score them based on specification clarity and historical cases; calculate the dynamic fusion of the co-occurrence frequency score and the context semantic attention weight score of entity pairs.
3. The method for extracting and strengthening road disease knowledge based on large model self-learning according to claim 2, wherein: The hierarchical relationship encoding rules include: the parent class encoding represents the macro relationship type, such as "disease-treatment process", and the subclass encoding refines the specific relationship, such as "transverse crack-grooving and grouting". The format is "parent class encoding.subclass serial number"; the encoding table stores fields such as the head and tail nodes of entities, relationship strength, update time, and confidence.
4. A method for extracting knowledge of road diseases based on self - learning of large models according to claim 3, characterized in that: Conflict detection and dynamic update include: when new data is input, calculate the deviation between the current relationship strength and the historical value. If the deviation is higher than the weight parameter, trigger expert review.
5. The method for extracting and strengthening road disease knowledge based on large model self-learning according to claim 4, wherein: The knowledge graph is applied to the recommendation of road disease treatment plans, specifically including: sorting and recommending priorities according to the relationship strength values, and the plans with strength values higher than specific weight parameters are used as the primary recommendations; supporting chain relationship reasoning, and combining the strengths of "disease-process" and "process-material" to generate a complete treatment chain.
6. An extraction system for the reinforcement road disease knowledge extraction method based on large model self-learning according to claim 5, characterized in that: It includes the road disease knowledge extraction model 1 BERT, the road disease knowledge extraction decision table, the self-learning road disease knowledge extraction model 2, the primary road disease detection and maintenance knowledge graph, and the decision effect detection result table.
7. An extraction system according to claim 6, characterized in that: The road disease knowledge extraction model 1 BERT: It consists of the road disease detection and maintenance prompt words and the BERT model. Domain experts input the road disease knowledge extraction objectives, requirements, and methods into the BERT model, and require BERT to interpret the task. Expert review supplements and corrects the model's thinking, and after guiding the model to be consistent with human thinking, the road disease knowledge extraction model 1 BERT is formed. The road disease knowledge extraction model 1 BERT realizes the extraction of entities and relationships from the standard documents in the field of road disease detection and maintenance.
8. An extraction system according to claim 7, wherein: Road disease knowledge extraction decision table: It includes graph relationship encoding, relationship strength quantification, and decision table weight self-update. The graph relationship encoding assigns unique codes to different types of road relationships. The "road-disease" relationship code is 1, the "disease-treatment process" relationship code is 2, and the "bridge-structural component" relationship code is 3, which facilitates subsequent processing and identification of relationships. Quantify the strength of the relationship, represented by a decimal between 0 and 1. The initial quantified relationship strength is generated by the road disease knowledge extraction model 1 BERT extraction standard specification file, and the subsequent relationship strength values are verified and calibrated by experts based on the decision effect detection result table.
9. An extraction system according to claim 8, characterized in that: Self-learning road disease knowledge extraction model 2: Using the road disease knowledge extraction decision table as the knowledge base, train the road disease knowledge extraction model 1 BERT to obtain the self-learning road disease knowledge extraction model 2. The road disease knowledge extraction decision table is updated after each knowledge extraction task, so the self-learning road disease knowledge extraction model 2 continuously self-learns to improve the extraction accuracy.
10. An extraction system according to claim 9, wherein: Primary knowledge graph for road disease detection and maintenance: Input the text knowledge sources of relevant papers and policy documents in the road field into the self-learning road disease knowledge extraction model 2. The road disease knowledge extraction model 2 extracts relevant entities and relationships according to the decision-making situation of the road disease knowledge extraction decision table to construct the primary knowledge graph for road disease detection and maintenance; Decision effect detection result table: Generated based on the data-driven strength calculation results. After the self-learning road disease knowledge extraction model 2 extracts the knowledge of the paper file and obtains the primary knowledge graph for road disease detection and maintenance, calculate the impact on the strength relationship after the paper extraction based on the primary knowledge graph to form the decision effect detection result table, and dynamically update the result value to the road disease knowledge extraction decision table.