A Method for Constructing a Knowledge Graph for Intelligent Airport Pavement Maintenance Based on Multi-Source Heterogeneous Data

CN122334436BActive Publication Date: 2026-09-01TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610814825.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-08
Publication Date
2026-09-01
Estimated Expiration
2046-06-08

AI Technical Summary

Technical Problem

[0004]为了解决针对机场道面养护的知识图谱构建方法对数据置信度的评估准确性较低的技术问题,本发明的目的在于提供一种面向多源异构数据的机场道面智能养护知识图谱构建方法,所采用的技术方案具体如下:

Benefits of technology

本发明依据机场道面的多源异构数据对同一位置同一属性的多源检测数据量化差异表现,差异愈大则置信度愈低。进一步引入知识图谱中邻域空间及关联参考数据,通过量化损伤演化表现检验当前目标源数据是否违背道面损伤的多源反馈规律,结合差异表现和损伤演化表现确定目标源的数据合理程度;再基于历史时序与合理演化的数据推测,进一步通过目标源数据的时序趋势综合确定置信度。本发明充分考虑到多源系统性误差的情况,有效识别多源数据冲突与异常反馈,降低误判风险,提高数据置信度量化的准确性,进而提升损伤检测的鲁棒性与可靠性,为机场道面维护提供高可信决策支持。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122334436B_ABST
    Figure CN122334436B_ABST
Patent Text Reader

Abstract

This invention relates to the field of electronic digital data processing technology, specifically to a method for constructing a knowledge graph for intelligent maintenance of airport pavement based on multi-source heterogeneous data. The method includes: determining the differences between target source data of a target data type at a target location and other source data; determining the target node corresponding to the target source and its adjacent nodes in the knowledge graph, and determining historical reference data of the target data type in the target node based on historical pavement monitoring data of the adjacent nodes; determining the damage evolution performance based on the target source data and historical reference data, and determining the data rationality of the target source based on the differences and damage evolution performance; using a linear regression model to predict the target data type at the target location, and determining the confidence level of the target source data based on the predicted data and the data rationality, and visualizing the confidence level. Through the technical solution of this invention, the risk of misjudgment is reduced, and the accuracy of data confidence quantification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electronic digital data processing technology, specifically to a method for constructing an intelligent maintenance knowledge graph for airport pavement based on multi-source heterogeneous data. Background Technology

[0002] Airport pavement maintenance involves multi-source heterogeneous data from detection, inspection, and sensors, which traditional methods struggle to effectively integrate and utilize. Knowledge graph technology offers a new approach to integrating multi-source information and supporting intelligent decision-making; however, existing methods for constructing knowledge graphs lack systematic adaptation to airport pavement scenarios, necessitating a dedicated knowledge graph construction method for this field.

[0003] Traditional methods only remove outlier data by observing the differences in values ​​obtained from different data sources for the same data type at the same location. This does not take into account the situation where multiple data sources can interfere with each other, meaning that the values ​​obtained from multiple data sources may deviate from the actual values. This situation can easily lead to incorrect assessments of the confidence levels of different data, which in turn affects subsequent data analysis. Summary of the Invention

[0004] To address the technical problem of low accuracy in assessing data confidence in knowledge graph construction methods for airport pavement maintenance, this invention aims to provide a method for constructing an intelligent knowledge graph for airport pavement maintenance based on multi-source heterogeneous data. The specific technical solution adopted is as follows: This invention provides a method for constructing an intelligent maintenance knowledge graph for airport pavement based on multi-source heterogeneous data, the method comprising: An initial knowledge graph is created based on pavement monitoring data from multiple sources at the target location to determine the differences between the current target source data and other source data for the target data type at the target location. Determine the target node and its neighboring nodes corresponding to the target source in the initial knowledge graph, and determine the historical reference data of the target data type in the target node based on the historical pavement monitoring data of the neighboring nodes; The target source reflects the pavement damage evolution based on the target source data and historical reference data. The reasonableness of the target source data is determined based on the differences and damage evolution. The predicted data of the target data type at the target location is obtained by using a linear regression model. The confidence level of the target source data is determined based on the predicted data and the reasonableness of the data, and the confidence level is visualized.

[0005] Furthermore, the attributes of the target data type are numeric attributes; The determination of the differences between the current target source data and other source data at the target location includes: Get the absolute value of the difference between the current source value of the target data type at the target location and other source values; The difference between the target source data and other source data is determined by summing the absolute values ​​of each difference.

[0006] Furthermore, the attribute of the target data type is a categorical attribute; The determination of the differences between the current target source data and other source data at the target location includes: Get the number of times the target data type at the target location appears in all source category results for the current target source category; Obtain the category result that appears most frequently among all source category results, and determine the difference between the target source data and other source data based on the difference between the target occurrence count and the most frequent occurrence count.

[0007] Furthermore, the attribute of the target data type is a hierarchical attribute; The determination of the differences between the current target source data and other source data at the target location includes: Get the absolute value of the difference between the current source level value of the target data type at the target location and other source level values; The difference between the target source data and other source data is determined by summing the absolute values ​​of the differences at each level.

[0008] Furthermore, the step of determining historical reference data of the target data type in the target node based on historical pavement monitoring data of adjacent nodes includes: Obtain historical pavement condition data for adjacent nodes at historical moments and current pavement condition data for the current moment; Based on historical pavement condition data and current pavement condition data, determine the historical reference data for the target data type in the target node.

[0009] Furthermore, the step of determining historical reference data of the target data type in the target node based on historical pavement condition data and current pavement condition data includes: Determine whether the historical pavement condition at a given time and the current pavement condition at the present time are similar based on historical pavement condition data and current pavement condition data. When the historical and current operating conditions are similar, determine the historical reference data corresponding to the target data type in the target node at the historical moment.

[0010] Furthermore, the step of determining the target source based on target source data and historical reference data to reflect the pavement damage evolution includes: Identify the differences between the target source data and historical reference data and use them as the target source to reflect the damage evolution of the pavement.

[0011] Furthermore, the determination of the data reasonableness of the target source based on differential performance and damage evolution performance includes: The data rationality of the target source is obtained by combining the differential performance and damage evolution performance, where both differential performance and damage evolution performance are negatively correlated with the data rationality.

[0012] Furthermore, the confidence level of the target source data is determined based on the predicted data and the reasonableness of the data, including: Determine the prediction deviation between the predicted data and the target source data, and determine the difference in reasonableness between the reasonableness of the target source data and the reasonableness of the historical maximum data. The confidence level of the target source data is obtained by combining the prediction bias and the difference in reasonableness.

[0013] Furthermore, the step of determining the confidence level of the target source data based on the predicted data and the reasonableness of the data, and visualizing the confidence level, includes: The confidence levels of each source data are compared, and the source data with the highest confidence level that is greater than a preset confidence threshold is taken as the application data. The initial knowledge graph is updated based on the application data to obtain the target knowledge graph. Each node in the target knowledge graph is assigned a corresponding node fill color and edge style based on its confidence level for visualization purposes.

[0014] The present invention has the following beneficial effects: This invention quantifies the differences in multi-source detection data with the same attribute at the same location based on multi-source heterogeneous data of airport pavements; the greater the difference, the lower the confidence level. It further introduces neighborhood space and associated reference data from a knowledge graph to examine whether the current target source data violates the multi-source feedback pattern of pavement damage by quantifying damage evolution performance. The reasonableness of the target source data is determined by combining the difference performance and damage evolution performance. Then, based on historical time series and reasonable evolution data, the confidence level is further determined by comprehensively considering the time series trend of the target source data. This invention fully considers the situation of multi-source systematic errors, effectively identifies multi-source data conflicts and abnormal feedback, reduces the risk of misjudgment, improves the accuracy of data confidence quantification, and thus enhances the robustness and reliability of damage detection, providing highly reliable decision support for airport pavement maintenance. Attached Figure Description

[0015] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 The flowchart shows the steps of a method for constructing an intelligent airport pavement maintenance knowledge graph based on multi-source heterogeneous data, as provided in one embodiment of the present invention. Figure 2 This is a detailed flowchart of step S1 in a method for constructing an intelligent airport pavement maintenance knowledge graph based on multi-source heterogeneous data, provided in an embodiment of the present invention. Figure 3 This is a detailed flowchart of step S2 in a method for constructing an intelligent airport pavement maintenance knowledge graph based on multi-source heterogeneous data, provided in an embodiment of the present invention. Figure 4 This is a detailed flowchart of step S22 in a method for constructing an intelligent airport pavement maintenance knowledge graph based on multi-source heterogeneous data, provided in an embodiment of the present invention. Figure 5 This is a detailed flowchart of step S4 in a method for constructing an intelligent airport pavement maintenance knowledge graph based on multi-source heterogeneous data, provided in an embodiment of the present invention. Figure 6 This is a schematic diagram of the hardware operating environment of the airport pavement intelligent maintenance knowledge graph construction device for multi-source heterogeneous data involved in the embodiments of the present invention. Detailed Implementation

[0017] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a method for constructing an intelligent airport pavement maintenance knowledge graph based on multi-source heterogeneous data proposed by the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0019] The following description, in conjunction with the accompanying drawings, details the specific scheme of the method for constructing an intelligent airport pavement maintenance knowledge graph based on multi-source heterogeneous data provided by the present invention.

[0020] Example 1: Please see Figure 1 , Figure 1 The flowchart illustrates the steps of a method for constructing an intelligent airport pavement maintenance knowledge graph based on multi-source heterogeneous data, according to an embodiment of the present invention.

[0021] The method for constructing an intelligent airport pavement maintenance knowledge graph for multi-source heterogeneous data includes the following steps: Step S1: Create an initial knowledge graph based on multi-source pavement monitoring data at the target location, and determine the differences between the current target source data and other source data of the target data type at the target location. The pavement monitoring data required for intelligent maintenance of airport pavements often includes periodic professional inspection data (including automated and manual inspections, mainly referring to automated inspections). Broadly speaking, this includes pavement condition index data, structural performance test data, etc. It also includes routine manual and intelligent inspection data, such as foreign object lists, sudden defect records, maintenance operation inspection data, etc. It also includes sensor network data, such as environmental data from temperature and humidity sensors, dynamic load data from piezoelectric / strain sensors, and pavement condition data from pavement condition sensors. These multi-source pavement monitoring data from professional inspections, daily manual inspections, sensors, and other data sources are common in this field, and will not be further illustrated here.

[0022] For the collection of pavement monitoring data for an intelligent airport pavement maintenance knowledge graph, the data structure includes: structured data (such as building information models and sensor databases) directly read via database interfaces or file export; semi-structured data (such as inspection Excel files and extensible markup language reports) parsing key fields using regular expression templates; unstructured text (such as manual inspection descriptions and maintenance logs) extracting entities and relationships based on a standard glossary; and unstructured images (such as photos of pavement defects) identifying defect types and locations using a pre-trained object detection model. All collected data is appended with pavement station numbers and timestamps, and collection logs are recorded for traceability. Here, pavement station numbers represent individual pavement locations, and the target location refers to any pavement location. For clarity, the following embodiments all use the same target location as an example.

[0023] After completing data collection and preprocessing (cleaning missing data, merging duplicates, and standardizing terminology), the graph construction phase begins. First, an ontology for airport pavement maintenance is constructed, defining core classes: pavement location (runway / taxiway / apron), defects (cracks / ruts / joint breakage), maintenance measures, environmental loads, and relationships (occurrence / endurance / adaptation, etc.). Then, the extracted entity-relationship-entity triples are imported into a graph database (such as Neo4j). Nodes are created for each entity and labeled with category and attributes; edges are created for relationships and assigned confidence levels and timestamps. Finally, indexes are built for frequently queried fields such as station number and defect type, forming an initial knowledge graph that can be queried and used for reasoning.

[0024] When acquiring airport pavement maintenance data (pavement monitoring data), the results from automated detection equipment, manual inspection, and sensors (collectively referred to as "three sources" or "multi-sources") will inevitably fluctuate due to differences in measurement principles, ambient lighting, or operational experience. If the detection value from one data source deviates significantly from the consensus range formed by other sources, it is considered an anomaly. Different data types with different attributes require different methods of measuring difference. For numerical attributes such as crack width, assuming the results from the three sources are 2.5 mm, 2.6 mm, and 3.9 mm, the latter is assigned a higher difference rating because it is far from the other values. For categorical attributes such as damage type, if two sources identify longitudinal cracks and another source reports joint breakage, deviation from the majority category is considered an anomaly. For graded attributes such as severity (mild, moderate, severe), when the majority of sources rate it as moderate and a data source gives severe, the deviation is calculated after mapping the graded values. The greater the deviation, the higher the score for the degree of difference, thus accurately identifying unreliable detections in the multi-source data.

[0025] Here, the target data type refers to any data type, and the target source data refers to data detected by any data source, such as sensor data.

[0026] In one embodiment, the attribute of the target data type is a numeric attribute; please refer to... Figure 2 ; Step S1, determining the difference between the current target source data and other source data at the target location, includes: Step S11: Obtain the absolute value of the difference between the current source value of the target data type at the target location and other source values; Step S12: Determine the difference between the target source data and other source data based on the sum of the absolute values ​​of each difference.

[0027] In this embodiment, for data types of numerical attributes (such as crack width and deflection value), at the current time i, all numerical values ​​of the target data type b (e.g., crack width) obtained from multi-source measurements are acquired for the target location. The numerical value of data type b obtained from a single data source a is then calculated. The magnitude of the target source value and the value of data type b obtained from other data sources c. The absolute value of the difference between (other source values) is then used to sum the absolute values ​​of the differences, which represents the difference between the value of data type b obtained from data source a and the values ​​from other data sources (after normalization). norm ( (This is the number of other sources for the available data type b, for example, 2). `norm` represents normalization, such as maximum and minimum value normalization. The maximum and minimum values ​​are taken from the historical maximum and minimum value differences (the original values ​​that are neither normalized). If there are extreme cases where values ​​exceed historical extremes, truncation can be performed: values ​​greater than the maximum are normalized to 1, and values ​​less than the minimum are normalized to 0.

[0028] In one embodiment, the attribute of the target data type is a categorical attribute; Step S1, determining the difference between the current target source data and other source data at the target location, includes: Get the number of times the target data type at the target location appears in all source category results for the current target source category; Obtain the category result that appears most frequently among all source category results, and determine the difference between the target source data and other source data based on the difference between the target occurrence count and the most frequent occurrence count.

[0029] In this embodiment, for data types with categorical attributes (such as disease types, including cracks, deformation, surface defects, etc.), all results obtained from multi-source detection for target data type b (e.g., crack) at target location i at the current time are counted. The number of times the target appears in the category results (target source category results) obtained from a single data source a among all data source detection results (all source category results) is also counted. .

[0030] At the current time i, count the number of all data sources that yield the category results for data type b. The maximum number of times each category result appears across all data sources is determined by comparison. This also determines the category result corresponding to the most frequent occurrence. The most frequent occurrence... Number of times the target appears The difference (number of times difference) and The ratio of the values ​​of the two values ​​represents the difference between the category result of data type b obtained from data source a and the category results of other sources. , , here Normally, the value is not 0; otherwise, all other parameters in the formula would also be 0, making it meaningless to calculate and analyze differences. If a special anomaly is detected... A value of 0 indicates that analysis of the current data type 'b' is stopped or ignored, or manual checks can be performed to ensure that... Further analysis is only necessary if the result is not zero.

[0031] In one embodiment, the attribute of the target data type is a hierarchical attribute; Step S1, determining the difference between the current target source data and other source data at the target location, includes: Get the absolute value of the difference between the current source level value of the target data type at the target location and other source level values; The difference between the target source data and other source data is determined by summing the absolute values ​​of the differences at each level.

[0032] In this embodiment, at the current time i, for data types with graded attributes (such as severity levels, including mild, moderate, and severe), all grade results obtained from multi-source detection for the target data type b are statistically analyzed. The grades are mapped to arithmetic grade values ​​(e.g., mild is 1, moderate is 2, and severe is 3). The grade value of data type b obtained from a single data source a is calculated. (Target source level value), and the level value of data type b obtained from other data sources c. The sum of the absolute values ​​of the level differences (of other source level values) ( (This is the number of other sources for the data type b that can be obtained, for example, 2).

[0033] Calculate the sum of the absolute values ​​of the differences With the level range (maximum level) and minimum value The difference The ratio of the two values ​​represents the difference in grade between data type b obtained from data source a and data type b obtained from other data sources. Based on preset level values, such as 1, 2, and 3, its maximum level is... and minimum value The difference between It is obvious that it is not 0. Even if other level values ​​are preset as needed, its maximum and minimum values ​​cannot be the same. Otherwise, there would be no need to identify the level results, that is, there would be no level distinction.

[0034] Step S2: Determine the target node and its neighboring nodes corresponding to the target source in the initial knowledge graph, and determine the historical reference data of the target data type in the target node based on the historical pavement monitoring data of the neighboring nodes. For the same target location, considering that all current data sources may exhibit systematic biases in the same direction (e.g., all detection equipment and sensors are uncalibrated), multi-source data may be highly consistent but collectively deviate from the true value, leading to misjudgment as reliable. Furthermore, when a data source exhibits significant differences from other data sources due to systematic biases, it becomes impossible to determine which data source is correct, lacking an accurate independent reference benchmark and making it difficult to distinguish between random and systematic errors. Therefore, this approach utilizes the associated data of neighboring nodes in a knowledge graph when their historical detection conditions are similar to the current working conditions to verify whether the current target source data of the target node conforms to the continuity and similarity patterns of pavement damage. Through spatial or semantic redundancy constraints, it effectively identifies isolated point anomalies or systematic biases, avoiding misjudgments caused by limited information from a single node and significantly improving the accuracy and robustness of data confidence assessment.

[0035] Specifically, please refer to Figure 3 ; Step S2, which determines historical reference data of the target data type in the target node based on historical pavement monitoring data of adjacent nodes, includes: Step S21: Obtain historical pavement condition data of adjacent nodes at historical times and current pavement condition data at the current time. Step S22: Based on historical pavement condition data and current pavement condition data, determine the historical reference data of the target data type in the target node.

[0036] More specifically, please refer to Figure 4 Step S22 includes: Step S221: Determine whether the historical working conditions at a historical moment and the current working conditions at the current moment are similar working conditions based on historical pavement working condition data and current pavement working condition data. Step S222: When the historical working conditions and the current working conditions are similar, determine the historical reference data corresponding to the target data type in the target node at the historical moment.

[0037] In this embodiment, the target node corresponding to the target source is first determined. Multiple nodes adjacent to the target node (e.g., within a fixed distance before and after it) are retrieved from the initial knowledge graph. For example, three adjacent nodes are selected. The same attribute value of these adjacent nodes, which is similar to the current pavement condition in historical detections, is extracted. Taking any adjacent node d as an example, the current pavement condition data of adjacent node d at the current time i is obtained. In the most recent historical detections, historical time i' and historical pavement condition data are found. This pavement condition data can include working conditions such as temperature, humidity, weather, and environmental load. The historical pavement condition data and the current pavement condition data are compared. For numerical attribute data, the absolute value of the numerical difference between the two can be obtained. Different similarity thresholds can be manually set according to different data types. The similarity threshold can be set in the following way: obtain a large amount of historical data of the corresponding data type, and then calibrate the large amount of historical data to obtain the corresponding similarity threshold. The numerical difference is compared with a similarity threshold. If the numerical difference is less than or equal to the similarity threshold, the data is considered similar. For example, if the historical temperature is 20℃ and the current temperature is 22℃, the similarity threshold for temperature data is 3℃. Therefore, the historical temperature and the current temperature can be considered similar. Of course, pavement condition data is generally a dataset that can include various condition data. It is also necessary to compare the historical and current data of other condition categories separately. When all historical pavement condition data and all current pavement condition data are similar according to their respective condition categories, then the historical pavement condition data and the current pavement condition data are considered similar, that is, the corresponding historical conditions and current conditions are similar.

[0038] For data with categorical and hierarchical attributes, semantic recognition can be used to determine whether all historical pavement condition data are identical to the corresponding current pavement condition data. If they are identical, then the historical and current pavement condition data are considered similar. If the pavement condition data includes numerical, categorical, and hierarchical data, numerical data is judged based on a similarity threshold, while categorical and hierarchical data are judged based on whether they are identical. Thus, if all are similar or all are identical, the similarity between historical and current pavement conditions can be identified.

[0039] When the historical and current operating conditions of an adjacent node d are found to be similar, the above operations can be used to further determine whether the historical and current operating conditions of other adjacent nodes are also similar. If the historical operating conditions of all adjacent nodes at historical time i' are similar to the current operating conditions at current time i, then historical time i' is determined. The historical reference data corresponding to the target data type b in the target node at historical time i' is extracted. This historical reference data is the historical target source data. Generally, there are multiple historical times i', and therefore multiple historical reference data. The number of historical times i' often depends on the duration of the most recent historical detection, that is, how long of historical pavement operating condition data to take. This is set according to actual needs, such as one week or 24 hours before the current time.

[0040] Step S3: Determine the damage evolution performance of the pavement reflected by the target source based on the target source data and historical reference data, and determine the rationality of the target source data based on the difference performance and damage evolution performance. Specifically, step S3, which determines the damage evolution of the pavement reflected by the target source based on the target source data and historical reference data, includes: Identify the differences between the target source data and historical reference data and use them as the target source to reflect the damage evolution of the pavement.

[0041] Based on the above implementation of the difference assessment, for data type b according to the corresponding attribute (numerical, categorical, or hierarchical), the similar difference assessment between the target source data and the historical reference data is obtained. For example, for numerical data, for data type b, calculate the absolute value of the difference between the target source value of data type b obtained from data source a and the value of each of its historical reference data, and then use the average of the absolute values ​​of each difference as the source difference performance of the target source data of data source a relative to the historical reference data. For example, for categorical data, for the target data type b, we can collect the category results of all historical reference data (which may include the target source data) and calculate the occurrence count of each category. The maximum value among the occurrences is determined as the maximum occurrence count; the ratio of the difference (numerator) between the maximum occurrence count and the target occurrence count corresponding to data source a, to the total number of category results (denominator), is calculated as the source data's homogeneity difference performance relative to historical reference data.

[0042] Specifically, when the historical reference data contains only one sample, the duration of historical detection is extended.

[0043] For graded data, for data type b, calculate the sum of the absolute values ​​of the grade difference between the grade value corresponding to the target source data of data source a and the grade value of the historical reference data. Then calculate the sum of the absolute values ​​of the grade differences and the grade range (maximum grade value). and minimum value The difference The ratio of the extreme values ​​of the grade values ​​(which can be obtained from historical reference data) is used as the source data a to represent the source data differences between the target source data and the historical reference data.

[0044] This homology difference manifests It reflects the damage evolution of the track surface detected by the target source. Under normal circumstances, the smaller the value of the damage evolution, the more it conforms to the actual situation of normal damage evolution. Therefore, it can reflect the confidence level of the current target source data. That is, the smaller the damage evolution value, the greater the confidence level.

[0045] Specifically, step S3, determining the reasonableness of the target source data based on the difference performance and damage evolution performance, includes: The data rationality of the target source is obtained by combining the differential performance and damage evolution performance, where both differential performance and damage evolution performance are negatively correlated with the data rationality.

[0046] Damage evolution manifestations The smaller the value, the greater the difference in performance obtained from the above embodiments. ( , , The smaller the value of the target source data (the data of type b obtained by data source a), the closer it is to the historical reference data corresponding to the most recent similar working conditions, and the smaller the difference between it and the data obtained from different data sources during the same period of monitoring. The fewer problems exist in the target source data, the greater its confidence level should be.

[0047] This allows us to determine the reasonableness of data of data type b obtained from data source a. Using the max-min normalization method to... After normalization, we get Its range is [0,1]. The maximum and minimum values ​​can be taken as the maximum and minimum reasonableness of the historical data of data type b obtained from data source a. The 0.01 in the denominator is a constant set to avoid the denominator being 0. At the same time, considering that the range of the above differences is in [0,1], 0.01 has a small impact on the calculation results.

[0048] Step S4: Use a linear regression model to predict the target data type at the target location, determine the confidence level of the target source data based on the predicted data and the reasonableness of the data, and visualize the confidence level.

[0049] Pavement damage also exhibits clear dynamic temporal evolution patterns, such as the slow increase in crack width with accumulated load and the phased recovery of performance after maintenance. Relying solely on spatial neighborhood or semantic similarity may not be sufficient to determine whether current data deviates from normal temporal trends. For example, if a crack width at a certain location was 2.0 mm last month but is reported as 0.5 mm this month, even if consistent with the surrounding data, it clearly violates the irreversible growth pattern of damage. Comparing historical timelines with evolution rates can identify such anomalies that contradict temporal logic, avoiding misjudgments due to systematic biases in neighborhood information, thus providing a more comprehensive basis for overall confidence levels.

[0050] Specifically, please refer to Figure 5 ; Step S4, determining the confidence level of the target source data based on the predicted data and the reasonableness of the data, includes: Step S41: Determine the prediction deviation between the predicted data and the target source data, and determine the difference in reasonableness between the reasonableness of the target source data and the reasonableness of the historical maximum data. Step S42: Combine the prediction bias and the difference in reasonableness to obtain the confidence level of the target source data.

[0051] Based on historical monitoring data from the same location, a linear regression model is used to predict future trends in pavement damage. Taking crack width as an example, width records for a specific pavement station at multiple points in the past year (e.g., 2.0 mm, 2.2 mm, 2.4 mm, 2.7 mm) are collected. A linear regression equation is fitted with time as the independent variable and width as the dependent variable, calculating an average monthly growth rate of approximately 0.07 mm, thus predicting that the width will reach 2.9 mm in three months. If maintenance interventions were observed in the historical data, piecewise linear regression is used, fitting different slopes before and after maintenance. For data with large fluctuations, an exponential smoothing model is used to assign higher weights to recent data, improving predictive sensitivity.

[0052] For categorical attributes, one-hot encoding is first performed on unordered categories (such as pavement type), while ordinal encoding is retained for ordered categories (such as damage level). Numerical features (pavement age, traffic volume, etc.) must be standardized to eliminate dimensions. After manually constructing cross terms such as "pavement type × pavement age," the training and test sets are strictly divided according to time order. A ridge regression model is trained, and L2 regularization is used to suppress multicollinearity, making it suitable for pavement deterioration scenarios with small samples and approximately linear relationships. When obtaining prediction results, the data to be predicted undergoes the same encoding, standardization, and cross-term processing, and is directly substituted into the trained model to calculate the predicted pavement condition index value. Finally, based on the sign and magnitude of each feature coefficient, its impact on the pavement condition index is directly interpreted, meeting the high interpretability requirements of engineering projects.

[0053] By using the historical pavement monitoring data of data type b at the same location as described above, the data for the current time i is predicted, resulting in the predicted data. .

[0054] Calculate prediction data The difference (prediction bias) between the actual data of data type b detected from data source a at time i and the actual data of data type b (target source data). (Numerical attribute data represents the numerical difference; categorical attribute data represents the difference preset by the user based on similarity; and graded attribute data represents the difference between grade values.) It should be noted that for categorical attribute data, the difference can be set manually based on the similarity between the two categories, with a value range of [0,1]. For example, the difference between longitudinal cracks and transverse cracks can be set to 0.4.

[0055] More specifically, when the target source detection category result is completely consistent with the predicted category result, ( (This refers to a general term for prediction deviations); when the two are inconsistent but belong to the same type of disease (such as "longitudinal cracks" and "diagonal cracks"), Values Empirical values ​​between them; when the mechanisms of the two are completely different (such as "cracks" and "chipping"), Values In this way, qualitative semantic logic is transformed into quantitative indicators that can be used in confidence calculations.

[0056] It should also be noted that the numerical difference or the level numerical difference are normalized to the maximum and minimum values ​​so that their value range is [0,1].

[0057] When the difference The smaller the value, the more reasonable the target source data of data type b detected by data source a at time i. Compared to the maximum value of the reasonableness of data during historical testing of data type b (i.e., the maximum reasonableness of historical data). The difference (the degree of reasonableness) The smaller the value, the higher the confidence level of the target source data. This indicates that the target source data of data type b detected by data source a at time i is of reasonable size under the comparison of multi-source heterogeneous data, and also meets its own data change requirements over time. If the value is less than 0, then... Set to 0.

[0058] Therefore, the confidence level of the target source data of data type b detected by data source a at time i can be obtained. : .

[0059] Using the maximum-minimum normalization method After normalization, we get ,Right now and The meaning is the same. yes The normalization result, The value range is [0,1]. The maximum and minimum values ​​used for normalization are taken from the historical confidence score extreme values ​​of data source a for data type b. If there are extreme cases where the values ​​exceed the historical extreme values, truncation can be performed, that is, values ​​greater than the maximum value are normalized to 1, and values ​​less than the minimum value are normalized to 0. The 0.01 in the denominator is a constant set to avoid the denominator being 0. Considering that the parameter range of each denominator is [0,1], 0.01 has a small impact on the calculation result. The confidence scores mentioned later are all normalized confidence scores.

[0060] Compare the confidence levels of data of data type b at the same location at time i across all data sources. The data with the highest confidence level (greater than 0.5, a preset confidence threshold) is selected as the final application data to ensure greater accuracy in the initial knowledge graph. If the confidence level of data from multiple data sources at the same location and time is less than 0.5, remeasurement and re-testing are required.

[0061] In one embodiment, step S4, determining the confidence level of the target source data based on the predicted data and the reasonableness of the data, and visually displaying the confidence level, includes: The confidence levels of each source data are compared, and the source data with the highest confidence level that is greater than a preset confidence threshold is taken as the application data. The initial knowledge graph is updated based on the application data to obtain the target knowledge graph. Each node in the target knowledge graph is assigned a corresponding node fill color and edge style based on its confidence level for visualization purposes.

[0062] To address the confidence levels of multi-source heterogeneous data, a hierarchical visual coding approach is employed in the knowledge graph visualization interface. The confidence level (from 0 to 1) of each knowledge point (entity node) is mapped to the node's fill color and edge style: nodes with a confidence level of 0.85 or higher are dark green with solid edges; 0.70 to 0.85 are light green; 0.60 to 0.70 are yellow with dashed edges; 0.50 to 0.60 are orange with dotted lines; and nodes with a confidence level below 0.50 are red with bold red dashed edges, overlaid with a cross or exclamation mark icon to clearly identify low-quality data. Small icons in the corners of nodes indicate the data source type (camera, radar, personnel, sensor), and hovering the mouse displays the specific confidence level value for each data source. The interface also provides a timeline slider, allowing playback of confidence level distributions at different historical moments. This visualization approach allows users to intuitively grasp the quality differences of multi-source data in the knowledge graph, improving decision-making efficiency and system interpretability.

[0063] This invention quantifies the differences in multi-source detection data with the same attribute at the same location based on multi-source heterogeneous data of airport pavements; the greater the difference, the lower the confidence level. It further introduces neighborhood space and associated reference data from a knowledge graph to examine whether the current target source data violates the multi-source feedback pattern of pavement damage by quantifying damage evolution performance. The reasonableness of the target source data is determined by combining the difference performance and damage evolution performance. Then, based on historical time series and reasonable evolution data, the confidence level is further determined by comprehensively considering the time series trend of the target source data. This invention fully considers the situation of multi-source systematic errors, effectively identifies multi-source data conflicts and abnormal feedback, reduces the risk of misjudgment, improves the accuracy of data confidence quantification, and thus enhances the robustness and reliability of damage detection, providing highly reliable decision support for airport pavement maintenance.

[0064] Example 2: This invention also proposes a device for constructing an intelligent knowledge graph for airport pavement maintenance based on multi-source heterogeneous data. The device can be a data processing device such as a computer or a server, or a combination of multiple devices.

[0065] like Figure 6 As shown, Figure 6 This is a schematic diagram of the hardware operating environment of the airport pavement intelligent maintenance knowledge graph construction device for multi-source heterogeneous data involved in the embodiments of the present invention.

[0066] like Figure 6As shown, the airport pavement intelligent maintenance knowledge graph construction device for multi-source heterogeneous data may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display or an input unit such as a control panel; the user interface 1003 may also include standard wired or wireless interfaces. The network interface 1004 may optionally include standard wired or wireless interfaces (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as a disk drive. The memory 1005 may also optionally be a storage device independent of the aforementioned processor 1001. The memory 1005, as a computer storage medium, may include an airport pavement intelligent maintenance knowledge graph construction program for multi-source heterogeneous data (hereinafter referred to as the "airport pavement intelligent maintenance knowledge graph construction program").

[0067] Those skilled in the art will understand that Figure 6 The hardware structure shown does not constitute a limitation on the device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0068] Continue to refer to Figure 6 , Figure 6 The memory 1005, which is a computer-readable storage medium, may include an operating system, a user interface module, a network communication module, and a knowledge graph construction program for intelligent maintenance of airport pavement for multi-source heterogeneous data.

[0069] exist Figure 6 In this embodiment, the network communication module is mainly used to connect to the server and can communicate with the server for data; while the processor 1001 can call the airport pavement intelligent maintenance knowledge graph construction program for multi-source heterogeneous data stored in the memory 1005 and execute the steps in the above embodiments.

[0070] The hardware structure of the airport pavement intelligent maintenance knowledge graph construction device based on the above-mentioned multi-source heterogeneous data is used to implement various embodiments of the airport pavement intelligent maintenance knowledge graph construction method based on multi-source heterogeneous data of the present invention.

[0071] Furthermore, the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a program for constructing an intelligent airport pavement maintenance knowledge graph for multi-source heterogeneous data. When executed by a processor, the program implements the steps of the method described above for constructing an intelligent airport pavement maintenance knowledge graph for multi-source heterogeneous data.

[0072] The method implemented when the airport pavement intelligent maintenance knowledge graph construction program for multi-source heterogeneous data is executed can be referred to in various embodiments of the airport pavement intelligent maintenance knowledge graph construction method for multi-source heterogeneous data of the present invention, and will not be repeated here.

[0073] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0074] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0075] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0076] The above description is only a preferred embodiment of the present invention and does not limit the scope of protection of the present invention. All equivalent structural / method transformations made under the inventive concept of the present invention using the contents of the present invention specification and drawings, or direct / indirect applications in other related technical fields, are included within the scope of protection of the present invention.

Claims

1. A method for constructing an intelligent maintenance knowledge graph for airport pavement based on multi-source heterogeneous data, characterized in that, The method includes: An initial knowledge graph is created based on pavement monitoring data from multiple sources at the target location to determine the differences between the current target source data and other source data for the target data type at the target location. Determine the target node and its neighboring nodes corresponding to the target source in the initial knowledge graph, and determine the historical reference data of the target data type in the target node based on the historical pavement monitoring data of the neighboring nodes; The target source reflects the pavement damage evolution based on the target source data and historical reference data. The reasonableness of the target source data is determined based on the differences and damage evolution. The predicted data of the target data type at the target location is obtained by using a linear regression model. The confidence level of the target source data is determined based on the predicted data and the reasonableness of the data, and the confidence level is visualized.

2. The method for constructing an intelligent airport pavement maintenance knowledge graph based on multi-source heterogeneous data as described in claim 1, characterized in that, The attribute of the target data type is a numeric attribute; The determination of the differences between the current target source data and other source data at the target location includes: Get the absolute value of the difference between the current source value of the target data type at the target location and other source values; The difference between the target source data and other source data is determined by summing the absolute values ​​of each difference.

3. The method for constructing an intelligent airport pavement maintenance knowledge graph based on multi-source heterogeneous data as described in claim 1, characterized in that, The attribute of the target data type is a categorical attribute; The determination of the differences between the current target source data and other source data at the target location includes: Get the number of times the target data type at the target location appears in all source category results for the current target source category; Obtain the category result that appears most frequently among all source category results, and determine the difference between the target source data and other source data based on the difference between the target occurrence count and the most frequent occurrence count.

4. The method for constructing an intelligent airport pavement maintenance knowledge graph based on multi-source heterogeneous data according to claim 1, characterized in that, The attribute of the target data type is a hierarchical attribute; The determination of the differences between the current target source data and other source data at the target location includes: Get the absolute value of the difference between the current source level value of the target data type at the target location and other source level values; The difference between the target source data and other source data is determined by summing the absolute values ​​of the differences at each level.

5. The method for constructing an intelligent airport pavement maintenance knowledge graph based on multi-source heterogeneous data as described in claim 1, characterized in that, The step of determining historical reference data of the target data type in the target node based on historical pavement monitoring data of adjacent nodes includes: Obtain historical pavement condition data for adjacent nodes at historical moments and current pavement condition data for the current moment; Based on historical pavement condition data and current pavement condition data, determine the historical reference data for the target data type in the target node.

6. The method for constructing an intelligent airport pavement maintenance knowledge graph based on multi-source heterogeneous data according to claim 5, characterized in that, The step of determining historical reference data for the target data type in the target node based on historical pavement condition data and current pavement condition data includes: Determine whether the historical pavement condition at a given time and the current pavement condition at the present time are similar based on historical pavement condition data and current pavement condition data. When the historical and current operating conditions are similar, determine the historical reference data corresponding to the target data type in the target node at the historical moment.

7. The method for constructing an intelligent airport pavement maintenance knowledge graph based on multi-source heterogeneous data according to claim 1, characterized in that, The determination of the target source, reflecting the pavement damage evolution, based on target source data and historical reference data includes: Identify the differences between the target source data and historical reference data and use them as the target source to reflect the damage evolution of the pavement.

8. The method for constructing an intelligent airport pavement maintenance knowledge graph based on multi-source heterogeneous data according to claim 1, characterized in that, The determination of the data reasonableness of the target source based on the differential performance and damage evolution performance includes: The data rationality of the target source is obtained by combining the differential performance and damage evolution performance, where both differential performance and damage evolution performance are negatively correlated with the data rationality.

9. The method for constructing an intelligent airport pavement maintenance knowledge graph based on multi-source heterogeneous data according to claim 1, characterized in that, The confidence level of the target source data is determined based on the predicted data and the reasonableness of the data, including: Determine the prediction deviation between the predicted data and the target source data, and determine the difference in reasonableness between the reasonableness of the target source data and the reasonableness of the historical maximum data. The confidence level of the target source data is obtained by combining the prediction bias and the difference in reasonableness.

10. The method for constructing an intelligent airport pavement maintenance knowledge graph based on multi-source heterogeneous data according to claim 1, characterized in that, The process of determining the confidence level of the target source data based on the predicted data and the reasonableness of the data, and visualizing the confidence level, includes: The confidence levels of each source data are compared, and the source data with the highest confidence level that is greater than a preset confidence threshold is taken as the application data. The initial knowledge graph is updated based on the application data to obtain the target knowledge graph. Each node in the target knowledge graph is assigned a corresponding node fill color and edge style based on its confidence level for visualization purposes.

Citation Information

Patent Citations

  • Target action prediction method based on time sequence knowledge graph

    CN120724294A

  • Data analysis method and system fusing knowledge graph and deep learning

    CN121327780A