A knowledge graph construction and reasoning method and system for fault location

CN122779243APending Publication Date: 2026-09-18SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611223509.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-13
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0005]本发明的目的在于:针对现有技术中多源故障知识噪声较多、知识图谱关联路径强弱难以表达、和相似故障原因难以区分的技术问题,通过本发明实施例提供的一种用于故障定位的知识图谱构建与推理方法及系统,通过机理/数据融合方式挖掘并筛选汽车故障动因,将多维可信度转化为知识图谱边权,再基于带权知识图谱执行逻辑推理、路径推理和概率推理的逐层协同诊断,从而提高汽车故障诊断的准确性、稳定性和可解释性,并形成可工程化部署的智能诊断系统

Benefits of technology

1、本发明通过故障树分析和关联规则挖掘分别获得表示机理驱动关系和数据驱动关系的故障动因,通过计算机理路径的机理支持强度并进行初步筛选,然后根据来源权威度、统计显著性、上下文相关性和时效性形成的综合权重对故障动因进一步筛选,使得候选知识用于构建知识图谱时,能够抑制缺乏机理支持、统计关联较弱、场景适配程度较低或时效适用性较低的关联关系,提高知识图谱中有效故障知识的占比及知识图谱的构建质量。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122779243A_ABST
    Figure CN122779243A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of knowledge graph and automobile intelligent operation and maintenance, and particularly relates to a knowledge graph construction and reasoning method and system for fault positioning. The method comprises the following steps: obtaining multi-source fault data at least comprising historical maintenance records, and obtaining a candidate knowledge set through heterogeneous data preprocessing; extracting fault causes representing mechanism-driven relationships through fault tree analysis, and extracting fault causes representing data-driven relationships through association rule mining; calculating comprehensive weights according to the source authority, statistical significance, context correlation and timeliness of the fault causes, and performing screening; constructing a weighted triple based on the screened fault causes, the comprehensive weights and the hierarchical structure and logic gates of the fault tree, and constructing a knowledge graph for fault diagnosis. The application can inhibit noise relationships and weak correlation relationships, make the knowledge graph represent the credibility and importance of different fault relationships, and improve the knowledge graph construction quality and fault diagnosis accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent operation and maintenance and fault diagnosis technology, specifically to a knowledge graph construction and reasoning method and system for fault location. Background Technology

[0002] Fault diagnosis is a core research direction in the field of intelligent operation and maintenance. Its core objective is to automatically identify and accurately diagnose equipment fault states (such as engine vibration, difficulty starting, and weak acceleration in automobiles) through artificial intelligence and computer technology, providing core support for safe equipment operation and intelligent maintenance. With the development of artificial intelligence and big data technologies, equipment fault diagnosis has gradually evolved from relying on single inference based on human experience and static rules to knowledge graph reasoning that integrates multi-source fault data.

[0003] However, current knowledge graph construction focuses primarily on the identification and storage of entities and relationships. Before knowledge is entered into the graph, there is insufficient consideration given to filtering entities and relationships based on fault propagation mechanisms, data source reliability, and scenario adaptability. This results in the retention of noisy relationships with co-occurrence or local statistical correlations. Furthermore, the lack of characterization of the strength differences in entity relationships makes it difficult to highlight key relationships supporting the diagnostic chain, leading to numerous redundant relationships in the knowledge graph and inconsistencies in relationship priority and credibility.

[0004] When using knowledge graph reasoning, the logical coherence verification based on path reasoning makes it difficult to uncover implicit relationships that propagate across layers. Although path reasoning can discover multi-hop associations, it is easily affected by redundant and weak relationships, resulting in problems such as an overly broad range of candidate faults and unclear key diagnostic paths in the process of diagnosing fault phenomena to fault causes. While probabilistic reasoning can perform uncertainty assessment, the stability of probability-based ranking results is low in the absence of fault propagation mechanism constraints. This leads to insufficient confidence differentiation of multiple similar fault causes in complex fault scenarios, reducing the accuracy and stability of fault cause diagnosis. Summary of the Invention

[0005] The purpose of this invention is to address the technical problems in existing technologies, such as excessive noise in multi-source fault knowledge, difficulty in expressing the strength of knowledge graph association paths, and difficulty in distinguishing similar fault causes. This invention provides a knowledge graph construction and reasoning method and system for fault localization. It mines and filters automotive fault causes through mechanism / data fusion, transforms multi-dimensional credibility into knowledge graph edge weights, and then performs layer-by-layer collaborative diagnosis based on the weighted knowledge graph, involving logical reasoning, path reasoning, and probabilistic reasoning. This improves the accuracy, stability, and interpretability of automotive fault diagnosis and forms an intelligent diagnostic system that can be deployed in an engineered manner.

[0006] To achieve the above-mentioned objectives, the present invention provides the following technical solution: A knowledge graph construction method for fault localization includes: Acquire multi-source fault data, including at least historical maintenance records, and perform heterogeneous data preprocessing according to preset entity types to obtain a candidate knowledge set that includes at least the entity types of fault phenomenon, fault component, fault cause, and contextual information. The fault tree analysis method is used to analyze the hierarchical causal relationship between fault phenomena, faulty components and fault causes, convert the candidate knowledge set into a fault tree, and extract the fault drivers that represent the mechanism driving relationship from the fault tree. Based on the co-occurrence relationship between fault phenomena and faulty components or causes in historical maintenance records, fault drivers representing data-driven relationships are mined from candidate knowledge sets using association rule mining algorithms. A comprehensive weight is calculated based on the source authority, statistical significance, contextual relevance, and timeliness of the failure cause; the failure causes are then screened based on the comprehensive weight to obtain failure causes with a comprehensive weight greater than or equal to a preset weight threshold. Weighted triples are constructed based on fault causes and their corresponding comprehensive weights; logic gate triples are constructed based on the hierarchical structure and logic gates in the fault tree; the logic gate triples and their corresponding comprehensive weights are combined into weighted triples; and a knowledge graph for fault diagnosis is constructed based on the weighted triples.

[0007] The present invention also provides a knowledge graph construction system for fault location, comprising: One or more processors; A memory that stores one or more programs, which, when executed by one or more processors, enable the one or more processors to implement the knowledge graph construction method for fault location.

[0008] This invention also provides a knowledge graph reasoning method for fault location, comprising: Obtain fault information, obtain the constructed knowledge graph, map the fault information to the corresponding entity nodes in the knowledge graph, and map it to the corresponding event nodes in the fault tree; Entity sets are extracted from fault information, and hierarchical node sets are extracted from the mapped fault tree mechanism paths. Continuous matching scores are determined based on the matching relationship between entity sets and mechanism paths. Consistency scores are calculated based on the continuous matching scores of multiple mechanism paths. Fault tree mechanism paths are then filtered based on the consistency scores to obtain a primary candidate fault set. Based on the initial candidate fault set, we search the knowledge graph for the associated paths from the entity node corresponding to the fault information to the fault cause. We calculate the path confidence based on the comprehensive weight of each relation edge in the associated path and the number of hops in the associated path. Based on the maximum path length and the path confidence, we filter the associated paths to obtain candidate paths and candidate fault causes. Based on historical maintenance records and fault information, the posterior probability of each candidate fault cause is determined; the consistency score, path confidence, and posterior probability corresponding to each candidate fault cause are weighted and summed to obtain the inference score; the multiple candidate fault causes are ranked according to the inference score, and the top-ranked preset number of candidate fault causes are obtained as the fault diagnosis result.

[0009] The present invention also provides a knowledge graph reasoning system for fault location, comprising: One or more processors; A memory that stores one or more programs, which, when executed by one or more processors, enable the one or more processors to implement the knowledge graph reasoning method for fault location.

[0010] Compared with the prior art, the beneficial effects of the present invention include: 1. This invention obtains fault drivers representing mechanism-driven and data-driven relationships through fault tree analysis and association rule mining, respectively. It performs preliminary screening by calculating the mechanism support strength of the logical path, and then further screens the fault drivers based on a comprehensive weight formed by source authority, statistical significance, contextual relevance, and timeliness. This enables the candidate knowledge to suppress associations that lack mechanism support, have weak statistical correlations, low scenario adaptability, or low timeliness applicability when used to construct a knowledge graph, thereby increasing the proportion of effective fault knowledge in the knowledge graph and improving the construction quality of the knowledge graph.

[0011] 2. This invention maps fault information to entity nodes in a knowledge graph and event nodes in a fault tree. Through logical reasoning, it filters primary candidate faults based on the continuous matching relationship between fault information and fault tree mechanism paths, as well as logic gate constraints. Then, through path reasoning, it filters candidate paths and candidate fault causes based on the comprehensive weight of relation edges and the number of path hops. Finally, it determines the posterior probability of each candidate fault cause through probabilistic reasoning and integrates consistency score, path confidence, and posterior probability to obtain a reasoning score. This allows fault diagnosis to gradually converge through path filtering and posterior probability calibration, reducing interference from noisy relationships and weakly associated paths, improving the ability to distinguish similar fault causes, and enhancing the accuracy and stability of fault diagnosis. Attached Figure Description

[0012] Figure 1 A schematic diagram illustrating the process of constructing a knowledge graph for this invention; Figure 2 This is a schematic diagram illustrating the fault reasoning process of this invention using knowledge graphs. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention. It should be noted that, in the absence of conflict, the embodiments and features and technical solutions in the embodiments of the present invention can be combined with each other. It should be noted that similar reference numerals and letters in the following drawings indicate similar items; therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. The terms "first," "second," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance.

[0014] It should be noted that the knowledge graph construction and reasoning method for fault diagnosis provided by this invention can be used for fault diagnosis in different mechanical fields, including but not limited to mechanical equipment and systems with power modules such as automobiles, motorcycles, and generators. In this embodiment, fault diagnosis in the automotive field is taken as an example.

[0015] According to a first embodiment of the present invention, for example Figure 1 The diagram shown is a flowchart illustrating a knowledge graph construction method for fault location, as claimed in an embodiment of the present invention. The method includes: Step A1: Obtain multi-source fault data including at least historical maintenance records, perform heterogeneous data preprocessing according to preset entity types, and obtain a candidate knowledge set that includes at least the entity types of fault phenomenon, fault component, fault cause, and contextual information.

[0016] Specifically, the input data of this invention consists of multi-source fault data related to automotive fault diagnosis. Specific data sources include, but are not limited to, automotive repair manuals, fault code databases, historical repair work orders, warranty claim records, repair technician experience records, user repair requests, vehicle basic information, environmental information, and historical repair results. Vehicle basic information includes, but is not limited to, vehicle model, series, engine model, vehicle age, and mileage; environmental information includes, but is not limited to, season, weather, region, temperature, and road conditions; historical repair results include, but are not limited to, faulty components, fault causes, diagnostic methods, solutions, and repair result feedback. All data sources are used to collectively describe the fault characteristics, vehicle scenario, actual fault causes, and repair results. Historical repair work orders, warranty claim records, repair technician experience records, user repair requests, and historical repair results together form the historical repair record.

[0017] Based on the data sources mentioned above, multi-source fault data is heterogeneous data, including structured text, semi-structured text, and unstructured text. Structured text includes data organized using fixed data structures, such as fault code databases, repair work order fields, or basic vehicle information. Semi-structured text includes data such as automotive repair manuals, basic vehicle information, and warranty claim records; these data have hierarchical headings or fixed column structures, but also contain manually entered data without a fixed data structure. Unstructured text includes data with no fixed data structure, such as repair technician experience records and user repair requests.

[0018] Candidate knowledge sets are constructed based on preset entity types to extract corresponding data or text from multi-source fault data according to entity types, forming structured data. Each group of candidate knowledge in this set contains entity types including at least fault phenomenon, fault code, vehicle model, faulty component, fault cause, repair plan, and contextual information. The fault cause acts on the faulty component, causing one or more faulty components or their sub-components to experience system failures or component anomalies, thus manifesting as a fault phenomenon. Contextual information may include environmental information based on repair time, vehicle age, vehicle model, mileage, season, region, temperature, etc., from historical repair records. Missing contextual information is marked with a NULL field or other preset fields. Different types of multi-source fault data are preprocessed, using the candidate knowledge set as a data container. Each group of candidate knowledge in the set describes a fault semantic chain composed of a fault phenomenon, a faulty component, and a fault cause. For the same fault phenomenon, according to the one-to-one correspondence between the fault phenomenon and different faulty components and causes, the data required for each group of candidate knowledge is extracted from the preprocessed multi-source fault data and filled into the corresponding candidate knowledge, resulting in a candidate knowledge set containing multiple fault semantic chains.

[0019] In one embodiment, preprocessing includes: for structured text, using the Pandas data processing library to unify field names into entity type names according to preset field mapping relationships, while unifying measurement units and codes, and uniformly marking missing fields as NULL fields; for semi-structured text, using a document parsing tool to parse document content, and extracting hierarchical titles, fault code descriptions, diagnostic steps, and repair suggestions based on preset title hierarchy rules; for unstructured text, using the HanLP natural language processing library for word segmentation and entity recognition, and performing stop word filtering, terminology normalization, and synonym merging based on a preset stop word list, automotive fault terminology dictionary, and thesaurus. Terminology normalization is used to merge different text expressions pointing to the same fault semantic chain into the same entity type. For example, the entity types corresponding to "engine vibration," "idle vibration," and "unstable idle" are normalized to fault phenomena; the entity types corresponding to "spark plug," "fuel injector," "battery," and "crankshaft position sensor" are normalized to faulty components; and the entity types corresponding to "carbon buildup," "clogging," "aging," and "signal abnormality" are normalized to fault causes. After preprocessing, a candidate knowledge set is formed. Each group of candidate knowledge Including fault symptoms Fault codes Model Faulty components Cause of the malfunction Repair plan and context information Record the generation time of the candidate knowledge set or the most recent verification time of the validity of the candidate knowledge set, and associate the candidate knowledge set with the data source, generation time, or verification time.

[0020] Step A2: Analyze the hierarchical causal relationship between fault phenomena, faulty components and fault causes using the fault tree analysis method, convert the candidate knowledge set into a fault tree, and extract the fault drivers that represent the mechanism driving relationship from the fault tree.

[0021] Specifically, hierarchical causal relationships indicate that the underlying fault cause first acts on the corresponding faulty component, causing it to experience functional or state abnormalities, which in turn leads to system or component malfunctions, and then manifests as corresponding fault phenomena through the fault propagation process. Mechanism-driven relationships represent the propagation association between fault phenomena and faulty components or fault causes supported by the fault propagation mechanism, including the association formed by the path of "fault phenomenon—faulty component—fault cause," where the faulty component can include multiple levels. Fault phenomena, faulty components, and fault causes are obtained from the candidate knowledge set. Fault tree analysis (FTA) is used to describe the hierarchical causal relationships between automotive fault phenomena, faulty components, and fault causes as a hierarchical structure consisting of top events, intermediate events, and bottom events. Logic gates are set according to the causal relationships of the lower-level events corresponding to the same upper-level event, thereby constructing a mechanistic path from the fault cause through one or more intermediate events to the fault phenomenon.

[0022] Furthermore, the fault phenomenon is taken as the top event, the system fault or component abnormality corresponding to the faulty component that has a direct causal relationship with the fault phenomenon is taken as the intermediate event, and the fault cause that no longer propagates through the next level faulty component is taken as the bottom event. Logic gates are set according to the constraint relationship that multiple lower-level events corresponding to the same upper-level event need to satisfy together or any one of them, to obtain the fault tree, and the fault driving force representing the mechanism driving relationship is extracted from the mechanism path from the bottom event through the intermediate event to the top event in the fault tree.

[0023] Specifically, multiple sets of candidate knowledge related to the same fault phenomenon are obtained from the candidate knowledge set, denoted as the first screening set. The fault phenomenon is taken as the top event, and based on the automotive fault propagation mechanism, faulty components with a direct causal relationship to the top event are identified from the first screening set. The system fault or component abnormality corresponding to the faulty component is taken as the first-level intermediate event. For any intermediate event, the candidate knowledge set is further screened for faulty components associated with the intermediate event, and the corresponding system fault or component abnormality is taken as the next-level intermediate event, until a fault cause that no longer propagates through the next-level faulty component is determined. This fault cause is taken as the bottom event.

[0024] Based on the constraint relationship that multiple lower-level events corresponding to the same upper-level event need to satisfy together or any one of them, the connection relationship between multiple lower-level events and the upper-level event is mapped to logic gates such as AND gates or OR gates, thereby forming a multi-level fault tree consisting of a top event, at least one layer of intermediate events and multiple bottom events.

[0025] Furthermore, extracting fault drivers representing mechanistic driving relationships from the fault tree includes: forming a mechanistic path along the fault tree from any bottom event through intermediate events at each level to the top event; extracting fault drivers representing mechanistic driving relationships from the mechanistic path; obtaining the degree of confirmation of fault drivers based on statistical analysis of multiple data sources; calculating the degree of support of logic gates in the fault tree for the connected lower-level event branches; and statistically analyzing the path hierarchy length of the mechanistic path; determining the path length attenuation value based on the path hierarchy length and a preset path length attenuation coefficient; determining the mechanistic support strength of the mechanistic path based on the degree of confirmation, the degree of support, and the path length attenuation value; obtaining mechanistic paths with mechanistic support strength greater than or equal to a preset mechanistic support threshold; and forming a fault driver set from the corresponding fault drivers.

[0026] Specifically, along the fault tree, a complete mechanism path is formed from any bottom event through intermediate events at each level to the top event. Adjacent events at each level and the logic gates between them constitute a fault driver representing a mechanism-driven relationship. Multiple fault drivers form a fault driver set. For example, for each set of candidate knowledge in the candidate knowledge set, taking "engine starting difficulty" as an example, based on the automotive fault propagation mechanism recorded in the repair manual, engine starting difficulty may be caused by spark plug malfunction or fuel injector malfunction. Spark plug malfunction may be caused by carbon buildup or aging, while fuel injector malfunction may be caused by blockage. Therefore, the candidate knowledge formed from multi-source fault data can include: the first set of candidate knowledge includes at least the fault phenomenon "engine starting difficulty," the faulty component "spark plug," and the fault cause "carbon buildup"; the second set of candidate knowledge includes at least the fault phenomenon "engine starting difficulty," the faulty component "spark plug," and the fault cause "aging"; and the third set of candidate knowledge includes at least the fault phenomenon "engine starting difficulty," the faulty component "fuel injector," and the fault cause "blockage." Each group of candidate knowledge can also include the corresponding fault code, vehicle model, repair plan, and contextual information.

[0027] The aforementioned candidate knowledge is input into the fault tree analysis software. "Engine starting difficulty" is designated as the top event, and "spark plug" and "fuel injector" are designated as the two first-level intermediate events. "Carbon buildup" and "aging" of the spark plug are designated as bottom events, as are "clogging" of the fuel injector. According to the automotive fault propagation mechanism, either the bottom event "carbon buildup" or "aging" can cause spark plug malfunction; therefore, the bottom events "carbon buildup" and "aging" are connected to the intermediate event "spark plug" via an OR gate. "Clogging" can cause fuel injector malfunction; therefore, the bottom event "clogging" is connected to the intermediate event "fuel injector." Either the intermediate events "spark plug" or "fuel injector" can cause engine starting difficulty; therefore, they are connected to the top event via an OR gate. In other implementations, when multiple lower-level events need to occur simultaneously to cause a corresponding upper-level event, the multiple lower-level events are connected to the upper-level event via an AND gate; or, based on the causal relationship between the lower-level and upper-level events, they are connected via logic gates such as XOR gates and NOT gates. This forms a mechanism path, such as "carbon buildup - spark plug malfunction - engine starting difficulty", corresponding to fault causes supported by the automotive fault propagation mechanism, such as "carbon buildup, OR gate, spark plug malfunction" and "spark plug malfunction, OR gate, engine starting difficulty".

[0028] By forming mechanistic paths based on fault propagation mechanisms and extracting fault drivers that represent mechanistic driving relationships from these paths, all mechanistic-driven fault drivers entering the subsequent evaluation process can have clear fault propagation chains, reducing candidate associations formed solely due to data co-occurrence that do not conform to fault propagation mechanisms.

[0029] Furthermore, statistical analysis based on multiple data sources is used to obtain the degree of confirmation of the fault cause; the degree of support of the logic gates in the fault tree for the connected lower-level event branches is calculated; and the path level length of the statistical mechanism path is calculated; the path length attenuation value is calculated based on the path level length and the preset path length attenuation coefficient; the degree of confirmation, the degree of support and the path length attenuation value are multiplied together to obtain the mechanism support strength of the corresponding mechanism path.

[0030] Specifically, for failure drivers that point from bottom events to top events, the computer theory supports strength. To quantify the degree to which the fault propagation mechanism supports the fault cause, the calculation formula is as follows: ; in, This indicates the degree of confirmation of the fault cause by multiple data sources such as expert rules or maintenance manuals. It is obtained by statistically analyzing the number of fault causes under the same fault phenomenon and the total number of all fault causes under the same fault phenomenon, and then calculating the ratio of the two. The degree to which a fault tree logic gate supports the fault cause can be expressed as the structural importance calculated by fault tree analysis software or its normalized result; or, after obtaining the probability of the occurrence of the bottom event (e.g., statistically analyzing the probability of the occurrence of the bottom event under the same fault phenomenon in historical maintenance), the cut set importance, probability importance, or critical importance can be calculated by fault tree analysis software. This represents the number of path levels from the j-th bottom event to the top event, and λ is the path length decay coefficient.

[0031] Mechanism paths with a support strength greater than or equal to the support threshold are identified. Higher support strength indicates more sufficient support for mechanism propagation and logical structure, and a shorter propagation hierarchy from fault cause to fault phenomenon. Therefore, fault drivers with more sufficient mechanistic basis and more direct propagation relationship can be extracted from the selected mechanism paths, reducing pseudo-associations that obviously do not conform to fault propagation logic.

[0032] Step A3: Based on the co-occurrence relationship between fault phenomena and faulty components or causes in historical maintenance records, use association rule mining algorithms to mine fault drivers that represent data-driven relationships from the candidate knowledge set.

[0033] Specifically, data-driven relationships represent the co-occurrence associations between fault phenomena and faulty components or causes, based on statistics from historical maintenance records, maintenance manuals, etc. Historical maintenance records contain multiple records where fault phenomena, faulty components, and fault causes correspond to each other. With extensive data analysis, varying degrees of co-occurrence relationships emerge among these fault phenomena, causes, and components.

[0034] Furthermore, the association rule mining algorithm is used to mine fault drivers representing data-driven relationships from the candidate knowledge set, including: extracting data or text from each historical maintenance record according to a preset entity type, and converting the historical maintenance record into a set of transaction items; mining candidate association rules from the set of transaction items, using a combination of fault phenomenon, fault code, and context information as the first antecedent and the faulty component or fault cause as the first consequent; determining the first support, first confidence, and first lift of each candidate association rule based on the frequency of occurrence of the first antecedent and the first consequent in the historical maintenance record; obtaining candidate association rules whose first support, first confidence, and first lift are all greater than the corresponding preset thresholds, and adding the obtained candidate association rules as fault drivers representing data-driven relationships to the fault driver set.

[0035] Specifically, based on this co-occurrence relationship, data or text is extracted from each historical maintenance record according to a preset entity type, and the historical maintenance record is converted into a set of transaction items. This set of transaction items has the same entity type structure as the candidate knowledge set. Association rule mining algorithms such as the Apriori frequent itemset mining algorithm are used to mine association rules for the set of transaction items, resulting in multiple candidate association rules. Among them, the candidate association rules use the combination of fault phenomenon, fault code, and context information as the first antecedent X and the faulty component or fault cause as the first consequent Y.

[0036] For candidate association rules, the first support, first confidence, and first lift are calculated separately using the following formulas: ; ; ; Where N represents the total number of historical maintenance records; first support Reflects the coverage of candidate association rules across all historical maintenance records; first confidence level The first lift (X→Y) reflects the conditional probability of the first consequent Y given the first antecedent X. The first lift reflects the degree to which this conditional probability is enhanced relative to the overall probability of the first consequent Y. When the first lift is greater than 1, the first antecedent X has a positive reinforcing effect on the first consequent Y. The preset support threshold is 0.02; the confidence threshold is 0.6; and the lift threshold is 1.2. Candidate association rules that simultaneously satisfy the first support, first confidence, and first lift being greater than the corresponding three thresholds are recorded as valid association rules. These valid association rules represent the fault drivers of data-driven relationships, and these fault drivers are added to the fault driver set.

[0037] In another embodiment, for candidate drivers that only have statistical co-occurrence but lack propagation mechanism support, the present invention reduces their weight in subsequent multidimensional weighting; for candidate relationships that have both propagation mechanism support and statistical co-occurrence support, they are preferentially retained as highly reliable failure drivers.

[0038] By mining association rules from historical maintenance records using association rule mining algorithms, fault knowledge associations that were not fully extracted in fault tree analysis but have stable co-occurrence relationships can be extracted from historical maintenance records. Then, based on statistical analysis of historical maintenance records, support, confidence, and lift are used to filter out accidental co-occurrence relationships with low coverage, weak conditional associations, or lack of reinforcing effects, thereby improving the reliability of the fault cause set.

[0039] Step A4: Calculate the comprehensive weight based on the source authority, statistical significance, contextual relevance, and timeliness of the fault causes in the fault cause set; filter the fault cause set based on the comprehensive weight to obtain a set of fault causes with a comprehensive weight greater than or equal to the preset weight threshold.

[0040] Further, a comprehensive weight is calculated, and the set of fault drivers is filtered based on the comprehensive weight, including: merging the set of fault drivers representing mechanism-driven relationships with the set of fault drivers representing data-driven relationships to obtain a set of candidate drivers; determining the source authority based on the knowledge sources associated with the candidate fault drivers; determining the statistical significance based on the support, confidence, and lift of the candidate fault drivers in historical maintenance records; determining the contextual relevance based on the degree of matching between the contextual information corresponding to the candidate fault drivers and the contextual information of the current diagnosis; and determining the timeliness based on the generation time or the interval between the candidate knowledge corresponding to the candidate fault drivers and the current time; weighting the source authority, statistical significance, contextual relevance, and timeliness according to the preset dimension weights to obtain the comprehensive weight of the candidate fault drivers; and obtaining the candidate fault drivers whose comprehensive weight is greater than or equal to the preset weight threshold to obtain the filtered set of fault drivers.

[0041] Specifically, for each fault cause, the authority of its source is determined based on its knowledge origin; its statistical significance is determined based on its support, confidence, and lift in historical repair records; its contextual relevance is determined based on the degree of matching between its corresponding vehicle model, engine model, mileage, season, region, temperature, or repair scenario and the corresponding attributes of the vehicle to be diagnosed; and its timeliness is determined based on the interval between the knowledge generation time or the most recent verification time and the current time. The source authority, statistical significance, contextual relevance, and timeliness are weighted and summed to obtain the comprehensive weight of each fault cause. Fault causes with a comprehensive weight greater than or equal to a preset weight threshold are selected; alternatively, the fault cause set is sorted from largest to smallest comprehensive weight, and a preset number of fault causes at the top of the sorted list are selected to obtain a new fault cause set.

[0042] By uniformly evaluating and screening the fault drivers corresponding to mechanism-driven and data-driven relationships based on source authority, statistical significance, contextual relevance, and timeliness, the impact of fault drivers with low source reliability, weak statistical support, low matching degree with the current diagnostic scenario, and low timeliness applicability on the knowledge graph construction results can be reduced, thereby improving the credibility of fault drivers.

[0043] Furthermore, the authority of the knowledge source is determined based on its authority score; statistical significance is calculated based on support, confidence, and lift; contextual relevance is determined based on the degree of matching between contextual information and the contextual information corresponding to the vehicle to be diagnosed; timeliness is determined based on the interval between the knowledge generation time or the most recent verification time and the current time; and the comprehensive weight of the candidate fault drivers is obtained by weighting the source authority, statistical significance, contextual relevance, and timeliness.

[0044] Specifically, the set of failure drivers representing mechanism-driven relationships and the set of failure drivers representing data-driven relationships are merged through a union operation to obtain a candidate driver set. For each candidate failure driver in the candidate driver set, source authority, statistical significance, contextual relevance, and timeliness are calculated. These factors are then weighted and summed according to preset dimension weights to obtain the comprehensive weight of each candidate failure driver. The calculation formula is as follows: ; Wherein, α, β, γ, and δ are preset dimension weights, and α+β+γ+δ=1, and α, β, γ, and δ are all greater than or equal to 0. W represents the comprehensive weight of the candidate fault causes; A is used to measure the reliability of the knowledge source; S is used to measure the statistical significance of the candidate fault causes in historical maintenance records; C is used to measure the degree of matching between the candidate fault causes and the current vehicle and usage environment; and T is used to measure the timeliness and applicability of the candidate fault causes.

[0045] The comprehensive weight simultaneously characterizes the reliability of knowledge sources, the degree of support from historical data for candidate fault drivers, the matching degree between candidate fault drivers and the vehicle to be diagnosed and its usage environment, and the timeliness and applicability of the knowledge. The candidate driver set is sorted according to the comprehensive weight, retaining those with a comprehensive weight greater than or equal to a preset comprehensive threshold, or retaining a preset number of candidate fault drivers at the top of the ranking. Therefore, candidate fault drivers that do not match vehicle model, engine model, mileage, season, region, or maintenance scenario, or those with weak statistical support, low source reliability, or early occurrence, will receive lower comprehensive weights, thus reducing their impact on the knowledge graph construction results.

[0046] Furthermore, the authoritative scores of the knowledge sources associated with the candidate fault causes are obtained. Using the highest authoritative score among the multiple knowledge sources participating in the evaluation as the benchmark, all authoritative scores are normalized to obtain the source authority. The calculation formula is as follows: ; in, An authoritative rating indicating the source of knowledge. This represents the highest authority score among multiple knowledge sources used in the evaluation. The authority score is obtained by technical experts scoring different data sources, including repair manuals, fault code databases, historical repair work orders, repair technician experience records, and user repair reports.

[0047] Furthermore, using the current fault phenomenon and contextual information of the vehicle to be diagnosed as the second antecedent, and the candidate fault cause as the second consequent, the co-occurrence relationship between the second antecedent and the second consequent is statistically analyzed based on historical maintenance records, and the second support, second confidence, and second lift are calculated. The second support and second lift are normalized, and the normalized second support, second confidence, and normalized second lift are weighted and summed according to preset statistical index weights to obtain statistical significance. The calculation formula is as follows: ; in, This represents the second support after normalization. This represents the second lift after normalization; , and The weights of the statistical indicators are represented, and the sum of the three is 1. By combining the second support, the second confidence, and the second lift, the statistical correlation strength of different candidate fault drivers can be distinguished. When using knowledge graphs for reasoning, this can be used to prevent high-frequency faults from being mistakenly considered to have a strong correlation simply because they occur frequently, and to suppress candidate fault drivers that only appear occasionally in a very small number of historical maintenance records.

[0048] Furthermore, for any candidate fault driver in the fusion candidate fault driver set, contextual information is obtained from the candidate knowledge associated with that candidate fault driver, and the current contextual information of the vehicle to be diagnosed at the time of the fault occurrence is also obtained. Corresponding contextual information such as vehicle model, engine model, mileage, season, region, temperature, and repair scenario are extracted from these two types of contextual information. The attribute similarity between corresponding contextual information is calculated. Using the importance of the contextual information as the attribute weight for attribute similarity, a weighted average of multiple attribute similarities is calculated to obtain the contextual relevance. The calculation formula is as follows: ; in, This represents the k-th context information corresponding to the candidate fault cause; This represents the k-th context information corresponding to the vehicle to be diagnosed; This represents an attribute similarity function, such as direct matching or semantic similarity. The importance of the k-th contextual information is indicated by technical experts based on the frequency of its occurrence in historical maintenance records. When constructing the knowledge graph, calculating the attribute similarity between the contextual information in candidate fault causes and the contextual information during the current fault diagnosis helps in selecting candidate fault causes that are more relevant to the current diagnosis and usage scenario.

[0049] In one implementation, different attribute similarity functions are used based on the different value types of the context information. For example, for discrete context information such as vehicle model, engine model, season, region, and maintenance scenario, direct matching is used. When the corresponding attribute values ​​are the same, the attribute similarity is set to one, and when the corresponding attribute values ​​are different, the attribute similarity is set to zero. For continuous context information such as mileage and temperature, the attribute similarity is determined by calculating the difference between two corresponding attribute values ​​using distance similarity. The smaller the difference, the greater the attribute similarity.

[0050] Furthermore, the generation time or most recent verification time of the candidate knowledge corresponding to the candidate fault cause is obtained, denoted as the first time, where the most recent verification time is the time when the candidate knowledge most recently passed validity verification; the time interval between the current time and the first time is calculated, and the timeliness T is calculated based on the time decay function, which reduces timeliness as the time interval increases. The formula for calculating the time decay function is as follows: ; Where Δt represents the time interval between the current time and the first time when the vehicle to be diagnosed is diagnosed; τ represents the time decay coefficient, which gives higher timeliness weight to recent maintenance records or recent verification knowledge and reduces the interference of outdated models and outdated fault modes on the current diagnostic results.

[0051] Step A5: Construct weighted triples based on the fault causes and their corresponding comprehensive weights in the fault cause set; construct logic gate triples based on the hierarchical structure and logic gates in the fault tree, and combine the logic gate triples and their corresponding comprehensive weights into weighted triples; construct a knowledge graph for fault diagnosis based on the weighted triples.

[0052] Furthermore, a knowledge graph for fault diagnosis is constructed based on weighted triples, including: extracting head entities, tail entities, and relation types connecting head and tail entities from candidate knowledge corresponding to fault causes; using the comprehensive weight of fault causes as the attribute of the corresponding relation edge to obtain weighted triples; using the top event, intermediate event, and bottom event in the fault tree as entity nodes; converting the logic gate connections between adjacent levels such as top event to intermediate event and intermediate event to bottom event into relation edges; using the logic gates as the attribute of the corresponding relation edge to obtain logic gate triples; using the comprehensive weight corresponding to the fault cause representing the mechanism-driven relationship as the attribute of the corresponding relation edge in the logic gate triple to convert the logic gate triple into weighted triples; and constructing a knowledge graph based on the entity nodes, relation edges, relation types, and comprehensive weights in each weighted triple.

[0053] Specifically, candidate fault drivers in the fault driver set are transformed into weighted triples e, which include a head entity, a relation type, and a tail entity structure, and the overall weight is used as one of the attributes of the relation edge. The knowledge graph G is represented as: G = (V, E, R, W); ; Where V represents the set of entity nodes, E represents the set of relation edges, R represents the set of relation types, and W represents the set of comprehensive weights; Indicates the head entity. Let 'r' represent the tail entity, 'r' represent the relation type, and 'w' represent the comprehensive weight corresponding to the relation edge. When constructing the knowledge graph, the comprehensive weight is written into an attribute of the corresponding relation edge, enabling the knowledge graph to not only represent whether there is a semantic connection between two entities, but also the credibility and diagnostic priority of the fault relationship. Entity nodes include fault symptoms, fault codes, vehicle models, faulty components, fault causes, solutions, and contextual information. Relationship types are extracted from historical repair records containing text describing the association between two entities. These types describe the connection between fault symptoms and faulty components, or faulty components and fault causes, or fault symptoms and fault causes, including: manifestation, possible cause, belonging to, associated, requiring diagnosis, corresponding solution, and environmental influence.

[0054] Simultaneously, the fault tree is also embedded into the knowledge graph, with top, intermediate, and bottom events as entity nodes. Hierarchical connections from top to intermediate events and from intermediate to bottom events are transformed into relational edges. Logic gates are converted into attributes of relational edges between different types of events. This transforms adjacent-level event nodes in the mechanism path into logic gate triples, for example... Here, p is the logical gate attribute of relation type r, and its values ​​include AND and OR gates. The fault cause based on the mechanism-driven relation is a path in the knowledge graph. The comprehensive weight of this fault cause is also written as the weight attribute of the relation edge, resulting in the weighted triple corresponding to the fault tree in the knowledge graph. After converting the fault tree to a knowledge graph, a weighted fault knowledge graph is formed that combines mechanistic constraints and relational weighting capabilities.

[0055] By incorporating the hierarchical connections of the fault tree, logic gate constraints, and comprehensive weights of fault causes into the knowledge graph, the same graph can simultaneously retain the fault propagation mechanism and the strength of relationship credibility, thereby guiding the direction of diagnostic path search and providing a consistent data foundation for logical reasoning, path reasoning, and probabilistic reasoning.

[0056] According to a second embodiment of the present invention, the present invention also claims protection for a knowledge graph construction system for fault location, comprising: One or more processors; A memory that stores one or more programs, which, when executed by one or more processors, enable one or more processors to implement a knowledge graph construction method for fault location.

[0057] According to a third embodiment of the present invention, for example Figure 2 The diagram shown is a flowchart illustrating a knowledge graph reasoning method for fault location, as claimed in an embodiment of the present invention. The method includes: Step B1: Obtain the fault information of the vehicle to be diagnosed, map the fault information to the corresponding entity node in the knowledge graph, and map it to the corresponding event node in the fault tree.

[0058] Specifically, the system acquires fault information about the vehicle to be diagnosed, including fault symptoms, fault codes, and contextual information composed of related data such as vehicle model, mileage, vehicle age, weather, and region. Using the same terminology normalization method as when constructing the knowledge graph, the fault symptoms, fault codes, and contextual information are mapped to corresponding entity nodes in the knowledge graph and corresponding event nodes in the fault tree.

[0059] Step B2: Extract the entity set from the fault information and extract the hierarchical node set from the mapped fault tree mechanism path. Determine the continuous matching score based on the matching relationship between the entity set and the mechanism path. Calculate the consistency score based on the continuous matching scores of multiple mechanism paths. Filter the fault tree mechanism path based on the consistency score to obtain the primary candidate fault set.

[0060] Further, a consistency score is calculated, and fault tree mechanism paths are screened based on the consistency score. This includes: extracting multiple entities from the fault information according to entity type to obtain an entity set; extracting multiple hierarchical nodes from each mapped fault tree mechanism path to obtain a corresponding hierarchical node set; determining the continuous matching score of each fault tree mechanism path based on the ratio of the intersection size of the entity set and each hierarchical node set to the size of the corresponding hierarchical node set; for AND gates in the fault tree, determining the consistency score based on the continuous matching scores of multiple fault tree mechanism paths that need to be satisfied; for OR gates in the fault tree, determining the consistency score based on the continuous matching score corresponding to at least one fault tree mechanism path; obtaining fault tree mechanism paths with a consistency score greater than or equal to a preset logical screening threshold, and forming a primary candidate fault set by combining the candidate fault drivers corresponding to the obtained fault tree mechanism paths.

[0061] Specifically, multiple entities are extracted from the current fault information according to entity type to obtain an entity set, including the fault phenomenon name or synonym obtained after terminology normalization, fault codes, and contextual information text. The mapped fault tree mechanism path is obtained from the fault tree, and multiple hierarchical nodes are extracted from the fault tree mechanism path to obtain a hierarchical node set, including the standard name or synonym of the fault tree node, the associated fault code in the corresponding candidate knowledge, and contextual information. The matching relationship between fault information and the fault tree mechanism path is quantified by the ratio of the intersection size of the entity set and the hierarchical node set to the size of the hierarchical node set, as shown in the following formula: ; Where m(e, v) represents the continuous matching score between the current fault information and the mapped mechanism path, and its value range is [0, 1]. Represents a set of entities; This represents the set of hierarchical nodes. When the text of the fault phenomenon in the entity set completely matches the node name, fault code, or context information similarity exceeds the threshold, the continuous matching score is closer to 1. If the text description of the fault phenomenon is incomplete, or only partially matches after terminology normalization, the continuous matching score is between 0 and 1. When the fault information, fault code, and context information do not match, the matching score is closer to 0.

[0062] In this embodiment, taking AND and OR gates in the fault tree as examples, the consistency score of the mapped multiple mechanism paths is calculated using the following formula: ; ; in, This represents the consistency score when multiple mechanistic paths under an AND constraint need to be satisfied together. A decrease in the consecutive matching score of any mechanistic path will decrease the consistency score of the AND constraint. A consistency score is formed for supporting the event at the next higher level if at least one mechanistic path under the OR gate constraint is satisfied. An increase in the consecutive matching score of any mechanistic path will improve the consistency score of the OR gate constraint. Fault tree mechanistic paths with a consistency score greater than or equal to a preset logical filtering threshold are selected. The candidate fault drivers corresponding to multiple mechanistic paths form a primary candidate fault set.

[0063] Step B3: Based on the initial candidate fault set, search the knowledge graph for the associated paths from the entity node corresponding to the fault information to the fault cause. Calculate the path confidence based on the comprehensive weight of each relation edge in the associated path and the number of hops in the associated path. Filter the associated paths based on the maximum path length and the path confidence to obtain candidate paths and corresponding candidate fault causes.

[0064] Furthermore, the process of filtering associated paths based on the maximum path length and path confidence includes: searching for associated paths from the entity node corresponding to the fault information to the fault cause corresponding to the primary candidate fault set; determining the geometric mean of the comprehensive weight of each relation edge in the associated path, and determining the path length penalty result based on the number of hops in the associated path; combining the geometric mean with the path length penalty result to obtain the path confidence; obtaining associated paths whose path length is less than or equal to the maximum path length and whose path confidence is greater than or equal to the preset path confidence threshold as candidate paths; and when the same candidate fault cause corresponds to multiple associated paths, obtaining the associated path with the highest path confidence.

[0065] Specifically, after obtaining a preliminary candidate fault set, a multi-hop association path from the fault phenomenon, fault code, or contextual information to the fault cause is searched in the knowledge graph based on the preliminary candidate fault set. In this embodiment, the path search can be implemented using breadth-first search, weighted shortest path search, or other graph search algorithms.

[0066] In a knowledge graph, an entity can be connected to multiple other entities through different relationships, and the number of associated paths increases rapidly with the search level. Paths containing edges with low overall weight relationships have low support for candidate fault causes and are considered weakly associated paths. To address the problems of the proliferation of knowledge-related paths and interference from weakly associated paths, a maximum path length limit strategy and a path confidence filtering strategy are introduced during path search. Let's consider an associated path... Path confidence of associated paths The calculation formula is as follows: ; in, This represents the number of hops in the associated path; For the first The overall weight of the edges; This represents the geometric mean of all the combined weights on the associated paths of the search, used to avoid excessive penalties for long paths due to the increase in relational edges caused by simple multiplication. The penalty term is the number of relation edges. This is the penalty coefficient. The path confidence is calculated using the above formula, ensuring that the path confidence of associated paths with higher overall weights and shorter paths is higher, while the path confidence of associated paths formed by weak relationships or excessively long hop counts is lower. In another implementation, for the same fault cause, if multiple associated paths exist, the path with the highest confidence is selected.

[0067] The hop count is the number of nodes in the associated path. Associated paths with a hop count less than or equal to the maximum path strength and a path confidence greater than or equal to the path confidence threshold are identified as candidate paths. The fault causes under these candidate paths are identified as candidate fault causes.

[0068] Step B4: Determine the posterior probability of each candidate fault cause based on historical maintenance records and fault information; perform a weighted summation of the consistency score, path confidence, and posterior probability corresponding to each candidate fault cause to obtain the inference score; rank the multiple candidate fault causes according to the inference score, and obtain the preset number of candidate fault causes ranked first as the fault diagnosis result.

[0069] Specifically, after obtaining candidate fault causes and their associated paths through logical reasoning and path reasoning, the probabilistic evaluation of candidate fault causes is performed by comprehensively considering features such as historical co-occurrence frequency, knowledge graph edge weights, logical consistency score, path confidence, context matching degree, and maintenance result feedback, calculating the posterior probability of each candidate fault cause. The prior probability of a candidate fault cause is determined based on the number of times it appears in historical maintenance records, using the following formula: ; in, This represents the number of times a candidate fault cause appears in historical maintenance records; N represents the total number of historical maintenance records. Given a set of fault information... The posterior probability of the candidate fault cause is updated according to the Bayesian formula, as follows: ; in, This indicates that fault information occurs when the candidate fault cause r is true. The conditional probability is calculated based on the co-occurrence frequency of corresponding fault information and candidate fault causes in historical maintenance records. The statistical estimate is then corrected based on the combined weight of the corresponding edges in the knowledge graph, the context matching score, or the support level of expert rules. For candidate fault causes with fewer historical maintenance records, the conditional probability is calculated using the following formula. : ; in, This represents the number of times fault information and candidate fault causes co-occur in historical maintenance records; count(r) represents the number of times a candidate fault cause occurs; |E| represents the number of elements in the fault information set; k is a smoothing coefficient used to avoid low-frequency fault causes, whose conditional probabilities are zero due to a small number of records. Multiple candidate fault causes with similar fault phenomena and small differences in the number of hops in their associated paths have similar path confidence, making accurate screening based solely on path confidence difficult. Updating the posterior probability of candidate fault causes using Bayes' theorem helps to further distinguish different candidate fault causes and improve the precision of fault diagnosis.

[0070] Furthermore, the consistency score obtained through logical reasoning, the path confidence obtained through path reasoning, and the posterior probability obtained through probabilistic reasoning are weighted and summed to obtain the reasoning score. Multiple candidate fault causes are then sorted from highest to lowest according to their reasoning scores. A predetermined number of the top-ranked candidate fault causes are selected as the fault diagnosis results. Simultaneously, the associated faulty component and reasoning path for each candidate fault cause are determined. The user is then presented with the top-ranked fault cause, a list of other candidate fault causes, path confidence, output results at each stage, and suggested repair solutions to support repair personnel in quickly diagnosing fault causes.

[0071] According to a fourth embodiment of the present invention, the present invention also claims protection for a knowledge graph reasoning system for fault location, comprising: One or more processors; A memory that stores one or more programs, which, when executed by one or more processors, enable one or more processors to implement a knowledge graph reasoning method for fault location.

[0072] The above embodiments are only used to illustrate the present invention and are not intended to limit the technical solutions described herein. Although the present invention has been described in detail with reference to the above embodiments, the present invention is not limited to the specific embodiments described above. Therefore, any modifications or equivalent substitutions to the present invention, as well as all technical solutions and improvements that do not depart from the spirit and scope of the invention, are covered within the scope of the claims of the present invention.

Claims

1. A knowledge graph construction method for fault location, characterized in that, include: Acquire multi-source fault data, including at least historical maintenance records, and perform heterogeneous data preprocessing according to preset entity types to obtain a candidate knowledge set that includes at least the entity types of fault phenomenon, fault component, fault cause, and contextual information. The fault tree analysis method is used to analyze the hierarchical causal relationship between fault phenomena, faulty components and fault causes, convert the candidate knowledge set into a fault tree, and extract the fault drivers that represent the mechanism driving relationship from the fault tree. Based on the co-occurrence relationship between fault phenomena and faulty components or causes in historical maintenance records, fault drivers representing data-driven relationships are mined from candidate knowledge sets using association rule mining algorithms. The comprehensive weight is calculated based on the authority of the source of the failure cause, statistical significance, contextual relevance, and timeliness. The fault causes are screened based on the comprehensive weight, and the fault causes with a comprehensive weight greater than or equal to the preset weight threshold are obtained. Weighted triples are constructed based on the fault causes and their corresponding comprehensive weights; Based on the hierarchical structure and logic gates in the fault tree, logic gate triples are constructed, and the logic gate triples and their corresponding comprehensive weights are combined into weighted triples; a knowledge graph for fault diagnosis is constructed based on the weighted triples.

2. The knowledge graph construction method for fault location according to claim 1, characterized in that, Extracting fault drivers representing mechanistic driving relationships from the fault tree includes: A mechanism path is formed by tracing the fault tree from any bottom event through intermediate events at each level to the top event, and the fault drivers representing the mechanism driving relationship are extracted from the mechanism path. The degree of confirmation of the fault cause is obtained through statistical analysis based on multiple data sources. The degree of support of the logic gates in the fault tree for the connected lower-level event branches is calculated, and the path hierarchy length of the mechanism path is statistically analyzed. The path length attenuation value is determined based on the path hierarchy length and the preset path length attenuation coefficient. The mechanism support strength of the mechanism path is determined based on the degree of confirmation, the degree of support, and the path length attenuation value. Obtain mechanistic paths whose mechanistic support strength is greater than or equal to a preset mechanistic support threshold, and form a set of corresponding fault causes.

3. The knowledge graph construction method for fault location according to claim 1, characterized in that, Fault drivers representing data-driven relationships are extracted from candidate knowledge sets using association rule mining algorithms, including: Extract data or text from each historical maintenance record according to the preset entity type, and convert the historical maintenance record into a set of transaction items; Using the combination of fault phenomenon, fault code and context information as the first antecedent and the faulty component or fault cause as the first consequent, candidate association rules are mined from the transaction item set. Based on the frequency of occurrence of the first antecedent and the first consequent in historical maintenance records, determine the first support, first confidence, and first lift of each candidate association rule; Candidate association rules with first support, first confidence, and first lift all greater than the corresponding preset thresholds are obtained, and the obtained candidate association rules are added to the fault cause set as fault causes representing data-driven relationships.

4. The knowledge graph construction method for fault location according to claim 1, characterized in that, The set of failure drivers is filtered based on a comprehensive weighting, including: The set of fault causes representing mechanism-driven relationships is merged with the set of fault causes representing data-driven relationships to obtain a candidate set of causes. The authority of the source is determined based on the knowledge source associated with the candidate fault cause; the statistical significance is determined based on the support, confidence and lift of the candidate fault cause in the historical maintenance records; the contextual relevance is determined based on the degree of matching between the contextual information corresponding to the candidate fault cause and the contextual information of the current diagnosis; and the timeliness is determined based on the generation time or the interval between the candidate knowledge corresponding to the candidate fault cause and the current time. The comprehensive weight of the candidate fault cause is obtained by weighting the source authority, statistical significance, contextual relevance and timeliness according to the preset dimension weights. Candidate fault causes with a comprehensive weight greater than or equal to a preset weight threshold are obtained, resulting in a set of filtered fault causes.

5. The knowledge graph construction method for fault location according to claim 1, characterized in that, A knowledge graph for fault diagnosis is constructed based on weighted triples, including: Extract the head entity, tail entity, and relation type connecting the head entity and tail entity from the candidate knowledge corresponding to the fault cause. Use the comprehensive weight of the fault cause as the attribute of the corresponding relation edge to obtain weighted triples. The top, middle, and bottom events in the fault tree are treated as entity nodes. The hierarchical connections from the top event to the middle event and from the middle event to the bottom event are converted into relational edges. The logic gates are used as attributes of the corresponding relational edges to obtain logic gate triples. The comprehensive weight corresponding to the fault cause representing the mechanism-driven relationship is used as the attribute of the corresponding relation edge in the logic gate triplet, and the logic gate triplet is converted into a weighted triplet. A knowledge graph is constructed based on the entity nodes, relation edges, relation types, and comprehensive weights in each weighted triple.

6. A knowledge graph construction system for fault location, characterized in that, include: One or more processors; A memory having stored one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the knowledge graph construction method for fault location as described in any one of claims 1-5.

7. A knowledge graph reasoning method for fault location, characterized in that, include: Obtain fault information, obtain the knowledge graph constructed according to any one of claims 1-6, and map the fault information to the corresponding entity node in the knowledge graph and the corresponding event node in the fault tree; Entity sets are extracted from fault information, and hierarchical node sets are extracted from the mapped fault tree mechanism paths. Continuous matching scores are determined based on the matching relationship between entity sets and mechanism paths. Consistency scores are calculated based on the continuous matching scores of multiple mechanism paths. Fault tree mechanism paths are then filtered based on the consistency scores to obtain a primary candidate fault set. Based on the initial candidate fault set, we search the knowledge graph for the associated paths from the entity node corresponding to the fault information to the fault cause. We calculate the path confidence based on the comprehensive weight of each relation edge in the associated path and the number of hops in the associated path. Based on the maximum path length and the path confidence, we filter the associated paths to obtain candidate paths and candidate fault causes. The posterior probability of each candidate fault cause is determined based on historical maintenance records and fault information. The consistency score, path confidence, and posterior probability corresponding to each candidate fault cause are weighted and summed to obtain the inference score; the multiple candidate fault causes are ranked according to the inference score, and the top-ranked preset number of candidate fault causes are obtained as the fault diagnosis result.

8. The knowledge graph reasoning method for fault location according to claim 7, characterized in that, Calculate the consistency score, and filter the fault tree mechanism path based on the consistency score, including: Multiple entities are extracted from the fault information according to entity type to obtain an entity set; multiple hierarchical nodes are extracted from each mapped fault tree mechanism path to obtain the corresponding hierarchical node set. The continuous matching score of each fault tree mechanism path is determined based on the ratio of the size of the intersection of the entity set and the node set at each level to the size of the corresponding node set. For AND gates in a fault tree, the consistency score is determined by the continuous matching scores of multiple fault tree mechanism paths that need to be satisfied together; for OR gates in a fault tree, the consistency score is determined by the continuous matching scores corresponding to at least one fault tree mechanism path. Obtain fault tree mechanism paths with consistency scores greater than or equal to a preset logical screening threshold, and form a primary candidate fault set by combining the candidate fault causes corresponding to the obtained fault tree mechanism paths.

9. The knowledge graph reasoning method for fault location according to claim 8, characterized in that, Filtering associated paths based on maximum path length and path confidence includes: Search for the associated path from the entity node corresponding to the fault information to the fault cause corresponding to the primary candidate fault set; Determine the geometric mean of the combined weights of each relation edge in the associated path, and determine the path length penalty result based on the number of hops in the associated path. Combine the geometric mean with the path length penalty result to obtain the path confidence. Retrieve associated paths whose path length is less than or equal to the maximum path length and whose path confidence is greater than or equal to the preset path confidence threshold as candidate paths; when the same candidate fault cause corresponds to multiple associated paths, retrieve the associated path with the highest path confidence.

10. A knowledge graph reasoning system for fault location, characterized in that, include: One or more processors; A memory having stored one or more programs, which, when executed by one or more processors, cause the one or more processors to implement a knowledge graph reasoning method for fault location as described in any one of claims 7-9.