Method and device for carrying out fault analysis by utilizing knowledge graph
By using knowledge graph technology to integrate multi-source heterogeneous data in equipment failure analysis, the problems of data dispersion and insufficient correlation are solved, and efficient and accurate fault analysis results are achieved.
Patent Information
- Application Number
- CN202311610107.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-05-27
AI Technical Summary
In equipment failure analysis, the dispersion and lack of correlation of multi-source heterogeneous data lead to difficulties in data integration and management, which in turn affects the efficiency and accuracy of fault analysis.
Knowledge graph technology is used to integrate multi-source heterogeneous data from different data domains, and data sharing and interaction are realized by building integrated knowledge graphs, thus supporting efficient failure analysis.
Through the application of knowledge graph technology, the complexity and time-consuming of the ETL process are reduced, semantic interoperability conflicts are avoided, and the efficiency and accuracy of fault analysis are significantly improved.
Smart Images

Figure CN120045366A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of data analysis, and more particularly to methods and apparatuses for fault analysis using a knowledge graph. Background Art
[0002] It has been noted that various faults may occur during the on-site application of a device, resulting in the device losing its specified function or operating with a degraded specified function. Such faults may also be referred to as on-site faults, on-site anomalies, and so on. For a device manufacturer, when the same or similar on-site faults occur repeatedly in the same type of device, it is necessary to collect various data associated with the on-site faults in a timely manner and perform fault analysis based on the collected data. Fault analysis aims to determine the causes of on-site faults, so as to formulate corresponding strategies to avoid the continuous occurrence of similar fault problems. Summary of the Invention
[0003] It is generally advantageous to take a large amount of relevant data into account during the fault analysis process, because more relevant data helps to analyze on-site faults more comprehensively and fully. However, these relevant data may include homogeneous and heterogeneous data from different data domains, making the data analysis for on-site faults face the severe challenge of integrating homogeneous and heterogeneous data together. Therefore, it is desirable to provide an improved mechanism to efficiently integrate homogeneous and heterogeneous data associated with on-site faults, so as to facilitate obtaining analysis results for on-site faults efficiently and accurately.
[0004] According to one aspect of the present disclosure, there is provided a method for fault analysis using a knowledge graph, including: obtaining original data associated with an on-site fault of a device, wherein the original data comes from multiple data domains; constructing an integrated knowledge graph for the multiple data domains; and performing fault analysis based on the integrated knowledge graph.
[0005] According to another aspect of the present disclosure, there is provided a system for fault analysis using a knowledge graph, including: a data collection and preprocessing module for obtaining original data associated with an on-site fault of a device, wherein the original data comes from multiple data domains; a data integration module based on a knowledge graph for constructing an integrated knowledge graph for the multiple data domains; and an application module for performing fault analysis based on the integrated knowledge graph.
[0006] According to still another aspect of the present disclosure, there is provided a device for fault analysis using a knowledge graph, including: a memory; and a processor. The processor is coupled to the memory and is configured to execute the method according to any one of the various embodiments of the present disclosure.
[0007] According to another aspect of the present disclosure, there is provided a computer-readable medium storing a computer program including instructions that, when executed by a processor, cause the processor to be configured to perform the method according to any one of the various embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The various embodiments of the claimed subject matter will now be described by way of example with reference to the accompanying drawings. In the different drawings, the same reference numerals are used to denote the same or similar components.
[0009] Figure 1 A schematic diagram showing the construction of an integrated knowledge graph for multiple data domains according to an embodiment of the present disclosure is shown.
[0010] Figure 2 A flowchart showing a method for constructing a corresponding domain knowledge graph for each data domain according to an embodiment of the present disclosure is shown.
[0011] Figure 3 A schematic diagram showing the formation of an integrated knowledge graph based on multiple constructed domain knowledge graphs according to an embodiment of the present disclosure is shown.
[0012] Figure 4 A flowchart showing a method for performing fault analysis using a knowledge graph according to an embodiment of the present disclosure is shown.
[0013] Figure 5 A flowchart showing a system for performing fault analysis using a knowledge graph according to an embodiment of the present disclosure is shown.
[0014] Figure 6 A schematic diagram showing the architecture of a computing device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0015] In the following description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure. However, those skilled in the relevant art will recognize that the present disclosure may be practiced without one or more of the specific details, or may be practiced using alternative methods, components, etc. In some instances, well-known structures, operations are not shown or described in detail so as not to unnecessarily obscure the present disclosure.
[0016] The present disclosure proposes an improved mechanism that utilizes a knowledge graph to integrate homogeneous and heterogeneous data from multiple data domains associated with on-site faults of a device to facilitate obtaining analysis results for on-site faults efficiently and accurately.
[0017] In the present disclosure, a device may refer to any product, component, sub-component, part, etc. that is designed to provide an intended function, and it is not limited to any field or industry. For example, the field of mechanical manufacturing, the field of transportation, the field of construction, etc. In the following detailed description, for ease of understanding, the steering gear of a vehicle is used as a specific example of the above device to describe the content related to constructing an integrated knowledge graph for an efficient fault analysis process.
[0018] The on-site failure of a device, which may also be referred to as the on-site anomaly of the device, generally refers to the situation where the device loses its specified function or operates with a degraded specified function during on-site application. When the same or similar on-site failures occur repeatedly for the same type of device, it is necessary to timely collect various data associated with the on-site failure and perform a fault analysis based on the collected data. In one example, the fault analysis in the present disclosure may specifically refer to root cause analysis (RCA), which aims to determine the root cause leading to the failure and can form a complete data chain describing the failure-cause-consequence to guide the formulation of corresponding countermeasures for the on-site failure.
[0019] It is generally beneficial to take a large amount of relevant data into account during the fault analysis process because more relevant data helps to analyze the on-site failure more comprehensively and sufficiently, so that various potential factors leading to the on-site failure can be accurately and comprehensively determined. The applicant notes that although some prior art solutions have proposed to perform data analysis for on-site failures based on specific types of data, however, these data types are not comprehensive or do not analyze the data in an interrelated manner. The present disclosure proposes to perform a fault analysis by comprehensively considering more data associated with the on-site failure, thereby improving the accuracy of the fault analysis result.
[0020] However, introducing a large amount of relevant data may face the problem of "data silos" during the data analysis process. The "data silos" problem is caused by the dispersion of data or the lack of correlation between data. In the field of data analysis, such dispersed data can also be referred to as multi-source heterogeneous data. Multi-source heterogeneous data can refer to data that is not unified in terms of storage format, access rights, semantic information, etc. due to different sources. Thus, such data usually poses great challenges in data integration, management, and analysis. In this disclosure, for ease of understanding and discussion, the above-mentioned multi-source heterogeneous data is described in combination with the concept of data domains. A data domain can refer to a division formed for the field to which the data belongs, thereby making each data domain form an obvious data boundary. In other words, data belonging to one data domain may have a higher similarity in certain dimensions compared to other data domains. The data domain can be divided from various perspectives, for example, the source of the data, the way, time, and location of data acquisition, the ownership of data permissions, the semantics of the data, and so on. Based on such a concept, in this disclosure, data from different data domains is referred to as multi-source heterogeneous data.
[0021] The multi-source heterogeneity of data may pose an obstacle to data analysis for on-site faults.
[0022] On the one hand, the multi-source heterogeneity of data may lead to a significant increase in data analysis overhead and latency. To perform data analysis, it is usually necessary to use technologies such as Extract, Transform, Load (ETL) to preprocess multi-source heterogeneous data. ETL technology is suitable for extracting various distributed and heterogeneous source data, cleaning "dirty" data content such as incomplete data, duplicate data, and error data according to pre-designed rules, so as to obtain "clean" data that meets the requirements. This usually requires a large amount of time, effort, and overhead. In particular, the time-consuming ETL process may result in a latency of up to several months, which is often unacceptable for urgent fault problems.
[0023] On the other hand, the multi-source heterogeneity of data may lead to semantic interoperability conflicts. Semantic interoperability conflicts refer to differences in the expression of equivalent concepts. For example, data in different data sources may use different words, terms, or symbols to represent the same concept or entity. Semantic interoperability conflicts not only pose a huge challenge to accurately understanding data semantics during the fault analysis process but are also another important factor leading to latency in data analysis for on-site faults.
[0024] The above problems of multi-source heterogeneous data can be solved through data integration. Data integration aims to integrate multi-source heterogeneous data from different data domains into a unified data model to achieve data sharing and interaction.
[0025] In the present disclosure, a knowledge graph technology is proposed to perform data integration on multi-source heterogeneous raw data from different data domains associated with on-site faults, thereby contributing to efficient fault analysis.
[0026] The ability of knowledge graph technology to achieve knowledge modeling, semantic data integration, and data standardization can be utilized to efficiently organize and manage multi-source heterogeneous data from different data domains. On the one hand, this can reduce the requirements of the fault analysis process for complex and time-consuming ETL processes. On the other hand, and more importantly, it can avoid the semantic interoperability conflict problems that are not well solved in existing technical solutions by promoting a common understanding of data semantics. In this way, the efficiency of fault analysis can be significantly improved, and the latency in the fault analysis process can be reduced.
[0027] Aspects of the present disclosure will be discussed in detail below with reference to the accompanying drawings.
[0028] Figure 1 FIG. 100 shows a schematic diagram of constructing an integrated knowledge graph for multiple data domains according to an embodiment of the present disclosure.
[0029] The raw data associated with the on-site fault of the device may come from multiple data domains. For example, Figure 1 FIG. shows four data domains associated with the on-site fault of the device: on-site data domain 102, production data domain 104, analysis data domain 106, and expert knowledge data domain 108. Corresponding domain knowledge graphs 112, 114, 116, and 118 can be constructed for multiple data domains respectively. Based on the constructed multiple domain knowledge graphs 112, 114, 116, and 118, an integrated knowledge graph 110 can be formed. The knowledge graph 110 is referred to as "integrated" in the present disclosure because it is constructed by integrating homogeneous and heterogeneous data from multiple data domains, rather than only for data from a specific data domain. Such an integrated knowledge graph 110 can maintain a large amount of data associated with the on-site fault of the device from multiple data domains. In other words, the integrated knowledge graph 110 can store homogeneous and heterogeneous data from multiple data domains in a data platform with a unified semantic model. Therefore, when performing fault analysis based on the integrated knowledge graph 110, for example, the above-mentioned RCA analysis, potential factors leading to the on-site fault can be comprehensively integrated and analyzed, so as to efficiently and accurately provide analysis results.
[0030] Corresponding raw data can be obtained from each data domain. For example, for Figure 1For the four data fields shown, field data 122 can be obtained from the field data field 102, production data 124 can be obtained from the production data field 104, analysis data 126 can be obtained from the analysis data field 106, and expert knowledge data 128 can be obtained from the expert knowledge data field 108.
[0031] Field data 122 from the field data field 102 can be preliminary data collected from service stations, repair stations, etc. and related to the fault problems that occur on the device at the site. It is beneficial to take such field data 122 into account in fault analysis because it usually provides important information related to device information, descriptions of fault phenomena, preliminary diagnoses of faults, etc.
[0032] In one example, the field data field 102 can be further divided to include at least one sub - data field ( Figure 1 not shown in the figure). For example, the field data field 102 can include a device data sub - field, a diagnostic data sub - field, a customer complaint data sub - field, and so on.
[0033] Device data from the device data sub - field can provide information related to the device itself. The device data can also provide information related to the product to which the device belongs. For example, for the example where the faulty device is a vehicle steering gear, since service stations, repair stations, etc. usually provide services for vehicles rather than specific components (e.g., steering gears), the device data can include vehicle data. In this example, the device data can include: a vehicle identification number, which can be the unique identification code of the vehicle; vehicle model or type; production platform information; engine model; production date; place of origin; sales date; place of sale; and so on.
[0034] Diagnostic data from the diagnostic data sub - field can provide information related to the preliminary diagnoses of technicians or repair personnel at service stations, repair stations, etc. for on - site faults. For example, for the example of the steering gear, the diagnostic data can be data fields in a fixed format that contain various preliminary diagnostic information such as diagnostic trouble codes, fault types, electronic control unit (ECU) information, time and date of the diagnostic session, etc.
[0035] Customer complaint data from the customer complaint data domain can provide information associated with complaints made by users regarding the device or the product to which the device belongs during the warranty period of the device. In one example, the description of a fault in the customer complaint data can be for a group of devices rather than for the device itself because users usually can only detect the fault phenomenon and do not have the ability to locate the faulty device. Technicians at service stations or repair stations, etc., can initially add indicators in the customer complaint data to indicate the initially diagnosed faulty device. The customer complaint data can also include direct comments from users or technicians on the fault problem, as well as various other information that helps technicians make a preliminary diagnosis. This information can be combined with the above indicators to identify the type of fault. In one example, the customer complaint data for a customer complaint can include the identification (e.g., number) of the customer complaint, the fault description, the repair date, the repair location, the engine serial number (e.g., for the example of the steering gear above), and the customer's description of the details of the fault site, etc. Introducing such customer complaint data in fault analysis is beneficial because such data is a direct and primary source of the fault problems that occur in the device during the warranty period.
[0036] In addition to the above device data, diagnostic data, and customer complaint data, the on-site data 122 can further include other data directly or indirectly associated with the fault site of the device.
[0037] Production data 124 from the production data domain 104 can include data associated with the production process of the device and can be collected from the production line of the device.
[0038] In one example, the production data domain 104 can be further divided to include at least one sub-data domain ( Figure 1 not shown in the figure). For example, the production data domain 104 can include a location data sub-domain, a flow data sub-domain, a component data sub-domain, etc. The location data from the location data sub-domain can include information associated with the production location of the device, such as information about the factory, production line, part number, etc. The flow data from the flow data sub-domain can include manufacturing parameters measured at each station in the production process of the device, such as temperature, pressure, etc. The component data from the component data sub-domain can provide information associated with the assembly relationship of the device, such as information about the upper-level component containing the device, information about the lower-level components that make up the device, etc. Integrating the production data 124 as part of the data to be considered in fault analysis in the knowledge graph 110 has significant benefits. For example, it enables the integrated knowledge graph 110 to trace the production process of the device, thereby having the ability to trace the root cause of potential problems in the production process that may lead to on-site faults.
[0039] The analysis data 126 from the analysis data domain 106 may include information associated with the characteristics of on-site failures. For example, the analysis data 126 may include the criticality level, priority, etc. of on-site failures, which may be determined based on metrics such as the scope of the impact that the failure may cause, the time required to repair the failure, etc., and can be presented in a visual manner. The analysis data 126 can facilitate users or technicians in specifying appropriate countermeasures for on-site failures. For example, deciding whether immediate repair operations need to be performed, etc.
[0040] The expert knowledge data 128 from the expert knowledge data domain 108 may include expert experience information provided by experts associated with equipment failures. For example, experts can, based on experience, judge that a specific failure may be caused by a certain potential cause and can be solved by a certain repair method. Such information can be used as the expert knowledge data 128. In one example, the expert knowledge data 128 can be stored in the form of a fault tree. A fault tree is a technique for reliability analysis and fault diagnosis, which deduces and analyzes the causes leading to the occurrence of a fault event through event nodes and logic gates connecting the event nodes, thereby guiding fault diagnosis and facilitating the proposal of solutions for fault events. In one example, in the integrated knowledge graph 110, the expert knowledge data 128 can be linked to the diagnostic data in the on-site data 122, thus contributing to the execution of fault diagnosis. In addition, given that the fault tree itself is a graph structure, it is more convenient to construct the corresponding domain knowledge graph 118 based on the expert knowledge data 128 in the form of a fault tree. For example, the entities (represented as nodes) in the domain knowledge graph 118 can be determined based on the event nodes in the fault tree, and the relationships (represented as directed edges connecting between the nodes) in the domain knowledge graph 118 can be determined based on the logic gates in the fault tree. This enables the expert knowledge data to be presented and stored in a machine-readable manner.
[0041] It should be noted that although the above Figure 1 describes four data domains associated with on-site failures of equipment and the corresponding raw data that can be obtained from each data domain, the present disclosure is not limited thereto, but may cover other data domains and data types related to on-site failures not discussed in detail above.
[0042] In addition, it has been discussed above that a data domain can be divided into including one or more sub-data domains. It can be understood that a finer-grained data domain division makes the association between the data in the data domain closer. Although Figure 1It is shown to construct a domain knowledge graph for each data domain respectively. However, in one example, a domain knowledge graph can also be constructed for a sub - data domain. For example, instead of constructing a knowledge graph 112 for the on - site data domain 102, three corresponding domain knowledge graphs can be constructed respectively for the device data sub - domain, the diagnostic data sub - domain, and the customer complaint data sub - domain. Constructing a domain knowledge graph based on a finer - grained sub - data domain can make the domain knowledge graph more flexible or portable.
[0043] The following refers to Figure 2 to discuss the flowchart 200 of the method for constructing a corresponding domain knowledge graph for each data domain.
[0044] In step S202, the raw data obtained from the corresponding data domain can be pre - processed to obtain pre - processed data. The pre - processing can include screening out useless or meaningless data, such as incomplete data, duplicate data, and obviously incorrect data, etc. The pre - processing can also include data cleaning techniques known in the art. Data cleaning is a process of re - examining and validating data, aiming to delete duplicate information, correct existing errors, and provide data consistency. The pre - processing can also include data transformation techniques known in the art to transform the data into a required data structure and form.
[0045] In step S204, data extraction can be performed on the pre - processed data. Data extraction aims to extract the pre - processed data in a graph format for filling into the knowledge graph. For example, when the pre - processed data is in a table format, data extraction can be achieved by using the SPARQL query language to convert the pre - processed data in table format into a graph format.
[0046] In step S206, a domain ontology corresponding to the data domain can be developed. A domain ontology can generally be understood as a conceptual or high - level representation and description of a specific domain, and can represent knowledge data that can be understood by technicians in that domain in terms of classes and relationships. The domain ontology can correspond to the data domain. For example, in combination with Figure 1 the example, a field domain ontology corresponding to the on - site data domain 102, a production domain ontology corresponding to the production data domain 104, an analysis domain ontology corresponding to the analysis data domain 106, or an expert knowledge domain ontology corresponding to the expert knowledge data domain 108, etc. can be developed. When developing a domain ontology, mappings can be written through the SPARQL query language to map the attributes of the extracted data to the classes of the developed domain ontology.
[0047] In one example, the domain ontology can be developed in an industry-standardized manner. For example, the domain ontology can be developed based on the Core Information Model for Manufacturing (CIMM). CIMM is a set of high-level standardized ontologies adapted to the entire manufacturing field. Thus, the domain ontology developed in an industry-standardized manner can have better scalability and can be applied to other products without modification.
[0048] In step S208, the domain knowledge graph can be constructed based on the extracted data and the domain ontology. The extracted data can be converted into the Resource Description Framework (RDF) triple format. RDF triples can organize data in the format of {subject - predicate - object}. Then, the RDF triples can be stored in the developed domain ontology to construct the knowledge graph. The process of populating the domain ontology with the extracted data to construct the domain knowledge graph as described above can also be referred to as the materialization of the knowledge graph. It should be noted that the process of constructing the domain knowledge graph described above is only an example, and other methods known in the art can be used to construct the knowledge graph based on data from a specific data domain.
[0049] The above combination Figure 2 describes the process of constructing the domain knowledge graph for a specific data domain. Combining Figure 1 , the domain knowledge graph constructed using the Figure 2 method can include the domain knowledge graph 112 constructed for the field data domain 102, the domain knowledge graph 114 constructed for the production data domain 104, the domain knowledge graph 116 constructed for the analysis data domain 106, or the domain knowledge graph 118 constructed for the expert knowledge data domain 108, and so on. Although in the above description for Figure 2 , the domain knowledge graph is constructed for a specific "data domain", however, such a process is equally applicable to the various "sub-data domains" described in combination with Figure 1 . That is, the domain knowledge graph can be constructed for a sub-data domain. For example, the raw data obtained from the corresponding sub-data domain can be preprocessed, data extraction can be performed on the preprocessed data, the corresponding sub-domain ontology can be developed, and the domain knowledge graph can be constructed based on the extracted data and the sub-domain ontology. In one example, the finally constructed integrated knowledge graph can include the domain knowledge graphs established for both the data domain and the sub-data domain. For example, the integrated knowledge graph can include the domain knowledge graphs constructed respectively for the analysis data domain, the production data domain, the equipment data sub-domain, the diagnostic data sub-domain, and the customer complaint data sub-domain, and these domain knowledge graphs can be constructed based on the analysis ontology, the production ontology, the equipment ontology, the diagnostic ontology, and the customer complaint ontology respectively.
[0050] The following will describe, in conjunction with Figure 3 how to form an integrated knowledge graph based on multiple constructed domain knowledge graphs. For example, in conjunction with Figure 1 the integrated knowledge graph 110 discussed.
[0051] In Figure 3 , four domain knowledge graphs 302, 304, 306, and 308 are exemplarily shown, which are respectively constructed for the four data domains discussed in conjunction with Figure 1 using the domain knowledge graph construction method discussed in conjunction with Figure 2 . For ease of description, the four domain knowledge graphs shown in Figure 3 are respectively referred to as the on-site data knowledge graph 302, the production data knowledge graph 304, the analysis data knowledge graph 306, and the expert data knowledge graph 308. Among them, each domain knowledge graph may include nodes corresponding to entities (represented by circles) and directed edges representing the relationships between entities (represented by connected lines with arrows).
[0052] An integrated knowledge graph can be formed by connecting any two of the multiple domain knowledge graphs. A pair of nodes can be connected together based on the semantic association between a certain node in one domain knowledge graph and the corresponding node in another domain knowledge graph (which can also be called a pair of nodes). The semantic association can indicate whether there is an association between a pair of nodes and what type of association exists. For example, if one node (for simplicity, called the first node) corresponds to the model of a specific device, and another node (for simplicity, called the second node) corresponds to the origin of the device of this specific model, then the semantic association between this pair of nodes can indicate that this specific device comes from this origin. Based on this semantic association, an edge pointing from the first node to the second node can be added between the first node and the second node to achieve the connection of these two nodes. In one example, the semantic association between two nodes can be determined based on expert experience. In another example, the semantic association between two nodes can be determined based on the analysis of the attributes of the entities represented by the nodes.
[0053] Then, by connecting one or more pairs of nodes together, the connection between two domain knowledge graphs can be achieved. As Figure 3As shown in the example, a new edge 312-1 can be added between node 310-1 in the in-situ data knowledge graph 302 and node 310-2 in the production data knowledge graph 304 to connect the in-situ data knowledge graph 302 and the production data knowledge graph 304. Similarly, the production data knowledge graph 304 and the analysis data knowledge graph 306 can be connected by adding an edge 312-2, and the production data knowledge graph 304 and the expert data knowledge graph 308 can be connected by adding an edge 312-3, thus forming the integrated knowledge graph 300.
[0054] The integrated knowledge graph 300 can link homogeneous and heterogeneous data from different data domains. For example, the in-situ data and the production data can be linked, and so on. This enables the integrated knowledge graph 300 to provide good traceability and tracking capabilities for the complete supply chain of the device from production to in-situ operation. Such traceability and tracking capabilities contribute to the early detection or early warning of in-situ device failures. For example, if it is inferred based on the integrated knowledge graph 300 that an error in a specific production process may ultimately lead to a specific failure of the device during in-situ operation, then the specific failure can be detected or warned in advance when the error in the specific production process is discovered.
[0055] Information related to a specific in-situ failure, such as the cause of the failure, the impact, the repair strategy, etc., can be inferred based on the integrated knowledge graph 300. The inference can be performed through a data access point that can call the integrated knowledge graph 300. In one example, the data access point can receive a query statement and return a query result based on the inference made in the integrated knowledge graph 300. For example, the query statement can be implemented using a graph query language (e.g., SPARQL query language). In one example, the data access point can be embodied in the form of an application programming interface (API) so that various external applications can efficiently access the integrated knowledge graph 300 using the API to obtain query or inference results.
[0056] In addition, various downstream applications can be further developed based on the integrated knowledge graph 300 for fault analysis. For example, the APIs discussed above can be utilized to implement the connection between the integrated knowledge graph 300 and the downstream applications. In one aspect, the downstream applications can include data visualization applications, which can implement data visualization based on the integrated knowledge graph 300. In one example, the data visualization application can be used to present a data statistical table to provide an overview associated with on-site faults of the device. Since the integrated knowledge graph 300 integrates homogeneous and heterogeneous data from different data domains and links these data together, such homogeneous and heterogeneous data and the link relationships therebetween can be visually presented in the data statistical table. In one example, based on the on-site data, production data, analysis data, and expert data maintained by the integrated knowledge graph 300 and the link relationships between these data, the data statistical table can display statistical information associated with on-site faults of the device in terms of, for example, the code of the on-site fault, the description of the fault phenomenon, the cause of the fault, the criticality level, etc.
[0057] In another aspect, the downstream applications can also include various AI algorithm-based applications. These AI algorithms can include symbolic AI algorithms, graph machine learning algorithms, graph deep learning algorithms, and other graph AI algorithms. The integrated knowledge graph 300 can be utilized to provide training data for training the AI algorithms. The integrated knowledge graph 300 provides device fault-related data containing standardized and rich semantics, and maintains homogeneous and heterogeneous data from different data domains and the link relationships between these data, thereby helping to provide comprehensive and accurate training data. The trained AI algorithms can be incorporated into the downstream applications to perform various prediction tasks associated with fault analysis of the device. For example, the associations between fault-related data can be predicted, whether an on-site fault will occur can be predicted, and so on.
[0058] Figure 4 The flowchart of a method 400 for fault analysis using a knowledge graph according to an embodiment of the present disclosure is shown.
[0059] In step S402, raw data associated with on-site faults of the device can be obtained, where the raw data can come from multiple data domains. For example, in combination with Figure 1, in this step, field data 122 from the field data domain 102 can be obtained, production data 124 from the production data domain 104 can be obtained, analysis data 126 from the analysis data domain 106 can be obtained, or expert knowledge data 128 from the expert knowledge data domain 108 can be obtained, and so on. More specifically, obtaining such raw data can include calling a database that stores the above data, for example, an Oracle database, a CSV database, or a database in other formats, and so on. In this case, the finally constructed integrated knowledge graph can fuse the raw data from databases in different formats, thus achieving efficient data source management.
[0060] In step S404, an integrated knowledge graph can be constructed for the multiple data domains. For each of the multiple data domains, a domain knowledge graph can be constructed respectively based on the raw data from each data domain. The domain knowledge graph can be constructed according to the process 200 discussed in Figure 2 For example, the raw data from this data domain can be preprocessed first to obtain the preprocessed data. Next, data extraction can be performed on the preprocessed data. Then, a domain ontology corresponding to this data domain can be established, and the domain knowledge graph can be constructed based on the extracted data and the domain ontology. The integrated knowledge graph can be formed based on the multiple domain knowledge graphs constructed respectively for the multiple data domains. Combining the discussion in Figure 3 , any two knowledge graphs in the multiple domain knowledge graphs can be connected together by adding new directed edges between the nodes based on the semantic associations between the nodes of the domain knowledge graphs.
[0061] In step S406, fault analysis can be performed based on this integrated knowledge graph. In one example, the integrated knowledge graph can be accessed by combining the data access points discussed in Figure 3 to obtain query or inference results regarding on-site faults of the device. In one example, this integrated knowledge graph can be provided to various downstream applications, such as the data visualization application and the AI algorithm-based application discussed above in Figure 3 .
[0062] Figure 5 FIG. shows a flowchart of a system 500 for fault analysis using a knowledge graph according to an embodiment of the present disclosure.
[0063] System 500 can include a data collection and preprocessing module 510, a knowledge graph-based data integration module 520, and an application module 530. In one example, the data collection and preprocessing module 510 can perform the operations discussed above in Figure 4The operations of step S402 discussed to obtain the raw data associated with the on-site failures of the device, wherein the raw data comes from multiple data domains. The knowledge graph-based data integration module 520 can perform the operations of step S404 discussed above Figure 4 to construct an integrated knowledge graph for the multiple data domains. The application module 530 can perform the operations of step S406 discussed above Figure 4 to perform failure analysis based on the integrated knowledge graph.
[0064] In one example, the knowledge graph-based data integration module 520 is used to construct an integrated knowledge graph by the following operations: for each of the multiple data domains, construct a domain knowledge graph based on the raw data from each data domain; and form the integrated knowledge graph based on the multiple domain knowledge graphs constructed respectively for the multiple data domains.
[0065] In one example, the knowledge graph-based data integration module 520 is used to construct a domain knowledge graph based on the raw data from each data domain by the following operations: preprocess the raw data from each data domain to obtain preprocessed data; perform data extraction on the preprocessed data; establish a domain ontology corresponding to the data domain; and construct the domain knowledge graph based on the extracted data and the domain ontology.
[0066] In one example, the knowledge graph-based data integration module 520 is used to form the integrated knowledge graph by the following operations: connect the first domain knowledge graph and the second domain knowledge graph among the multiple domain knowledge graphs. The connection of the first domain knowledge graph and the second domain knowledge graph can be performed based on the semantic association between the first node in the first domain knowledge graph and the second node in the second domain knowledge graph. The first domain knowledge graph and the second domain knowledge graph can be connected together by adding a new directed edge between the first node and the second node based on the semantic association.
[0067] In one example, the multiple data domains include at least one of the following: on-site data domain; production data domain; analysis data domain; expert knowledge data domain.
[0068] In one example, the multiple domain knowledge graphs include: on-site data knowledge graph for the on-site data domain; production data knowledge graph for the production data domain; analysis data knowledge graph for the analysis data domain; and expert data knowledge graph for the expert knowledge data domain.
[0069] In one example, one of the multiple data domains is further divided into at least one sub-data domain; and at least one of the multiple domain knowledge graphs is constructed for the at least one sub-data domain.
[0070] In one example, the sub-data domain includes at least one of the following: device data sub-domain; diagnostic data sub-domain; customer complaint data sub-domain.
[0071] In one example, the application module 530 is used to perform fault analysis by the following operations: receiving a query statement using a data access point and returning a query result based on the reasoning performed in the integrated knowledge graph, where the query result includes information associated with fault analysis; and the query statement is based on a graph query language.
[0072] In one example, the application module 530 includes: a data visualization application that provides an overview associated with the on-site faults of the device based on the integrated knowledge graph; and / or an AI algorithm-based application that performs a prediction task associated with the fault analysis of the device, where the AI algorithm is trained using the integrated knowledge graph.
[0073] In one example, the device includes a steering gear of a vehicle.
[0074] Figure 6 A schematic diagram showing the architecture of a computing device according to an embodiment of the present disclosure is shown. The method for using a knowledge graph to perform fault analysis according to the present disclosure can be implemented on the computing device 600.
[0075] The exemplary computing device 600 includes an internal communication bus 602 and a processor (e.g., a central processing unit (CPU)) 604 connected to the internal communication bus 602. The processor 604 is used to execute instructions stored in the memory 606 to implement the method for using a knowledge graph to perform fault analysis described in detail above. The memory 606 is adapted to tangibly embody computer program instructions and data, and may include various forms of memory, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM disks, and so on. The computing device 600 may also include an input / output (I / O) interface 608, so that various I / O devices (e.g., a cursor control device such as a mouse, a keyboard, etc.) can be coupled to the computing device 600 through the I / O interface 608 to allow a user to apply various commands and input data. The computing device 600 may also include a display unit 610 for displaying a graphical user interface.
[0076] A computer program may include instructions executable by a computer for causing a processor 604 of a computing device 600 to execute the method of the present disclosure for fault analysis using a knowledge graph. The program may be recorded on any data storage medium including a memory. For example, the program may be implemented in digital electronic circuits, or in computer hardware, firmware, software, or in combinations thereof. The process / method steps described in the present disclosure may be performed by a programmable processor executing program instructions to perform the methods, steps, operations by operating on input data and generating output.
[0077] In addition to what has been described herein, various modifications may be made to the disclosed embodiments and implementations without departing from the scope of the disclosed embodiments and implementations. Accordingly, the description and examples herein should be construed as illustrative and not in a limiting sense. The scope of the present disclosure should be measured only by reference to the claims.
Claims
1. A method for fault analysis using a knowledge graph, comprising: obtaining original data associated with on-site faults of a device, wherein the original data comes from multiple data domains; constructing an integrated knowledge graph for the multiple data domains; and performing fault analysis based on the integrated knowledge graph.
2. The method according to claim 1, wherein, constructing an integrated knowledge graph for the multiple data domains includes: for each of the multiple data domains, respectively constructing a domain knowledge graph based on the original data from each data domain; and forming the integrated knowledge graph based on the multiple domain knowledge graphs respectively constructed for the multiple data domains.
3. The method according to claim 2, wherein, respectively constructing a domain knowledge graph based on the original data from each data domain includes: preprocessing the original data from each data domain to obtain preprocessed data; performing data extraction on the preprocessed data; establishing a domain ontology corresponding to the data domain; and constructing the domain knowledge graph based on the extracted data and the domain ontology.
4. The method according to claim 2, wherein, forming the integrated knowledge graph based on the multiple domain knowledge graphs respectively constructed for the multiple data domains includes: connecting a first domain knowledge graph and a second domain knowledge graph among the multiple domain knowledge graphs together.
5. The method according to claim 4, wherein, performing the connection of the first domain knowledge graph and the second domain knowledge graph based on the semantic association between a first node in the first domain knowledge graph and a second node in the second domain knowledge graph.
6. The method according to claim 5, wherein, connecting the first domain knowledge graph and the second domain knowledge graph together by adding a new directed edge between the first node and the second node based on the semantic association.
7. The method according to claim 2, wherein, the multiple data domains include at least one of the following: on-site data domain; production data domain; analysis data domain; expert knowledge data domain.
8. The method according to claim 7, wherein, the multiple domain knowledge graphs include: an on-site data knowledge graph for the on-site data domain; a production data knowledge graph for the production data domain; an analysis data knowledge graph for the analysis data domain; and an expert data knowledge graph for the expert knowledge data domain.
9. The method according to claim 2, wherein: one of the multiple data domains is further divided into at least one sub-data domain; and at least one of the multiple domain knowledge graphs is constructed for the at least one sub-data domain.
10. The method according to claim 9, wherein, the sub-data domain includes at least one of the following: device data sub-domain; diagnostic data sub-domain; customer complaint data sub-domain.
11. The method according to claim 1, wherein, performing fault analysis based on the integrated knowledge graph includes: Receiving a query statement using a data access point and returning a query result based on reasoning performed in the integrated knowledge graph, where the query result includes information associated with fault analysis; wherein The query statement is based on a graph query language.
12. The method according to claim 1, wherein Performing fault analysis based on the integrated knowledge graph includes providing the integrated knowledge graph to a downstream application, where the downstream application includes: A data visualization application that provides an overview associated with on-site faults of the device based on the integrated knowledge graph; and / or, An AI algorithm-based application that performs a prediction task associated with fault analysis of the device, where the AI algorithm is trained using the integrated knowledge graph.
13. The method according to claim 1, wherein The device includes a steering gear of a vehicle.
14. A system for performing fault analysis using a knowledge graph, comprising: A data collection and preprocessing module for obtaining raw data associated with on-site faults of a device, where the raw data comes from multiple data domains; A knowledge graph-based data integration module for constructing an integrated knowledge graph for the multiple data domains; and An application module for performing fault analysis based on the integrated knowledge graph.
15. The system according to claim 14, wherein The knowledge graph-based data integration module is used to construct an integrated knowledge graph by: For each of the multiple data domains, constructing a domain knowledge graph based on the raw data from each data domain; and Forming the integrated knowledge graph based on the multiple domain knowledge graphs constructed for the multiple data domains.
16. The system according to claim 15, wherein The multiple domain knowledge graphs include: An on-site data knowledge graph for the on-site data domain; A production data knowledge graph for the production data domain; An analysis data knowledge graph for the analysis data domain; and An expert data knowledge graph for the expert knowledge data domain.
17. The system according to claim 14, wherein The application module is used to perform fault analysis by: Receiving a query statement using a data access point and returning a query result based on reasoning performed in the integrated knowledge graph, where the query result includes information associated with fault analysis; wherein The query statement is based on a graph query language.
18. The system according to claim 14, wherein The application module includes: A data visualization application that provides an overview associated with on-site faults of the device based on the integrated knowledge graph; and / or, An AI algorithm-based application that performs a prediction task associated with fault analysis of the device, where the AI algorithm is trained using the integrated knowledge graph.
19. A device for performing fault analysis using a knowledge graph, comprising: A memory; A processor, coupled to the memory, the processor being configured to execute the method according to any one of claims 1-13.
20. A computer-readable medium storing a computer program comprising instructions that, when executed by a processor, cause the processor to be configured to execute the method according to any one of claims 1-13.