Equipment operation and maintenance knowledge graph construction method and system based on AI
By analyzing the semantic coherence and noise characteristics of equipment operation and maintenance data, identifying and eliminating noise triplets, the problem of noise influence in the equipment operation and maintenance knowledge graph is solved, and the quality of the knowledge graph and operation and maintenance efficiency are improved.
Patent Information
- Application Number
- CN202511279075.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-09
AI Technical Summary
When constructing equipment operation and maintenance knowledge graphs, existing technologies are unable to accurately mine abnormal features of semantic coherence due to the presence of noisy triples caused by fuzzy semantics in historical operation and maintenance data, resulting in poor accuracy of knowledge extraction and affecting the efficiency and quality of equipment operation and maintenance.
By obtaining the historical operation and maintenance data of the operation and maintenance objects, extracting the text vectors and word vectors of the triple data, analyzing the semantic coherence and local outlier features, and combining the noise semantic feature values to identify and eliminate noise triplets, a knowledge graph for equipment operation and maintenance is constructed.
The construction quality of the equipment operation and maintenance knowledge graph is improved, important fault operation and maintenance information is accurately retained, error filtering is avoided, and the efficiency and quality of equipment operation and maintenance are improved.
Smart Images

Figure CN120806100A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of knowledge graph construction, in particular to an AI-based device operation and maintenance knowledge graph construction method and system. BACKGROUND
[0002] At present, the construction of a device operation and maintenance knowledge graph is an important method for improving operation and maintenance efficiency and quality. A device operation and maintenance knowledge graph constructed by using AI artificial intelligence technology can systematically display device operation and maintenance related knowledge and information, including fault entities, entity relationships and operation and maintenance events of various operation and maintenance objects, and the use of a device operation and maintenance knowledge graph helps a team to more quickly locate the root cause of a problem and quickly find a corresponding solution, thereby more efficiently achieving operation and maintenance management of various operation and maintenance objects. Therefore, as an important tool in the field of device operation and maintenance, an operation and maintenance knowledge graph has important use value for improving the efficiency and quality of device operation and maintenance and reducing operation and maintenance costs.
[0003] In the prior art, historical operation and maintenance data of various operation and maintenance objects are collected, and a DeepKE open source knowledge extraction tool is used to extract knowledge from the historical operation and maintenance data, and then an AI artificial intelligence technology is used to construct a device operation and maintenance knowledge graph based on the triple information extracted by the knowledge extraction, for improving the efficiency and quality of device operation and maintenance. However, due to the existence of noise triples caused by fuzzy semantics in the historical operation and maintenance data, the semantic coherence between different triple data may be abnormal, and the prior art uses the DeepKE open source knowledge extraction tool to extract knowledge from the historical operation and maintenance data, which cannot accurately mine the abnormal features of the semantic coherence between different triple data, resulting in poor accuracy of noise filtering on the triple information extracted by the knowledge extraction, thereby reducing the quality of the construction of the device operation and maintenance knowledge graph and affecting the efficiency and quality of the device operation and maintenance. SUMMARY
[0004] In order to solve the above technical problems, the purpose of the present application is to provide an AI-based device operation and maintenance knowledge graph construction method and system, and the technical solutions adopted are as follows: The present application provides an AI-based device operation and maintenance knowledge graph construction method, comprising the following steps: Obtain historical operation and maintenance data of each operation and maintenance object, and perform knowledge extraction to obtain each triple data, and then extract a text vector and a text word vector of each triple data; Analyze the text word vector similarity and semantic coherence between different triple data to obtain the operation and maintenance feature coherence of each triple data, and combine the local outlier features of the operation and maintenance feature coherence of each triple data to obtain the semantic disjoint confidence of each triple data; According to the semantic similarity degree between different word segmentation in the text vector of each triple data and the random change of the semantic similarity degree, the fuzzy interference degree of each triple data is obtained, and then the noise semantic feature value of each triple data is obtained by combining the semantic disjoint confidence; Based on the noise semantic feature value, the noise triple is identified and removed, and then the device operation and maintenance knowledge graph is constructed by using the retained triple data in the historical operation and maintenance data.
[0005] Preferably, during knowledge extraction, the triple data contains the subject, object and the relationship between the subject and object.
[0006] Preferably, the similarity between the text word vector of each triple data and the text word vector of each remaining triple data is calculated, and the similarity corresponding to each triple data and all the remaining triple data is clustered, and the cluster with the maximum mean value of all elements in the cluster is taken as the highly associated cluster of each triple data.
[0007] Preferably, the acquisition of the operation and maintenance feature coherence of each triple data is further: ; In the formula, is the operation and maintenance feature coherence of the i-th triple data, is a normalization function, is the number of elements in the highly associated cluster of the i-th triple data, is the sum of the elements in the highly associated cluster of the i-th triple data, is the dispersion degree of the elements in the highly associated cluster of the i-th triple data, is a constant to avoid the denominator being 0.
[0008] Preferably, the reciprocal of the sum of the operation and maintenance feature coherence of each triple data and the normalized result of the local density of the operation and maintenance feature coherence is taken as the semantic disjoint confidence of each triple data, wherein the local density of the operation and maintenance feature coherence of each triple data is obtained by the density peak clustering algorithm.
[0009] Preferably, the text vector of each triple data is subjected to word segmentation processing, and all the word segmentation forms a word segmentation data set of each triple data.
[0010] Preferably, the acquisition of the fuzzy interference degree of each triple data is further: ; In the formula, is the fuzzy interference degree of the i-th triple data, is an exponential function with a natural constant as the base number, a mean value of normalized Google distances between all segmented words in the segmented data set of the i-th triple data, an information entropy of normalized Google distances between all segmented words in the segmented data set of the i-th triple data.
[0011] Preferably, the noise semantic feature value of each triple data is a sum of the semantic disjointedness confidence and the fuzzy interference degree of each triple data.
[0012] The application further provides an AI-based device operation and maintenance knowledge graph construction system, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, and the processor implements the steps of the AI-based device operation and maintenance knowledge graph construction method described in any of the above.
[0013] As can be seen from the above, the AI-based device operation and maintenance knowledge graph construction method and system provided by the application has at least the following beneficial effects: The application extracts highly correlated clusters of triple data through semantic similarity analysis, accurately measures the coherence of operation and maintenance features between triple data through data features in the highly correlated clusters, more clearly shows the correlation between device operation and maintenance information, retains triple data information with good operation and maintenance feature coherence, thereby avoiding mistakenly filtering out important fault operation and maintenance information; Further, the application accurately reflects the abnormal degree of semantic coherence of the triple data in all triple data by extracting the local density feature of the operation and maintenance feature coherence of the triple data, and then constructs a semantic disjointedness confidence in combination with the operation and maintenance feature coherence and the local density feature of the triple data, thereby clearly showing the problem of semantic disjointedness or semantic incoherence of the semantic relationship between the triple data and other triple data, for more accurately filtering out noise triple data caused by fuzzy semantics in the future; The application accurately measures the noise semantic features of each triple data through the reliability of the semantic disjointedness of each triple data and in combination with the interference degree of each triple data caused by fuzzy semantics, and then accurately identifies and filters out noise triple data by using an anomaly detection algorithm, thereby improving the quality of device operation and maintenance knowledge graph construction and avoiding affecting the efficiency and quality of subsequent device operation and maintenance. BRIEF DESCRIPTION OF DRAWINGS
[0014] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the application or the prior art, the drawings needed in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0015] Figure 1 A step flow chart of an AI-based device operation and maintenance knowledge graph construction method provided by the present application is provided. Figure 2 A device operation and maintenance knowledge graph construction system block diagram provided by the embodiments of the present application is provided. DETAILED DESCRIPTION
[0016] In order to further illustrate the technical means and effects adopted by the present application to achieve the predetermined invention purpose, the specific implementation, structure, features and effects of the AI-based device operation and maintenance knowledge graph construction method and system according to the present application are described in detail as follows. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0017] Unless otherwise defined and limited, such as the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the circuit structure, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such article or device. Without more limitation, the element limited by the statement "including one" does not exclude the presence of another identical element in the article or device including the element. In addition, the term "and / or" used herein includes any and all combinations of one or more related listed items. All technical and scientific terms used herein have the same meaning as understood by those skilled in the art of the technology to which the present application belongs.
[0018] The specific scheme of the AI-based device operation and maintenance knowledge graph construction method and system provided by the present application is described in detail below with reference to the accompanying drawings.
[0019] Please refer to Figure 1 which shows a step flow chart of an AI-based device operation and maintenance knowledge graph construction method provided by one embodiment of the present application, including the following steps: Step 1: Obtain historical operation and maintenance data of each operation and maintenance object, and perform knowledge extraction to obtain each triple data, and then extract text vector and text word vector of each triple data.
[0020] In order to more efficiently realize device operation and maintenance management, it is necessary to accurately mine the abnormal features of semantic coherence between different triple data in the knowledge extraction result, so as to accurately filter out the noise triple data caused by ambiguous semantics, thereby improving the quality of device operation and maintenance knowledge graph construction.
[0021] Therefore, in the embodiment, first, historical operation and maintenance data of various operation and maintenance objects are collected, the operation and maintenance objects include CPU servers, GPU servers, storage hard disks, and network devices, and the historical operation and maintenance data are text data of historical fault cases and solutions in a process, wherein common operation and maintenance faults include CPU server overheating, GPU core fault (power reduction or GPU card drop), storage hard disk head damage, and network device port fault.
[0022] The historical operation and maintenance data are subjected to knowledge extraction by installing and configuring a DeepKE open-source knowledge extraction tool on a computer, wherein the knowledge extraction task includes named entity recognition, relationship extraction, attribute extraction, and extraction of triple data, all triple data in the historical operation and maintenance data are obtained, the triple data include a subject, an object, and a relationship between the subject and the object, and the knowledge extraction by DeepKE is a known technology and will not be described in detail.
[0023] Further, the triple data extracted by the AI artificial intelligence technology are preprocessed, specifically, text vectors of each triple data in the historical operation and maintenance data are converted to obtain text vectors of each triple data, and the text vectors of each triple data are input into a Word2Vec word vector algorithm to obtain text word vectors of each triple data by using the Word2Vec word vector algorithm, wherein the text vector conversion and the Word2Vec word vector algorithm are known technologies in the AI artificial intelligence field and will not be described in detail.
[0024] Step 2: analyze the text word vector similarity and semantic coherence degree between different triple data to obtain the operation and maintenance feature coherence degree of each triple data, and combine the local outlier features of the operation and maintenance feature coherence of each triple data to obtain the semantic disjoint confidence of each triple data.
[0025] Since there are noise triple data due to fuzzy semantics in the historical operation and maintenance data, the prior art uses the DeepKE open-source knowledge extraction tool to extract knowledge from the historical operation and maintenance data, which cannot accurately mine the abnormal features of the semantic coherence between different triple data, resulting in that the noise triple data due to fuzzy semantics cannot be accurately filtered out. Therefore, in order to improve the quality of the device operation and maintenance knowledge graph, it is necessary to fully mine the abnormal features of the semantic coherence between different triple data.
[0026] In order to analyze the correlation characteristics between different triple data in the historical operation and maintenance data, taking the i-th triple data as an example, the similarity between the text word vector of the i-th triple data and the text word vector of each of the remaining triple data is calculated. The similarity can be measured by cosine similarity or Jaccard similarity coefficient. In this embodiment, the cosine similarity is used to measure the similarity. The greater the similarity, the higher the similarity between the text word vectors. The measurement of the cosine similarity is a known technology, and the specific process will not be described here.
[0027] Further, in order to further analyze the semantic coherence characteristics between each triple data and other triple data, the similarity corresponding to the i-th triple data and all the remaining triple data is taken as the input of the K-means clustering algorithm. The elbow rule is used to obtain the optimal number of clustering clusters K, and the K-means clustering algorithm is used to obtain K clustering clusters composed of the similarity. The mean value of all elements in each clustering cluster is calculated, and the clustering cluster corresponding to the maximum mean value is taken as the highly associated cluster of the i-th triple data. The K-means clustering algorithm and the elbow rule are both known technologies, and the specific process will not be described here.
[0028] Generally, the more the number of elements in the highly associated cluster of a certain triple data, the more information related to the triple data in the historical operation and maintenance data, and the less likely the triple data belongs to the noise triple data generated by the fuzzy semantics.
[0029] At the same time, the greater the sum of the elements in the highly associated cluster and the smaller the dispersion degree, the better the operation and maintenance feature coherence between part of the triple data and the triple data in the historical operation and maintenance data, and the more important the triple data information for the construction of the operation and maintenance knowledge graph. For example, the triple data has a device operation fault entity, such as a damaged mechanical hard disk head or a packet loss of a switch port, and there are more triple data information closely associated with the device operation fault entity in the corresponding solution.
[0030] Based on the above analysis, the operation and maintenance feature coherence of each triple data is calculated. ; In the formula, is the operation and maintenance feature coherence of the i-th triple data, is a normalization function, is the number of elements in the highly associated cluster of the i-th triple data, is the sum of the elements in the highly associated cluster of the i-th triple data, is the dispersion degree of the elements in the highly associated cluster of the i-th triple data, In order to avoid the constant with the denominator of 0, the value is taken in the smaller data range (0.01, 0.1), the influence on the calculation result can be ignored, and the adjustment parameter is taken as 0.05 in the embodiment.
[0031] The measurement method of the dispersion degree can be variance, standard deviation or dispersion coefficient, and the dispersion coefficient is used to measure the dispersion degree in the embodiment.
[0032] According to the above calculation process, it can be understood that the operation and maintenance feature continuity reflects the continuity of the operation and maintenance feature between the i-th triple data and other triple data. The greater the operation and maintenance feature continuity, the better the continuity of the operation and maintenance feature between the triple data and other triple data in the historical operation and maintenance data. Then, the triple data is less likely to be a noise triple data generated by fuzzy semantics, and should be retained when constructing the device operation and maintenance knowledge graph.
[0033] Further, in order to accurately extract the abnormal features of the semantic continuity between different triple data, the operation and maintenance feature continuity of all triple data in the historical operation and maintenance data is taken as the input of the DPC density peak clustering algorithm (density peaks clustering, DPC). The selection rule of the cut-off distance is that the number of data points in the cut-off distance range of each data point accounts for 1-2% of the total number of data points. In the embodiment, the cut-off distance when the number of data points in the cut-off distance range of each data point accounts for 2% of the total number of data points is taken as the preset cut-off distance in the algorithm. The local density of the operation and maintenance feature continuity of each triple data is obtained by the DPC density peak clustering algorithm. The smaller the local density of the operation and maintenance feature continuity, the more significant the outlying features of the corresponding triple data in the historical operation and maintenance data, that is, the higher the abnormal degree of the semantic continuity of the triple data among all triple data, which means that the triple data is more likely to be a noise triple data generated by fuzzy semantics, so the triple data should be filtered out when constructing the device operation and maintenance knowledge graph, thereby avoiding affecting the quality of the device operation and maintenance knowledge graph construction.
[0034] Therefore, in order to more accurately filter out the noise triple data generated by fuzzy semantics, the local density of the operation and maintenance feature continuity of all triple data in the historical operation and maintenance data is maximum normalized, and the reciprocal of the sum of the operation and maintenance feature continuity and the normalized result of the local density of the operation and maintenance feature continuity of each triple data is taken as the semantic disconnection confidence of each triple data.
[0035] It can be understood that the semantic disconnection confidence reflects the credibility of the loss of semantic coherence between the triple data and other triple data in the historical operation and maintenance data. The higher the credibility of the loss of semantic coherence, the more likely the triple data is a noise triple generated by fuzzy semantics, which will affect the quality of the construction of the device operation and maintenance knowledge graph, so it is necessary to filter out the noise triple generated by fuzzy semantics.
[0036] Step 3: According to the semantic similarity between different word segmentation in the text vector of each triple data and the random change of the semantic similarity, the fuzzy interference degree of each triple data is obtained, and then the noise semantic feature value of each triple data is obtained by combining the semantic disconnection confidence.
[0037] Further, in order to extract the features of fuzzy semantics in the text vector of the noise triple, the text vector of each triple data is taken as the input of the jieba word segmentation technology, and the text vector of each triple data is processed by the jieba word segmentation technology. The set composed of all word segmentation after word segmentation is taken as the word segmentation data set of each triple data, wherein the jieba word segmentation technology is a known technology, and the specific process will not be repeated.
[0038] Generally, the semantic similarity between different word segmentation in the word segmentation data set of normal triple data is relatively high, while in the noise triple generated by noise or fuzzy semantics, the semantic similarity between different word segmentation in the word segmentation data set of the noise triple is relatively low due to the interference of fuzzy semantics, and the semantic similarity between different word segmentation is relatively chaotic. When constructing the device operation and maintenance knowledge graph, the noise triple data should be filtered out as much as possible to avoid affecting the quality of the construction of the device operation and maintenance knowledge graph. Therefore, in the embodiment, the fuzzy interference degree of each triple data is calculated according to the average level of the Google distance between different word segmentation of the triple data and the random change of the Google distance: In the formula, is the fuzzy interference degree of the i-th triple data, is an exponential function with a natural constant as the base, is the mean value of the normalized Google distance between all word segmentation in the word segmentation data set of the i-th triple data, is the information entropy of the normalized Google distance between all word segmentation in the word segmentation data set of the i-th triple data.
[0039] According to the above calculation process, it can be understood that the fuzzy interference degree reflects the interference degree of each triple data by fuzzy semantics. The greater the fuzzy interference degree, the more likely the subject, object and the relationship between the subject and the object in the corresponding triple data are interfered by fuzzy semantics, and the more likely the corresponding triple data is a noise triple generated by fuzzy semantics.
[0040] Further, by combining the semantic disconnection feature and the fuzzy interference feature of the triple data, the noise semantic feature of each triple data is measured, so that subsequent noise triples generated by fuzzy semantics can be accurately filtered out. Specifically, the sum of the semantic disconnection confidence and the fuzzy interference degree of each triple data in the historical operation and maintenance data is taken as the noise semantic feature value of each triple data. The noise semantic feature value reflects the degree of influence of each triple data by noise. The higher the abnormal degree of the noise semantic feature value of a certain triple data in all noise semantic feature values of triple data, the more serious the influence of the triple data by noise, and the more the abnormal triple data should be filtered out, so as to avoid affecting the quality of the device operation and maintenance knowledge graph construction.
[0041] Step 4: Based on the noise semantic feature value, noise triples are identified and removed, and then the device operation and maintenance knowledge graph is constructed by using the retained triple data in the historical operation and maintenance data.
[0042] Further, in order to accurately filter out noise triples generated by fuzzy semantics, all noise semantic feature values of triple data in the historical operation and maintenance data are taken as the input of the LOF anomaly detection algorithm (Local Outlier Factor). The preset neighborhood parameter in the algorithm is 10. The abnormal data in all noise semantic feature values are detected by the LOF anomaly detection algorithm. The LOF anomaly detection algorithm is a known technology, and the specific process is not described again.
[0043] In this embodiment, for all triple data, the noise semantic feature values of all triple data can be detected by the above process, and the abnormal data in the noise semantic feature values are extracted. The triple data corresponding to the abnormal data in all noise semantic feature values are all recorded as noise triples. It should be noted that the noise triple is the triple data corresponding to the abnormal data in the anomaly detection result of the noise semantic feature value.
[0044] The noise triplets are filtered out by removing the noise triplets in all triplet data, thereby removing the noise triplets generated by ambiguous semantics, to obtain the triplet data in the filtered historical operation and maintenance data, and the triplet data in the filtered historical operation and maintenance data is fused and a knowledge graph is constructed by using the protege open source software, to obtain the device operation and maintenance knowledge graph. The knowledge fusion and knowledge graph construction using the protege open source software are known technologies, and the specific process will not be described again.
[0045] Further, in the embodiment, the device operation and maintenance knowledge graph is stored in the Neo4j graph database in the form of graph data, wherein the Neo4j graph database is used as a graph database engine to store knowledge in the form of graph data, the fault entity of the device operation and maintenance is represented as a node, and the relationship exists in the form of an edge. The storage mode of such a graph database provides efficient device operation and maintenance fault query and fault correlation analysis capabilities, enabling the device operation and maintenance personnel to easily perform complex device operation and maintenance operations, thereby better understanding the device operation and maintenance faults and the solution processing method of the faults.
[0046] Based on the same inventive concept as the above method, the embodiments of the present application also provide an AI-based device operation and maintenance knowledge graph construction system, which includes a memory, a processor, and a computer program stored in the memory and running on the processor, and the processor implements the steps of the AI-based device operation and maintenance knowledge graph construction method according to any one of the above embodiments when executing the computer program.
[0047] Preferably, in the embodiment, the device operation and maintenance knowledge graph construction system includes a data processing module, a coherence analysis module, a noise analysis module, and a knowledge graph construction module. The data processing module is used to obtain historical operation and maintenance data of each operation and maintenance object, perform knowledge extraction to obtain each triplet data, and then extract the text vector and text word vector of each triplet data. The coherence analysis module is used to analyze the text word vector similarity and semantic coherence degree between different triplet data, to obtain the operation and maintenance feature coherence degree of each triplet data, and obtain the semantic disjoint confidence of each triplet data in combination with the local outlier feature of the operation and maintenance feature coherence of each triplet data. The noise analysis module is used to obtain the fuzzy interference degree of each triplet data according to the semantic similarity degree between different word segmentation in the text vector of each triplet data and the random change of the semantic similarity degree, and obtain the noise semantic feature value of each triplet data in combination with the semantic disjoint confidence. The knowledge graph construction module is used to identify and remove noise triplets based on the noise semantic feature value, and then construct the device operation and maintenance knowledge graph using the retained triplet data in the historical operation and maintenance data. Specifically, the device operation and maintenance knowledge graph construction system block diagram provided by the embodiment is shown in Figure 2
[0048] It can be understood that the above-mentioned embodiments of the application are only for description, and do not represent the advantages and disadvantages of the embodiments. And the above describes specific embodiments of the specification. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or can be advantageous.
[0049] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments.
[0050] The above is only an embodiment of the application, and is not used to limit the scope of the application. Any equivalent structure or equivalent process transformation using the content of the specification and the drawings, or direct or indirect application in other related technical fields, is also included in the protection scope of the application.
Claims
1. A method for constructing an AI-based equipment operation and maintenance knowledge graph, characterized in that: The following steps are involved: Obtain historical operation and maintenance data for each operation and maintenance object, perform knowledge extraction to obtain triple data, and then extract text vectors and text word vectors for each triple data; Analyze the similarity of text word vectors and the degree of semantic coherence between different triple data to obtain the operational feature coherence of each triple data. Combined with the local outlier features of the operational feature coherence of each triple data, the semantic disconnection confidence of each triple data is obtained. According to the semantic similarity between different word segments in the text vector of each triple data and the random change of the semantic similarity, the fuzzy interference degree of each triple data is obtained, and then the noise semantic feature value of each triple data is obtained by combining the semantic disconnection confidence. Noise triplets are identified and eliminated based on noise semantic feature values, and then the equipment operation and maintenance knowledge graph is constructed using the triple data retained in the historical operation and maintenance data.
2. The method for constructing an AI-based equipment operation and maintenance knowledge graph according to claim 1, wherein: When performing knowledge extraction, triple data contains the subject, the object, and the relationship between the subject and the object.
3. The method for constructing an AI-based equipment operation and maintenance knowledge graph according to claim 1, wherein: The similarity between the text word vector of each triple data and the text word vector of each other triple data is counted, and the similarity corresponding to each triple data and all other triple data is clustered. The cluster corresponding to the maximum mean of all elements in the cluster is taken as the highly correlated cluster of each triple data.
4. The method for constructing an AI-based equipment operation and maintenance knowledge graph according to claim 3, wherein: The acquisition of the operational feature coherence of each triple data is further as follows: Where, is the operational feature coherence of the i-th triple data, is the normalization function, is the number of elements in the highly correlated cluster of the i-th triple data, is the sum of the elements in the highly correlated cluster of the i-th triple data, is the degree of dispersion of the elements in the highly correlated cluster of the i-th triple data, To avoid constants with denominators equal to 0.
5. The method for constructing an AI-based equipment operation and maintenance knowledge graph according to claim 1, wherein: The inverse of the sum of the normalized results of the operation and maintenance feature coherence of each triple data and the local density of the operation and maintenance feature coherence is taken as the semantic disconnection confidence of each triple data, wherein the local density of the operation and maintenance feature coherence of each triple data is obtained by the density peak clustering algorithm.
6. The method for constructing an AI-based equipment operation and maintenance knowledge graph according to claim 1, wherein: The text vector of each triple data is segmented, and all the segmented words are combined into a segmentation dataset for each triple data.
7. The method for constructing an AI-based equipment operation and maintenance knowledge graph according to claim 6, wherein: The fuzzy interference degree of each triplet data is further obtained as follows: Where, is the fuzzy interference degree of the i-th triplet data, is an exponential function with a natural constant as base, is the mean of the normalized Google distances between all the segmentations in the segmentation dataset of the i-th triple data, It is the information entropy of the normalized Google distance between all the segmentations in the segmentation dataset of the i-th triple data.
8. The method for constructing an AI-based equipment operation and maintenance knowledge graph according to claim 1, wherein: The noise semantic feature value of each triple data is the sum of the semantic disjointness confidence and the fuzzy interference degree of each triple data.
9. The method for constructing an AI-based equipment operation and maintenance knowledge graph according to claim 1, wherein: Anomaly detection is performed on the noise semantic feature values of all triple data, and the triple data corresponding to the abnormal data is regarded as the noise triple.
10. An AI-based equipment operation and maintenance knowledge graph construction system, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method for constructing an AI-based equipment operation and maintenance knowledge graph as described in any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Method and system for constructing knowledge graph by extracting triples from large model, and medium
CN117273132A
Operation and maintenance knowledge graph construction method, track operation and maintenance method and related device
CN119227798A
Knowledge exchange method and system based on knowledge graph
CN120278252A
Knowledge extraction algorithm suitable for knowledge graph in intelligent operation and maintenance field
CN120542534A
Method and system for pattern discovery and real-time anomaly detection based on knowledge graph
US20200073932A1