Aero-engine quality tracing method based on knowledge graph

By acquiring and cleaning aircraft engine production data, using the combination of active learning algorithms and RB-CasRel model to extract and fusion knowledge, a knowledge map of aircraft engine quality traceability was constructed, solving the problems of difficulty in building ontology and diverse data types, and achieving efficient quality traceability.

CN120163489APending Publication Date: 2025-06-17BEIJING INFORMATION SCI & TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510229629.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The aircraft engine mass traceability method based on knowledge graph faces the problems of difficulty in building the ontology, diverse data types and high confidentiality, which leads to the increased difficulty in building the knowledge graph.

Method used

By obtaining structured and unstructured data for the entire process of aircraft engine production, data cleaning and knowledge extraction are carried out separately, the unstructured data is extracted by combining active learning algorithms and RB-CasRel model, and knowledge fusion is carried out through cosine similarity and editing distance algorithms to build an aircraft engine mass traceability knowledge map.

Benefits of technology

It realizes efficient integration and management of complex aero engine production data, reduces the problems of data diversity and complexity, and improves the accuracy of knowledge extraction and the efficiency and accuracy of quality traceability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163489A_ABST
    Figure CN120163489A_ABST
Patent Text Reader

Abstract

The invention discloses an aero-engine quality tracing method based on a knowledge graph, and relates to the technical field of artificial intelligence, and the method comprises the steps: constructing an aero-engine quality tracing knowledge graph through introducing a knowledge graph technology, and achieving the quality tracing of a whole life cycle from design to after-sales service; the method specifically comprises the steps that structured data and unstructured data are acquired based on aero-engine production whole-process data records, and an ontology framework is generated; respectively performing knowledge extraction on the unstructured data and the structured data; performing knowledge fusion on vocabularies output by knowledge extraction of the structured data and the unstructured data to obtain a triple; based on the triple, constructing an aero-engine quality tracing knowledge graph; through the method, the accuracy and generalization ability of knowledge extraction are improved, rapid and accurate positioning and traceability of quality problems are realized, and the performance of the aero-engine is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and specifically to a method for quality traceability of aero-engines based on a knowledge graph. Background Technique

[0002] With the acceleration of the industrial digital and intelligent transformation, the knowledge graph has shown great potential in the field of intelligent decision-making. By constructing a structured knowledge network, the knowledge graph can effectively integrate and manage complex data resources, and achieve in-depth data mining and application. At present, the knowledge graph has been widely used in many fields. Especially when introducing the knowledge graph into the field of aero-engine production and manufacturing, it can quickly and accurately locate and trace the quality problems generated during the whole life cycle of aero-engine production.

[0003] However, the method for quality traceability of aero-engines based on the knowledge graph faces the following challenges: the ontology construction for aero-engine quality traceability is difficult and requires high professional knowledge and in-depth participation of experts in related fields; the production process of aero-engines involves a large amount of structured and unstructured data, and it is necessary to design knowledge extraction methods separately to construct the knowledge graph, making it difficult to achieve effective integration of the graph; in addition, due to the confidential nature of aero-engines, the available data resources are relatively limited, which further increases the difficulty of constructing the knowledge graph. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for quality traceability of aero-engines based on a knowledge graph to solve the problems raised in the prior art.

[0005] To achieve the above purpose, the present invention provides the following technical solution. A method for quality traceability of aero-engines based on a knowledge graph, the aero-engine quality traceability method includes:

[0006] Step S100: Based on the full-process data records of aero-engine production, obtain structured data and unstructured data, and generate an ontology framework;

[0007] Step S200: Perform knowledge extraction on unstructured data and structured data respectively;

[0008] Step S300: Perform knowledge fusion on the vocabulary output from the knowledge extraction of structured data and unstructured data to obtain triples;

[0009] Step S400: Based on the triples, construct a knowledge graph for aero-engine quality traceability.

[0010] Further, step S100 includes:

[0011] Step S101: Extract structured data from the full-process data records of aero-engine production; the structured data includes raw material warehousing record data, processing process record data, manufacturing execution record data, and assembly process record data;

[0012] Step S102: Extract unstructured data from the full-process data records of aero-engine production; the unstructured data includes engine fault record data, engine maintenance record data, engine inspection record data, non-conformance review form data, defect report data, and rework and repair form data;

[0013] Step S103: Perform data cleaning on the obtained structured data and unstructured data respectively to generate an ontology framework;

[0014] In the above steps, the structured data includes but is not limited to data extracted from raw material warehousing records (warehousing receipts, batch records), processing process records (processing process documents, processing inspection records), manufacturing execution records, assembly process records, etc.; the unstructured data includes but is not limited to data extracted from fault records, engine maintenance records, engine overhaul records, non-conformance review forms, defect reports, rework and repair forms, etc.; for the data and documents involved in each link, experts clearly define the entity relationship types contained in the quality-related data during the aero-engine production cycle, and generate an ontology framework accordingly.

[0015] Further, step S200 includes: performing knowledge extraction on the unstructured data by combining the active learning algorithm with the RB-CasRel model; the active learning algorithm uses uncertainty sampling to label the data points in the unstructured data; constructing the RB-CasRel model based on the CasRel model and using the RoBERTa model as the encoder; the RB-CasRel model uses a deep learning network and combines the Attention mechanism to extract features in the feature extraction part;

[0016] In view of the confidentiality and professionalism of the data in the field of aero-engines in the above steps, it is difficult to obtain data, and the annotation work is time-consuming and costly; when the amount of training data is limited, the adaptability of the model in actual applications will be restricted; the present invention introduces an active learning algorithm to intelligently select the most valuable samples for annotation, thereby maximizing the performance of the model within a limited budget; specifically, active learning ensures good training of the model globally by selecting data points that are helpful for improving the model performance, using uncertainty sampling, diversity sampling, or a hybrid method based on uncertainty and diversity;

[0017] The above steps take fault records as an example. This type of data belongs to typical unstructured data. In the fault records of an aero-engine fuel control system, the annotation of entities covers components of the fuel control system (such as fuel pumps, fuel metering devices, etc.), fault types (such as unstable fuel supply, main pump cavitation, pipeline leakage, etc.); the annotation of relationships includes the connection relationships between components, the causal relationships between faults, etc.; in addition, the attribute annotation of entities is the performance parameters of components (such as temperature, pressure, etc.).

[0018] Based on the encoding layer of the RoBERTa model, the above steps introduce a deep learning network (bidirectional long short-term memory network BiLSTM) and an attention mechanism (Attention) to further enhance the model's learning ability for semantic information. The bidirectional long short-term memory network BiLSTM can capture bidirectional dependencies in the text, and the attention mechanism can help the model focus on the semantic information relevant to the current task, thereby improving the model's ability to identify entities and relationships.

[0019] Furthermore, the RB-CasRel model is divided into three layers, including an encoding layer, a head entity recognition layer, and a relationship and tail entity recognition layer.

[0020] Encoding layer: The RoBERTa model is used as an encoder to perform semantic encoding on the input unstructured data.

[0021] Head entity recognition layer: A binary classifier is used to predict the probabilities that each character in the encoding output by the encoding layer belongs to the start position and the end position of the head entity respectively. Based on the probabilities that each character belongs to the start position and the end position of the head entity, the position of the head entity is determined.

[0022] Relationship and tail entity recognition layer: A multi-level pointer network is used for decoding. For each head entity recognized by the head entity recognition layer, the RoBERTa model judges the probabilities that each character belongs to the start position and the end position of the tail entity for the relationship category corresponding to the head entity. Based on the probabilities that the character belongs to the start position and the end position of the tail entity, the position of the tail entity is determined.

[0023] The encoding and feature extraction parts of the CasRel model are improved in the RB-CasRel model proposed in the above steps. As a method for jointly extracting entity relationships based on parameter sharing, the CasRel model can effectively solve the problem of overlapping entity relationships. However, in response to the specific requirements of the aero-engine quality traceability field, the CasRel model is improved, making it show higher accuracy and generalization ability when dealing with complex text data with overlapping relationships, and thus being more suitable for the extraction of aero-engine quality traceability knowledge.

[0024] Further, step S200 includes: if the structured data is tabular data, use Python tools to extract knowledge from the tabular data; if the structured data is data in a relational database, use D2R tools to extract knowledge from the data in the relational database;

[0025] Taking the structured data in the above steps as an example of a parts procurement form, the entities extracted include the following: part number, part name, specification model, supplier name, warehousing time, batch number, etc.; these structured data can be preprocessed and cleaned through Python scripts to ensure data consistency and accuracy; the D2R tool is used to map the data in the relational database to the RDF (Resource Description Framework) format, thus realizing the conversion from structured data to a knowledge graph.

[0026] Further, step S300 includes:

[0027] Step S301 includes: input each vocabulary output in the knowledge extraction of structured data and unstructured data into the RoBERTa model to generate a word vector corresponding to each vocabulary;

[0028] Step S302: Calculate the cosine similarity between every two word vectors, and calculate the edit distance score between every two word vectors through the edit distance algorithm;

[0029] Step S303: Perform a weighted average on the cosine similarity and edit distance score corresponding to every two word vectors to obtain an entity similarity score;

[0030] Step S304: If the entity similarity score is lower than the set score threshold, determine that the two vocabularies do not match; if the entity similarity score is higher than the set similarity score threshold, determine that the two vocabularies match, and generate a triple according to the relationship between the two vocabularies;

[0031] In the structured data and unstructured data of aeroengines, the same entity often has multiple different description methods, which brings challenges to knowledge fusion; the above steps can effectively solve the semantic difference problem between entity descriptions through the combination of cosine similarity and the edit distance algorithm, thus realizing the precise alignment and fusion of entities in the knowledge graph, providing a solid foundation for subsequent quality traceability analysis; the above edit distance refers to the minimum number of edit operations required to transform one string into another string.

[0032] Further, step S400 stores the triples in a graph database to generate an aeroengine quality traceability knowledge graph;

[0033] The above steps store the extracted triples in Neo4j, construct an aero-engine quality traceability knowledge graph, and realize the quality traceability of the aero-engine production process; in the process of constructing the knowledge graph, as a high-performance graph database, Neo4j can efficiently store and manage a large amount of triple data; by using the Cypher query language of Neo4j, data query and analysis can be flexibly carried out, so as to realize the efficient management and application of the aero-engine quality traceability knowledge graph.

[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0035] Efficiently process complex data: The production data of aero-engines is diverse in form and strong in professionalism. The present invention comprehensively collects structured and unstructured data, screens and cleans it to ensure data quality; at the same time, an ontology is constructed by defining entity and relationship types based on the whole-process production data, laying a solid foundation for subsequent processing; this method standardizes data management, solves the problems of data diversity and complexity, makes the originally scattered and chaotic data orderly, and is convenient for in-depth mining and utilization.

[0036] Accurately extract and integrate knowledge: For structured data, Python and D2R tools are flexibly used to accurately extract data and convert it into the format required by the knowledge graph; when dealing with unstructured data, the active learning algorithm is innovatively combined with the RB-CasRel model; the active learning algorithm reduces the annotation cost, solves the problem of limited training data, and improves the generalization ability of the model; the RB-CasRel model is optimized from the CasRel model. With the RoBERTa encoder, BiLSTM model and Attention mechanism, it can accurately identify entities and relationships in complex texts; in the entity alignment link, the cosine similarity is combined with the edit distance algorithm to effectively solve the semantic differences of different descriptions of the same entity, realize the accurate integration of knowledge, and greatly improve the accuracy and reliability of knowledge extraction.

[0037] Realize fast and accurate quality traceability: Store the extracted triples in the Neo4j graph database to construct a knowledge graph. With the help of the Cypher query language, data can be quickly and flexibly queried and analyzed; when quality problems occur in aero-engines, the root causes of the problems can be quickly located, traced back to all links such as design, production, assembly, and inspection, providing key evidence for solving quality problems, and greatly improving the efficiency and accuracy of quality traceability. Description of the Drawings

[0038] Figure 1 It is a schematic flowchart of the method for an aero-engine quality traceability method based on a knowledge graph according to the present invention;

[0039] Figure 2Example diagram for constructing the ontology of unstructured data in a method for quality traceability of aero-engines based on a knowledge graph according to the present invention;

[0040] Figure 3 Schematic diagram of the RB-CasRel entity-relationship joint extraction model structure in a method for quality traceability of aero-engines based on a knowledge graph according to the present invention;

[0041] Figure 4 Example diagram of the entity-relationship extraction result in a method for quality traceability of aero-engines based on a knowledge graph according to the present invention. Detailed implementation manners

[0042] Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0043] Embodiment: As Figures 1 - 4 shown, the present invention provides a technical solution, a method for quality traceability of aero-engines based on a knowledge graph, and the aero-engine quality traceability method includes:

[0044] Step S100: Based on the full-process data records of aero-engine production, obtain structured data and unstructured data, and generate an ontology framework;

[0045] Among them, step S100 includes:

[0046] Step S101: Extract structured data from the full-process data records of aero-engine production; the structured data includes raw material warehousing record data, processing process record data, manufacturing execution record data, and assembly process record data;

[0047] Step S102: Extract unstructured data from the full-process data records of aero-engine production; the unstructured data includes engine fault record data, engine maintenance record data, engine inspection record data, non-conformance review form data, defect report data, and rework and repair form data;

[0048] Step S103: Perform data cleaning on the obtained structured data and unstructured data respectively, and generate an ontology framework;

[0049] As Figure 2Taking the fault record text shown as an example, the ontology defined in the present invention at least includes an aero-engine system, subsystems, components, component numbers, component names, specifications and models, supplier names, storage times, production batch numbers, technical parameters, production dates, processing personnel, processing equipment, material specifications, materials, quantities, inspection statuses, operating personnel, operating times, delivery times, fault modes, fault phenomena, fault causes, detection methods, detection tools, maintenance measures, maintenance dates, maintenance personnel, solution plans, etc.; the relationships at least include: design-related relationships, process-related relationships, raw material procurement relationships, component processing relationships, quality inspection relationships, fault-related relationships, after-sales service relationships, etc.;

[0050] Step S200: Perform knowledge extraction on unstructured data and structured data respectively;

[0051] Among them, the unstructured data is subjected to knowledge extraction by combining an active learning algorithm with an RB-CasRel model; the active learning algorithm uses uncertainty sampling to label data points in the unstructured data; an RB-CasRel model is constructed based on the CasRel model and using the RoBERTa model as an encoder; the RB-CasRel model uses a deep learning network and combines an Attention mechanism in the feature extraction part to extract features;

[0052] In an embodiment of the present invention, the maximum entropy uncertainty metric method in uncertainty sampling is used to identify data points in the structured data that are difficult for the model to distinguish, and the identification method is as follows:

[0053] Uncertainty(x) = -∑ i P(y i |x)log(P(y i |x));

[0054] Among them, P(y i |x) is the i-th predicted probability distribution of the model for the sample x;

[0055] Use a small amount of labeled data as a training set to train an initial RB-CasRel model. The purpose of the initial training is to let the model learn the basic entity relationship extraction ability on a small amount of data; hand over the data samples that the model identifies as difficult to distinguish to experts for annotation; after the annotation is completed, add the newly annotated data to the training set and retrain the RB-CasRel model; repeat the above steps until the model performance reaches the expectation or the annotation budget is exhausted; in each iteration, the model will learn from the newly annotated data and gradually improve the performance;

[0056] Such as Figure 3As shown, the RB-CasRel model is divided into three layers, including an encoding layer, a head entity recognition layer, and a relation and tail entity recognition layer;

[0057] In the encoding layer, the RoBERTa model is used as an encoder to perform semantic encoding on the input unstructured data;

[0058] In the embodiment of the present invention, in the design of the encoding layer, the RoBERTa model is used as an encoder to perform deep semantic encoding on the input text; the RoBERTa model captures complex semantic information and long-distance dependency relationships in the text through its powerful context modeling ability;

[0059] In the head entity recognition layer, a binary classifier is used to predict the probability that each character in the encoding output by the encoding layer belongs to the start position and end position of the head entity; based on the probability that each character belongs to the start position and end position of the head entity, the position of the head entity is determined;

[0060] In the embodiment of the present invention, the head entity recognition layer is the first decoding stage of the model, and its goal is to decode the start position and end position of the head entity from the output of the encoding layer; for this purpose, two independent binary classifiers are constructed to predict the probability that each character belongs to the start position and end position of the head entity respectively; the calculation formula is as follows:

[0061] p start (i) = σ(W start h i + b start );

[0062] p end (i) = σ(W end h i + b end );

[0063] Among them, p start (i) represents the probability that character i is the start position of the entity; represents the probability that character i is the end position of the entity; W start and W end respectively represent the weight matrices corresponding to the binary classifiers for predicting the probabilities of the start position and end position; b start and b end are bias terms; σ represents the activation function; h i represents the i-th character;

[0064] By optimizing the following likelihood function, the model can determine the accurate position of the entity:

[0065]

[0066] Among them, y start (i) and yend (i) are the true labels indicating whether character i is the start position and end position of the head entity respectively;

[0067] For the relation and tail entity recognition layer, a multi-level pointer network is used for decoding. For each head entity recognized by the head entity recognition layer, the RoBERTa model judges the probability that each character belongs to the start position and end position of the tail entity for the relation category corresponding to the head entity; the tail entity position is determined through the probabilities that the character belongs to the start position and end position of the tail entity.

[0068] In the embodiments of the present invention, the relation tail entity is decoded using a multi-level pointer network. For each recognized head entity, the model respectively judges the probabilities that each character belongs to the start position and end position of the tail entity for its corresponding relation category; specifically, for relation r, the probabilities of the start position and end position of the tail entity are calculated as follows:

[0069]

[0070] where is the representation of the i-th character combined with the head entity feature; and respectively represent the probability of the start position and end position of the tail entity; and respectively represent the weight matrices of the binary classifiers for judging the probabilities of the start position and end position for relation r; and are bias terms;

[0071] In the tail entity extraction stage, the start and end positions of the tail entity are determined by optimizing the following likelihood function:

[0072]

[0073] where and are the true labels indicating whether Tokeni is the start position and end position of the tail entity of relation r respectively, and n is the length of the sequence;

[0074] where step S200 includes: if the structured data is tabular data, use Python tools to perform knowledge extraction on the tabular data; if the structured data is data in a relational database, use D2R tools to perform knowledge extraction on the data in the relational database;

[0075] such as Figure 4As shown in the figure, an example diagram of an entity relationship extraction structure is provided. For the unstructured text "The spline of the fuel pump shaft is worn, resulting in the engine running out of fuel and stalling", the entities are extracted as fuel pump, shaft spline, engine, wear, and running out of fuel and stalling, and the relationships are included, occurred, and caused. The extracted triples are [fuel pump, includes, shaft spline], [shaft spline, occurred, wear], [engine, occurred, running out of fuel and stalling], [wear, caused, running out of fuel and stalling];

[0076] Step S300: Perform knowledge fusion on the vocabulary output by structured data and unstructured data knowledge extraction to obtain triples;

[0077] Step S301 includes: Input each vocabulary output by structured data and unstructured data knowledge extraction into the RoBERTa model to obtain the word vector corresponding to each vocabulary;

[0078] Step S302: Calculate the cosine similarity between every two word vectors, and calculate the edit distance score between every two word vectors through the edit distance algorithm;

[0079] In the embodiment of the present invention, the cosine similarity calculation formula of word embedding is:

[0080]

[0081] For string x and string y, their edit distance calculation is as shown in the formula:

[0082]

[0083] where, lev x,y (i, j) represents the edit distance between the first i characters of x and the first j characters of y; the first formula in the min operation represents the delete character operation, the second formula represents the insert character operation, and the third formula represents the replace operation; is the indicator function, when x i ≠y j , the value is 1, otherwise it is 0; take the reciprocal of the edit distance as the edit distance score to ensure that the more similar the text, the higher its score, that is

[0084] Step S303: Perform weighted average on the cosine similarity and edit distance score corresponding to every two word vectors to obtain the entity similarity score;

[0085] In the embodiment of the present invention, the similarity sim(x, y) calculation formula of entity x and entity y is as follows:

[0086]

[0087] Step S304: If the entity similarity score is lower than the set score threshold, it is determined that the two words do not match; if the entity similarity score is higher than the set similarity score threshold, it is determined that the two words match, and a triple is generated according to the relationship between the two words.

[0088] Step S400: Based on the triple, construct an aero-engine quality traceability knowledge graph.

[0089] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for tracing the quality of aircraft engines based on knowledge graph, characterized in that: The method comprises: Step S100: Based on the data records of the entire process of aircraft engine production, structured data and unstructured data are acquired, and an ontology framework is generated; Step S200: extracting knowledge from unstructured data and structured data respectively; Step S300: performing knowledge fusion on the vocabulary output by knowledge extraction of structured data and unstructured data to obtain triples; Step S400: storing the triples in a graph database to generate an aircraft engine quality traceability knowledge graph.

2. The method for tracing quality of aircraft engines based on knowledge graph according to claim 1 is characterized in that: Step S100 includes: Step S101: extracting structured data from the data records of the entire process of aircraft engine production; the structured data includes raw material storage record data, processing record data, manufacturing execution record data, and assembly process record data; Step S102: extracting unstructured data from the data records of the entire process of aircraft engine production; the unstructured data includes engine failure record data, engine maintenance record data, engine inspection record data, non-conformity review data, defect report data, and rework and repair data; Step S103: clean the acquired structured data and unstructured data respectively to generate an ontology framework.

3. The method for tracing quality of aircraft engines based on knowledge graph according to claim 1 is characterized in that: The step S200 includes: extracting knowledge from unstructured data by combining an active learning algorithm with the RB-CasRel model; the active learning algorithm uses uncertainty sampling to label data points in the unstructured data; constructing the RB-CasRel model based on the CasRel model and using the RoBERTa model as an encoder; the RB-CasRel model uses a deep learning network in the feature extraction part and combines the Attention mechanism to extract features.

4. The method for tracing quality of aircraft engines based on knowledge graph according to claim 3 is characterized by: The RB-CasRel model is divided into three layers including encoding layer, head entity recognition layer, relation and tail entity recognition layer; The encoding layer uses the RoBERTa model as an encoder to semantically encode the input unstructured data; The head entity recognition layer predicts the probability of each character in the code output by the coding layer belonging to the starting position and the ending position of the head entity through a binary classifier; and determines the position of the head entity through the probability of each character belonging to the starting position and the ending position of the head entity; The relationship and tail entity recognition layer uses a multi-level pointer network for decoding. For each head entity recognized by the head entity recognition layer, the RoBERTa model determines the probability of each character belonging to the starting position and the ending position of the tail entity based on the relationship category corresponding to the head entity; the tail entity position is determined by the probability that the character belongs to the starting position and the ending position of the tail entity.

5. The method for tracing quality of aircraft engines based on knowledge graph according to claim 1, characterized in that: The step S200 includes: if the structured data is tabular data, using Python tools to extract knowledge from the tabular data; if the structured data is data in a relational database, using D2R tools to extract knowledge from the data in the relational database.

6. The method for tracing quality of aircraft engines based on knowledge graph according to claim 1, characterized in that: Step S300 includes: Step S301: input each word output in the structured data and unstructured data knowledge extraction into the RoBERTa model to generate a word vector corresponding to each word; Step S302: Calculate the cosine similarity between every two word vectors, and calculate the edit distance score between every two word vectors using the edit distance algorithm; Step S303: performing weighted average of the cosine similarity and edit distance scores corresponding to every two word vectors to obtain an entity similarity score; Step S304: if the entity similarity score is lower than the set score threshold, it is determined that the two words do not match; if the entity similarity score is higher than the set similarity score threshold, it is determined that the two words match, and a triple is generated according to the relationship between the two words.