Aerospace field quality problem return-to-zero report automatic generation method based on large model

By generating zero-reports of aerospace quality issues using a large-model-based approach, the problems of low generation efficiency and poor accuracy in existing technologies have been solved. This enables rapid and accurate analysis of quality issues and automatic generation of effective measures, promoting the accumulation of aerospace technology knowledge and improving management levels.

CN121543569APending Publication Date: 2026-02-17BEIJING JINGHANG COMPUTING & COMM RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511665397.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

The existing methods for generating zero-report quality issues in the aerospace field are inefficient, making it difficult to analyze problems accurately and timely and formulate effective measures. Furthermore, they lack consistency and standardization, which affects the transmission of knowledge.

Method used

A large model-based approach is adopted to obtain user-input fault phenomenon and cause analysis text, and generate a quality problem zeroing report using structured knowledge triples and fault trees, including the automatic generation of fault location, mechanism analysis, corrective and preventive measures.

Benefits of technology

It enables the rapid generation of zero-out reports for quality issues, improves the accuracy and consistency of analysis, reduces human error, supports timely improvement measures, and promotes knowledge sharing and inheritance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121543569A_ABST
    Figure CN121543569A_ABST
Patent Text Reader

Abstract

The invention discloses a method for automatically generating a spaceflight field quality problem return-to-zero report based on a large model. The method comprises the following steps: acquiring a spaceflight product fault phenomenon and reason analysis text input by a user; screening a historical fault candidate set from the historical fault data, and obtaining a historical fault association set from the historical fault knowledge graph based on the candidate set; performing report analysis and information extraction on the fault phenomenon and reason analysis input by the user and the historical fault association set to obtain a structured knowledge triple, and generating a fault tree fusing a product physical structure; semantic integration is carried out based on the fault tree, and a problem positioning text and a mechanism analysis text are generated; respectively generating a corrective measure text, a preventive measure text and a first-third measure text; and finally outputting a return-to-zero report conforming to the format specification. The method aims at overcoming the defects of an existing spaceflight field quality problem return-to-zero report generation mode, and rapid, accurate and standard generation of the quality problem return-to-zero report is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of aerospace engineering and quality control technology, and in particular to an automatic generation method for zeroing out quality problems in the aerospace field based on a large model. Background Technology

[0002] As a highly complex industry with extremely high requirements for reliability and safety, the aerospace field is where even the slightest quality problem can lead to serious consequences, even causing the failure of the entire space mission, resulting in incalculable economic losses and adverse international impact. Therefore, the zero-quality problem report plays a crucial role in the aerospace field.

[0003] A quality issue zero-tolerance report is an important document that comprehensively analyzes and summarizes quality problems that arise during the design, production, testing, and service of aerospace products. It delves into the root causes of quality problems from both technical and managerial perspectives, and formulates corresponding corrective and preventative measures to avoid recurrence of similar problems in subsequent space missions. At the same time, the quality issue zero-tolerance report also serves as an important vehicle for aerospace companies to accumulate knowledge and pass on experience, providing valuable reference for future aerospace projects and promoting the continuous improvement of aerospace technology and management levels.

[0004] Currently, the main methods for generating zero-quality issues reports in the aerospace field include traditional manual writing and simple template application.

[0005] In the traditional manual documentation approach, when quality issues arise, relevant technical and management personnel need to invest a significant amount of time and effort in collecting various information related to the quality problems, including product design documents, production records, test reports, and descriptions of malfunctions. Then, leveraging their professional knowledge and experience, they analyze and organize this information, attempting to pinpoint the exact location of the problem, analyze its root cause, formulate corresponding corrective measures, and consider how to learn from this experience to prevent similar problems from recurring.

[0006] While traditional manual drafting of quality issue zero-out reports can fully leverage the expertise of technical and managerial personnel, it has several drawbacks. Firstly, manual drafting is time-consuming and inefficient. Quality issues in aerospace products often involve multiple systems and processes, with complex and time-consuming information collection and analysis. This results in a lengthy report generation cycle, hindering timely support for subsequent improvement measures. For example, in the zero-out process of a complex aerospace system, technicians spent several weeks completing information collection and preliminary analysis, severely impacting the timeliness of problem resolution. Secondly, manual analysis is prone to errors. Due to the complexity of quality issues and the influence of human factors, technicians may make subjective judgment errors or omit key information during analysis, leading to inaccurate problem identification and incomplete mechanism analysis, thus affecting the effectiveness of corrective and preventative measures. Furthermore, differences in writing styles and professional levels among personnel make it difficult to ensure the standardization and consistency of reports, hindering knowledge sharing and transmission.

[0007] Using a simple template simplifies the report generation process to some extent. Companies can pre-define a quality issue zero-out report template, which includes the basic framework and content requirements for each section. When a quality issue arises, staff simply need to fill in the relevant information about the specific quality issue in the corresponding field of the template to generate a quality issue zero-out report.

[0008] While using simple templates improves report generation efficiency, it also has significant drawbacks. This approach lacks specificity; template content is often generic and fails to fully adapt to diverse and complex quality issues. For some unique and novel quality problems, simply applying a template may not accurately reflect the essence and key aspects of the problem, resulting in a superficial report that fails to truly address the quality issue. Furthermore, template-based approaches can limit staff thinking, leading to over-reliance on templates and a lack of in-depth reflection and innovative analysis, hindering the identification of potential quality risks and improvement opportunities. Summary of the Invention

[0009] Based on the above analysis, this invention aims to provide an automatic generation method for zero-out reports of quality problems in the aerospace field based on a large model. This method overcomes the shortcomings of existing methods for generating zero-out reports of quality problems in the aerospace field and utilizes the powerful data analysis and natural language processing capabilities of the large model to achieve rapid, accurate, and standardized generation of zero-out reports of quality problems.

[0010] On one hand, embodiments of the present invention provide an automatic generation method for zeroing out quality issues in the aerospace field based on a large model. This method includes the following steps:

[0011] Obtain user-inputted text describing the malfunction phenomena and causes of aerospace products;

[0012] Faults that are associated with the fault phenomena entered by the user are filtered out from the historical fault data in the database to form a historical fault association set;

[0013] The system performs report parsing and information extraction on the user-input fault phenomena, cause analysis, and the historical fault association set to obtain structured knowledge triples stored in the form of "entity-relationship-entity".

[0014] Based on the structured knowledge triples and the tree structure describing the physical composition of the product, a fault tree that integrates the physical structure of the product is generated.

[0015] Based on the fault tree, historical faults represented by leaf nodes of the fault tree are extracted, semantic integration is performed, and problem location text is generated; based on the traversal of the fault tree, multiple causal chains are formed from bottom to top, and mechanism analysis text is generated.

[0016] Based on the fault tree and the structured knowledge triples, corrective action text, preventive action text, and inference-based action text are generated respectively.

[0017] Integrate the problem location text, mechanism analysis text, corrective action text, preventive action text, and lessons learned and similar measures text into the aerospace zeroing report standard template, and output a zeroing report that conforms to the format specifications.

[0018] Furthermore, the step of filtering out faults associated with the fault phenomena input by the user from the historical fault data in the database includes: filtering out a candidate set of historical faults from the historical fault data in the database by calculating the semantic similarity of text, and further filtering the faults in the candidate set by calculating the component correlation degree based on the historical fault knowledge graph to obtain a set of historical fault associations.

[0019] The text semantic similarity calculation includes:

[0020] The user-inputted text describing the fault phenomenon and its cause is used as input. A BERT model, fine-tuned based on a corpus of aerospace fault data, is then used to output a context-aware text vector V. text The non-textual information in the user-input description of the fault phenomenon and cause is encoded to obtain a structured vector V. struct ;

[0021] The text vector V is processed using a gating fusion mechanism. text With structured vector V struct Dynamic fusion is performed to obtain the fusion vector V. fusion ;

[0022] The same method is used to convert the historical fault data in the database into fusion vectors of the same dimension. The cosine similarity between the fusion vector corresponding to the user input fault and the fusion vector corresponding to the historical fault data is calculated. Historical faults with similarity greater than a threshold are selected to form a candidate set of historical faults.

[0023] Furthermore, the component correlation calculation includes:

[0024] Assign corresponding weighting factors to the relationship types between different entities in the historical fault knowledge graph;

[0025] The user-input fault is used as the target node, and each fault in the historical fault candidate set is used as the starting node. The weighted path distance between each starting node and the target node is calculated using the improved Dijkstra algorithm.

[0026] The joint correlation score is calculated by combining text semantic similarity and weighted path distance. The joint correlation scores are ranked, and a preset number of historical faults with the highest rankings are selected to form a historical fault association set.

[0027] Furthermore, the report parsing and information extraction include:

[0028] The user-input text describing the fault phenomenon and its causes, along with the text from the historical fault association set, are converted into low-dimensional dense vectors. These vectors are then input into an LSTM network to obtain a vector representation containing contextual semantics. Finally, the vector representation is input into a conditional random field model to obtain scores for different entity label sequences. The entity label sequence with the highest score among all possible entity label sequences is selected as the identified entity.

[0029] For the identified entities, a neural network-based relation classifier is used to determine the relationships between entities. The relationship with the highest probability is selected as the identified relationship, forming a structured knowledge triple of "entity-relationship-entity".

[0030] Furthermore, the fault tree generation includes:

[0031] Based on the physical hierarchy of the product structure tree, the entities representing fault phenomena in the structured knowledge triples are divided according to the physical hierarchy of the product structure tree to obtain the faults at the corresponding levels.

[0032] The faults at each level are traversed from top to bottom. For faults at adjacent levels, faults that have corresponding relationships in the structured knowledge triples are connected to construct a preliminary fault association network.

[0033] The fault phenomenon input by the user is taken as the top event. A subtree with the top event as the core is pruned from the preliminary fault association network to obtain a complete fault tree that integrates the physical structure of the product.

[0034] Furthermore, the fusion vector is calculated using the following formula:

[0035] V fusion =σ(W1·V text +W2·V struct +b)⊙V text +(1-σ(W1·V text +W2·V struct +b))⊙V struct

[0036] In the formula, σ is the sigmoid activation function, W1 and W2 are learnable weight matrices that determine the respective "contribution" of text features and structured features in the fusion process, and ⊙ is the element-wise product.

[0037] Furthermore, the joint correlation score is calculated using the following formula:

[0038]

[0039] In the formula, S semantic The semantic similarity of the text is represented by w(u,v), which represents the weighted path distance between the starting node and the target node.

[0040] Furthermore, the conditional random field model is as follows:

[0041]

[0042] In the formula, Z(x) is the normalization factor, and ψ i (y i ,x) is the state characteristic function, ψ i,i+1 (y i ,y i+1 ) is the transition characteristic function, x i y represents the word vector at a single position i in the input fault text sequence x; y represents the sequence of entity labels assigned to x.

[0043] Furthermore, the calculation formula for the relation classifier is as follows:

[0044]

[0045] Among them, R(e i ,e j ) represents the entity e predicted by the model. i and e j The probability of the relationship between them, where W and b are the weight matrix and bias vector, respectively, and Emb(w) k ) represents the k-th word in the sentence. k The word vectors, Pos(w k ,e i ,ej ) represents the k-th word relative to the target entity e. i and e j The positional feature vector, N is the number of input word vectors, LSTM(·) is a Long Short-Term Memory network, Softmax(·) represents converting the mapped vector into a probability distribution of each relation category, and argmax is the positional feature vector. r∈R (·) indicates that the relation with the highest probability is selected as the final judgment result from the set of all possible relation categories R.

[0046] Furthermore, the corrective action text, preventive action text, and inference-based action text are formed by extracting the historical faults corresponding to the leaf nodes and their upper-level nodes of the faulty tree, obtaining the corresponding measures associated with the faults of each leaf node and its upper-level node from the structured knowledge triples, and semantically integrating the three types of measures using large model technology.

[0047] The beneficial effects of this technical solution are:

[0048] This invention's automatic generation method based on a large model significantly reduces report generation time, greatly improves work efficiency, and effectively avoids delays in problem resolution caused by slow report generation. By learning from a large amount of historical data on zero-report quality issues in the aerospace field, it can more accurately analyze problems and provide targeted measures, effectively reducing human error and improving the accuracy and reliability of problem resolution.

[0049] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description

[0050] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts.

[0051] Figure 1 This is a flowchart illustrating the overall method of an embodiment of the present invention;

[0052] Figure 2 This is a flowchart illustrating the structured knowledge triple generation process according to an embodiment of the present invention. Detailed Implementation

[0053] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.

[0054] A specific embodiment of the present invention discloses an automatic generation method for zeroing out quality problems in the aerospace field based on a large model. The overall process of the method is as follows: Figure 1 As shown, the specific steps include:

[0055] Obtain user-inputted text describing the malfunction phenomena and causes of aerospace products;

[0056] Faults that are associated with the fault phenomena entered by the user are filtered out from the historical fault data in the database to form a historical fault association set;

[0057] The system performs report parsing and information extraction on the user-input fault phenomena, cause analysis, and the historical fault association set to obtain structured knowledge triples stored in the form of "entity-relationship-entity".

[0058] Based on the structured knowledge triples and the tree structure describing the physical composition of the product, a fault tree that integrates the physical structure of the product is generated.

[0059] Based on the fault tree, historical faults represented by leaf nodes of the fault tree are extracted, semantic integration is performed, and problem location text is generated; based on the traversal of the fault tree, multiple causal chains are formed from bottom to top, and mechanism analysis text is generated.

[0060] Based on the fault tree and the structured knowledge triples, corrective action text, preventive action text, and inference-based action text are generated respectively.

[0061] Integrate the problem location text, mechanism analysis text, corrective action text, preventive action text, and lessons learned and similar measures text into the aerospace zeroing report standard template, and output a zeroing report that conforms to the format specifications.

[0062] During implementation, when a fault occurs, the information system receives parameters input by the user. These parameters include text describing the fault phenomenon and cause provided by the user, as well as non-textual structured information (such as the time of the fault, ambient temperature, and equipment model), providing an initial data foundation for future tasks.

[0063] Furthermore, the step of filtering out faults associated with the fault phenomena input by the user from the historical fault data in the database includes: filtering out a candidate set of historical faults from the historical fault data in the database by calculating the semantic similarity of text, and further filtering the faults in the candidate set by calculating the component correlation degree based on the historical fault knowledge graph to obtain a set of historical fault associations.

[0064] The text semantic similarity calculation includes:

[0065] The user-inputted text describing the fault phenomenon and its cause is used as input. A BERT model, fine-tuned based on a corpus of aerospace fault data, is then used to output a context-aware text vector V. text The non-textual information in the user-input description of the fault phenomenon and cause is encoded to obtain a structured vector V. struct ;

[0066] The text vector V is processed using a gating fusion mechanism. text With structured vector V struct Dynamic fusion is performed to obtain the fusion vector V. fusion ;

[0067] The same method is used to convert the historical fault data in the database into fusion vectors of the same dimension. The cosine similarity between the fusion vector corresponding to the user input fault and the fusion vector corresponding to the historical fault data is calculated. Historical faults with similarity greater than a threshold are selected to form a candidate set of historical faults.

[0068] Specifically, the goal of this step is to quickly retrieve a set of potential historical fault candidates from massive historical fault data that are similar to the fault phenomena and cause descriptions input by the user. First, the historical fault information contained in the database of user-input fault phenomena and cause descriptions is converted into high-dimensional vectors. Then, based on a BERT model fine-tuned using domain corpus, the input fault phenomenon description is used to output a context-aware text vector V. text The semantic information of the fault text is represented, and then non-textual information (such as fault occurrence time, ambient temperature, and equipment model) is encoded into a structured vector V. struct The semantic information of the structured dimension in the fault scenario is represented, and the fusion vector describing the input fault phenomenon is obtained by dynamically fusing it with the text vector through a gating mechanism.

[0069] Specifically, the fusion vector is calculated using the following formula:

[0070] V fusion =σ(W1·V text +W2·V struct +b)⊙V text +(1-σ(W1·V text +W2·V struct +b))⊙V struct

[0071] In the formula, σ is the sigmoid activation function, W1 and W2 are learnable weight matrices that determine the respective "contribution" of text features and structured features in the fusion process, and ⊙ is the element-wise product.

[0072] Specifically, to reduce the decay or expansion of gradients during propagation, the following formula is used to form the initial weight matrix W.

[0073]

[0074] Where, n in For the input dimension, i.e., V text Or V struct The dimension value of n. out For the output dimension, this solution adopts V. text Or V struct The dimensional value with a relatively large dimension.

[0075] After generating the vectors for user-input faults, the same operation is performed on historical faults, converting all existing historical faults into fused vectors of the same dimension. Then, semantic relevance is evaluated by calculating the cosine similarity between the fused vectors of user-input faults and historical faults. The calculation method is as follows:

[0076]

[0077] Where A and B represent the fusion vectors of user input faults and historical faults, respectively; A i B i represents the component values ​​of vectors A and B in the i-th dimension, respectively; n represents the dimension of vectors A and B.

[0078] The above method is executed repeatedly until the cosine similarity between all user input faults and historical faults is calculated. Historical faults with a similarity greater than the threshold are selected and set to form a candidate set of historical faults.

[0079] Preferably, the threshold is set to 0.8.

[0080] Furthermore, the component correlation calculation includes:

[0081] Assign corresponding weighting factors to the relationship types between different entities in the historical fault knowledge graph;

[0082] The user-input fault is used as the target node, and each fault in the historical fault candidate set is used as the starting node. The weighted path distance between each starting node and the target node is calculated using the improved Dijkstra algorithm.

[0083] The joint correlation score is calculated by combining text semantic similarity and weighted path distance. The joint correlation scores are ranked, and a preset number of historical faults with the highest rankings are selected to form a historical fault association set.

[0084] Specifically, the goal of this step is to use the structured relationships in the historical fault knowledge graph to further refine and filter the historical faults in the obtained historical fault candidate set, forming a historical fault association set filtered by the historical fault knowledge graph. This can solve the problem that relying solely on text semantic similarity may not be able to capture the deep structured relationships between faults.

[0085] Specifically, the historical fault knowledge graph describes the relationships between models, products, and faults. Its node types include model, product, and fault, and its relationship types include structural inclusion relationships between models and products, functional dependencies between different products, causal relationships between fault phenomena, and correlation relationships. Nodes of type "model" represent the model that experienced the fault; nodes of type "product" represent the product name that experienced the fault; and nodes of type "fault" represent previously occurring faults. In engineering, it is generally believed that a fault in a product is more likely to be caused by a fault in a product that it contains than by a fault in a product it does not contain. Because the historical fault knowledge graph describes the structural inclusion relationships between models and products, as well as the relationships between products and faults, the similarity between two faults is represented in the historical fault knowledge graph by the distance of their paths.

[0086] Specifically, this invention maps relationship types (such as "contains", "cause", and "depends") in the knowledge graph to weighting factors, and associates the user-input fault as a fault phenomenon node with the product node that caused the fault. Then, using this node as the target node, the weighted path distance between the target node and the nodes corresponding to each fault in the candidate set is calculated.

[0087] Furthermore, the relationships contained in the knowledge graph and their corresponding weighting factors are as follows:

[0088] Relationship type Semantic description Weighting factor w(u,v) causation u directly causes v to fail. 0.1 Functional Dependencies u's functionality depends on v. 0.3 Structural inclusion relationship u is a component of v 0.6 Relationship There is an indirect relationship between u and v. 0.8

[0089] During the calculation, an improved Dijkstra algorithm is used. The historical fault knowledge graph, target nodes (user-input faults), starting nodes (all faults in the candidate set), and weighting factors (the weighting factors corresponding to each type of relation in the path type weighting) are input into the algorithm. First, the distances from all nodes in the historical fault knowledge graph to the starting node are initialized to infinity, except for the starting node, whose distance is 0, and these are placed in a priority queue. Then, the node with the smallest current distance is iterated from the queue and explored. All its adjacent nodes are traversed, and the new distance to that adjacent node via the current node is calculated (the current node distance plus the weighting factor of the connecting edge). If the new distance is smaller, the distance record of the connected node is updated, and the relation from which the current node reached the node is recorded. The priority queue is also updated. This process continues until the target node is removed or the queue is empty. The Dijkstra algorithm is repeated until the distances between the nodes corresponding to each fault in the candidate set and the target node are calculated, resulting in the weighted distance w(u,v) between the nodes corresponding to all faults in the candidate set and the target node.

[0090] Furthermore, a joint relevance score is calculated by combining text semantic similarity and weighted path distance. The joint relevance score is calculated using the following formula:

[0091]

[0092] In the formula, S semantic The semantic similarity of the text is represented by w(u,v), which represents the weighted path distance between the starting node and the target node.

[0093] Furthermore, the correlation scores are sorted, and a preset number of historical fault records with the highest rankings are selected as the historical fault correlation set.

[0094] Preferably, the preset number of the top-ranked items is 70%.

[0095] Furthermore, the report parsing and information extraction include:

[0096] The user-input text describing the fault phenomenon and its causes, along with the text from the historical fault association set, are converted into low-dimensional dense vectors. These vectors are then input into an LSTM network to obtain a vector representation containing contextual semantics. Finally, the vector representation is input into a conditional random field model to obtain scores for different entity label sequences. The entity label sequence with the highest score among all possible entity label sequences is selected as the identified entity.

[0097] For the identified entities, a neural network-based relation classifier is used to determine the relationships between entities. The relationship with the highest probability is selected as the identified relationship, forming a structured knowledge triple of "entity-relationship-entity".

[0098] Specifically, during the report parsing and information extraction process, considering the randomness of the size of the historical fault association set, and to avoid the impact of excessive input data on the efficiency and accuracy of the report parsing and information extraction process, the historical fault association set and the faults input by the user are parsed and information extracted one by one. The process of generating structured knowledge triples is as follows: Figure 2 As shown.

[0099] Specifically, named entity recognition (NER) is first performed to identify key elements in the fault (such as model number, product, fault phenomenon, cause, solution, and generalization methods). During the recognition process, each word in the input is converted into a low-dimensional dense vector to capture the semantic features of the word itself. Next, an LSTM network is used to analyze the contextual information of the words, generating a vector representation for each word that contains complete contextual semantics. Finally, a Conditional Random Field (CRF) model is used to output the final identified entity results. The CRF model represents the input sequence as x = (x1, x2, ..., x...). n When calculating the scores of different entity label sequences (where one of the label sequences is y = (y1, y2, ..., y...), the score is calculated. n The entity label sequence with the highest score is selected as the final recognition result.

[0100] Furthermore, the conditional random field model is as follows:

[0101]

[0102] In the formula, Z(x) is the normalization factor, and ψ i (y i ,x) is the state characteristic function, ψ i,i+1 (y i ,y i+1 ) is the transition characteristic function, x i y represents the word vector at a single position i in the input fault text sequence x; y represents the sequence of entity labels assigned to x.

[0103] Furthermore, after identifying all entities, a neural network-based relation classifier is used to extract relations to determine the relationships between these entities. Entities include elements such as model name, product name, fault phenomenon, cause, solution, and generalization method; relations include expressions such as "caused by," "originates from," and "is a phenomenon of something." First, a fault containing identified entities is input. Each word in the fault is converted into a word vector, and the relative distance of each word to two entities is calculated and encoded as a position vector. The word vectors and position vectors are concatenated to form a distributed representation of each word, which serves as the input to the neural network. Next, sentence features are extracted using a Long Short-Term Memory (LSTM) network, and these features are input to a fully connected layer and the relation classifier to calculate the probability of relationships between entities. Finally, the probability R(e) is selected. i ,e j The highest-ranking relation is used as the identified relation, forming an "entity-relationship-entity" triple. The above method is executed iteratively until all entity relations contained in all faults are extracted.

[0104] Specifically, the calculation formula for the relation classifier is as follows:

[0105]

[0106] Among them, R(e i ,e j ) represents the entity e predicted by the model. i and e j The probability of the relationship between them, where W and b are the weight matrix and bias vector, respectively, and Emb(w) k ) represents the k-th word in the sentence. k The word vectors, Pos(w k ,e i ,e j ) represents the k-th word relative to the target entity e. i and e j The positional feature vector, N is the number of input word vectors, LSTM(·) is a Long Short-Term Memory network, Softmax(·) represents converting the mapped vector into a probability distribution of each relation category, and argmax is the positional feature vector. r∈R (·) indicates that the relation with the highest probability is selected as the final judgment result from the set of all possible relation categories R.

[0107] Specifically, after completing the entity relationship identification process for each fault in the user input fault and historical fault association set, considering that entity relationships may change over time due to technological advancements and improved personnel skills, leading to inconsistencies in relationships between the same entity pairs, a time-weighted relationship strength calculation method is used to determine whether to accept the relationship between entities. For each constructed entity relationship, a time decay factor and a frequency enhancement factor are introduced to dynamically update the relationship weight W.r :

[0108] W r (t)=W r0 ·exp(-λ·(t now -t r ))·(1+α·log(1+f r ))

[0109] Among them, W r0 As the initial relation weights (based on the confidence level extracted from a single report), t r t represents the fault formation time. now Let f be the current time, λ be the decay coefficient, and f be the current time. r Let α be the frequency of co-occurrence of a relation in subsequent failures (e.g., how many failures a relation pair appears in), and let α be the frequency sensitivity.

[0110] Preferably, the initial relationship weight is 0.7, the attenuation coefficient is 0.05, the co-occurrence frequency is 50, and the frequency sensitivity is 0.1.

[0111] Based on the above algorithm, the weight of frequently occurring entity relationships is increased, while the weight of older, less frequent relationships is decreased. For instances where the relationships between identical entity pairs are inconsistent, the relationship with the higher weight between the entity pairs is adopted.

[0112] Furthermore, the extracted entity types include model, product, fault phenomenon, cause, solution, and lessons learned. Among them, the three types of entities—model, product, and fault phenomenon—and their relationships are used for fault tree generation, while the four types of entities—fault phenomenon, cause, solution, and lessons learned—and their relationships are used for generating corrective measures and lessons learned measures in the subsequent zeroing report.

[0113] Furthermore, the fault tree generation includes:

[0114] Based on the physical hierarchy of the product structure tree, the entities representing fault phenomena in the structured knowledge triples are divided according to the physical hierarchy of the product structure tree to obtain the faults at the corresponding levels.

[0115] The faults at each level are traversed from top to bottom. For faults at adjacent levels, faults that have corresponding relationships in the structured knowledge triples are connected to construct a preliminary fault association network.

[0116] The fault phenomenon input by the user is taken as the top event. A subtree with the top event as the core is pruned from the preliminary fault association network to obtain a complete fault tree that integrates the physical structure of the product.

[0117] Specifically, the goal of this step is to effectively align unstructured fault knowledge with the rigorous physical structure of the product, thereby generating a concrete fault tree model that conforms to both logical reasoning and engineering practice. A hierarchical knowledge distillation framework is proposed, forming a fault logic hierarchy (top event, intermediate event, bottom event) based on the product structure hierarchy (model, subsystem, equipment, component, part). Each model has its own unique model name, which corresponds to the model name in the triplet. Subsystems, equipment, components, and parts each have their own unique product names, which correspond to the product names in the triplet. The entities representing fault phenomena in the structured knowledge triplets are divided according to the physical hierarchy of the product structure tree to obtain the faults at the corresponding levels.

[0118] Furthermore, the physical hierarchy of the product structure tree and the corresponding fault representation in the structured knowledge triple are as follows:

[0119]

[0120] Specifically, firstly, based on the hierarchical division results, hierarchical attributes are added to each entity representing a fault phenomenon. Then, following the hierarchy, from top to bottom, fault phenomena at each level are progressively searched for relationships with their subordinate levels. For faults at adjacent levels, faults with corresponding relationships in the structured knowledge triples are connected: starting with model-level faults, all model-level and subsystem-level faults are traversed; if a relationship exists between a model-level fault and a subsystem-level fault, they are connected. Then, all subsystem-level and equipment-level faults are traversed, and similarly, related subsystem-level and equipment-level faults are connected. This process continues until the connections between component-level and part-level faults are completed, constructing a preliminary fault association network. Finally, the user-input fault phenomenon is found and used as a branch of the fault tree for the top node, pruned for use as the fault tree for analyzing that fault.

[0121] Furthermore, based on the fault tree structure, historical faults represented by its leaf nodes are extracted. Large-scale modeling techniques are used to semantically integrate historical fault phenomena, forming problem localization text, thus addressing the shift from "knowing where the fault occurred" to "knowing how the fault was caused."

[0122] Furthermore, by utilizing the logical connections provided by the fault tree as an analytical framework, starting from the leaf nodes, multiple bottom-up causal chains are formed through traversal methods. These causal chains define the propagation path of the fault, demonstrating how a leaf node, through a series of intermediate events, ultimately leads to the occurrence of the top event, thus forming a mechanistic analysis text and solving the problem of moving from "knowing where it is broken" to "explaining why it is broken and how it is broken."

[0123] Furthermore, the corrective action text, preventive action text, and inference-based action text are formed by extracting the historical faults corresponding to the leaf nodes and their upper-level nodes of the faulty tree, obtaining the corresponding measures associated with the faults of each leaf node and its upper-level node from the structured knowledge triples, and semantically integrating the three types of measures using large model technology.

[0124] Furthermore, based on the five principles of "accurate positioning, clear mechanism, problem reproduction, effective measures, and learning from one case to apply to others" for technical zeroing, the existing analysis text is filled into the corresponding chapters of the report template to form a zeroing report. After the zeroing report is generated, users can download, view, modify, and delete it.

[0125] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.

[0126] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for automatically generating zero-out reports on quality issues in the aerospace field based on a large model, characterized in that, The method includes the following steps: Obtain user-inputted text describing the malfunction phenomena and causes of aerospace products; Faults that are associated with the fault phenomena entered by the user are filtered out from the historical fault data in the database to form a historical fault association set; The system performs report parsing and information extraction on the user-input fault phenomena, cause analysis, and the historical fault association set to obtain structured knowledge triples stored in the form of "entity-relationship-entity"; Based on the structured knowledge triples and the tree structure describing the physical composition of the product, a fault tree that integrates the physical structure of the product is generated. Based on the fault tree, historical faults represented by leaf nodes of the fault tree are extracted, semantic integration is performed, and problem location text is generated; based on the traversal of the fault tree, multiple causal chains are formed from bottom to top, and mechanism analysis text is generated. Based on the fault tree and the structured knowledge triples, corrective action text, preventive action text, and inference-based action text are generated respectively. Integrate the problem location text, mechanism analysis text, corrective action text, preventive action text, and lessons learned and similar measures text into the aerospace zeroing report standard template, and output a zeroing report that conforms to the format specifications.

2. The method for automatically generating zero-out reports on quality issues in the aerospace field based on a large model, as described in claim 1, is characterized in that... The step of filtering out faults associated with user-input fault phenomena from historical fault data in the database includes: filtering out a candidate set of historical faults from historical fault data in the database through text semantic similarity calculation, and further filtering the faults in the candidate set by calculating the component correlation degree based on the historical fault knowledge graph to obtain a historical fault association set. The text semantic similarity calculation includes: The user-inputted text describing the fault phenomenon and its cause is used as input. A BERT model, fine-tuned based on a corpus of aerospace fault data, is then used to output a context-aware text vector V. text The non-textual information in the user-input description of the fault phenomenon and cause is encoded to obtain a structured vector V. struct ; The text vector V is processed using a gating fusion mechanism. text With structured vector V struct Dynamic fusion is performed to obtain the fusion vector V. fusion ; The same method is used to convert the historical fault data in the database into fusion vectors of the same dimension. The cosine similarity between the fusion vector corresponding to the user input fault and the fusion vector corresponding to the historical fault data is calculated. Historical faults with similarity greater than a threshold are selected to form a candidate set of historical faults.

3. The method for automatically generating zero-out reports on quality issues in the aerospace field based on a large model, as described in claim 2, is characterized in that... The component correlation calculation includes: Assign corresponding weighting factors to the relationship types between different entities in the historical fault knowledge graph; The user-input fault is used as the target node, and each fault in the historical fault candidate set is used as the starting node. The weighted path distance between each starting node and the target node is calculated using the improved Dijkstra algorithm. The joint correlation score is calculated by combining text semantic similarity and weighted path distance. The joint correlation scores are ranked, and a preset number of historical faults with the highest rankings are selected to form a historical fault association set.

4. The method for automatically generating zero-report of quality problems in the aerospace field based on a large model according to claim 1, characterized in that, The report parsing and information extraction include: The user-input text describing the fault phenomenon and its causes, along with the text from the historical fault association set, are converted into low-dimensional dense vectors. These vectors are then input into an LSTM network to obtain a vector representation containing contextual semantics. Finally, the vector representation is input into a conditional random field model to obtain scores for different entity label sequences. The entity label sequence with the highest score among all possible entity label sequences is selected as the identified entity. For the identified entities, a neural network-based relation classifier is used to determine the relationships between entities. The relationship with the highest probability is selected as the identified relationship, forming a structured knowledge triple of "entity-relationship-entity".

5. The method for automatically generating zero-out reports on quality issues in the aerospace field based on a large model, as described in claim 1, is characterized in that... The fault tree generation includes: Based on the physical hierarchy of the product structure tree, the entities representing fault phenomena in the structured knowledge triples are divided according to the physical hierarchy of the product structure tree to obtain the faults at the corresponding levels. The faults at each level are traversed from top to bottom. For faults at adjacent levels, faults that have corresponding relationships in the structured knowledge triples are connected to construct a preliminary fault association network. The fault phenomenon input by the user is taken as the top event. A subtree with the top event as the core is pruned from the preliminary fault association network to obtain a complete fault tree that integrates the physical structure of the product.

6. The method for automatically generating zero-out reports on quality issues in the aerospace field based on a large model, as described in claim 2, is characterized in that... The fusion vector is calculated using the following formula: V fusion =σ(W1·V text +W2·V struct +b)⊙V text +(1-σ(W1·V text +W2·V struct +b))⊙V struct In the formula, σ is the sigmoid activation function, W1 and W2 are learnable weight matrices that determine the "contribution" of text features and structured features in the fusion process, and ⊙ is the element-wise product.

7. The method for automatically generating zero-out reports on quality issues in the aerospace field based on a large model, as described in claim 3, is characterized in that... The joint correlation score is calculated using the following formula: In the formula, S semantic The semantic similarity of the text is represented by w(u,v), which represents the weighted path distance between the starting node and the target node.

8. The method for automatically generating zero-report of quality problems in the aerospace field based on a large model according to claim 4, characterized in that, The conditional random field model is: In the formula, Z(x) is the normalization factor, and ψ i (y i ,x) is the state characteristic function, ψ i,i+1 (y i ,y i+1 ) is the transition characteristic function, x i y represents the word vector at a single position i in the input fault text sequence x; y represents the sequence of entity labels assigned to x.

9. The method for automatically generating zero-out reports on quality issues in the aerospace field based on a large model, as described in claim 4, is characterized in that... The calculation formula for the relation classifier is as follows: Among them, R(e i ,e j ) represents the entity e predicted by the model. i and e j The probability of the relationship between them, where W and b are the weight matrix and bias vector, respectively, and Emb(w) k ) represents the k-th word in the sentence. k The word vectors, Pos(w k ,e i ,e j ) represents the k-th word relative to the target entity e. i and e j The positional feature vector, N is the number of input word vectors, LSTM(·) is a Long Short-Term Memory network, Softmax(·) represents converting the mapped vector into a probability distribution of each relation category, and argmax is the positional feature vector. r∈R (·) indicates that the relation with the highest probability is selected as the final judgment result from the set of all possible relation categories R.

10. The automatic generation system for zero-out reports of quality problems in the aerospace field based on a large model as described in claim 1, characterized in that, The corrective measures text, preventive measures text, and inference-based measures text are formed by extracting the historical faults corresponding to the leaf nodes and their upper-level nodes of the faulty tree, obtaining the corresponding measures associated with the faults of each leaf node and its upper-level node from the structured knowledge triples, and semantically integrating the three types of measures using large model technology.