Nuclear power unit quality defect source analysis and total value chain tracing method
Through multimodal data feature extraction and dynamic knowledge graph construction, the problem of difficult source calibration and low traceability efficiency of nuclear power equipment quality defects is solved, and the rapid and accurate traceability of the entire value chain is achieved to ensure the safe and reliable operation of the nuclear power unit.
Patent Information
- Application Number
- CN202510727140.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The source calibration of nuclear power equipment quality defects is difficult and the traceability efficiency is low. Traditional methods rely on expert experience, resulting in a long traceability process, high cost and low efficiency, and the rapid and accurate traceability of the entire value chain cannot be achieved.
Multimodal data feature extraction and dynamic knowledge graph construction, combined with BERT, BiGRU-Attention and CRF sequence labeling technologies, a hybrid database architecture is built, defect entity labeling and association rule mining is carried out, sample balance is optimized using chi-square test and NSGAII algorithm, and knowledge graphs of multiple types of nodes are designed for intelligent traceability.
It realizes rapid and accurate traceability of quality defects from design to operation and maintenance, improves the safe and reliable operation of nuclear power units, and reduces time and human resources costs.
Smart Images

Figure CN120258632A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of quality control of nuclear power equipment, and relates to a method for analyzing the sources of quality defects and full-value chain traceability of nuclear power units. Background Art
[0002] As a clean energy source, nuclear power occupies an important position in the global energy structure. The safe and stable operation of nuclear power units is of crucial significance for ensuring energy supply and reducing environmental pollution. However, nuclear power equipment has significant characteristics such as a giant system, a long service life, and full operating conditions. Its production and manufacturing belong to a semi-continuous and semi-discrete hybrid manufacturing industry, involving multiple complex processes such as design, manufacturing, construction, and commissioning. The complexity of the control process leads to problems such as difficulty in accurately calibrating the sources of quality defects and difficulty in tracing the quality control.
[0003] Traditional quality traceability methods have limitations. Because traditional quality traceability methods for nuclear power equipment usually involve forming an expert group and conducting step-by-step investigations according to the established quality traceability process of the enterprise. This process highly relies on expert experience and lacks intelligent analysis means. In actual operation, the resources of each platform within the nuclear power system are independent of each other, and it is extremely difficult to communicate with each other, making the traceability process lengthy. In terms of time cost, the entire traceability cycle may be as long as more than 3 months, consuming a large amount of human and material resources, with low traceability efficiency and high costs.
[0004] Quality defects will have a significant impact on nuclear power safety and benefits. During the operation of nuclear power equipment, any quality defect and failure will not only reduce production efficiency but may also trigger serious safety accidents. For example, the failure of key equipment in a nuclear power unit may lead to the shutdown of the nuclear reactor, affecting normal power generation and causing huge economic losses. At the same time, once a nuclear safety accident occurs, it will bring immeasurable harm to the surrounding environment and public health. In addition, quality problems will also affect the progress of nuclear power plant construction, increase investment costs, reduce the social and economic benefits of nuclear power projects, and pose a serious threat to the sustainable development of the nuclear power industry.
[0005] With the rapid development of information technology, intelligent means have been widely used in various industries. In the field of nuclear power, building a cyber-physical fusion system based on the traceability of abnormal nuclear power product quality, using intelligent technologies to integrate and refine quality traceability-related data resources, establishing a structured information expression system, and realizing fast and accurate traceability of quality defects have become an urgent need to ensure nuclear power quality safety and improve economic benefits. Through intelligent traceability methods, the efficiency and accuracy of quality traceability can be effectively improved, and measures can be taken in a timely manner to prevent or mitigate the occurrence and development of accidents, ensuring the safe and reliable operation of nuclear power units.
[0006] To solve the above problems, various methods have been proposed in existing research. For example, in the patent CN114969267A "A Method for Analyzing the Causes of Nuclear Power Quality Defects", natural language processing technology is used to extract text keywords to form standardized defect targets and causes. Chi-square test is used to screen features, a random forest algorithm is combined to train a model, and a genetic algorithm is used to optimize parameters to improve the accuracy of defect cause analysis. This method can handle the standardization problem of nuclear power text data and optimize the model to improve the accuracy, but it only focuses on defect cause analysis and has deficiencies in the full value chain traceability, not covering the traceability of the entire process from design to operation and maintenance. In the patent CN117973512A "A Method for Tracing Quality Defects in Engineering Excavation Based on Knowledge Graph Reasoning", a knowledge graph is constructed by integrating multi-source data and expert knowledge, and a graph database is used for storage and query. Combining knowledge graph reasoning and probability judgment to identify the root causes of quality defects and provide solutions. This method has advantages in knowledge integration and reasoning, but when applied to the nuclear power field, it needs to be adjusted according to the particularity of nuclear power equipment, such as the high sensitivity of nuclear power data and complex technical processes. These factors have not been fully considered at present.
[0007] In summary, the existing methods have certain limitations in solving the problems of nuclear power quality defect source analysis and full value chain traceability. To meet the requirements of ensuring nuclear power quality and safety and improving economic benefits, the present invention aims to combine various intelligent technologies to construct a comprehensive and efficient method for analyzing the sources of quality defects in nuclear power units and full value chain traceability, realizing rapid and accurate traceability of quality defects in the entire process from design, manufacturing to operation and maintenance, improving the efficiency and accuracy of quality traceability, and ensuring the safe and reliable operation of nuclear power units. Summary of the Invention
[0008] The present invention provides a method for analyzing the sources of quality defects in nuclear power units and full value chain traceability, which solves the problems such as difficult calibration of defect sources and low traceability efficiency in the quality control of nuclear power equipment.
[0009] To solve the above problems, the technical solution adopted by the invention is: A method for analyzing the sources of quality defects in nuclear power units and full value chain traceability includes the following steps: S01 Multimodal Data Feature Extraction and Data Resource Pool Construction: Collect structured quality data, semi-structured log files, unstructured text reports, and sensor time-series data throughout the entire value chain of nuclear power equipment design, manufacturing, construction, and commissioning. Construct a multimodal dataset, use the BIO annotation method to mark defect entities in unstructured text, and combine the BERT bidirectional encoding model to fuse semantic, positional, and paragraph features to generate dynamic context vector representations. Extract key semantic features through the BiGRU-Attention network, use bidirectional gated recurrent units to capture text context dependencies, introduce a CRF conditional random field and combine it with the Viterbi algorithm to perform sequence annotation on equipment fault entities and output the optimal label path. Store the information processed by multimodality in a hybrid architecture of a graph database and a relational database to construct a data resource pool. S02 Multi-Dimensional Defect Cause Analysis and Sample Balance Processing: Execute a dynamic pruning strategy on the normalized quality data. By compressing infrequent item sets, pre-pruning single-element low-frequency item sets, and introducing lift to quantify the positive and negative correlations of association rules, use the chi-square test to calculate the correlation between features and defects. For unbalanced datasets, use the non-dominated sorting genetic algorithm to generate differentiated training subsets. S03 Dynamic Knowledge Graph Construction and Intelligent Tracing: Knowledge Graph Modeling: Design multiple types of nodes including quality problem manifestations, equipment objects, causes, stages, and responsible parties, define composition relationships, causal relationships, and stage attribution relationships, and assign thresholds to the nodes. After inputting a quality event, match the graph nodes through entity linking technology, combine regular matching and probabilistic path search to trace the defect source backward, and output the tracing chain in descending order of probability.
[0010] The principle and advantages of this solution are as follows: For heterogeneous data generated throughout the entire value chain of nuclear power equipment, such as structured quality data, semi-structured logs, unstructured text, and sensor time-series data, use the BIO annotation method to accurately mark defect entities in the text, such as equipment names and fault phenomena. Combine the BERT bidirectional encoding model to fuse semantic, positional, and paragraph features, and convert the text into dynamic context vectors to solve the problem that traditional word vectors cannot represent polysemy. Capture text context dependencies through the BiGRU-Attention network, combine the attention mechanism to screen key semantic features, and then use the CRF conditional random field combined with the Viterbi algorithm for sequence annotation to output the optimal label path of equipment fault entities, realizing the accurate extraction of structured entity information from unstructured text. Finally, uniformly store the multimodal data in a hybrid architecture of a graph database and a relational database to construct a data resource pool containing entity vectors, time-series features, and association relationships, providing multi-dimensional data support for subsequent analysis.
[0011] In the multi-dimensional defect cause analysis and sample balance processing stage, the dynamic pruning strategy is first implemented on the normalized data: by compressing non-frequent item sets and pre-pruning single-element low-frequency item sets, redundant calculations are reduced; the positive and negative correlations of the quantitative association rules are improved, strong association rules are screened, and the defects of the traditional Apriori algorithm that cannot distinguish between positive and negative correlations are solved. The chi-square test is used to calculate the correlation between features and defects, and key influencing factors are screened. Combined with the random forest algorithm, a visual tree diagram is constructed to intuitively display the defect propagation path. In view of the sample imbalance problem of defects in nuclear power data, a non-dominated sorting genetic algorithm is used to generate differentiated training subsets, optimize the random forest classification boundary, and improve the recognition accuracy of minority defects.
[0012] In the stage of dynamic knowledge graph construction and intelligent tracing, multi-type nodes including quality problem manifestations, equipment objects, causes, stages, and responsible parties are designed, and multiple edge types such as "composition relationship", "causal relationship", and "stage attribution relationship" are defined. Attributes such as quality standard thresholds and historical occurrence probabilities are given to nodes to build a dynamic knowledge graph. When a quality event is input, the graph nodes are matched through entity linking technology, and the source of the defect is traced back in reverse by combining regular matching and probabilistic path search. High-probability traceability chains are output preferentially. At the same time, the stage to which the defect belongs is locked through time series correlation analysis, and the responsible party nodes are associated to form a complete traceability report.
[0013] Further, in S01 BIO Marked with B, I, O The beginning, middle or end of the entity, and non-entity are marked respectively, which is used to annotate the nuclear power defect information text.
[0014] Furthermore, the BERT text vector conversion in S01 is to transform the annotated text into BERT Encoding, fusing semantic, position and paragraph features to generate a comprehensive feature vector , and then Transformer The encoder obtains the corresponding vector of the text , its simplified formula is:
[0015] Given a nuclear power quality text description sentence ,through BERT After processing, we get the corresponding vector of the text description sentence ,in Representative quality text description sentence i Words, Representative i The word vector of n words, n is the number of words in the quality text description sentence.
[0016] Further, in S01 BiGRU-Attention of BiGRUThis part concatenates the forward and backward outputs of the GRU network to extract text semantic features. The GRU calculation process is as follows:
[0017] Among them, is t the input word vector at time step represents the hidden state vector, represents the time step the hidden state vector, at time step t the candidate hidden state vector, input to the update gate weight matrix, represents the hidden state to the update gate weight matrix, input to the update gate weight matrix, represents the hidden state to the reset gate weight matrix, represents the input to the candidate hidden state weight matrix, reset gate bias vector, is the update gate bias vector, represents the update gate, controlling the information to enter the next state; represents the reset gate, determining the selection of information; * represents the Hadamard product, and GRU the forward and backward concatenation of the network output results in BiGRU a single-character feature output. The formula is:
[0018] Among them, represents the feature output at the forward time step GRU , represents the feature output at the backward time step. Merging all the character information at each time step results in a text semantic feature matrix. Attention The mechanism filters the features related to the equipment quality defects according to the formula A
[0019] and outputs the key semantic feature matrix. Among them Query vector ( Q) , Key vector(K) , Value The vector (V) is the new matrix after the projection change of the feature matrix, is the transpose matrix of K, is the dimension of the input vector, is the activation function.
[0020] Furthermore, in the S02, based on the improved Apriori association mining of quality defects in nuclear power equipment, the original data is first normalized to obtain the cause and phenomenon representations. Apriori The algorithm generates frequent item sets by scanning the transaction database, self-joining operations, and threshold screening, and then obtains frequent 2-item sets through support degree calculation and screening. It recursively continues until the candidate 4-item set is an empty set. The improvement measures include: compressing the database and deleting infrequent item sets; pruning in advance for item sets with the number of single-element occurrences less than k times of k ; the lift formula is where , the probability that event X and event Y occur simultaneously, is the probability that event X occurs, is the probability that event Y occurs, represents confidence, represents support degree. When lift is less than 1, the transaction and Y are negatively correlated, that is, the occurrence of Y will cause lift not to occur; when lift is approximately 1, the transaction and are independent of each other and have nothing to do with each other; when Y is greater than 1, the transaction and
[0022] are positively correlated, and the larger the L value, the higher the correlation between the two, and the higher the probability of their co-occurrence.
[0023] where is the actual frequency, is the theoretical frequency, The larger the k value, the stronger the correlation between the two variables. Select the top
[0024] Further, in the S02, for the analysis of the causes of nuclear power equipment defects under sample imbalance based on NSGALL-RF , the use of NSGAII generates training subsets with large differences and high classification accuracy. NSGAII The algorithm first randomly generates an initial population, and generates offspring populations through selection, crossover, and mutation, repeating until the stopping criterion is met.
[0025] Further, the S03 collects structured, semi-structured, and unstructured data from the nuclear power intelligent construction platform, construction management platform, and Internet of Things platform system. The node design includes the manifestations of quality problems, occurrence objects, and causes; the relationship design includes multiple relationships composed of equipment and responsible parties, and equipment and sub-components; the attribute design includes quality standards and node probabilities.
[0026] Further, in the S03, during quality traceability, entity, relationship, and phenomenon information are identified, a standardized query-driven knowledge graph search is generated, and the traceability is performed by combining an inference model with a regular matching method. The path probability is calculated using the node probability attribute, and the traceability results are output in the order of probability. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a flowchart of the method of the present invention; Figure 2 is the representation of the dataset in NSGAII; Figure 3 is a detailed flowchart of NSGAII-RF fault classification. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] Example 1, as Figure 1 shown, a method for analyzing the sources of quality defects of nuclear power units and full value chain traceability includes the following steps: S01 Multi-modal data feature extraction and data resource pool construction: Collect structured quality data, semi-structured log files, unstructured text reports, and sensor time series data of nuclear power equipment in the entire value chain of design, manufacturing, construction, and commissioning, construct a multi-modal dataset, use the BIO annotation method to mark defect entities in unstructured text, and combine the BERT bidirectional encoding model to fuse semantic, position, and paragraph features to generate a dynamic context vector representation; extract key semantic features through the BiGRU-Attention network, use the bidirectional gated recurrent unit to capture the context dependence of the text, introduce the CRF conditional random field and combine it with the Viterbi algorithm to perform sequence annotation on equipment fault entities, and output the optimal label path; store the information processed by multi-modal in a hybrid architecture of a graph database and a relational database to construct a data resource pool; S02 Multi-dimensional Defect Cause Analysis and Sample Balancing Processing: Execute a dynamic pruning strategy on the standardized quality data. By compressing infrequent item sets, pre-pruning single-element low-frequency item sets, and introducing Lift to quantify the positive and negative correlations of association rules, calculate the correlation between features and defects using chi-square test. For unbalanced data sets, use the non-dominated sorting genetic algorithm to generate differentiated training subsets. S03 Dynamic Knowledge Graph Construction and Intelligent Tracing: Knowledge Graph Modeling: Design multiple types of nodes including quality problem manifestations, equipment objects, causes, stages, and responsible parties, define composition relationships, causal relationships, and stage attribution relationships, and assign thresholds to the nodes. After inputting quality events, match the graph nodes through entity linking technology, and combine regular matching and probabilistic path search to trace the defect source backward, and output the tracing chain in descending order of probability.
[0029] For heterogeneous data generated in the entire value chain of nuclear power equipment, such as structured quality data, semi-structured logs, unstructured text, and sensor time-series data, use the BIO annotation method to accurately mark defect entities in the text, such as equipment names and fault phenomena. Use the BERT bidirectional encoding model to fuse semantic, position, and paragraph features, and convert the text into dynamic context vectors to solve the problem that traditional word vectors cannot represent polysemy. Capture the text context dependence relationship through the BiGRU-Attention network, screen key semantic features using the attention mechanism, and then use the CRF conditional random field combined with the Viterbi algorithm for sequence annotation to output the optimal label path of the equipment failure entity, realizing the accurate extraction of structured entity information from unstructured text. Finally, uniformly store the multi-modal data in a hybrid architecture of a graph database and a relational database, and construct a data resource pool containing entity vectors, time-series features, and association relationships to provide multi-dimensional data support for subsequent analysis.
[0030] In the stage of multi-dimensional defect cause analysis and sample balancing processing, first execute a dynamic pruning strategy on the standardized data: By compressing infrequent item sets, pre-pruning single-element low-frequency item sets, reduce redundant calculations; introduce Lift to quantify the positive and negative correlations of association rules, screen strong association rules, solve the defect that the traditional Apriori algorithm cannot distinguish positive and negative associations, calculate the correlation between features and defects using chi-square test, screen key influencing factors, and combine the random forest algorithm to construct a visual tree diagram to intuitively display the defect propagation path. For the problem of sample imbalance of defects in nuclear power data, use the non-dominated sorting genetic algorithm to generate differentiated training subsets, optimize the random forest classification boundary, and improve the recognition accuracy of minority-class defects.
[0031] In the dynamic knowledge graph construction and intelligent traceability stage, design multiple types of nodes including quality problem manifestations, equipment objects, causes, stages, and responsible parties, define multiple edge types such as "composition relationship", "causal relationship", "stage attribution relationship", etc., and assign attributes such as quality standard thresholds and historical occurrence probabilities to the nodes to construct a dynamic knowledge graph. When a quality event is input, match the graph nodes through entity linking technology, combine regular matching and probabilistic path search to trace back the defect source reversely, preferentially output high-probability traceability chains, and at the same time lock the stage to which the defect belongs through temporal correlation analysis and associate the responsible party nodes to form a complete traceability report.
[0032] In the above S01 BIO Annotate with B 、 I 、 O to mark the beginning, middle or end, and non-entity of the entity respectively, for annotating the nuclear power defect information text. The information text in the nuclear power field contains a large number of professional terms and complex technical descriptions, with a complex structure. BIO The annotation is through clear B 、 I 、 O markings, which can accurately define the starting, middle and ending positions of each entity in the text. Taking "On xx date of xx month of xx year, Wang XX found that there was a spark in the steam turbine of Unit xx, resulting in the generator speed exceeding xxx value, with xxx risk." as an example, with the help of BIO annotation, it can be clearly marked as " x / B-date xx year xx month xx / I-date day / I-date” (Time entity), "Wang / B-person Mou / I-person Mou / I-person” (Person entity), "x / B-crew x machine / I-crew group / I-cerw” (unit entity), etc. This accurate annotation effectively avoids the ambiguity or error of entity recognition caused by complex text expressions, ensures that when extracting quality defect-related information from the text later, the key entities can be accurately locked, greatly improving the accuracy of information extraction, and providing a reliable data basis for defect source analysis and traceability; when extracting information related to quality defects from the text, BIO the annotation can help the subsequent processing model quickly and accurately locate the key information. For example, when performing text vector transformation and key semantic feature extraction, the model can focus on the parts marked as entities according to the annotation results, effectively excluding the interference of non-entity information, and more accurately capturing the text content directly related to the quality defect, thereby improving the accuracy and efficiency of the entire information extraction process, and making the extracted information more reliable and usable.
[0033] In the above S01 BERT Text vector transformation is to encode the annotated text through BERT to fuse semantic, position and paragraph features to generate a comprehensive feature vector , and then obtain the corresponding vector of the text through Transformer encoder , and its simplified formula is:
[0034] Given the nuclear power quality text description sentence , after BERT processing, the corresponding vector of the text description sentence is obtained , where represents the i th character of the quality text description sentence, represents the i th character vector, and n is the number of characters in the quality text description sentence.
[0035] In the above S01 BiGRU-Attention of BiGRU part, the forward and backward outputs of the network are concatenated to extract the text semantic features. The GRU calculation process is as follows: GRU Among them,
[0036] where, is the input character vector at t time step; represents the hidden state vector, represents the hidden state vector at time step , at time step t is the candidate hidden state vector, input to the update gate weight matrix, represents the weight matrix from the hidden state to the update gate weight matrix, input to the update gate weight matrix, represents the weight matrix from the hidden state to the reset gate weight matrix, represents the weight matrix from the input to the candidate hidden state weight matrix, reset gate bias vector, is the bias vector of the update gate , represents the update gate, which controls the information entering the next state; represents the reset gate, which determines the information selection; * represents the Hadamard product. Concatenating the forward and backward outputs of the GRU network gives BiGRU a single character feature output, and the formula is:
[0037] where, Represents the forward moment GRU of the characteristic output Represents the characteristic output of the reverse moment. Combine all the text information of the time steps to obtain the text semantic feature matrix Attention The mechanism is based on the formula A
[0038] Screen the features related to the equipment quality defects and output the key semantic feature matrix, where Query Vector ( Q) , Key Vector (K) 、 Value Vector (V) Is the new matrix after the projection change of the feature matrix Is the transpose matrix of K Is the dimension of the input vector Is the activation function
[0039] The above scheme BERT The text vector transformation generates a comprehensive feature vector by fusing semantic, position and paragraph features, which can comprehensively capture various key information in the nuclear power quality text description sentence. The semantic features ensure that the model understands the actual meaning of the text, the position features record the position information of each word in the sentence, and the paragraph features consider the context environment of the text where it is located. This makes the transformed vector not only contain the meaning of the words themselves, but also incorporate their context information in the sentence and paragraph, thus more accurately representing the connotation of the text. For example, in the text describing the nuclear power equipment failure, not only can the name of the failed equipment be identified, but also the relevant scenarios and conditions of the failure can be accurately understood by combining its position and paragraph background in the text Attention The mechanism can accurately screen out the key information from the text semantic feature matrix. When processing the nuclear power quality text, the text may contain a large amount of information irrelevant to the quality defects Attention The mechanism can focus on the features directly related to the equipment quality defects, ignore the secondary information and highlight the key points. For example, in the text describing various operating parameters of the nuclear power equipment Attention The mechanism can accurately identify the parameter change information related to the current quality defect, improving the pertinence and accuracy of the analysis
[0040] In the above S02, based on the improved Apriori Association mining of nuclear power equipment quality defects, first normalize the original data to obtain the cause and phenomenon characterization AprioriThe algorithm generates frequent item sets by scanning the transaction database, self-joining operations, and threshold screening, and then obtains frequent 2-item sets through support calculation screening. It recursively continues until the candidate 4-item set is an empty set. The improvement measures include: compressing the database and deleting infrequent item sets; pre-pruning, pre-pruning item sets with the number of single-element occurrences less than k times k ; the lift formula is where , the probability of the event X and the event Y occurring simultaneously, is the probability of the event X occurring, is the probability of the event Y occurring, represents confidence, represents support. When lift is less than 1, the transaction and Y are negatively correlated, that is, the occurrence of one leads to the non-occurrence of Y ; when lift is approximately 1, the transaction and are independent of each other and have nothing to do with each other; when lift is greater than 1, the transaction and Y are positively correlated, and L the larger the value, the higher the correlation between the two, and the higher the probability of their co-occurrence.
[0042] The above scheme first normalizes the original data to obtain the cause and phenomenon representations, which can make the complex and diverse nuclear power quality data more organized and standardized, provide a clear data basis for subsequent mining work, and facilitate the accurate analysis of the internal relationships between data.
[0043] Apriori The algorithm uses the method of generating frequent item sets by scanning the transaction database, self-joining operations, and threshold screening to systematically and comprehensively mine the potential associations in the data. By continuously recursing until the candidate 4-item set is an empty set, it can gradually explore the frequent relationships between different item sets from simple to complex, mine multi-level and multi-dimensional association information, and provide a comprehensive perspective for analyzing the complex relationship between quality defect phenomena and causes.
[0044] And in the improvement measures, compressing the database and deleting infrequent item sets can effectively reduce the data processing volume and storage space occupation. In the case of a large amount of data in the nuclear power field, reducing the processing of unnecessary infrequent item sets can greatly improve the algorithm operation efficiency, avoid redundant calculations, and make the mining process more efficient and fast. Pre-pruning, for item sets with the number of single-element occurrences less than k times kThe item set is pre-pruned, further optimizing the search space, reducing the number of candidate item sets, avoiding a large amount of calculations for invalid item sets, accelerating the joining speed, saving computing resources and time, and improving the mining efficiency.
[0045] The concept of lift is introduced and the relevance of transactions is judged according to its formula, providing a quantitative index for analyzing the relationship between quality defect phenomena and causes. It can help engineers clearly understand the degree of association between different factors, accurately judge which factors are the key factors truly affecting quality defects, and which factors are independent of each other or have a negative correlation. For example, when analyzing a certain fault of nuclear power equipment, the lift can be used to determine which component faults are highly correlated with this fault and which seemingly relevant factors actually have little association, providing a scientific basis for on-site judgment and subsequent precise formulation of solutions, assisting engineers to analyze quality problems more efficiently, and also providing a solid theoretical foundation and reliable analysis basis for subsequent algorithm classification model analysis. In the above S02, the chi-square value between a single feature and the target variable is calculated using the chi-square test, and the formula is
[0046] where is the actual frequency, is the theoretical frequency, The larger the value indicates the stronger the correlation between the two variables. Select the top k important features according to the chi-square value ranking. Samples are drawn through the bootstrap sampling method to form multiple data subsets, and decision trees are trained to construct a random forest. The sample category is discriminated by the majority voting strategy, and the optimal tree is selected from the random forest through visualization to form a visualized tree diagram for defect cause analysis; calculate the chi-square value using the chi-square test and select the top kAn important feature is that it can accurately screen out the key features with the strongest correlation with the target variable (such as quality defects) from a large number of complex nuclear power quality data features. In the nuclear power field, the data dimension is high and the information is complex. Not all features are equally important for quality defect analysis. Through chi-square test, the degree of association between each feature and quality defects can be quantified, avoiding interference from irrelevant or weakly correlated features, ensuring that subsequent analysis focuses on the factors that truly affect quality defects, and improving the pertinence and accuracy of the analysis. For example, when analyzing the faults of the nuclear power unit cooling system, it can accurately identify key influencing factors such as coolant flow and temperature, ignore some secondary environmental factors, and improve the fault diagnosis efficiency; the self-sampling method extracts samples to form multiple data subsets to train decision trees to construct a random forest. This method effectively reduces the dependence of the model on a single data sample and improves the stability of the model. Since each sampling is random and with replacement, the data features included in different data subsets are different, and the trained decision trees also have their own characteristics. The random forest formed by integrating many decision trees synthesizes the judgment results of multiple decision trees, reduces the errors caused by individual sample biases, and makes the model prediction more robust and reliable. When facing nuclear power quality data under different working conditions, the random forest can more stably output accurate analysis results and reduce the risk of misjudgment.
[0047] The analysis of the causes of nuclear power equipment defects under sample imbalance in S02 is based on NSGALL-RF to generate training subsets with large differences and high classification accuracy. NSGAII The algorithm first randomly generates an initial population, and through selection, crossover, and mutation, generates offspring populations, repeating until the stopping criterion is met. There is a sample imbalance in nuclear power equipment quality data, and minority class samples, such as information on rare quality defect types, are easily ignored in traditional analysis. NSGAII By generating training subsets with large differences and high classification accuracy, the information in minority class samples can be fully mined, avoiding the omission of important data. The process of randomly generating an initial population and generating offspring populations through selection, crossover, and mutation enables the algorithm to explore different regions of the sample space, increase the learning opportunities for the characteristics of various samples, improve the classification ability for minority class samples, and thus more comprehensively and accurately analyze the causes of nuclear power equipment defects in the case of sample imbalance; NSGAII The generated training subsets have large differences, and the decision trees trained based on these subsets are also different. The finally constructed random forest contains diverse decision trees. This diversity enables the random forest to analyze and judge from multiple angles when facing different nuclear power quality data, enhancing the generalization ability of the model. Even when encountering new or incompletely covered quality defect situations, the model can give relatively reasonable analysis results relying on the diverse combination of decision trees, improving the adaptability of the model to the complex and changeable nuclear power equipment quality data. NSGAII
[0048] The S03 collects structured, semi-structured, and unstructured data from the nuclear power intelligent construction platform, construction management platform, and Internet of Things platform system. The node design includes the manifestations of quality problems, the objects involved, and the causes. The relationship design contains multiple relationships composed of equipment and responsible parties, and equipment and sub-components. The attribute design includes quality standards and node probabilities. By collecting structured, semi-structured, and unstructured data from the nuclear power intelligent construction platform, construction management platform, and Internet of Things platform system, rich information throughout the entire life cycle of nuclear power units can be obtained. Structured data such as construction progress and equipment parameters provides accurate quantitative information and can be used for direct analysis and comparison. Semi-structured data such as construction logs and quality inspection reports contains detailed process records and helps to understand the background of quality problems. Unstructured data such as on-site pictures and operation manuals supplements information that is difficult to express in a fixed format and can uncover potential quality influencing factors. The multi-platform data fusion can comprehensively restore the scenarios of quality events, avoid information loss, and provide sufficient data support for subsequent traceability and analysis.
[0049] In the S03, during quality traceability, entity, relationship, and phenomenon information is identified to generate a standardized query-driven knowledge graph search. The traceability is carried out by combining an inference model with regular matching. The path probability is calculated using the node probability attribute, and the traceability results are output in probability order. By determining the key entities in a quality event, such as the equipment and components involved, clarifying the relationships between entities, and understanding the specific phenomena of quality problems, the core elements of quality problems can be quickly locked. Taking the leakage of the flange of the main steam filter as an example, entities such as "main steam filter", "flange bolts", and "flange gaskets" are identified, as well as the compositional relationship, causal relationship between them, and the phenomenon of "leakage", which helps to directly cut into the core of the problem, avoid blind investigation among a large amount of irrelevant information, accurately locate the key links that may cause the problem, and improve the accuracy and pertinence of traceability. Generating a standardized query-driven knowledge graph search makes the traceability process more efficient and orderly. The standardized query can quickly match relevant information in the knowledge graph, avoiding aimless traversal and greatly shortening the search time. The knowledge graph stores and organizes data in a structured manner, can quickly locate the nodes and relationships related to quality problems, and improve the speed of information acquisition. In the face of complex nuclear power systems and a large amount of quality data, this efficient search method can quickly screen out useful information, reduce the time cost of traceability, respond to quality problems in a timely manner, and ensure the stable operation of nuclear power units.
[0050] Embodiment 2 S1: Utilize the large amount of quality event data and quality characteristic phrases accumulated in the entire value chain process of nuclear power equipment design, manufacturing, construction, and commissioning to further extract key information from the quality data, and extract relevant information containing quality defects to form a data resource pool. When similar quality problems occur in the future, based on the constructed data resource pool, comparative analysis can be carried out to quickly locate and trace the defect source.
[0051] S2: During the entire life cycle of nuclear power equipment construction and operation, various problems will occur at each stage. Three methods for analyzing the causes of quality defects are adopted, namely, the correlation mining of nuclear power equipment quality defects based on improvement Apriori the cause analysis of nuclear power equipment defects based on feature engineering, and the cause analysis of nuclear power equipment defects under sample imbalance based on NSGALL-RF .
[0052] S3: For the construction and application of the knowledge graph, construct a general design method for the knowledge graph in the field of nuclear power equipment quality traceability, and use the knowledge graph as the underlying layer to design the product quality anomaly traceability process. Apply the knowledge graph technology to the entire value chain traceability field of nuclear power equipment, including the design of the quality traceability knowledge system and the application of downstream tasks.
[0053] Specifically, S1 is as follows: S11: Since there are few categories of nuclear power defect information entity extraction, BIO is used for annotation. BIO annotation uses B to represent the beginning of an entity, I to represent the middle or end of an entity, and O to represent a non-entity. For example, for the annotation of "On xx / xx / xxxx, Wang XX discovered that the steam turbine of Unit xx generated sparks, resulting in the generator speed exceeding xxx value, with xxx risk existing.", the example is as follows: x / B-date xx year xx month xx / I-date day / I-date, O Wang / B-person Mou / I-person Mou / I-person found / O x / B-crew x machine / I-crew group / I-cerw steam / B-Equiment turbine / I-Equiment generated / O fire / B- phenomenon flower / I-phenomenon, resulting / O in / B-Equiment motor / I-Equiment speed exceeding xxx value , O xxx risk exists.
[0054] S12: To construct a model for extracting nuclear power equipment quality entities, it is first necessary to convert the annotated key information text into a text vector. The vector representation of text is essentially to represent words, sentences, articles, entire books, etc. with a vector, and there are many different representation learning methods. We use BERT method to perform text vectorization. BERT is a representation model for encoding text, which can convert a piece of text into a set of vector representations, and the vectors integrate the global semantic information of the text.
[0055] After fusing semantic features, position features, and paragraph features, a comprehensive feature vector is obtained, and then the fused feature vector is input into the Transformer encoder for encoding, and finally the vector representation corresponding to the text is output , and its formula is shown as follows:
[0056] Given the text description sentence of nuclear power quality , after BERT , the corresponding vector of the text description sentence is obtained , where represents the i th character of the quality text description sentence, represents the word vector of the i th character.
[0057] S121: After the conversion of the text vector, extract the key semantic features related to the faulty equipment and related descriptions in the text from BiGRU-Attention to narrow the interpretation range. Then, through the Attention mechanism, screen out the feature information related to the equipment quality defect from the extracted semantic information and ignore the secondary information, and output the key semantic feature matrix. BiGRU is the forward and reverse splicing output by the GRU network, and the calculation process is shown as follows: GRU
[0058] Among them, t is the input word vector at the moment; represents the hidden state vector, represents the hidden state vector at the time step , t At the time step the candidate hidden state vector, input to the update gate weight matrix, represents the weight matrix from the hidden state to the update gate weight matrix, input to the update gate weight matrix, represents the weight matrix from the hidden state to the reset gate weight matrix, represents the input to the candidate hidden state weight matrix, reset gate bias vector, is the bias vector of the update gate , represents the update gate, which controls the information to enter the next state; Indicates the reset gate, which determines whether information is retained or discarded. The update gate and the reset gate jointly determine the hidden state output, and * represents the Hadamard product. The output of the GRU network is concatenated in the forward and reverse directions to obtain BiGRU a single character to extract the corresponding feature output. The formula is as follows:
[0059] where represents the feature output at the forward time step GRU , represents the feature output at the reverse time step. The information of all time step characters is merged to obtain the extracted text semantic feature matrix. The formula is as follows: H
[0060] S122: After obtaining the text semantic feature matrix H, each word in the sentence is compared with every other word to calculate the similarity, obtaining the mutual relationship between words, extracting the information related to the faulty device, and weakening the secondary text information to obtain the key semantic information matrix. The calculation formula is as follows:
[0061] where Query vector (Q) , Key vector (K) , Value vector (V) are all new matrices after the projection transformation of the feature matrix, representing three different matrices formed after the feature matrix is mapped, is the dimension of the input vector.
[0062] S123: The quality text vector is BiGRU-Attention extracted to obtain the key semantic feature vector of the quality text H
[0063] S124: After extracting the key semantic feature vector of the quality text, CRF is used to calculate the faulty device.
[0064] S125: The optimal path planning method solves the dependency relationship between adjacent labels and calculates the transition probability between adjacent labels. By increasing the transition probability within the entity label, the best entity label sequence is predicted. CRF The calculation formula is as follows:
[0065] where:
[0066]
[0067]
[0068] Among them:
[0069] where represents the input data sequence to be observed, represents a state sequence, and are feature functions, and are the corresponding weights, represents the conditional probability of the output sequence under the given input sequence predicted, that is, under the given quality text, by observing the transition relationship between adjacent words, calculating the transition probability between adjacent words, and predicting the probability that the word sequence belongs to the faulty equipment entity.
[0070] S126: Optimize the objective function through the Viterbi algorithm to obtain the best prediction label:
[0071] wherein, represents the best prediction label obtained after optimizing the objective function through the Viterbi algorithm. According to the above formula, the optimal prediction sequence is calculated, and the label probability corresponding to each word is predicted. The word combination with the maximum label probability is the extracted faulty equipment.
[0072] S13: The nuclear power quality traceability relationship extraction model is the same as the nuclear power equipment quality entity extraction model. First, through BERT, the labeled text is transformed into text vectors, and then through BiGRU-Attention the key semantic features are extracted .
[0073] S131: Output the key semantic features as input and use to perform fine-tuning.
[0074] S132: Set the MLP output variable as a 4D vector, representing the probabilities of the four stages of design, procurement, construction, and commissioning of the nuclear power business process. Take the highest output probability value as the stage to which the final determined quality defect belongs, forming a classification problem based on .
[0075] S133: MLP Compress and non-linearly fuse the extracted key semantic feature information, comprehensively consider the contribution of each character to the output, and the weight size of each layer represents the proportion of the influence of each character feature on the output. The specific calculation process is as follows: The quality text description sentence passes through BERT-BiGRU-Attention After calculation, a key semantic feature vector is obtained ATT, ATT Dimensionality reduction processing makes each character correspond to a numerical value, and a text vector after dimensionality reduction is obtained , and use as MLP input, MLP For perform non-linear fusion, representing the contribution of each character to the output. The calculation formula is as follows: , where the output is a four-dimensional vector, representing the probabilities corresponding to the four stages. The serial number corresponding to the maximum value among the four dimensions is taken as the classification result, is MLP parameters for training.
[0076] S134: Extract the feature information of the quality text through the above information annotation strategy, text vector conversion technology and feature extraction technology, obtain the key information corresponding to the quality defect, and use the database to structurally store the quality text-defect information to form a data resource pool to provide support for subsequent defect source calibration and traceability.
[0077] The specific content of the said S2 is as follows: S21: A large amount of defect text data is recorded in the nuclear power quality event form, including information such as defect phenomena and cause analysis. Introduce the association relationship mining method to automatically obtain the association relationship between the nuclear power equipment quality defect and the cause in the event form, providing a basis for constructing a knowledge-based defect classification model.
[0078] S211: According to the characteristics of complex and redundant data information in the nuclear power quality event form, normalize the original data to obtain cause and phenomenon representations.
[0079] S212: Through improving Apriori algorithm, mine the association relationships between defect phenomena and causes, and between causes and causes.
[0080] S213: Use Apriori algorithm to mine the association rules of nuclear power equipment defect causes. Its core idea is to find the frequent item sets in the data through two-step mining and multiple recursive iterations, generate k + 1 item sets through k-item set iteration until no maximum item set is generated, and finally mine out the transaction with the largest association relationship in the data set.
[0081] S214: Describe the generation steps of frequent item sets between causes and phenomena for the normalized quality event list data, and set the minimum support and confidence thresholds.
[0082] S215: Scan the transaction database to obtain candidate 1-item sets C1: {A}, {B}, {C}, {D}, {E} , where if the support of an item set {D} is less than the set threshold and the support of the remaining item sets is greater than the set support threshold, then the frequent 1-item set is {A}, {B}, {C}, {E} .
[0083] S216: Perform a self-join operation to obtain candidate 2-item sets : {A, B}, {A, C}, {A, E}, {B, C}, {B, E}, {C, E} , where all subsets of an item set are included in . After calculating the support, the item sets that meet the threshold are {A, C}, {B, C}, {B, E} Therefore, they are classified as frequent 2-item sets .
[0084] S217: Similarly, perform a recursive self-join operation and filter through the threshold to obtain frequent 3-item sets which are {B, C, E} .
[0085] S218: Finally, perform a self-join by to generate an empty set of candidate 4-item sets, and the algorithm execution ends. Calculate the confidence of the frequent item sets and compare it with the set threshold to obtain the association rules.
[0086] S219: Based on the generation rules and screening criteria of frequent item sets, the following improvements are made: (1) To address the problem that the search process for generating frequent item sets wastes a large amount of time and database space, a method of compressing the database is proposed to reduce the use of data space. In each scan, the set determined to be a non-frequent item set is directly deleted from the database to avoid repeated screening and searching when traversing the database next time, reducing the waste of space and time.
[0087] (2) Use the properties of frequent item sets for screening and adopt a method of pruning in advance to delete unnecessary item sets, that is, calculate the number of occurrences of single elements in a certain k item set. If the number of occurrences is less than k times, then the item set containing this element does not belong to the frequent k+1For the item set, perform a pre-pruning operation on the item set containing this element. Only search for specific minimum item transaction branch nodes instead of repeatedly searching the entire database, which can speed up the connection speed and reduce the number of candidate item sets generated.
[0088] (3) To reduce the generation, analysis, and interference of redundant rules, the concept of lift is introduced to assist in the determination of association rules. For a certain frequent k item set, the lift is defined as:
[0089] Among them, , the probability that event X and event Y occur simultaneously, is the probability that event X occurs, is the probability that event Y occurs, represents confidence, represents support. From the calculation formula, when lift is less than 1, the transaction and Y are negatively correlated, that is, the occurrence of one leads to the non-occurrence of Y ; when lift is approximately 1, the transaction and are independent of each other and have nothing to do with each other; when lift is greater than 1, the transaction and Y are positively correlated, and L the larger the value, the higher the correlation between the two, and the higher the probability of their co-occurrence.
[0090] By introducing the concept of lift, the correlation degree between the defect phenomenon and the cause, and between the cause and the cause can be obtained. The larger the lift value, the higher the correlation between the two. Therefore, it can better understand the correlation and influence degree between the defect phenomenon and the cause in the quality event list, which is helpful for the defect association relationship of nuclear power equipment and the on-site judgment of engineers. At the same time, it provides a theoretical basis and analysis basis for the subsequent algorithm classification model analysis.
[0091] S22: After obtaining the fault information, extract the characteristic information of the quality text through the information annotation strategy, text vector conversion technology, and feature extraction technology, obtain the key information corresponding to the quality defect, and structurally store the quality text - defect information in the database to form a data resource pool to provide support for subsequent defect source calibration and traceability.
[0092] S221: After obtaining the extracted information, we need to construct a random forest algorithm model. The random forest uses the bootstrap sampling method to randomly extract the data set, so as to ensure that the probability of each sample being selected is the same. After obtaining the sample set, construct a decision tree. When finding features for splitting at the node, randomly extract a part of the features from the features, find the optimal solution among these features, and then apply it to the node for splitting. In this way, the random forest can avoid overfitting.
[0093] S222: The decision tree constructs a classification model in the form of a tree. Each of its nodes represents an attribute. After division according to the attribute, it enters the leaf nodes of the intuitive nodes until the leaf nodes. Each leaf node represents a certain category. Through this method, classification is ultimately achieved. The finally formed tree can clearly show the formation process from features to target classification and is interpretable. Because of this characteristic, the propagation path of defect sources can be formed, and then the defect causes can be analyzed.
[0094] S223: The chi-square test is used to measure the dispersion of the occurrence distribution of features independent of the category values, and calculate the chi-square value between a single feature map and the target variable. The chi-square value calculation formula is as follows:
[0095] In the formula is the actual frequency, is the theoretical frequency, is the chi-square value. The larger the , the greater the deviation, indicating a stronger correlation between the two variables. Therefore, feature ranking is performed according to the chi-square value. Select the first k important features, and select the k features as the input of the random forest model. Through the bootstrap sampling method, samples are randomly drawn with replacement to form multiple data subsets. When training several decision trees using each drawn data subset, randomly select attributes as the node splitting attributes until no further splitting is possible, and establish a large number of decision trees to form a forest. Using the idea of integration, the strategy of majority voting is used to determine the category to which the sample belongs. Finally, visualize the random forest and select the optimal tree to form a visual tree diagram for defect cause analysis. Steps for defect cause analysis based on feature engineering.
[0096] S23: To solve the drawbacks of losing important sample information, time-varying data feature information in traditional sampling methods and low accuracy of sub-models in the integrated classification method, NSGAII-RF is adopted to solve the fault diagnosis under the problem of sample imbalance. This algorithm uses an evolutionary method to select subsets of features and samples. On the basis of taking into account sampling to retain important sample and feature information, NSGAII is used to optimize both diversity and accuracy simultaneously to generate diverse and high-performance feature subsets and sample subsets, improving the diversity between decision trees while improving accuracy. This method uses the optimal solution set obtained by the multi-objective evolutionary optimization algorithm as different training subsets to train decision trees. In addition, the number of obtained solutions determines the number of decision trees required to build a random forest, avoiding redundant classifiers. The performance of the RF algorithm is improved by creating diverse and high-performance classifiers and determining the optimal number of classifiers.
[0097] S231: By NSGAII generate a training subset with large differences and high classification accuracy, NSGAII and the obtained Pareto solution is the optimized training subset.
[0098] S232: Based on NSGAII the generated training set, train a decision tree and form a random forest.
[0099] In NSGAII, each chromosome represents a training subset and is represented by a binary vector.
[0100] Each chromosome consists of samples and features, which can be regarded as obtaining certain samples and certain features. The following figure describes NSGAII the population and chromosome representations used in Figure 2 as shown. NSGAII The steps of the algorithm are as follows: First, generate an initial population, which is usually randomly generated; generate a child population from the current population through selection, crossover, and mutation, and the child is composed of the current generation and the offspring; repeat the process of generating the child until the stopping criterion is met.
[0101] S233: According to NSGAII the number of generated training sets, determine the number of decision trees, and based on the diversity and good performance of the generated training sets, train diverse and accurate decision trees. Finally, form a random forest.
[0102] Based on NSGAII The essence of improving the random forest is to incorporate the goal of generating a high-diversity and high-accuracy decision tree into NSGAII , and through NSGAII the evolutionary optimization algorithm, generate the corresponding training set to train the decision tree, so as to reduce the randomness when forming the decision tree by random sampling under the imbalanced data set. The overall diagnosis process will be described below, and the detailed flowchart is as Figure 3 shown.
[0103] The specific content of S3 is as follows: S31: According to the quality status evolution map of the nuclear power equipment throughout the life cycle and the characteristics of the quality data resources used for quality traceability, design an intelligent quality traceability process for nuclear power equipment based on the knowledge graph for the quality traceability problem of equipment in the nuclear power field.
[0104] S311: The relevant data for quality traceability can be obtained from the nuclear power intelligent construction platform within the nuclear power company, such as NICE platform, construction management platform, Internet of Things platform and other systems to obtain data sources (including structured data, semi-structured data, and unstructured data).
[0105] S312: Integrate the quality characteristic data, quality state evolution map, and quality plan data of nuclear power equipment with the three-dimensional model, extract the traceability-related knowledge such as quality key characteristic information and quality state transfer, and make a useful supplement to the quality traceability map.
[0106] S313: Node design includes quality problem manifestation nodes, quality problem occurrence object nodes, quality problem cause nodes, nuclear power equipment / sub-component involved stage nodes, nuclear power equipment component nodes, and responsible entity nodes. Among them, the quality problem occurrence object node and the nuclear power equipment / nuclear power equipment sub-component node can coincide. When they coincide, this node has two node labels.
[0107] S314: Relationship design includes: the relationship between equipment / component / involved stage and the responsible party, the composition relationship between equipment and equipment sub-components, the association relationship between quality problem manifestation and quality problem occurrence object, the relationship between quality problems and the causes leading to the occurrence of quality problems, and the flow relationship of equipment / component manufacturing and installation processes.
[0108] S315: Attribute design includes: the quality standards and normal state attributes of equipment / components, the probability attributes of nodes appearing in the total event space, and supplementary descriptions of nodes or relationships, etc.
[0109] S32: When a new nuclear power equipment quality traceability task is generated, through intelligent analysis of the input quality events, identify the entities, entity relationships, and quality problem phenomenon information of the quality events, and generate a standardized query event to drive the knowledge graph search. Locate the nodes, attributes, and relationships corresponding to the quality problem state manifestation in the quality traceability knowledge graph as anchor points, trace back the object where the quality problem occurs along the quality formation path of nuclear power equipment in reverse, and conduct step-by-step comparison and correlation analysis based on relevant historical quality problem occurrence locations, quality problem state manifestations and causes, and equipment normal state information, give the reasoning result of the quality problem cause, and judge the stage to which the quality problem belongs, forming a traceability chain of quality defect generation - phenomenon - object - involved stage.
[0110] For example: For the on-site quality event recorded in a certain quality event form: "The flange of the main steam filter leaks. Reasons: (1) Through on-site inspection, it is found that the bolt torque of the flange of the main steam filter is insufficient, resulting in the leakage phenomenon; (2) The gasket of the steam filter flange is a serrated gasket, which has high installation requirements, and any slight deviation in the gasket itself and the installation process is likely to cause leakage." S321: According to the design principles of the quality traceability knowledge system, the identified named entities are: "main steam filter", "main steam filter flange", "flange bolt", "flange gasket", "serrated gasket". The identified quality problem manifestations are: "air leakage". The object where the quality problem occurs is: "main steam filter flange". The identified causes are: "insufficient torque", "installation deviation". The identified relationships are: The relationships between "main steam filter" and "main steam filter flange", "main steam filter flange" and "flange bolt", "main steam filter flange" and "flange gasket" belong to the composition relationships between equipment and its sub-components. The relationship between "steam leakage" and "main steam filter flange" belongs to the association relationship between the quality problem manifestation and the object where the quality problem occurs. The relationships between "steam leakage" and "insufficient torque", "installation deviation" belong to the relationships between the quality problem and the causes leading to the occurrence of the quality problem. The identified attribute is: The flange gasket is a "serrated gasket".
[0111] S322: According to the nodes, relationships, and attributes located in the knowledge graph, traverse and query along the associated facts in the knowledge graph, and continuously compare whether there are differences between the states of each equipment node and the normal state of the equipment. Calculate the probability of each path according to parameters such as probability attributes, and finally give the quality traceability query results based on the knowledge graph in the order of occurrence probability. The above-mentioned associated facts can be understood as the quality formation path of nuclear power equipment.
[0112] S323: Suppose a quality event of air leakage in the main steam filter flange occurs at the commissioning site of nuclear power equipment. In the input on-site quality event form, only the phenomenon of air leakage in the main steam filter is recorded, but the location and cause of the air leakage are not specifically investigated. Using the quality traceability method for nuclear power equipment based on the knowledge graph, according to the node, relationship, and attribute information extracted from the input information, relevant initial nodes, attributes, and relationships are matched in the knowledge graph. The nodes matched to the knowledge graph from the input information are "main steam filter" and "steam leakage". According to the existing node and relationship graph in the knowledge graph, start traversing and searching along the relationship path from the initial node: S3231: First, there is an intermediate node "Main Steam Filter Flange" between the "Main Steam Filter" and "Steam Leakage" nodes, and "Main Steam Filter Flange" is a sub-component of "Main Steam Filter". Therefore, it can be inferred that the specific component that causes the "Steam Leakage" of the "Main Steam Filter" equipment node is "Main Steam Filter Flange". From the "Gas Leakage" node, continue along the relationship between the quality problem and the cause of the quality problem to reach the "Insufficient Torque" node and the "Installation Deviation" node. From the "Insufficient Torque" node, trace back to the "Flange Bolt" node, and from the "Installation Deviation" node, trace back To the "flange gasket" node, according to the knowledge graph, "flange bolts" and "flange gasket" are both sub-components of the "main steam filter flange", so the traversal of all paths is completed. The probability of each path is calculated according to the probability attribute of the node. Assuming that in the total event space, the number of "insufficient torque" occurs more than the number of "installation deviation", then the query results returned in order of occurrence probability from high to low are: the most likely location causing the main steam filter leakage is the "main steam filter flange", and the possible reasons are: insufficient torque of the flange bolts and installation deviation of the flange gasket.
[0113] S32: In the process of building the knowledge graph, the index information of the equipment is stored in the knowledge graph. For quality events that can obtain specific status information of the equipment on site, these attribute values can be compared when performing quality traceability based on the normal status of the nodes or quality standard attributes stored in the knowledge graph to find abnormal nodes or exclude normal nodes. For example, the normal status of flange bolts is stored in the nuclear power equipment quality traceability knowledge graph as torque greater than or equal to A The engineer recorded the measured torque of the main steam filter flange bolts in the on-site quality incident sheet as B In the result of this extraction of input information, the "Flange Bolt" node contains the attribute "Torque =B In the derivation process based on the knowledge graph, when reaching the "flange bolt" node, the attribute "torque ≥ A" of the "flange bolt" node is compared with the attribute "torque = B" of the "flange bolt" node in the input information. If B ≥ A, it is judged that the "flange bolt" node is not the location where the quality problem causing the main steam filter leakage occurs, and the "insufficient torque" connected to it is not the cause of the "steam leakage". Therefore, the final traceability result is: the most likely location causing the main steam filter leakage is the "main steam filter flange", and the possible cause is: flange gasket installation deviation. We can further index the construction unit when the flange gasket is installed, thereby completing cross-enterprise quality traceability.
[0114] The above are only embodiments of the present invention. Specific structures and common knowledge such as characteristics that are well-known in the art are not described in detail herein. Those of ordinary skill in the art know all the common general technical knowledge in the technical field to which the invention pertains before the filing date or the priority date, can know all the prior art in this field, and have the ability to apply the conventional experimental means before this date. Those of ordinary skill in the art can, under the inspiration given by this application, combine their own abilities to complete and implement this solution. Some typical well-known structures or well-known methods should not become an obstacle for those of ordinary skill in the art to implement this application. It should be noted that for those skilled in the art, without departing from the structure of the present invention, several deformations and improvements can still be made, and these should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent. The protection scope claimed in this application shall be subject to the content of its claims, and the specific implementation manners and the like recorded in the specification can be used to interpret the content of the claims.
Claims
1. A method for analyzing the quality defect sources of a nuclear power unit and tracing the entire value chain, characterized in that, It includes the following steps: S01 Multimodal data feature extraction and data resource pool construction: Collect structured quality data, semi-structured log files, unstructured text reports, and sensor time-series data throughout the entire value chain of nuclear power equipment design, manufacturing, construction, and commissioning. Construct a multimodal dataset, use the BIO annotation method to mark defect entities in unstructured text, and combine the BERT bidirectional encoding model to fuse semantic, location, and paragraph features to generate dynamic context vector representations. Extract key semantic features through the BiGRU-Attention network, use bidirectional gated recurrent units to capture text context dependencies, introduce the CRF conditional random field and combine with the Viterbi algorithm to perform sequence annotation on equipment fault entities, and output the optimal label path. Store the information processed by multimodality in a hybrid architecture of a graph database and a relational database to construct a data resource pool; S02 Multidimensional defect cause analysis and sample balance processing: Execute a dynamic pruning strategy on the normalized quality data, compress infrequent item sets, pre-prune single-element low-frequency item sets, and introduce lift to quantify the positive and negative correlations of association rules. Use the chi-square test to calculate the correlation between features and defects; For unbalanced datasets, use the non-dominated sorting genetic algorithm to generate differentiated training subsets; S03 Dynamic knowledge graph construction and traceability: Knowledge graph modeling: Design multiple types of nodes including quality problem manifestations, equipment objects, causes, stages, and responsible parties, define composition relationships, causal relationships, and stage attribution relationships, and assign thresholds to the nodes. After inputting quality events, match the graph nodes through entity linking technology, combine regular matching and probabilistic path search to trace the defect source backward, and output the traceability chain in descending order of probability.
2. The method for analyzing the quality defect sources of a nuclear power unit and the full value chain traceability according to claim 1, characterized in that, In the above S01 BIO Annotated with B, I, O Mark the beginning, middle or end of the entity and non-entity respectively, for annotating the text of nuclear power defect information.
3. The method for analyzing the quality defect sources of a nuclear power unit and tracing the entire value chain according to claim 1, characterized in that In the above S01 BERT The text vector transformation is to process the labeled text through BERT encoding, fusing semantic, position, and paragraph features to generate a comprehensive feature vector , and then through Transformer the encoder to obtain the corresponding vector of the text , and its simplified formula is: Given the nuclear power quality text description sentence , after BERT processing, the corresponding vector of the text description sentence is obtained , where represents the -th character of the quality text description sentence, represents the word vector of the -th character, and n is the number of characters in the quality text description sentence.
4. The method for analyzing the quality defect source of a nuclear power unit and tracing the entire value chain according to claim 1, characterized in that In the above S01 BiGRU-Attention of BiGRU part, the forward and backward outputs of the GRU network are concatenated to extract text semantic features. The GRU calculation process is as follows: Among them, is t the input word vector at a moment; represents the hidden state vector, represents the time step the hidden state vector, at the time step t the candidate hidden state vector, the input to the update gate the weight matrix, represents the hidden state to the update gate the weight matrix, the input to the update gate the weight matrix, represents the hidden state to the reset gate the weight matrix, represents the input to the candidate hidden state the weight matrix, the reset gate the bias vector, is the bias vector of the update gate ; represents the update gate, which controls the information entering the next state; represents the reset gate, which decides the selection of information; * represents the Hadamard product, which GRU concatenates the forward and backward network outputs to obtain BiGRU a single-character feature output, and the formula is: Among them, represents the characteristic output at the forward moment GRU , and represents the characteristic output at the reverse moment. Merging all the text information of time steps to obtain the text semantic feature matrix, Attention The mechanism is based on the formula A Screen for features related to equipment quality defects and output a key semantic feature matrix, where Query vector( Q) , Key vector (K) 、 Value vector (V) is the new matrix after the projection change of the feature matrix is the transpose matrix of K, is the dimension of the input vector, is the activation function.
5. The method for analyzing the quality defect sources of a nuclear power unit and the full value chain traceability according to claim 1, wherein, In the above-mentioned S02, based on the improved Apriori association mining of quality defects in nuclear power equipment, the original data is first normalized to obtain the cause and phenomenon representations. The Apriori algorithm generates frequent item sets by scanning the transaction database, self-joining operations, and threshold screening, and then obtains frequent 2-item sets through support calculation screening. It recursively continues until the candidate 4-item set is an empty set. The improvement measures include: compressing the database and deleting non-frequent item sets; performing pre-pruning on item sets with the number of single-element occurrences less than k times k ; the lift formula is Among them, , the probability that event X and event Y occur simultaneously, is the probability that event X occurs, is the probability that event Y occurs, represents the confidence level, represents the support degree. When lift is less than 1, the transaction and Y are negatively correlated, that is, the occurrence of Y will cause lift not to occur; when Y is approximately 1, the transaction and lift are independent of each other and have nothing to do with each other; when Y is greater than 1, the transaction and L are positively correlated, and the larger the L value, the higher the correlation between the two, and the higher the probability of their co-occurrence.
6. The method for analyzing the quality defect source and full value chain traceability of a nuclear power unit according to claim 1, wherein In the above S02, the chi-square value between a single feature and the target variable is calculated using the chi-square test, and the formula is Among them is the actual frequency is the theoretical frequency The larger the value, the stronger the correlation between the two variables. Select the top k important features according to the chi-square value ranking. Extract samples through the bootstrap sampling method to form multiple data subsets, train decision trees to construct a random forest, use the majority voting strategy to discriminate the sample category, and visualize the random forest to select the optimal tree to form a visual tree diagram for defect cause analysis.
7. The method for analyzing the quality defect source and full value chain traceability of a nuclear power unit according to claim 1, characterized in that, In the S02, for the analysis of the causes of nuclear power equipment defects under sample imbalance based on NSGALL-RF , by using NSGAII to generate training subsets with large differences and high classification accuracy, NSGAII The algorithm first randomly generates an initial population, and generates offspring populations through selection, crossover, and mutation, repeating until the stopping criterion is met.
8. The method for analyzing the quality defect source of a nuclear power unit and tracing the entire value chain according to claim 1, characterized in that, In the above S03, structured, semi-structured, and unstructured data are collected from the nuclear power intelligent construction platform, construction management platform, and Internet of Things platform system. The node design includes quality problem manifestations, occurrence objects, and causes; the relationship design contains multiple relationships composed of equipment and responsible parties, and equipment and sub-components; the attribute design includes quality standards and node probabilities.
9. The method for analyzing the quality defect sources of a nuclear power unit and tracing the entire value chain according to claim 1, wherein In the above S03, during quality traceability, entity, relationship, and phenomenon information are identified, a standardized query-driven knowledge graph search is generated, traced using an inference model combined with regular matching, and the path probability is calculated using the node probability attribute, and the traceability results are output in probability order.
Citation Information
Patent Citations
BERT-BiGRU-IDCNN-CRF named entity identification method based on attention mechanism
CN112733541A
Nuclear power equipment quality tracing method and system, computer equipment and medium
CN113487211A
Hydraulic engineering potential safety hazard description association rule mining method based on improved Apriori algorithm
CN114756656A
Nuclear power quality defect cause analysis method
CN114969267A
Power equipment defect early warning method and system based on improved association rule, and medium
CN115271263A
Cited By
Welding process tracing method based on multi-source data fusion
CN120725545A
Steel structure construction component tracking and tracing method based on Internet of Things
CN120744401A
Steel structure construction component tracking and tracing method based on internet of things
CN120744401B
Employment matching method and equipment based on data analysis and medium
CN120952730A