Knowledge distillation method and apparatus for reactor nuclear accident deduction and fault diagnosis
By using a PINN-based multi-task network and a multi-task proportional knowledge distillation method, the problems of high computational resource consumption and insufficient real-time performance in nuclear accident fault diagnosis are solved, achieving efficient and accurate fault diagnosis and deduction, which is applicable to real-time diagnosis of nuclear reactors.
Patent Information
- Application Number
- CN202510922854.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2026-03-17
- Estimated Expiration
- 2045-07-04
AI Technical Summary
Existing technologies for nuclear accident fault diagnosis suffer from high computational resource consumption, slow diagnosis speed, difficulty in meeting real-time requirements, and the distillation method cannot handle multiple tasks of inference and diagnosis simultaneously. Furthermore, large models are prone to catastrophic forgetting when they need to be retrained.
We adopt a multi-task network architecture based on Physical Information Neural Network (PINN) and combine it with a multi-task proportional knowledge distillation method. We transfer knowledge from the large model to the student model through a dynamic weight strategy, use kernel space mapping relationship to assist knowledge distillation, and design a dynamic weight function to balance distillation loss and task loss to achieve continuous learning.
It improves the accuracy and extrapolation capabilities of nuclear accident fault diagnosis, reduces computational resource consumption, ensures improved model performance under new data, avoids catastrophic amnesia, and meets the needs of real-time diagnosis.
Smart Images

Figure CN120745746B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of knowledge distillation technology, specifically to a knowledge distillation method for reactor nuclear accident simulation and fault diagnosis. Background Technology
[0002] In the field of knowledge distillation, from the perspective of knowledge type, there are three main distillation methods:
[0003] 1) Distillation based on output characteristics
[0004] This method utilizes the output information of the last layer of neurons in the teacher model to guide the training of the student model. Based on different processing locations of the output feature knowledge, it is further subdivided into distillation based on probability distribution output and distillation based on Logits output. The probability distribution output is the output of the neural network after the softmax layer; the Logits output is the output of the neural network before the last softmax layer.
[0005] For distillation based on probability distribution output, the core idea is to have the student model mimic the output probability distribution of a well-trained teacher model after a softmax layer. Distance metrics (such as KL divergence and cross-entropy) are used to measure the difference in probability distributions between the teacher and student models, and the student model parameters are optimized to minimize this difference, resulting in highly reliable results. Hinton et al.'s distillation concept utilizes a temperature coefficient to soften labels, emphasizing the role of dark knowledge, allowing the student model to learn the softened probability distribution of the teacher model. For distillation based on Logits output, the principle is to have the student model learn the original output of the teacher model before the softmax layer. These Logits contain the original information of the teacher model before softmax normalization. Distance metrics such as Euclidean distance are used to measure the difference in Logits outputs between the teacher and student models, guiding the student model's training. Ba et al. proposed using simple networks to mimic the Logits output of complex networks to achieve model compression.
[0006] 2) Distillation based on intermediate characteristics
[0007] This method leverages knowledge from the intermediate feature maps of a neural network to guide student model training. Typically, specific projection layers need to be designed to align the semantic information of the teacher and student model feature maps, enabling the student model to learn better from the intermediate features of the teacher model. Different intermediate layers of a neural network contain different levels of feature information, which is crucial for the student model's learning. By designing projection layers, the intermediate features of the teacher and student models are mapped to the same semantic space. Then, distance metrics (Manhattan distance, Euclidean distance, maximum mean difference, etc.) are used to measure the difference between the intermediate feature maps of the teacher and student models. Minimizing this difference is the optimization objective, gradually bringing the student model's intermediate features closer to those of the teacher model, thereby improving the student model's performance. Adriana et al. first proposed transferring intermediate knowledge by selecting a hidden layer from the teacher model as a cue layer and a hidden layer from the student model as a guide layer. The guide layer learns from the cue layer, achieving knowledge transfer from the teacher model to the student model.
[0008] 3) Distillation based on relational features
[0009] This method is based on knowledge of relational features, typically referring to the relationships between features of different layers or different input samples. For inter-layer relationships, it focuses on the associations between features of different layers; for inter-sample relationships, it focuses on the associations between features of different input samples. Regarding inter-layer relationships, it argues that the feature relationships between different layers of the teacher model contain the problem-solving process and logic, and the student model can better understand the teacher model's knowledge system by learning these inter-layer relationships. Regarding inter-sample relationships, it believes that similar samples should have similar feature representations or activation patterns in the teacher model. By introducing graph structures, similarity matrices, and other methods, it captures the relational information between samples and uses this information to guide the optimization of the student model parameters.
[0010] In the nuclear energy field, fault diagnosis of reactor nuclear accidents is crucial, as it directly relates to the safe and stable operation of nuclear power plants and the safety of the surrounding environment and personnel. However, current nuclear accident fault diagnosis faces many severe challenges, urgently requiring a method capable of rapid diagnosis and real-time inference.
[0011] Nuclear accidents are characterized by their suddenness and complexity. Nuclear reactor systems are inherently complex, comprising numerous subsystems and devices that are interconnected and mutually influential. Once an accident occurs, the manifestations of the fault are diverse, and the causes may involve multiple links and factors, making the fault diagnosis process exceptionally complex. Traditional fault diagnosis methods often require substantial computational resources and time to analyze various data and signals, making it difficult to provide accurate diagnostic results within the short timeframe following an accident, thus failing to meet real-time requirements.
[0012] With the continuous development of nuclear energy technology, the requirements for the accuracy and reliability of nuclear accident fault diagnosis are becoming increasingly stringent. In the event of a nuclear accident, every second is crucial; accurate diagnostic results provide critical information for emergency response, helping operators take timely and correct measures to prevent further escalation of the accident. Therefore, a method capable of providing high-precision diagnostic results in a short time is needed.
[0013] In recent years, deep learning technology has made significant progress in the field of fault diagnosis. Large models, with their powerful learning capabilities and rich parameters, have performed exceptionally well in handling complex data and tasks, learning complex fault characteristics and patterns from large amounts of nuclear accident data, providing new insights for nuclear accident fault diagnosis. However, large models also have significant limitations. They typically have a massive parameter scale, high computational complexity, and demanding hardware resource requirements, making them difficult to deploy in resource-constrained environments, such as embedded devices or mobile terminals at nuclear power plant sites. Furthermore, the inference speed of large models is relatively slow, failing to meet the real-time requirements of nuclear accident fault diagnosis.
[0014] To overcome these limitations of large models while fully leveraging their powerful learning capabilities, knowledge distillation has emerged. Knowledge distillation transfers knowledge learned by the teacher model to the student model, enabling the student model with fewer parameters to achieve near-large model performance while maintaining a smaller scale and lower computational complexity. In nuclear accident fault diagnosis scenarios, knowledge distillation can transfer the rich fault features and diagnostic knowledge learned by the large model in processing nuclear accident data to the small model, giving it the ability for rapid diagnosis and real-time inference. This allows the small model to be deployed on various equipment at the nuclear power plant site, quickly and accurately diagnosing nuclear accidents and providing timely and effective support for emergency response, thereby effectively ensuring the safe operation of the nuclear power plant. Therefore, researching a knowledge distillation method based on a multi-task model for reactor nuclear accident simulation and fault diagnosis has significant practical implications and application value.
[0015] In summary, existing distillation methods have the following problems:
[0016] 1) Distillation is often targeted at a specific task and cannot be carried out in a multi-task manner. To address this issue, nuclear accident processes require both simulation and diagnosis, necessitating the design of a multi-task model that integrates simulation and diagnosis.
[0017] 2) Whether it is a distillation method based on output features, intermediate features or relation, it only aligns from one aspect, such as the feature map of classification or the distribution of regression, and tends to let the model learn knowledge from one aspect. However, fault diagnosis is not only related to classification and regression results, but also to physical processes. Therefore, this invention application designs a knowledge distillation method that targets both physics and data.
[0018] 3) Existing distillation methods require large models to be retrained and re-distilled once the data changes. To address the above problems, this invention proposes a continuous distillation learning method. By employing a progressively infiltrated continuous distillation approach, it ensures the preservation of original knowledge, avoids catastrophic forgetting, and learns new knowledge simultaneously. Summary of the Invention
[0019] The present invention proposes a knowledge distillation method, equipment and storage medium for reactor nuclear accident simulation and fault diagnosis, which can at least solve one of the technical problems in the background art.
[0020] To achieve the above objectives, the present invention adopts the following technical solution:
[0021] A knowledge distillation method for reactor nuclear accident simulation and fault diagnosis includes the following steps:
[0022] Step 1: Obtain fault data from the nuclear energy simulation device;
[0023] Step 2: Set up a multi-task network;
[0024] Step 3: Take the dataset obtained in Step 1 D The network is divided into a training set, a validation set, and a test set. The training set is used to pre-train the multi-task network built in step 2.
[0025] Step 4: After pre-training is completed, fix the parameters in the multi-task network, that is, freeze the large-parameter pre-trained model of the multi-task network as the teacher model.
[0026] Step 5: Use a multi-task, proportional knowledge distillation method to transfer knowledge from the teacher model to the student model;
[0027] Step 6: Input the real-time monitoring data of the target pipeline to be diagnosed into the student model trained in Step 5, and use the student model to infer the physical quantity deduction results and fault results of the target pipeline.
[0028] In another aspect, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method described above.
[0029] In another aspect, the present invention also discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the method described above.
[0030] In summary, this invention proposes a method for reactor nuclear accident simulation and fault diagnosis based on multi-task model knowledge distillation. By constructing a multi-task network based on Physical Information Neural Network (PINN), it can simultaneously handle multiple fault diagnosis and parameter simulations related to reactor nuclear accidents, fully utilizing the correlation between different tasks to improve the accuracy of fault diagnosis and the ability to simulate the development process of nuclear accidents. After pre-training on a large-scale fault dataset, the large-parameter pre-trained PINN model is frozen, and a multi-task, proportionally infiltrated knowledge distillation method is used to transfer the knowledge from the large model to the student model. This significantly reduces computational resource consumption and improves computational efficiency while ensuring diagnostic and simulation performance.
[0031] This method can diagnose reactor faults promptly and accurately and predict the development of nuclear accidents, providing decision support for nuclear power plant operators. It helps to take effective countermeasures in the early stages of an accident, reducing its severity and ensuring the safe and stable operation of the nuclear reactor. Through knowledge distillation technology, computational resource consumption is reduced without affecting model performance, lowering diagnostic costs and improving resource utilization efficiency, thus aligning with the requirements of sustainable development.
[0032] Compared with existing reactor nuclear accident simulation and fault diagnosis technologies, this invention has the following significant advantages and positive effects:
[0033] 1) Improve fault diagnosis accuracy and nuclear accident simulation capabilities. This invention, by constructing a multi-task network architecture based on Physical Information Neural Network (PINN), can simultaneously handle fault diagnosis and parameter simulation tasks. In contrast, traditional methods typically rely on expert experience and simple models, making it difficult to comprehensively and accurately identify fault types and simulate accident development processes. PINN integrates physical laws into neural network training and utilizes kernel space mapping relationships to assist in knowledge distillation, making the model more consistent with actual physical laws when simulating physical quantities, thus further improving the accuracy of the simulation.
[0034] 2) Reduce computational resource consumption and improve computational efficiency. This invention pre-trains the PINN pre-trained model on a large-scale fault dataset, freezes the model with a large number of parameters, and employs a multi-task, proportional knowledge distillation method to transfer knowledge from the large model to a student model with a simpler structure and fewer parameters. This significantly reduces computational resource consumption and improves computational efficiency while ensuring diagnostic and inference performance. Traditional methods often face problems of low computational efficiency and low diagnostic accuracy when processing large-scale fault data, while this invention effectively solves this problem through knowledge distillation technology.
[0035] 3) The dynamic weighting strategy prioritizes task performance on new data while retaining a certain degree of original knowledge. During knowledge distillation, this invention designs a dynamic weighting function based on data similarity and training progress. This allows for greater emphasis on preserving original knowledge in the early stages of training, gradually reducing reliance on original knowledge as training progresses and focusing more on task performance on new data. This dynamic adjustment mechanism not only improves the learning efficiency of the student model but also avoids the problem of catastrophic forgetting. Attached Figure Description
[0036] Figure 1 This is a flowchart of an embodiment of the present invention;
[0037] Figure 2 This is a flowchart illustrating the multi-task, proportionally infiltrated knowledge distillation process in an embodiment of the present invention.
[0038] Figure 3 This is a schematic diagram illustrating the principle of multi-task knowledge distillation in an embodiment of the present invention. Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0040] In the nuclear energy field, the safe and stable operation of reactors is of paramount importance. A nuclear accident can cause enormous economic losses and pose serious threats to the environment and public health. Therefore, timely and accurate reactor nuclear accident simulation and fault diagnosis are crucial for ensuring nuclear safety. Traditional reactor fault diagnosis methods often rely on expert experience and simple models. These methods have limitations when dealing with complex and ever-changing nuclear accident scenarios, making it difficult to comprehensively and accurately identify fault types and simulate accident development. With the continuous development of nuclear energy technology, reactor systems are becoming increasingly complex, generating fault data that is large-scale and high-dimensional. Existing diagnostic methods face problems of low computational efficiency and low diagnostic accuracy when processing large-scale fault data. Furthermore, nuclear accident simulation requires comprehensive consideration of multiple related physical quantities and parameters to better understand the fault formation process and factors. Therefore, this invention proposes a method capable of efficiently processing large-scale fault data and accurately diagnosing faults and simulating nuclear accidents, which is an urgent need in the field of nuclear safety.
[0041] Based on the above background, this invention aims to propose a method for reactor nuclear accident simulation and fault diagnosis based on multi-task model knowledge distillation, such as... Figure 1 As shown, the specific steps are as follows:
[0042] Step 1: Acquire fault data from the nuclear energy simulation device. The nuclear energy simulation device is monitored in real time using temperature sensors, pressure sensors, neutron flux detectors, etc., within the nuclear energy setup. Relevant physical parameter data for key components of the reactor are collected under different fault scenarios. Let the collected dataset be... ,in It is the input feature vector, which includes temperature T, pressure P, and neutron flux. Physical quantities such as coolant flow rate Q, i.e. ; It is the corresponding output label vector, including the fault type F, i.e. N is the number of data samples.
[0043] Step 2: Set up a multi-task network for PINN
[0044] A multi-task network architecture based on Physics-Informed Neural Networks (PINN) is constructed. This network consists of an input layer, multiple hidden layers, and an output layer. Let the network input be... x The output is y The forward propagation process of a network can be represented as
[0045] ;
[0046] Where W and b are the weight matrix and bias vector of the l-th layer, respectively, σ is the ReLU activation function, and L is the number of layers in the network. In a multi-task network, different tasks share some hidden layers to take advantage of the correlation between tasks, while the output layer has its own output nodes for different tasks.
[0047] Step 3: Pre-train the nuclear energy multi-task model on a large-scale fault dataset. Unlike ordinary fault diagnosis models, this model's training is divided into two parts: firstly, classifying the causes of accidents using the classification results; and secondly, performing a regression task to deduce unknown quantities using known quantities. The dataset obtained in Step 1 will be used... D The dataset is divided into training, validation, and test sets. The multi-task network built in step 2 is pre-trained using the training set. A loss function L is defined, which is a combination of the loss functions of multiple tasks. For the fault type identification task, the cross-entropy loss function can be used. For parameter extrapolation and evaluation tasks, the mean squared error loss function can be used. The total loss function can then be expressed as:
[0048] ;
[0049] Here, α and β are weighting coefficients that balance the losses from different tasks. By minimizing the loss function L, the network parameters are updated using the stochastic gradient descent optimization algorithm and its variant Adam, enabling the network to learn patterns and features from the data on the training set.
[0050] Step 4: Freeze the large-parameter pre-trained model of PINN. After pre-training, the parameters in the multi-task network are fixed, i.e., the large-parameter pre-trained model of PINN is frozen. The advantage of doing this is that it preserves the rich knowledge and feature representations already learned by the pre-trained model, avoiding the destruction of this knowledge during subsequent knowledge distillation, and reducing the computational cost of subsequent training.
[0051] Step 5: Multi-task proportional knowledge distillation. A student model with 3 layers is designed, which has a simpler structure and fewer parameters than the teacher model (i.e., the pre-trained model frozen in Step 4). A multi-task proportional knowledge distillation method is used to transfer knowledge from the teacher model to the student model. Let the output of the teacher model be... The output of the student model is The loss function for knowledge distillation is defined as:
[0052] ;
[0053] in, This is the distillation loss, used to measure the difference between the outputs of the teacher model and the student model. In this embodiment of the invention, KL divergence is used.
[0054] ;
[0055] This represents the task loss of the student model across various tasks. Similar to the loss function in step 3, λ is a hyperparameter used to balance the ratio of distillation loss to task loss. During training, the value of λ is gradually adjusted so that the student model can learn the knowledge of the teacher model at different stages while maintaining its performance on each task.
[0056] Step 6: Utilize the student model to deduce the physical quantity projections and fault results for the target pipeline. Input the real-time monitoring data of the target pipeline to be diagnosed into the student model trained in Step 5. Based on the learned knowledge and feature representations, the student model outputs the predicted values of relevant physical quantities of the target pipeline, such as temperature and pressure, as well as the fault diagnosis result and fault type. This provides decision support for nuclear power plant operators, helping them to promptly detect and handle faults in the reactor, ensuring the safe and stable operation of the nuclear energy system.
[0057] The specific process of multi-task proportional knowledge distillation mentioned in step 5 is as follows: Figure 2 As shown.
[0058] The following are the steps of the distillation process for the multi-task, proportionally infiltrated knowledge distillation method:
[0059] Step S50: Design a student model with a relatively simple structure and a small number of parameters. This ensures that it has the ability to handle multiple tasks, but with lower computational complexity than larger models. For the student model... The parameters are used for Kaiming initialization.
[0060] Step S51: Freeze the teacher model and initially distill the student model. Set the initial distillation loss weight λ (to balance distillation loss and task loss), as well as the learning rate, batch size, etc. Then, freeze the large model... Freeze and use the initial training dataset D on the student model. Perform initial knowledge distillation. During training, calculate the distillation loss KL divergence. and mission losses According to the total loss Update the parameters of the student model. After a certain number of training rounds, the student model after initial distillation is obtained. .
[0061] Step S52, New Data Collection and Integration. Continuously monitor the operation of the nuclear energy system. When new fault data or changes in system status are detected, collect new data samples and construct a new dataset.
[0062] Step S53: Gradually adjust the distillation loss weights according to the number of training rounds. A progressive distillation method is used to gradually adjust the distillation loss weights λ. The value of λ is gradually changed during training according to similarity and model training progress, so that task performance on new data is emphasized in the later stages of training, while retaining a certain degree of original knowledge. This embodiment of the invention designs a dynamic weight function based on data similarity and training progress. Let the number of training rounds be t, the total number of training rounds be T, and the similarity between the new data and the initial data be s (0≤s≤1, obtained by calculating the L2 norm distance of the data distribution). The dynamic adjustment function of λ is defined as follows:
[0063] ;
[0064] Where α is a hyperparameter controlling the decay rate of the function, and γ is a weighting coefficient adjusting the influence of similarity. The first term of this function... It is an exponentially decaying term, whose value gradually decreases as the number of training epochs t increases. In the early stages of training, the distillation loss has a larger weight, and the model focuses more on retaining the original knowledge obtained from distillation of the initial large model; as training progresses, the dependence on the original knowledge gradually decreases, and more attention is paid to the task performance on new data.
[0065] Step S54: Continuous Distillation Learning. Freeze the large model and use the updated dataset D to refine the student model. Continuous distillation training is performed. During the training process, each training round t is calculated based on the aforementioned dynamic weight function. The distillation loss and task loss are calculated, and the parameters of the student model are updated based on the total loss. As training progresses, the student model gradually learns knowledge from new data while dynamically retaining the original knowledge obtained from the distillation of the initial large model.
[0066] like Figure 3This paper introduces the principle of knowledge distillation algorithm incorporating a proportional approach across multiple tasks. When handling multi-task scenarios such as reactor nuclear accident classification, a teacher network and a student network based on Physical Information Neural Network (PINN) are first constructed using a nuclear accident classification dataset. The teacher network has a large number of parameters and high performance, while the student network has a simple structure. During training, the teacher network learns the fault classification results and parameter probability distribution from the data; simultaneously, the student network learns from the same data. A dynamic distillation loss weight adjustment mechanism is designed. Initially, a larger distillation loss weight is assigned, allowing the student network to learn more from the teacher network and retain original knowledge. As training progresses, the distillation loss weight is gradually reduced, increasing the proportion of task loss for the student network, making it more focused on task performance on new data. Simultaneously, by combining the core physical parameter derivation task, physical laws are integrated into the training, and kernel space mapping relationships are used to assist knowledge distillation. Ultimately, this enables the student network to effectively learn new knowledge while avoiding catastrophic forgetting, achieving knowledge transfer and fusion under multiple tasks.
[0067] The following are alternative solutions to the above embodiments:
[0068] 1) Model Alternatives. The distillation method described in this embodiment of the invention is designed for two tasks in the field of nuclear safety: parameter derivation and fault diagnosis. For the model part, the PINN model is used. However, its specific structure can be a new model such as the Transformer model, Mamba architecture, or XLSTM.
[0069] 2) The order of the distillation process. The student model can be initialized first, or the initial distillation can be performed first without model initialization. The difference is that initialization can speed up convergence and allow the loss in the first few rounds to converge smoothly, so as not to fail to learn the initial information and cause gradient explosion or gradient vanishing.
[0070] As can be seen from the above, the core technical points of this invention application are as follows:
[0071] 1) Construct a multi-task network based on Physical Information Neural Network (PINN).
[0072] Unlike single-function fault diagnosis models, which first obtain physical parameters and then use thresholding or direct classification to achieve fault diagnosis, the multi-task PINN network proposed in this patent can simultaneously handle fault diagnosis and parameter extrapolation tasks in reactor nuclear accidents. By integrating physical laws into the neural network training using PINN, and aiding knowledge distillation through kernel space mapping relationships, the accuracy of fault diagnosis and the ability to extrapolate nuclear accidents are improved.
[0073] 2) Multi-task proportional knowledge distillation method
[0074] This patent's knowledge distillation is multi-task-based, designing loss functions and aligning features across two tasks: nuclear accident fault classification and parameter extrapolation regression. It employs a multi-task, proportionally infiltrated knowledge distillation method to transfer knowledge from the teacher model (a pre-trained model with a large number of parameters) to the student model, which has a simpler structure and fewer parameters. By progressively adjusting the distillation loss weights, it emphasizes preserving original knowledge in the early stages of training, while gradually increasing focus on performance on new data tasks as training progresses, thus avoiding catastrophic forgetting.
[0075] 3) Design of dynamic weight function
[0076] Unlike ordinary distillation functions with fixed weights, this patent proposes a proportionally varying loss function, a dynamic weight function based on data similarity and training progress, used to balance distillation loss and task loss. This function assigns a larger weight to the distillation loss in the early stages of training, allowing the student model to learn more from the teacher model; as training progresses, the weight of the distillation loss is gradually reduced, increasing the proportion of the student model's own task loss, making it more focused on task performance on new data.
[0077] 4) Continuous distillation learning mechanism
[0078] Continuously monitor the operation of the nuclear energy system. When new fault data or changes in system status are detected, collect new data and construct a new dataset. Use the updated dataset to continuously distill and train the student model, ensuring that the student model can dynamically learn new knowledge while retaining the original knowledge obtained from the initial large model distillation.
[0079] In another aspect, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method described above.
[0080] In another aspect, the present invention also discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the method described above.
[0081] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the knowledge distillation methods for reactor nuclear accident simulation and fault diagnosis described in the above embodiments.
[0082] It is understood that the systems, devices, and storage media provided in the embodiments of the present invention correspond to the methods provided in the embodiments of the present invention, and the explanations, examples, and beneficial effects of the relevant content can be referred to the corresponding parts of the above methods.
[0083] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).
[0084] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0085] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0086] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A knowledge distillation method for reactor nuclear accident progression and fault diagnosis, characterized in that, The method comprises the following steps, Step 1, obtaining nuclear energy simulation device fault data; Step 2: building a multi-task network, the multi-task network adopts a PINN model; Step 3, dividing the data set obtained in step 1 into a training set, a validation set and a test set, and using the training set to pre-train the multi-task network built in step 2 D Step 3, dividing the data set obtained in step 1 into a training set, a validation set and a test set, and using the training set to pre-train the multi-task network built in step 2 Step 4, after pre-training is completed, parameters in the multi-task network are fixed, that is, a large-parameter pre-training model of the multi-task network is frozen as a teacher model; Step 5, a multi-task proportional penetration knowledge distillation method is adopted to migrate the knowledge of the teacher model to a student model; Step 6, real-time monitoring data of a target pipeline to be diagnosed are input into the student model trained in step 5, and physical quantity deduction results and fault results of the target pipeline are obtained by inference of the student model; Step 5 specifically comprises, Let the output of the teacher model be , and the output of the student model be , the loss function of knowledge distillation is defined as: ; where, is the distillation loss, measuring the difference between the teacher and student model outputs using the KL divergence: ; is the task loss of the student model on each task; λ is a hyperparameter balancing the ratio of the distillation loss and the task loss; during the training process, the value of λ is adjusted step by step, so that the student model can learn the knowledge of the teacher model and maintain the performance on each task at different stages; Step 5 involves a multi-task proportional penetration knowledge distillation method, which comprises, Step S50, designing a student model , ensuring that it has the ability to handle multi-task processing, but the computational complexity is lower than that of large models, and the parameters of the student model are initialized by Kaiming Step S51, freeze the teacher model, and initial distillation of the student model; set the initial distillation loss weight λ for balancing the distillation loss and the task loss, and the learning rate, batch size; freeze the large model , use the initial training data set D to train the student model , perform initial knowledge distillation, in the training process, calculate the distillation loss KL divergence and the task loss , update the parameters of the student model according to the total loss , after a set of training, get the initial distilled student model ; Step S52, new data collection and integration; Step S53, adjusting the distillation loss weight step by step according to the training round; a continuous distillation method of proportional penetration is adopted to adjust the distillation loss weight λ; the value of λ is gradually changed in the training process according to the set increasing or decreasing rule, so that the performance of the task on the new data is paid more attention to in the later training period, and the original knowledge meeting the requirements is reserved; Step S54, continuous distillation learning; Step S52 comprises, The running condition of the nuclear energy system is continuously monitored, and when new fault data or system state changes are found, new data samples are collected to construct a new data set; Step S53 comprises, A dynamic weight function based on data similarity and training progress is designed; the training round is t, the total training round is T, the similarity of new data and initial data is s, 0≤s≤1, which is obtained by calculating the L2 norm distance of data distribution, and the dynamic adjustment function of λ is defined as ; Where α is a hyperparameter controlling the decay rate of the function, and γ is a weighting coefficient adjusting the influence of similarity; the first term of this function It is an exponentially decaying term, and its value gradually decreases as the number of training epochs t increases. In the early stage of training, the distillation loss has a larger weight, and the model focuses more on retaining the original knowledge obtained from the distillation of the initial large model. As training progresses, the dependence on the original knowledge gradually decreases, and more attention is paid to the task performance on new data. Step S54 comprises, Freezing the large model, using the updated dataset to train the student model Performing continuous distillation training; during the training process, each training round t is calculated according to the dynamic weight function , calculate the distillation loss and task loss, and update the parameters of the student model according to the total loss; as the training progresses, the student model gradually learns the knowledge in the new data, while dynamically retaining the original knowledge distilled from the initial large model.
2. The knowledge distillation method for reactor core accident progression and fault diagnosis according to claim 1, characterized in that: Step 1 comprises, The nuclear energy simulation device is monitored in real time through temperature sensors, pressure sensors and neutron flux detectors in the nuclear energy setting, and relevant physical parameter data of each key part of the reactor under different fault scenarios are collected.
3. The knowledge distillation method for reactor core accident progression and fault diagnosis according to claim 2, characterized in that: Step 1 further comprises, Let the collected dataset be where is the input feature vector, containing the physical quantities of temperature T, pressure P, neutron flux Φ, coolant flow rate Q, i.e. ; is the corresponding output label vector, including the fault type F, i.e. , and N is the number of data samples.
4. The knowledge distillation method for reactor core accident progression and fault diagnosis of claim 1, wherein: Step 2 comprises, A multi-task network architecture based on a physical information neural network PINN is constructed. The network consists of an input layer, a plurality of hidden layers, and an output layer, where the input of the network is x , the output is y , and the forward propagation process of the network is represented by ; Wherein, W and b are the weight matrix and bias vector of the lth layer respectively, σ is the activation function ReLU function, and L is the number of network layers; in the multi-task network, different tasks share part of the hidden layer to utilize the correlation between tasks, and each output node is provided for different tasks in the output layer.
5. The knowledge distillation method for reactor core accident progression and fault diagnosis of claim 1, wherein: Step 3 comprises, The model training is divided into two parts, one is to classify the causes of the accident according to the classification result, and the other is to use the known quantity to deduce the unknown quantity.
6. The knowledge distillation method for reactor core accident progression and fault diagnosis according to claim 5, characterized in that: Step 3 comprises, The loss function L is defined by combining the loss functions of multiple tasks, and a cross-entropy loss function is used for the fault type identification task A mean square error loss function is used for the parameter deduction evaluation task The total loss function is represented as: ; Wherein, α and β are weight coefficients for balancing the loss of different tasks; by minimizing the loss function L, the parameters of the network are updated using the optimization algorithm stochastic gradient descent method and its variant Adam, so that the network learns the patterns and features in the data on the training set.
7. The knowledge distillation method for reactor core accident progression and fault diagnosis of claim 1, wherein: Step S6 comprises, The student model outputs the prediction value of the related physical quantity temperature and pressure of the target pipeline and the fault diagnosis result fault type according to the learned knowledge and feature representation, provides decision support for the operators of the nuclear power plant, helps them to discover and handle the faults in the reactor in time, and guarantees the safe and stable operation of the nuclear energy system.
8. The knowledge distillation method for reactor core accident progression and fault diagnosis of claim 1, wherein: The order of the distillation process of the method of multi-task proportional penetration of knowledge distillation in step 5 is to first perform initialization of the student model, or to first perform initial distillation without model initialization.
9. A computer-readable storage device storing a computer program, wherein the computer program comprises instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1-8. The computer program, when executed by the processor, causes the processor to perform the steps of the method of any one of claims 1 to 8. 10.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-9. The computer program, when executed by the processor, causes the processor to perform the steps of the method of any one of claims 1 to 8.
Citation Information
Patent Citations
Vehicle re-identification method and system based on multi-task learning and knowledge distillation
CN114022697A
Knowledge distillation method based on parameter efficient module and multi-teacher knowledge distillation
CN118747507A