Nuclear power software detection method and device based on generative training model

By applying the M-BERT of the generative training model in nuclear power software review, the problems of low efficiency and strong subjectivity of traditional review methods are solved, efficient and accurate automated review is achieved, and the safety and stability of the nuclear power system are improved.

CN119988208APending Publication Date: 2025-05-13STATE POWER INVESTMENT CORPORATION RESEARCH INSTITUTE +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311484732.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-08
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Traditional nuclear power software review methods are inefficient, subjective, and prone to errors, making it difficult to continuously monitor and respond to software version changes.

Method used

Using a nuclear power software detection method based on a generative training model, the natural language processing model M-BERT is constructed, and the nuclear power software data is analyzed using pre-training and fine-tuning modules to realize automated review.

Benefits of technology

It significantly shortens the review cycle, improves the accuracy and efficiency of the review, reduces subjectivity and error rates, ensures the sustainability and scalability of the review, and improves the safety of the nuclear power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988208A_ABST
    Figure CN119988208A_ABST
Patent Text Reader

Abstract

The invention discloses a nuclear power software detection method and device based on a generative training model, and the method comprises the steps: obtaining a training data set and a test data set based on nuclear power software review data; constructing a natural language processing model M-BERT suitable for the nuclear power software review data; wherein the M-BERT comprises a fine tuning module and a pre-training module; training the M-BERT by using the training data set through a pre-training module to obtain a pre-trained M-BERT, and inputting the preprocessed nuclear power software review data into the pre-trained M-BERT through a fine tuning module to perform model fine tuning on the pre-trained M-BERT to obtain a fine-tuned M-BERT; and inputting the test data set into the fine-tuned M-BERT to carry out nuclear power software quality detection so as to obtain a nuclear power software quality detection result. According to the method, the quality detection of the nuclear power software can be efficiently, accurately and automatically carried out, and powerful technical support is provided for development and operation of the nuclear power software.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of nuclear power software review, and in particular to a nuclear power software detection method and device based on a generative training model. Background Art

[0002] As a form of clean energy, nuclear energy plays an important role in energy supply and environmental protection. However, the high complexity and risk of the nuclear energy field make the operation, management and decision-making of nuclear power plants challenging. In order to ensure the safety, stability and efficiency of nuclear power plants, accurate data analysis and prediction tools are needed to make wise decisions in complex environments.

[0003] The nuclear power software review method based on generative training models aims to solve the problems of low efficiency, strong subjectivity, and easy errors in traditional nuclear power software review methods. By introducing generative training model technology, efficient review of nuclear power software can be achieved, thereby improving the safety, stability and development efficiency of nuclear power systems.

[0004] Specifically, the present invention aims to solve the following technical problems: Traditional manual review methods require a lot of time and human resources, resulting in a slow review process. The present invention aims to provide an efficient review method, which uses a generative training model to achieve rapid review of nuclear power software and significantly shorten the review cycle. Manual review is easily affected by factors such as individual experience and subjective judgment, which may lead to inconsistency in review results. The present invention uses a generative training model to reduce the influence of subjective judgment and improve the objectivity and accuracy of the review by learning and analyzing a large amount of data. There may be problems such as fatigue, repetition and lack of concentration in manual review, which may easily lead to wrong judgments. The present invention avoids these problems and reduces the error rate that may occur in the review by automating the review process. Traditional review methods are difficult to achieve continuous, 7×24-hour monitoring, and are also difficult to cope with different software versions and technical changes. Based on a generative training model, the present invention has the ability to continuously learn and adapt to new environments, ensuring the sustainability and scalability of the review method. Manual review may ignore or miss some potential safety hazards, resulting in safety risks for nuclear power systems during operation. Summary of the invention

[0005] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.

[0006] To this end, the present invention proposes a nuclear power software detection method based on a generative training model. By learning and analyzing a large amount of nuclear power software data, efficient review of nuclear power software is achieved, which effectively reduces the workload of manual review and improves the accuracy and efficiency of the review.

[0007] Another object of the present invention is to propose a nuclear power software detection device based on a generative training model.

[0008] To achieve the above object, the present invention provides a method for measuring a signal based on the characteristics of paroxysmal chaos, comprising:

[0009] Obtain training and testing data sets based on nuclear power software review data;

[0010] Constructing a natural language processing model M-BERT suitable for the nuclear power software review data; wherein the M-BERT includes a fine-tuning module and a pre-training module;

[0011] The pre-training module uses the training data set to train the M-BERT to obtain a pre-trained M-BERT, and the fine-tuning module inputs the pre-processed nuclear power software review data into the pre-trained M-BERT to perform model fine-tuning on the pre-trained M-BERT to obtain a fine-tuned M-BERT;

[0012] The test data set is input into the fine-tuned M-BERT to perform nuclear power software quality detection to obtain the nuclear power software quality detection result.

[0013] The nuclear power software detection method based on the generative training model in the embodiment of the present invention may also have the following additional technical features:

[0014] In one embodiment of the present invention, the data type of the input nuclear power software review data is determined. If the data type is a non-text data type, M-BERT-P is selected for data processing. If the data type is a text data type, M-BERT-T is selected for data processing.

[0015] In one embodiment of the present invention, a large amount of nuclear power software review data is obtained, including at least source code, design documents, test data and error reports.

[0016] In one embodiment of the present invention, a large amount of nuclear power software review data is screened according to preset screening conditions to obtain nuclear power software review data that meets the model requirements.

[0017] In one embodiment of the present invention, preprocessing the nuclear power software review data that meets the model requirements includes:

[0018] Remove duplicate data, handle missing values, label data, and convert data formats.

[0019] In one embodiment of the present invention, the data annotation includes:

[0020] Marking the code data of the non-text data type: marking key structures, variables and functions of the code;

[0021] The text data of the text data type is annotated: technical terms and key information in the text are annotated.

[0022] In one embodiment of the present invention, the pre-training module includes at least an optimizer, and the fine-tuning module includes a generator and a discriminator; the method further includes:

[0023] Generate fake samples based on random noise using a generator, and obtain real samples from real nuclear power software review data;

[0024] Inputting the false sample and the real sample into a discriminator to output a classification probability of the real sample, updating the parameters of the discriminator and adjusting the loss function of the discriminator according to the classification probability to train the discriminator;

[0025] Training the generator by updating the parameters of the generator through back propagation and adjusting the loss function of the generator according to the classification of the generated false samples by the discriminator;

[0026] The fine-tuned M-BERT is obtained by minimizing the loss function through the optimizer to fine-tune the pre-trained generator and discriminator.

[0027] In one embodiment of the present invention, the method further includes:

[0028] Get a training dataset of text data type;

[0029] Extracting input sequences of different lengths from a training data set of the text data type;

[0030] The input sequences of different lengths are input into the M-BERT to automatically select to truncate long sentences or to fill short sentences in the input sequence based on a dynamic masking technique.

[0031] To achieve the above object, the present invention further provides a nuclear power software detection device based on a generative training model, comprising:

[0032] A data set acquisition module is used to acquire training data sets and test data sets based on nuclear power software review data;

[0033] A model building module, used to build a natural language processing model M-BERT suitable for the nuclear power software review data; wherein the M-BERT includes a fine-tuning module and a pre-training module;

[0034] A model training module, used to train the M-BERT using the training data set through the pre-training module to obtain a pre-trained M-BERT, and input the pre-processed nuclear power software review data into the pre-trained M-BERT through the fine-tuning module to perform model fine-tuning on the pre-trained M-BERT to obtain a fine-tuned M-BERT;

[0035] The data detection module is used to input the test data set into the fine-tuned M-BERT to perform nuclear power software quality detection to obtain a nuclear power software quality detection result.

[0036] The nuclear power software detection method and device based on the generative training model of the embodiment of the present invention can quickly identify potential problems in the software, can timely identify and correct safety hazards in the software, and improve the safety of the nuclear power system.

[0037] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0039] Figure 1 is a flow chart of a nuclear power software detection method based on a generative training model according to an embodiment of the present invention;

[0040] Figure 2 is a flowchart of network model training according to an embodiment of the present invention;

[0041] Figure 3 is a schematic diagram of data preparation and preprocessing according to an embodiment of the present invention;

[0042] Figure 4 It is a structural schematic diagram of a nuclear power software detection device based on a generative training model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0043] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0044] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0045] The following describes a nuclear power software detection method and device based on a generative training model according to an embodiment of the present invention with reference to the accompanying drawings.

[0046] Figure 1 It is a flow chart of a nuclear power software detection method based on a generative training model according to an embodiment of the present invention.

[0047] like Figure 1 As shown, the method includes but is not limited to the following steps:

[0048] S1, obtain training data set and test data set based on nuclear power software review data;

[0049] S2, constructing a natural language processing model M-BERT suitable for the nuclear power software review data; wherein the M-BERT includes a fine-tuning module and a pre-training module;

[0050] S3, training the M-BERT using the training data set through the pre-training module to obtain a pre-trained M-BERT, and inputting the pre-processed nuclear power software review data into the pre-trained M-BERT through the fine-tuning module to perform model fine-tuning on the pre-trained M-BERT to obtain a fine-tuned M-BERT;

[0051] S4, input the test data set into the fine-tuned M-BERT to perform nuclear power software quality detection to obtain the nuclear power software quality detection results.

[0052] In one embodiment of the present invention, before applying the generative training model, a large amount of nuclear power software data, including source code, design documents, test data, etc., needs to be prepared. At the same time, the data is preprocessed, including data cleaning and standardization, to ensure the training effect of the model. Figure 2 shown.

[0053] It is understandable that before inputting the input data into the pre-trained language representation model, the collected data needs to be pre-processed.

[0054] Preferably, a large amount of nuclear power software data is collected, including but not limited to source code: source code files of various software modules. Design documents: software design specifications, flow charts and other documents. Test data: data sets used to test software functions and performance. Error reports: reports of known software defects and problems.

[0055] Preferably, the detailed steps of data preparation and preprocessing are as follows: Determine where to obtain nuclear power software related data. Data sources may include: Nuclear power software sample data provided by nuclear power plant operators or developers. Nuclear power software safety-related documents issued by regulatory agencies or standardization organizations. Data related to nuclear power software contained in academic research papers and reports.

[0056] Preferably, data acquisition and screening: obtain relevant data from the determined data source. This may involve accessing databases, downloading files, etc. Perform preliminary screening on the acquired data to ensure that it is relevant to the nuclear power software review.

[0057] Preferably, data cleaning and preprocessing: After collecting data, data cleaning is required to ensure data quality and consistency. Including: Removing duplicate data: If there is identical or similar data in the data set, it needs to be removed to reduce duplicate information during training. Handling missing values: Detect and handle missing values ​​to ensure data integrity. Format unification: For data in different formats, convert them into a unified format so that the model can process them.

[0058] It is understandable that the acquired data is cleaned to remove irrelevant information, noise or erroneous data. The data is labeled to clarify the type, attributes and other information of each sample. The data format is converted to ensure that the data is suitable for subsequent model training.

[0059] Preferably, data annotation: the data is annotated according to the specific needs of the review. The annotation can be information about software features, safety issues, etc. For nuclear power software data, labeling or annotation is required so that the model can identify potential problems. This may include:

[0060] Label the code data: annotate the key structures, variables, functions, etc. of the code so that the model can understand the logical structure of the code.

[0061] Annotate text data: Mark technical terms, key information, etc. in the text to help the model understand the meaning of the text.

[0062] Preferably, feature extraction: extract key features from the annotated data. This can include code structure, terms in the text, logical relationships, etc. For different types of data (such as code, text, etc.), corresponding feature extraction is required. For code data, features such as the grammatical structure and variables of the code can be extracted; for text data, natural language processing technology can be used to extract features such as keywords and sentence structure.

[0063] Preferably, the data set is divided into training set, validation set and test set. The training set is used to train the model, the validation set is used to tune the model's hyperparameters and prevent overfitting, and the test set is used to evaluate the model's performance. Ensure that the data distribution and sample size of each set can reflect the characteristics of the review task.

[0064] Preferably, data storage and backup: store the prepared data, and it is recommended to establish a good data management system to ensure the security and reliability of the data.

[0065] Preferably, data quality control: regularly check the quality of the data set, and promptly update and fix any problems that may arise.

[0066] In summary, the above data processing process ensures the quality, accuracy and applicability of the data from data acquisition to data preparation through the above steps.

[0067] In the present invention, a generative model suitable for processing complex software data is selected, which is an improved pre-trained language representation model (Modification Bidirectional Encoder Representation from Transformers, M-BERT).

[0068] In one embodiment of the present invention, Figure 3 As shown, the model training process of the present invention adopts a generative training model as the core technology, and realizes a rapid review of the software by learning and analyzing a large amount of nuclear power software data. The generative training model can simulate the process of human learning and reasoning, and automatically extract features, identify potential problems, and make corresponding judgments through a large amount of training data. It can be known that the generative training model is a type of machine learning model. Its core idea is to simulate the distribution characteristics of the data by learning a large amount of training data, so as to generate new data with similar characteristics to the training data. In the nuclear power software review method, the present invention uses a generative training model to automatically learn the characteristics and problem patterns of the software.

[0069] It is understandable that M-BERT is a natural language processing model based on the transformer architecture, which can learn rich semantic information. During the review process, the M-BERT model can be used to deeply understand and analyze the text information of nuclear power software to identify potential problems.

[0070] It is understandable that M-BERT is an improved pre-trained natural language processing model. According to different processing scenarios, the M-BERT model can be divided into two branches: M-BERT-P (Modification Bidirectional Encoder Representation from Transformers for Photograph) and M-BERT-T. (Modification Bidirectional Encoder Representation from Transformers for Text) M-BERT-P and M-BERT-T are further evolutions of the M-BERT model.

[0071] In one embodiment of the present invention, for nuclear power software review materials such as pictures, tables, and codes, it is suitable to select the M-BERT-P model to perform specialized processing on pictures, tables, and codes. For nuclear power software review materials that process text data, the M-BERT-T model based on the transformer architecture can be selected.

[0072] In one embodiment of the present invention, the operation of the M-BERT model consists of four important links: generator, discriminator, pre-training module and fine-tuning module.

[0073] It can be understood that the generator is responsible for generating data for simulating nuclear power software. It is a deep neural network whose input is a random noise vector and output is a simulated nuclear power software data. The discriminator is responsible for determining whether the input data is real nuclear power software data or simulated data generated by the generator. It can be a binary classifier that outputs a probability value, indicating the probability that the input is real data. Momentum optimizer: The optimizer is an important component in the deep learning model, which is responsible for adjusting the weights of the model to minimize the loss function. The momentum optimizer uses stochastic gradient descent and introduces momentum terms to accelerate convergence, so that the model converges to the optimal solution during the training process. Pre-training module: The pre-training module consists of an optimizer and a training data set. The M-BERT model is pre-trained using large-scale text data to enable it to learn to understand the contextual relationship of the language. The data set used is the regulatory data set for model training. Fine-tuning module: The fine-tuning module takes the pre-processed nuclear power software review document as input, covering the generator and the discriminator. The pre-trained model is combined with data related to nuclear power software review, and the model is fine-tuned to adapt to the specific semantics of nuclear power software data, ensuring that the M-BERT model is significantly different from the traditional BERT model.

[0074] In one embodiment of the present invention, the selected M-BERT model is trained and fine-tuned based on the training dataset to obtain a specific M-BERT model.

[0075] Preferably, the generator and the discriminator are randomly initialized.

[0076] Preferably, the generator generates simulated nuclear power software data using random noise, receives a random noise vector as input, and generates a false sample, while obtaining a real sample from a real data set.

[0077] Preferably, the discriminator receives as input the generated fake samples and real samples, and then outputs the probability that they are real. The parameters of the discriminator are updated so that it can more accurately distinguish between real data and simulated data.

[0078] Preferably, the discriminator's loss function is adjusted based on how well it classifies real and fake samples. The goal is to minimize the probability that the discriminator makes a mistake.

[0079] Preferably, the parameters of the generator are updated through back propagation so that the simulated data generated by it is closer to the real nuclear power software data. The goal of the generator is to generate as many false samples as possible that are sufficient to deceive the discriminator. The generator learns by minimizing the probability that the generated samples are judged as false.

[0080] Preferably, the loss function of the generator is adjusted according to the classification of the generated samples by the discriminator. Then it enters the momentum optimizer optimization, with the goal of maximizing the probability of the generator successfully deceiving.

[0081] Repeated iterations: The above steps are performed alternately in multiple iterations, so that the performance of the generator and discriminator gradually improves.

[0082] In one embodiment of the present invention, during the process of model training, the input data of the model is processed based on the dynamic masking technology to achieve the optimal conditions for model training.

[0083] It is understandable that the dynamic masking technique allows the model to accept input sequences of different lengths, not just fixed lengths. This allows the model to maintain performance under different tasks and sentence lengths. Using dynamic masks to adjust the weights of the self-attention mechanism allows the model to better capture information on long sentences. Based on dynamic masks, you can automatically choose to truncate long sentences or pad short sentences, thereby alleviating the problem of information loss or padding redundancy.

[0084] It can be understood that in M-BERT, the embedding layer (E) of each representation of the input layer and the implicit feature scale (H) of the transformation layer are directly connected, which leads to the fact that the feature scale of the embedding layer and the implicit feature scale can be proportionally modeled, and the two are equivalent in special cases.

[0085] It is understandable that in order to map the input representation size, i.e., the vocabulary feature (V), to the hidden layer, a mapping matrix must be established, and the time complexity of matrix operations is O(V, H). When the vocabulary size increases, the output dimension of the embedding layer will also increase relatively synchronously, thereby significantly increasing the review workload, which is contrary to the vision of replacing manual review of nuclear power software with big data.

[0086] It is understandable that the reason why M-BERT is superior is mainly because the hidden layer can learn the relationship between "contexts", and the word vector matrix itself is relatively sparse (because the input is a point connection, only the part corresponding to the word index will be activated), so the dimension of the byte embedding layer matrix does not need to be consistent with the matrix dimension of the transmission layer and the hidden layer, and the dimension of the embedding layer matrix must be much smaller than the matrix dimension of the hidden layer.

[0087] When H=512,E=64,V=10000:

[0088] O(V,H)=10000×512=5120000;

[0089] O(V,E)+O(E,H)=10000×64+64×512=672768;

[0090] It can be seen that the difference between the two is nearly an order of magnitude.

[0091] Therefore, the M-BERT of the present invention is improved and optimized based on the BERT model. Compared with BERT, the M-BERT model performs specific training steps on the application scenarios of nuclear power software review during the training process, thereby achieving the purpose of fine-tuning the model; dynamic masking technology is adopted; and based on the word vector factorization technology, the next scene prediction is abandoned, thereby balancing the contradiction between response time and training amount, effectively avoiding duplication of training data, and improving the model operation efficiency.

[0092] In summary, the generative training model can automatically extract features and identify potential problems by learning a large amount of nuclear power software data. For example, for code data, the model can learn to identify common code structures, error patterns, etc. For text data, the model can understand the technical terms, logical relationships, etc. Through repeated training and optimization, the misjudgment rate of the model can be reduced. This can be achieved by adjusting model parameters, increasing training data, improving model architecture, etc. The generative training model has the ability to learn in real time and can continuously receive new data during operation, thereby updating the cognitive ability of the model. This enables the review system to adapt to the ever-changing nuclear power software environment. Through the comprehensive application of the above technical principles, the present invention realizes efficient review of nuclear power software, improves the accuracy and efficiency of the review, and provides important technical support for the safe and stable development of the nuclear power industry.

[0093] Furthermore, according to the usage, the M-BERT model trained based on the embodiment of the present invention has good applicability for graphics application scenarios and is mainly used in the following scenarios:

[0094] Scenario 1: Generate realistic radiation field simulation images to help evaluate the radiation environment in nuclear power plants, and synthesize environmental monitoring data to help evaluate the radiation level in the environment around nuclear power plants.

[0095] Scenario 2: Generate images or data simulating nuclear power plant accidents for training and evaluation of accident emergency plans.

[0096] Scenario 3: Generate a virtual nuclear power plant environment for operator training and simulation of operations under different scenarios.

[0097] Scenario 4: Generate simulation data under different fuel states to evaluate the stability and safety of reactor operation.

[0098] Scenario 5: Generate data distribution under normal conditions to help detect abnormal or fault conditions in nuclear power plant operations.

[0099] In one embodiment of the present invention, the regulatory data set for the supply optimizer to train the model has the following main sources, relying on the current laws and regulations of the People's Republic of China, China Nuclear Safety Regulations (HAF), China Nuclear Safety Guidelines (HAD) and the International Atomic Energy Agency (IAEA) to provide a series of regulatory documents and guidelines related to nuclear safety. Including but not limited to the following laws and regulations:

[0100] Such as: Law of the People's Republic of China on the Prevention and Control of Radioactive Pollution, Environmental Protection Law of the People's Republic of China, Environmental Impact Assessment Law of the People's Republic of China, Earthquake Disaster Prevention and Mitigation Law of the People's Republic of China, Atmospheric Pollution Prevention and Control Law of the People's Republic of China, Water Pollution Prevention and Control Law of the People's Republic of China, Marine Environment Protection Law of the People's Republic of China, Law of the People's Republic of China on the Prevention and Control of Environmental Pollution by Solid Waste, Occupational Disease Prevention and Control Law of the People's Republic of China, Road Traffic Safety Law of the People's Republic of China, Work Safety Law of the People's Republic of China, Administrative Regulations on Nuclear and Radiation Safety, Regulations on Safety Supervision and Management of Civilian Nuclear Facilities, Regulations on Emergency Management of Nuclear Accidents in Nuclear Power Plants, Regulations on Nuclear Material Control, Regulations on Supervision and Management of Civilian Nuclear Safety Equipment, Regulations on Safety Management of Transport of Radioactive Articles, Regulations on Safety and Protection of Radioactive Isotopes and Radiation Devices, Administrative Regulations on Nuclear and Radiation Safety, Regulations on Export Control of Nuclear Dual-Use Items and Related Technologies, and Regulations on Nuclear Export Control.

[0101] Furthermore, based on the model training, the present invention inputs the nuclear power software to be reviewed into the model system. The trained generative training model is used to analyze the software data and extract key features. The model is used to identify potential problems in the software, such as safety hazards, performance bottlenecks, etc. The review results are output in the form of a report, including detailed information such as problem description and location.

[0102] Furthermore, the review results are fed back to the nuclear power software development team, allowing them to make corresponding repairs and improvements. After the repairs, they are reviewed again to ensure that the problems have been properly resolved.

[0103] Furthermore, the training data is updated regularly and the model is optimized to ensure its adaptability under different software versions and technological changes.

[0104] According to the nuclear power software detection method based on the generative training model of the embodiment of the present invention, the generative training model is adopted to realize the automated review of nuclear power software, greatly improve the review efficiency and save human resources. It can deeply understand the characteristics of nuclear power software and identify potential problems. Compared with the traditional manual review method, it has higher accuracy. The demand for professional reviewers is reduced, costs are saved, and the reviewer's resources can be used for higher-level tasks. The generative training model has the ability to continuously learn and adapt to new environments, and can cope with the ever-changing nuclear power software environment. By timely discovering and correcting problems in the software, the quality of nuclear power software is improved and potential safety risks are reduced. Compared with the traditional manual review method, the automated review method can complete the review task in a shorter time. By discovering and solving potential problems, the security and stability of nuclear power software are enhanced, which helps to ensure the normal operation of nuclear power plants. This method is not only applicable to nuclear power software review, but can also be extended to software review in other fields, and has strong versatility.

[0105] In order to implement the above embodiment, Figure 4 As shown, this embodiment also provides a nuclear power software detection device 10 based on a generative training model, and the device 10 includes a data set acquisition module 100, a model construction module 200, a model training module 300 and a data detection module 400;

[0106] The data set acquisition module 100 is used to acquire a training data set and a test data set based on nuclear power software review data;

[0107] A model building module 200 is used to build a natural language processing model M-BERT suitable for nuclear power software review data; wherein M-BERT includes a fine-tuning module and a pre-training module;

[0108] The model training module 300 is used to train M-BERT using the training data set through the pre-training module to obtain a pre-trained M-BERT, and input the pre-processed nuclear power software review data into the pre-trained M-BERT through the fine-tuning module to perform model fine-tuning on the pre-trained M-BERT to obtain a fine-tuned M-BERT;

[0109] The data detection module 400 is used to input the test data set into the fine-tuned M-BERT to perform nuclear power software quality detection to obtain the nuclear power software quality detection result.

[0110] Furthermore, the data type of the input nuclear power software review data is determined. If the data type is a non-text data type, M-BERT-P is selected for data processing. If the data type is a text data type, M-BERT-T is selected for data processing.

[0111] According to the nuclear power software detection device based on the generative training model of the embodiment of the present invention, the generative training model is adopted to realize the automated review of nuclear power software, greatly improve the review efficiency and save human resources. It can deeply understand the characteristics of nuclear power software and identify potential problems, and has higher accuracy than the traditional manual review method. It reduces the demand for professional reviewers, saves costs, and can use the reviewer's resources for higher-level tasks. The generative training model has the ability to continuously learn and adapt to new environments, and can cope with the ever-changing nuclear power software environment. By timely discovering and correcting problems in the software, the quality of nuclear power software is improved and potential safety risks are reduced. Compared with the traditional manual review method, the automated review method can complete the review task in a shorter time. By discovering and solving potential problems, the safety and stability of nuclear power software are enhanced, which helps to ensure the normal operation of nuclear power plants. This method is not only applicable to nuclear power software review, but can also be extended to software review in other fields, and has strong versatility.

[0112] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0113] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

Claims

1. A nuclear power software detection method based on a generative training model, characterized in that: include: Obtain training and testing data sets based on nuclear power software review data; Constructing a natural language processing model M-BERT suitable for the nuclear power software review data; wherein the M-BERT includes a fine-tuning module and a pre-training module; The pre-training module uses the training data set to train the M-BERT to obtain a pre-trained M-BERT, and the fine-tuning module inputs the pre-processed nuclear power software review data into the pre-trained M-BERT to perform model fine-tuning on the pre-trained M-BERT to obtain a fine-tuned M-BERT; The test data set is input into the fine-tuned M-BERT to perform nuclear power software quality detection to obtain the nuclear power software quality detection result.

2. The method according to claim 1, characterized in that Determine the data type of the input nuclear power software review data. If the data type is a non-text data type, select M-BERT-P for data processing. If the data type is a text data type, select M-BERT-T for data processing.

3. The method according to claim 1, characterized in that Obtain a large amount of the nuclear power software review data, including at least source code, design documents, test data and error reports.

4. The method according to claim 3, characterized in that A large amount of nuclear power software review data is screened according to preset screening conditions to obtain nuclear power software review data that meets the model requirements.

5. The method according to claim 4, characterized in that Preprocessing the nuclear power software review data that meets the model requirements includes: Remove duplicate data, handle missing values, label data, and convert data formats.

6. The method according to claim 5, characterized in that The data annotation includes: Marking the code data of the non-text data type: marking key structures, variables and functions of the code; The text data of the text data type is annotated: technical terms and key information in the text are annotated.

7. The method according to claim 6, characterized in that The pre-training module at least includes an optimizer, and the fine-tuning module includes a generator and a discriminator; the method further includes: Generate fake samples based on random noise using a generator, and obtain real samples from real nuclear power software review data; Inputting the false sample and the real sample into a discriminator to output a classification probability of the real sample, updating the parameters of the discriminator and adjusting the loss function of the discriminator according to the classification probability to train the discriminator; Training the generator by updating the parameters of the generator through back propagation and adjusting the loss function of the generator according to the classification of the generated false samples by the discriminator; The fine-tuned M-BERT is obtained by minimizing the loss function through the optimizer to fine-tune the pre-trained generator and discriminator.

8. The method according to claim 7, characterized in that The method further comprises: Get a training dataset of text data type; Extracting input sequences of different lengths from a training data set of the text data type; The input sequences of different lengths are input into the M-BERT to automatically select to truncate long sentences or to fill short sentences in the input sequence based on a dynamic masking technique.

9. A nuclear power software detection device based on a generative training model, characterized in that: include: A data set acquisition module is used to acquire training data sets and test data sets based on nuclear power software review data; A model building module, used to build a natural language processing model M-BERT suitable for the nuclear power software review data; wherein the M-BERT includes a fine-tuning module and a pre-training module; A model training module, used to train the M-BERT using the training data set through the pre-training module to obtain a pre-trained M-BERT, and input the pre-processed nuclear power software review data into the pre-trained M-BERT through the fine-tuning module to perform model fine-tuning on the pre-trained M-BERT to obtain a fine-tuned M-BERT; The data detection module is used to input the test data set into the fine-tuned M-BERT to perform nuclear power software quality detection to obtain a nuclear power software quality detection result.

10. The device according to claim 9, characterized in that Determine the data type of the input nuclear power software review data. If the data type is a non-text data type, select M-BERT-P for data processing. If the data type is a text data type, select M-BERT-T for data processing.