Key nuclide prediction device based on multi-channel cavity convolution attention model
Through the key nuclide prediction device based on the multi-channel hollow convolutional attention model, the problems of inaccurate manual detection and aging of nuclear detectors in the nuclear power field are solved, and high-precision and rapid prediction of key nuclides are achieved.
Patent Information
- Application Number
- CN202510060436.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-15
AI Technical Summary
In the field of nuclear power, manual detection and analysis methods are inaccurate and time-consuming. Nuclear detectors are prone to aging in high temperature and humidity and high radiation environments, resulting in reduced accuracy of monitoring data and it is difficult to achieve high-precision and rapid prediction of key nuclides.
The key nuclide prediction device based on the multi-channel hollow convolutional attention model is adopted, including the sodium iodide spectrometer data acquisition module, database and host computer. The sodium iodide spectrometer data is trained and predicted through the multi-channel hollow convolutional attention prediction model, and the multi-channel hollow convolutional attention prediction model is automatically learned and updated in real time.
High-precision and rapid prediction of key nuclides are achieved, reducing the influence of human factors, and improving the accuracy and automation of monitoring data.
Smart Images

Figure CN120069169A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of critical nuclide prediction, and particularly relates to a critical nuclide prediction device based on a multi-channel dilated convolutional attention model. Background Art
[0002] With the rapid development of the Chinese national economy, the large demand for energy in China has made the vigorous development of nuclear energy an inevitable requirement. In the field of nuclear power, nuclear safety is of utmost importance in nuclear power operation. All along, the radiation environment monitoring automatic stations in China have adopted the method of manual detection and analysis. This method is affected by expert human factors. At the same time, the manual method is very inaccurate and requires a long prediction time. Therefore, accurately predicting critical nuclides with high precision is an urgent problem and challenge to be solved at present.
[0003] In recent years, intelligent monitoring in the industrial field has begun to receive attention. Nuclear detectors are important devices in radiation safety protection monitoring and nuclear power plant safety monitoring. They work in an environment of high temperature, high humidity and high radiation intensity for a long time, which easily leads to phenomena such as aging, decline in working performance or partial functional failures, and ultimately reduces the accuracy of measurement data in the monitoring site. Instrumentation technology is widely used in industrial processes. The gradually developing science and technology has also improved the automation level of industrial process production management. The instrumentation measurement data of the actual industrial process is stored in the database through on-site equipment. A large amount of data contains a lot of process information, which can be used for industrial process monitoring. Due to the huge amount of historical data accumulated in the nuclear power field, it makes the application of machine learning and deep learning methods in this field possible. Therefore, it is of great significance to realize the rapid and accurate prediction of the ionization chamber dose rate of critical nuclides based on the data of sodium iodide spectrometers, so as to carry out predictive maintenance of the nuclear power process. Summary of the Invention
[0004] The purpose of the present invention is to provide a critical nuclide prediction device based on a multi-channel dilated convolutional attention model in view of the deficiencies of the prior art.
[0005] The purpose of the present invention is achieved through the following technical solutions: A critical nuclide prediction device based on a multi-channel dilated convolutional attention model, the device is successively composed of a sodium iodide spectrometer data acquisition module, a database and a host computer; the host computer includes a data division module, a multi-channel dilated convolutional attention prediction model modeling module, a multi-channel dilated convolutional attention prediction module and a prediction result output module;
[0006] The sodium iodide spectrometer data acquisition module is used to collect sodium iodide spectrometer data E and upload it to the database;
[0007] The database is used to save the sodium iodide spectrometer data E collected by the sodium iodide spectrometer data acquisition module and upload it to the host computer;
[0008] The host computer is used to train the multi-channel dilated convolution attention prediction model with the sodium iodide spectrometer data E, and an optimized multi-channel dilated convolution attention prediction model is obtained through training. Subsequently, the optimized multi-channel dilated convolution attention prediction model is used to predict the sodium iodide spectrometer data to be measured to obtain a predicted value.
[0009] Further, the sodium iodide spectrometer data E is E = {e 1 , e 2 , …, e g , …, e G}, where eg is the g-th spectrometer vector in the sodium iodide spectrometer data E, and the dimension of each spectrometer vector e g is d, G is the length of the sodium iodide spectrometer data E, and g = 1, 2, …, g, …, G.
[0010] Further, the host computer trains the multi-channel dilated convolution attention prediction model with the sodium iodide spectrometer data E, and an optimized multi-channel dilated convolution attention prediction model is obtained through training. Subsequently, the optimized multi-channel dilated convolution attention prediction model is used to predict the sodium iodide spectrometer data to be measured to obtain a predicted value, specifically:
[0011] (a.1) First, the uploaded sodium iodide spectrometer data E is divided into a training set E 1 , a validation set E 2 , and a test set E 3 by the data division module in the host computer, and the training set E 1 and the validation set E 2 are uploaded to the multi-channel dilated convolution attention prediction model modeling module, and the test set E 3 is uploaded to the multi-channel dilated convolution attention prediction module;
[0012] (a.2) Subsequently, the multi-channel dilated convolution attention prediction model modeling module trains the multi-channel dilated convolution attention prediction model with the training set E 1 , and an optimized multi-channel dilated convolution attention prediction model is obtained through training and uploaded to the multi-channel dilated convolution attention prediction module;
[0013] (a.3) The multi-channel dilated convolution attention prediction module predicts the sodium iodide spectrometer data to be measured through the optimized multi-channel dilated convolution attention prediction model to obtain the predicted value of the ionization chamber dose rate of the key nuclide corresponding to the data;
[0014] (a.4) The predicted value of the ionization chamber dose rate of the key nuclide corresponding to the data is output through the prediction result output module.
[0015] Further, the step (a.2) specifically includes the following sub-steps:
[0016] (a.2.1) The multi-channel dilated convolution attention prediction model includes dilated convolution feature extraction channels with n different dilation rates, a bidirectional recurrent feature extraction module, an attention mechanism module, and a dropout module;
[0017] (a.2.2) First, use the dilated convolution feature extraction channels with n different dilation rates to extract n output vectors from the training set E 1 respectively: X 1 , X 2 , …, X s , …, X n , where X s represents the output vector extracted by the s-th dilated convolution feature extraction channel, s = 1, 2, …, s, …, n;
[0018] (a.2.3) Subsequently, any one of the output vectors X s passes through the bidirectional recurrent feature extraction module with l hidden neurons to obtain the output B s ;
[0019] (a.2.4) Subsequently, the output B s passes through the attention mechanism module to obtain the output A s ;
[0020] (a.2.5) The dropout module repeats steps (a.2.3) - (a.2.4) for each output vector X s to obtain n outputs: A 1 , A 2 , …, A s , …, A n ;
[0021] Subsequently, the n outputs A 1 , A 2 , …, A s , …, A n are concatenated to obtain A multiple : A multiple = concat(A 1 , A 2 , …, A s , …, A n );
[0022] Some units are randomly discarded from the network temporarily with a certain probability, and the output D multiple after dropout is:
[0023] D multiple = DropOut(Amultiple , dr);
[0024] Among them, dr represents the probability of discarding a neural network unit;
[0025] (a.2.6) And the training set E 1 Through the bidirectional cyclic feature extraction module, the attention mechanism module, and the dropout module, the output D is obtained original ;
[0026] (a.2.7) The output D multiple and the output D original are integrated to obtain the output F: F = D multiple + D original ;
[0027] (a.2.8) Finally, the output F and the learnable transformation weight W O are multiplied to obtain the prediction vector of the training set E 1
[0028] (a.2.9) Through the training set E 1 and the prediction vector A loss function is constructed, and the multi-channel dilated convolutional attention prediction model is trained through the loss function. The optimized multi-channel dilated convolutional attention prediction model is obtained through training and uploaded to the multi-channel dilated convolutional attention prediction module.
[0029] Furthermore, the extraction process of the output vector X s is specifically as follows:
[0030] The s-th dilated convolutional extraction feature channel first uses a dilated convolution with a dilation rate of r s to extract features from the training set E 1 , and uses the hyperbolic linear unit as the activation function. This process can be expressed as:
[0031]
[0032] Among them, X s represents the output vector; W c represents the weight of the dilated convolution kernel; ξ(·) represents the hyperbolic linear unit; α is the hyperparameter in the hyperbolic linear unit;
[0033] The training set E 1 is Among them, is the t-th spectrometer vector in the training set E 1 , k is the length of the training set E 1 , t = 1, 2,..., t,..., k; the output vector X s is Among them, represents the output vector X s of the t-th dilated convolution module in
[0034] Furthermore, the step (a.2.3) is specifically:
[0035] The bidirectional recurrent feature extraction module includes a forward unidirectional recurrent feature extraction module and a backward unidirectional recurrent feature extraction module. The forward unidirectional recurrent feature extraction module and the backward unidirectional recurrent feature extraction module each consist of a forget gate, an input gate, and an output gate;
[0036] The output f of the forget gate t can be expressed as:
[0037] f t = σ(W f [h t-1 , x t + b f );
[0038] Among them, f t represents the output of the t-th forget gate, W f represents the weight of the forget gate, h t-1 represents the output of the (t - 1)-th unidirectional recurrent feature extraction module unit, x t represents the input of the t-th unidirectional recurrent feature extraction module unit, b f represents the bias term of the forget gate, and σ(·) represents the Sigmoid function; the forget gate can determine the information to be forgotten through h t-1 and x t ;
[0039] The output i of the input gate t can be expressed as:
[0040] i t = σ(W i [h t-1 , x t + b i );
[0041] Among them, i t represents the output of the t-th input gate, W i represents the weight of the input gate, and b i represents the bias term of the input gate;
[0042] The input gate simultaneously updates the temporary cell state
[0043]
[0044] Among them, Represents a temporary cell state, W c Represents the cell state update weight, b c Represents the bias term for cell state update, C t-1 Represents the (t - 1)-th cell state; the input gate determines which information needs to be updated through h t-1 and x t and then obtains the new cell state based on the outputs of the forget gate and the input gate;
[0045] The output of the output gate can be expressed as:
[0046] o t = σ(W o [h t-1 , x t +b o );
[0047] where, o t represents the output of the output gate, W o represents the weight of the output gate, b o represents the bias term of the output gate; the output gate can obtain the output judgment condition through h t-1 and x t The output h t of the single-directional cyclic feature extraction module unit is:
[0048] h t = o t * tanh(C t );
[0049] The output B s of the bidirectional cyclic feature extraction module is expressed as:
[0050]
[0051] where represents the output of the forward single-directional cyclic feature extraction module unit for the output vector X s , represents the output of the backward single-directional cyclic feature extraction module unit for the output vector X s .
[0052] Furthermore, the step (a.2.4) is specifically:
[0053] The attention mechanism module assigns different weights to the output A s to obtain three matrices Q s , K s and V s ; Subsequently, according to the matrices Q s , K s and Vs , obtain output A s , and the calculation formula is as follows:
[0054] Q s = B s W Q ;
[0055] K s = B s W K ;
[0056] V s = B s W V ;
[0057]
[0058] Among them, W Q , W K and W V represent different weight matrices, d B represents the dimension of B s , and softmax(·) converts the input value into a probability distribution ranging from [0,1] and summing to 1. The calculation formula is as follows:
[0059]
[0060] Among them, z i represents the i-th input value, and C represents the number of input nodes.
[0061] The beneficial effects of the present invention are:
[0062] 1) The features of sodium iodide spectrometer data can be efficiently extracted through the multi-channel dilated convolution method;
[0063] 2) The multi-channel dilated convolution attention prediction model can be updated in real time by using the newly input sodium iodide spectrometer data;
[0064] 3) The key nuclide prediction device based on the multi-channel dilated convolution attention model can automatically learn according to the training data, has strong intelligence, and is less affected by human factors. Description of the Drawings
[0065] Figure 1 is a structural diagram of a key nuclide prediction device based on a multi-channel dilated convolution attention model;
[0066] Figure 2 is a structural diagram of the host computer;
[0067] In the figure, 1 - Sodium iodide spectrometer data acquisition module; 2 - Database; 3 - Host computer; 4 - Data division module; 5 - Multi-channel dilated convolutional attention prediction model building module; 6 - Multi-channel dilated convolutional attention prediction module; 7 - Prediction result output module. Detailed implementation mode
[0068] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0069] Embodiment 1
[0070] As Figure 1 shown, the present invention provides a key nuclide prediction device based on a multi-channel dilated convolutional attention model. The device is successively composed of a sodium iodide spectrometer data acquisition module 1, a database 2 and a host computer 3. The host computer 3 includes a data division module 4, a multi-channel dilated convolutional attention prediction model building module 5, a multi-channel dilated convolutional attention prediction module 6 and a prediction result output module 7.
[0071] The sodium iodide spectrometer data acquisition module 1 is used to collect sodium iodide spectrometer data E and upload it to the database 2.
[0072] The sodium iodide spectrometer data E is E = {e 1 , e 2 , …, e g , …, e G}, where eg is the g-th spectrometer vector in the sodium iodide spectrometer data E, and the dimension of each spectrometer vector e g is d, G is the length of the sodium iodide spectrometer data E, and g = 1, 2, …, g, …, G.
[0073] The database 2 is used to save the sodium iodide spectrometer data E collected by the sodium iodide spectrometer data acquisition module 1 and upload it to the host computer 3.
[0074] The host computer 3 is used to train the multi-channel dilated convolutional attention prediction model with the sodium iodide spectrometer data E. The features of the sodium iodide spectrometer data can be efficiently extracted by the multi-channel dilated convolutional method, and an optimized multi-channel dilated convolutional attention prediction model is obtained through training. Subsequently, the optimized multi-channel dilated convolutional attention prediction model is used to predict the sodium iodide spectrometer data to be measured to obtain a predicted value. Specifically:
[0075] (a.1) First, divide the uploaded sodium iodide spectrometer data E into a training set E, a validation set E, and a test set E according to the ratio of 7:1:2 through the data division module in the host computer 3, and upload the training set E and the validation set E to the multi-channel dilated convolutional attention prediction model modeling module 5, and upload the test set E to the multi-channel dilated convolutional attention prediction module 6. 1 and validation set E 2 and test set E 3 , and upload the training set E 1 and validation set E 2 to the multi-channel dilated convolutional attention prediction model modeling module 5, and upload the test set E 3 to the multi-channel dilated convolutional attention prediction module 6.
[0076] (a.2) Subsequently, the multi-channel dilated convolutional attention prediction model modeling module 5 trains the multi-channel dilated convolutional attention prediction model through the training set E, and uploads the optimized multi-channel dilated convolutional attention prediction model to the multi-channel dilated convolutional attention prediction module 6. 1 The above step (a.2) specifically includes the following sub-steps:
[0077] The above step (a.2) specifically includes the following sub-steps:
[0078] (a.2.1) The multi-channel dilated convolutional attention prediction model includes dilated convolutional feature extraction channels with n different dilation rates, a bidirectional recurrent feature extraction module, an attention mechanism module, and a dropout module.
[0079] (a.2.2) First, use the dilated convolutional feature extraction channels with n different dilation rates to extract n output vectors: X 1 , X 2 , …, X s , …, X n from the training set E, where X s represents the output vector extracted by the s-th dilated convolutional feature extraction channel, and s = 1, 2, …, s, …, n. 1 extract to n output vectors: X 1 ,X 2 ,…,X s ,…,X n , where, X s represents the output vector extracted by the s-th dilated convolutional feature extraction channel, s = 1, 2, …, s, …, n.
[0080] The extraction process of the output vector X s is specifically as follows:
[0081] The s-th dilated convolutional feature extraction channel first uses a dilated convolution with a dilation rate of r s to extract features from the training set E 1 and uses a hyperbolic linear unit as the activation function. This process can be expressed as:
[0082]
[0083] where, X s represents the output vector; W c represents the weight of the dilated convolution kernel; ξ(·) represents the hyperbolic linear unit; α is the hyperparameter in the hyperbolic linear unit;
[0084] The training set E 1 is wherein, is the t-th spectrometer vector in the training set E 1 , k is the length of the training set E 1 , t = 1, 2, …, t, …, k; the output vector X s is wherein, represents the output of the t-th dilated convolution module in the output vector X s .
[0085] (a.2.3) Subsequently, any one of the output vectors X s passes through a bidirectional recurrent feature extraction module with l hidden neurons to obtain the output B s .
[0086] The specific steps of (a.2.3) are as follows:
[0087] The bidirectional recurrent feature extraction module includes a forward unidirectional recurrent feature extraction module and a backward unidirectional recurrent feature extraction module. The forward unidirectional recurrent feature extraction module and the backward unidirectional recurrent feature extraction module are each composed of a forget gate, an input gate, and an output gate, which can automatically learn according to the training data, have strong intelligence, and are less affected by human factors;
[0088] The output f of the forget gate t can be expressed as:
[0089] f t = σ(W f [h t-1 , x t + b f );
[0090] wherein, f t represents the output of the t-th forget gate, W f represents the weight of the forget gate, h t-1 represents the output of the (t - 1)-th unidirectional recurrent feature extraction module unit, x t represents the input of the t-th unidirectional recurrent feature extraction module unit, b f represents the bias term of the forget gate, and σ(·) represents the Sigmoid function; the forget gate can determine the information to be forgotten through h t-1 and x t ;
[0091] The output i of the input gate t can be expressed as:
[0092] i t = σ(W i[h t-1 ,x t +b i );
[0093] Among them, i t represents the output of the t-th input gate, W i represents the weight of the input gate, b i represents the bias term of the input gate;
[0094] The input gate updates the temporary cell state simultaneously
[0095]
[0096] Among them, represents the temporary cell state, W c represents the cell state update weight, b c represents the bias term of the cell state update, C t-1 represents the (t - 1)-th cell state; The input gate can determine which information needs to be updated through the input gate based on h t-1 and x t , and then obtain the new cell state according to the output of the forget gate and the output of the input gate;
[0097] The output of the output gate can be expressed as:
[0098] o t = σ(W o [h t-1 ,x t +b o );
[0099] Among them, o t represents the output of the output gate, W o represents the weight of the output gate, b o represents the bias term of the output gate; The output gate can obtain the output judgment condition through h t-1 and x t , and the output h t of the unidirectional cyclic feature extraction module unit is:
[0100] h t = o t * tanh(C t );
[0101] The output B s of the bidirectional cyclic feature extraction module is expressed as:
[0102]
[0103] Among them Indicates the output of the forward unidirectional cyclic feature extraction module unit for the output vector X s of Indicates the output of the backward unidirectional cyclic feature extraction module unit for the output vector X s of
[0104] (a.2.4) Subsequently, output B s passes through the attention mechanism module to obtain output A s .
[0105] The specific steps of (a.2.4) are as follows:
[0106] The attention mechanism module assigns different weights to output A s to obtain three matrices Q s , K s and V s ; Subsequently, based on matrices Q s , K s and V s , output A s is obtained, and the calculation formula is as follows:
[0107] Q s = B s W Q ;
[0108] K s = B s W K ;
[0109] V s = B s W V ;
[0110]
[0111] where W Q , W K and W V represent different weight matrices, d B represents the dimension of B s , and softmax(·) converts the input value into a probability distribution ranging from [0,1] and summing to 1. The calculation formula is as follows:
[0112]
[0113] where z i represents the i-th input value, and C represents the number of input nodes.
[0114] (a.2.5) The reverse dropout module applies dropout to each output vector X sRepeat steps (a.2.3) - (a.2.4) to obtain n outputs: A 1 , A 2 , …, A s , …, A n .
[0115] Subsequently, concatenate the n outputs A 1 , A 2 , …, A s , …, A n to obtain A multiple : A multiple = concat(A 1 , A 2 , …, A s , …, A n ).
[0116] Randomly discard some units from the network temporarily with a certain probability, and the output D multiple after inverse random inactivation is:
[0117] D multiple = DropOut(A multiple , dr);
[0118] where dr represents the probability of discarding neural network units.
[0119] (a.2.6) Pass the training set E 1 through the bidirectional cyclic feature extraction module, the attention mechanism module, and the inverse random inactivation module to obtain the output D original .
[0120] (a.2.7) Integrate the output D multiple and the output D original to obtain the output F: F = D multiple + D original .
[0121] (a.2.8) Finally, multiply the output F by the learnable transformation weight W O to obtain the prediction vector of the training set E 1
[0122] (a.2.9) Construct a loss function based on the training set E 1 and the prediction vector , and train the multi-channel dilated convolutional attention prediction model through the loss function. After training, obtain the optimized multi-channel dilated convolutional attention prediction model and upload it to the multi-channel dilated convolutional attention prediction module 6. Using the newly input sodium iodide spectrometer data, the multi-channel dilated convolutional attention prediction model can be updated in real time.
[0123] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A key nuclide prediction device based on a multi-channel dilated convolutional attention model, characterized in that: The device is composed of a sodium iodide spectrometer data acquisition module, a database and a host computer in sequence; the host computer includes a data partitioning module, a multi-channel dilated convolution attention prediction model modeling module, a multi-channel dilated convolution attention prediction module and a prediction result output module; The sodium iodide spectrometer data acquisition module is used to acquire sodium iodide spectrometer data E and upload it to a database; The database is used to store the sodium iodide spectrometer data E collected by the sodium iodide spectrometer data acquisition module and upload it to the host computer; The host computer is used to train the multi-channel dilated convolution attention prediction model with the sodium iodide spectrometer data E to obtain an optimized multi-channel dilated convolution attention prediction model, and then predict the sodium iodide spectrometer data to be tested according to the optimized multi-channel dilated convolution attention prediction model to obtain a predicted value.
2. The key nuclide prediction device based on the multi-channel dilated convolutional attention model according to claim 1 is characterized in that: The sodium iodide spectrometer data E is E={e1, e2, ..., e g ,…,e G }, where eg is the g-th spectrometer vector in the sodium iodide spectrometer data E, and each spectrometer vector e g The dimension of is d, G is the length of the sodium iodide spectrometer data E, g = 1, 2,…, g,…, G.
3. The key nuclide prediction device based on the multi-channel dilated convolutional attention model according to claim 2 is characterized in that: The host computer trains the multi-channel dilated convolution attention prediction model with the sodium iodide spectrometer data E to obtain an optimized multi-channel dilated convolution attention prediction model, and then predicts the sodium iodide spectrometer data to be tested according to the optimized multi-channel dilated convolution attention prediction model to obtain a predicted value, specifically: (a.1) First, the uploaded sodium iodide spectrometer data E is divided into a training set E1, a validation set E2 and a test set E3 through the data partitioning module in the host computer, and the training set E1 and the validation set E2 are uploaded to the multi-channel dilated convolution attention prediction model modeling module, and the test set E3 is uploaded to the multi-channel dilated convolution attention prediction module; (a.2) Then, the multi-channel dilated convolution attention prediction model modeling module trains the multi-channel dilated convolution attention prediction model through the training set E1, obtains the optimized multi-channel dilated convolution attention prediction model after training, and uploads it to the multi-channel dilated convolution attention prediction module; (a.3) The multi-channel dilated convolution attention prediction module predicts the sodium iodide spectrometer data to be measured through the optimized multi-channel dilated convolution attention prediction model to obtain the predicted value of the key nuclide ionization chamber dose rate corresponding to the data; (a.4) The predicted value of the key nuclide ionization chamber dose rate corresponding to the data is output through the prediction result output module.
4. The key nuclide prediction device based on the multi-channel dilated convolutional attention model according to claim 3 is characterized in that: The step (a.2) specifically includes the following sub-steps: (a.2.1) The multi-channel dilated convolutional attention prediction model includes n dilated convolutional feature extraction channels with different dilation rates, a bidirectional recurrent feature extraction module, an attention mechanism module, and a reverse random dropout module; (a.2.2) First, we use n dilated convolutions with different dilation rates to extract feature channels and extract n output vectors from the training set E1: X1, X2, …, X s ,…,X n , where X s It means the output vector is extracted by the s-th hole convolution feature channel, s = 1, 2, ..., s, ..., n; (a.2.3) Then any output vector X s Through the bidirectional cyclic feature extraction module with l hidden layer neurons, the output B is obtained. s ; (a.2.4) Then output B s Through the attention mechanism module, we get the output A s ; (a.2.5) The reverse random dropout module converts each output vector X s Repeat steps (a.2.3)-(a.2.4) to get n outputs: A1, A2, …, A s ,…,A n ; Then the n outputs A1, A2, ..., A s ,…,A n Perform vector concatenation to obtain A multiple : A multiple =concat(A1,A2,…,A s ,…,A n ); According to a certain probability, some units are randomly discarded from the network temporarily, and the output D after reverse random inactivation is multiple for: D multiple =DropOut(A multiple ,dr); Among them, dr represents the probability of discarding a neural network unit; (a.2.6) The training set E1 is passed through the bidirectional loop feature extraction module, the attention mechanism module and the reverse random inactivation module to obtain the output D original ; (a.2.7) will output D multiple and output D original Integrate and get the output F: F = D multiple +D original ; (a.2.8) Finally, the output F and the learnable transformation weight W O Multiply them to get the prediction vector of training set E1 (a.2.9) Through the training set E1 and the prediction vector Construct a loss function and use it to train the multi-channel dilated convolutional attention prediction model. The optimized multi-channel dilated convolutional attention prediction model is trained and uploaded to the multi-channel dilated convolutional attention prediction module.
5. The key nuclide prediction device based on the multi-channel dilated convolutional attention model according to claim 4 is characterized in that: The output vector X s The specific extraction process is as follows: The sth dilated convolution extracts the feature channel by first using a dilation rate of r s The hole convolution extracts features from the training set E1 and uses the hyperbolic unit as the activation function. The process can be expressed as: Among them, X s represents the output vector; W c represents the weight of the dilated convolution kernel; ξ(·) represents the hyperbolic unit; α is the hyperparameter in the hyperbolic unit; The training set E1 is in, is the t-th spectrometer vector in the training set E1, k is the length of the training set E1, t = 1, 2, ..., t, ..., k; the output vector X s for in, Represents the output vector X s The output of the t-th atrous convolutional module in .
6. The key nuclide prediction device based on the multi-channel dilated convolutional attention model according to claim 5 is characterized in that: The step (a.2.3) is specifically: The bidirectional cycle feature extraction module includes a forward unidirectional cycle feature extraction module and a backward unidirectional cycle feature extraction module, wherein the forward unidirectional cycle feature extraction module and the backward unidirectional cycle feature extraction module are respectively composed of a forget gate, an input gate and an output gate; The output f of the forget gate t It can be expressed as: f t =σ(W f [h t-1 ,x t ]+b f ); Among them, f t represents the output of the tth forget gate, W f represents the weight of the forget gate, h t-1 represents the output of the t-1th unidirectional loop feature extraction module unit, x t represents the input of the t-th unidirectional loop feature extraction module unit, b f represents the bias term of the forget gate, σ(·) represents the Sigmoid function; the forget gate is passed through h t-1 and x t You can determine what information needs to be forgotten; The output of the input gate i t It can be expressed as: i t =σ(W i [h t-1 ,x t ]+b i ); Among them, i t represents the output of the tth input gate, W i represents the weight of the input gate, b i represents the bias term of the input gate; The input gate also updates the temporary cell state in, represents the temporary cell state, W c represents the cell state update weight, b c represents the bias term for updating the cell state, C t-1 Represents the state of the t-1th cell; the input gate is passed through h t-1 and x t It can determine which information needs to be updated through the input gate, and then get the new cell state based on the output of the forget gate and the output of the input gate; The output of the output gate can be expressed as: the t =σ(W o [h t-1 ,x t ]+b o ); Among them, t represents the output of the output gate, W o represents the weight of the output gate, b o Represents the bias term of the output gate; the output gate can be t-1 and x t Obtain the output judgment condition, the output h of the one-way loop feature extraction module unit t for: h t =o t *tanh(C t ); Output B of the bidirectional loop feature extraction module s It is expressed as: in Represents the output vector X of the forward unidirectional loop feature extraction module unit pair s The output, Represents the output vector X of the backward unidirectional loop feature extraction module unit pair s Output.
7. The key nuclide prediction device based on the multi-channel dilated convolutional attention model according to claim 6 is characterized in that: The step (a.2.4) is specifically: The attention mechanism module outputs A s Assign different weights and get three matrices Q s , K s and V s ; Then according to the matrix Q s , K s and V s , and get the output A s , the calculation formula is as follows: Q s =B s W Q ; K s =B s W K ; V s =B s W V ; Among them, W Q , W K and W V Represents different weight matrices, d B Indicates B s , softmax(·) converts the input value into a probability distribution in the range [0,1] and with a sum of 1. The calculation formula is as follows: where z i represents the i-th input value, and C represents the number of input nodes.
Citation Information
Patent Citations
Mineral spectrum classification method based on dilated convolutional neural network
CN113420795A
Road defect detection method based on deep learning
CN118918551A
Method for predicting remaining useful life of railway train bearing based on can-lstm
US20230153608A1