Nuclide identification method based on HHT band energy characteristics and convolutional neural network
Through the HHT band energy characteristics and convolutional neural network method, the problems of low accuracy and poor adaptability of nuclide identification in low counting rates and strong interference in the existing technology are solved, and fast and accurate nuclide identification is achieved, with an average recognition rate of 99.93%.
Patent Information
- Application Number
- CN202210283309.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-22
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2042-03-22
AI Technical Summary
Existing nuclide identification methods based on energy spectrum data have low recognition accuracy and poor adaptability under low counting rates and strong interference conditions, making it difficult to achieve fast and accurate nuclide identification in a short time.
A method based on HHT frequency band energy characteristics and convolutional neural network is adopted to eliminate noise through discrete wavelet transform, and the time-frequency characteristics of nuclear pulse signals are extracted using EMD decomposition and Hilbert transform. Then, a one-dimensional convolutional neural network is combined for nuclide identification, and the complete information of nuclear pulse signals is used for feature extraction and classification.
It achieves fast and accurate nuclide identification under time signals of arbitrary length, improves the recognition speed and accuracy, enhances the adaptability of the method, and achieves an average recognition rate of 99.93%.
Smart Images

Figure CN115034254B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to nuclide identification, in particular to a nuclide identification method based on HHT band energy characteristics and convolutional neural network. Background Art
[0002] It's well known that radionuclides play a vital role in national production, life, and economic development. Therefore, their identification is particularly crucial in the field of nuclear safety. Rapid radionuclide identification requires a qualitative analysis within a short timeframe (a few seconds), determining the type of radionuclide or the presence of a particular nuclide. Traditional radionuclide identification involves analyzing the energy spectrum of radionuclides and identifying nuclides based on the matching degree of characteristic peaks. However, the statistical properties of energy spectrum data indicate that, under conditions of low count rates and strong interference, identification time and count rate are significantly conflicting. Furthermore, factors such as energy resolution and background noise make energy spectrum analysis difficult, resulting in low nuclide identification accuracy. Therefore, new approaches are needed to address the limitations of using energy spectrum data for data analysis. In recent years, several researchers have proposed nuclide identification methods based on energy spectrum data, using algorithms such as artificial neural networks, machine learning, and convolutional neural networks. These methods have demonstrated the ability to identify nuclides by training neural networks. However, these methods remain constrained by the concept of energy spectrum and statistical laws, failing to significantly improve system decision time.
[0003] In 2009, Candy et al. at the Los Alamos National Laboratory in the United States proposed a nuclide identification method based on sequential Bayesian data analysis. This method uses a Bayesian algorithm to process acquired nuclear pulse signals, discriminating between event sequences based on the amplitude and temporal probability distribution of the nuclear pulses, and thus identifying nuclides. This method can identify nuclides even with short measurement times and low photon counts, demonstrating the feasibility of directly analyzing nuclear pulse signals for nuclide identification. Domestic researchers have also conducted further research and application of sequential Bayesian analysis based on this approach. However, this method only utilizes the amplitude and temporal information of the nuclear pulse signal, lacking an examination of the complete nuclear signal. Furthermore, changes in measurement conditions necessitate adjustment of analysis parameters, resulting in poor adaptability. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method that utilizes the complete information of a nuclear pulse signal sequence, is capable of processing nonlinear and non-stationary signals, and after time-frequency conversion of the signal, can process and identify time signals of any length without constantly adjusting analysis parameters, thereby improving the adaptability of the nuclide identification method; it can avoid the disadvantage of previous pulse-based nuclide identification that only uses partial signal features, and secondly, it has a fast recognition speed and high accuracy. The nuclide identification method based on HHT band energy features and convolutional neural network.
[0005] The technical solution adopted by the present invention to solve the technical problem is: a nuclide identification method based on HHT band energy characteristics and convolutional neural network, comprising the following steps:
[0006] S1. Collect nuclear pulse signals;
[0007] S2, eliminating the kernel pulse signal noise by discrete wavelet transform, comprising the following steps:
[0008] S21, select wavelet and wavelet decomposition level, calculate the wavelet decomposition of the original signal to the nth level; use the Coiflet (coifN) wavelet system;
[0009] S22. Select a threshold value for each layer of high-frequency coefficients and correct the high-frequency coefficients. Use an improved threshold function to correct the high-frequency coefficients. The improved threshold function is as follows:
[0010]
[0011] thr=max(x j )(2)
[0012]
[0013] Where thr is the threshold; k is the empirical coefficient, 0≤k≤1, when k=0 it is equivalent to the hard threshold function, when k=1 it is equivalent to the soft threshold function; x jt and η jt are the t-th high-frequency coefficients of the j-th layer before and after correction respectively; sign is the sign function;
[0014] S23, performing wavelet reconstruction of the signal based on the low-frequency coefficients of the nth layer and the corrected high-frequency coefficients from the 1st layer to the nth layer;
[0015] S3, feature extraction based on HHT;
[0016] For the nuclear pulse signal sequence x(t), it can be decomposed into a finite number of intrinsic mode functions IMF through EMD, that is, the original signal is the superposition of these intrinsic mode functions:
[0017]
[0018] Where, IMF refers to the intrinsic mode function; n is the order of the IMF component, c i (t) is the i-th order IMF component, r n (t) is the trend term, which represents the average trend or mean of the signal;
[0019] The steps of the EMD decomposition algorithm are as follows:
[0020] S31, initialization, set r1(t)=x(t), i=1, k=0;
[0021] S32, obtaining the n-th order intrinsic mode function, comprising the following steps:
[0022] S321, initialization, set h1(t) = r1(t);
[0023] S322. Find all the maximum and minimum points of hk(t);
[0024] S323, fitting the maximum point and the minimum point respectively by using an interpolation function to obtain the upper and lower envelopes: e+(t) and e-(t);
[0025] S324. Calculate the mean of the upper and lower envelopes mk(t)=(e+(t)+e-(t)) / 2;
[0026] S325. Calculate hk+1(t)=hk(t)-mk(t);
[0027] S326, judgment: 0.2≤ε≤0.3, if it holds, then c i (t) = h k (t), otherwise, set k=k+1 and return to S322;
[0028] S33, rk+1(t)=rk(t)-ck+1(t), judging whether the remainder is a monotonic function or a constant, if so, the decomposition process ends;
[0029] After S34 and x(t) are decomposed by EMD, n intrinsic mode functions (IMFs) with frequencies from high to low are obtained. The kernel signal after EMD decomposition is then subjected to Hilbert transform for feature extraction. The process is as follows:
[0030] S341, performing EMD decomposition on the denoised nuclear pulse signal to obtain a series of IMFs;
[0031] S342, performing Hilbert transform on each IMF component to obtain the instantaneous frequency, amplitude function and corresponding Hilbert spectrum of each IMF;
[0032] S343. Summarize the Hilbert spectra of all IMFs to obtain the Hilbert spectrum H(w,t) of the original random kernel signal. This is a function of frequency w and time t, and can well demonstrate the joint time-frequency distribution relationship of the random kernel signal.
[0033] S344, dividing the instantaneous frequency marginal spectrum of the nuclear signal into n instantaneous frequency intervals at a certain frequency interval, and respectively calculating the energy Ei in each frequency band interval, where i = 1, 2, ..., n;
[0034] S345. Let the total energy be:
[0035]
[0036] Define the normalized band energy:
[0037]
[0038] S346. The characteristic vector of the signal after HHT is finally obtained:
[0039] T=[e1,e2,e3,…,en](7)S4. Nuclide identification is completed through a one-dimensional convolutional neural network. Specifically, the following steps are included:
[0040] S41, determining the objective function;
[0041] First, the pulse signal eigenvector is defined as a set of random variables X={x0,…,x n}, Belongs to m discrete labels, The recursive network establishes a probability distribution model Q(X|θ,I) on random variables, where θ is the parameter of the network. This distribution is usually modeled as a signal feature edge Q(X|θ,I)=∏ i q i (x i |θ,I), where each edge represents the maximum similarity probability:
[0042]
[0043] in The function f is the partition function for each signal sampling point i. i A real-valued score representing the neural network. The higher the score, the greater the probability.
[0044] remember is the vectorized form of the network output Q(X|θ,I). The framework of the deep neural network can be optimized as follows:
[0045] Finding θ
[0046] Constrained by
[0047] Among them A I ∈R k×nm , Implement k separate linear constraints on the output distribution of signal I;
[0048] S42, setting network structure and training;
[0049] The network structure is input vector → convolution layer → feature mapping signal layer → continuous pooling layer → feature mapping signal layer → full connection → Gaussian connection; the continuous pooling layer refers to the data being subjected to the pooling layer → feature mapping signal layer → convolution layer → feature mapping signal layer → pooling layer, and repeated operations are performed so that the data fed into the fully connected layer reaches a suitable dimension; the suitable dimension means that the number of high-dimensional features after multiple feature extractions is less than or equal to two orders of magnitude of the number of categories of nuclide identification.
[0050] Specifically, the original data X, with a data length of m, is fed into the neural network; first, it is convolved with n different convolution kernels W, with a length of k∈[1, m], to obtain n feature maps Y, i.e., Y=X*W i , i∈[1,n];
[0051] Convolution kernel W i The dimension of is smaller than that of the original data, and a moving average is performed;
[0052] The moving average is to correspond the convolution kernel of length k to the first k elements of the original data one by one, multiply the corresponding elements in sequence, accumulate the product values of k pairs of elements by M, and let the average value of the sum of the products M / k be the result of the first convolution operation; before performing the next convolution operation, the convolution kernel needs to be shifted to the right by t unit lengths. After the shift is completed, the subsequent convolution operations are completed in the same way as the first convolution operation until the tail element of the convolution kernel is moved to align with the tail element of the original data, completing the last convolution operation;
[0053] Pooling is performed on the convolutional data; pooling refers to downsampling each region to obtain a value as a summary of this region; after obtaining the feature map, the feature map is subjected to maximum pooling operation;
[0054] The maximum pooling refers to selecting the maximum value in a region data, that is,
[0055] The pooled data is further convolved and pooled, ultimately making the data fed into the fully connected layer reach an appropriate dimension; the appropriate dimension means that the number of high-dimensional features obtained by multiple convolutions and pooling processes is less than or equal to two orders of magnitude of the number of categories for nuclide identification;
[0056] The Rectified Linear Unit (ReLU), also known as the rectifier function, is used in the fully connected layer. It is the main activation function used in deep neural networks. ReLU is actually a ramp function, defined as shown in formula (10):
[0057] ReLU(x)=max(0,x)(10)
[0058] The last layer is the output layer, which uses the softmax activation function to obtain the number of neurons C with the same nuclide type (single nuclide, mixed nuclide) of the input data; for multi-class problems, the category label y∈{1,2,…,C}; given a sample x, the conditional probability of belonging to category c predicted by softmax regression is shown in Equation (11), forcing the sum of all C output values of the neural network to be 1; the output value will represent the probability of each category in these C categories, and the one with the largest probability is the model prediction value;
[0059]
[0060] S5. Perform weight migration;
[0061] Input the unlearned experimental data into the network. For a given category c, let To train the feature object detection weights in the last layer of the network, let Detect weights for features in the network to be trained; Use the universal weight transfer function T to parameterize it instead of training it as a model parameter:
[0062]
[0063] Where θ is the parameter of the transfer learning network;
[0064] If the same transfer function T is intended to be applied to the classification of any feature c, θ is set to generalize the transfer function T to features not observed during training;
[0065] S6. Validate and evaluate the model;
[0066] The area under the ROC curve (AUC) was used as an indicator for model evaluation;
[0067] Among them, the horizontal axis is the false positive rate (FPR);
[0068]
[0069] Represents the probability of all negative samples being incorrectly predicted as positive samples, the false alarm rate; where: FP represents the number of samples whose true category is other nuclides and the model incorrectly identifies them as nuclides c; TN represents the number of samples whose true category is other nuclides and the model correctly identifies them as other nuclides;
[0070] The vertical axis is the True Positive Rate (TPR);
[0071]
[0072] Represents the probability of correct prediction among all positive samples, the hit rate; where: TP represents the number of samples whose true category is c and the model correctly identifies them as nuclide c; TN is the same as above;
[0073] AUC is the area under the ROC curve; the closer the AUC is to 1, the better the performance of the trained network model.
[0074] Specifically, in step 1, the instrument is used to simulate data and collect data; the nuclide library of the DT5800D digital detector simulator is used to simulate the pulse signal;
[0075] By using the nuclide library preset by DT5800D, gamma rays that are sufficient to serve as characteristics of the specified nuclide are selected to form a standard energy spectrum; then the pulse amplitude, rate, rise time, fall time, noise and baseline drift are set, and finally the pulse signal visualization window is displayed. 60 Co nuclear pulse signal.
[0076] Specifically, the steps of modal decomposition using the EMD algorithm in step S3 are as follows:
[0077] Sa1. Find all the local maximum and minimum points of the nuclear pulse signal, and perform interpolation operations on all the maximum and minimum points respectively to obtain the upper envelope u(t) and lower envelope l(t), whose mean m1(t) = [u(t) + l(t)] / 2, and calculate a sequence h 10 =x(t)-m1(t);
[0078] Sa2, for h 10 This sequence, if it meets the following two conditions: one is h 10 The number of all extreme points and zero crossing points must be the same or differ by at most one point; secondly, h 1o The upper and lower envelopes are symmetrical about the time axis, so c1=h 10 , that is, the first IMF is decomposed; if h 10 It is not an IMF. Repeat (1) k times as the original signal until h 1k Satisfy IMF conditions and become the first IMF;
[0079] Sa3. Subtract c1 from x(t) to get r1=x(t)-c1. Repeat steps (1) and (2) to get several c n (n=1,2,3,…); and so on, when r nWhen the monotonic function can no longer extract IMF components, the loop ends and n IMF components are obtained. where r n is the residual function, representing the average trend of the signal.
[0080] Specifically, the feature extraction based on the Hilbert marginal spectrum in step S3 includes the following steps:
[0081] For the decomposed IMF component c k (t) is subjected to Hilbert transform, and we get make:
[0082]
[0083] where a k (t),ω k (t) are c k (t) is the instantaneous amplitude and instantaneous frequency, from which the Hilbert spectrum H(t,ω) is obtained as:
[0084]
[0085] The Hilbert spectrum H(t,ω) is a function of frequency ω and time t, which can well show the time-frequency joint distribution relationship of random kernel signals. The Hilbert marginal spectrum is defined as the integral of the Hilbert spectrum on the time axis:
[0086] In the process of pulse signal feature extraction, the instantaneous frequency marginal spectrum of the nuclear signal is divided into n instantaneous frequency intervals at a certain frequency interval, and the energy E in each frequency band interval is calculated respectively. i ,i=1,2,3,…n; the total energy is After eliminating the dimension of energy, The final eigenvector of the signal after HHT is T = [e1, e2, e3, ..., e n ] and input into the convolutional neural network as the feature of the kernel pulse for identification.
[0087] The beneficial effects of the present invention are as follows: the nuclide identification method based on HHT band energy characteristics and convolutional neural network of the present invention utilizes the complete information of the pulse signal sequence, which is different from the shortcomings of the previous pulse-based nuclide identification that only uses partial signal characteristics, and has a fast recognition speed and high accuracy;
[0088] Nuclear pulse signals are non-stationary signals. HHT is the best choice for processing nonlinear and non-stationary signals. After performing time-to-frequency conversion on the signal, it can process and identify time signals of any length without the need to constantly adjust analysis parameters, thus improving the adaptability of the nuclide identification method. Drawing on the latest achievements of previous nuclide identification based on energy spectra, deep learning is used to further extract and classify features. Direct recognition of nuclear pulse signals is similar to time series classification. One-dimensional convolutional neural networks are currently the most widely used network architecture for time series classification problems due to their robustness and reduced training time.
[0089] And, it has the following beneficial effects:
[0090] (1) The HHT method is used to analyze nuclear pulse signals. The energy trend characteristics of the Hilbert spectra corresponding to different nuclides in different frequency bands are quite different, which is beneficial as an identification feature for nuclide identification research;
[0091] (2) Using the above features as the input features of the neural network, various nuclides can be effectively identified. When using simulated data, the average recognition rate reaches 99.93%, which lays a theoretical foundation for the subsequent design of a pulse-based portable nuclide identifier. BRIEF DESCRIPTION OF THE DRAWINGS
[0092] Figure 1 is a pulse signal simulated by the instrument in the embodiment of the present invention;
[0093] Figure 2 The IMF components obtained by EMD decomposition in the embodiment of the present invention;
[0094] Figure 3 In the embodiment of the present invention 60 Co eigenvector;
[0095] Figure 4 Schematic diagram of DWT eliminating nuclear pulse signal noise in an embodiment of the present invention;
[0096] Figure 5 Schematic diagram of the nuclear pulse signal feature extraction process based on HHT in an embodiment of the present invention;
[0097] Figure 6 This is a flow chart of a one-dimensional convolutional neural network in an embodiment of the present invention;
[0098] Figure 7 1. A one-dimensional convolutional neural network structure diagram according to an embodiment of the present invention;
[0099] Figure 8 Schematic diagram of one-dimensional convolution in an embodiment of the present invention;
[0100] Figure 9Schematic diagram of ROC curve in an embodiment of the present invention;
[0101] Figure 10 is the accuracy of the training set in the embodiment of the present invention;
[0102] Figure 11 is the training set loss value in the embodiment of the present invention;
[0103] Figure 12 is the accuracy of the verification set in the embodiment of the present invention;
[0104] Figure 13 is the validation set loss value in the embodiment of the present invention. DETAILED DESCRIPTION
[0105] The present invention will be further described below with reference to the accompanying drawings and examples.
[0106] Combined with attachment Figures 1 to 13 As shown, the nuclide identification method based on a one-dimensional convolutional neural network of a nuclear pulse sequence of the present invention comprises the following steps:
[0107] S1. Collect nuclear pulse signals;
[0108] S2, eliminating the kernel pulse signal noise by discrete wavelet transform, comprising the following steps:
[0109] S21, select wavelet and wavelet decomposition level, calculate the wavelet decomposition of the original signal to the nth level; use the Coiflet (coifN) wavelet system;
[0110] S22. Select a threshold value for each layer of high-frequency coefficients and correct the high-frequency coefficients. Use an improved threshold function to correct the high-frequency coefficients. The improved threshold function is as follows:
[0111]
[0112] thr=max(x j )(2)
[0113]
[0114] Where thr is the threshold; k is the empirical coefficient, 0≤k≤1, when k=0 it is equivalent to the hard threshold function, when k=1 it is equivalent to the soft threshold function; x jt and η jt are the t-th high-frequency coefficients of the j-th layer before and after correction respectively; sign is the sign function;
[0115] S23, performing wavelet reconstruction of the signal based on the low-frequency coefficients of the nth layer and the corrected high-frequency coefficients from the 1st layer to the nth layer;
[0116] S3, feature extraction based on HHT;
[0117] For the nuclear pulse signal sequence x(t), it can be decomposed into a finite number of intrinsic mode functions IMF through EMD, that is, the original signal is the superposition of these intrinsic mode functions:
[0118]
[0119] Where, IMF refers to the intrinsic mode function; n is the order of the IMF component, c i (t) is the i-th order IMF component, r n (t) is the trend term, which represents the average trend or mean of the signal;
[0120] The steps of the EMD decomposition algorithm are as follows:
[0121] S31, initialization, set r1(t)=x(t), i=1, k=0;
[0122] S32, obtaining the n-th order intrinsic mode function, comprising the following steps:
[0123] S321, initialization, set h1(t) = r1(t);
[0124] S322. Find all the maximum and minimum points of hk(t);
[0125] S323, fitting the maximum point and the minimum point respectively by using an interpolation function to obtain the upper and lower envelopes: e+(t) and e-(t);
[0126] S324. Calculate the mean of the upper and lower envelopes mk(t)=(e+(t)+e-(t)) / 2;
[0127] S325. Calculate hk+1(t)=hk(t)-mk(t);
[0128] S326, judgment: 0.2≤ε≤0.3, if it holds, then c i (t) = h k (t), otherwise, set k=k+1 and return to S322;
[0129] S33, rk+1(t)=rk(t)-ck+1(t), judging whether the remainder is a monotonic function or a constant, if so, the decomposition process ends;
[0130] After S34 and x(t) are decomposed by EMD, n intrinsic mode functions (IMFs) with frequencies from high to low are obtained. The kernel signal after EMD decomposition is then subjected to Hilbert transform for feature extraction. The process is as follows:
[0131] S341, performing EMD decomposition on the denoised nuclear pulse signal to obtain a series of IMFs;
[0132] S342, performing Hilbert transform on each IMF component to obtain the instantaneous frequency, amplitude function and corresponding Hilbert spectrum of each IMF;
[0133] S343. Summarize the Hilbert spectra of all IMFs to obtain the Hilbert spectrum H(w,t) of the original random kernel signal. This is a function of frequency w and time t, and can well demonstrate the joint time-frequency distribution relationship of the random kernel signal.
[0134] S344, dividing the instantaneous frequency marginal spectrum of the nuclear signal into n instantaneous frequency intervals at a certain frequency interval, and respectively calculating the energy Ei in each frequency band interval, where i = 1, 2, ..., n;
[0135] S345. Let the total energy be:
[0136]
[0137] Define the normalized band energy:
[0138]
[0139] S345. The characteristic vector of the signal after HHT is finally obtained:
[0140] T=[e1,e2,e3,…,en](7)S4. Nuclide identification is completed through a one-dimensional convolutional neural network. Specifically, the following steps are included:
[0141] S41, determining the objective function;
[0142] First, the pulse signal eigenvector is defined as a set of random variables X={x0,…,x n}, Belongs to m discrete labels, The recursive network establishes a probability distribution model Q(X|θ,I) on random variables, where θ is the parameter of the network. This distribution is usually modeled as a signal feature edge Q(X|θ,I)=∏ i q i (x i |θ,I), where each edge represents the maximum similarity probability:
[0143]
[0144] in The function f is the partition function for each signal sampling point i. iA real-valued score representing the neural network. The higher the score, the greater the probability.
[0145] remember is the vectorized form of the network output Q(X|θ,I). The framework of the deep neural network can be optimized as follows:
[0146] Finding θ
[0147] Constrained by
[0148] Among them A I ∈R k×nm , Implement k separate linear constraints on the output distribution of signal I;
[0149] S42, setting network structure and training;
[0150] The network structure is input vector → convolution layer → feature mapping signal layer → continuous pooling layer → feature mapping signal layer → full connection → Gaussian connection; the continuous pooling layer refers to performing a pooling layer → feature mapping signal layer → convolution layer → feature mapping signal layer → pooling layer on the data, and repeating the operation so that the data fed into the fully connected layer reaches a suitable dimension; the suitable dimension refers to the number of high-dimensional features after multiple feature extractions, which cannot be more than two orders of magnitude more than the number of categories of nuclide identification. Too large a dimension will increase the computational complexity of the fully connected layer, and too small a dimension will lead to poor classification effect of the fully connected layer;
[0151] Specifically, the original data X (data length: m) is sent to the neural network; first, it is convolved with n different convolution kernels W (length: k∈[1, m]) to obtain n feature maps Y, that is, Y=X*W i , i∈[1,n];
[0152] Convolution kernel W i The dimension of the convolution kernel is smaller than that of the original data, and a moving average is performed; the moving average refers to a one-to-one correspondence between the convolution kernel of length k and the first k elements of the original data, and the corresponding elements are multiplied in turn, and the product values of the k pairs of elements are accumulated M, and the average value of the product sum M / k is taken as the result of the first convolution operation; before the next convolution operation, the convolution kernel needs to be translated to the right by t unit lengths. After the movement is completed, the subsequent convolution operations are completed in the same way as the first convolution operation until the tail element of the convolution kernel is moved to align with the tail element of the original data, and the last convolution operation is completed; the data after convolution is pooled; the pooling refers to downsampling each area to obtain a value as a summary of this area; after obtaining the feature map, the feature map is subjected to a maximum pooling operation;
[0153] The maximum pooling refers to selecting the maximum value in a region data, that is,
[0154] The pooled data is further convolved and pooled, ultimately ensuring that the data fed into the fully connected layer reaches an appropriate dimension. Appropriate dimension refers to the number of high-dimensional features obtained through multiple convolutions and pooling processes. It cannot be more than two orders of magnitude greater than the number of nuclide categories to be identified. Too large a dimension will increase the computational load of the fully connected layer, while too small a dimension will result in poor classification performance.
[0155] The Rectified Linear Unit (ReLU), also known as the rectifier function, is used in the fully connected layer. It is the main activation function used in deep neural networks. ReLU is actually a ramp function, defined as shown in formula (10):
[0156] ReLU(x)=max(0,x)(10)
[0157] The last layer is the output layer, which uses the softmax activation function to obtain the number of neurons C with the same nuclide type (single nuclide, mixed nuclide) of the input data; for multi-class problems, the category label y∈{1,2,…,C}; given a sample x, the conditional probability of belonging to category c predicted by softmax regression is shown in Equation (11), forcing the sum of all C output values of the neural network to be 1; the output value will represent the probability of each category in these C categories, and the one with the largest probability is the model prediction value;
[0158]
[0159] S5. Perform weight migration;
[0160] Input the unlearned experimental data into the network. For a given category c, let To train the feature object detection weights in the last layer of the network, let Detect weights for features in the network to be trained; Use the universal weight transfer function T to parameterize it instead of training it as a model parameter:
[0161]
[0162] Where θ is the parameter of the transfer learning network;
[0163] If the same transfer function T is intended to be applied to the classification of any feature c, θ is set to generalize the transfer function T to features not observed during training;
[0164] S5. Validate and evaluate the model;
[0165] The area under the ROC curve (AUC) was used as an indicator for model evaluation;
[0166] Among them, the horizontal axis is the false positive rate (FPR);
[0167]
[0168] Represents the probability of all negative samples being incorrectly predicted as positive samples, the false alarm rate; where: FP represents the number of samples whose true category is other nuclides and the model incorrectly identifies them as nuclides c; TN represents the number of samples whose true category is other nuclides and the model correctly identifies them as other nuclides;
[0169] The vertical axis is the True Positive Rate (TPR);
[0170]
[0171] Represents the probability of correct prediction among all positive samples, the hit rate; where: TP represents the number of samples whose true category is c and the model correctly identifies them as nuclide c; TN is the same as above;
[0172] AUC is the area under the ROC curve; the closer the AUC is to 1, the better the performance of the trained network model.
[0173] Specifically, in step 1, the instrument is used to simulate data and collect data; the nuclide library of the DT5800D digital detector simulator is used to simulate the pulse signal;
[0174] By using the nuclide library preset by DT5800D, gamma rays that are sufficient to serve as characteristics of the specified nuclide are selected to form a standard energy spectrum; then the pulse amplitude, rate, rise time, fall time, noise and baseline drift are set, and finally the pulse signal visualization window is displayed. 60 Co nuclear pulse signal.
[0175] Specifically, the steps of modal decomposition using the EMD algorithm in step S3 are as follows:
[0176] Sa1. Find all the local maximum and minimum points of the nuclear pulse signal, and perform interpolation operations on all the maximum and minimum points respectively to obtain the upper envelope u(t) and lower envelope l(t), whose mean m1(t) = [u(t) + l(t)] / 2, and calculate a sequence h 10 =x(t)-m1(t);
[0177] Sa2, for h10 This sequence, if it meets the following two conditions: one is h 10 The number of all extreme points and zero crossing points must be the same or differ by at most one point; secondly, h 10 The upper and lower envelopes are symmetrical about the time axis, so c1=h 10 , that is, the first IMF is decomposed; if h 10 It is not an IMF. Repeat (1) k times as the original signal until h 1k Satisfy IMF conditions and become the first IMF;
[0178] Sa3, subtract c1 from x(t) to get r1=x(t)-c1. Repeat steps Sa1 and Sa2 to get several c n (n=1,2,3,…); and so on, when r n When the monotonic function can no longer extract IMF components, the loop ends and n IMF components are obtained. where r n is the residual function, representing the average trend of the signal.
[0179] Specifically, the feature extraction based on the Hilbert marginal spectrum in step S3 includes the following steps:
[0180] For the decomposed IMF component c k (t) is subjected to Hilbert transform, and we get make:
[0181]
[0182] where a k (t),ω k (t) are c k (t) is the instantaneous amplitude and instantaneous frequency, from which the Hilbert spectrum H(t,ω) is obtained as:
[0183]
[0184] The Hilbert spectrum H(t,ω) is a function of frequency ω and time t, which can well show the time-frequency joint distribution relationship of random kernel signals. The Hilbert marginal spectrum is defined as the integral of the Hilbert spectrum on the time axis:
[0185] In the process of pulse signal feature extraction, the instantaneous frequency marginal spectrum of the nuclear signal is divided into n instantaneous frequency intervals at a certain frequency interval, and the energy E in each frequency band interval is calculated respectively. i ,i=1,2,3,…n; the total energy is After eliminating the dimension of energy, The final eigenvector of the signal after HHT is T = [e1, e2, e3, ..., e n ] and input into the convolutional neural network as the feature of the kernel pulse for identification.
[0186] Example
[0187] The nuclide identification method based on HHT band energy characteristics and convolutional neural network includes the following steps:
[0188] S1. Collect nuclear pulse signals;
[0189] The nuclear pulse signal data is derived from a nuclide library simulation using the DT5800D digital detector simulator produced by CAEN. Using instrument-simulated data is driven by safety considerations, as long-term data collection in a radioactive environment is inherently harmful to the human body. Furthermore, due to cost and limited experimental conditions, establishing a nuclear pulse signal database for more than ten nuclides (single and mixed sources) would be both expensive and challenging to manage and approve. Therefore, this project utilizes the DT5800D digital detector simulator to collect simulated nuclear pulse signals.
[0190] The DT5800D, a multi-channel digital instrument developed jointly by nuclear instrumentation companies CAEN, SrL, and the Politecnico di Milano, is designed for simulating radiation detection systems. The processor is initialized with a reference pulse shape, which exhibits a statistical distribution of amplitude and time. Based on this information, the device generates a series of events that can be optionally summed to simulate pile-up phenomena. Arbitrary noise and baseline deviations can be superimposed on each pulse. Therefore, the digital detector simulator is a unique random pulse generator and also a simulator of radiation detector signals, with configurable energy and time distribution. The simulated signal stream becomes a statistical sequence of pulses reflecting the programmed input characteristics. When the simulation process is reset, the generator core can be reinitialized with new random data to ensure a consistently different sequence, or it can be stored to reproduce the same sequence multiple times. Therefore, using the DT5800D digital detector simulator offers advantages over simulating pulse signals in-house and reduces the complexity of data acquisition.
[0191] By using the nuclide library preset in DT5800D, we select gamma rays that are characteristic enough for the specified nuclide to form a standard energy spectrum. Then we set the pulse amplitude, rate, rise time, fall time, noise and baseline drift, and finally get the following in the pulse signal visualization window: Figure 1 The 60Co nuclear pulse signal (partial) is shown.
[0192] S2, eliminating the noise of the kernel pulse signal by discrete wavelet transform;
[0193] When measuring low-energy weak radioactivity, rapid identification is required. The number of counts in a short period of time is small and the impact of noise is large, so the impact of noise needs to be reduced as much as possible.
[0194] Discrete Wavelet Transform (DWT) is a signal processing method developed on the basis of Fourier transform, which has multi-resolution characteristics. Its basic idea is to project the signal into subspaces with different frequencies, process the signal in the frequency space, and finally reconstruct the signal. The nuclear signal data obtained by the nuclear pulse signal acquisition equipment is actually the superposition of noise (radioactive statistical fluctuations and electronic noise) and characteristic signals. The noise and characteristic signal are in different frequency bands, and the noise is a high-frequency component relative to the characteristic signal. By removing the high-frequency noise component and then restoring the signal, the signal can be filtered. Therefore, it is ideal to use the multi-resolution characteristics of DWT for signal denoising. The applicant has previously studied the application of DWT denoising to α energy spectrum and achieved good results.
[0195] The method eliminates the noise of the kernel pulse signal based on discrete wavelet transform, including the following steps:
[0196] S21. Select wavelets and wavelet decomposition levels, and calculate the wavelet decomposition of the original signal up to the nth level. This project intends to use the Coiflet (coifN) wavelet family, which is widely used in signal processing. Numerous experiments have shown that more decomposition levels remove more high-frequency information, but also reduce the true information contained in the low-frequency portion. This manifests as a peak suppression in the denoised pulse signal. Therefore, a limited number of decomposition levels is recommended.
[0197] S22. Select a threshold for each layer of high-frequency coefficients and correct the high-frequency coefficients. Since the high-frequency coefficients contain both noise and signal high frequencies, the threshold functions include hard threshold functions, soft threshold functions, and some improved threshold functions. This project intends to use an improved threshold function to correct the high-frequency coefficients. The function form is as follows:
[0198]
[0199] thr=max(x j )(2)
[0200]
[0201] Where thr is the threshold; k is the empirical coefficient, 0≤k≤1, when k=0 it is equivalent to the hard threshold function, when k=1 it is equivalent to the soft threshold function; x jt and η jt are the t-th high-frequency coefficients of the j-th layer before and after correction respectively; sign is the sign function;
[0202] S23, performing wavelet reconstruction of the signal based on the low-frequency coefficients of the nth layer and the corrected high-frequency coefficients from the 1st layer to the nth layer; Figure 4 shown.
[0203] S3, feature extraction based on HHT;
[0204] In actual nuclear signal digital sampling, the sampling frequency ranges from 5MHz to 100MHz. Even at the lowest sampling rate, a large amount of sampled data can be obtained per unit time.
[17] . If the entire nuclear pulse signal sampling data is directly used as the input feature of the deep neural network, it will inevitably result in a huge input layer, which is easy to cause the model to overfit, and the network models established for different signal lengths will also be different, resulting in poor generalization of the trained network model. If the signal length is trimmed in order to input signals of different lengths into the neural network, it will inevitably lead to data loss, increase or deviation from the true value. In order to make the simulation experiment as close to the real experiment as possible, this article also considers the problem of large data volume and inconsistent signal length when conducting simulation experiments. We collect the signal length, and theoretically we can ensure the consistency of the length. However, in the data trimming, we remove a large number of sampling points with an amplitude of 0 in the signal, which makes the length of each nuclear pulse signal inconsistent, which is close to the situation of inconsistent signal length in reality.
[0205] The Hilbert-Huang Transform (HHT) is a time-frequency analysis algorithm proposed by Norden E. Huang in 1998. It consists of empirical mode decomposition (EMD) and Hilbert spectrum analysis (HSA) and is very suitable for processing non-stationary nonlinear nuclear pulse signals.
[0206] Based on the local properties of the signal itself, EMD can adaptively decompose the signal into several intrinsic mode functions (IMFs) and a trend term. Each IMF obtained by decomposition is Hilbert transformed and summarized by HSA to obtain a Hilbert spectrum represented as time-frequency-amplitude. The Hilbert marginal spectrum is then obtained by time integration, and the energy of each frequency band is calculated by frequency band division. The frequency band energy trend is used as the eigenvector of the pulse signal. [18,19,20] .
[0207] The purpose of time-frequency conversion of the signal is to enable the identification of nuclear pulse signals of any length without constantly adjusting the analysis parameters to adapt to the requirement of consistent input data size of the neural network input layer. This increases the adaptability of the new nuclide identification method and solves the problem of difficulty in training neural networks due to inconsistent signal lengths when analyzing based on time series signals.
[0208] For the nuclear pulse signal sequence x(t), it can be decomposed into a finite number of intrinsic mode functions IMF through EMD, that is, the original signal is the superposition of these intrinsic mode functions:
[0209]
[0210] Where IMF refers to the intrinsic mode function; n is the order of the IMF component, ci(t) is the i-th order IMF component, and rn(t) is the trend term, which represents the average trend or mean of the signal.
[0211] The steps of the EMD decomposition algorithm are as follows:
[0212] S31, initialization, set r1(t)=x(t), i=1, k=0;
[0213] S32, obtaining the n-th order intrinsic mode function, comprising the following steps:
[0214] S321, initialization, set h1(t) = r1(t);
[0215] S322. Find all the maximum and minimum points of hk(t);
[0216] S323, fitting the maximum point and the minimum point respectively by using an interpolation function to obtain the upper and lower envelopes: e+(t) and e-(t);
[0217] S324. Calculate the mean of the upper and lower envelopes mk(t)=(e+(t)+e-(t)) / 2;
[0218] S325. Calculate hk+1(t)=hk(t)-mk(t);
[0219] S326, judgment: 0.2≤ε≤0.3. If so, ci(t)=hk(t). Otherwise, set k=k+1 and return to S322.
[0220] S33, rk+1(t)=rk(t)-ck+1(t), judging whether the remainder is a monotonic function or a constant, if so, the decomposition process ends;
[0221] After S34 and x(t) are decomposed by EMD, n intrinsic mode functions (IMFs) with frequencies from high to low are obtained. The kernel signal after EMD decomposition is then subjected to Hilbert transform for feature extraction. The process is as follows:
[0222] S341, performing EMD decomposition on the denoised nuclear pulse signal to obtain a series of IMFs;
[0223] S342, performing Hilbert transform on each IMF component to obtain the instantaneous frequency, amplitude function and corresponding Hilbert spectrum of each IMF;
[0224] S343. Summarize the Hilbert spectra of all IMFs to obtain the Hilbert spectrum H(w,t) of the original random kernel signal. This is a function of frequency w and time t, and can well demonstrate the joint time-frequency distribution relationship of the random kernel signal.
[0225] S344, dividing the instantaneous frequency marginal spectrum of the nuclear signal into n instantaneous frequency intervals at a certain frequency interval, and respectively calculating the energy Ei in each frequency band interval, where i = 1, 2, ..., n;
[0226] S345. Let the total energy be:
[0227]
[0228] Define the normalized band energy:
[0229]
[0230] S346. The characteristic vector of the signal after HHT is finally obtained:
[0231] T=[e1,e2,e3,…,en](7)
[0232] Specifically, in actual nuclear signal digital sampling, the sampling frequency ranges from 5MHz to 100MHz. Even at the lowest sampling rate, a large amount of sampled data can be obtained per unit time.
[17] . If the entire nuclear pulse signal sampling data is directly used as the input feature of the deep neural network, it will inevitably result in a huge input layer, which is easy to cause the model to overfit, and the network models established for different signal lengths will also be different, resulting in poor generalization of the trained network model. If the signal length is trimmed in order to input signals of different lengths into the neural network, it will inevitably lead to data loss, increase or deviation from the true value. In order to make the simulation experiment as close to the real experiment as possible, this article also considers the problem of large data volume and inconsistent signal length when conducting simulation experiments. We collect the signal length, and theoretically we can ensure the consistency of the length. However, in the data trimming, we remove a large number of sampling points with an amplitude of 0 in the signal, which makes the length of each nuclear pulse signal inconsistent, which is close to the situation of inconsistent signal length in reality.
[0233] The Hilbert-Huang Transform (HHT) is a time-frequency analysis algorithm proposed by Norden E. Huang in 1998. It consists of empirical mode decomposition (EMD) and Hilbert spectrum analysis (HSA) and is very suitable for processing non-stationary nonlinear nuclear pulse signals.
[0234] Based on the local properties of the signal itself, EMD can adaptively decompose the signal into several intrinsic mode functions (IMFs) and a trend term. Each IMF obtained by decomposition is Hilbert transformed and summarized by HSA to obtain a Hilbert spectrum represented as time-frequency-amplitude. The Hilbert marginal spectrum is then obtained by time integration, and the energy of each frequency band is calculated by frequency band division. The frequency band energy trend is used as the eigenvector of the pulse signal. [18,19,20] .
[0235] The purpose of time-frequency conversion of the signal is to enable the identification of nuclear pulse signals of any length without constantly adjusting the analysis parameters to adapt to the requirement of consistent input data size of the neural network input layer. This increases the adaptability of the new nuclide identification method and solves the problem of difficulty in training neural networks due to inconsistent signal lengths when analyzing based on time series signals.
[0236] EMD principle: EMD is essentially a process of adaptive signal screening. Assume that the original nuclear pulse signal is
[0237] x(t), the basic steps of modal decomposition using the EMD algorithm are:
[0238] (1) Find all the local maximum and minimum points of the nuclear pulse signal, and perform interpolation operations on all the maximum and minimum points respectively to obtain the upper envelope u(t) and lower envelope l(t), whose mean m1(t) = [u(t) + l9t)] / 2, and calculate a sequence h 10 =x(t)-m1(t).
[0239] (2) For h 10 This sequence, if it meets the following two conditions: one is h 10 The number of all extreme points and zero crossing points must be the same or differ by at most one point; secondly, h 10 The upper and lower envelopes are symmetrical about the time axis, so c1=h 10 , that is, decomposition to obtain the first IMF. If h 10 It is not an IMF. Repeat (1) k times as the original signal until h 1k Meet the IMF conditions and become the first IMF.
[0240] (3) Subtract c1 from x(t) to get r1 = x(t) - c1. Repeat steps Sa1 and Sa2 to get several c n (n=1,2,3,…). And so on, when r n When the monotonic function can no longer extract IMF components, the loop ends and n IMF components are obtained. where r n is the residual function, representing the average trend of the signal. For the nuclear pulse signal x(t), the result of the EMD process is as follows Figure 2 shown.
[0241] Feature extraction based on Hilbert marginal spectrum
[0242] For the decomposed IMF component c k 9t) Perform Hilbert transform and get make:
[0243]
[0244]
[0245] where a k (t),ω k (t) are c k (t) is the instantaneous amplitude and instantaneous frequency, from which the Hilbert spectrum H(t,ω) is obtained as:
[0246]
[0247] The Hilbert spectrum H(t,ω) is a function of frequency ω and time t, which can well show the time-frequency joint distribution relationship of random kernel signals. The Hilbert marginal spectrum is defined as the integral of the Hilbert spectrum on the time axis:
[0248] In the process of pulse signal feature extraction, the instantaneous frequency marginal spectrum of the nuclear signal is divided into n instantaneous frequency intervals at a certain frequency interval, and the energy E in each frequency band interval is calculated respectively. i ,i=1,2,3,…n. The total energy is After eliminating the dimension of energy, The final eigenvector of the signal after HHT is T = [e1, e2, e3, ..., e n ] and input into the convolutional neural network as the feature of the kernel pulse for identification.
[0249] After eliminating dimension 60 Co nuclear signal eigenvector is as follows Figure 3 shown.
[0250] S4, complete nuclide identification through one-dimensional convolutional neural network;
[0251] Deep learning is the most important technological breakthrough and research hotspot in the era of machine learning. This type of neural network with multiple hidden layers has powerful learning capabilities, and this learning ability will be further improved as the depth of the hidden layer or the number of neurons increases.
[0252] Deep Convolutional Neural Networks (DCNNs) are trained using stochastic gradient descent and backpropagation algorithms under supervised learning. Training involves calculating the gradient of each neuron parameter and iteratively updating the parameters using parameter sensitivity until a stopping criterion is met. DCNNs integrate feature extraction and classification into a single learning process, allowing them to learn optimized features directly from the raw input during training.
[0253] Typically, two-dimensional deep convolutional neural networks (2DCNNs) are widely used in processing two-dimensional (2D) data such as images and videos. However, when processing signal-related one-dimensional (1D) data, compact and adaptive one-dimensional deep convolutional neural networks (1DCNNs) have fewer hidden layers and neurons than 2DCNNs, making them easier to train and implement. Furthermore, 1DCNNs do not require matrix operations, only simple array operations, which means that the computational complexity of 1DCNNs is significantly lower than that of 2DCNNs. Due to their low computational requirements, compact 1DCNNs are well-suited for real-time and low-cost applications, especially on mobile or handheld devices.
[0254] The specific steps include:
[0255] S41, determining the objective function;
[0256] First, the pulse signal eigenvector is defined as a set of random variables X={x0,…,x n}, Belongs to m discrete labels, The recursive network establishes a probability distribution model Q(X|θ,I) on random variables, where θ is the parameter of the network. This distribution is usually modeled as a signal feature edge Q(X|θ,I)=∏ i q i (x i |θ,I), where each edge represents the maximum similarity probability:
[0257]
[0258] in The function f is the partition function for each signal sampling point i. i A real-valued score representing the neural network. The higher the score, the greater the probability.
[0259] remember is the vectorized form of the network output Q(X|θ,I). The framework of the deep neural network can be optimized as follows:
[0260] Finding θ
[0261] Constrained by
[0262] Among them A I ∈R k×nm , Implement k separate linear constraints on the output distribution of signal I;
[0263] S42, setting network structure and training;
[0264] The network structure is input vector → convolution layer → feature mapping signal layer → continuous pooling layer → feature mapping signal layer → full connection → Gaussian connection; the continuous pooling layer refers to performing a pooling layer → feature mapping signal layer → convolution layer → feature mapping signal layer → pooling layer on the data, and repeating the operation so that the data fed into the fully connected layer reaches a suitable dimension; the suitable dimension refers to the number of high-dimensional features after multiple feature extractions, which cannot be more than two orders of magnitude more than the number of categories of nuclide identification. Too large a dimension will increase the computational complexity of the fully connected layer, and too small a dimension will lead to poor classification effect of the fully connected layer;
[0265] Specifically, the original data X (data length: m) is sent to the neural network; first, it is convolved with n different convolution kernels W (length: k∈[1, m]) to obtain n feature maps Y, that is, Y=X*W i , i∈[1,n];
[0266] Convolution kernel W i The dimension of is smaller than that of the original data, so a moving average is performed; the moving average means: a convolution kernel of length k is matched one by one with the first k elements of the original data, the corresponding elements are multiplied in turn, the product values of k pairs of elements are accumulated M, and the average value of the sum of the products M / k is taken as the result of the first convolution operation; before the next convolution operation, the convolution kernel is shifted to the right by t unit lengths. After the movement is completed, the subsequent convolution operations are completed in the same way as the first convolution operation until the tail element of the convolution kernel is moved to align with the tail element of the original data, completing the last convolution operation;
[0267] Pooling is performed on the convolutional data; pooling refers to downsampling each region to obtain a value as a summary of this region; after obtaining the feature map, the feature map is subjected to maximum pooling operation;
[0268] The maximum pooling refers to selecting the maximum value in a region data, that is,
[0269] The pooled data is further convolved and pooled, ultimately ensuring that the data fed into the fully connected layer reaches an appropriate dimension. Appropriate dimension refers to the number of high-dimensional features obtained through multiple convolutions and pooling processes. It cannot be more than two orders of magnitude greater than the number of nuclide categories to be identified. Too large a dimension will increase the computational load of the fully connected layer, while too small a dimension will result in poor classification performance.
[0270] The Rectified Linear Unit (ReLU), also known as the rectifier function, is used in the fully connected layer. It is the main activation function used in deep neural networks. ReLU is actually a ramp function, defined as shown in formula (10):
[0271] ReLU(x)=max(0,x)(10)
[0272] The last layer is the output layer, which uses the softmax activation function to obtain the number of neurons C with the same nuclide type (single nuclide, mixed nuclide) of the input data; for multi-class problems, the category label y∈{1,2,…,C}; given a sample x, the conditional probability of belonging to category c predicted by softmax regression is shown in Equation (11), forcing the sum of all C output values of the neural network to be 1; the output value will represent the probability of each category in these C categories, and the one with the largest probability is the model prediction value;
[0273]
[0274] S5. Perform weight migration;
[0275] Input the unlearned experimental data into the network. For a given category c, let To train the feature object detection weights in the last layer of the network, let Detect weights for features in the network to be trained; Use the universal weight transfer function T to parameterize it instead of training it as a model parameter:
[0276]
[0277] Where θ is the parameter of the transfer learning network;
[0278] If the same transfer function T is intended to be applied to the classification of any feature c, θ is set to generalize the transfer function T to features not observed during training;
[0279] S6. Validate and evaluate the model;
[0280] The area under the ROC curve (AUC) was used as an indicator for model evaluation;
[0281] Among them, the horizontal axis is the false positive rate (FPR);
[0282]
[0283] Represents the probability of all negative samples being incorrectly predicted as positive samples, the false alarm rate; where: FP represents the number of samples whose true category is other nuclides and the model incorrectly identifies them as nuclides c; TN represents the number of samples whose true category is other nuclides and the model correctly identifies them as other nuclides;
[0284] The vertical axis is the True Positive Rate (TPR);
[0285]
[0286] Represents the probability of correct prediction among all positive samples, the hit rate; where: TP represents the number of samples whose true category is c and the model correctly identifies them as nuclide c; TN is the same as above;
[0287] AUC is the area under the ROC curve; the closer the AUC is to 1, the better the performance of the trained network model.
[0288] This paper divides the Hilbert marginal spectrum into 200 frequency bins based on frequency. The energy value of each frequency bin is used as the input layer neuron of the 1DCNN, meaning that the input layer has a total of 200 neurons. The verification experiment designed in this paper is to identify a specific nuclide from 60Co, 137Cs, 152Eu, 60Co+137Cs, 60Co+152Eu, and 137Cs+152Eu. Therefore, the number of neurons in the output layer of the 1DCNN is set to 6. 200 data points are collected for each nuclide, for a total of 1200 data points for the six nuclide combinations, of which 70% are used as the training set and 30% as the test set.
[0289] The experimental environment parameters of this paper are shown in the following table.
[0290] Table 1 Experimental platform parameters
[0291]
[0292] The whole process is trained for 100 epochs, and the batch size is 64. During the training process, the changes in the loss value and accuracy of the training set, and the changes in the loss value and accuracy of the validation set are as follows: Figure 10-13 shown.
[0293] The network model trained on the training and validation sets was used to test the test set data to verify the effectiveness of the new method. A total of 8 experiments were conducted on the validation set, and the experimental results are shown in the following table.
[0294] Table 2 Test set accuracy
[0295]
Claims
1. A nuclide identification method based on HHT band energy characteristics and convolutional neural network, characterized in that: The following steps are involved: S1. Collect nuclear pulse signals; S2, eliminating the kernel pulse signal noise by discrete wavelet transform, comprising the following steps: S21, select wavelet and wavelet decomposition level, calculate the wavelet decomposition of the original signal to the nth level; use the Coiflet (coifN) wavelet system; S22. Select a threshold value for each layer of high-frequency coefficients and correct the high-frequency coefficients. Use an improved threshold function to correct the high-frequency coefficients. The improved threshold function is as follows: thr=max(x j ) (2) Where thr is the threshold; k is the empirical coefficient, 0≤k≤1, when k=0 it is equivalent to the hard threshold function, when k=1 it is equivalent to the soft threshold function; x jt and η jt are the t-th high-frequency coefficients of the j-th layer before and after correction respectively; sign is the sign function; S23, performing wavelet reconstruction of the signal based on the low-frequency coefficients of the nth layer and the corrected high-frequency coefficients from the 1st layer to the nth layer; S3, feature extraction based on HHT; For the nuclear pulse signal sequence x(t), it can be decomposed into a finite number of intrinsic mode functions IMF through EMD, that is, the original signal is the superposition of these intrinsic mode functions: Where, IMF refers to the intrinsic mode function; n is the order of the IMF component, c i (t) is the i-th order IMF component, r n (t) is the trend term, which represents the average trend or mean of the signal; The steps of the EMD decomposition algorithm are as follows: S31, initialization, set r1(t)=x(t), i=1, k=0; S32, obtaining the n-th order intrinsic mode function, comprising the following steps: S321, initialization, set h1(t) = r1(t); S322. Find all the maximum and minimum points of hk(t); S323, fitting the maximum point and the minimum point respectively by using an interpolation function to obtain the upper and lower envelopes: e+(t) and e-(t); S324. Calculate the mean of the upper and lower envelopes mk(t)=(e+(t)+e-(t)) / 2; S325. Calculate hk+1(t)=hk(t)-mk(t); S326, judgment: If established, then c i (t) = h k (t), otherwise, set k=k+1 and return to S322; S33, rk+1(t)=rk(t)-ck+1(t), judging whether the remainder is a monotonic function or a constant, if so, the decomposition process ends; After S34 and x(t) are decomposed by EMD, n intrinsic mode functions (IMFs) with frequencies from high to low are obtained. The kernel signal after EMD decomposition is then subjected to Hilbert transform for feature extraction. The process is as follows: S341, performing EMD decomposition on the denoised nuclear pulse signal to obtain a series of IMFs; S342, performing Hilbert transform on each IMF component to obtain the instantaneous frequency, amplitude function and corresponding Hilbert spectrum of each IMF; S343. Summarize the Hilbert spectra of all IMFs to obtain the Hilbert spectrum H(w,t) of the original random kernel signal; it is a function of frequency w and time t, and can well demonstrate the joint time-frequency distribution relationship of the random kernel signal; S344, dividing the instantaneous frequency marginal spectrum of the nuclear signal into n instantaneous frequency intervals at a certain frequency interval, and respectively calculating the energy Ei in each frequency band interval, where i = 1, 2, ..., n; S345. Let the total energy be: Define the normalized band energy: S346. The characteristic vector of the signal after HHT is finally obtained: T = [e1, e2, e3, …, en] (7) S4. Nuclide identification is completed through a one-dimensional convolutional neural network. Specifically, the following steps are included: S41, determining the objective function; First, the pulse signal eigenvector is defined as a set of random variables Belongs to m discrete labels, The recursive network establishes a probability distribution model Q(X|θ,I) on random variables, where θ is the parameter of the network; this distribution is usually modeled as a signal feature edge Q(X|θ,I)=∏ i q i (x i |θ,I), where each edge represents the maximum similarity probability: in is the partition function for each signal sampling point i; function f i represents the real-valued score of the neural network; the higher the score, the greater the probability; remember is the vectorized form of the network output Q(X|θ,I); the framework of the deep neural network can be optimized as follows: Finding θ Constrained by Among them A I ∈R k×nm , Implement k separate linear constraints on the output distribution of signal I; S42, setting network structure and training; The network structure is input vector → convolution layer → feature mapping signal layer → continuous pooling layer → feature mapping signal layer → full connection → Gaussian connection; the continuous pooling layer refers to performing a pooling layer → feature mapping signal layer → convolution layer → feature mapping signal layer → pooling layer on the data, and repeating the operation so that the data fed into the fully connected layer reaches an appropriate dimension; the appropriate dimension means that the number of high-dimensional features after multiple feature extractions is less than or equal to two orders of magnitude of the number of categories of nuclide identification; Specifically, the original data X, with a data length of m, is fed into the neural network; first, it is convolved with n different convolution kernels W, with a length of k∈[1, m], to obtain n feature maps Y, i.e., Y=X*W i , i∈[1,n]; Convolution kernel W i The dimension of the convolution kernel is smaller than that of the original data, and a moving average is performed. The moving average is to correspond the convolution kernel of length k to the first k elements of the original data one by one, multiply the corresponding elements in turn, and accumulate the product values of k pairs of elements by M. The average value of the sum of the products M / k is used as the result of the first convolution operation. Before the next convolution operation, the convolution kernel needs to be shifted to the right by t unit lengths. After the shift is completed, the subsequent convolution operations are completed in the same way as the first convolution operation until the tail element of the convolution kernel is moved to align with the tail element of the original data, completing the last convolution operation. Pooling is performed on the convolutional data; pooling refers to downsampling each region to obtain a value as a summary of this region; after obtaining the feature map, the feature map is subjected to the maximum pooling operation; The maximum pooling refers to selecting the maximum value in a region data, that is, The pooled data is further convolved and pooled, ultimately making the data fed into the fully connected layer reach an appropriate dimension; the appropriate dimension means that the number of high-dimensional features obtained by multiple convolutions and pooling processes is less than or equal to two orders of magnitude of the number of categories for nuclide identification; The Rectified Linear Unit (ReLU), also known as the rectifier function, is used in the fully connected layer. It is the main activation function used in deep neural networks. ReLU is actually a ramp function, defined as shown in formula (10): ReLU(x) = max(0, x) (10) The last layer is the output layer, which uses the softmax activation function to obtain the number of neurons C with the same nuclide type as the input data; for multi-class problems, the category label y∈{1,2,…,C}; given a sample x, the conditional probability of belonging to category c predicted by softmax regression is shown in Equation (11), forcing the sum of all C output values of the neural network to be 1; the output value will represent the probability of each category in these C categories, and the one with the largest probability is the model prediction value; S5. Perform weight migration; Input the unlearned experimental data into the network. For a given category c, let To train the feature object detection weights in the last layer of the network, let Detect weights for features in the network to be trained; Use the universal weight transfer function T to parameterize it instead of training it as a model parameter: Where θ is the parameter of the transfer learning network; If the same transfer function T is intended to be applied to the classification of any feature c, θ is set to generalize the transfer function T to features not observed during training; S6. Validate and evaluate the model; The area under the ROC curve (AUC) was used as an indicator for model evaluation; Among them, the horizontal axis is the false positive rate (FPR); Represents the probability of all negative samples being incorrectly predicted as positive samples, the false alarm rate; where: FP represents the number of samples whose true category is other nuclides and the model incorrectly identifies them as nuclides c; TN represents the number of samples whose true category is other nuclides and the model correctly identifies them as other nuclides; The vertical axis is the True Positive Rate (TPR); Represents the probability of correct prediction among all positive samples, the hit rate; where: TP represents the number of samples whose true category is c and the model correctly identifies them as nuclide c; TN is the same as above; AUC is the area under the ROC curve; the closer the AUC is to 1, the better the performance of the trained network model.
2. The nuclide identification method based on HHT band energy characteristics and convolutional neural network according to claim 1, characterized in that: In step S1, the instrument simulates data and collects it; the nuclide library of the DT5800D digital detector simulator simulates the pulse signal; By using the nuclide library preset by DT5800D, gamma rays that are sufficient to be characteristic of the specified nuclide are selected to form a standard energy spectrum; then the pulse amplitude, rate, rise time, fall time, noise and baseline drift are set, and finally the pulse signal visualization window is obtained. 60 Co nuclear pulse signal.
3. The nuclide identification method based on HHT band energy characteristics and convolutional neural network according to claim 2, characterized in that: The steps of modal decomposition using the EMD algorithm in step S3 are as follows: Sa1. Find all the local maximum and minimum points of the nuclear pulse signal, and perform interpolation operations on all the maximum and minimum points respectively to obtain the upper envelope u(t) and lower envelope l(t), whose mean m1(t) = [u(t) + l(t)] / 2, and calculate a sequence h 10 =x(t)-m1(t); Sa2, for h 10 This sequence, if it meets the following two conditions: one is h 10 The number of all extreme points and zero crossing points must be the same or differ by at most one point; secondly, h 10 The upper and lower envelopes are symmetrical about the time axis, so c1=h 10 , that is, the first IMF is decomposed; if h 10 It is not an IMF. Repeat (1) k times as the original signal until h 1k Satisfy IMF conditions and become the first IMF; Sa3, subtract c1 from x(t) to get r1=x(t)-c1; repeat steps Sa1 and Sa2 to get several c n (n=1,2,3,…); and so on, when r n When the monotonic function can no longer extract the IMF component, the loop ends and n IMF components are obtained; at this time, where r n is the residual function, representing the average trend of the signal.
4. The nuclide identification method based on HHT band energy characteristics and convolutional neural network according to claim 3, characterized in that: In step S3, feature extraction based on Hilbert marginal spectrum includes the following steps: For the decomposed IMF component c k (t) is subjected to Hilbert transform, and we get make: where a k (t),ω k (t) are c k (t) is the instantaneous amplitude and instantaneous frequency, from which the Hilbert spectrum H(t,ω) is obtained as: The Hilbert spectrum H(t,ω) is a function of frequency ω and time t, which can well show the time-frequency joint distribution relationship of random kernel signals; the Hilbert marginal spectrum is defined as the integral of the Hilbert spectrum on the time axis: In the process of pulse signal feature extraction, the instantaneous frequency marginal spectrum of the nuclear signal is divided into n instantaneous frequency intervals at a certain frequency interval, and the energy E in each frequency band interval is calculated respectively. i ,i=1,2,3,…n; the total energy is After eliminating the dimension of energy, The final eigenvector of the signal after HHT is T = [e1, e2, e3, ..., e n ] and input into the convolutional neural network as the feature of the kernel pulse for identification.