Nuclear database management system based on nuclear dynamic detection technology
By deploying high-precision sensors in the nuclear facility, combining noise suppression and drift compensation algorithms, using DTW, LSTM autoencoder and MerkleTree hash chain verification to detect abnormalities, using Kriging interpolation and SVR models to correct wrong data, and combining distributed storage and indexing technology, the incompleteness and unreliability of nuclear data acquisition and storage in the nuclear facility are solved, and efficient and reliable nuclear data management is achieved.
Patent Information
- Application Number
- CN202510417063.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the environment of strong radiation, electromagnetic interference and temperature and humidity fluctuations, the nuclear data collected by high-precision sensors are susceptible to noise and drift, resulting in data interruption, errors and storage media damage. The existing technology has not been effectively identified and corrected, resulting in incomplete and unreliable nuclear data.
The data acquisition and processing module is used to deploy high-precision sensors, combining noise suppression and drift compensation algorithms for pre-processing; the data quality monitoring module uses DTW to combine LSTM autoencoder, improved 3σ criterion algorithm and MerkleTree hash chain verification to detect abnormalities; the data repair and reconstruction module interpolates missing data and support vector regression SVR model to correct wrong data through the Kriging spatial interpolation algorithm; the data storage management module adopts distributed disaster recovery storage and compressed storage technology; the data query analysis module uses B+ tree index and R-tree radiation field spatial index for multi-dimensional query.
It improves the accuracy and reliability of nuclear data collection, promptly discovers and repairs abnormal data, ensures data integrity and consistency, reduces storage resource consumption, improves data storage security and availability, and supports rapid retrieval and efficient management.
Smart Images

Figure CN120336304A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of nuclear data management, and specifically to a nuclear database management system for nuclear dynamic detection technology. Background Art
[0002] Nuclear dynamic detection technology relies on high-precision sensors such as neutron detectors and γ-ray monitors to collect parameters of the operating power, temperature, and radiation field intensity of a nuclear reactor in real time, and stores and analyzes them through a nuclear database management system (Nuclear database management system, NDMS). The safe operation of nuclear facilities is the significance of nuclear dynamic detection, and the nuclear data obtained through nuclear dynamic detection is an important basis for evaluating and analyzing nuclear reaction experiments. Therefore, NDMS needs to have high security and high reliability to promote the development of nuclear science.
[0003] However, in special environments such as strong radiation, electromagnetic interference, and fluctuating temperature and humidity, nuclear facilities may cause the following problems:
[0004] The prior art has the following deficiencies: When high-precision sensors collect nuclear data, environmental interference such as low-frequency electromagnetic noise and temperature and humidity drift may cause the acquisition signal to be interrupted or the storage medium to be damaged, and may also cause incorrect data to be output. However, if the database fails to identify or correct it in time, abnormal loss and abnormal jump of nuclear data will occur, resulting in the incompleteness and unreliability of nuclear data and increasing the management difficulty.
[0005] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0006] The purpose of the present invention is to provide a nuclear database management system for nuclear dynamic detection technology. The present invention detects abnormal data by setting a data quality monitoring module and using dynamic time warping (DTW) combined with a long short-term memory (LSTM) autoencoder, an improved 3σ criterion algorithm, and a Merkle Tree hash chain verification algorithm. By setting a data repair and reconstruction module, the Kriging spatial interpolation algorithm is used to interpolate missing data, and the support vector regression (SVR) model is used to correct incorrect data, successfully repairing problematic nuclear data to solve the problems in the above background art.
[0007] To achieve the above object, the present invention provides the following technical solutions: A nuclear database management system for nuclear dynamic detection technology, including a data acquisition and processing module: Deploy high-precision sensors on nuclear facilities to collect raw data of dynamic detection nuclear data and environmental parameter data of nuclear facilities in real time. Considering the influence of environmental parameter data on dynamic detection nuclear data, use noise suppression and drift compensation algorithms to preprocess the raw data, and combine abnormal detection and processing of storage media to generate preprocessed nuclear data, and then transmit it to the data quality monitoring module;
[0008] Data quality monitoring module: Read reference nuclear data from the historical database, receive the preprocessed nuclear data in the data acquisition and processing module, use dynamic time warping (DTW) combined with LSTM autoencoder, improved 3σ criterion algorithm and MerkleTree hash chain verification algorithm to perform real-time detection of abnormal states of nuclear data on the preprocessed nuclear data, mark the problematic nuclear data, and transmit it to the data repair and reconstruction module;
[0009] Data repair and reconstruction module: Receive the marked problematic nuclear data in the data quality monitoring module and trigger the repair mechanism, including using the Kriging spatial interpolation algorithm to interpolate missing data in the problematic nuclear data, and using the support vector regression (SVR) model to correct the error data in the problematic nuclear data, obtaining standardized nuclear data that is consistent with the reference nuclear data after recovery, and transmitting it to the data storage and management module;
[0010] Data storage and management module: Receive the repaired standardized nuclear data in the data repair and reconstruction module, use distributed disaster tolerance storage and compression storage technology to minimize the resource storage of the standardized nuclear data while achieving high-reliability storage and fast retrieval functions, as archived nuclear data;
[0011] Data query and analysis module: Nuclear technicians use the B+ tree index combined with the R-tree radiation field spatial index strategy to perform multi-dimensional conditional queries on the archived nuclear data, can directly retrieve relevant nuclear data, and generate radiation reports through visualization analysis technology, thereby assisting in nuclear facility management decision-making analysis.
[0012] Optionally, the preprocessing steps of the raw data are as follows:
[0013] The sensor collects raw data in real time and calibrates it as Rnd(t), including dynamic detection nuclear data and environmental parameter data of nuclear facilities. Perform noise suppression processing on the raw data, using wavelet decomposition combined with an adaptive threshold processing algorithm. Among them, decompose the raw data Rnd(t) into multi-scale wavelet coefficients, perform threshold filtering on the high-frequency noise coefficients, and then reconstruct the denoised signal with the filtered coefficients to generate denoised data. The calculation formula for noise suppression is And T(W i,j ) = sign(Wi,j ) * max(|W i,j | - λ, 0), where Rnd'(t) represents the denoised data, and W i,j represents the wavelet coefficients obtained by decomposing the original data Rnd(t) at scale i and translation j. I represents the maximum decomposition level, and T(W i,j ) represents the adaptive threshold function, ψ i,j (Rnd(t)) represents the Daubechies wavelet basis function, λ represents the soft threshold parameter, and sign represents the sign function for extracting the sign information of the wavelet coefficients of the input value W i,j ;
[0014] In the environment where the nuclear facility is located, a temperature and humidity drift compensation model is established based on the offset of the data collected by the sensor. The original data Rnd(t) is iteratively processed using Kalman filtering. The true state is estimated through the prediction-update step, and the original data is corrected according to the estimated value of the temperature and humidity state to obtain the compensated data.
[0015] Then the calculation formula for temperature and humidity drift compensation is Rnd”(t) = Rnd(t) - α·ΔT - β·ΔH, and ΔT = T t - T0, ΔH = H t - H0, where Rnd”(t) represents the compensated data, α represents the sensitivity coefficient of the experimentally calibrated temperature to the sensor, β represents the sensitivity coefficient of the experimentally calibrated humidity to the sensor, ΔT represents the temperature change, ΔH represents the humidity change, T t represents the temperature value at the current time t after Kalman filter prediction-update, H t represents the humidity value at the current time t after Kalman filter prediction-update, T0 represents the actual measured temperature value of the sensor, and H0 represents the actual measured humidity value of the sensor.
[0016] Optionally, the abnormal detection processing steps of the storage medium are as follows:
[0017] Extract the characteristic data of the read-write rate, error rate, and response time from the storage medium log, calibrate it as Fd, and construct an isolation forest;
[0018] Randomly generate multiple isolation trees according to the characteristic data, and isolate the samples by randomly splitting the feature space;
[0019] Statistically calculate the average path length of the data points in the isolation tree, and calculate the anomaly score. Then the calculation formula for the anomaly score is and Wherein, AS(Fd) represents the anomaly score, E(d(Fd)) represents the mean of the path lengths in multiple isolation trees, d(Fd) represents the path length of the feature data point Fd in the isolation tree, c(n) represents the normalization factor, P(n) represents the harmonic number, and n represents the number of training samples of the feature data Fd in the isolation forest after random partitioning;
[0020] Set the anomaly threshold, designated as At. Judge by comparing the anomaly score with the anomaly threshold. When the anomaly score exceeds the anomaly threshold AS(Fd)>At, trigger the migration of nuclear data to the healthy storage node. Otherwise, store it normally. Among them, the anomaly threshold At = 0.75.
[0021] Specifically, the generation steps of the preprocessed nuclear data are as follows:
[0022] Align and fuse the denoised data Rnd'(t) after noise suppression processing and the compensated data Rnd”(t) after temperature and humidity drift compensation processing of the original data Rnd(t). Then the expression for data fusion is D Rnd = ω1Rnd'(t) + ω2Rnd”(t), and ω1 + ω2 = 1. In the formula, D Rnd represents the preprocessed nuclear data after fusing the denoised data Rnd'(t) and the compensated data Rnd”(t), and ω1 and ω2 respectively represent the weight coefficients for dynamically adjusting the sensor confidence corresponding to the denoised data Rnd'(t) and the compensated data Rnd”(t);
[0023] During the data alignment process, use the CRC check method to check the consistency of the data timestamps and eliminate the invalid data with timestamp conflicts. Then the expression for CRC check is CRC = CRC32(D Rnd (t)||T Rnd (t)), where CRC represents the check code, CRC32 represents the 32-bit check code for data concatenation, and T Rnd (t) represents the data aligned according to the timestamp t;
[0024] Associate the data block with the health status label of the storage medium to ensure that the data is only written into the healthy storage medium;
[0025] Encapsulate it into a structured data packet and transmit it to the quality monitoring module.
[0026] Optionally, the steps for detecting the abnormal state of the nuclear data are as follows:
[0027] Obtain the preprocessed nuclear data D Rnd and the reference nuclear data, designated as D Ref ;
[0028] The steps of using Dynamic Time Warping (DTW) combined with Long Short-Term Memory (LSTM) autoencoder to detect data interruption are as follows:
[0029] A1. Align the time series of the preprocessed core data D Rnd and the reference core data D Ref . Assume the length of the preprocessed core data D Rnd is u, then the preprocessed core data sequence is D Rnd,u = {D Rnd,1 , D Rnd,2 , …, D Rnd,u}. Assume the length of the reference core data D Ref is v, then the reference core data sequence is D Ref,v = D Ref,1 , D Ref,2 , …, D Ref,v}. Then calculate the DTW distance matrix to eliminate the influence of time axis offset. The calculation formula of the distance matrix is In the formula, DTW(u, v) represents the DWT distance matrix, || || represents the Euclidean distance, D Rnd,u represents the data at the u-th time point of the preprocessed core data D Rnd , and D Ref,v represents the data at the v-th time point of the reference core data D Ref ;
[0030] A2. Use the LSTM autoencoder for modeling. Among them, the encoder compresses the input preprocessed core data sequence D Rnd,u into a low-dimensional feature vector, and the decoder reconstructs the original sequence from the low-dimensional feature vector as
[0031] A3. Compare the input sequence D Rnd,u of the preprocessed core data with the reconstructed sequence to calculate the reconstruction error. The calculation formula of the reconstruction error is In the formula, e Rnd represents the reconstruction error value of the LSTM autoencoder, and U0 represents the total length of the time series;
[0032] The steps of improving the 3σ criterion algorithm to detect data jumps are as follows:
[0033] B1. Define a sliding window on the preprocessed core data D Rnd , and the window size is l. Calculate the mean and standard deviation within the window. The calculation formula of the mean is In the formula, represents the mean of the preprocessed core data D Rnd within the window, represents the data point corresponding to the preprocessed core data D Rnd at the current t0 moment, Denoted as the (t0 - l)-th data point within the window, where l represents the sliding window length; the calculation formula for the standard deviation is In the formula, σ t0 Denoted as the standard deviation of the preprocessed core data D Rnd within the window.
[0034] B2. The condition for determining a jump anomaly is
[0035] The steps of the MerkleTree hash chain verification algorithm for detecting data integrity are as follows:
[0036] C1. Divide the preprocessed core data D Rnd into m blocks, and generate a hash value H m (D Rnd ) for each block, and gradually merge them to generate the root hash. Then the hash value expression for the entire block is RH = SHA - 256(H(H1(D Rnd ) || H2(D Rnd ) ||... || H m (D Rnd ))). In the formula, RH is denoted as the hash value at the top of the blockchain - like structure, SHA - 256 is denoted as the SHA - 256 algorithm that stores the hash value calculated for each block in the block header, H is denoted as the hash function, and || is denoted as the data block concatenation operation.
[0037] C2. Compare and analyze the root hash RH calculated in real - time with the stored reference root hash, denoted as RH0, to generate a comparison result RH'. When RH' = 1, that is, when RH ≠ RH0, the comparison result is inconsistent, and it is determined that the data has been tampered with. When RH' = 0, that is, when RH = RH0, the comparison result is consistent, and it is determined that the data has not been tampered with.
[0038] Optionally, the process of marking the problem core data is as follows:
[0039] Perform weighted fusion processing on the anomaly detection results of the preprocessed core data by DTW combined with LSTM auto - encoder, 3σ criterion algorithm, and MerkleTree hash chain to calculate the anomaly score. The calculation formula for the anomaly score is And In the formula, C f is denoted as the anomaly score of the preprocessed core data D Rnd after the anomaly detection results, respectively denote the weight coefficients corresponding to the anomaly detection results of the preprocessed core data by DTW combined with LSTM auto - encoder, 3σ criterion algorithm, and MerkleTree hash chain;
[0040] Set the confidence level of the anomaly detection, denoted as C0, and perform comparison and analysis based on the anomaly score and the confidence level C0. When Cf When C0, if the data is determined to be abnormal, it is marked as problematic data;
[0041] Add a metadata label as marked data at the abnormal data point where the problematic data is detected, encapsulate the marked data in JSON format, and sort it from high to low confidence. Prioritize the transmission of the problem core data with high confidence, calibrated as D' Rnd .
[0042] Optionally, the steps for imputing missing data in the problem core data are as follows:
[0043] Extract the valid data points {Z(s Rnd )|k = 1, 2,..., K} around the missing point from the marked problem core data D' k , where s k is the position coordinate of the k-th in the nuclear facility;
[0044] Calculate the semivariogram of the valid data points Z(s k ) to describe the change of data spatial correlation with distance. The calculation formula of the semivariogram is and k, k' ∈ {1, 2,..., K}. In the formula, γ(d) represents the semivariance value between the valid data points Z(s k ) and Z(s k′ ), d represents the Euclidean distance, log d Z(s k ) represents the number of pairs of data points at a distance of d, Z(s k ) represents the value of the valid data point with spatial coordinate s k , and Z(s k′ ) represents the value of the valid data point with spatial coordinate s k′ ;
[0045] Construct the Kriging equations to solve the weight coefficients, minimizing the interpolation error and satisfying the unbiasedness constraint. Then, the expression of the Kriging equations is and In the formula, λ k′ represents the weight coefficient corresponding to the semivariogram γ(d), δ represents the Lagrange multiplier, and s0 represents the spatial coordinate of the missing point;
[0046] Use the weights and valid data points to calculate the estimated value of the missing point and perform data interpolation. The calculation formula of the estimated value of the missing point is In the formula, represents the estimated value of the missing point s0.
[0047] Optionally, the steps for correcting incorrect data in the problem core data are as follows:
[0048] According to the reference nuclear data D Ref Extract normal data as the training set {(A Rnd,u , D Rnd,u ),} and train the SVR model. Among them, A Rnd,u represents the feature data of input time, temperature and humidity, and D Rnd,u represents the data at the u-th time point of the preprocessed nuclear data D Rnd ;
[0049] Solve the support SVR vectors and regression hyperplane by optimizing the SVR objective function. The expression of the SVR objective function is In the formula, W represents the weight of the regression hyperplane, B represents the bias of the regression hyperplane, ζ and ξ both represent the slack variables of the SVR model, χ represents the penalty factor for balancing the complexity and error tolerance of the SVR model, u and U respectively represent the u-th time point and the total number of time points;
[0050] Define the constraint conditions to limit the change range of the predicted value. The expression of the constraint conditions is And In the formula, φ represents the kernel function, ∈ represents the width of the insensitive region, which controls the regression accuracy, represents the average value of the u feature data A Rnd,u , || ||2 2 represents the square of the 2-norm, and T represents the transpose;
[0051] Input the problem nuclear data D' Rnd value of the detected anomaly into the SVR model and output the corrected value. The calculation formula of the corrected value is Y Rnd = W T ·φ(D′ Rnd ) + B. In the formula, Y Rnd represents the corrected value output by the SVR model.
[0052] Optionally, the acquisition logic of the standardized nuclear data is as follows:
[0053] Calculate the global mean μ Ref and standard deviation σ Ref according to the reference nuclear data D Ref ;
[0054] Repair the calculated missing data and the corrected value Y Rnd to obtain the repaired nuclear data, calibrated as R Rnd , and align it with the reference data D Ref by timestamp and spatial position;
[0055] Perform Z-score normalization on the repaired data to eliminate the dimension difference and ensure consistency with the reference data. The expression for the normalized data is In the formula, Z Rnd represents the normalized core data after Z-score normalization processing, μ Ref represents the global mean of the reference core data D Ref , σ Ref represents the standard deviation of the reference core data D Ref , and R Rnd represents the repaired core data;
[0056] Perform consistency check analysis on the normalized core data Z Rnd and the reference core data D Ref to determine whether to perform secondary repair.
[0057] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the above-mentioned nuclear database management system based on nuclear dynamic detection technology.
[0058] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the above-mentioned nuclear database management system based on nuclear dynamic detection technology.
[0059] In the above technical solution, the technical effects and advantages provided by the present invention are:
[0060] The present invention collects raw data by deploying high-precision sensors in nuclear facilities, performs preprocessing by combining noise suppression and drift compensation algorithms with storage medium anomaly detection processing, ensuring the accuracy and reliability of nuclear data collection, effectively reducing the influence of environmental factors and noise on the data. By reading historical reference nuclear data and using dynamic time warping DTW combined with LSTM autoencoders, improved 3σ criterion algorithm and MerkleTree hash chain verification algorithm to detect data anomalies, it can timely and accurately discover anomalies in nuclear data, providing a basis for data repair and ensuring data quality; using the Kriging spatial interpolation algorithm to interpolate missing data and the support vector regression SVR model to correct incorrect data, successfully repairing problematic nuclear data and restoring the integrity and accuracy of the data, making the data consistent with the reference data; and applying distributed disaster-tolerant storage and compression storage technologies, combined with using B+ tree indexing and R-tree radiation field spatial indexing strategies for multi-dimensional queries, reducing the resource consumption of data storage, improving the security and availability of data storage, and simultaneously achieving fast retrieval of data. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0062] Figure 1 This is a block diagram of a nuclear database management system based on nuclear dynamic detection technology for the present invention. Detailed implementation manners
[0063] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these example embodiments are provided so that this disclosure will be more thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art.
[0064] The present invention provides a nuclear database management system based on nuclear dynamic detection technology as shown in Figure 1 which includes a data acquisition and processing module: deploying high-precision sensors on nuclear facilities to collect raw data of dynamic detection nuclear data and environmental parameter data of nuclear facilities in real time. In view of the influence of environmental parameter data on dynamic detection nuclear data, noise suppression and drift compensation algorithms are used to preprocess the raw data to obtain relatively stable and reliable nuclear data, reduce data generated due to problems such as interrupted, incorrect acquisition signals or damaged storage media caused by environmental interference, provide high-quality input for subsequent modules, and combine anomaly detection processing of the storage media to generate preprocessed nuclear data, and then transmit it to the data quality monitoring module;
[0065] Specifically, the preprocessing steps of the raw data are as follows:
[0066] The sensor collects raw data in real time and calibrates it as Rnd(t), including dynamic detection nuclear data and environmental parameter data of nuclear facilities. Noise suppression processing is performed on the raw data using wavelet decomposition combined with an adaptive threshold processing algorithm. Among them, the raw data Rnd(t) is decomposed into multi-scale wavelet coefficients, threshold filtering processing is performed on the high-frequency noise coefficients, and then the filtered coefficients are used to reconstruct the denoised signal to generate denoised data. The calculation formula for noise suppression is and T(W i,j ) = sign(W i,j ) * max(|W i,j | - λ, 0), where Rnd'(t) represents the denoised data, W i,j represents the wavelet coefficient corresponding to decomposing the raw data Rnd(t) to scale i and translation j, I represents the maximum decomposition layer, and T(W i,j) is represented as an adaptive threshold function, ψ i,j (Rnd(t)) is represented as a Daubechies wavelet basis function, λ is represented as a soft threshold parameter, and sign is represented as a sign function for extracting the wavelet coefficient W of the input value i,j of the symbol information;
[0067] In the environment where the nuclear facility is located, a temperature and humidity drift compensation model is established based on the offset of the data collected by the sensor. The original data Rnd(t) is processed iteratively using Kalman filtering. The true state is estimated through the prediction-update step, and the original data is corrected according to the estimated values of the temperature and humidity states to obtain the compensated data,
[0068] Then the calculation formula for temperature and humidity drift compensation is Rnd”(t) = Rnd(t) - α·ΔT - β·ΔH, and ΔT = T t - T0, ΔH = H t - H0. In the formula, Rnd”(t) is represented as the compensated data, α is represented as the sensitivity coefficient of the experimentally calibrated temperature to the sensor, β is represented as the sensitivity coefficient of the experimentally calibrated humidity to the sensor, ΔT is represented as the temperature change, ΔH is represented as the humidity change, T t represents the temperature value at the current time t after Kalman filter prediction-update, H t represents the humidity value at the current time t after Kalman filter prediction-update, T0 represents the actual measured temperature value of the sensor, and H0 represents the actual measured humidity value of the sensor;
[0069] The steps for establishing the temperature and humidity drift compensation model are as follows:
[0070] Obtain and define the state variables of temperature, humidity, and sensor offset;
[0071] Use the Kalman filter algorithm to initialize the state estimate and the estimated error covariance matrix, and sequentially determine the state transition matrix, control input matrix, process noise covariance matrix, observation matrix, and observation noise covariance matrix. Among them, the state transition matrix describes the relationship between temperature, humidity, and sensor offset over time, and the observation matrix establishes the relationship between the state and the observed value;
[0072] The prediction step includes calculating the prior state estimate at the current time based on the posterior state estimate at the previous time, and calculating the prior estimated error covariance at the current time based on the posterior estimated error covariance at the previous time;
[0073] The update steps include calculating the Kalman gain, updating the state, and updating the error covariance. Among them, the Kalman gain is calculated based on the prior estimated error covariance, the observation matrix, and the observation noise covariance matrix; the posterior state estimate at the current moment is obtained by combining the prior state estimate, the Kalman gain, and the current observation value; the posterior estimated error covariance at the current moment is updated according to the Kalman gain, the observation matrix, and the prior estimated error covariance.
[0074] The temperature and humidity change amounts are calculated using the temperature and humidity in the updated state estimate, and the original measurement values of the sensor are compensated to obtain the compensated measurement values.
[0075] Specifically, the abnormal detection processing steps of the storage medium are as follows:
[0076] Extract the characteristic data of the read / write rate, error rate, and response time from the storage medium log, label it as Fd, and construct an isolation forest.
[0077] Randomly generate multiple isolation trees according to the characteristic data, and isolate the samples by randomly dividing the feature space.
[0078] The average path length of the data points in the isolation tree is statistically calculated, and the anomaly score is calculated. The calculation formula of the anomaly score is And In the formula, AS(Fd) represents the anomaly score, E(d(Fd)) represents the mean of the path lengths in multiple isolation trees, d(Fd) represents the path length of the characteristic data point Fd in the isolation tree, c(n) represents the normalization factor, P(n) represents the harmonic number, and n represents the number of training samples of the characteristic data Fd in the isolation forest after random division.
[0079] Set the anomaly threshold, label it as At, and make a judgment by comparing the anomaly score with the anomaly threshold. When the anomaly score exceeds the anomaly threshold AS(Fd)>At, trigger the nuclear data migration to the healthy storage node. Otherwise, store it normally. Among them, the anomaly threshold At = 0.75.
[0080] Specifically, the generation steps of the preprocessed nuclear data are as follows:
[0081] Align and fuse the denoised data Rnd'(t) after noise suppression processing and the compensated data Rnd”(t) after temperature and humidity drift compensation processing of the original data Rnd(t). The expression of data fusion is D Rnd = ω1Rnd'(t)+ω2Rnd”(t), and ω1+ω2 = 1. In the formula, D RndIt represents the preprocessed core data after the fusion process of the denoised data Rnd'(t) and the compensated data Rnd”(t). ω1 and ω2 respectively represent the weight coefficients of the sensor confidence dynamic adjustment corresponding to the denoised data Rnd'(t) and the compensated data Rnd”(t).
[0082] During the data alignment process, the CRC check method is used to check the data timestamp consistency and eliminate the invalid data with timestamp conflicts. Then the expression of the CRC check is CRC = CRC32(D Rnd (t)||T Rnd (t)), where CRC represents the check code, CRC32 represents the 32-bit check code for data concatenation, and T Rnd (t) represents the data aligned according to the timestamp t.
[0083] Associate the data block with the health status label of the storage medium to ensure that the data is only written into the healthy storage medium.
[0084] Encapsulate it into a structured data packet and transmit it to the quality monitoring module.
[0085] Furthermore, the environment of nuclear facilities is complex, and the data collected by sensors is easily affected by noise interference and environmental factors, resulting in inaccurate data. Noise will cause data fluctuations and mask the true signal characteristics. Therefore, wavelet denoising and Kalman filtering are important measures to eliminate the signal distortion caused by environmental interference. The accuracy and stability of the preprocessed core data are significantly improved, which can more truly reflect the operating status and environmental parameters of nuclear facilities, providing a reliable basis for subsequent data analysis and decision-making. Establish an isolated forest to monitor the health status of the storage medium in real time, avoid writing data to faulty nodes, and thus ensure the integrity and continuity of the data; ensure the physical rationality and logical integrity of the preprocessed core data through data fusion and verification.
[0086] Data quality monitoring module: Read the reference core data in the historical database, receive the preprocessed core data in the data acquisition and processing module, and use the dynamic time warping DTW combined with the LSTM autoencoder, improved 3σ criterion algorithm, and MerkleTree hash chain verification algorithm to detect the abnormal state of the preprocessed core data in real time and mark the problematic core data, so as to timely discover the interruption, jump, and tampering abnormal situations in the data, avoid abnormal and incorrect data from entering the database, ensure the continuity and rationality of the data, improve the integrity and reliability of the data, and transmit it to the data repair and reconstruction module;
[0087] Specifically, the steps for detecting the abnormal state of nuclear data are as follows:
[0088] Obtain the preprocessed core data D Rnd and the reference core data, calibrated as D Ref ;
[0089] The steps of using Dynamic Time Warping (DTW) combined with Long Short-Term Memory (LSTM) autoencoder to detect data interruption are as follows:
[0090] A1. Align the time series of the preprocessed kernel data D Rnd and the reference kernel data D Ref . Suppose the length of the preprocessed kernel data D Rnd is u, then the preprocessed kernel data sequence is D Rnd,u = {D Rnd,1 , D Rnd,2 , …, D Rnd,u}. Suppose the length of the reference kernel data D Ref is v, then the reference kernel data sequence is D Ref,v = D Ref,1 , D Ref,2 , …, D Ref,v}. Then calculate the DTW distance matrix to eliminate the influence of time axis offset. The calculation formula of the distance matrix is In the formula, DTW(u, v) represents the DWT distance matrix, || || represents the Euclidean distance, D Rnd,u represents the data at the u-th time point of the preprocessed kernel data D Rnd , and D Ref,v represents the data at the v-th time point of the reference kernel data D Ref ;
[0091] A2. Use the LSTM autoencoder for modeling. Among them, the encoder compresses the input preprocessed kernel data sequence D Rnd,u into a low-dimensional feature vector, and the decoder reconstructs the original sequence from the low-dimensional feature vector as
[0092] A3. Compare the input sequence D Rnd,u of the preprocessed kernel data with the reconstructed sequence to calculate the reconstruction error. The calculation formula of the reconstruction error is In the formula, e Rnd represents the reconstruction error value of the LSTM autoencoder, and U0 represents the total length of the time series;
[0093] The steps of using the improved 3σ criterion algorithm to detect data jumps are as follows:
[0094] B1. Define a sliding window on the preprocessed kernel data D Rnd , and the window size is l. Calculate the mean and standard deviation within the window. The calculation formula of the mean is In the formula, μ t0 represents the mean of the preprocessed kernel data D Rnd within the window, represents the preprocessed kernel data D at the current t0 momentRnd The corresponding data point is expressed as the (t0 - l)-th data point within the window, where l represents the sliding window length; the calculation formula for the standard deviation is In the formula,[[]] represents the standard deviation of the preprocessed core data D within the window Rnd ;
[0095] B2. The condition for determining a jump anomaly is
[0096] The steps of the MerkleTree hash chain verification algorithm for detecting data integrity are as follows:
[0097] C1. Divide the preprocessed core data D Rnd into m blocks, and generate a hash value H m (D Rnd ) for each block. Then, merge them step by step to generate the root hash. The hash value expression for the entire block is RH = SHA - 256(H(H1(D Rnd ) || H2(D Rnd ) ||... || H m (D Rnd ))). In the formula, RH represents the hash value at the top of the blockchain - like structure, SHA - 256 represents the SHA - 256 algorithm that stores the hash value calculated for each block in the block header, H represents the hash function, and || represents the data block concatenation operation
[0098] C2. Compare and analyze the root hash RH calculated in real - time with the stored reference root hash, designated as RH0, to generate a comparison result RH'. When RH' = 1, that is, when RH ≠ RH0, the comparison result is inconsistent, indicating data tampering; when RH' = 0, that is, when RH = RH0, the comparison result is consistent, indicating that the data has not been tampered with.
[0099] Specifically, the process of marking the problem core data is as follows:
[0100] Perform weighted fusion processing on the anomaly detection results of the preprocessed core data by DTW combined with the LSTM auto - encoder, the 3σ criterion algorithm, and the MerkleTree hash chain, and calculate the anomaly score. The calculation formula for the anomaly score is And In the formula, C f represents the anomaly score of the preprocessed core data D Rnd after the anomaly detection results, respectively represent the weight coefficients corresponding to the anomaly detection results of the preprocessed core data by DTW combined with the LSTM auto - encoder, the 3σ criterion algorithm, and the MerkleTree hash chain;
[0101] Set the confidence level for anomaly detection as C0, and conduct a comparative analysis based on the anomaly score and the confidence level C0. When C f > C0, it is determined that the data is abnormal, and it is marked as problem data;
[0102] Add metadata tags to the abnormal data points of the detected problem data as marked data, encapsulate the marked data in JSON format, and sort it in descending order of confidence level. Give priority to transmitting the problem core data with high confidence level, calibrated as D' Rnd .
[0103] Furthermore, by using the Dynamic Time Warping (DTW) algorithm, the preprocessed core data is compared with the reference core data to effectively handle the stretching and bending problems of data on the time axis, thereby capturing the similarity features between data. Then, combined with the LSTM autoencoder, feature learning and reconstruction are performed on the core data. By comparing the differences between the input data and the reconstructed data, potential abnormal patterns are detected; the improved 3σ criterion algorithm can more accurately identify the abnormal values in the core data, improving the sensitivity and accuracy of anomaly detection; the MerkleTree hash chain verification algorithm is used to verify the integrity of the core data to prevent data from being tampered with during transmission and storage. The synergistic effect of these fusion algorithms on anomaly detection of preprocessed core data can accurately mark the abnormal status and problem data in the core data, not only improving the efficiency and accuracy of core data quality monitoring, ensuring the reliability and security of core data, but also providing a high-quality data foundation for subsequent data analysis and decision-making, thus enhancing the performance and stability of the entire core data processing system.
[0104] Data repair and reconstruction module: Receive the marked problem core data in the data quality monitoring module and trigger the repair mechanism, including using the Kriging spatial interpolation algorithm to interpolate the missing data in the problem core data, and using the Support Vector Regression (SVR) model to correct the error data in the problem core data, obtaining the standardized core data that is consistent with the reference core data after recovery, which is used to timely repair and reconstruct the detected abnormal data, make up for the loss of data and correct the error data, restore the integrity and physical consistency of the data, ensure the reliability of the database, reduce the management complexity, and transmit it to the data storage and management module;
[0105] Specifically, the steps for interpolating the missing data in the problem core data are as follows:
[0106] Extract the valid data points {Z(s Rnd )|k = 1, 2,..., K} around the missing points from the marked problem core data D' k , where s k is the position coordinate of the kth in the nuclear facility;
[0107] Calculate the semivariogram of the effective data points Z(s k ), which describes the change of data spatial correlation with distance. Then the calculation formula of the semivariogram is and k, k' ∈ {1, 2, …, K}. In the formula, γ(d) represents the semivariance value between the effective data points Z(s k ) and Z(s k′ ), d represents the Euclidean distance, log d Z(s k ) represents the number of data point pairs with a distance of d, Z(s k ) represents the value of the effective data point with the spatial coordinate s k , and Z(s k′ ) represents the value of the effective data point with the spatial coordinate s k′ ;
[0108] Construct the Kriging equations to solve the weight coefficients to minimize the interpolation error and satisfy the unbiasedness constraint. Then the expression of the Kriging equations is and In the formula, λ k′ represents the weight coefficient corresponding to the semivariogram γ(d), which is used to minimize the interpolation error and satisfy the unbiasedness constraint. δ represents the Lagrange multiplier for the unbiasedness constraint, and s0 represents the spatial coordinate of the missing point;
[0109] Use the weights and effective data points to calculate the estimated value of the missing point and perform data interpolation. Then the calculation formula of the estimated value of the missing point is In the formula, represents the estimated value of the missing point s0, and fills the missing value through spatial correlation to ensure the integrity of the continuous data in the nuclear data space.
[0110] Specifically, the steps for correcting the wrong data in the problem nuclear data are as follows:
[0111] Extract the normal data from the reference nuclear data D Ref as the training set {(A Rnd,u , D Rnd,u )} to train the SVR model. Among them, A Rnd,u represents the characteristic data of input time, temperature and humidity, and D Rnd,u represents the data at the u-th time point of the preprocessed nuclear data D Rnd ;
[0112] Solve the support vectors and regression hyperplane of the SVR by optimizing the SVR objective function. Then the expression of the SVR objective function is Wherein, W represents the weight of the regression hyperplane, B represents the bias of the regression hyperplane, ζ and ξ both represent the slack variables of the SVR model, χ represents the penalty factor for balancing the complexity and error tolerance of the SVR model, u and U respectively represent the u-th time point and the total number of time points;
[0113] Define the constraint conditions to limit the change range of the predicted value, then the expression of the constraint conditions is And Wherein, φ represents the kernel function, ∈ represents the width of the insensitive region, which controls the regression accuracy, Represents the average value of u feature data A Rnd,u The square of the 2-norm is represented by || ||2 2 T represents the transpose;
[0114] Input the problem kernel data D' Rnd Value to the SVR model and output the corrected value. Then the calculation formula of the corrected value is Y Rnd = W T ·φ(D′ Rnd ) + B. Wherein, Y Rnd Represents the corrected value output by the SVR model. By using the rules and constraint conditions of the reference kernel data in the historical database, the jump or error data is corrected to restore its physical reasonableness.
[0115] Specifically, the acquisition logic of the standardized kernel data is as follows:
[0116] According to the reference kernel data D Ref Calculate the global mean μ Ref And the standard deviation σ Ref ;
[0117] Repair the calculated missing data And the corrected value Y Rnd To obtain the repaired kernel data, calibrated as R Rnd , Align with the reference data D Ref According to the timestamp and spatial position;
[0118] Perform Z-score standardization on the repaired data to eliminate the dimension difference and ensure consistency with the reference data distribution. Then the expression of the standardized data is Wherein, Z Rnd Represents the standardized kernel data after Z-score standardization processing, μ Ref Represents the global mean of the reference kernel data D Ref σ Ref Represents the standard deviation of the reference kernel data D Ref R Rnd Represents the repaired kernel data;
[0119] Standardize the nuclear data Z Rnd with the reference nuclear data D Ref Perform a consistency check analysis and determine whether to perform secondary repair.
[0120] Furthermore, by applying the Kriging spatial interpolation algorithm, a semi-variogram is constructed to describe the spatial correlation of the data, the Kriging equations are solved to obtain the weight coefficients, and then these weights and valid data points are used to estimate and interpolate the missing points to fill the data gaps. For the error data in the problem nuclear data, the support vector regression SVR model is used for correction, and then the repaired nuclear data is Z-score standardized to eliminate the dimension difference and ensure consistency with the reference data distribution. At the same time, a consistency check is performed. This not only ensures the integrity and accuracy of the nuclear data, restores the physical rationality of the data, but also provides a high-quality data foundation for subsequent data storage, analysis, and application, improving the reliability and availability of the entire nuclear data processing system.
[0121] Data storage management module: Receive the repaired standardized nuclear data in the data repair and reconstruction module, and use distributed disaster-tolerant storage and compressed storage technologies to achieve high-reliability storage and fast retrieval functions while minimizing the resource storage of the standardized nuclear data. As the archived nuclear data, it is used to provide high-reliability and anti-damage data storage services, prevent data loss, ensure data security and accessibility, and optimize the utilization rate of storage resources;
[0122] Specifically, the storage processing steps of the standardized nuclear data are as follows:
[0123] Divide the standardized nuclear data Z Rnd into q original data blocks, and the q original data blocks are represented as Z Rnd,1 、Z Rnd,2 、…、Z Rnd,q};
[0124] Use the RS erasure code algorithm to generate p redundant check blocks, and the p redundant check blocks are {RS1, RS2, …, RS p}, and the total number of storage blocks g = q + p. Among them, the expression of the RS erasure code algorithm is In the formula, (′Z Rnd,1 , ′Z Rnd,2 , …, ′Z Rnd,g ) represents the encoded data blocks including the original data blocks and the redundant check blocks;
[0125] Disperse the g data blocks and store them on different disk array nodes to ensure that any q blocks can recover the original data;
[0126] Periodically check the survival status of storage nodes for disaster recovery verification. If the number of failed nodes is greater than p, trigger dynamic data block migration;
[0127] For the standardized nuclear data Z Rnd Perform wavelet decomposition to obtain multi-scale coefficients;
[0128] Retain the low-frequency feature components, discard the high-frequency noise coefficients below the wavelet compression threshold, and only store the wavelet coefficients and position indexes that are greater than the wavelet compression threshold and retained;
[0129] Error check between the decompressed data and the original data;
[0130] And use a B+ tree index combined with an R-tree radiation field spatial index to establish a multi-dimensional index, attach labels to the data blocks, generate a hash value for each data block, and associate it with the index to generate archived nuclear data. While optimizing the fast index, ensure retrieval integrity.
[0131] Furthermore, the distributed disaster recovery storage technology disperses the data storage on multiple nodes, avoiding the risk of data loss caused by single-point failures and ensuring the high reliability of the data. The compression storage technology processes the standardized nuclear data, removes redundant information in the data, and stores the data in a more compact form, thus effectively reducing the occupancy of data storage space. The combination of the two has a fast retrieval function, realizes the storage of standardized nuclear data in the way of minimizing resource occupancy, and at the same time ensures the high-reliability storage and fast retrieval of the data. Taking it as archived nuclear data not only reduces the data storage cost, improves the security and availability of the data, but also provides efficient support for subsequent data access and analysis, making the entire nuclear data management system more stable and efficient.
[0132] Data query and analysis module: Nuclear technicians use a B+ tree index combined with an R-tree radiation field spatial index strategy to perform multi-dimensional conditional queries on the archived nuclear data, can directly retrieve relevant nuclear data, and generate radiation reports through visualization analysis technology, and then assist in nuclear facility management decision-making analysis, providing users with efficient data query and visualization display and analysis functions, enabling users to timely discover problems in the data, and providing decision support for nuclear facility safety management, accident traceability, and nuclear data management.
[0133] Specifically, the visualization analysis steps of the radiation report are as follows:
[0134] The visualized radiation report includes a radiation heat map, which is calculated and generated by the kernel density estimation algorithm;
[0135] Nuclear technicians index the archived nuclear data and extract the radiation intensity and the spatial coordinates of the nuclear facilities from the nuclear data query results;
[0136] The Silverman rule is used to calculate the optimal bandwidth to balance smoothness and resolution;
[0137] A Gaussian kernel function is superimposed on each spatial coordinate point to calculate the radiation density value;
[0138] Rendering technology is used to map the radiation density value into a color gradient to represent the radiation intensity, thereby visually displaying the radiation leakage hotspots;
[0139] By dynamically analyzing the time series trend and the time - space correlation, the change trend of the radiation intensity is traced and sudden anomalies are identified.
[0140] Furthermore, by using the B+ tree index combined with the R - tree radiation field spatial index strategy, multi - dimensional conditional queries can be efficiently processed. It can speed up the retrieval of archived nuclear data and accurately retrieve relevant nuclear data. By applying visualization analysis technology, the nuclear data is presented in an intuitive heat map. This not only clearly shows key information such as the distribution and change trend of nuclear data, but also makes it easier for nuclear technology personnel to understand and analyze the data, thus helping nuclear technology personnel to more accurately grasp the operating conditions and radiation situation of nuclear facilities, and then make scientific and reasonable management decisions, improving the scientific nature and effectiveness of nuclear facility management.
[0141] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The preset parameters in the formulas are set by technicians in this field according to the actual situation.
[0142] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains a set of one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, or a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0143] It should be understood that in various embodiments of the present application, the sequence numbers of the above processes do not indicate the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0144] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0145] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all such changes or substitutions should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A nuclear database management system for nuclear dynamic detection technology, characterized in that, It includes a data acquisition and processing module: High-precision sensors are deployed on nuclear facilities to collect the original data of dynamic detection nuclear data and environmental parameter data of nuclear facilities in real time. Considering the influence of environmental parameter data on dynamic detection nuclear data, noise suppression and drift compensation algorithms are used to preprocess the original data, combined with the anomaly detection and processing of the storage medium, to generate preprocessed nuclear data, and then transmit it to the data quality monitoring module; Data quality monitoring module: Read the reference nuclear data in the historical database, receive the preprocessed nuclear data in the data acquisition and processing module, and use dynamic time warping (DTW) combined with LSTM autoencoders, improved 3σ criterion algorithm and MerkleTree hash chain verification algorithm to detect the abnormal state of the preprocessed nuclear data in real time, mark the problematic nuclear data, and transmit it to the data repair and reconstruction module; Data repair and reconstruction module: Receive the marked problematic nuclear data in the data quality monitoring module and trigger the repair mechanism, including using the Kriging spatial interpolation algorithm to interpolate the missing data in the problematic nuclear data, and using the support vector regression (SVR) model to correct the error data in the problematic nuclear data, to obtain the standardized nuclear data that is consistent with the reference nuclear data after recovery, and transmit it to the data storage and management module; Data storage and management module: Receive the repaired standardized nuclear data in the data repair and reconstruction module, and use distributed disaster recovery storage and compression storage technologies to store the standardized nuclear data with minimal resources while achieving high-reliability storage and fast retrieval functions, as the archived nuclear data; Data query and analysis module: Nuclear technicians use the B+ tree index combined with the R-tree radiation field spatial index strategy to perform multi-dimensional conditional queries on the archived nuclear data, can directly retrieve relevant nuclear data, and generate radiation reports through visualization analysis technology, thereby assisting in the decision-making analysis of nuclear facility management.
2. The nuclear database management system for nuclear dynamic detection technology according to claim 1, characterized in that, The preprocessing steps of the original data are as follows: The sensor collects raw data in real time and calibrates it as Rnd(t), including the dynamic detection nuclear data and environmental parameter data of nuclear facilities. The raw data is subjected to noise suppression processing using the wavelet decomposition combined with the adaptive threshold processing algorithm. Among them, the raw data Rnd(t) is decomposed into multi-scale wavelet coefficients, the high-frequency noise coefficients are subjected to threshold filtering processing, and then the filtered coefficients are used to reconstruct the denoised signal to generate denoised data. The calculation formula for noise suppression is And T(W i,j ) = sign(W i,j ) * max(|W i,j | - λ, 0). In the formula, Rnd'(t) represents the denoised data, W i,j represents the wavelet coefficient corresponding to decomposing the raw data Rnd(t) to scale i and translation j. I represents the maximum decomposition level, T(W i,j ) represents the adaptive threshold function, ψ i,j (Rnd(t)) represents the Daubechies wavelet basis function, λ represents the soft threshold parameter, and sign represents the sign function that extracts the sign information of the input value wavelet coefficient W i,j ; In the environment where the nuclear facility is located, a temperature and humidity drift compensation model is established according to the offset of the data collected by the sensor, and the original data Rnd(t) is processed by Kalman filter iteration. The true state is estimated through the prediction-update step, and the original data is corrected according to the estimated value of the temperature and humidity state to obtain the compensated data. The calculation formula for temperature and humidity drift compensation is Rnd”(t) = Rnd(t) - α·ΔT - β·ΔH, and ΔT = T t - T0, ΔH = H t - H0. In the formula, Rnd”(t) represents the compensated data, α represents the sensitivity coefficient of the sensor to the experimentally calibrated temperature, β represents the sensitivity coefficient of the sensor to the experimentally calibrated humidity, ΔT represents the temperature change, ΔH represents the humidity change, T t represents the temperature value at the current time t after Kalman filter prediction - update, H t represents the humidity value at the current time t after Kalman filter prediction - update, T0 represents the actual measured temperature value of the sensor, and H0 represents the actual measured humidity value of the sensor.
3. The nuclear database management system for nuclear dynamic detection technology according to claim 2, characterized in that, The anomaly detection and processing steps of the storage medium are as follows: Extract the characteristic data of read / write rate, error rate, and response time from the storage medium log, label it as Fd, and construct an isolation forest; Randomly generate multiple isolation trees according to the characteristic data, and isolate the samples by randomly dividing the feature space; Statistically calculate the average path length of data points in the isolation tree and compute the anomaly score. The formula for the anomaly score is and where AS(Fd) represents the anomaly score, E(d(Fd)) represents the mean of the path lengths in multiple isolation trees, d(Fd) represents the path length of the feature data point Fd in the isolation tree, c(n) represents the normalization factor, P(n) represents the harmonic number, and n represents the number of training samples of the feature data Fd in the isolation forest after random partitioning; Set the anomaly threshold, label it as At, and make a judgment by comparing the anomaly score with the anomaly threshold. When the anomaly score exceeds the anomaly threshold AS(Fd)>At, trigger the migration of nuclear data to a healthy storage node, otherwise, store it normally, where the anomaly threshold At = 0.
75. Specifically, the generation steps of the preprocessed nuclear data are as follows: Align and fuse the denoised data Rnd'(t) after noise suppression processing and the compensated data Rnd”(t) after temperature and humidity drift compensation processing of the original data Rnd(t). Then the expression for data fusion is D Rnd = ω1Rnd'(t) + ω2Rnd”(t), and ω1 + ω2 = 1. In the formula, D Rnd represents the preprocessing core data after the fusion processing of the denoised data Rnd'(t) and the compensated data Rnd”(t). ω1 and ω2 respectively represent the weight coefficients corresponding to the dynamic adjustment of the sensor confidence for the denoised data Rnd'(t) and the compensated data Rnd”(t); During the data alignment process, the CRC check method is used to check the consistency of the data timestamps, and the invalid data with timestamp conflicts is eliminated. Then the CRC check expression is CRC = CRC32(D Rnd (t)||T Rnd (t)), where CRC represents the check code, CRC32 represents the 32-bit check code for data splicing, and T Rnd (t) represents the data aligned according to the timestamp t; Associate the data block with the health status label of the storage medium to ensure that the data is only written into the healthy storage medium; Encapsulate it into a structured data packet and transmit it to the quality monitoring module.
4. The nuclear database management system for nuclear dynamic detection technology according to claim 3, characterized in that The steps for detecting the abnormal state of the nuclear data are as follows: Obtain preprocessed nuclear data D Rnd and reference nuclear data, calibrated as D Ref ; The steps of detecting data interruption by combining Dynamic Time Warping (DTW) with LSTM autoencoder are as follows: A1. Align the preprocessed nuclear data D Rnd and the reference nuclear data D Ref in time series. Assume the length of the preprocessed nuclear data D Rnd is u, then the preprocessed nuclear data sequence is D Rnd,u ={D Rnd,1 , D Rnd,2 , …, D Rnd,u}. Assume the length of the reference nuclear data D Ref is v, then the reference nuclear data sequence is D Ref,v ={D Ref,1 , D Ref,2 , …, D Ref,v}. Then calculate the DTW distance matrix to eliminate the influence of the time - axis offset. The calculation formula of the distance matrix is In the formula, DTW(u, v) represents the DWT distance matrix, |||| represents the Euclidean distance, D Rnd,u represents the data at the u - th time point of the preprocessed nuclear data D Rnd , and D Ref,v represents the data at the v - th time point of the reference nuclear data D Ref ; A2. Model using an LSTM autoencoder, where the encoder compresses the input preprocessed core data sequence D Rnd,u into a low-dimensional feature vector, and the decoder reconstructs the original sequence from the low-dimensional feature vector as A3. Compare the input sequence D of the preprocessed nuclear data Rnd,u with the reconstructed sequence and calculate the reconstruction error. The calculation formula for the reconstruction error is where e Rnd represents the reconstruction error value of the LSTM autoencoder, and U0 represents the total length of the time series; The steps of detecting data jumps by improving the 3σ criterion algorithm are as follows: B1. Define a sliding window on the preprocessed nuclear data D Rnd with a window size of l, and calculate the mean and standard deviation within the window. The formula for the mean is where represents the mean of the preprocessed nuclear data D Rnd within the window, represents the data point corresponding to the preprocessed nuclear data D Rnd at the current time t0, represents the (t0 - l)-th data point within the window, and l represents the length of the sliding window. The formula for the standard deviation is where represents the standard deviation of the preprocessed nuclear data D Rnd within the window; B2. The condition for determining a jump anomaly is The steps of detecting data integrity by MerkleTree hash chain verification algorithm are as follows: C1. Divide the preprocessed nuclear data D Rnd into m blocks, and generate a hash value H for each block m (D Rnd ), and merge them step by step to generate the root hash. Then the hash value expression of the entire block is RH = SHA-256(H(H1(D Rnd ) || H2(D Rnd ) || … || H m (D Rnd ))). In the formula, RH represents the hash value at the top of the blockchain structure, SHA-256 represents the SHA-256 algorithm that stores the hash value calculated for each block in the block header, H represents the hash function, and || represents the data block concatenation operation C2. Compare and analyze the root hash RH calculated in real time with the stored reference root hash, designated as RH0, to generate a comparison result RH'. When RH' = 1, that is, when RH ≠ RH0, the comparison result is inconsistent, and it is determined that data has been tampered with; when RH' = 0, that is, when RH = RH0, the comparison result is consistent, and it is determined that the data has not been tampered with.
5. The nuclear database management system for nuclear dynamic detection technology according to claim 4, characterized in that, The process of marking the problem core data is as follows: The anomaly detection results of the preprocessed nuclear data by combining DTW with LSTM autoencoder, 3σ criterion algorithm and MerkleTree hash chain are weighted and fused to calculate the anomaly score. The calculation formula of the anomaly score is And In the formula, C f represents the anomaly score of the preprocessed nuclear data D Rnd after the anomaly detection results, respectively represent the weight coefficients corresponding to the anomaly detection results of the preprocessed nuclear data by combining DTW with LSTM autoencoder, 3σ criterion algorithm and MerkleTree hash chain; Set the confidence level for anomaly detection, designated as C0, and conduct a comparative analysis based on the anomaly score and the confidence level C0. When C f > C0, it is determined that the data is abnormal, and it is marked as problematic data; Add a metadata tag as marked data to the abnormal data points where the problem data is detected, encapsulate the marked data in JSON format, and sort it from high to low confidence. Prioritize the transmission of problem core data with high confidence and calibrate it as D' Rnd 。 6. The nuclear database management system for nuclear dynamic detection technology according to claim 5, characterized in that, The steps of imputing missing data in the problem core data are as follows: Extract the valid data points {Z(s Rnd )|k = 1, 2, …, K} around the missing point from the labeled problem core data D' k , where s k is the position coordinate of the k-th in the nuclear facility; Calculate the semivariogram of the effective data points Z(s k ), which describes the change of data spatial correlation with distance. The calculation formula of the semivariogram is and k, k' ∈ {1, 2, …, K}. In the formula, γ(d) represents the semivariance value between the effective data points Z(s k ) and Z(s k′ ), d represents the Euclidean distance, log d Z(s k ) represents the number of data point pairs with a distance of d, Z(s k ) represents the value of the effective data point with the spatial coordinate s k , and Z(s k′ ) represents the value of the effective data point with the spatial coordinate s k′ ; Construct the Kriging equations, solve for the weight coefficients to minimize the interpolation error and satisfy the unbiasedness constraint. Then, the expression of the Kriging equations is and where λ k′ represents the weight coefficient corresponding to the semivariogram γ(d), δ represents the Lagrange multiplier, and s0 represents the spatial coordinates of the missing point; Calculate the estimated value of the missing point using weights and valid data points and perform data interpolation. The calculation formula for the estimated value of the missing point is In the formula, represents the estimated value of the missing point s0.
7. The nuclear database management system for nuclear dynamic detection technology according to claim 6, characterized in that, The steps of correcting incorrect data in the problem core data are as follows: According to the reference nuclear data D Ref Extract the normal data as the training set {(A Rnd,u , D Rnd,u ), and train the SVR model. Among them, A Rnd,u represents the characteristic data of input time, temperature and humidity, and D Rnd,u represents the data at the u-th time point of the preprocessed nuclear data D Rnd ; By optimizing the solution of the SVR objective function to obtain the support vectors and the regression hyperplane, the expression of the SVR objective function is where W represents the weight of the regression hyperplane, B represents the bias of the regression hyperplane, ζ and ξ both represent the slack variables of the SVR model, χ represents the penalty factor for balancing the complexity of the SVR model and the error tolerance, u and U represent the u-th time point and the total number of time points respectively; Define the constraint conditions to limit the range of variation of the predicted value. Then the expression of the constraint conditions is and where φ represents the kernel function, ∈ represents the width of the insensitive region, which controls the regression accuracy, represents the average value of u feature data A Rnd,u The || ||2 2 represents the square of the 2-norm, and T represents the transpose; Problem core data D' of input detection anomaly Rnd is input to the SVR model and the corrected value is output. The calculation formula for the corrected value is Y Rnd = W T ·φ(D′ Rnd ) + B, where Y Rnd represents the corrected value output by the SVR model.
8. The nuclear database management system for nuclear dynamic detection technology according to claim 7, characterized in that, The acquisition logic of the standardized core data is as follows: According to the reference nuclear data D Ref Calculate the global mean μ Ref and the standard deviation σ Ref ; The calculated missing data and the correction value Y Rnd are repaired to obtain the repaired nuclear data, calibrated as R Rnd , and aligned with the reference data D Ref according to the timestamp and spatial position; Perform Z-score normalization on the repaired data to eliminate the dimensional difference and ensure consistency with the reference data distribution. The expression for the normalized data is where Z Rnd represents the normalized core data after Z-score normalization, μ Ref represents the global mean of the reference core data D Ref and σ Ref represents the standard deviation of the reference core data D Ref and R Rnd represents the core data after repair; Standardize the nuclear data Z Rnd and perform a consistency check analysis with the reference nuclear data D Ref to determine whether to perform secondary repair.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the core database management system for nuclear dynamic detection technology according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the core database management system for nuclear dynamic detection technology according to any one of claims 1 to 8.