A method and system for optimizing a radioactive tank dismantling scheme based on digital twinning

By combining digital twin technology with a system-human risk coupling model, the problem of separating system risk and human risk in traditional methods has been solved, enabling comprehensive pre-control of complex risks during the dismantling of radioactive tanks and improving safety and reliability.

CN120996586BActive Publication Date: 2026-03-20SICHUAN ENVIRONMENTAL PROTECTION ENG CO LTD CNNC +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Traditional methods for dismantling large radioactive containers separate systemic risks from human-caused risks, ignoring the interaction and coupling evolution between the two. This results in incomplete risk prevention and control, difficulty in dealing with complex risk events, and potential safety hazards.

Method used

A digital twin-based approach is adopted to generate a digital twin model by acquiring multi-condition operating data and operator behavior data of a large radioactive tank. Using a system critical state assessment model, an error mode library, and a system-human risk coupling model, the interaction between system risk status and operational error types is analyzed to generate risk prevention and control intervention strategies, including system parameter adjustment and operational behavior guidance.

Benefits of technology

It enables quantitative analysis of the interaction between the system's critical state and human error, improving the accuracy of risk identification and the ability to prevent and control human-related risks, effectively preventing complex risk events, and enhancing the safety and reliability of dismantling operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996586B_ABST
    Figure CN120996586B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of digital twin technology and radioactive waste treatment, and discloses a method and system for optimizing a radioactive large tank dismantling scheme based on digital twin, wherein the method comprises the following steps: obtaining multi-working-condition operation data and operator behavior data to generate a digital twin model basic data set; using a system critical state evaluation model to generate system risk state quantitative indexes; identifying potential operation error types based on an error mode library; using a system-human risk coupling model to generate a risk coupling index; and generating a risk pre-control intervention strategy based on the risk coupling index. The present application solves the technical problem of traditional methods that separate system risks and human risks for processing, and achieves the collaborative prevention of compound risks.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of digital twin technology, radioactive waste treatment technology and risk control technology, more particularly, it relates to a radioactive large tank dismantling scheme optimization method based on digital twin. BACKGROUND

[0002] At present, the traditional risk prevention and control method is mainly used for radioactive large tank dismantling. These methods usually treat system risk assessment and human operation risk analysis as two independent modules. System risk assessment focuses on equipment state monitoring and parameter threshold control, and human risk analysis mainly focuses on operation procedure execution and personnel training management.

[0003] However, the traditional risk prevention and control method has obvious defects: the system risk and human risk are treated separately, and the interaction and coupled evolution law between the two are ignored. In the actual dismantling process, the system critical state will affect the psychological state and operation difficulty of the operator, and the human operation error will in turn exacerbate the system risk, and the two will interact and may form a risk amplification effect. This kind of risk treatment method leads to incomplete risk prevention and control, and it is difficult to deal with complex risk events, which brings safety hazards to the dismantling operation of radioactive large tank. SUMMARY

[0004] The present application provides a radioactive large tank dismantling scheme optimization method based on digital twin, which solves the technical problem that the system risk and human risk are treated separately and the interaction and coupled evolution law between the two are ignored in the related art.

[0005] The present application discloses a radioactive large tank dismantling scheme optimization method based on digital twin, comprising the following steps: obtaining the multi-condition operation data and the operator behavior data of the radioactive large tank, and generating the basic data set of the digital twin model; analyzing the multi-condition operation data by using a system critical state evaluation model to generate system risk state quantitative indicators; analyzing the operator behavior data based on an error mode library to identify potential operation error types; using a system-human risk coupling model to analyze the interaction between system risk state and operation error type to generate a risk coupling index; generating a risk prevention and control intervention strategy based on the risk coupling index and the system state prediction result; wherein the system risk state quantitative indicators are calculated by accumulating the weighted risk values of each system parameter, and the risk value of each parameter is converted into a normalized risk value by a risk conversion function.

[0006] The application discloses a system for implementing the above-mentioned digital-twin-based radioactive large tank dismantling scheme optimization method, comprising: a data acquisition module for acquiring multi-working-condition operation data and operator behavior data of a radioactive large tank; a system risk assessment module for generating a system risk state quantitative index; an operation error identification module for identifying potential operation error types; a risk coupling analysis module for generating a risk coupling index; and a pre-control strategy generation module for generating a risk pre-control intervention strategy.

[0007] Further, in the data acquisition step, the acquired data is preprocessed, and the preprocessing includes: abnormal value detection and processing, data normalization processing, operation behavior data coding, and time series alignment.

[0008] Further, the system critical state assessment model is an adaptive multi-layer parameter analysis model, and the structure of the model comprises a feature extraction layer, a parameter correlation layer and a risk quantification layer; the risk quantification layer quantifies the influence of each parameter on system stability into a risk index, and obtains the system risk state quantitative index through a weighted sum of the calculation results of the risk conversion functions of the parameters.

[0009] Further, the operation personnel behavior data is analyzed based on an error mode library, and a mode recognition algorithm based on semantic distance is used to calculate the similarity between the current operation behavior and the error modes in the library; when the similarity exceeds a preset threshold, the operation is determined as a potential error.

[0010] Further, the system-human factor risk coupling model is constructed based on an improved Bayesian network, and the risk coupling index is calculated by considering the product of the system risk index, the human risk index and the system-human factor coupling coefficient.

[0011] Further, the method further comprises: processing the operation behavior monitoring data by using a multi-sensor data fusion algorithm to identify abnormal operation behaviors in real time; the data fusion algorithm adopts a hierarchical adaptive fusion strategy, including data layer fusion, feature layer fusion and decision layer fusion.

[0012] Further, the method further comprises: processing the operation error data and the system parameters by using an error consequence propagation simulator to predict the system state changes caused by the error operation; the error consequence propagation simulator is composed of a state representation module, a disturbance propagation module and a state prediction module.

[0013] Further, the method further comprises: analyzing the historical data and the current state by using a risk coupling evolution prediction algorithm to identify potential risk amplification paths; the risk coupling evolution prediction algorithm adopts a spatiotemporal graph neural network model, and the future possible risk evolution paths are predicted by using a Monte Carlo simulation method.

[0014] Further, the risk pre-control intervention strategy includes system parameter adjustment suggestions and operation behavior guidance suggestions, which are generated by a multi-objective adaptive optimization algorithm based on weights, and the optimization objective function considers the weighted sum of the risk coupling index and the intervention cost.

[0015] The method for optimizing a radioactive large tank dismantling scheme based on digital twinning provided by the application realizes quantitative analysis of the interaction between the system critical state and human operation errors by establishing a system-human risk coupling model, solves the technical problem that traditional risk prevention and control methods separate system risks and human risks and ignore the interaction and coupling evolution law between the two, and achieves the following technical effects:

[0016] The system critical state evaluation model under multiple working conditions realizes accurate quantification of the risk of the radioactive large tank system, and improves the accuracy of risk identification;

[0017] The early identification of potential operation errors is realized through the construction of an error mode library and a mode recognition algorithm based on semantic distance, and the human risk pre-control capability is enhanced;

[0018] The influence of the system state on the operation difficulty and the influence of the operation error on the system state are revealed by introducing the system-human risk coupling model, the coupling evaluation of the risk is realized, and the problem that the traditional method ignores the risk interaction is overcome;

[0019] Through the risk pre-control intervention strategy generator, targeted system adjustment and operation guidance suggestions are provided to effectively prevent the occurrence of complex risk events and significantly improve the safety and reliability of the dismantling operation;

[0020] Through the combination of data-driven and model analysis, comprehensive pre-control of complex risks in the radioactive large tank dismantling process is realized, and safe and efficient technical support is provided for the decommissioning of nuclear facilities and the treatment of radioactive waste. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 is the overall flowchart of the method for optimizing a radioactive large tank dismantling scheme based on digital twinning of the application, which shows the main steps from data acquisition to risk pre-control intervention strategy generation; DETAILED DESCRIPTION

[0022] The method of the embodiment includes the following steps:

[0023] Step 100: Obtain the multi-working condition operation data and the operation personnel behavior data of the radioactive large tank, and generate the basic data set of the digital twinning model;

[0024] In step 100, the system parameter data of the radioactive tank in different working states, operating environments and dismantling stages are acquired, including but not limited to: internal pressure, temperature distribution, radiation level, structural stress and other system parameters; at the same time, the behavior data of the operator in the dismantling process are acquired, including operation trajectory, operation time, operation sequence and other information. The acquired data will be used for subsequent system critical state evaluation and human risk analysis.

[0025] Before using the acquired raw data, the following data preprocessing operations are required:

[0026] Outlier detection and processing: 3-sigma rule is used to detect outliers, and moving average is used to replace or remove abnormal data points;

[0027] Further, the implementation details of the 3-sigma rule are as follows: for any system parameter sequence {x_1, x_2, …,x_n}, where 、 、…、 respectively represent the 1st, 2nd, …, n-th system parameter value, first calculate the mean μ and standard deviation σ, then for any data point x_i, if |x_i - μ| > 3σ, it is considered as an outlier. The moving average window size used for replacement is 5, that is, (x_{i-2} + x_{i-1} + x_{i+1} + x_{i+2}) / 4 replaces the abnormal value x_i;

[0028] Data normalization: since each system parameter has different physical dimensions (such as pressure unit Pa, temperature unit ℃, radiation level unit Sv / h, etc.), all system parameter data are processed by Min-Max normalization;

[0029] Operation behavior data coding: non-numeric operation behavior data (such as operation type, operation object, etc.) are converted into numeric data vectors through one-hot encoding (One-hot encoding) for subsequent model processing;

[0030] Further, for the complex data coding of operation behavior, the implementation details are as follows: first, construct an operation type dictionary D={operation type 1, operation type 2, …, operation type m}, where operation type 1, operation type 2, …, operation type m represent the 1st, 2nd, …, m-th different operation behavior type, for any operation type i, encode it into an m-dimensional vector v, where the i-th element is 1 and the rest are 0. For multiple attribute operation behavior, different attribute one-hot encoding is used to splice joint feature vectors;

[0031] Time series alignment: time synchronization processing is performed on sensor data of different sampling frequencies, and a cubic spline interpolation method is used to align all data to a unified timestamp sequence.

[0032] Step 200: Analyze multi-condition operation data by using system critical state evaluation model to generate system risk state quantitative indicators;

[0033] In step 200, the multi-condition operation data obtained in step 100 is input into the system critical state evaluation model to calculate the system risk state quantitative indicators of the radioactive large tank under different conditions. The system critical state evaluation model is an adaptive multi-layer parameter analysis model, which consists of a feature extraction layer, a parameter correlation layer and a risk quantification layer.

[0034] The mathematical representation of the feature extraction layer is:

[0035] ;

[0036] Wherein, is the input multi-condition operation data, is the weight matrix of the feature extraction layer, is the bias vector, is the activation function (such as ReLU function).

[0037] Further, the network structure of the system critical state evaluation model is specifically configured as follows: the feature extraction layer adopts a three-layer fully connected neural network, the number of neurons in each layer is 128, 256 and 128 respectively, and the activation function is ReLU. In order to prevent overfitting, a Dropout layer is added between each layer, and the dropout rate is set to 0.2. The dimension of the input layer is determined according to the number of system parameters, which is usually 50-100 dimensions; the initialization method adopts Xavier method, which makes the variance of each layer input and output similar, which is helpful for stable signal transmission; batch processing is used during training, the batch size is 64, and the initial learning rate is 0.001;

[0038] The mathematical representation of the parameter correlation layer is:

[0039] ;

[0040] Wherein, is the weight matrix of the parameter correlation layer, is the bias vector.

[0041] Further, the parameter correlation layer adopts a two-layer fully connected neural network, with the number of neurons in each layer being 64 and 32, respectively, and the activation function being Leaky ReLU (with a negative slope of 0.01). The advantage of this activation function is that it avoids the "dead neuron" problem of ReLU in the negative value region. The parameter correlation layer also introduces an attention mechanism to calculate the correlation weight matrix between different system parameters , enabling the network to adaptively learn the interaction between parameters: , where and are the learnable query and key matrices, respectively, and

[0042] represents the transpose. This structure design enables the model to capture complex nonlinear dependencies between different system parameters;

[0043] ;

[0044] where is the number of system parameters, is the th system parameter, is the weight coefficient of the parameter (the weight coefficient satisfies and ), and is the risk conversion function corresponding to the parameter.

[0045] Further, the determination method of the parameter weight coefficient is through the analytic hierarchy process (AHP), i.e.: first, the importance of each parameter is compared by domain experts, and a judgment matrix A is constructed; then the largest eigenvalue of matrix A and the corresponding eigenvector v are calculated; finally, the weight coefficient is obtained by normalizing the eigenvector v. At the same time, consistency check is performed, requiring a consistency ratio CR < 0.1, where CR = CI / RI, , RI is the random consistency index; since each system parameter has different physical dimensions, the risk conversion function converts it into a dimensionless normalized risk value, defined as:

[0046] ;

[0047] where is the minimum safe value of the parameter, is the parameter safety threshold, is the parameter critical threshold. All ​​The dimensionless values of the value range of [0, 1] represent the risk degree from the safe state (0) to the critical state (1), ensuring that the parameters of different physical quantities are dimensionally consistent when added.

[0048] Further, the parameter safety threshold and the parameter critical threshold The determination method is: for each system parameter , based on historical operation data and national relevant safety standards (such as GB / T 1254-2017 “Regulations for Safe Transportation of Radioactive Substances”), the upper limit value of the normal working range of the parameter is set as , and the critical value of the parameter that begins to cause system instability is set as . Specifically, for the pressure parameter, is 80% of the design working pressure, is 95% of the design working pressure; for the temperature parameter, is 85% of the maximum allowable working temperature, is 97% of the maximum allowable working temperature; for the radiation level parameter, is 70% of the annual dose limit, is 90% of the annual dose limit.

[0049] The training of the system critical state evaluation model uses a supervised learning method, and the loss function consists of two parts: mean square error loss and safety constraint loss:

[0050] ;

[0051] The mean square error loss calculates the difference between the predicted risk index and the labeled risk index:

[0052] ;

[0053] wherein, is the number of training samples, is the true risk index of the th sample, is the risk index predicted by the model.

[0054] The safety constraint loss is used to ensure the conservatism of risk assessment and avoid underestimation of risk:

[0055] ;

[0056] wherein, is the allowable error threshold, which punishes risk underestimation and ensures the safety and reliability of the evaluation results. The model parameters are updated through back propagation and Adam optimizer, and the learning rate is dynamically adjusted using the cosine annealing strategy.

[0057] Further, the error tolerance threshold is set to 0.05, i.e. the risk index predicted by the model is allowed to be at most 0.05 (5%) higher than the real risk index, but not allowed to underestimate the risk more than this threshold. This value is determined by Monte Carlo simulation and expert evaluation, which guarantees safety margin while avoiding excessive conservatism leading to reduced operational efficiency;

[0058] It should be noted that the calculation of the system risk state quantitative index described above can also be based on the structural features and physical properties of the radioactive tank, and a state space representation can be used:

[0059] ;

[0060] wherein, is the system state vector at time t, is the control input vector, is the state transition matrix, is the control matrix, is the system noise. The system risk state quantitative index can be determined by measuring the distance of the state vector from the preset safety domain.

[0061] Step 300: Analyze the operator behavior data based on the error mode library to identify potential operation error types;

[0062] In step 300, the operator behavior data obtained in step 100 is matched and analyzed with the pre-established error mode library to identify potential operation error types. The error mode library contains typical operation error modes that can occur during the radioactive tank dismantling process, and each error mode includes error description, trigger condition, occurrence probability and severity, etc.

[0063] The error mode matching process uses a semantic distance-based pattern recognition algorithm to calculate the similarity between the current operation behavior and the error modes in the library :

[0064] The aforementioned semantic distance-based pattern recognition algorithm converts the operation behavior data and error mode data into feature vectors, and matches by calculating the distance in the semantic space. Specifically, the input of the algorithm is the operation behavior feature vector and each error mode feature vector in the error mode library, and the output is the similarity score and the matched possible error mode. The semantic distance-based pattern recognition algorithm uses a weighted similarity calculation method, which gives higher weight to key features to enhance the accuracy of matching.

[0065] ;

[0066] wherein,​ for the behavior feature dimension, for the current operation, feature value, for the error mode, feature value, for the feature weight, for the similarity calculation function.

[0067] Further, the determination method of the feature weight is based on information entropy theory, and the information gain of each feature is determined: wherein is the information gain of the feature , and the calculation formula is , is the entropy of the error type, is the conditional entropy of the error type under the condition of the known feature . For key features such as operation sequence and safety interlock state, the weight range is [0.15, 0.25]; for secondary features such as operation speed and operation force, the weight range is [0.05, 0.15]; in order to ensure the comparability between different types of features, the operation behavior data needs to be preprocessed before similarity calculation:

[0068] Numerical behavior features (such as operation duration, operation speed, etc.) are standardized by Z-score:

[0069] ;

[0070] wherein is the mean value of the feature, is the standard deviation;

[0071] Categorical features (such as operation type, operation sequence, etc.) are converted into vector representation by one-hot encoding;

[0072] Sequence features (such as operation trajectory) are calculated by dynamic time warping (DTW) algorithm to calculate the distance between sequences, and converted into similarity value by Gaussian kernel function:

[0073] ;

[0074] wherein denotes the exponential function,

[0075] Further, the standard deviation parameter in the Gaussian kernel function is determined by an adaptive method based on the distribution characteristics of the sequence distance in the training data set: wherein denotes the sequence distance value of the i-th training sample in the set, For the number of training samples, The median operation is represented, and this method can make the similarity function have better adaptability to different feature scales;

[0076] When the preset threshold is exceeded , it is determined as a potential error. The preset threshold is set to 0.75, which is determined by ROC curve analysis, and at this threshold, the sensitivity of error detection is 0.92, the specificity is 0.88, and the F1 score is 0.90, which can control the false positive rate within an acceptable range while ensuring the detection rate;

[0077] In the embodiments of the present application, in order to improve the real-time performance of operation error identification, step 300 can further include:

[0078] Step 310: processing the operation behavior monitoring data using a multi-sensor data fusion algorithm to identify abnormal operation behavior in real time;

[0079] In step 310, a plurality of sensors (including visual sensors, motion capture sensors, force feedback sensors, etc.) are deployed to collect real-time behavior data of the operator, and a data fusion algorithm is used to integrate the multi-source heterogeneous data to form a unified operation behavior representation, and compare it with the standard operation procedure to identify abnormal operation behavior in real time. The data fusion uses a hierarchical adaptive fusion algorithm, including data layer fusion, feature layer fusion and decision layer fusion:

[0080] The aforementioned hierarchical adaptive fusion algorithm uses time synchronization and data standardization methods at the data layer to convert heterogeneous data from different sensors to a unified time reference and numerical range. The feature layer fusion uses a deep autoencoder to extract effective features from the standardized data to generate a data representation in a unified feature space. The decision layer fusion combines the decision results of each sensor through an adaptive weighting mechanism to generate the final judgment.

[0081] The algorithm input is a heterogeneous data stream from multiple sensors, and the output is a unified operation behavior representation and an abnormal operation detection result. The algorithm uses an adaptive learning method to dynamically adjust the fusion weight according to the data quality of different sensors, improving the freshness and reliability of data fusion.

[0082] ;

[0083] wherein, is the fusion result, , , are the data layer, feature layer and decision layer fusion results, , , are layer weight coefficients, is a nonlinear mapping function.

[0084] Further, the layer weight coefficients , , are determined using an adaptive weight allocation strategy, which dynamically adjusts based on the quality indicators of each sensor data: wherein is the quality indicator of the i-th layer fusion, which is calculated by combining the signal-to-noise ratio, data integrity, and fusion accuracy. The initial weight is set to , , and the weight coefficient is continuously updated adaptively as the system runs and information accumulates; the nonlinear mapping function is implemented using a multi-layer perception structure and is defined as:

[0085] ;

[0086] wherein , is the weight matrix, , is the bias vector, and ReLU is the rectified linear unit activation function. The multi-layer perception structure nonlinear mapping function can capture the nonlinear relationship between different levels of fusion results, improving the expression ability of multi-sensor data fusion.

[0087] Step 400: Analyze the interaction between system risk state and operation error type using the system-human risk coupling model, and generate a risk coupling index;

[0088] In step 400, the system risk state quantitative indicators obtained in step 200 and the potential operation error types identified in step 300 are input into the system-human risk coupling model to analyze the interaction between the system state and human operation, and generate a risk coupling index. The system-human risk coupling model is based on an improved Bayesian network, which describes the influence of system state on operation difficulty and the influence of operation error on system state.

[0089] The improved Bayesian network is composed of a system state node set , an operation behavior node set , and a coupling relationship edge set , and its probability graph model is represented as wherein is the node set. Each system state node represents a system parameter, and each operation behavior node represents an operation behavior type. The directed edges between nodes denotes the conditional dependency.

[0090] The joint probability distribution of the system-human risk coupling model is:

[0091] ;

[0092] where, denotes the parent node set of node , is the conditional probability distribution.

[0093] The training of the system-human risk coupling model adopts the maximum likelihood estimation method, and the loss function is the negative log-likelihood:

[0094] ;

[0095] where, is the training data set, is the node value of the data sample . The optimization adopts the gradient descent method to update the model parameters.

[0096] The calculation formula of the risk coupling index is:

[0097] ;

[0098] where, is the time parameter, is the system risk index at time , is the human risk index at time , both of which are dimensionless normalized values with a value range of [0, 1], is the system-human coupling coefficient at time . The calculation of takes into account the influence factor of system state on operation difficulty and the influence factor of operation error on system state :

[0099] ;

[0100] Further, the time evolution characteristics of the risk coupling index are shown in the following aspects:

[0101] The time sequence calculation of the system risk index :

[0102] , where denotes the base number of the natural exponential function, is the system risk decay coefficient (with a value of 0.05), is the time step (0.5 hour), is the system risk value calculated based on real-time data at the current time;

[0103] Human risk indicator Considering the fatigue effect of operators: where is the basic human risk value, is the starting time of the current shift of the operator, is the standard shift length (8 hours), is the fatigue coefficient (0.3), and the formula shows that the human risk gradually increases with the increase of the continuous working time of the operator;

[0104] System-human coupling coefficient reflects the dynamic interaction between system state and human operation, and is calculated using a sliding time window method: where is the time window width (6), is the sampling time interval (10 minutes), is the instantaneous coupling coefficient at the time;

[0105] where, , , is the weight coefficient, satisfying and . Influence factors and are obtained by standardizing the system state parameters and operation error characteristics, both of which are dimensionless values and range between [0, 1], ensuring consistent dimensions when combined. The risk coupling index calculated finally is also a dimensionless normalized value, with a value range of [0, 1], representing the coupling risk degree from low risk (0) to high risk (1).

[0106] Further, the weight coefficients , and are determined by historical accident data analysis and expert experience, with initial values set to , , , indicating that the influence of operation errors on system state is slightly greater than the influence of system state on operation difficulty, and the weight of the interaction effect is relatively small but not negligible. The specific calculation method of influence factors and is: where is the influence function of system parameters on operation difficulty; ,in Number of operation error types Operational error Influence functions on system parameters. These influence functions are constructed based on piecewise linear mappings, converting inputs of different dimensions into dimensionless outputs within the interval [0,1].

[0107] In this embodiment of the application, in order to more accurately predict the impact of operational errors on the system state, step 400 may further include:

[0108] Step 410: Utilize the error consequence propagation simulator to process operational error data and system parameters, and predict system state changes caused by erroneous operations;

[0109] In step 410, an error consequence propagation simulator is used to simulate and analyze the impact path and extent of potential operational errors identified in step 300 on various system parameters, and to predict the trajectory of system state changes. The error consequence propagation simulator consists of a dynamic influence graph network, including a state representation module, a disturbance propagation module, and a state prediction module.

[0110] The state characterization module encodes system parameters into state vectors:

[0111] ;

[0112] in, for The system state vector at time t. For encoding functions, This represents the implicit state representation. The encoding function of the state representation module is implemented using a bidirectional long short-term memory (BiLSTM) network, defined as follows:

[0113] ;

[0114] in, This is the parameter set of the BiLSTM network. This encoding function can extract the temporal features and key attributes of the system state and generate a compact state representation.

[0115] Further, the specific structure configuration of the BiLSTM network is as follows: the input dimension is equal to the number of system state parameters (usually 20-30 dimensions), the hidden layer dimension is 128, the number of layers is 2, and the forward and backward LSTM outputs are merged by concatenation. The network also adopts a residual connection structure to connect the input directly to the output to alleviate the gradient disappearance problem; batch normalization is used to accelerate network convergence and improve stability; a temporal attention mechanism is used to highlight the features of key time points, and the attention weight calculation formula is: wherein , and are learnable parameters, denotes transposition. This structure design can effectively capture the long-term dependence and short-term fluctuation characteristics of the system state;

[0116] The disturbance propagation module calculates the propagation effect of the disturbance introduced by the operation error in the system:

[0117] ;

[0118] wherein, is the operation error vector, is the propagation function, is the disturbance effect vector. The propagation function is implemented by a gated graph convolution network (GGCN), which is defined as:

[0119] ;

[0120] wherein, denotes the concatenation of state representation and error vector, is the system parameter dependency adjacency matrix, is the graph convolution network parameter, and the gated graph convolution network (GGCN) propagation function can simulate the transmission effect of the disturbance in the system parameter network.

[0121] Further, the specific implementation details of the gated graph convolution network (GGCN) are as follows: the network contains 3 layers of graph convolution, and the feature dimensions of each layer are 64, 32, and 64 respectively; the mathematical form of each layer of graph convolution is wherein is the degree matrix, is the adjacency matrix, is the node feature of the th layer, is the weight matrix; the gating mechanism is realized by introducing update gate and reset gate , and the calculation formula is: , The output is ,in Candidate activation values, This represents element-wise multiplication. The gating mechanism of Gated Graph Convolutional Networks (GGCNs) enables the network to adaptively control the flow of information and selectively update node states, thereby more accurately simulating the complex interactions between system parameters and the effects of disturbance propagation.

[0122] The state prediction module predicts the future system state based on the current state and the disturbance effect:

[0123] ;

[0124] in, To control the input vector, For the prediction function, The predicted future state. The prediction function is implemented using a multi-layer feedforward neural network and is defined as:

[0125] ;

[0126] in, This represents the concatenation of state representation, disturbance effects, and control inputs. , This is the weight matrix. , As the bias vector, the multilayer feedforward neural network prediction function can comprehensively consider the current state, disturbance effects, and control input to accurately predict the future state of the system.

[0127] The error consequence propagation simulator is trained using a supervised learning method, with the loss function being the mean squared error between the predicted and actual states.

[0128] ;

[0129] in, This represents the number of training samples. Model parameters are updated via backpropagation and the Adam optimizer.

[0130] The aforementioned simulation of error consequence propagation is based on a system dynamics model and employs state transition equations:

[0131] ;

[0132] in, for The system state vector at time t. For normal control input, Disturbances introduced by operational errors The system dynamics equation is: Through numerical integration of the above equation, the evolution process of the system state after the operation error can be predicted.

[0133] Further, the definition and value range of the time parameter in the state transition equation are as follows:

[0134] The physical meaning of the time variable represents the cumulative time from the start of the demolition work, and the unit is hour, and the value range is , wherein is the total estimated duration of the entire demolition work (usually hundreds to thousands of hours);

[0135] The setting of the time step : The basic time step is set to 10 minutes (i.e. hour), which is used for general state evolution prediction; for critical state regions, an adaptive time step is used, wherein is a sensitivity coefficient (with a value of 10), and this strategy ensures that a smaller time step is used when the state changes rapidly to improve the prediction accuracy;

[0136] The integral time window : It represents the time range of state prediction, and the prediction window length is usually set to hour, which can be dynamically adjusted according to different prediction tasks;

[0137] Discretization implementation of the system dynamics equation : The fourth-order Runge-Kutta method (RK4) is used for numerical integration, that is:

[0138] ;

[0139] ;

[0140] ;

[0141] ;

[0142] ;

[0143] The truncation error of this integration method is , which can provide high-precision state prediction results;

[0144] Step 500: Based on the risk coupling index and the system state prediction result, a risk pre-control intervention strategy is generated;

[0145] In step 500, based on the risk coupling index calculated in step 400 and the system state prediction results, targeted risk prevention and intervention strategies are generated. These strategies include two aspects: system parameter adjustment suggestions and operational behavior guidance suggestions.

[0146] The pre-control intervention strategy is generated using a weighted multi-objective adaptive optimization algorithm, with the objective function being:

[0147] ;

[0148] in, As the intervention action vector, The time parameter indicates that at Intervention actions implemented at all times In order to be in Intervention actions should be taken at all times. The subsequent risk coupling index, The cost function for the intervention action (converted to a dimensionless value through normalization). and These are the weighting coefficients.

[0149] The aforementioned weighted multi-objective adaptive optimization algorithm solves the problem by transforming the two objectives of risk coupling index and intervention cost into a weighted single-objective problem. The input of the weighted multi-objective adaptive optimization algorithm is the current system state, the risk coupling index, and the set of possible intervention operations, and the output is the parallel vector of the optimal intervention strategy.

[0150] The algorithm employs an adaptive weight adjustment strategy, dynamically adjusting the weight ratios of the risk index and cost in the objective function based on the current risk level. When the risk level reaches a severe level, the weight of the risk index is increased, while the impact of cost factors is reduced; when the risk is within an acceptable range, minimizing intervention costs is considered. The algorithm uses an improved particle swarm optimization method to solve the problem, rapidly converging to the optimal solution that satisfies the constraints through a combination of global search and local refinement.

[0151] Furthermore, optimizing the time dimension of the objective function is reflected in the following aspects:

[0152] Dynamic risk coupling index: Over time Dynamic changes reflect the time evolution characteristics of the system state, and the calculation formula is: ,in express The change in risk at any given moment;

[0153] Time-series differentiated costs: intervention costs ,in This represents the base of the natural exponential function. The total duration of the demolition task. and The weights are 0.5 and 0.2 respectively. This function makes the weights decrease as the task nears its end. near Interventions implemented at lower costs encourage early preventative measures.

[0154] Time lag in intervention effects: The effects of intervention actions have a time lag, as shown by the effect function. Modeling, meaning in Intervention actions implemented at all times exist The effect of time, in which time delay Time unit;

[0155] Furthermore, weighting coefficients and Based on the current risk coupling index Horizontal adaptive adjustment: when (In the high-risk zone) , Prioritize risk reduction; when (In the medium-risk range) , Balancing risks and costs; when (In the low-risk range) , Risks and costs are considered equally. Intervention cost function. The weighted sum, including economic cost, time cost, and operational complexity, is converted into a dimensionless value in the [0,1] interval using Min-Max normalization. The constraints of the optimization problem include:

[0156] ;

[0157] Furthermore, the completeness of the constraints in the optimization problem is reflected in the following five aspects:

[0158] Risk and security constraints: Requires the entire prediction time domain Within this range, the risk coupling index does not exceed the safety threshold. ,in To predict the length of the time domain (usually set to 24 hours);

[0159] Action space constraints: The range of values ​​for each component in the intervention action vector was defined;

[0160] Motion smoothness constraints: To ensure that intervention actions do not change drastically between adjacent time points, This is set to the maximum permissible range of change (usually 0.2), which helps to avoid system oscillations and operational chaos.

[0161] Resource budget constraints: The total resource consumption of the intervention was limited, among which For the first The unit resource consumption coefficient of each intervention action This is the upper limit of the total resource budget;

[0162] Time constraints: The time limit for intervention actions was restricted. The range, of which and These are the minimum and maximum allowed execution times, respectively.

[0163] in, This represents an acceptable threshold for the risk coupling index. and These are the lower and upper bound vectors for the intervention action, respectively.

[0164] Furthermore, the acceptable threshold for the risk coupling index. Based on the safety level settings for the dismantling of radioactive storage tanks, for dismantling work at the standard safety level... For high-security-level demolition work, For demolition work at an extremely high safety level, Upper and lower bound vectors of intervention actions and Based on the system design parameters and operational safety specifications, the adjustment range of system parameters is limited to 70% to 90% of the design safety boundary; the scope of operational behavior guidance is determined based on standard operating procedures and the qualifications and capabilities of operators.

[0165] Furthermore, intervention action vectors It is A dimensional vector containing two parts: a subvector for adjusting system parameters. and operational behavior guidance subvector ,in This represents the total dimension of the intervention action vector. Each dimension of the system parameter adjustment sub-vector corresponds to an adjustable system parameter (such as pressure, temperature, ventilation rate, etc.), typically a normalized value within the range of [0, 1]. Each dimension of the operational behavior guidance sub-vector corresponds to the strength or priority of an operational guidance strategy, also represented by a normalized value within the range of [0, 1]. The timestamp attribute of the intervention action vector. The execution time of the action is defined, and its value range is: wherein is the current time, is the prediction time horizon (usually 8 hours in the future). The vector also contains the execution duration , which represents the effective action time of the intervention action, and the value range is hours;

[0166] The abstract intervention strategy output by the optimization algorithm needs to be decoded and converted into specific executable operation instructions. The decoding process includes:

[0167] System parameter adjustment decoding: map the parameter adjustment vector obtained by optimization to specific system control instructions, such as pressure regulation value, temperature control set point, radiation protection measures, etc., in the form of:

[0168] ;

[0169] wherein, represents the specific control instruction for adjusting the system parameter to the target value ;

[0170] Operation behavior guidance decoding: convert the operation behavior optimization vector into standard operation procedures (SOP) and warning prompts, in the form of:

[0171] ;

[0172] wherein, is the operation step description, is the warning information for specific risk points. The decoded intervention strategy is presented to the operator through the human-computer interaction system, guiding the safe execution of subsequent demolition work.

[0173] In the embodiments of the present application, in order to improve the forward-looking of the risk pre-control strategy, step 500 can further include:

[0174] Step 510: analyze historical data and current state using a risk coupling evolution prediction algorithm to identify potential risk amplification paths;

[0175] In step 510, based on the risk event cases in the historical data and the current system-human coupling state, a risk coupling evolution prediction algorithm is applied to identify the evolution path that may lead to risk amplification. The risk coupling evolution prediction algorithm uses a spatio-temporal graph neural network model to construct a risk state transition network.

[0176] The spatio-temporal graph neural network model is composed of a time embedding layer, a spatial graph convolution layer, and an evolution prediction layer.

[0177] Furthermore, the specific parameter configuration of the spatiotemporal graph neural network model is as follows: The network consists of three main components, and the parameter configuration of each component is optimized according to the characteristics of the input data and the model requirements. For the input data, the historical window size... The system is set to 10 time steps, analyzing the risk status within the most recent 10 time units; the batch size is 32; training uses the Adam optimizer with an initial learning rate of 0.0005, employing a learning rate scheduling strategy that decays to 80% of the original rate every 20 epochs; the dropout rate is 0.15 to prevent overfitting; the model uses an early stopping strategy, stopping training if performance does not improve for 5 consecutive epochs on the validation set; to address the imbalanced sample problem, oversampling is used for high-risk samples to increase their proportion in the training data, making the ratio of high, medium, and low-risk samples approximately 1:2:2.

[0178] The temporal embedding layer encodes historical risk state sequences into temporal features:

[0179] ;

[0180] in, for The risk status at all times, For the size of the history window, For time coding functions, This is the temporal feature vector. The temporal encoding function is implemented using a temporal attention mechanism and is defined as follows:

[0181] ;

[0182] in, For extracting time features, Attention weights are calculated as follows:

[0183] ;

[0184] in, Represents an exponential function.

[0185] in, Here is the attention parameter matrix. As an attention vector, the temporal attention mechanism's temporal encoding function can adaptively extract temporal features based on the importance of the state.

[0186] Furthermore, the specific parameters of the temporal embedding layer are set as follows: the input state dimension is... Time feature extraction matrix The dimension is This involves mapping the input state to a 64-dimensional latent space; attention parameter matrix The dimension is Attention vector The dimension is This layer also employs positional encoding to enhance the model's ability to perceive temporal order. Positional encoding is implemented using a sine-cosine function. , ,in For location index, For dimensional indexing, For model dimensions;

[0187] Spatial graph convolutional layers capture the spatial dependencies between system parameters and operational behavior:

[0188] ;

[0189] in, For system parameter set, A collection of events performed by human operators. It is a spatial convolution function. This represents the spatiotemporal feature vector. The spatial convolution function is implemented using a heterogeneous graph convolutional network (HGCN), defined as:

[0190] ;

[0191] in, The constructed heterogeneous graph consists of nodes including system parameter nodes and operation event nodes, with edges representing the interaction relationships between them. and These are the feature matrices for system parameter nodes and operation event nodes, respectively. As parameters for graph convolutional networks, the spatial convolution function of Heterogeneous Graph Convolutional Networks (HGCN) can effectively capture the complex interaction relationships between different types of nodes.

[0192] Furthermore, the specific implementation method of Heterogeneous Graph Convolutional Network (HGCN) is as follows: First, the heterogeneous graph contains two types of nodes (system parameter nodes and operation event nodes) and three types of edges (parameter-parameter edges, event-event edges, and parameter-event edges); different transformation matrices are defined for each type of edge to form message passing of specific relationships; the message passing formula is, where Indicates the relation type, It is a relationship-specific transformation matrix with dimension . (Input feature dimensions) Output feature dimension Then, for each node, aggregate messages from different relations: where Indicates the relationship Connect to node The neighbor set; finally, node feature updates employ a gating mechanism: The GRU is a gated recurrent unit used to adaptively fuse current features and aggregated neighbor messages; the HGCN network is stacked with 3 layers, and batch normalization and residual connections are added after each layer to improve training stability and alleviate the gradient vanishing problem.

[0193] The evolutionary prediction layer generates the risk state transition probabilities for the next time step:

[0194] ;

[0195] in, The prediction function is defined as follows: The evolutionary prediction layer prediction function is implemented using a variational autoencoder (VAE) structure.

[0196] ;

[0197] in, As latent variables, generated by the encoder: ,in and According to the encoder and The calculation yielded:

[0198] ;

[0199] Variational autoencoder (VAE) prediction functions can model the uncertainty of risk state transitions and generate probabilistic prediction results.

[0200] The model is trained using the conditional likelihood maximization method, with the loss function being the negative log-likelihood.

[0201] ;

[0202] in, The sequence length is given. Optimization uses the RMSprop algorithm, and gradient clipping is employed to prevent gradient explosion. After model training, the risk evolution path is predicted using Monte Carlo simulation.

[0203] ;

[0204] in, for The risk status at all times, For system parameter set, A collection of events performed by human operators. Let be the state transition function. The state transition function for risk evolution path prediction, based on a trained spatiotemporal graphical neural network model, is defined as:

[0205] ;

[0206] wherein, is the number of integrated models, is the weight coefficient of the th model, is the spatio-temporal feature vector of the th model. The state transition function of risk evolution path prediction improves the robustness and accuracy of prediction through model integration. Through the Monte Carlo simulation method, the possible future risk evolution paths are predicted, and the probability of each path is calculated to identify high-risk paths.

[0207] The application of the aforementioned Monte Carlo simulation method in the present embodiment adopts the importance sampling technique to improve the simulation efficiency. The input of the importance sampling Monte Carlo simulation algorithm is the conditional probability distribution generated by the evolution prediction model and the current risk state , and the output is a series of possible risk evolution paths and their occurrence probabilities.

[0208] The algorithm first generates multiple state transition samples based on the current state , and then recursively generates transition samples from each newly generated state to form a state transition tree. For each path, the cumulative probability and risk index are calculated. In order to improve the calculation efficiency, the algorithm adopts the importance sampling technique, which preferentially explores the paths with higher risk indexes, and normalizes the path probabilities to ensure the correctness of the calculation results.

[0209] In a decommissioning project of a nuclear power plant, a waste liquid storage tank needs to be removed. The tank is 6 meters in diameter, 8 meters in height, and has a volume of about 200 cubic meters. The inside of the tank is left with waste liquid containing radioactive substances such as Cs-137 and Sr-90. The removal work includes three main stages: waste liquid extraction, equipment disassembly, and tank cutting. The operating environment is complex, and there is significant interaction between system state and personnel operation errors.

[0210] Step 100: Obtain multi-condition operation data;

[0211] During the waste liquid extraction stage, the system parameters and operator behavior data are obtained through the deployed sensor network, and after data preprocessing, the basic data set is formed, as shown in Table 1:

[0212] Table 1: Real-time monitoring data of system parameters;

[0213] ;

[0214] The operator behavior data is shown in Table 1A:

[0215] Table 1A: Operator behavior data;

[0216] ;

[0217] Step 200 implements: System Critical State Assessment;

[0218] The multi-condition operation data is input into the system critical state assessment model, and the risk state quantitative indicators at each time point are calculated, as shown in Table 2:

[0219] Table 2 System risk state quantitative results;

[0220] ;

[0221] Step 300 implements: Operation Error Identification;

[0222] Based on the error mode library, the operation personnel behavior data is analyzed, and the potential operation error types are identified, as shown in Table 3:

[0223] Table 3 Operation error identification results;

[0224] ;

[0225] Step 400 implements: Risk Coupling Analysis;

[0226] The system-human risk coupling model is used to analyze the interaction between system risk state and operation error, as shown in Table 4:

[0227] Table 4 Risk coupling index calculation results;

[0228] ;

[0229] Step 410 implements: Error Consequence Propagation Simulation;

[0230] Using the error consequence propagation simulator, for the potential operation errors identified in step 300, the system state change trajectory is predicted, as shown in Table 4A:

[0231] Table 4A Prediction results of the influence of operation errors on system state;

[0232] ;

[0233] Step 500 implements: Risk Pre-control Intervention Strategy Generation;

[0234] Based on the risk coupling index and system state prediction results, a weight-based multi-objective adaptive optimization algorithm is used to generate intervention strategies, as shown in Table 5:

[0235] Table 5 Risk pre-control intervention strategy;

[0236] .

Claims

1. A method for optimizing the dismantling of large radioactive tanks based on digital twins, characterized in that, Includes the following steps: Acquire multi-condition operating data and operator behavior data of large radioactive tanks to generate the basic dataset for digital twin models; The system critical state assessment model is used to analyze multi-condition operating data and generate quantitative indicators of system risk status. Based on error pattern analysis, analyze operator behavior data to identify potential operational error types; The interaction between system risk status and operational error type is analyzed using a system-human factor risk coupling model to generate a risk coupling index. Based on the risk coupling index and system state prediction results, risk prevention and control intervention strategies are generated. Among them, the system risk status quantification index is calculated by accumulating the weighted risk values ​​of each system parameter, and the risk value of each parameter is converted into a normalized risk value through a risk transformation function; Risk Coupling Index The calculation formula is: ; in, For time parameters, for System risk indicators at any given time for Both are dimensionless normalized values, representing human-caused risk indicators at any given time, and their values ​​range from [0,1]. for The system-human coupling coefficient at any given time. The calculation considers the influence of system state on operational difficulty. Factors influencing the system state and operational errors : ; Furthermore, the temporal evolution characteristics of the risk coupling index are manifested in the following aspects: System risk indicators Timing calculation: ,in This represents the base of the natural exponential function. This is the system risk attenuation coefficient. For time step, This represents the system risk value calculated based on real-time data at the current moment. Human-made risk indicators Considering operator fatigue: ,in Based on the risk value of the basic person, This is the start time of the operator's current shift. For standard shift duration, The fatigue coefficient is used in this formula, which shows that as the continuous working time of operators increases, the risk of human error gradually rises. System-human coupling coefficient This reflects the dynamic interaction between system state and human operation, calculated using the sliding time window method: ,in The width of the time window. The sampling time interval, for The instantaneous coupling coefficient at time t; in, , , For the weighting coefficients, satisfying and .

2. The method according to claim 1, characterized in that, In the data acquisition step, the acquired data is preprocessed, including outlier detection and handling, data normalization, operation behavior data encoding, and time series alignment.

3. The method according to claim 1, characterized in that, The system critical state assessment model is an adaptive multi-level parameter analysis model, which consists of a feature extraction layer, a parameter correlation layer, and a risk quantification layer. The risk quantification layer quantifies the impact of each parameter on system stability into risk indicators, and obtains the system risk state quantification index by weighting the calculation results of the risk transformation function of each parameter.

4. The method according to claim 1, characterized in that, The analysis of operator behavior data based on an error pattern library employs a semantic distance-based pattern recognition algorithm. By calculating the similarity between the current operation and the error patterns in the library, a potential error is identified when the similarity exceeds a preset threshold.

5. The method according to claim 1, characterized in that, The system-human risk coupling model is constructed based on an improved Bayesian network. The risk coupling index is calculated by considering the product of three factors: system risk index, human risk index, and system-human coupling coefficient.

6. The method according to claim 1, characterized in that, The method also includes: using a multi-sensor data fusion algorithm to process operational behavior monitoring data and identify abnormal operational behaviors in real time; the data fusion algorithm adopts a hierarchical adaptive fusion strategy, including data layer fusion, feature layer fusion and decision layer fusion.

7. The method according to claim 1, characterized in that, The method also includes: using an error consequence propagation simulator to process operational error data and system parameters, and predicting system state changes caused by erroneous operations; the error consequence propagation simulator consists of a state characterization module, a disturbance propagation module, and a state prediction module.

8. The method according to claim 1, characterized in that, The method also includes: using a risk coupling evolution prediction algorithm to analyze historical data and current state to identify potential risk amplification paths; the risk coupling evolution prediction algorithm uses a spatiotemporal graph neural network model and uses Monte Carlo simulation to predict possible future risk evolution paths.

9. The method according to claim 1, characterized in that, Risk prevention and control intervention strategies include system parameter adjustment suggestions and operational behavior guidance suggestions, which are generated by a weighted multi-objective adaptive optimization algorithm. The optimization objective function takes into account the weighted sum of risk coupling index and intervention cost.

10. A system for executing the optimization method for the dismantling of radioactive large tanks based on digital twins as described in any one of claims 1-9, characterized in that, include: The data acquisition module is used to acquire multi-condition operating data and operator behavior data of the radioactive tank; The system risk assessment module is used to generate quantitative indicators of system risk status. Operation error identification module, used to identify potential operation error types; The risk coupling analysis module is used to generate risk coupling indices; The risk prevention and control strategy generation module is used to generate risk prevention and control intervention strategies.

Citation Information

Patent Citations

  • Construction method of tank field digital twinborn model and corresponding digital twinborn model

    CN116796808A

  • Hazardous chemical substance transportation risk prediction system and method based on big data analysis

    CN120634225A