Industrial safety intelligent agent whole-process autonomous management method for high-risk operation
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 西安圣瞳科技有限公司
- Filing Date
- 2026-05-19
- Publication Date
- 2026-08-07
AI Technical Summary
[0002]现有高风险工业作业安全管理工作,多依托工业现场布设的单一类型传感器采集数据,通过人工梳理作业人员、设备及环境相关信息,搭配常规风险评估模型完成安全状态判断,再依靠人工或基础管控程序下达安全干预指令,部分技术方案仅对传感数据做简单筛选处理,未构建标准化的状态数据处理流程
对初始状态数据集合开展时空对齐与缺失值补偿处理,能够规整多源异构传感数据的时空维度,填补数据采集过程中产生的信息空缺,生成的标准时空状态序列可维持数据的连贯性与完整性,对标准时空状态序列执行同步特征提取,能够完整抓取作业人员生理、设备运行、环境参数及作业流程的多维信息,多维度状态特征向量可完整保留各类传感数据的关联信息,数据处理环节的信息损耗被控制在更低范围,状态特征的呈现形式更贴合后续安全分析的使用需求。
Smart Images

Figure CN122222403B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of industrial safety management and control technology, specifically a method for full-process autonomous management of industrial safety intelligent agents for high-risk operations. Background Technology
[0002] Current safety management of high-risk industrial operations largely relies on data collected by single-type sensors deployed at industrial sites. Safety status is determined by manually sorting through information related to workers, equipment, and the environment, combined with conventional risk assessment models. Safety intervention instructions are then issued manually or through basic control procedures. Some technical solutions only perform simple screening of sensor data and have not established a standardized status data processing workflow.
[0003] Conventional industrial safety management technologies are prone to spatiotemporal mismatch and data gaps when processing multi-source heterogeneous mixed sensor data. They are difficult to generate regular spatiotemporal state sequences, and there are information omissions and dimensional deviations in the synchronous feature extraction process. Safety situation simulations rely on preset fixed rules and cannot be combined with professional knowledge to achieve dynamic simulations. The decision-making model is trained using traditional strategy optimization algorithms, and the output risk intervention instructions do not match the actual risk state on site.
[0004] The standardization of hybrid sensor data streams has loopholes that directly affect the accuracy of state feature extraction. Insufficient optimization of safety situation simulation methods and decision-making models can lead to risk intervention commands being unable to adapt to the full-process control needs of high-risk operations, making it difficult for industrial safety intelligent agents to achieve autonomous full-process safety management. Summary of the Invention
[0005] This invention aims to solve at least one of the technical problems existing in the prior art; To address this, the present invention proposes a full-process autonomous management method for industrial safety intelligent agents in high-risk operations, comprising: By deploying multi-source heterogeneous sensors in industrial sites, a mixed sensor data stream containing physiological signals of workers, operating status of work equipment, on-site environmental parameters, and work process execution steps is collected in real time to form an initial state data set. The initial state data set is subjected to spatiotemporal alignment and missing value compensation processing to generate a standard spatiotemporal state sequence; Synchronous feature extraction is performed on the standard spatiotemporal state sequence to obtain a multi-dimensional state feature vector; Based on a preset security knowledge graph, the multi-dimensional state feature vector is used to perform real-time security situation simulation and generate the current security risk status assessment result. The current security risk status assessment result is input into the decision model trained based on the improved strategy optimization algorithm, and the corresponding risk intervention instruction set is output. By executing the set of risk intervention instructions through the actuators in the industrial field, autonomous safety management of high-risk work processes can be achieved.
[0006] Furthermore, synchronous feature extraction is performed on the standard spatiotemporal state sequence to obtain a multi-dimensional state feature vector, including: For the components characterizing the physiological signals of workers in the standard spatiotemporal state sequence, a time-frequency joint analysis method is used to extract the frequency domain features of heart rate variability, the trend features of skin conductance response, and the features of body surface temperature gradient change, thus forming a sub-vector of personnel state features; For the components characterizing the operating state of the equipment in the standard spatiotemporal state sequence, a dynamic mode decomposition method is used to extract the energy distribution characteristics of the main vibration mode, the temperature change rate characteristics of key components, and the current harmonic distortion characteristics of the equipment, thus forming a sub-vector of equipment state characteristics. For the components characterizing the on-site environmental parameters in the standard spatiotemporal state sequence, a combination of spatial interpolation and field strength analysis is used to extract the spatial gradient characteristics of toxic gas concentration, the distribution and aggregation characteristics of inhalable particulate matter, and the main peak characteristics of the environmental noise spectrum, thus forming an environmental state feature sub-vector. For the components representing the execution steps of the work process in the standard spatiotemporal state sequence, the sequence pattern mining method is used to extract the time deviation features of the standard operation steps, the abnormal execution order features of key actions, and the tool usage interval features to form a process state feature sub-vector. The personnel state feature vector, the equipment state feature vector, the environment state feature vector, and the process state feature vector are concatenated and normalized to generate the multi-dimensional state feature vector.
[0007] Furthermore, the real-time security situation simulation based on the preset security knowledge graph and the multi-dimensional state feature vectors to generate the current security risk status assessment result includes: The multi-dimensional state feature vector is mapped to the entity attribute nodes of the preset security knowledge graph to activate the associated entities and relationship paths. The probability of risk transmission is calculated along the activated relational paths in the security knowledge graph, and the probability of risk transmission is calculated considering the coupling enhancement and inhibition effects between different risk factors. Based on the calculated risk transmission probability, and combined with the predefined risk event triggering conditions in the security knowledge graph, a risk event triggerability assessment is performed. By combining the calculated risk transmission probability with the risk event triggering assessment, the current security risk status assessment result, which includes risk level, location of major risk sources, and prediction of risk evolution trend, is derived.
[0008] Furthermore, by combining the calculation results of the risk transmission probability with the results of the risk event triggering assessment, the current security risk status assessment result, which includes risk level, location of major risk sources, and prediction of risk evolution trends, is derived, including: A risk level quantification mapping table is established, and the calculation results of the risk transmission probability of different paths and the risk event triggering assessment results of different risk events are integrated through a weighted aggregation function to obtain a comprehensive risk quantification value. Based on the distribution of the comprehensive risk quantification value within a preset threshold range, a discrete risk level is determined; Tracing back the original risk transmission paths and risk events whose contribution exceeds a preset threshold during the weighted aggregation process, the source entities in the corresponding security knowledge graph are marked as the main risk source locations; Using time series forecasting methods, the changes in the comprehensive risk quantification value and its constituent elements at the current moment over a future period are predicted, thus forming the risk evolution trend prediction.
[0009] Furthermore, the improved policy optimization algorithm employs an improved near-end policy optimization algorithm, the working principle of which includes: Based on the objective function of the standard near-end policy optimization algorithm, an adaptive trust domain constraint based on the current security risk state assessment result is introduced. The boundary size of the adaptive trust domain constraint is negatively correlated with the risk level in the current security risk status assessment result; that is, the higher the risk level, the stricter the constraint on policy updates. During each policy iteration update, the difference in action distribution between the new policy and the old policy is calculated, and the difference is compared with the boundary of the adaptive trust domain constraint. If the difference is less than the boundary of the adaptive trust domain constraint, then the policy parameters are updated according to the gradient direction of the standard near-end policy optimization algorithm. If the difference is greater than or equal to the boundary of the adaptive trust domain constraint, the gradient of the standard near-end policy optimization algorithm is pruned, and the policy parameters are updated along the pruned gradient direction to ensure that the policy update step size is always within the range allowed by the adaptive trust domain constraint.
[0010] Furthermore, the specific method for determining the boundary size of the adaptive trust domain constraint includes: Extract the quantitative value corresponding to the risk level from the current security risk status assessment results; The quantified value of the risk level is input into a preset boundary mapping function, which is a monotonically decreasing piecewise linear function. The boundary mapping function outputs a positive real number, which serves as the boundary threshold for the adaptive trust domain constraint in this policy iteration. As the current security risk status assessment results are updated over time, the boundary thresholds of the adaptive trust domain constraint are also dynamically adjusted.
[0011] Furthermore, the process by which the decision-making model outputs the corresponding set of risk intervention instructions includes: The decision model receives the current security risk status assessment result as input status; The value network within the decision-making model evaluates the long-term expected cumulative safety benefits of different alternative intervention action sequences under the input state. The policy network within the decision-making model generates a probability distribution of alternative intervention actions based on the evaluation results of the value network and combined with the immediate risk features extracted from the input state. Sample from the probability distribution of the candidate intervention actions, or select the intervention action with the highest probability to form a preliminary sequence of intervention actions; The initial sequence of intervention actions is logically consistent with the preset work safety procedure knowledge base, and intervention actions that violate hard safety rules are replaced or deleted. The verified and corrected sequence of intervention actions is converted into a standardized control command format that can be recognized and executed by the execution devices in the industrial field, forming the risk intervention instruction set.
[0012] Furthermore, the preliminary intervention sequence is logically consistent with a pre-set work safety procedure knowledge base, including: Analyze each intervention action in the preliminary intervention action sequence to identify its action type, target, and expected parameters; Retrieve all safety constraint rules related to the current high-risk work scenario from the preset work safety procedure knowledge base; The identified intervention actions, their target objects, and expected parameters are matched and logically deduced one by one with the retrieved safety constraint rules. If an intervention action directly conflicts with or implicitly contradicts any safety constraint rule, the corresponding intervention action is deemed to be logically inconsistent. For all intervention actions deemed logically inconsistent, within the permitted scope of the work safety procedure knowledge base, find functionally equivalent or lower-risk alternative actions to replace them. If no suitable alternative action is found, the corresponding intervention action is directly deleted.
[0013] Furthermore, the process of converting the verified and corrected intervention action sequence into a standardized control command format that can be recognized and executed by the actuators in the industrial field includes: Establish a mapping dictionary from intervention actions to equipment control commands, wherein the mapping dictionary defines one or more underlying equipment control command templates corresponding to each type of intervention action; Based on the specific parameters of each intervention action in the verified and corrected intervention action sequence, fill the variable placeholders in the corresponding equipment control instruction template to generate specific equipment control instructions. Based on the physical layout and communication protocol of the industrial field actuators, add target device address code and communication protocol header information to each generated specific device control command; Based on the logical sequence of action execution and parallel execution relationship in the intervention action sequence, the timing of all device control instructions with added address and protocol information is arranged to generate the final risk intervention instruction set containing complete control timing information.
[0014] Furthermore, after achieving autonomous safety management of high-risk work processes by executing the set of risk intervention instructions through actuators at the industrial site, the method further includes: During and after the execution of the risk intervention instruction set, feedback sensing data streams continue to be collected through the multi-source heterogeneous sensors; The feedback sensing data stream is processed using the same procedure as the initial state data set to obtain the feedback security risk status assessment result. The feedback security risk status assessment result is compared and analyzed with the current security risk status assessment result before the execution of the risk intervention instruction set, and the risk status change measure is calculated. By using the risk state change metric, the parameters of the decision model are updated online through incremental learning, thereby optimizing the decision model's ability to make decisions on similar future risk states.
[0015] Compared with the prior art, the beneficial effects of the present invention are: Spatiotemporal alignment and missing value compensation processing of the initial state data set can regularize the spatiotemporal dimensions of multi-source heterogeneous sensor data, fill the information gaps generated during data acquisition, and generate a standard spatiotemporal state sequence that can maintain the continuity and integrity of the data. Performing synchronous feature extraction on the standard spatiotemporal state sequence can fully capture multi-dimensional information on the physiological state of operators, equipment operation, environmental parameters, and work processes. The multi-dimensional state feature vector can fully retain the correlation information of various sensor data, and the information loss in the data processing stage is controlled to a lower range. The presentation form of the state features is more in line with the needs of subsequent safety analysis.
[0016] Based on a pre-set safety knowledge graph, real-time safety situation simulation is performed on multi-dimensional state feature vectors. Combined with professional safety knowledge, dynamic analysis of risk status is completed. The risk status assessment results can be aligned with the actual working conditions of high-risk operations in industrial sites. The assessment results are input into the decision model trained by the improvement strategy optimization algorithm, which can optimize the instruction output logic of the decision model. The risk intervention instruction set can be accurately matched with the current on-site safety risk status. When the execution device carries out control operations according to the instructions, it can be aligned with the control rhythm of the entire high-risk operation process. The autonomous control and execution process of the safety intelligent agent is more in line with the complex on-site operation scenarios. Attached Figure Description
[0017] Figure 1 This is a state diagram of the full-process autonomous management method for industrial safety intelligent agents for high-risk operations as described in this invention. Figure 2 A flowchart for obtaining multi-dimensional state feature vectors through synchronous feature extraction; Figure 3 This is a flowchart for real-time security situation simulation based on a security knowledge graph. Detailed Implementation
[0018] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] See Figure 1This invention provides a method for full-process autonomous management of industrial safety intelligent agents for high-risk operations. The overall implementation scheme is as follows: A multi-source heterogeneous sensor network is deployed in the industrial field. These sensors collect real-time mixed sensor data streams, including physiological signals of workers, operating status of equipment, environmental parameters, and execution steps of the work process. This data is aggregated to form an initial state data set. Spatiotemporal alignment and missing value compensation are performed on this initial state data set to address the synchronization and integrity issues between data from different sources and with different sampling rates, generating a standardized spatiotemporal state sequence aligned in both time and space dimensions. Synchronization features are extracted from this standardized spatiotemporal state sequence, mining information that characterizes the essence of the system state from data components of different dimensions, forming a multi-dimensional state feature vector. Based on a pre-defined safety knowledge graph, the multi-dimensional state feature vector is used as input for real-time safety situation simulation. This simulation process simulates the propagation and coupling of risk factors, thereby generating a current safety risk status assessment result containing information such as risk level and risk source. The current safety risk assessment result is input into a decision model, which is trained based on an improved strategy optimization algorithm and can output a set of risk intervention instructions for the current risk status. The execution devices in the industrial field receive and execute this set of risk intervention instructions, thereby completing closed-loop autonomous safety management of high-risk operation processes.
[0020] In one embodiment of the present invention, see [reference] Figure 2The process involves synchronous feature extraction of a standard spatiotemporal state sequence to obtain a multi-dimensional state feature vector. Specifically, the process is as follows: For the components representing the physiological signals of workers in the standard spatiotemporal state sequence, a time-frequency joint analysis method is used to extract the frequency domain features of heart rate variability, skin conductance response trends, and body surface temperature gradient changes. These features collectively constitute the personnel state feature sub-vector. For the components representing the operating state of the equipment in the standard spatiotemporal state sequence, a dynamic mode decomposition method is used to extract the energy distribution features of the equipment vibration master modes, the temperature change rate features of key components, and current harmonic distortion features. These features collectively constitute the equipment state feature sub-vector. For the components representing the environmental parameters in the standard spatiotemporal state sequence, a combination of spatial interpolation and field strength analysis is used to extract the spatial gradient features of toxic gas concentrations, the distribution and aggregation features of inhalable particulate matter, and the main peak features of the environmental noise spectrum. These features collectively constitute the environmental state feature sub-vector. For the components representing the execution steps of the work process in the standard spatiotemporal state sequence, a sequence pattern mining method is used to extract the time deviation features of standard operation steps, the abnormal execution sequence features of key actions, and the tool usage interval features. These features together constitute the process state feature sub-vector. Finally, the obtained personnel state feature sub-vector, equipment state feature sub-vector, environmental state feature sub-vector, and process state feature sub-vector are concatenated and normalized to generate a unified multi-dimensional state feature vector.
[0021] In practical implementation, the full-process autonomous management method for industrial safety intelligent agents in high-risk operations involves synchronous feature extraction of standard spatiotemporal state sequences to obtain multi-dimensional state feature vectors. In an example scenario of high-voltage substation equipment inspection, deployed multi-source heterogeneous sensors continuously generate standard spatiotemporal state sequences. These sequences are regularized data streams containing multiple dimensional components, processed through spatiotemporal alignment and missing value compensation. The synchronous feature extraction operation operates in parallel on the four logical components of the standard spatiotemporal state sequence, extracting information characterizing the essential state of the system from the raw sensor data.
[0022] In practice, the components representing the physiological signals of workers in the standard spatiotemporal state sequence are processed using a time-frequency joint analysis method. The monitoring equipment worn by the workers provides equally spaced sampling sequences of raw electrocardiogram (ECG), electrodermal signal, and body surface temperature signals. The time-frequency joint analysis first performs R-wave detection on the ECG signal and calculates the instantaneous heart rate sequence, then analyzes heart rate variability. Fourier transform or autoregressive models are applied to the heart rate variability signal to extract its power spectral density in the low-frequency range (e.g., 0.04-0.15Hz) and high-frequency range (e.g., 0.15-0.4Hz), and the low-frequency to high-frequency power ratio is calculated as the frequency domain feature of heart rate variability. The electrodermal signal is smoothed and subjected to first-order difference calculation, and its linear fitting slope within a certain time window is extracted as the trend feature of the electrodermal response. The body surface temperature signal is calculated for its spatial (e.g., between multiple monitoring points) or temporal differences to form a temperature gradient vector, and the statistics of this vector (e.g., mean and variance) are calculated as the feature of body surface temperature gradient change. Finally, the features extracted from different physiological signals are combined to form a sub-vector of human state features. In some embodiments, heart rate variability analysis can employ continuous wavelet transform to obtain better time-frequency locality, and its time-frequency joint analysis process can be described by the following formula: in: The scale parameter is represented as Translation parameters are The continuous wavelet transform coefficients at time, This represents the time-series data of the physiological signals to be analyzed. This represents the selected wavelet basis function. express The complex conjugate of the coefficient matrix, ∫[-∞,+∞]...dt, represents the integral operation with respect to time t from negative infinity to positive infinity, and dt represents the time derivative in the integral operation. By integrating or statistically analyzing the coefficient matrix over specific frequency bands and time intervals, the time-frequency energy distribution characteristics of the signal can be quantified.
[0023] In practical implementation, the components characterizing the operating state of the equipment in the standard spatiotemporal state sequence are processed using a dynamic mode decomposition method. Vibration acceleration sensors, infrared temperature sensors, and current transformers deployed on key equipment such as transformers and circuit breakers provide multi-channel vibration signals, temperature signals, and current signals during equipment operation. The dynamic mode decomposition method treats the multi-channel time series data as an observation of a dynamic system. By constructing a high-dimensional state snapshot matrix and performing singular value decomposition and low-order approximation on it, dynamic modes describing the dominant oscillation modes of the equipment's operating state are extracted. From these dynamic modes, the energy proportion of each mode can be calculated as the energy distribution characteristics of the equipment's main vibration mode. The temperature change rate of key components (such as transformer windings) can be extracted as the temperature change rate characteristics of key components. Spectral analysis of the current signal is performed to calculate the total harmonic distortion rate or the content of specific harmonics as the current harmonic distortion characteristics. These features together constitute the equipment state feature sub-vector. Optionally, the dynamic mode decomposition can employ variants such as exact dynamic mode decomposition or optimal dynamic mode decomposition to improve the stability of mode extraction.
[0024] In practical implementation, the components characterizing the on-site environmental parameters in the standard spatiotemporal state sequence are processed using a combination of spatial interpolation and field strength analysis. A network of multiple gas sensors, particulate matter sensors, and acoustic sensors distributed throughout the work area provides data on toxic gas concentrations, inhalable particulate matter (PM2.5 / PM10) concentrations, and environmental noise levels at discrete locations. The spatial interpolation method uses these discrete point measurements to generate a continuous concentration or sound pressure level distribution surface covering the entire work area. The spatial gradient of this distribution surface is calculated, and the maximum value of the gradient vector's modulus or the regional average value is extracted as the spatial gradient feature of the toxic gas concentration. Cluster analysis is performed on the particulate matter concentration distribution surface to identify the number, area, and central concentration value of high-concentration clusters, forming the distribution and clustering characteristics of inhalable particulate matter. A fast Fourier transform is performed on the environmental noise signal to identify the frequency and amplitude corresponding to the peak with the highest energy in the spectrum, which is used as the main peak feature of the environmental noise spectrum. This constitutes the environmental state feature sub-vector. In some embodiments, spatial interpolation can also employ an inverse distance weighting method, and field strength analysis can be combined with vector field visualization techniques to aid in understanding the spatial distribution pattern.
[0025] In practical implementation, the components representing the execution steps of the work process in the standard spatiotemporal state sequence are processed using sequence pattern mining methods. UWB positioning tags worn by operators, RFID or inertial measurement units on tools, and fixed cameras collectively generate discrete event sequences describing personnel movement, tool picking and putting down, and equipment operation actions. Sequence pattern mining first defines a reference event sequence template for the standard work process. By dynamically warping or calculating the edit distance between the real-time acquired event sequences and the reference template, the deviation of the start and end times of each operation from the standard time is quantified, and the time deviation features of the standard operation steps are extracted. A sub-sequence matching algorithm is used to detect whether prohibited reverse, skip, or redundant operations appear in the actual operation event sequence, and these abnormal patterns are identified, forming abnormal features of the execution sequence of key actions. By analyzing the time interval between two consecutive use events of the same tool and comparing it with the minimum cooling or resting time specified in the safety procedures, the tool use interval duration feature is extracted. These process compliance features constitute a process state feature sub-vector. Optionally, sequence pattern mining can use the PrefixSpan algorithm or the GSP algorithm to discover frequently occurring abnormal operation sequence patterns.
[0026] In practice, the personnel state feature vectors, equipment state feature vectors, environmental state feature vectors, and process state feature vectors are concatenated and normalized to generate a multi-dimensional state feature vector. The concatenation operation links the four sub-vectors end-to-end in each dimension, forming a high-dimensional comprehensive feature vector. For each dimension of the concatenated vector, the normalization fusion employs min-max normalization or Z-score standardization to scale feature values with different physical meanings and dimensions to similar numerical ranges. Min-max normalization linearly maps the original values to the [0,1] interval; Z-score standardization processes the data based on the mean and standard deviation of the features, ensuring the processed data conforms to a standard normal distribution. This normalization fusion ensures a balanced contribution of features from different sources and scales to the subsequent safety situation simulation model, generating a multi-dimensional state feature vector representing the current comprehensive and standardized state perceived by the industrial safety agent.
[0027] In one embodiment of the present invention, see [reference] Figure 3Based on a pre-defined security knowledge graph, real-time security situation simulation is performed on multi-dimensional state feature vectors to generate a current security risk status assessment result. Specifically, this process involves mapping each feature value in the multi-dimensional state feature vector to the corresponding entity attribute node in the pre-defined security knowledge graph. This mapping operation activates the associated entity nodes and the relationship paths between entities. Along the activated relationship paths in the security knowledge graph, the probability of risk transmission between different entities is calculated. The risk transmission probability calculation process considers the potential coupling enhancement or inhibition effects between different risk factors. Simultaneously, based on the calculated risk transmission probability and combined with the predefined triggering conditions for various risk events in the security knowledge graph, the likelihood of each potential risk event being triggered in the current state is assessed. The final current security risk status assessment result is derived by combining the calculated risk transmission probability with the risk event triggering assessment results. When generating this result, a risk level quantification mapping table is established. A weighted aggregation function integrates the risk transmission probability calculation results of different paths with the risk event triggering assessment results of different risk events to obtain a comprehensive risk quantification value. Based on the preset threshold range into which this comprehensive risk quantification value falls, a discrete risk level is determined. By tracing back the original risk transmission paths and risk events whose contribution exceeded a preset threshold during the weighted aggregation process, the source entities corresponding to these events in the security knowledge graph are marked as primary risk sources. Using time series forecasting methods, the changing trends of the current comprehensive risk quantification value and its constituent elements over a future period are predicted, forming a risk evolution trend prediction. The risk level, primary risk source locations, and risk evolution trend prediction together constitute the current security risk status assessment result.
[0028] In practical implementation, based on a pre-defined safety knowledge graph, real-time safety situation simulation is performed on multi-dimensional state feature vectors to generate current safety risk status assessment results. The implementation process uses the maintenance operation of a chemical reactor in a confined space as an example scenario. The multi-dimensional state feature vectors include quantitative features extracted from personnel physiological states, equipment operating parameters, ambient gas concentrations, and compliance of operational procedures. The simulation process begins by mapping the multi-dimensional state feature vectors to entity attribute nodes in the pre-defined safety knowledge graph. This pre-defined safety knowledge graph is a large graph-structured knowledge base that stores domain safety rules and causal relationships in the form of "entity-relationship-attribute" triples. The mapping operation, based on predefined matching rules, binds feature values to attribute slots of corresponding entity nodes in the knowledge graph. For example, a higher "skin conductance trend feature" value is mapped to the "stress level" attribute node of the "operator A" entity. This binding operation activates the "operator A" entity node and, based on predefined relationship paths such as "may lead to" and "associated with" in the graph, propagates activation to the "operational error" event node and the "reactor R1 inlet valve" equipment node.
[0029] In practical implementation, the probability of risk transmission is calculated along the activated relational paths in the safety knowledge graph. The relational edges in the safety knowledge graph are accompanied by conditional probability weights. The risk transmission probability calculation considers the coupling enhancement and inhibition effects between different risk factors. For example, the abnormal "combustible gas concentration spatial gradient characteristics" indicated by the "environmental state feature sub-vector" and the increased "static electricity accumulation probability" indicated by the "equipment state feature sub-vector" both point to the "deflagration risk" event node in the knowledge graph through an "AND" logical relationship. Their coupling effect makes the joint risk probability higher than the linear superposition of the risk probabilities of a single factor. The calculation process starts from the set of activated source entity nodes, uses a graph algorithm based on random walks to simulate the propagation of risk in the graph network, and iteratively updates the risk probability values of each node. The core iterative process of the risk transmission probability calculation can be described by the following formula: in: Indicates the process After the iteration, the risk probability distribution vector formed by all nodes in the security knowledge graph is... It is the normalized adjacency matrix of the security knowledge graph, whose elements Encoded from node To the node The intensity of risk transmission, It is a restart probability parameter between 0 and 1. It is the initial risk probability distribution vector obtained based on the mapping of the current multi-dimensional state feature vector. Representation matrix The transpose of the formula simulates risk in terms of probability. Propagation along the edge, with probability The process of returning to the initial state.
[0030] In practice, risk event triggerability assessment is conducted based on the calculated risk transmission probability and predefined triggering conditions for various risk events in the safety knowledge graph. The safety knowledge graph predefines risk event nodes such as "personnel poisoning," "fire," and "explosion," with each node associated with a logical triggering condition expression. The risk event triggerability assessment parses these logical expressions, comparing and performing logical operations on the calculated risk probability values and attribute values of the relevant entity nodes with the thresholds in the expressions. For example, the triggering condition for a "fire" event might be quantified as "('flammable gas concentration' > threshold)". AND('probability of potential ignition source risk' > threshold) The assessment process will compare the real-time "flammable gas concentration" attribute value with the threshold. Compare and match the risk probability value of the "potential ignition source" node with the threshold. The results are compared and then a logical AND operation is performed to output a trigger confidence level between 0 and 1, which serves as the result of the risk event trigger assessment.
[0031] In practical implementation, by combining the calculated risk transmission probability with the risk event triggering assessment results, a current safety risk status assessment result including risk level, location of major risk sources, and prediction of risk evolution trends is derived. The integration process establishes a risk level quantification mapping table, which defines the mapping interval from the comprehensive risk quantification value to discrete risk levels. The calculated risk transmission probability results for different paths are then vectorized. Risk event trigger assessment result vector with different risk events Through weighted aggregation functions Integrate and weighted aggregate functions By performing a weighted linear combination or nonlinear transformation on each input vector, a comprehensive risk quantification value is obtained. It's understandable that the weight vector... This can be learned through domain expert knowledge or historical data. Based on a comprehensive risk quantification value. Within the preset threshold range The distribution within the range determines a discrete risk level. During the backtracking weighted aggregation process, if the contribution exceeds a preset threshold... The original risk transmission paths and risk events are identified, and the source entities in the corresponding security knowledge graph are marked as the main risk sources. Using time series forecasting methods, a comprehensive risk quantification value sequence is generated for the current and a series of historical moments. Modeling is performed to predict values within a future time window. This leads to a prediction of risk evolution trends. In some embodiments, the time series forecasting method may employ exponential smoothing. Optionally, the risk evolution trend prediction may be modified by incorporating the temporal logic rules of event triggering in the knowledge graph. Discrete risk levels, the location of major risk sources, and the risk evolution trend prediction together constitute a structured assessment result of the current security risk status.
[0032] In one embodiment of the present invention, the improved policy optimization algorithm employs an improved near-end policy optimization algorithm. Its working principle is specifically implemented as follows: based on the objective function of the standard near-end policy optimization algorithm, an adaptive trust domain constraint based on the current security risk state assessment result is introduced. The boundary size of this adaptive trust domain constraint is not fixed, but is negatively correlated with the risk level in the current security risk state assessment result; the higher the risk level, the smaller the boundary of the constraint, and the stricter the restriction on policy updates. During each policy iteration update, a difference metric between the new and old policies in action distribution is calculated. This difference metric is compared with the boundary threshold of the adaptive trust domain constraint within the current iteration period. If the difference is less than the boundary threshold, the parameters of the policy network are updated according to the gradient direction of the standard near-end policy optimization algorithm. If the difference is greater than or equal to the boundary threshold, the calculated policy gradient is pruned, and the policy parameters are updated along the pruned gradient direction, thereby ensuring that the policy update step size is always within the range allowed by the adaptive trust domain constraint. Specifically, the boundary threshold of the adaptive trust domain constraint is determined by extracting the quantized value corresponding to the risk level from the current security risk state assessment result. The quantified value is input into a preset boundary mapping function, which is a monotonically decreasing piecewise linear function. The boundary mapping function outputs a positive real number, which is set as the boundary threshold of the adaptive trust domain constraint in this policy iteration. As the system runs, the current security risk status assessment result is dynamically updated, and the boundary threshold of the adaptive trust domain constraint is also dynamically adjusted accordingly.
[0033] In practical implementation, the improved policy optimization algorithm employs an improved proximal policy optimization algorithm. In a safety decision-making scenario for high-altitude welding operations in large tank farms, the improved proximal policy optimization algorithm is used to train the decision-making model of the industrial safety agent. The working principle of the improved proximal policy optimization algorithm involves introducing an adaptive trust domain constraint based on the current safety risk state assessment result, on top of the objective function of the standard proximal policy optimization algorithm. The objective function of the standard proximal policy optimization algorithm aims to maximize the expected cumulative reward, while limiting the difference between the old and new policies through pruning or penalty terms to ensure training stability. The adaptive trust domain constraint further dynamizes this constraint; its constraint boundary is not a fixed value but is negatively correlated with the risk level in the current safety risk state assessment result. The higher the risk level, the stricter the constraint on policy updates, i.e., the smaller the allowed policy update step size. During each policy iteration update, it is necessary to calculate the difference in action distribution between the new and old policies. This difference is typically measured using metrics such as KL divergence or total variational distance. In practical implementation, the network parameters of the new policy are calculated. Compared with the old strategy network parameters The KL divergence of the output action probability distribution under the same input state. As a measure of difference .
[0034] In practical implementation, the difference measurement value Boundary threshold of adaptive trust region constraint within the current iteration period Compare them. If the difference measure... Less than the boundary threshold of the adaptive trust region constraint Then, the policy parameters are updated according to the gradient direction of the standard proximal policy optimization algorithm. The gradient of the standard proximal policy optimization algorithm typically contains the product of the advantage function estimate and the policy probability ratio term, and the update magnitude is ensured through a pruning operation. If the difference metric... The boundary threshold is greater than or equal to the adaptive trust region constraint. Then, the calculated policy gradient needs to be pruned, and the policy parameters updated along the pruned gradient direction. The pruning operation ensures that the policy update step size is always within the range allowed by the adaptive trust domain constraint, thereby forcing a more conservative and smaller-amplitude policy update under high-risk conditions, avoiding unpredictable security intervention commands due to drastic policy changes. In some embodiments, the pruning operation can be implemented by scaling the gradient vector so that the difference between the old and new policies after parameter update is just close to but does not exceed the boundary threshold. Understandably, this mechanism allows the decision-making model to explore and learn more quickly under low-risk conditions, while under high-risk conditions, it focuses more on the stability and reliability of the strategy.
[0035] In practical implementation, the boundary threshold of the adaptive trust domain constraint The specific determination method involves extracting the quantified values corresponding to the risk levels from the current security risk status assessment results. Risk levels are typically predefined as discrete levels, such as "low risk = 1", "medium risk = 2", "high risk = 3", and "emergency risk = 4". These quantified values serve as inputs to the boundary mapping function. The quantified values of the risk levels are input into the preset boundary mapping function, which is a monotonically decreasing piecewise linear function, meaning that the function's output value decreases as the risk level increases. The boundary mapping function outputs a positive real number, which serves as the boundary threshold for the adaptive trust domain constraint in this policy iteration. As the current security risk assessment results are dynamically updated over time, the risk level may change, affecting the boundary thresholds of the adaptive trust domain constraint. It also adjusts dynamically accordingly. An example of a boundary mapping function is as follows: in: This indicates that when the risk level is The adaptive trust domain boundary threshold is calculated in time. It is the basic boundary threshold. It is the attenuation coefficient. It is a preset minimum boundary threshold to ensure basic training stability. The function ensures that the output value is not less than All symbols in the formula are positive real numbers. In some embodiments, the boundary mapping function can also be implemented using a lookup table, directly obtained from a predefined table of correspondence between risk levels and boundary thresholds. Optionally, a simplified mapping table is shown in Table 1: Table 1: Mapping Table between Risk Level and Adaptive Trust Domain Boundary Threshold In practical implementation, after determining the boundary threshold of the adaptive trust region constraint according to the above method, the policy update rule of the improved near-end policy optimization algorithm can be formally expressed. It can be understood that the improved objective function is an improvement over the standard near-end policy optimization algorithm objective function. Building upon this, a hard constraint on policy update differences has been added. Policy parameters The update needs to meet the following conditions Under the given conditions, optimize the expected cumulative safety gain. If the new strategy calculated in the gradient update step causes the KL divergence to exceed the boundary threshold... Then, the strategy parameters are adjusted to the nearest point that satisfies the constraints through methods such as projection or backtracking search. Through this adaptive trust domain mechanism that is tightly coupled with the real-time security risk status assessment results, the decision model can adaptively balance the benefits of exploring new strategies with the need to maintain strategy stability during training and execution, based on the urgency of the on-site risks.
[0036] In one embodiment of the present invention, the process of the decision model outputting the corresponding risk intervention instruction set is specifically implemented as follows: The decision model receives the current safety risk status assessment result as the input state. The value network inside the decision model evaluates the long-term expected cumulative safety benefits that different alternative intervention action sequences can bring under this input state. The strategy network inside the decision model generates a probability distribution of all alternative intervention actions based on the evaluation result of the value network and combined with the immediate risk features extracted from the input state. Sampling is performed from this probability distribution, or the intervention action with the highest probability is selected to form a preliminary intervention action sequence. The preliminary intervention action sequence is logically consistent with a preset work safety procedure knowledge base. This verification process parses each intervention action in the preliminary intervention action sequence, identifies its action type, target, and expected parameters. All safety constraint rules related to the current high-risk work scenario are retrieved from the work safety procedure knowledge base. The identified intervention action parameters are matched and logically reasoned with the retrieved safety constraint rules one by one. If an intervention action directly conflicts with or implicitly contradicts any safety constraint rule, the intervention action is determined to be logically inconsistent. For all intervention actions deemed logically inconsistent, within the limits of the work safety procedure knowledge base, equivalent or lower-risk alternative actions are sought for replacement. If no suitable alternative is found, the intervention action is deleted. The verified and corrected sequence of intervention actions is converted into a standardized control command format recognizable and executable by industrial field actuators. This conversion process is achieved through a predefined mapping dictionary of intervention actions to equipment control instructions, which defines the underlying equipment control instruction template corresponding to each type of intervention action. Based on the specific parameters of each action in the verified sequence of intervention actions, variable placeholders in the corresponding equipment control instruction template are filled in to generate specific equipment control instructions. According to the physical layout and communication protocol of the industrial field actuators, target device address codes and communication protocol header information are added to each generated equipment control instruction. Finally, all instructions are time-sequentially arranged according to the logical order of action execution and possible parallel execution relationships in the sequence of intervention actions to generate the final set of risk intervention instructions.
[0037] In practical implementation, the process of the decision model outputting the corresponding set of risk intervention instructions includes the decision model receiving the current safety risk status assessment result as the input state. In the scenario of high-altitude installation of large steel structures, the current safety risk status assessment result may include information such as "high risk level," "the main risk sources are located in abnormal braking torque of the hoisting equipment and excessive wind speed at high altitudes," and "the risk evolution trend is predicted to continue to rise." The value network inside the decision model evaluates the long-term expected cumulative safety benefits of different alternative intervention action sequences under the input state. The value network is a trained deep neural network whose input is the state description and the embedded representation of the alternative action sequence, and its output is a scalar value used to estimate the discounted total safety benefits that can be obtained in the future after executing a specific intervention action sequence under the input state. The policy network inside the decision model generates a probability distribution of alternative intervention actions based on the evaluation results of the value network and combined with the immediate risk features extracted from the input state. The policy network is also a deep neural network, and its output layer usually uses the Softmax function to assign a selection probability to each possible discrete intervention action, or output a parameterized probability distribution for continuous actions. Sampling is performed from the probability distribution of the alternative intervention actions, or the intervention action with the highest probability is selected to form a preliminary sequence of intervention actions. For example, a preliminary sequence of actions such as "suspend hoisting operation", "activate area audible and visual alarm", and "issue personnel evacuation order" may be generated.
[0038] In practice, the initial sequence of intervention actions is logically consistent with a pre-defined operational safety procedure knowledge base. This knowledge base, stored as a rule engine or ontology, contains mandatory safety regulations, operational prohibitions, and best practices for the work scenario. The logical consistency check parses each intervention action in the initial sequence, identifying its action type, target, and expected parameters. For example, the action "reduce hoisting speed" is classified as "equipment control," the target is the "main crane," and the expected parameter is the "target speed value." The pre-defined operational safety procedure knowledge base retrieves all safety constraint rules related to the current high-risk work scenario, filtering based on contextual information such as work type, equipment model, and environmental conditions. The identified intervention action's action type, target, and expected parameters are then matched and logically deduced against the retrieved safety constraint rules one by one. This matching and logical deduction process substitutes the action parameters into the rule's conditional part for calculation. If an intervention action directly conflicts with or implicitly contradicts any safety constraint rule, the corresponding intervention action is deemed logically inconsistent. For example, a safety rule might stipulate that "any hoisting operation is prohibited when the wind speed exceeds level X," while the preliminary sequence includes the action of "reducing hoisting speed," creating a direct conflict. For all intervention actions deemed logically inconsistent, within the limits of the work safety procedure knowledge base, equivalent or lower-risk alternative actions are searched for and replaced. For example, "reducing hoisting speed" is replaced with "stop hoisting and lock the boom." If no suitable alternative action is found, the corresponding intervention action is deleted.
[0039] In practice, the verified and corrected sequence of intervention actions is converted into a standardized control command format that can be recognized and executed by the actuators in the industrial field. The conversion process establishes a mapping dictionary from intervention actions to equipment control commands. This mapping dictionary defines one or more underlying equipment control command templates corresponding to each type of intervention action. The command template is a command string or binary frame structure containing variables, as shown in Table 2. Table 2: Mapping Table of Intervention Actions to Equipment Control Commands Based on the specific parameters of each intervention action in the verified and corrected intervention action sequence, the variable placeholders in the corresponding equipment control instruction template are filled in to generate specific equipment control instructions. For example, for the "Adjust Lifting Speed" action with the target "Main Crane CRANE_01" and the parameter "Target Speed" being 0, the instruction "CMD:SET;DEV=CRANE;SPEED=0" is generated. In some embodiments, long-term expected cumulative safety benefits The value network assessment can be formally expressed by the following formula: in: Indicates the state Execution sequence The long-term expected cumulative safety benefits; Represents the mathematical expectation; This represents the summation operation, starting from the current decision time. From the beginning to infinity in the future, all moments Accumulate; The discount factor is a constant between 0 and 1. Indicates any point in the future; Indicates the current decision-making moment; Indicates at time In state Next action The immediate security benefits gained; Indicates the decision model at time 10:00 The received current state; : Indicates a future moment The state; Indicates a future moment A single intervention action performed; The part after the vertical line indicates the condition, that is, "at the current moment". The state is And the sequence of actions taken is Under this condition, calculate the subsequent expected cumulative return; Indicates exponentiation, discount factor The index represents future moments. With the current moment The goal of the value network is to accurately estimate this expected value, which is the difference between the two.
[0040] In practical implementation, based on the physical layout and communication protocols of the industrial field actuators, target device address codes and communication protocol header information are added to each generated specific device control command. For example, the IP address and Modbus TCP protocol transaction identifier are added to commands sent to alarms in a specific area. All device control commands with added address and protocol information are time-series orchestrated according to the logical order of action execution and possible parallel execution relationships in the intervention action sequence. The time-series orchestration determines the sending time, interval, and synchronous or asynchronous execution relationship of the commands. Optionally, commands without strict time dependencies can be marked as parallel groups. A final set of risk intervention commands containing complete control timing information is generated. This risk intervention command set is typically a structured list or script containing command frames executed sequentially or in parallel and their timestamps. In some embodiments, time-series orchestration can manage the dependencies between actions using a directed acyclic graph. It can be understood that this transformation and orchestration from high-level intervention actions to low-level standardized control commands ensures that the abstract safety strategy output by the decision model can be reliably and accurately understood and executed by the diverse actuators in the industrial field.
[0041] In one embodiment of the present invention, after implementing a set of risk intervention commands through an execution device at the industrial site to achieve autonomous safety management of high-risk work processes, the method further includes the following steps: During and for a period of time after the execution of the risk intervention command set, feedback sensor data streams are continuously collected through multi-source heterogeneous sensors deployed on-site. The feedback sensor data stream undergoes the same process as the initial state data set, including spatiotemporal alignment, feature extraction, and safety situation deduction, to obtain a feedback safety risk status assessment result. This feedback safety risk status assessment result is compared and analyzed with the current safety risk status assessment result generated before triggering the execution of this risk intervention command set, and a risk status change metric is calculated, which reflects the effectiveness of the intervention measures. Using the calculated risk status change metric, the parameters of the decision model are updated through online incremental learning, thereby optimizing the decision model's ability to make decisions about similar risk states that may occur in the future.
[0042] In practical implementation, after executing risk intervention command sets through on-site actuators to achieve autonomous safety management of high-risk work processes, the full-process autonomous management method of industrial safety intelligent agents for high-risk operations also includes subsequent feedback and optimization steps. In work scenarios involving hydrostatic testing of large pressure vessels, the risk intervention command set may include control commands such as "shut down the booster pump," "open safety valve A," and "issue audible and visual alarms," which the actuators execute. During the execution of the risk intervention command set and for a preset monitoring period after execution, feedback sensor data streams continue to be collected through multi-source heterogeneous sensors deployed on-site. The feedback sensor data streams have the same source and type as the initial state data set and continuously include real-time information such as pressure, flow rate, valve position status, personnel location, and ambient sound.
[0043] In practice, the feedback sensor data stream undergoes the same processing flow as the initial state data set to obtain the feedback security risk status assessment result. The feedback sensor data stream first undergoes spatiotemporal alignment and missing value compensation processing to generate a standard spatiotemporal state sequence aligned with historical data. Next, synchronous feature extraction is performed on this sequence to obtain a new multi-dimensional state feature vector. Then, based on the same pre-defined security knowledge graph, real-time security situation simulation is performed on this new multi-dimensional state feature vector. The simulation process is consistent with the process used to generate the current security risk status assessment result, ultimately outputting a feedback security risk status assessment result reflecting the system's security status after intervention. It can be understood that the feedback security risk status assessment result has the same structure as the previously generated current security risk status assessment result, including elements such as risk level, location of major risk sources, and prediction of risk evolution trends.
[0044] In practice, the feedback security risk status assessment results are compared and analyzed with the current security risk status assessment results before the execution of this set of risk intervention instructions, and a risk status change metric is calculated. The comparative analysis focuses on the differences in quantitative indicators between the two assessment results. The risk status change metric is used to objectively evaluate the execution effect of the risk intervention instruction set. In one specific calculation method, the risk status change metric... It can be defined using the following formula: in: and These represent the quantitative values of the risk level in the pre-intervention safety risk assessment and the post-intervention safety risk assessment, respectively. and These represent the numbers before and after the intervention, respectively. Risk probability values of the activated main risk source nodes It is a symbolic function. and These represent the overall risk quantification values before and after the intervention, respectively. These are the weighting coefficients for each item. The first term of the formula measures the change in risk level, the second term measures the overall change in the probability of each risk source, and the third term uses a sign function to determine the upward or downward trend of the comprehensive risk value. Risk Status Change Measurement It is a scalar value; the larger the value, the more significant the improvement in risk status.
[0045] In practice, risk state change metrics are used to perform online incremental learning and updates to the parameters of the decision-making model. Risk State Change Metrics It is used as an immediate reward signal or a basis for strategy evaluation. The online incremental learning update process of the decision model stores the decision experience (including the state before intervention, the set of risk intervention instructions output by the decision model, the state after intervention, and the calculated risk state change measure) in the experience replay buffer. Periodically, or when the experience replay buffer reaches a certain capacity, a batch of experience data is sampled from the buffer to update the parameters of the policy network and value network inside the decision model. In some embodiments, the update process can employ a gradient-based temporal difference learning method to measure the risk state change. By combining the predicted value of a state with the value network of the decision model, the policy gradient is calculated, thereby fine-tuning the model parameters in the direction of improving the measure of future expected risk state changes. Optionally, for decision models trained based on an improved proximal policy optimization algorithm, the risk state change measure can be directly used to calculate the advantage function estimate and then participate in the policy gradient update. Online incremental learning updates and optimizes the decision model's ability to make decisions on similar risk states that will appear in the future, enabling the decision model to continuously learn from the actual intervention effects and constantly adjust its policy to generate a better set of risk intervention instructions. In some embodiments, the learning rate can be adaptively adjusted according to the magnitude of the risk state change measure. When the risk state change measure indicates a significant intervention effect (positive or negative), a larger learning rate is used to accelerate learning; when the change is not significant, a smaller learning rate is used to maintain policy stability. Optionally, the experience replay buffer can adopt a priority replay mechanism, assigning sampling priority according to the absolute value of the risk state change measure, so that experiences with significant intervention effects (whether good or bad) are used more frequently for learning.
[0046] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.
Claims
1. A method for full-process autonomous management of industrial safety intelligent agents for high-risk operations, characterized in that: The method includes: By deploying multi-source heterogeneous sensors in industrial sites, a mixed sensor data stream containing physiological signals of workers, operating status of work equipment, on-site environmental parameters, and work process execution steps is collected in real time to form an initial state data set. The initial state data set is subjected to spatiotemporal alignment and missing value compensation processing to generate a standard spatiotemporal state sequence; Synchronous feature extraction is performed on the standard spatiotemporal state sequence to obtain a multi-dimensional state feature vector; Based on a preset security knowledge graph, the multi-dimensional state feature vector is used to perform real-time security situation simulation and generate the current security risk status assessment result. The current security risk status assessment result is input into the decision model trained based on the improved strategy optimization algorithm, and the corresponding risk intervention instruction set is output. By executing the set of risk intervention instructions through the actuators in the industrial field, autonomous safety management of high-risk work processes can be achieved; The process by which the decision model outputs the corresponding set of risk intervention instructions includes: The decision model receives the current security risk status assessment result as input status; The value network within the decision-making model evaluates the long-term expected cumulative safety benefits of different alternative intervention action sequences under the input state; The policy network within the decision-making model generates a probability distribution of alternative intervention actions based on the evaluation results of the value network and combined with the immediate risk features extracted from the input state. Sample from the probability distribution of the candidate intervention actions, or select the intervention action with the highest probability to form a preliminary sequence of intervention actions; The initial sequence of intervention actions is logically consistent with the preset work safety procedure knowledge base, and intervention actions that violate hard safety rules are replaced or deleted. The verified and corrected sequence of intervention actions is converted into a standardized control command format that can be recognized and executed by the actuators in the industrial field, forming the risk intervention instruction set; The improved policy optimization algorithm employs an improved near-end policy optimization algorithm, and its working principle includes: Based on the objective function of the standard near-end policy optimization algorithm, an adaptive trust domain constraint based on the current security risk state assessment result is introduced. The boundary size of the adaptive trust domain constraint is negatively correlated with the risk level in the current security risk status assessment result; that is, the higher the risk level, the stricter the constraint on policy updates. During each policy iteration update, the difference in action distribution between the new policy and the old policy is calculated, and the difference is compared with the boundary of the adaptive trust domain constraint. If the difference is less than the boundary of the adaptive trust domain constraint, then the policy parameters are updated according to the gradient direction of the standard near-end policy optimization algorithm. If the difference is greater than or equal to the boundary of the adaptive trust domain constraint, the gradient of the standard near-end policy optimization algorithm is pruned, and the policy parameters are updated along the pruned gradient direction to ensure that the policy update step size is always within the range allowed by the adaptive trust domain constraint.
2. The method for full-process autonomous management of industrial safety intelligent agents for high-risk operations according to claim 1, characterized in that, Synchronous feature extraction is performed on the standard spatiotemporal state sequence to obtain a multi-dimensional state feature vector, including: For the components representing the physiological signals of workers in the standard spatiotemporal state sequence, a time-frequency joint analysis method is used to extract the frequency domain features of heart rate variability, the trend features of skin conductance response, and the features of body surface temperature gradient change, thus forming a sub-vector of personnel state features; For the components characterizing the operating state of the equipment in the standard spatiotemporal state sequence, a dynamic mode decomposition method is used to extract the energy distribution characteristics of the main vibration mode, the temperature change rate characteristics of key components, and the current harmonic distortion characteristics of the equipment, thus forming a sub-vector of equipment state characteristics. For the components characterizing the on-site environmental parameters in the standard spatiotemporal state sequence, a combination of spatial interpolation and field strength analysis is used to extract the spatial gradient characteristics of toxic gas concentration, the distribution and aggregation characteristics of inhalable particulate matter, and the main peak characteristics of the environmental noise spectrum, thus forming an environmental state feature sub-vector. For the components representing the execution steps of the work process in the standard spatiotemporal state sequence, the sequence pattern mining method is used to extract the time deviation features of the standard operation steps, the abnormal execution order features of key actions, and the tool usage interval duration features to form a process state feature sub-vector. The personnel state feature vector, the equipment state feature vector, the environment state feature vector, and the process state feature vector are concatenated and normalized to generate the multi-dimensional state feature vector.
3. The method for full-process autonomous management of industrial safety intelligent agents for high-risk operations according to claim 1, characterized in that, The method, based on a preset security knowledge graph, performs real-time security situation simulation on the multi-dimensional state feature vectors to generate a current security risk status assessment result, including: The multi-dimensional state feature vector is mapped to the entity attribute nodes of the preset security knowledge graph to activate the associated entities and relationship paths. The probability of risk transmission is calculated along the activated relational paths in the security knowledge graph, and the probability of risk transmission is calculated considering the coupling enhancement and inhibition effects between different risk factors. Based on the calculated risk transmission probability, and combined with the predefined risk event triggering conditions in the security knowledge graph, a risk event triggerability assessment is performed. By combining the calculated risk transmission probability with the risk event triggering assessment, the current security risk status assessment result, which includes risk level, location of major risk sources, and prediction of risk evolution trend, is derived.
4. The method for full-process autonomous management of industrial safety intelligent agents for high-risk operations according to claim 3, characterized in that, By combining the calculated risk transmission probability with the risk event triggering assessment, the current security risk status assessment result, which includes risk level, location of major risk sources, and prediction of risk evolution trends, is derived, including: A risk level quantification mapping table is established, and the calculation results of the risk transmission probability of different paths and the risk event triggering assessment results of different risk events are integrated through a weighted aggregation function to obtain a comprehensive risk quantification value. Based on the distribution of the comprehensive risk quantification value within a preset threshold range, a discrete risk level is determined; Tracing back the original risk transmission paths and risk events whose contribution exceeds a preset threshold during the weighted aggregation process, the source entities in the corresponding security knowledge graph are marked as the main risk source locations; Using time series forecasting methods, the changes in the comprehensive risk quantification value and its constituent elements at the current moment over a future period are predicted, thus forming the risk evolution trend prediction.
5. The method for full-process autonomous management of industrial safety intelligent agents for high-risk operations according to claim 4, characterized in that, The specific method for determining the boundary size of the adaptive trust region constraint includes: Extract the quantitative value corresponding to the risk level from the current security risk status assessment results; The quantified value of the risk level is input into a preset boundary mapping function, which is a monotonically decreasing piecewise linear function. The boundary mapping function outputs a positive real number, which serves as the boundary threshold for the adaptive trust domain constraint in this policy iteration. As the current security risk status assessment results are updated over time, the boundary thresholds of the adaptive trust domain constraint are also dynamically adjusted.
6. The method for full-process autonomous management of industrial safety intelligent agents for high-risk operations according to claim 5, characterized in that, The preliminary intervention sequence is logically consistent with the preset work safety procedure knowledge base, including: Analyze each intervention action in the preliminary intervention action sequence to identify its action type, target, and expected parameters; Retrieve all safety constraint rules related to the current high-risk work scenario from the preset work safety procedure knowledge base; The identified intervention actions, their target objects, and expected parameters are matched and logically deduced one by one with the retrieved safety constraint rules. If an intervention action directly conflicts with or implicitly contradicts any safety constraint rule, the corresponding intervention action is deemed to be logically inconsistent. For all intervention actions deemed logically inconsistent, within the permitted scope of the work safety procedure knowledge base, find functionally equivalent or lower-risk alternative actions to replace them. If no suitable alternative action is found, the corresponding intervention action is directly deleted.
7. The method for full-process autonomous management of industrial safety intelligent agents for high-risk operations according to claim 6, characterized in that, The process of converting the verified and corrected intervention action sequence into a standardized control command format that can be recognized and executed by the actuators in the industrial field includes: Establish a mapping dictionary from intervention actions to equipment control commands, wherein the mapping dictionary defines one or more underlying equipment control command templates corresponding to each type of intervention action; Based on the specific parameters of each intervention action in the verified and corrected intervention action sequence, fill the variable placeholders in the corresponding equipment control instruction template to generate specific equipment control instructions. Based on the physical layout and communication protocol of the industrial field actuators, add target device address code and communication protocol header information to each generated specific device control command; Based on the logical sequence of action execution and parallel execution relationship in the intervention action sequence, all device control instructions with added address and protocol information are time-sequenced to generate the final risk intervention instruction set containing complete control timing information.
8. The method for full-process autonomous management of industrial safety intelligent agents for high-risk operations according to claim 7, characterized in that, After achieving autonomous safety management of high-risk work processes by executing the set of risk intervention instructions through actuators in the industrial field, the method further includes: During and after the execution of the risk intervention instruction set, feedback sensing data streams continue to be collected through the multi-source heterogeneous sensors; The feedback sensing data stream is processed using the same procedure as the initial state data set to obtain the feedback security risk status assessment result. The feedback security risk status assessment result is compared and analyzed with the current security risk status assessment result before the execution of the risk intervention instruction set, and the risk status change measure is calculated. By using the risk state change metric, the parameters of the decision model are updated online through incremental learning, thereby optimizing the decision model's ability to make decisions on similar future risk states.
Citation Information
Patent Citations
Dynamic spectrum autonomous collaborative optimization method and system based on large language model
CN120474648A
Security risk assessment and operation management integrated method and system based on intelligent agent
CN121836383A