Occupational health comprehensive risk management evaluation system based on reinforcement learning

By constructing an ideal benchmark model based on standard operating procedures and a reinforcement learning agent, and combining dual-track differential computation and feature space similarity analysis, the problem of difficulty in identifying weak cumulative risks in existing technologies has been solved, achieving high-precision occupational health risk assessment and rapid response.

CN122066239APending Publication Date: 2026-05-19FOSHAN NUOWA ANPING DETECTION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FOSHAN NUOWA ANPING DETECTION CO LTD
Filing Date
2026-03-02
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing occupational health risk assessment methods are unable to accurately separate weak cumulative risk characteristics from complex environmental background noise, and lack an active identification mechanism for unknown compound risks, resulting in insufficient accuracy and timeliness of assessments.

Method used

Multi-source data acquisition terminals are used to acquire multi-source heterogeneous monitoring data. An ideal benchmark model based on standard operating procedures is constructed, and virtual risk simulation data is generated using reinforcement learning agents. Combined with dual-track differential calculation and feature space similarity analysis, accurate identification and assessment of occupational health risks can be achieved.

Benefits of technology

It effectively removes high-amplitude background signals generated during normal equipment operation, enabling precise capture of weak risk signals, reducing false alarm rates, improving the system's risk identification accuracy and adaptability in complex industrial environments, and shortening the recognition and response time for emergencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122066239A_ABST
    Figure CN122066239A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of occupational health safety monitoring and artificial intelligence data processing, in particular to an occupational health comprehensive risk management evaluation system based on reinforcement learning, which comprises a multi-source data acquisition step: acquiring multi-source heterogeneous monitoring data of a target operation scene; an ideal benchmark construction and simulation step: constructing an ideal benchmark model based on a standard operation program, and generating a virtual risk simulation data vector by using a reinforcement learning agent; a double-track difference calculation step: respectively calculating a real residual feature vector and a theoretical residual feature vector based on the ideal reference model; a similarity analysis and optimization step: performing feature space similarity analysis on the two to generate a risk assessment result and a strategy update reward signal; the problem that cumulative occupational hazards and transient environment fluctuation are difficult to distinguish is solved, and generalization ability and risk identification precision in a complex industrial environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of occupational health and safety monitoring and artificial intelligence data processing technology, specifically to an occupational health comprehensive risk management and assessment system based on reinforcement learning. Background Technology

[0002] Occupational health risk assessment refers to the process of identifying potential occupational hazards by collecting environmental parameters and personnel data from the workplace. Current occupational health risk assessment methods include: real-time alarm method based on sensor thresholds, periodic manual sampling and analysis method, and trend prediction method based on historical statistical data.

[0003] When conducting risk assessments based on existing technologies, it is difficult to accurately isolate early, weak, cumulative risk characteristics from complex environmental background noise. Furthermore, due to the variability of working conditions in industrial scenarios, there is a lack of mechanisms for proactively practicing and identifying unknown complex risks. Both of these factors reduce the accuracy and timeliness of risk assessments. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention provides a comprehensive occupational health risk management and assessment system based on reinforcement learning. Specifically, the technical solution of this invention includes:

[0005] Multi-source data acquisition terminals, data processing servers, and risk management terminals;

[0006] Among them, the multi-source data acquisition terminal is used to acquire multi-source heterogeneous monitoring data of the target operation scenario and send the multi-source heterogeneous monitoring data to the data processing server;

[0007] The data processing server is configured to: construct an ideal benchmark model based on preset standard operating procedure data, and generate virtual risk simulation data vectors using reinforcement learning agents;

[0008] Based on multi-source heterogeneous monitoring data and an ideal benchmark model, the actual residual feature vector is calculated.

[0009] Based on virtual risk simulation data vectors and ideal benchmark models, calculate theoretical residual feature vectors;

[0010] Feature space similarity analysis is performed on the actual residual feature vector and the theoretical residual feature vector to generate risk assessment results, and policy update reward signals for optimizing reinforcement learning agents are generated based on the risk assessment results.

[0011] The risk management terminal is used to receive and display risk assessment results.

[0012] Preferably, when constructing an ideal benchmark model, the data processing server is configured to perform the following operations: acquire preset standard operating procedure digitized data and theoretical physical constraint parameters;

[0013] Based on the digital data from standard operating procedures and theoretical physical constraint parameters, theoretical time-series waveform data under interference-free conditions are reconstructed.

[0014] The theoretical time-series waveform data is used as an ideal benchmark model.

[0015] Preferably, when generating virtual risk simulation data vectors, the data processing server is configured to perform the following operations:

[0016] Extract risk factor parameters from a pre-set expert knowledge base;

[0017] Transform risk factor parameters into computable numerical perturbation operators;

[0018] The control reinforcement learning agent combines different numerical perturbation operators and superimposes the combined numerical perturbation operators onto the data sequence of the ideal benchmark model to generate virtual risk simulation data vectors.

[0019] Preferably, the data processing server is configured to perform the following operations when calculating the actual residual eigenvector and the theoretical residual eigenvector:

[0020] Calculate the numerical difference sequence between multi-source heterogeneous monitoring data and the ideal benchmark model, and determine the numerical difference sequence as the actual residual feature vector;

[0021] Calculate the numerical difference sequence between the virtual risk simulation data vector and the ideal benchmark model, and determine the numerical difference sequence as the theoretical residual feature vector;

[0022] Among them, the real residual feature vector contains the real risk signal component and the environmental noise signal component, while the theoretical residual feature vector contains the pure risk feature component generated by the reinforcement learning agent.

[0023] Preferably, when the data processing server performs feature space similarity analysis to generate risk assessment results, it is configured to perform the following operations: calculate the spatial similarity value between the actual residual feature vector and the theoretical residual feature vector in the multidimensional feature space;

[0024] Obtain the preset similarity threshold; compare the spatial similarity value with the similarity threshold;

[0025] If the spatial similarity value is greater than the similarity judgment threshold, it is determined that there is a definite occupational health risk in the target work scenario, and a risk assessment result including the risk type is generated.

[0026] Preferably, the data processing server is further configured as follows:

[0027] In response to a spatial similarity value being less than or equal to a similarity determination threshold, and the signal amplitude of the detected real residual feature vector being greater than a preset noise amplitude threshold, it is determined that there is environmental noise from non-occupational factors in the target work scene.

[0028] Generate risk assessment results that include noise removal instructions.

[0029] Preferably, the data processing server is configured to perform the following operations when generating a reward signal:

[0030] Acquire preset historical real accident feature data; calculate the feature deviation value between the virtual risk simulation data vector and the historical real accident feature data;

[0031] In response to a feature deviation value being less than a preset matching tolerance threshold, a positive reward signal is generated to update the generation strategy parameters of the reinforcement learning agent.

[0032] Preferably, the multi-source data acquisition terminal includes: an environmental sensing module, a work behavior monitoring module, and a physiological feature acquisition module;

[0033] The environmental sensing module is configured to collect time-series data of environmental parameters.

[0034] The work behavior monitoring module is configured to collect equipment operation log data and protective equipment wearing status data;

[0035] The physiological characteristic acquisition module is configured to collect physiological health indicator data of the workers.

[0036] Preferably, the risk management terminal is also used to: receive risk assessment results;

[0037] Based on the risk type in the risk assessment results, the corresponding emergency response plan data is retrieved; the risk type and emergency response plan data are then visualized and rendered.

[0038] Compared with the prior art, the present invention has the following beneficial effects:

[0039] 1. This invention constructs an ideal benchmark model based on standard operating procedure data and theoretical physical constraint parameters, and reconstructs the theoretical time-series waveform under interference-free conditions; by calculating the numerical difference sequence between multi-source heterogeneous monitoring data and the ideal benchmark model, it effectively removes high-amplitude background signals generated during normal equipment operation, making minute abnormal fluctuations stand out; this solves the technical problem that traditional methods cannot distinguish cumulative occupational hazards from complex background noise, and achieves accurate capture of weak risk signals;

[0040] 2. This invention achieves qualitative identification of risks by comparing the feature space similarity between the actual residual feature vector and the theoretical residual feature vector, combined with a noise amplitude threshold judgment mechanism. When the similarity is low but the signal amplitude is large, it is judged as environmental noise from non-occupational factors and a rejection instruction is generated. This dual verification mechanism effectively distinguishes between real occupational health risks and transient environmental fluctuations, significantly reduces the false alarm rate under non-steady-state working conditions, and improves the anti-interference capability of the system.

[0041] 3. This invention utilizes a reinforcement learning agent to extract risk factors from an expert knowledge base and transform them into numerical perturbation operators. By combining these operators and superimposing them onto an ideal model, virtual risk simulation data is generated. This mechanism not only simulates known faults but also covers rare composite risk patterns. Combined with a reward signal feedback mechanism based on historical real accident feature data, the generation strategy of the agent is continuously optimized, realizing the model's adaptive learning and generalization capabilities for the variable risk features in complex industrial environments.

[0042] 4. The multi-source data acquisition terminal of this invention integrates environmental sensing, work behavior monitoring and physiological feature acquisition modules, and constructs a three-dimensional monitoring network integrating humans, machines and environment to ensure the multidimensional completeness of data; the risk management terminal not only displays the assessment results, but can also call and visualize emergency plan data according to the risk type; this closed-loop management from holographic perception to intuitive decision support greatly shortens the time for on-site personnel to recognize and respond to sudden occupational health events. Attached Figure Description

[0043] The present invention will be further explained below with reference to the accompanying drawings and embodiments:

[0044] Figure 1 This is a structural diagram of the system of the present invention. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.

[0046] Example 1:

[0047] Please see Figure 1 A comprehensive occupational health risk management and assessment system based on reinforcement learning includes: a multi-source data acquisition terminal, a data processing server, and a risk control terminal; wherein, the multi-source data acquisition terminal is used to acquire multi-source heterogeneous monitoring data of the target work scenario and send the multi-source heterogeneous monitoring data to the data processing server;

[0048] The data processing server is configured to: construct an ideal benchmark model based on preset standard operating procedure data, and generate virtual risk simulation data vectors using reinforcement learning agents; calculate the real residual feature vector based on multi-source heterogeneous monitoring data and the ideal benchmark model;

[0049] Based on virtual risk simulation data vectors and ideal benchmark models, theoretical residual feature vectors are calculated; feature space similarity analysis is performed on the actual residual feature vectors and theoretical residual feature vectors to generate risk assessment results, and policy update reward signals for optimizing reinforcement learning agents are generated based on the risk assessment results; the risk management terminal is used to receive and display the risk assessment results.

[0050] This embodiment details the overall architecture and core processing logic of the aforementioned system, aiming to solve the technical challenge of distinguishing between cumulative occupational hazards and transient environmental fluctuations in existing occupational health monitoring. Multi-source data acquisition terminals, deployed as the sensing layer in target scenarios such as chemical workshops or mining operations, collect multi-source heterogeneous monitoring data, including environmental parameters, equipment status, and personnel vital signs, around the clock. And transmit the synchronized data stream to the cloud via industrial Ethernet;

[0051] The data processing server utilizes a high-performance computing cluster to execute synthetic analysis logic. Its core lies in constructing a theoretically undisturbed reference frame, i.e., an ideal benchmark model. This allows weak anomalous signals to be separated from complex background noise. Furthermore, instead of directly predicting the future, the server drives a reinforcement learning agent to act as a disruptor, actively generating virtual risk simulation data vectors with specific fault characteristics. ;

[0052] The system performs dual-track difference calculations to obtain the real residual feature vectors that combine real risk and noise. and theoretical residual eigenvectors containing only pure risk characteristics. By comparing the similarity of the feature spaces of the two, the system realizes the derivation of unknown risk identification using known information, and sends the identification results to the risk management terminal for visualization. Through a dual-track coupling mechanism, the system utilizes the trial-and-error learning of RL agents to cover multiple modes from common faults to rare composite risks, which greatly improves the system's generalization ability and risk identification accuracy in complex industrial environments.

[0053] Example 2:

[0054] When constructing an ideal benchmark model, the data processing server is configured to perform the following operations: acquire preset standard operating procedure digitized data and theoretical physical constraint parameters; reconstruct theoretical time-series waveform data under interference-free conditions based on the standard operating procedure digitized data and theoretical physical constraint parameters; and use the theoretical time-series waveform data as the ideal benchmark model.

[0055] This embodiment further defines the construction process of the ideal benchmark model. This process abandons the traditional method of simply relying on historical averages and instead adopts a physical reconstruction strategy based on first principles. The server reads digital data from standard operating procedures, such as the rated speed of the wind turbine. Cross-sectional area of ​​exhaust duct And theoretical physical constraint parameters, which are explicitly limited here to the effective volume of the workspace. With pollutant generation rate Remove the diffusion coefficient from the original description that was not actually used in subsequent differential equations. This ensures a strict correspondence between variables and formulas; it utilizes a built-in physics simulation engine to calculate theoretical time-series waveform data under conditions of strict adherence to safety protocols and no external interference; and it addresses the issue of toxic gas concentrations. The system's physical model, based on the law of conservation of mass, is constructed using a first-order differential equation, and its calculation formula is as follows:

[0056]

[0057] To provide a theoretical reference, let's first consider the ventilation volume. Under the ideal simplification condition of constant values, the analytical solution of this equation is:

[0058]

[0059] in, The pollutant generation rate under standard operating conditions; This is a time variable, representing the duration calculated from the starting point of the physical model calculation; This refers to the theoretical ventilation volume; specifically, the SOP digital data includes preset equipment operation control sequences (such as the curve of fan frequency changing over time). The system converts the curve into a real-time ventilation volume function based on the fan similarity law. The specific conversion formula is as follows:

[0060]

[0061] in, The rated ventilation volume specified on the equipment nameplate. This is the rated power supply frequency; however, considering the actual operating conditions... Since these are often time-varying parameters, the analytical solutions described above cannot be directly applied; therefore, the system only uses the analytical solutions for initial steady-state verification.

[0062] The specific steps for this verification are as follows: Before starting the time-varying simulation, the system temporarily locks the input parameters to their rated constant values. The specific steps for this verification are as follows: During the system initialization phase or when the device is detected to be in steady-state operation, the system temporarily locks the input parameters to their rated constant values; the theoretical steady-state values ​​are calculated using analytical formulas. And perform pre-calculation using the RK4 solver;

[0063] The judgment criterion is: the numerical solution calculated at the end of the pre-simulation and... The relative error is determined to be less than the preset threshold. If the error is less than the preset threshold, the verification is considered to be successful. The processing logic after the verification fails is as follows: if the error exceeds the limit, the system determines that the numerical integration step size is too large or the parameters are abnormal. It automatically reduces the time step size to 0.1 times the original setting and retryes. If the retry fails three times in a row, a model building interruption exception is thrown to the terminal.

[0064] After the verification was passed, the fourth-order Runge-Kutta method was used to numerically integrate the physical model in the form of the above differential equations during actual calculations. The time step was set to be consistent with the sensor sampling frequency, and the initial conditions of the differential equations were set. The background value is usually set to 0, thereby generating a high-precision discretized theoretical sequence; The volume of the workspace;

[0065] The system uses the calculated sequence as a basis. In this step, to eliminate the imbalance of feature space weights caused by differences in dimensions and orders of magnitude between different monitoring dimensions, the data processing server must perform standardization processing on the theoretical data of each dimension before performing tensor splicing. Specifically, the modified Z-score standardization formula is used:

[0066]

[0067] in, and These are the baseline mean and standard deviation of the physical parameter, obtained statistically based on standard operating procedures. For example, a preset numerical stability constant. Its dimensions are set to be the same as those of the two quantities. The same applies to satisfy the physical meaning of addition, and is used to avoid program runtime errors caused by division by zero when a certain physical parameter is a constant value;

[0068] The specific statistical calculation method is as follows: A complete work cycle after the equipment enters steady-state operation, as defined in the standard operating procedure, is selected as the statistical window. The specific duration of this cycle is... The period duration of the metadata field is directly read from the SOP data package. For example, for continuous production equipment, the default setting is 24 hours to cover the complete diurnal temperature variation.

[0069] For batch production equipment, the process time from material feeding to material discharge is set. The theoretical time-series waveform data generated within this window is discretely sampled, and its arithmetic mean is calculated as... Calculate its population standard deviation as ;

[0070] To ensure an ideal benchmark model Multi-source heterogeneous monitoring data To ensure consistency across the vector dimension, the system extracts corresponding nominal values ​​based on SOPs and occupational health standards, constructs constants or nominal time series sequences, and performs the same standardization process as described above before comparing them with those obtained from physical calculations. The sequence is concatenated into tensors to ultimately generate a sequence with... A multidimensional ideal benchmark model with perfectly matched dimensions .

[0071] Example 3:

[0072] When generating virtual risk simulation data vectors, the data processing server is configured to perform the following operations: extract risk factor parameters from a preset expert knowledge base; convert the risk factor parameters into computable numerical perturbation operators; control the reinforcement learning agent to combine different numerical perturbation operators, and superimpose the combined numerical perturbation operators onto the data sequence of the ideal benchmark model to generate virtual risk simulation data vectors.

[0073] This embodiment details the specific mechanism by which a reinforcement learning agent generates virtual risk data; the server extracts specific risk factor parameters from an expert knowledge base, such as the resistance coefficient growth rate for the risk of dust filter clogging. To accommodate the discrete action space of the DQN agent, the system will use continuous parameters. Hierarchical discretization is performed, and the parameter configuration corresponding to each discrete value is encapsulated as an independent numerical perturbation operator; the above parameters are transformed into mathematical transformation functions, i.e., numerical perturbation operators. Taking filter clogging as an example, the operator is defined as a linear decay transformation of ideal data, and its calculation formula is:

[0074]

[0075] in, The ideal reference value at time t; The growth rate of the resistance coefficient; It is a non-negative physical constraint function;

[0076] Building upon this, reinforcement learning agents are specifically constructed as deep learning-based... network Decision-making architecture; state space of the agent Defined as the feature vector of the actual residual within the current time window That is, the difference between the actual monitoring data and the ideal benchmark model, or defined as the standardized actual monitoring data. The channels of the aligned ideal baseline model Mb are cascaded together; if cascading is used, the input layer dimension is adjusted accordingly. ;

[0077] If only the actual residual feature vector is used Then the input layer dimension is kept to be ,in To monitor the number of channels, The time window length is defined; by introducing a sequence containing real-world disturbance information as the state, the agent is ensured to select the most suitable numerical perturbation operator based on the current abnormal pattern.

[0078] The system is processing the state matrix. Before inputting into the network, a Flatten operation is performed to convert it into a one-dimensional vector to ensure that subsequent fully connected layers can capture the joint distribution features across channels.

[0079] The agent consists of two fully connected hidden layers and one output layer. To ensure the stability of gradient propagation in the early stages of training, the weight parameters of the fully connected layers are initialized using a He normal distribution to match the nonlinear characteristics of the ReLU activation function and prevent gradient vanishing.

[0080] The agent based on the current - Greedy exploration strategy, input state And output the action value This allows for the selection of either the optimal or exploratory combination of operators; in this strategy, the exploration probability... The specific initialization and decay scheme is as follows: the initial exploration rate is set to This ensures a full traversal of the operator combination space; after each training episode, it is updated according to the following formula, which is calculated as follows:

[0081]

[0082] Among them, minimum exploration rate attenuation coefficient To address the computational logic issues associated with the superposition of multiple operators, the system employs a cascaded function mapping mechanism to perform combination and superposition operations. The calculation formula is as follows:

[0083]

[0084] in, For the above filter clogging operator, For voltage fluctuation operators; parameters The angular frequency representing the voltage fluctuation of the power grid is calculated using the following formula:

[0085]

[0086] in, The power grid fluctuation frequency is used to characterize the instability of the power supply system; for The operator, the system explicitly constructs a mathematical model based on additive perturbation to ensure code-level executability, and its calculation formula is as follows:

[0087]

[0088] in, The preset dimensionless fluctuation amplitude coefficient is set to a value range of [0.1, 0.5]. This is based on the operating object... For standardized dimensionless data, the coefficients retain their dimensionless properties to match the data space; In each simulation step A random phase sampled uniformly within the interval is used to simulate asynchronous power grid noise details; in the formula... The physical time within the current time window is calculated using the following formula:

[0089]

[0090] in, For sampling point index, The sampling frequency is set to ensure dimensional consistency in angular frequency calculations. Furthermore, to construct a complete set of single operators and avoid missing action space definitions, the system also defines a sensor drift operator. Its mathematical definition is:

[0091]

[0092] in, The preset linear drift coefficients are used; at this point, the system explicitly defines the basic set of single operators as follows: This cascaded processing ensures that fault characteristics from different physical mechanisms can be mathematically and correctly coupled into the same data stream; the results after the operator is applied are used as virtual risk simulation data vectors. ;

[0093] The system pre-constructs an action space lookup table, which lists actions based on sets. The table lists all single operators and their allowed pairwise combinations based on the set. All single operators and their allowed pairwise combinations are defined as follows: the allowed pairwise combinations refer to the effective fault set after eliminating mutually exclusive faults based on the physical compatibility principle, and each item is assigned a unique integer index. The number of output layer nodes is equal to the length of the lookup table, thus ensuring a deterministic mapping between discrete action selection and continuous mathematical operations.

[0094] Example 4:

[0095] When calculating the actual residual feature vector and the theoretical residual feature vector, the data processing server is configured to perform the following operations: calculate the numerical difference sequence between the multi-source heterogeneous monitoring data and the ideal benchmark model, and determine the numerical difference sequence as the actual residual feature vector;

[0096] The numerical difference sequence between the virtual risk simulation data vector and the ideal benchmark model is calculated, and the numerical difference sequence is determined as the theoretical residual feature vector. The real residual feature vector contains the real risk signal component and the environmental noise signal component, and the theoretical residual feature vector contains the pure risk feature component generated by the reinforcement learning agent.

[0097] This embodiment clarifies the calculation logic of the dual-track differential mechanism, which aims to amplify minute abnormal changes through differential operations. It is worth noting that before performing numerical differential operations, in order to eliminate the unavoidable phase deviation between the actual operation sequence of the operator and the preset sequence of the standard operation procedure, the server performs a time synchronization calibration operation.

[0098] Specifically, the system uses a sliding cross-correlation algorithm to calculate multi-source heterogeneous monitoring data. Device current characteristic path and ideal reference model The correlation between theoretical energy consumption channels is analyzed, and the time delay corresponding to the maximum correlation coefficient is searched. The calculation formula is as follows:

[0099]

[0100] And based on this, The time axis is shifted and compensated to generate an aligned baseline model. The calculation formula is as follows:

[0101]

[0102] During this time translation process, for In cases where the index may exceed the array index range, the system executes a zero-padding strategy: when calculating the index... satisfy or At that time, direct assignment ; Time window length here Associated with the sensor sampling frequency set in claim 2, i.e. Each time step corresponds to a physical duration. This ensures that the time scale of the model input matches the physical process;

[0103] because It has been Z-score standardized, with a value of 0 representing the statistical mean. This operation prevents program crashes and avoids introducing human bias into the data.

[0104] To ensure the theoretical residual eigenvector To ensure purity and avoid introducing artificial phase errors, the system will align the reference model. rather than original As state input to the reinforcement learning agent of claim 3; the agent in Numerical perturbation operators are superimposed on the data to generate data that is similar to real-world data. Virtual risk simulation data vectors strictly aligned on the time axis ;

[0105] In this step, the inconsistency in dimensions between multi-source monitoring data and the ideal benchmark model must be addressed; before performing the difference calculation, the server calls the standardized parameters stored when constructing the ideal benchmark model, i.e., the benchmark mean. and standard deviation Multi-source heterogeneous monitoring data Perform isomorphic Z-score standardization to generate dimensionless standardized monitoring sequences. This step ensures that real-world data is mapped to the same numerical space as the ideal baseline model, thus avoiding numerical logic errors caused by direct subtraction.

[0106] The server performs two differential operations in parallel: the first operation... and The difference is used to generate the real residual feature vector. Second path calculation and The difference is used to generate the theoretical residual eigenvector. The calculation formula is as follows:

[0107]

[0108] in, For time series indexing; through this closed-loop processing logic of first aligning, then standardizing, then generating, and then differencing, the system effectively filters out pseudo residuals caused by inconsistent operation times, and accurately strips away high-amplitude background signals under normal operating conditions, making minute abnormal fluctuations stand out.

[0109] Example 5:

[0110] When the data processing server performs feature space similarity analysis to generate risk assessment results, it is configured to perform the following operations: calculate the spatial similarity value between the actual residual feature vector and the theoretical residual feature vector in the multidimensional feature space; and obtain the preset similarity judgment threshold.

[0111] The spatial similarity value is compared with the similarity judgment threshold; in response to the spatial similarity value being greater than the similarity judgment threshold, it is determined that there is a definite occupational health risk in the target work scenario, and a risk assessment result including the risk type is generated;

[0112] The data processing server is also configured to: determine that there is environmental noise from non-occupational factors in the target work scene when the spatial similarity value is less than or equal to the similarity judgment threshold and the signal amplitude of the detected real residual feature vector is greater than the preset noise amplitude threshold; and generate a risk assessment result containing noise removal instructions.

[0113] This embodiment details the risk assessment and false alarm elimination logic based on feature space similarity; the server uses a cosine similarity algorithm to calculate... and Spatial similarity value To prevent the cosine similarity calculation from crashing or generating errors due to a zero denominator, even when the system is running perfectly. The system introduces a logic branch to prevent division by zero, and its calculation formula is as follows:

[0114]

[0115] in, For example, a preset minimum threshold value. This logic ensures that the similarity is forced to 0 when there is no signal energy, thus correctly reflecting the state without specific risk.

[0116] Regarding the similarity threshold The specific method for obtaining this information involves the system performing statistical boundary delineation during the initialization phase: selecting a sample set of historically confirmed incidents. Compared with historical normal operating condition noise sample set ;calculate Mean of similarity distribution between each sample and the theoretical risk characteristics ,as well as Mean of similarity distribution with theoretical risk characteristics To balance precision and recall, The weighted average of the two distribution centers is used as the calculation formula:

[0117]

[0118] in, The weighting coefficient is usually taken as... This indicates that equal importance is given to the distribution centers of positive and negative samples; this weight can also be adjusted according to the preference for false negative and false positive rates in actual business operations; noise amplitude threshold. Set as:

[0119]

[0120] and The method for obtaining the data is as follows: During the initial deployment of the system or the periodic maintenance phase, select a clean background time window where no work activities are confirmed and the equipment is in standby mode, and calculate the actual residual feature vector during this period. Find the root mean square sequence and calculate the arithmetic mean of the RMS sequence. with standard deviation This serves as the statistical benchmark for background noise in the environment;

[0121] The system determines that the current real-world residual highly matches the specific risks simulated by the agent, and generates a definite risk assessment result; in response to And detected root mean square amplitude The system determines that the fluctuation is environmental noise from non-occupational factors, such as air pressure fluctuations caused by sudden weather, and generates a noise removal instruction.

[0122] This embodiment achieves a risk assessment with high robustness and interpretability through explicit threshold calculation logic and anomaly handling mechanism.

[0123] Example 6:

[0124] When generating a reward signal, the data processing server is configured to perform the following operations: acquire preset historical real accident feature data; calculate the feature deviation value between the virtual risk simulation data vector and the historical real accident feature data; and generate a positive reward signal in response to the feature deviation value being less than a preset matching tolerance threshold, so as to update the generation strategy parameters of the reinforcement learning agent.

[0125] This embodiment describes the self-evolution and calibration mechanism of a reinforcement learning agent; based on the risk type determined in the risk assessment results, such as centrifugal pump cavitation, the server retrieves corresponding, manually verified historical accident characteristic data from the database. This selective acquisition mechanism ensures that reward signals can guide the agent to perform high-fidelity fitting for specific risk patterns.

[0126] In calculating the characteristic deviation value At that time, in order to solve the problem of virtual data generation Compared with historical real data To address the nonlinear time warping problem caused by the different evolution speeds, such as the same leakage accident where the actual process may be slower than the simulated process, the system abandons the Euclidean distance with strict alignment requirements and instead adopts the Dynamic Time Warping (DTW) algorithm.

[0127] In constructing the cumulative distance matrix During the process, to prevent the algorithm from introducing excessive time distortions that violate physical causality in pursuit of the minimum mathematical distance—for example, matching a slow leak lasting several minutes as a transient disturbance on the order of milliseconds—the system applies a Sakoe-Chiba global path constraint window; specifically, the path search is restricted to a strip-shaped region near the diagonal. Inside, among which The preset time-flexible window width, for example, 10% of the sequence length;

[0128] Under this constraint, the system constructs a cumulative distance matrix. Algorithm initialization: Set its calculation formula as follows:

[0129]

[0130] And for the boundary The calculation formula is set as follows:

[0131]

[0132]

[0133] in, , The sequence lengths are used to establish the recursive boundaries; in this recursive calculation, the subscripts are... Specifically representing virtual risk simulation data vectors The first in index of time steps subscript Specifically representing historical real accident characteristics data The first in index of time steps The shortest path is found using dynamic programming, and the formula for its calculation is as follows:

[0134]

[0135] In the formula, the symbol Explicitly specified as Euclidean distance The norm, calculated for scalar data as the absolute difference and for vector data as the square root of the sum of squares of the differences in each dimension, ensures the non-negativity of the distance metric and the properties of the triangle inequality. After completing the full matrix calculation, the system obtains the cumulative distance to the destination. And perform path length normalization processing, the calculation formula is:

[0136]

[0137] Finally, the normalized result is determined as the eigenvalue. Based on this, the system executes the reward determination logic, the calculation formula of which is:

[0138]

[0139] In addition, to enable the system to detect unknown compound risks, a supplementary reward rule is set: in response to the feature deviation value The data vector is greater than or equal to the preset matching tolerance threshold, but a virtual risk simulation data vector is detected. When a key indicator exceeds a preset physical safety limit, an exploratory reward signal of +0.5 is generated; this mechanism ensures that the agent can autonomously explore unknown fault modes with high risk while fitting historical accidents.

[0140] Matching tolerance threshold The method for obtaining this data is as follows: calculate the intra-class average DTW distance of the historical accident dataset. The calculation formula is set as follows:

[0141]

[0142] in, The standard deviation of the within-class distance distribution; in response to the bias Less than the threshold The system generates a positive reward signal;

[0143] To utilize this reward signal for specific policy updates, the system constructed a complete DQN training loop; the system is configured with a capacity of An experience replay buffer is used to store state transition samples generated by the agent during the trial and error process. ,in As the ideal reference segment for the current moment, For the selected combination of operators, The reward value is calculated above; when the buffer data size meets the training requirements, the system random sampling size is... Mini-batch; construct the mean squared error (MSE) loss function, the calculation formula of which is:

[0144]

[0145] in, The target value is the action value predicted by the current network. According to the formula calculate, The value is set to 0.9; the system uses the Adam optimizer to minimize this loss function to update the network parameters. Meanwhile, target network parameters Every 100 training steps and A hard synchronization is performed; this mechanism ensures that the RL agent can not only learn the waveform characteristics of the risk, but also adapt to the accident evolution patterns at different time scales, achieving effective fitting of complex historical data.

[0146] Example 7:

[0147] The multi-source data acquisition terminal includes: an environmental sensing module, a work behavior monitoring module, and a physiological characteristic acquisition module; wherein, the environmental sensing module is configured to collect time-series data of environmental parameters; the work behavior monitoring module is configured to collect equipment operation log data and protective equipment wearing status data; and the physiological characteristic acquisition module is configured to collect physiological health indicator data of workers.

[0148] This embodiment specifies the hardware configuration of the multi-source data acquisition terminal and constructs a holographic perception network; the environmental sensing module integrates a laser dust sensor and a PID photoionization gas sensor to collect time-series data of environmental parameters such as PM2.5 and VOCs in real time.

[0149] The work behavior monitoring module reads the status of electronic tags on workers' protective equipment through RFID readers deployed at key nodes, and synchronously collects the start-up, shutdown and speed logs of production equipment through the PLC interface; at the same time, the physiological characteristic acquisition module uses an industrial-grade smart wristband to monitor the heart rate variability (HRV) and blood oxygen saturation of workers in real time.

[0150] Through the coordinated work of the three modules mentioned above, this embodiment achieves three-dimensional monitoring integrating humans, machines, and the environment, ensuring the multidimensionality and completeness of the input data, and providing a solid physical data foundation for subsequent complex correlation analysis by the server.

[0151] Example 8:

[0152] The risk management terminal is also used for: receiving risk assessment results; calling up corresponding emergency response plan data based on the risk type in the risk assessment results; and visualizing and rendering the risk type and emergency response plan data.

[0153] This embodiment illustrates the interactive functions of the risk management terminal, which aims to improve on-site decision-making efficiency; the terminal receives risk assessment results from the server in real time via a wireless network;

[0154] Based on the risk type clearly identified in the results, such as chlorine leakage trends, the system automatically retrieves and calls up pre-stored digital emergency response plan data, including evacuation route maps and lists of first aid measures.

[0155] By using digital twin technology, the location of risk sources is highlighted on a 3D factory model displayed on the terminal screen, and the execution process of the emergency plan is dynamically demonstrated. This embodiment transforms abstract risk data into an intuitive emergency action guide, which greatly shortens the cognition and decision-making response time of on-site managers in the event of a sudden occupational health incident, and effectively improves the standardization and efficiency of emergency response.

[0156] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A comprehensive occupational health risk management and assessment system based on reinforcement learning, characterized in that, include: Multi-source data acquisition terminals, data processing servers, and risk management terminals; The multi-source data acquisition terminal is used to acquire multi-source heterogeneous monitoring data of the target operation scenario and send the multi-source heterogeneous monitoring data to the data processing server. The data processing server is configured to: construct an ideal benchmark model based on preset standard operating procedure data, and generate virtual risk simulation data vectors using reinforcement learning agents; Based on the multi-source heterogeneous monitoring data and the ideal benchmark model, calculate the actual residual feature vector; Based on the virtual risk simulation data vector and the ideal benchmark model, calculate the theoretical residual feature vector; A feature space similarity analysis is performed on the actual residual feature vector and the theoretical residual feature vector to generate a risk assessment result, and a policy update reward signal for optimizing the reinforcement learning agent is generated based on the risk assessment result. The risk management terminal is used to receive and display the risk assessment results.

2. The occupational health comprehensive risk management and assessment system based on reinforcement learning according to claim 1, characterized in that, When constructing an ideal benchmark model, the data processing server is configured to perform the following operations: acquire preset standard operating procedure digitized data and theoretical physical constraint parameters; Based on the digital data of the standard operating procedure and the theoretical physical constraint parameters, the theoretical time-series waveform data under interference-free conditions is reconstructed. The theoretical time-series waveform data is used as the ideal benchmark model.

3. The occupational health comprehensive risk management and assessment system based on reinforcement learning according to claim 1, characterized in that, When generating virtual risk simulation data vectors, the data processing server is configured to perform the following operations: Extract risk factor parameters from a pre-set expert knowledge base; The risk factor parameters are converted into computable numerical perturbation operators; The reinforcement learning agent is controlled to combine different numerical perturbation operators, and the combined numerical perturbation operators are superimposed on the data sequence of the ideal benchmark model to generate the virtual risk simulation data vector.

4. The occupational health comprehensive risk management and assessment system based on reinforcement learning according to claim 1, characterized in that, The data processing server is configured to perform the following operations when calculating the actual residual feature vector and the theoretical residual feature vector: Calculate the numerical difference sequence between the multi-source heterogeneous monitoring data and the ideal benchmark model, and determine the numerical difference sequence as the actual residual feature vector; Calculate the numerical difference sequence between the virtual risk simulation data vector and the ideal benchmark model, and determine the numerical difference sequence as the theoretical residual feature vector; The real residual feature vector contains real risk signal components and environmental noise signal components, while the theoretical residual feature vector contains pure risk feature components generated by the reinforcement learning agent.

5. The occupational health comprehensive risk management and assessment system based on reinforcement learning according to claim 1, characterized in that, When performing feature space similarity analysis to generate risk assessment results, the data processing server is configured to perform the following operations: calculate the spatial similarity value between the actual residual feature vector and the theoretical residual feature vector in the multidimensional feature space; Obtain the preset similarity threshold; The spatial similarity value is compared with the similarity determination threshold; In response to the spatial similarity value being greater than the similarity determination threshold, it is determined that the target work scenario has a definite occupational health risk, and a risk assessment result including the risk type is generated.

6. The occupational health comprehensive risk management and assessment system based on reinforcement learning according to claim 5, characterized in that, The data processing server is also configured to: In response to the spatial similarity value being less than or equal to the similarity determination threshold, and the detected signal amplitude of the real residual feature vector being greater than a preset noise amplitude threshold, it is determined that there is environmental noise from non-occupational factors in the target work scene; Generate the risk assessment results that include noise removal instructions.

7. The occupational health comprehensive risk management and assessment system based on reinforcement learning according to claim 1, characterized in that, When generating a reward signal, the data processing server is configured to perform the following operations: Acquire preset historical real accident feature data; calculate the feature deviation value between the virtual risk simulation data vector and the historical real accident feature data; In response to the feature deviation value being less than a preset matching tolerance threshold, a positive reward signal is generated to update the generation strategy parameters of the reinforcement learning agent.

8. A comprehensive occupational health risk management and assessment system based on reinforcement learning according to any one of claims 1-7, characterized in that, The multi-source data acquisition terminal includes: an environmental sensing module, a work behavior monitoring module, and a physiological feature acquisition module; The environmental sensing module is configured to collect time-series data of environmental parameters. The work behavior monitoring module is configured to collect equipment operation log data and protective equipment wearing status data; The physiological characteristic acquisition module is configured to collect physiological health indicator data of the workers.

9. A comprehensive occupational health risk management and assessment system based on reinforcement learning according to any one of claims 1-7, characterized in that, The risk management terminal is also used to: receive the risk assessment results; Based on the risk type in the risk assessment results, the corresponding emergency response plan data is retrieved; the risk type and the emergency response plan data are then visualized and rendered.