Laboratory safety anomaly prediction system based on big data

By designing an abnormal sample compensation module, an identity-sensitive modulation layer, and a dynamic weighted fusion layer, combined with a spectral normalization mechanism and a conditional Wasserstein distance adversarial objective function, the quality of generated samples is optimized, solving the problems of sample missing, identity difference, and time-series dependency in traditional laboratory safety anomaly prediction systems, and achieving highly accurate and personalized safety risk prediction.

CN120430519BActive Publication Date: 2025-09-30SHIJIAZHUANG HIGH-TECH ZONE MOLI ART TRAINING SCHOOL CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510890237.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-30
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

Traditional laboratory safety anomaly prediction systems have problems such as a serious lack of laboratory safety anomaly samples, failure to consider the identity differences of laboratory personnel, ignoring the internal temporal dependencies of time series when generating anomaly samples, and unreasonable loss function design. As a result, the model has difficulty identifying potential safety risks during training, and the prediction results lack accuracy and personalization.

Method used

An abnormal sample compensation module was designed, and an identity grouping abnormal sample generation model, identity sensitivity modulation layer and dynamic weighted fusion layer were introduced. The spectral normalization mechanism and the conditional Wasserstein distance adversarial objective function were combined to optimize the quality of generated samples. The risk prediction loss function with identity modulation weight was adopted to achieve personalized risk identification.

Benefits of technology

The model's ability to identify and predict potential laboratory safety risks has been significantly improved, meeting the needs of personalized laboratory safety risk prediction in complex dynamic experimental environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120430519B_ABST
    Figure CN120430519B_ABST
Patent Text Reader

Abstract

The present invention discloses a laboratory safety anomaly prediction system based on big data, comprising a raw data acquisition module, a data optimization module, an abnormal sample compensation module, a safety anomaly prediction model construction module, and a laboratory safety anomaly prediction module. The present invention relates to the field of laboratory management data processing technology, and specifically to a laboratory safety anomaly prediction system based on big data. This solution innovatively designs an abnormal sample compensation module, achieving structured and conditional supplementation of scarce abnormal samples; innovatively introduces laboratory identity feature vectors, achieving personalized risk identification and accurate prediction at the laboratory personnel level; innovatively designs a conditional Wasserstein distance adversarial objective function based on laboratory identity feature vectors, significantly improving the authenticity and diversity of synthetic abnormal samples; and innovatively designs an identity safety risk prediction loss function, improving the model's prediction accuracy in complex dynamic laboratory environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of laboratory management data processing, and in particular to a laboratory safety anomaly prediction system based on big data. Background Art

[0002] The laboratory safety anomaly prediction system based on big data refers to the use of big data analysis technology to standardize the collection and analysis of multi-source heterogeneous data generated in the daily operation of the laboratory, identify potential safety risks, and issue early warning signals before the risks occur, thereby effectively improving the laboratory's risk perception ability and overall safety management efficiency, achieving early warning of laboratory safety incidents, and providing intelligent decision-making support for laboratory safety management.

[0003] However, there is a technical problem in the traditional laboratory safety anomaly prediction system that there is a serious lack of laboratory safety anomaly samples, which makes it difficult for the model to learn anomaly patterns during training, thereby reducing the model's ability to identify potential safety risks; the traditional laboratory safety anomaly prediction system does not take into account the identity differences of experimental personnel, resulting in the laboratory safety risk prediction results ignoring individual differences and lacking accuracy; the generation of abnormal samples in the traditional laboratory safety anomaly prediction system has technical problems such as the inability to effectively constrain discrete variables such as identity and the neglect of internal temporal dependencies of time series, resulting in significant differences in the distribution of generated samples and real abnormal samples, making it difficult to meet the demand for high-quality training samples in complex dynamic experimental environments, thereby limiting the effective construction and performance improvement of safety risk prediction models; the existing laboratory safety anomaly prediction models have a technical problem of using a unified error measurement standard in the loss function design, resulting in the model being unable to effectively distinguish the risk sensitivity differences reflected by different experimental personnel identities during training, thereby affecting the accuracy of the prediction results. Summary of the Invention

[0004] In view of the above situation, in order to overcome the defects of the existing technology, the present invention provides a laboratory safety anomaly prediction system based on big data. In view of the technical problem that there is a serious lack of laboratory safety anomaly samples in traditional laboratory safety anomaly prediction systems, which makes it difficult for the model to learn abnormal patterns during training, thereby reducing the model's ability to identify potential safety risks. This solution innovatively designs an abnormal sample compensation module. By modeling abnormal samples of identity grouping and constructing an identity-aware abnormal sample generation model, it realizes the structured and conditional supplement of scarce abnormal samples. The generated samples not only increase the number of abnormal data, but also maintain the distribution balance and representativeness under different identity conditions, effectively alleviating the constraints of the scarcity of real abnormal data on model training, and significantly improving the model's ability to identify potential laboratory safety risks; In view of the fact that the identity differences of experimental personnel are not taken into account in traditional laboratory safety anomaly prediction systems, which leads to the neglect of individual differences and lack of accuracy in laboratory safety risk prediction results, this solution innovatively introduces laboratory identity feature vectors, introduces identity sensitivity modulation layer and identity dynamic weighted fusion layer into the prediction model, realizes differentiated dynamic adjustment of local behavior risks and remote dependency risks, accurately captures the interactive logic between identity, behavior and risk, and enhances the model's ability to identify potential laboratory safety risks. The ability to perceive individual behavioral risks enables personalized risk identification and accurate prediction at the laboratory personnel level, significantly improving the model's generalization and prediction accuracy in scenarios with multiple and heterogeneous laboratory personnel. Traditional laboratory safety anomaly prediction systems often generate anomaly samples in a way that effectively constrains discrete variables such as identity and ignores the temporal dependencies within time series. This results in significant distribution differences between generated and real anomaly samples, making it difficult to meet the demand for high-quality training samples in complex dynamic experimental environments. This, in turn, limits the effective construction and performance improvement of safety risk prediction models. This innovative approach introduces a spectral normalization mechanism and a conditional input structure to enhance the discriminator's ability to discriminate temporal structure consistency, improving the stability and accuracy of adversarial training. A conditional Wasserstein distance adversarial objective function based on laboratory identity feature vectors, combined with a gradient penalty strategy, optimizes the quality of generated samples while effectively constraining the dynamic control of temporal sequence structure and identity attributes during the generation process. This enables the generation of high-fidelity anomaly samples with multi-identity and temporal preservation properties, significantly improving the authenticity and diversity of synthetic samples. Ultimately, this effectively improves the training stability and prediction accuracy of the safety risk prediction model under multiple identities.To address the technical issue of using a unified error metric in the loss function design of existing laboratory safety anomaly prediction models, which results in the model being unable to effectively distinguish the risk sensitivity differences reflected by different laboratory personnel during training, thereby affecting the accuracy of prediction results, this solution innovatively designs an identity-based safety risk prediction loss function. By introducing identity modulation weights and combining them with a conditional risk variance estimation function based on identity and behavioral characteristics, this approach achieves personalized adjustment and dynamic weighting of prediction errors. This significantly enhances the model's ability to identify anomalies for heterogeneous individuals with multiple identities, improves the loss function's modeling of risk uncertainty, and ultimately improves the model's prediction accuracy in complex and dynamic laboratory environments, effectively meeting the practical needs of personalized laboratory safety risk prediction.

[0005] The technical solution adopted by the present invention is as follows: the laboratory safety anomaly prediction system based on big data provided by the present invention includes a raw data acquisition module, a data optimization module, an abnormal sample compensation module, a safety anomaly prediction model construction module and a laboratory safety anomaly prediction module;

[0006] The raw data acquisition module specifically collects data through the laboratory safety management platform to obtain laboratory safety raw data;

[0007] The data optimization module specifically performs data preprocessing, laboratory identity feature vector construction and data label definition on the laboratory safety raw data to obtain laboratory safety optimization data;

[0008] The abnormal sample compensation module is specifically constructed by grouping abnormal samples by identity, generating an identity abnormal feature vector set, and obtaining an enhanced training data set by constructing and training an abnormal sample generation model for identity grouping;

[0009] The security anomaly prediction model construction module specifically adopts a multi-layer structure model including a local risk perception layer, a remote dependency risk extraction layer, an experimenter identity sensitivity modulation layer, an identity dynamic weighted fusion layer and a risk prediction output layer to construct a security anomaly prediction model, and uses the enhanced training data set as input and adopts the identity-modulated risk prediction loss function to perform model training to obtain a trained security anomaly prediction model;

[0010] The laboratory safety anomaly prediction module specifically inputs real-time data into the trained safety anomaly prediction model to obtain laboratory safety risk prediction results.

[0011] Furthermore, the raw data acquisition module specifically collects the raw data required for laboratory safety anomaly prediction through the laboratory safety management platform to obtain laboratory safety raw data; the laboratory safety raw data includes historical laboratory safety data and real-time laboratory safety data; the historical laboratory safety data and real-time laboratory safety data both include laboratory personnel information data, laboratory equipment operation behavior data, experimental operation record data, laboratory equipment data and laboratory environment data; the historical laboratory safety data also includes historical laboratory safety results; the historical laboratory safety results include normal laboratory safety and abnormal laboratory safety.

[0012] Furthermore, the data optimization module specifically includes the following steps:

[0013] Data preprocessing, specifically data cleaning, standardization conversion and time series construction of raw data;

[0014] Laboratory identity feature vector construction is used to model the differences in different personnel identities in dynamic risk prediction. Specifically, the multi-dimensional identity fields in the laboratory personnel information data are encoded into a low-dimensional structured laboratory identity feature vector through the autoencoder algorithm;

[0015] Data label definition, specifically assigning labels to each time segment sample based on historical laboratory safety results, marking them as normal laboratory safety or abnormal laboratory safety.

[0016] Furthermore, the abnormal sample compensation module specifically includes identity grouping abnormal sample construction, identity grouping abnormal sample generation model construction, model training and acquisition of enhanced training data set, including the following steps:

[0017] The construction of identity-grouped abnormal samples involves classifying and extracting abnormal samples based on the identity type of the laboratory personnel based on labeled laboratory safety optimization data. For each type of abnormal sample, a sliding statistical analysis method is used to extract features from the time series data of the abnormal samples to generate a time series behavior feature vector. This feature vector is then jointly encoded with the corresponding laboratory identity feature vector to form an identity abnormality feature vector. Abnormal samples under all identity types are then merged to form a set of identity abnormality feature vectors.

[0018] Constructing a model for generating abnormal samples by identity grouping includes the following steps:

[0019] Design a generator to generate synthetic anomaly samples that match identity characteristics from random noise. Specifically, a multi-layer fully connected neural network based on a residual structure is used, consisting of an input layer, three sequentially connected residual blocks, and an output layer. The data passes through the multi-layer fully connected neural network to obtain a synthetic anomaly sample vector.

[0020] Design a discriminator, specifically a multi-layer perceptron structure with spectral normalization, consisting of an input layer, three sequentially connected hidden layers, and an output layer. The data is processed by the multi-layer perceptron structure to obtain a continuous real-valued score;

[0021] The overall adversarial objective function is designed. Specifically, the overall adversarial objective function is constructed by combining the conditional Wasserstein distance, introducing the laboratory identity feature vector y, and using the gradient penalty mechanism. The formula used is as follows:

[0022] ;

[0023] Where, represents the overall adversarial objective function, represents the real sample, represents the noise distribution, Indicates identity information The synthetic samples generated by the generator below, represents the authenticity score given by the discriminator to the sample x, represents the gradient penalty coefficient, Indicates sampling from the true abnormal sample distribution, represents sampling from a noise distribution, Indicates sampling from the intermediate samples obtained by linear interpolation between the real samples and the generated samples; represents the gradient norm at the middle sample, represents the objective maximization of the discriminator, The goal of the generator is to minimize Represents the authenticity score of the generated sample by the discriminator. Under the conditions; represents the laboratory identity feature vector, z represents the random noise vector, Represents the real sample vector, which is divided into synthetic abnormal sample vector and identity abnormal feature vector, represents the intermediate sample vector, represents the mathematical expectation function;

[0024] Define the loss function, specifically the discriminator loss function Designed as the opposite of the overall adversarial objective function, the generator loss function is designed to be ;

[0025] Model training, specifically, takes the identity anomaly feature vector set as input data, trains the identity grouping anomaly sample generation model, and minimizes the generator loss function by alternating and maximize the discriminator loss function , until convergence, and obtain the trained identity grouping abnormal sample generation model;

[0026] An enhanced training dataset is obtained by inputting the identity anomaly feature vector set and the random noise vector into the trained identity grouping anomaly sample generation model, and using the generator to generate synthetic anomaly samples that meet various identity conditions; then, an anomaly scoring function is used to evaluate the quality of the generated samples, and low-quality samples that are significantly different from the real anomaly samples or do not conform to the behavioral patterns of the identity are eliminated; the retained high-quality synthetic anomaly samples are fused with the identity anomaly feature vector set and normal samples to construct an enhanced training dataset with identity differentiation and balanced anomaly distribution.

[0027] Furthermore, the safety anomaly prediction model construction module specifically includes the following steps:

[0028] The local risk perception layer uses a multi-layer bidirectional long short-term memory network and performs a maximum pooling operation in the time dimension on the hidden state output of the last layer to obtain the local risk characteristics under the time window. ;

[0029] The remote dependency risk extraction layer specifically adopts a multi-layer Transformer encoding network based on the self-attention mechanism. The Transformer encoding network includes a multi-head attention mechanism and a position feedforward network, and is equipped with residual connections and layer normalization processing; and the final output of the multi-layer Transformer is average pooled to obtain the remote risk features. ;

[0030] The experimenter identity sensitivity modulation layer is specifically to enhance the representation ability of identity features by introducing a multi-head self-attention mechanism to obtain the identity enhancement vector , and input them into two multilayer perceptron network channels with the same structure but independent parameters, and nonlinear mapping generates local risk modulation factors and remote risk modulators The multi-layer perceptron network channel includes a shared identity feature extraction network and two sets of independent output branches; the risk modulation factor is used to dynamically reflect the sensitivity adjustment requirements of identity differences to different risk sources;

[0031] Identity dynamic weighted fusion layer, specifically based on the local risk modulation factor and remote risk modulators Perform weighted summation on local risk features and remote risk features to obtain identity-aware risk features ;

[0032] The risk prediction output layer is specifically to transform the identity perception risk characteristics Input the fully connected neural network layer and pass the Sigmoid activation function to obtain the laboratory safety anomaly prediction results. ;

[0033] Design identity security risk prediction loss function, specifically combining identity modulation weights The risk sensitivity of the prediction errors of different roles in the laboratory operation process is modulated with the conditional risk variance, and the identity-differentiated loss function is constructed. The formula used is as follows:

[0034] ;

[0035] Where, represents the identity security risk prediction loss function value, n represents the number of training samples, represents the true risk value of the i-th sample, represents the risk prediction value of the i-th sample, represents the input time series behavior characteristics of the i-th sample, represents the laboratory identity feature vector corresponding to the i-th sample, represents the natural logarithm function, represents the risk variance estimate of the i-th sample predicted by the model, represents the identity modulation weight of the i-th sample, which is obtained by learning the laboratory identity feature vector through a multi-layer perceptron;

[0036] Construct and train the model, specifically through the local risk perception layer, the remote dependency risk extraction layer, the experimenter identity sensitivity modulation layer, the identity dynamic weighted fusion layer and the risk prediction output layer, to construct a security anomaly prediction model, based on the enhanced training data set as training data, use the identity security risk prediction loss function to train the model as the optimization target, and obtain the trained security anomaly prediction model.

[0037] Furthermore, the laboratory safety anomaly prediction module specifically uses the real-time laboratory safety data in the laboratory safety optimization data as input data of the trained safety anomaly prediction model to obtain laboratory safety anomaly prediction results, and based on the laboratory safety anomaly prediction results, realizes early perception of potential abnormal events in the laboratory.

[0038] The beneficial effects achieved by the present invention using the above scheme are as follows:

[0039] (1) In order to solve the technical problem of serious lack of laboratory safety abnormality samples in traditional laboratory safety abnormality prediction systems, which makes it difficult for the model to learn abnormal patterns during training, thereby reducing the model's ability to identify potential safety risks, this solution innovatively designs an abnormal sample compensation module. By modeling abnormal samples of identity grouping and constructing an identity-aware abnormal sample generation model, it realizes the structured and conditional supplementation of scarce abnormal samples. The generated samples not only increase the number of abnormal data, but also maintain the distribution balance and representativeness under different identity conditions, effectively alleviating the constraints of the scarcity of real abnormal data on model training, and significantly improving the model's ability to identify potential laboratory safety risks.

[0040] (2) In view of the fact that the traditional laboratory safety anomaly prediction system does not take into account the identity differences of laboratory personnel, resulting in the laboratory safety risk prediction results ignoring individual differences and lacking accuracy, this scheme innovatively introduces laboratory identity feature vectors, introduces identity sensitivity modulation layer and identity dynamic weighted fusion layer into the prediction model, realizes differentiated dynamic adjustment of local behavior risk and remote dependency risk, accurately captures the interaction logic between identity, behavior and risk, enhances the model's ability to perceive individual behavior risks, realizes personalized risk identification and accurate prediction at the laboratory personnel level, and significantly improves the model's generalization ability and prediction accuracy in scenarios with multi-identity heterogeneous laboratory personnel.

[0041] (3) In view of the technical problems of the inability to effectively constrain discrete variables such as identity and the neglect of the internal temporal dependencies of time series in the generation of abnormal samples in traditional laboratory safety anomaly prediction systems, there are significant differences in the distribution of generated samples and real abnormal samples, which makes it difficult to meet the demand for high-quality training samples in complex dynamic experimental environments, thereby limiting the effective construction and performance improvement of safety risk prediction models. This scheme innovatively introduces a spectral normalization mechanism and a conditional input structure to enhance the discriminator's ability to distinguish the consistency of temporal structure, thereby improving the stability and accuracy of adversarial training; a conditional Wasserstein distance adversarial objective function based on the laboratory identity feature vector is designed, and combined with a gradient penalty strategy, while optimizing the quality of generated samples, it effectively constrains the dynamic regulation of temporal sequence structure and identity attributes in the generation process, and realizes the generation of high-fidelity abnormal samples with multi-identity and temporal preservation characteristics, significantly improving the authenticity and diversity of synthetic samples, and ultimately effectively improving the model training stability and prediction accuracy of the safety risk prediction model under multiple identities.

[0042] (4) In order to solve the technical problem of using a unified error metric standard in the loss function design of the existing laboratory safety anomaly prediction model, which results in the model being unable to effectively distinguish the risk sensitivity differences reflected by different laboratory personnel identities during training, thereby affecting the accuracy of the prediction results, this scheme innovatively designs an identity safety risk prediction loss function. By introducing identity modulation weights and combining the conditional risk variance estimation function based on identity and behavioral characteristics, it realizes personalized adjustment and dynamic weighting of the prediction error, significantly enhances the model's ability to identify anomalies of heterogeneous individuals with multiple identities, improves the loss function's modeling effect on risk uncertainty, and ultimately improves the model's prediction accuracy in complex dynamic laboratory environments, effectively meeting the actual needs of personalized laboratory safety risk prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 This is a module diagram of the laboratory safety anomaly prediction system based on big data provided by the present invention;

[0044] Figure 2 This is a flow chart of the abnormal sample compensation module;

[0045] Figure 3 A flowchart for constructing an abnormal sample generation model for identity grouping in the abnormal sample compensation module;

[0046] Figure 4 Schematic diagram of the process of building modules for the security anomaly prediction model;

[0047] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention. DETAILED DESCRIPTION

[0048] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0049] In the description of the present invention, it should be understood that terms such as "up", "down", "front", "back", "left", "right", "top", "bottom", "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the system or element referred to must have a specific direction, be constructed and operated in a specific direction. Therefore, they should not be understood as limiting the present invention.

[0050] Example 1, see Figure 1 The laboratory safety anomaly prediction system based on big data provided by the present invention includes a raw data acquisition module, a data optimization module, an abnormal sample compensation module, a safety anomaly prediction model construction module and a laboratory safety anomaly prediction module;

[0051] The raw data acquisition module specifically collects data through the laboratory safety management platform, obtains laboratory safety raw data, and sends the data to the data optimization module;

[0052] The data optimization module receives the data sent by the original data acquisition module, obtains laboratory safety optimization data by performing data preprocessing, laboratory identity feature vector construction and data label definition on the laboratory safety original data, and sends the data to the abnormal sample compensation module;

[0053] The abnormal sample compensation module receives the data sent by the data optimization module, specifically generates an identity anomaly feature vector set by constructing an abnormal sample of identity grouping, and obtains an enhanced training data set by constructing and training an abnormal sample generation model of identity grouping, and sends the data to the security anomaly prediction model construction module;

[0054] The safety anomaly prediction model construction module receives data sent by the abnormal sample compensation module, specifically adopts a multi-layer structure model including a local risk perception layer, a remote dependency risk extraction layer, an experimenter identity sensitivity modulation layer, an identity dynamic weighted fusion layer and a risk prediction output layer to construct a safety anomaly prediction model, and uses the enhanced training data set as input, adopts the identity-modulated risk prediction loss function to perform model training, obtains the trained safety anomaly prediction model, and sends the data to the laboratory safety anomaly prediction module;

[0055] The laboratory safety anomaly prediction module specifically inputs real-time data into the trained safety anomaly prediction model to obtain laboratory safety risk prediction results.

[0056] By performing the above operations, this solution innovatively designs an abnormal sample compensation module to address the technical problem of a serious lack of laboratory safety abnormality samples in traditional laboratory safety abnormality prediction systems, which makes it difficult for the model to learn abnormal patterns during training, thereby reducing the model's ability to identify potential safety risks. By modeling abnormal samples based on identity grouping and constructing an identity-aware abnormal sample generation model, it achieves structured and conditional supplementation of scarce abnormal samples. The generated samples not only increase the number of abnormal data, but also maintain the distribution balance and representativeness under different identity conditions, effectively alleviating the constraints on model training caused by the scarcity of real abnormal data, and significantly improving the model's ability to identify potential laboratory safety risks.

[0057] Example 2, see Figure 1 This embodiment is based on the above embodiment. Specifically, the raw data acquisition module collects raw data required for laboratory safety anomaly prediction through the laboratory safety management platform to obtain laboratory safety raw data; the laboratory safety raw data includes historical laboratory safety data and real-time laboratory safety data; the historical laboratory safety data and real-time laboratory safety data both include laboratory personnel information data, laboratory equipment operation behavior data, experimental operation record data, laboratory equipment data and laboratory environment data; the historical laboratory safety data also includes historical laboratory safety results; the historical laboratory safety results include normal laboratory safety and abnormal laboratory safety;

[0058] The laboratory personnel information data includes name, identity type, years of laboratory experience, laboratory working hours, relevant experimental operation training and experimental equipment operation frequency; the identity types are divided into interns, junior researchers, intermediate researchers and senior researchers; the laboratory equipment operation behavior data includes the operator's name, operation type, operation duration and operation time; the experimental operation record data includes the operator's name, experiment type, experiment name, experiment cycle and category of chemicals used; the laboratory equipment data includes equipment operating status, equipment load, equipment current, equipment voltage, equipment power consumption, equipment age and equipment maintenance records; the laboratory environment data includes temperature, humidity, gas concentration, noise intensity, dust concentration and environmental parameter change trends.

[0059] Example 3, see Figure 1 This embodiment is based on the above embodiment, and the data optimization module specifically includes the following steps:

[0060] Data preprocessing, specifically data cleaning, standardization conversion and time series construction of raw data;

[0061] The data cleaning specifically involves processing missing items, outliers, and duplicate records in the original data, including field unification, time format standardization, and null value filling;

[0062] The standardized conversion specifically includes performing one-hot encoding and normalization processing on the categorical field and the continuous field respectively;

[0063] The time series construction is specifically to perform time slicing on the laboratory equipment operation behavior data, experimental operation record data, laboratory equipment data and laboratory environment data in the laboratory safety original data according to a sliding window of fixed length;

[0064] Laboratory identity feature vector construction is used to model the differences in different personnel identities in dynamic risk prediction. Specifically, the multi-dimensional identity fields in the laboratory personnel information data are encoded into a low-dimensional structured laboratory identity feature vector through the autoencoder algorithm;

[0065] Data label definition, specifically assigning labels to each time segment sample based on historical laboratory safety results, marking them as normal laboratory safety or abnormal laboratory safety.

[0066] Example 4, see Figure 1 、 Figure 2 and Figure 3 This embodiment is based on the above embodiment. The abnormal sample compensation module specifically includes constructing abnormal samples by identity grouping, constructing an abnormal sample generation model by identity grouping, model training, and obtaining an enhanced training data set, including the following steps:

[0067] The construction of identity-grouped abnormal samples involves classifying and extracting abnormal samples based on the identity type of the laboratory personnel based on labeled laboratory safety optimization data. For each type of abnormal sample, a sliding statistical analysis method is used to extract features from the time series data of the abnormal samples to generate a time series behavior feature vector. This feature vector is then jointly encoded with the corresponding laboratory identity feature vector to form an identity abnormality feature vector. Abnormal samples under all identity types are then merged to form a set of identity abnormality feature vectors.

[0068] Constructing a model for generating abnormal samples by identity grouping includes the following steps:

[0069] Design a generator to generate synthetic anomaly samples that match identity characteristics from random noise. Specifically, a multi-layer fully connected neural network based on a residual structure is used, consisting of an input layer, three sequentially connected residual blocks, and an output layer. The data passes through the multi-layer fully connected neural network to obtain a synthetic anomaly sample vector.

[0070] The residual block includes a fully connected layer, a normalization layer, a nonlinear activation function and a skip connection; the output layer is used to generate a synthetic abnormal sample vector consistent with the original abnormal sample structure;

[0071] Design a discriminator, specifically a multi-layer perceptron structure with spectral normalization, consisting of an input layer, three sequentially connected hidden layers, and an output layer. The data is processed by the multi-layer perceptron structure to obtain a continuous real-valued score;

[0072] The hidden layer is composed of a fully connected layer, a LeakyReLU activation function, and a spectral normalization layer in sequence, which is used to enhance the nonlinear modeling capability and control the Lipschitz constant of each layer;

[0073] The output layer is a fully connected layer that outputs a continuous real-valued score reflecting its authenticity;

[0074] The overall adversarial objective function is designed. Specifically, the overall adversarial objective function is constructed by combining the conditional Wasserstein distance, introducing the laboratory identity feature vector y, and using the gradient penalty mechanism. The formula used is as follows:

[0075] ;

[0076] Where, represents the overall adversarial objective function, represents the real sample, represents the noise distribution, Indicates identity information The synthetic samples generated by the generator below, represents the authenticity score given by the discriminator to the sample x, represents the gradient penalty coefficient, Indicates sampling from the true abnormal sample distribution, represents sampling from a noise distribution, Indicates sampling from the intermediate samples obtained by linear interpolation between the real samples and the generated samples; represents the gradient norm at the middle sample, represents the objective maximization of the discriminator, The goal of the generator is to minimize Represents the authenticity score of the generated sample by the discriminator. Under the conditions; represents the laboratory identity feature vector, z represents the random noise vector, Represents the real sample vector, which is divided into synthetic abnormal sample vector and identity abnormal feature vector, represents the intermediate sample vector, represents the mathematical expectation function;

[0077] Define the loss function, specifically the discriminator loss function Designed as the opposite of the overall adversarial objective function, the generator loss function is designed to be ; The formula used is as follows:

[0078] ;

[0079] Model training, specifically, takes the identity anomaly feature vector set as input data, trains the identity grouping anomaly sample generation model, and minimizes the generator loss function by alternating and maximize the discriminator loss function , until convergence, and obtain the trained identity grouping abnormal sample generation model;

[0080] An enhanced training dataset is obtained by inputting the identity anomaly feature vector set and the random noise vector into the trained identity grouping anomaly sample generation model, and using the generator to generate synthetic anomaly samples that meet various identity conditions; then, an anomaly scoring function is used to evaluate the quality of the generated samples, and low-quality samples that are significantly different from the real anomaly samples or do not conform to the behavioral patterns of the identity are eliminated; the retained high-quality synthetic anomaly samples are fused with the identity anomaly feature vector set and normal samples to construct an enhanced training dataset with identity differentiation and balanced anomaly distribution.

[0081] By performing the above operations, the technical problems of abnormal sample generation in traditional laboratory safety anomaly prediction systems, such as the inability to effectively constrain discrete variables such as identity and the neglect of internal temporal dependencies of time series, result in significant differences in the distribution of generated samples and real abnormal samples, making it difficult to meet the demand for high-quality training samples in complex dynamic experimental environments, thereby limiting the effective construction and performance improvement of safety risk prediction models. This solution innovatively introduces a spectral normalization mechanism and a conditional input structure to enhance the discriminator's ability to distinguish temporal structure consistency, thereby improving the stability and accuracy of adversarial training. A conditional Wasserstein distance adversarial objective function based on the laboratory identity feature vector is designed, combined with a gradient penalty strategy. While optimizing the quality of generated samples, it effectively constrains the dynamic regulation of temporal sequence structure and identity attributes during the generation process, achieving high-fidelity abnormal sample generation with multi-identity and temporal sequence preservation characteristics, significantly improving the authenticity and diversity of synthetic samples, and ultimately effectively improving the model training stability and prediction accuracy of the safety risk prediction model under multiple identities.

[0082] Example 5, see Figure 1 and Figure 4 This embodiment is based on the above embodiment, and the security anomaly prediction model construction module specifically includes the following steps:

[0083] The local risk perception layer is used to model the local dynamic changes of the recent operating behavior of the experimenters and the experimental equipment environment, and to capture the potential safety risk signs in a short period of time. Specifically, it uses a multi-layer bidirectional long short-term memory network and performs a maximum pooling operation in the time dimension on the hidden state output of the last layer to obtain the local risk characteristics under the time window. ; The formula used is as follows:

[0084] ;

[0085] Where, represents the maximum pooling operation, represents the bidirectional long short-term memory network unit of the lth layer, Indicates the The hidden state output of the layer;

[0086] The remote dependency risk extraction layer is used to mine non-local risk dependencies across time spans in laboratory operation sequences and capture potential risk factors with lags and cross-behavior coupling. Specifically, it uses a multi-layer Transformer encoding network based on a self-attention mechanism. The Transformer encoding network contains a multi-head attention mechanism and a position feedforward network, and is equipped with residual connections and layer normalization processing. The final output of the multi-layer Transformer is average pooled to obtain remote risk features. ; The formula used is as follows:

[0087] ;

[0088] ;

[0089] Where, represents the output of the lth layer Transformer, Indicates the Layer Transformer output, represents the multi-head attention function, represents the position feedforward network, Representation layer normalization operation, represents the average pooling operation;

[0090] The experimenter identity sensitivity modulation layer is used to convert the experimenter identity information into a dynamic adjustment factor for local risk characteristics and remote risk characteristics, and realize the identity-based differentiated risk modeling capability. Specifically, the multi-head self-attention mechanism is introduced to enhance the representation ability of identity characteristics and obtain the identity enhancement vector , and input them into two multilayer perceptron network channels with the same structure but independent parameters, and nonlinear mapping generates local risk modulation factors and remote risk modulators The multi-layer perceptron network channel includes a shared identity feature extraction network and two independent output branches. The risk modulation factor is used to dynamically reflect the sensitivity adjustment requirements of identity differences to different risk sources. The formula used is as follows:

[0091] ;

[0092] ;

[0093] ;

[0094] Where, represents the laboratory identity feature vector, represents the Sigmoid activation function, and Represents the shared fully connected mapping parameters, and represents the independent weights and biases of the local risk modulation factor output branches, and represents the independent weights and biases of the remote risk modulation factor output branches;

[0095] The identity dynamic weighted fusion layer is used to realize the personalized dynamic integrated representation of multi-source risk characteristics under different identity conditions; specifically, it is based on the local risk modulation factor and remote risk modulators Perform weighted summation on local risk features and remote risk features to obtain identity-aware risk features ; The formula used is as follows:

[0096] ;

[0097] The risk prediction output layer is specifically to transform the identity perception risk characteristics Input the fully connected neural network layer and pass the Sigmoid activation function to obtain the laboratory safety anomaly prediction results. ; The formula used is as follows:

[0098] ;

[0099] Where, and Represents the weight and bias parameters of the risk prediction output layer;

[0100] Design identity security risk prediction loss function, specifically combining identity modulation weights The risk sensitivity of the prediction errors of different roles in the laboratory operation process is modulated with the conditional risk variance, and the identity-differentiated loss function is constructed. The formula used is as follows:

[0101] ;

[0102] Where, represents the identity security risk prediction loss function value, n represents the number of training samples, represents the true risk value of the i-th sample, represents the risk prediction value of the i-th sample, represents the input time series behavior characteristics of the i-th sample, represents the laboratory identity feature vector corresponding to the i-th sample, represents the natural logarithm function, represents the risk variance estimate of the i-th sample predicted by the model, represents the identity modulation weight of the i-th sample, which is obtained by learning the laboratory identity feature vector through a multi-layer perceptron;

[0103] Construct and train the model, specifically through the local risk perception layer, the remote dependency risk extraction layer, the experimenter identity sensitivity modulation layer, the identity dynamic weighted fusion layer and the risk prediction output layer, to construct a security anomaly prediction model, based on the enhanced training data set as training data, use the identity security risk prediction loss function to train the model as the optimization target, and obtain the trained security anomaly prediction model.

[0104] By performing the above operations, in order to solve the problem that the traditional laboratory safety anomaly prediction system does not take into account the identity differences of experimental personnel, resulting in the laboratory safety risk prediction results ignoring individual differences and lacking accuracy, this solution innovatively introduces laboratory identity feature vectors, introduces identity sensitivity modulation layer and identity dynamic weighted fusion layer into the prediction model, realizes differentiated dynamic adjustment of local behavior risk and remote dependency risk, accurately captures the interaction logic between identity, behavior and risk, enhances the model's perception of individual behavior risk, realizes personalized risk identification and accurate prediction at the experimental personnel level, and significantly improves the model's generalization ability and prediction accuracy in scenarios with multi-identity heterogeneous experimental personnel; for the existing laboratory safety anomaly prediction system There is a technical problem in the use of a unified error measurement standard in the loss function design of the prediction model, which results in the model being unable to effectively distinguish the risk sensitivity differences reflected by different experimental personnel identities during the training process, thereby affecting the accuracy of the prediction results. This scheme innovatively designs an identity security risk prediction loss function. By introducing identity modulation weights and combining it with the conditional risk variance estimation function based on identity and behavioral characteristics, it realizes personalized adjustment and dynamic weighting of the prediction error, significantly enhances the model's ability to identify anomalies of multi-identity heterogeneous individuals, and improves the loss function's modeling effect on risk uncertainty, ultimately improving the model's prediction accuracy in complex dynamic laboratory environments, effectively meeting the actual needs of personalized laboratory safety risk prediction.

[0105] Example 6, see Figure 1 This embodiment is based on the above embodiment. The laboratory safety anomaly prediction module specifically uses the real-time laboratory safety data in the laboratory safety optimization data as input data of the trained safety anomaly prediction model to obtain laboratory safety anomaly prediction results. Based on the laboratory safety anomaly prediction results, early perception of potential abnormal events in the laboratory is achieved.

[0106] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0107] While the embodiments of the present invention have been shown and described, it will be apparent to those skilled in the art that various changes, modifications, substitutions, and alterations can be made to the embodiments without departing from the principles and spirit of the invention.

[0108] The present invention and its embodiments are described above. This description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by this and, without departing from the purpose of the present invention, designs structures and embodiments similar to this technical solution without inventiveness, they shall fall within the scope of protection of the present invention.

Claims

1. A laboratory safety anomaly prediction system based on big data, characterized by: It includes raw data acquisition module, data optimization module, abnormal sample compensation module, safety anomaly prediction model construction module and laboratory safety anomaly prediction module; The raw data acquisition module specifically collects multi-source raw data from the laboratory operation process through the laboratory safety management platform to obtain laboratory safety raw data; The data optimization module is used to systematically process the laboratory safety raw data, specifically to perform data preprocessing on the laboratory safety raw data, generate laboratory identity feature vectors based on laboratory personnel information, and define sample labels based on historical safety results to obtain laboratory safety optimization data; The abnormal sample compensation module is used to solve the problem of the scarcity of laboratory safety abnormal samples. Specifically, it generates a set of identity abnormal feature vectors by constructing identity grouping abnormal samples, and constructs an identity grouping abnormal sample generation model by designing a generator, a discriminator that introduces spectral normalization, and an overall adversarial objective function that integrates the conditional Wasserstein distance of the laboratory identity feature vector and introduces a gradient penalty strategy. The model is then trained and the data is input into the trained identity grouping abnormal sample generation model to obtain an enhanced training data set. The security anomaly prediction model construction module specifically adopts a multi-layer structure model including a local risk perception layer, a remote dependency risk extraction layer, an experimenter identity sensitivity modulation layer, an identity dynamic weighted fusion layer and a risk prediction output layer to construct a security anomaly prediction model, and uses an enhanced training dataset as input and an identity-modulated risk prediction loss function to perform model training to obtain a trained security anomaly prediction model; the module includes the following steps: The local risk perception layer uses a multi-layer bidirectional long short-term memory network and performs a maximum pooling operation in the time dimension on the hidden state output of the last layer to obtain the local risk characteristics under the time window. ; The remote dependency risk extraction layer specifically adopts a multi-layer Transformer encoding network based on the self-attention mechanism. The Transformer encoding network includes a multi-head attention mechanism and a position feedforward network, and is equipped with residual connections and layer normalization processing; and the final output of the multi-layer Transformer is average pooled to obtain the remote risk features. ; The experimenter identity sensitivity modulation layer is specifically to enhance the representation ability of identity features by introducing a multi-head self-attention mechanism to obtain the identity enhancement vector , and input them into two multilayer perceptron network channels with the same structure but independent parameters, and nonlinear mapping generates local risk modulation factors and remote risk modulators The multi-layer perceptron network channel includes a shared identity feature extraction network and two sets of independent output branches; the risk modulation factor is used to dynamically reflect the sensitivity adjustment requirements of identity differences to different risk sources; Identity dynamic weighted fusion layer, specifically based on the local risk modulation factor and remote risk modulators Perform weighted summation on local risk features and remote risk features to obtain identity-aware risk features ; The risk prediction output layer is specifically to transform the identity perception risk characteristics Input the fully connected neural network layer and pass the Sigmoid activation function to obtain the laboratory safety anomaly prediction results. ; Design identity security risk prediction loss function, specifically combining identity modulation weights The risk sensitivity of the prediction errors of different roles in the laboratory operation process is modulated with the conditional risk variance, and the identity-differentiated loss function is constructed. The formula used is as follows: ; Where, represents the identity security risk prediction loss function value, n represents the number of training samples, represents the true risk value of the i-th sample, represents the risk prediction value of the i-th sample, represents the input time series behavior characteristics of the i-th sample, represents the laboratory identity feature vector corresponding to the i-th sample, represents the natural logarithm function, represents the risk variance estimate of the i-th sample predicted by the model, represents the identity modulation weight of the i-th sample, which is obtained by learning the laboratory identity feature vector through a multi-layer perceptron; Constructing and training a model, specifically, constructing a security anomaly prediction model through the local risk perception layer, the remote dependency risk extraction layer, the experimenter identity sensitivity modulation layer, the identity dynamic weighted fusion layer, and the risk prediction output layer, based on the enhanced training data set as training data, using the identity security risk prediction loss function to train the model as the optimization target, and obtaining a trained security anomaly prediction model; The laboratory safety anomaly prediction module specifically inputs real-time data into the trained safety anomaly prediction model to obtain laboratory safety risk prediction results.

2. The laboratory safety anomaly prediction system based on big data according to claim 1 is characterized by: The data optimization module specifically includes the following steps: Data preprocessing, specifically data cleaning, standardization conversion and time series construction of raw data; Laboratory identity feature vector construction is used to model the differences in different personnel identities in dynamic risk prediction. Specifically, the multi-dimensional identity fields in the laboratory personnel information data are encoded into a low-dimensional structured laboratory identity feature vector through the autoencoder algorithm; Data label definition, specifically assigning labels to each time segment sample based on historical laboratory safety results, marking them as normal laboratory safety or abnormal laboratory safety.

3. The laboratory safety anomaly prediction system based on big data according to claim 1, characterized in that: The abnormal sample compensation module specifically includes the following steps: The construction of identity-grouped abnormal samples involves classifying and extracting abnormal samples based on the identity type of the laboratory personnel based on labeled laboratory safety optimization data. For each type of abnormal sample, a sliding statistical analysis method is used to extract features from the time series data of the abnormal samples to generate a time series behavior feature vector. This feature vector is then jointly encoded with the corresponding laboratory identity feature vector to form an identity abnormality feature vector. Abnormal samples under all identity types are then merged to form a set of identity abnormality feature vectors. Construct an identity grouping abnormal sample generation model; Model training, specifically, takes the identity anomaly feature vector set as input data, trains the identity grouping anomaly sample generation model, and alternately minimizes the generator loss function and maximize the discriminator loss function , until convergence, and obtain the trained identity grouping abnormal sample generation model; An enhanced training dataset is obtained by inputting the identity anomaly feature vector set and the random noise vector into the trained identity grouping anomaly sample generation model, and using the generator to generate synthetic anomaly samples that meet various identity conditions; then, an anomaly scoring function is used to evaluate the quality of the generated samples, and low-quality samples that are significantly different from the real anomaly samples or do not conform to the behavioral patterns of the identity are eliminated; the retained high-quality synthetic anomaly samples are fused with the identity anomaly feature vector set and normal samples to construct an enhanced training dataset with identity differentiation and balanced anomaly distribution.

4. The laboratory safety anomaly prediction system based on big data according to claim 1, characterized in that: The construction of the identity grouping abnormal sample generation model specifically includes the following steps: Design a generator to generate synthetic anomaly samples that match identity characteristics from random noise. Specifically, a multi-layer fully connected neural network based on a residual structure is used, consisting of an input layer, three sequentially connected residual blocks, and an output layer. The data passes through the multi-layer fully connected neural network to obtain a synthetic anomaly sample vector. Design a discriminator, specifically a multi-layer perceptron structure with spectral normalization, consisting of an input layer, three sequentially connected hidden layers, and an output layer. The data is processed by the multi-layer perceptron structure to obtain a continuous real-valued score; The overall adversarial objective function is designed. Specifically, the overall adversarial objective function is constructed by combining the conditional Wasserstein distance, introducing the laboratory identity feature vector y, and using the gradient penalty mechanism. The formula used is as follows: ; Where, represents the overall adversarial objective function, represents the real sample, represents the noise distribution, Indicates identity information The synthetic samples generated by the generator below, represents the authenticity score given by the discriminator to the sample x, represents the gradient penalty coefficient, Indicates sampling from the true abnormal sample distribution, represents sampling from a noise distribution, Indicates sampling from the intermediate samples obtained by linear interpolation between the real samples and the generated samples; represents the gradient norm at the middle sample, represents the objective maximization of the discriminator, The goal of the generator is to minimize Represents the authenticity score of the generated sample by the discriminator, given Under the conditions; represents the laboratory identity feature vector, z represents the random noise vector, Represents the real sample vector, which is divided into synthetic abnormal sample vector and identity abnormal feature vector, represents the intermediate sample vector, represents the mathematical expectation function; Define the loss function, specifically the discriminator loss function Designed as the opposite of the overall adversarial objective function, the generator loss function is designed to be .

5. The laboratory safety anomaly prediction system based on big data according to claim 1 is characterized in that: The laboratory safety anomaly prediction module specifically uses the real-time laboratory safety data in the laboratory safety optimization data as input data of the trained safety anomaly prediction model to obtain laboratory safety anomaly prediction results, and based on the laboratory safety anomaly prediction results, realizes early perception of potential abnormal events in the laboratory.

6. The laboratory safety anomaly prediction system based on big data according to claim 1 is characterized by: The original data acquisition module specifically obtains laboratory safety original data by performing data acquisition; the laboratory safety original data includes historical laboratory safety data and real-time laboratory safety data; the historical laboratory safety data and real-time laboratory safety data both include laboratory personnel information data, laboratory equipment operation behavior data, experimental operation record data, laboratory equipment data and laboratory environment data; the historical laboratory safety data also includes historical laboratory safety results; the historical laboratory safety results include normal laboratory safety and abnormal laboratory safety.

Citation Information

Patent Citations

  • Method for enhancing security of picture verification code system

    CN114218549A

  • Customs laboratory full-process safety management intelligent decision-making system and method

    CN120069538A