A safety work monitoring system and method for a workman

By acquiring and analyzing facial videos and historical data of workers to be monitored, combined with physiological data analysis and safety risk assessment models, the problems of cumbersome regular physical examinations and untimely monitoring have been solved, realizing convenient safety operation monitoring and reducing operational risks.

CN122155394APending Publication Date: 2026-06-05ELECTRIC POWER RES INST OF GUANGDONG POWER GRID CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ELECTRIC POWER RES INST OF GUANGDONG POWER GRID CO LTD
Filing Date
2026-02-25
Publication Date
2026-06-05

Smart Images

  • Figure CN122155394A_ABST
    Figure CN122155394A_ABST
Patent Text Reader

Abstract

The application discloses a kind of safety operation monitoring system and method of operating worker, belong to safety operation monitoring technical field, above-mentioned system obtains the original face video of worker to be monitored before operation, historical medical data, operation area position and identity ID by data acquisition module;Through face video analysis module, the physiological condition data of worker is identified;Finally, through safety operation monitoring module, based on physiological condition data, historical medical data, operation area position and identity ID and preset safety risk assessment model, obtain the safety operation monitoring result of worker to be monitored. Through implementation of the present application, it can solve the problem that the safety operation monitoring is carried out by regular physical examination to worker in prior art, and the monitoring method is complicated, and the operation risk is increased due to the problem that monitoring is not timely.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of safe operation monitoring technology, and in particular to a safe operation monitoring system and method for workers. Background Technology

[0002] In the outdoor working environment of power grid workers, they often face multiple complex risks such as high altitude, high voltage, strong electromagnetic environment and extreme weather. Furthermore, modern medical research shows that high voltage electric field environment may cause autonomic nervous system dysfunction such as heart rate variability and blood pressure fluctuation. Long-term work pressure can easily lead to psychological changes such as anxiety and decreased decision-making ability. Therefore, the development of safety operation monitoring technology is crucial to ensuring worker safety.

[0003] Current monitoring methods mainly rely on regular medical examinations for workers. This method is cumbersome, and if there is a long gap between the medical examination date and the formal work date, the medical examination results are difficult to reflect the actual condition of the workers in real time before the work begins. This can lead to increased operational risks due to untimely monitoring. Summary of the Invention

[0004] This invention provides a safety monitoring system and method for workers, which can solve the problems of existing technologies that rely on regular medical examinations to monitor workers' safety, which are cumbersome and increase operational risks due to untimely monitoring.

[0005] One embodiment of the present invention provides a safety monitoring system for workers, comprising: Data acquisition module, facial video analysis module, and safe operation monitoring module; The aforementioned data acquisition module is used to acquire the worker's original facial video, historical physical examination data, work area location, and identity ID before the operation. The aforementioned facial video analysis module is used to input the aforementioned raw facial video into a preset physiological data analysis model to obtain the physiological condition data of the worker to be monitored; wherein, the aforementioned raw facial video is input into the input layer built into the preset physiological data analysis model to enhance the image quality of the aforementioned raw facial video to obtain a first facial video; Based on the built-in spatiotemporal feature extraction backbone network and encoder of the preset physiological data analysis model, feature extraction is performed on the first facial video to obtain the temporal feature matrix and mid-level spatiotemporal features. The aforementioned temporal feature matrix and intermediate spatiotemporal features were subjected to feature transformation and aggregation processes to obtain the physiological status data of the workers to be monitored. The aforementioned safe operation monitoring module is used to input the aforementioned physiological condition data, historical physical examination data, work area location, and identity ID into a preset safety risk assessment model, so that the preset safety risk assessment model can perform work risk assessment and safety qualification assessment on the aforementioned worker to be monitored, and obtain the safe operation monitoring results of the aforementioned worker to be monitored.

[0006] Furthermore, the aforementioned original facial video is input into the input layer built into the preset physiological data analysis model to enhance the image quality of the original facial video, resulting in a first facial video, including: Based on each video frame of the original facial video, the facial key points of each video frame are detected. Based on the above facial key points, affine transformation alignment is performed to obtain several facial images of a preset standard size; Each face image is deblurred to obtain several deblurred face images; For each face image, extract the luminance channel image and color channel image of the current face image, and divide the current luminance channel image into several image blocks; Calculate the histogram of each image block, and crop and equalize the histogram according to a preset contrast limit to obtain the processed histogram. The processed histograms are stitched together to obtain the processed luminance channel image; The processed luminance channel image and the corresponding color channel image are merged to obtain an enhanced face image, which in turn yields the aforementioned first facial video.

[0007] Furthermore, the aforementioned first facial video is subjected to feature extraction based on the spatiotemporal feature extraction backbone network and encoder built into the preset physiological data analysis model, resulting in a temporal feature matrix and intermediate spatiotemporal features, including: The aforementioned spatiotemporal feature extraction backbone network performs hierarchical convolution and feature transformation processing on the video frame sequence of the aforementioned first facial video to obtain local spatiotemporal features of facial motion features and mid-level spatiotemporal features of facial deformation features; wherein, the aforementioned local spatiotemporal features are obtained from the output of the last layer of the aforementioned spatiotemporal feature extraction backbone network, and the aforementioned mid-level spatiotemporal features are obtained from the output of the middle layer of the aforementioned spatiotemporal feature extraction backbone network. The aforementioned local spatiotemporal features are input into the built-in encoder to add time position encoding to the spatial features in the aforementioned local spatiotemporal features, thereby obtaining the aforementioned temporal feature matrix.

[0008] Furthermore, the aforementioned local spatiotemporal features are input into the built-in encoder to add temporal location encoding to the spatial features within the local spatiotemporal features, resulting in the aforementioned temporal feature matrix, including: The encoder described above flattens out the spatial features in the aforementioned local spatiotemporal features to obtain a spatial vector; After adding temporal position encoding to the above spatial vector, a long-term temporal dependency is modeled on the spatial vector with added temporal position encoding through a self-attention mechanism, resulting in the above temporal feature matrix used to characterize the global temporal context.

[0009] Furthermore, the aforementioned physiological data include: heart rate, blood oxygen, blood pressure, body temperature, and mood; The aforementioned time-series feature matrix and intermediate-level spatiotemporal features are subjected to feature transformation and aggregation processes to obtain the physiological status data of the workers to be monitored, including: Based on the attention mechanism, feature sequences for characterizing skin color changes are extracted from the above temporal feature matrix; The above feature sequences are subjected to causal convolution to generate PPG waveform signals; The PPG waveform signal is input into the built-in LSTM network so that the LSTM network can perform periodic analysis on the PPG waveform signal to obtain the heart rate and blood oxygen. The global feature vector is obtained by performing average pooling on the above temporal feature matrix along the time dimension. Obtain the vascular morphology features output by the preset shallow layer of the backbone network extracted from the above spatiotemporal features; The above global feature vector and blood vessel morphology features are concatenated and nonlinearly transformed to obtain a fused feature vector. Regression operation is performed on the above fused feature vectors to obtain the above blood pressure and the above body temperature; Spatial attention focusing processing is applied to the above mid-level spatiotemporal features to obtain several weighted feature maps; Global average pooling is performed on all weighted feature maps in the spatial dimension to obtain a micro-expression feature sequence for representing facial micro-expressions; The above micro-expression feature sequence is input into the built-in GRU network so that the GRU network can capture the facial emotion transformation process in the above micro-expression feature sequence and obtain the micro-expression global feature vector. The global feature vectors of the micro-expressions are sequentially subjected to fully connected operations and normalization to obtain the probability distributions corresponding to each emotion type, and the emotion type with the highest probability is taken as the emotion.

[0010] Furthermore, the aforementioned physiological condition data, historical medical examination data, work area location, and identity ID are input into a preset safety risk assessment model, so that the preset safety risk assessment model can perform work risk assessment and safety qualification assessment on the workers to be monitored, and obtain the work safety monitoring results of the workers to be monitored, including: The aforementioned physiological condition data, historical medical examination data, work area location, and identity ID are input into a preset safety risk assessment model. This model retrieves the risk information corresponding to the work area location from a preset area risk database based on the work area location. The model then obtains a work risk assessment result based on the risk information, physiological condition data, and historical medical examination data. This work risk assessment result indicates whether the worker being monitored is suitable for the job. The aforementioned preset safety risk assessment model retrieves the validity period of the safety qualifications of the workers to be monitored from the preset database based on the aforementioned identity ID, and compares the validity period of the aforementioned safety qualifications with the preset safety validity period to obtain the safety qualification assessment result of the workers to be monitored. If the above-mentioned work risk assessment result is unqualified or the safety qualification assessment result is unqualified, the safety operation monitoring result of the above-mentioned worker to be monitored shall be deemed unqualified; otherwise, the safety operation monitoring result of the above-mentioned worker to be monitored shall be deemed qualified.

[0011] Furthermore, the training of the aforementioned pre-defined physiological data analysis model includes: Acquire several facial video samples with a first real label; wherein, the first real label is used to represent the real physiological condition information of the worker corresponding to the facial video sample; the real physiological condition information includes: real heart rate, real blood oxygen, real blood pressure, real body temperature and real emotion; The above-mentioned face video samples are input into the physiological data analysis model to be trained for iterative training until the first loss function converges, and the above-mentioned preset physiological data analysis model is obtained. In each iteration of training, the current predicted physiological condition data is generated based on the current face video samples; the current first loss function is calculated based on the current predicted physiological condition data and the corresponding first true label, and it is determined whether the current first loss function has converged; if the current first loss function has converged, the current physiological data analysis model is used as the above-mentioned preset physiological data analysis model; otherwise, the model parameters in the current physiological data analysis model are adjusted and training continues.

[0012] Furthermore, the training of the aforementioned pre-set security risk assessment model includes: Acquire several safety risk assessment sample data with a second authenticity label; wherein, the safety risk assessment sample includes: historical physiological condition data of several workers, historical physical examination data corresponding to the historical physiological condition data, historical work area location and identity ID; the second authenticity label is used to represent the real work risk assessment result and the real safety qualification assessment result of the above safety risk assessment sample data; The above security risk assessment sample data is input into the security risk assessment model to be trained for iterative training until the second loss function converges, thereby generating the above security risk assessment model. In each iteration, the current predicted safety operation monitoring result is generated based on the current safety risk assessment sample data; the current second loss function is calculated based on the current predicted safety operation monitoring result and the corresponding second true label, and it is determined whether the current second loss function has converged; if the current second loss function has converged, the current safety risk assessment model is used as the above-mentioned preset safety risk assessment model; otherwise, the model parameters in the current safety risk assessment model are adjusted and training continues.

[0013] Furthermore, it also includes: Physiological condition monitoring module; The aforementioned physiological condition monitoring module is used to acquire facial video data of the worker to be monitored at preset monitoring intervals after the worker to be monitored starts working. The above-mentioned facial video data is input into the above-mentioned preset physiological data analysis model to obtain the working heart rate, working blood oxygen, working blood pressure, working body temperature and working mood of the worker to be monitored during work. If any of the above-mentioned work heart rate, work blood oxygen, work blood pressure and work body temperature do not meet the preset safe work threshold, or if the work mood does not meet the preset safe work mood category, a safety risk warning will be issued.

[0014] Based on the above-described apparatus embodiments, the present invention provides corresponding method embodiments; This invention provides a method for monitoring the safe operation of workers, applicable to the facial video analysis module in the above-described worker safety monitoring system; The safe operation monitoring method includes: The original facial video of the worker to be monitored before the operation is input into a preset physiological data analysis model, so that the input layer built into the preset physiological data analysis model can enhance the image quality of the original facial video to obtain a first facial video; wherein, the original facial video is obtained by the data acquisition module; Based on the built-in spatiotemporal feature extraction backbone network and encoder of the preset physiological data analysis model, feature extraction is performed on the first facial video to obtain the temporal feature matrix and mid-level spatiotemporal features; The time-series feature matrix and the intermediate-level spatiotemporal features are subjected to feature transformation and aggregation processes respectively to obtain the physiological status data of the worker to be monitored; The physiological condition data is sent to the safe operation monitoring module, which then inputs the physiological condition data, historical medical examination data, work area location, and identity ID into a preset safety risk assessment model. This allows the preset safety risk assessment model to perform an operation risk assessment and safety qualification assessment on the worker to be monitored, thereby obtaining the safe operation monitoring results for the worker. The historical medical examination data, work area location, and identity ID are obtained by the data acquisition module.

[0015] The embodiments of the present invention have the following beneficial effects: This invention provides a safety monitoring system and method for workers, the system comprising: a data acquisition module, a facial video analysis module, and a safety monitoring module; the data acquisition module is used to acquire the worker's original facial video, historical physical examination data, work area location, and identity ID before work; the facial video analysis module is used to input the original facial video into a preset physiological data analysis model to obtain the worker's physiological condition data; wherein, the original facial video is input into an input layer built into the preset physiological data analysis model to enhance the image quality of the original facial video, obtaining a first facial video; based on... The first facial video is used to extract features from the pre-defined physiological data analysis model using its built-in spatiotemporal feature extraction backbone network and encoder, resulting in a temporal feature matrix and intermediate spatiotemporal features. These features are then transformed and aggregated to obtain the physiological condition data of the worker to be monitored. The safety operation monitoring module inputs the physiological condition data, historical medical examination data, work area location, and identity ID into a pre-defined safety risk assessment model. This model then performs a work risk assessment and safety qualification assessment on the worker to be monitored, yielding the safety operation monitoring results. Therefore, in this invention, physiological condition analysis using recorded raw facial videos achieves convenient, contactless safety operation monitoring. Since the raw facial video is acquired before work begins, the results more closely reflect the worker's current health condition, significantly reducing operational risks. Attached Figure Description

[0016] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of the structure of a safety monitoring system for workers provided in an embodiment of the present invention.

[0018] Figure 2 This is a schematic diagram illustrating the principle of imaging photoplethysmography provided in an embodiment of the present invention.

[0019] Figure 3 This is a flowchart illustrating a method for monitoring the safe operation of workers according to an embodiment of the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.

[0022] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.

[0023] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0024] In the description of the embodiments in this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.

[0025] In the description of the embodiments of this application, the term "multiple" refers to two or more (including two), similarly, "multiple sets" refers to two or more (including two sets), and "multiple pieces" refers to two or more (including two pieces).

[0026] In the description of the embodiments of this application, unless otherwise expressly specified and limited, technical terms such as "installation," "connection," "joining," and "fixing" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. For those skilled in the art, the specific meaning of the above terms in the embodiments of this application can be understood according to the specific circumstances.

[0027] See Figure 1 To address the problem that existing technologies, which rely on regular medical examinations to monitor worker safety, may lead to increased operational risks due to untimely monitoring, an embodiment of the present invention provides a worker safety monitoring system, comprising: Data acquisition module, facial video analysis module, and safe operation monitoring module; The aforementioned data acquisition module is used to acquire the worker's original facial video, historical physical examination data, work area location, and identity ID before the operation. Specifically, the aforementioned raw facial video is recorded by the worker to be monitored before starting work, and the video can be thirty seconds long. Therefore, before starting work, the worker to be monitored records a raw facial video of a preset duration, and at the same time, the worker's historical physical examination data, the location of the work area to be carried out, and the worker's unique identity ID are obtained.

[0028] Preferably, the safety monitoring system for workers of the present invention is also applicable to monitoring operations in outdoor high-temperature environments.

[0029] Preferably, a wearable smart badge can be designed that takes into account the adaptability to high-temperature outdoor and electric field environments, and can ensure stable operation under various working conditions. The badge uses Beidou + GPS + WiFi + LBS multi-positioning technology to collect the location of the working area and transmit it to the system of this invention for corresponding analysis and monitoring.

[0030] The aforementioned facial video analysis module is used to input the aforementioned raw facial video into a preset physiological data analysis model to obtain the physiological condition data of the worker to be monitored; wherein, the aforementioned raw facial video is input into the input layer built into the preset physiological data analysis model to enhance the image quality of the aforementioned raw facial video to obtain a first facial video; Specifically, the aforementioned pre-defined physiological data analysis model adopts a hybrid neural network architecture with multi-task learning, combining temporal analysis and spatial feature extraction. The core modules include: a video preprocessing module, a spatiotemporal feature extraction backbone network (including 3D CNN and Transformer), a multi-branch physiological parameter prediction head, and a sentiment analysis branch.

[0031] Preferably, by using a pre-set physiological data analysis model, workers' physiological data can be collected without contact before work begins, providing them with comprehensive health analysis and risk monitoring. This method reduces interference with workers' work while improving the convenience and real-time nature of data collection.

[0032] In a preferred embodiment, the above-mentioned input of the original facial video to the input layer built into the preset physiological data analysis model, and the image quality enhancement of the original facial video to obtain a first facial video, includes: Based on each video frame of the original facial video, the facial key points of each video frame are detected. Specifically, the aforementioned facial key points are location points used to represent facial features (such as contours and facial features). In this invention, they can be the locations of the facial features within a video frame. Illustratively, facial key points can be detected using the model's built-in MTCNN (Multi-Task Convolutional Neural Network).

[0033] Based on the above facial key points, affine transformation alignment is performed to obtain several facial images of a preset standard size; Specifically, in order to eliminate geometric differences in the video caused by factors such as head rotation, tilt, and distance, and to provide stable input for subsequent accurate feature extraction, the detected face regions with key points need to be standardized into images of a fixed size (e.g., 112×112) with a frontal facial pose through geometric transformation.

[0034] Specifically, in this invention, affine transformation is used to preprocess the faces in the video frame. This process is existing technology and will not be described in detail here.

[0035] Each face image is deblurred to obtain several deblurred face images; Specifically, in order to eliminate image blur caused by camera shake and slight worker movement, and to correct uneven lighting (such as shadows and local overexposure), and to lay a high-quality image foundation for subsequent accurate extraction of skin color changes and subtle features such as micro-expressions, it is necessary to perform deblurring and lighting equalization processing on facial images.

[0036] Preferably, in this invention, the built-in and trained DeblurGAN-v2 model is used to perform deblurring. The DeblurGAN-v2 model is a generative adversarial network. Its generator is a U-Net-style encoder-decoder structure containing multiple downsampling and upsampling layers with skip connections in between. The network analyzes the blurring pattern of the image layer by layer (usually motion blur) and reconstructs the lost high-frequency details (such as clear skin texture, eyebrows and hair, pupil edges) in the feature space. Finally, the output is consistent with the input and is a number of deblurred face images.

[0037] For each face image, extract the luminance channel image and color channel image of the current face image, and divide the current luminance channel image into several image blocks; Calculate the histogram of each image block, and crop and equalize the histogram according to a preset contrast limit to obtain the processed histogram. The processed histograms are stitched together to obtain the processed luminance channel image; The processed luminance channel image and the corresponding color channel image are merged to obtain an enhanced face image, which in turn yields the aforementioned first facial video.

[0038] Specifically, in this invention, the CLAHE algorithm is used to achieve balanced illumination. Through the above process, the algorithm outputs an image with uniform illumination, where shadows on the face are faded, overexposed areas are suppressed, and the overall contrast is moderate, resulting in a more consistent appearance of the skin area.

[0039] In this preferred embodiment, image quality enhancement processing of the original facial video is achieved through the input layer.

[0040] Based on the built-in spatiotemporal feature extraction backbone network and encoder of the preset physiological data analysis model, feature extraction is performed on the first facial video to obtain the temporal feature matrix and mid-level spatiotemporal features. Specifically, the aforementioned spatiotemporal feature extraction backbone network mainly consists of four 3D ResNet blocks (3×3×3 convolutional kernels, 1×2×2 stride), capable of capturing subtle movements such as facial micro-expressions and vascular pulsations in the input first facial video. The encoder is a temporal Transformer encoder with six Transformer Encoder layers (8 heads, 256 hidden layers), which transforms the features extracted by the spatiotemporal feature extraction backbone network into temporal features capable of understanding the global context and long-range dependencies of the entire video sequence.

[0041] In a preferred embodiment, the above-mentioned feature extraction of the first facial video based on the spatiotemporal feature extraction backbone network and encoder built into the preset physiological data analysis model to obtain a temporal feature matrix and intermediate spatiotemporal features includes: The aforementioned spatiotemporal feature extraction backbone network performs hierarchical convolution and feature transformation processing on the video frame sequence of the aforementioned first facial video to obtain local spatiotemporal features of facial motion features and mid-level spatiotemporal features of facial deformation features; wherein, the aforementioned local spatiotemporal features are obtained from the output of the last layer of the aforementioned spatiotemporal feature extraction backbone network, and the aforementioned mid-level spatiotemporal features are obtained from the output of the middle layer of the aforementioned spatiotemporal feature extraction backbone network. Specifically, the aforementioned spatiotemporal feature extraction backbone network is composed of four stacked 3D ResNets. Each layer performs convolution processing on the input. The aforementioned local spatiotemporal features are the final output of the entire spatiotemporal feature extraction backbone network, while the aforementioned intermediate spatiotemporal features are the output of the intermediate layer of the aforementioned spatiotemporal feature extraction backbone network. In this invention, they can be the output of the second layer.

[0042] Specifically, for the aforementioned spatiotemporal feature extraction backbone network, in this invention, the size of the video frame sequence of the first facial video input can be 900 (time steps) × 112 (height) × 112 (width) × 3 (channels).

[0043] The aforementioned local spatiotemporal features are input into the built-in encoder to add time position encoding to the spatial features in the aforementioned local spatiotemporal features, thereby obtaining the aforementioned temporal feature matrix.

[0044] Specifically, the encoder flattens the spatial features in the local spatiotemporal features into a 3136-dimensional vector, adds learnable temporal location encoding, and then models long temporal dependencies (such as heart rate cycles) through a self-attention mechanism, finally outputting a 900×256 temporal feature matrix.

[0045] In this preferred embodiment, a temporal feature matrix and intermediate spatiotemporal features are extracted from the first facial video through a spatiotemporal feature extraction backbone network and an encoder.

[0046] In another preferred embodiment, the aforementioned local spatiotemporal features are input to a built-in encoder to add temporal location encoding to the spatial features within the local spatiotemporal features, resulting in the aforementioned temporal feature matrix, including: The encoder described above flattens out the spatial features in the aforementioned local spatiotemporal features to obtain a spatial vector; After adding temporal position encoding to the above spatial vector, a long-term temporal dependency is modeled on the spatial vector with added temporal position encoding through a self-attention mechanism, resulting in the above temporal feature matrix used to characterize the global temporal context.

[0047] Specifically, since the Transformer itself does not have the ability to process sequence order, it is necessary to explicitly inject positional information, i.e., the aforementioned temporal positional encoding, so that the model can clearly know the position of each feature vector on the time axis of the video frame sequence. Subsequently, these spatial vectors are weighted and concatenated through a multi-head self-attention layer in the encoder to obtain an abstract understanding of the entire first facial video input, and thus obtain the aforementioned temporal feature matrix.

[0048] In this preferred embodiment, the temporal feature matrix of the first facial video is extracted by an encoder.

[0049] The aforementioned temporal feature matrix and intermediate spatiotemporal features were subjected to feature transformation and aggregation processes to obtain the physiological status data of the workers to be monitored. Specifically, the physiological parameter prediction head obtains heart rate, blood oxygen, blood pressure, and body temperature data, while the emotion analysis branch detects emotions. Heart rate, blood oxygen, blood pressure, body temperature, and emotions together constitute the aforementioned physiological condition data.

[0050] In a preferred embodiment, the aforementioned physiological data includes: heart rate, blood oxygen, blood pressure, body temperature, and mood; The aforementioned time-series feature matrix and intermediate-level spatiotemporal features are subjected to feature transformation and aggregation processes to obtain the physiological status data of the workers to be monitored, including: Based on the attention mechanism, feature sequences for characterizing skin color changes are extracted from the above temporal feature matrix; Specifically, in the physiological parameter prediction head (composed of a 1D convolution, LSTM network, and PPG signal decoder) used for heart rate and blood oxygenation detection, an attention mechanism is used to filter out the subset of features most relevant to the minute color fluctuations on the skin surface caused by changes in blood volume from the entire temporal feature matrix, resulting in the aforementioned feature sequence. A schematic diagram illustrating the principle of imaging-based photoplethysmography is shown below. Figure 2 As shown, from Figure 2As can be seen, when light emitted from a light source shines on the skin surface, it is divided into two parts: the first part is the direct reflection (spectral reflection) of the light in the outermost layer of the skin (epidermis). This reflection does not penetrate the skin and only reflects the surface luster and texture, without carrying subcutaneous physiological information. The second part is the diffuse reflection of light that penetrates the epidermis, enters the dermis and subcutaneous tissue, and is reflected again after multiple scatterings. This part of the light interacts with substances such as hemoglobin in the blood, carrying key physiological information such as the blood volume, blood oxygen concentration, and periodically changing heart rate of the larger blood vessels and capillaries in the subcutaneous tissue (hypodermis). This is the source of the PPG waveform signal. Therefore, the diffuse and specular reflected light is received by the camera sensor to obtain the raw facial video containing the heart rate and blood oxygen information of the worker being monitored.

[0051] The above feature sequences are subjected to causal convolution to generate PPG waveform signals; Specifically, the aforementioned PPG waveform signal is the waveform signal of the photoplethysmography (PPG). Causal convolution is performed using 1D convolution to effectively compress the feature sequence and convert it into a scalar value, thus obtaining the aforementioned PPG waveform signal.

[0052] The PPG waveform signal is input into the built-in LSTM network so that the LSTM network can perform periodic analysis on the PPG waveform signal to obtain the heart rate and blood oxygen. Specifically, in order to capture the periodicity of the PPG waveform signal, it is input into an LSTM network for periodic analysis to learn the interpeak period between adjacent pulse peaks in the signal, obtain a context vector of fixed length, and finally regress and output heart rate and blood oxygen.

[0053] The global feature vector is obtained by performing average pooling on the above temporal feature matrix along the time dimension. Specifically, in the physiological parameter prediction head (composed of a fully connected layer and Gaussian process regression) used for blood pressure and body temperature detection, the global feature vector is first extracted through time-averaged pooling.

[0054] Obtain the vascular morphology features output by the preset shallow layer of the backbone network extracted from the above spatiotemporal features; Specifically, since the physical properties of blood vessels (such as hardening and elasticity) directly affect blood pressure, and body temperature also affects the vasodilation state of blood vessels on the skin surface, this information may have been partially lost in the highly abstract deep local spatiotemporal features. It is more effective to extract it directly from the shallow visual features. Therefore, the above-mentioned vascular morphology features can be obtained from the output of the first or second ResNet block.

[0055] The above global feature vector and blood vessel morphology features are concatenated and nonlinearly transformed to obtain a fused feature vector. Specifically, by performing feature concatenation and nonlinear transformation through a fully connected layer, a high-level feature vector that has undergone deep fusion and abstraction is obtained, namely the aforementioned fused feature vector.

[0056] Regression operation is performed on the above fused feature vectors to obtain the above blood pressure and the above body temperature; Specifically, the blood pressure and body temperature values ​​are obtained through Gaussian process regression prediction. This process is existing technology and will not be described in detail here.

[0057] Spatial attention focusing processing is applied to the above mid-level spatiotemporal features to obtain several weighted feature maps; Specifically, in the sentiment analysis branch (composed of a CNN and GRU hybrid), spatial attention modulation is used to enhance the features of the facial regions most relevant to emotional expression (such as the area between the eyebrows, the corners of the eyes, and the corners of the mouth) on the mid-level spatiotemporal features obtained by the CNN, while suppressing irrelevant regions (such as the background and hair), resulting in several weighted feature maps.

[0058] Global average pooling is performed on all weighted feature maps in the spatial dimension to obtain a micro-expression feature sequence for representing facial micro-expressions; Specifically, global average pooling is used to obtain the global statistical features of facial micro-expressions in the entire original facial video, namely the micro-expression feature sequence mentioned above.

[0059] The above micro-expression feature sequence is input into the built-in GRU network so that the GRU network can capture the facial emotion transformation process in the above micro-expression feature sequence and obtain the micro-expression global feature vector. Specifically, by using update and reset gates in a multi-layer GRU network, long-term information is selectively memorized and irrelevant information is forgotten, allowing for the learning of more complex emotional dynamics and the acquisition of global feature vectors for micro-expressions.

[0060] The global feature vectors of the micro-expressions are sequentially subjected to fully connected operations and normalization to obtain the probability distributions corresponding to each emotion type, and the emotion type with the highest probability is taken as the emotion.

[0061] Specifically, to map the global feature vector of micro-expressions to specific emotion categories, a fully connected operation and normalization process are used to output the probability distribution of various emotion types. These emotion types can include: anger, fear, happiness, sadness, and surprise, among others. The emotion type with the highest probability value is then used as the final emotion recognition result.

[0062] In this preferred embodiment, the heart rate, blood oxygen, blood pressure, body temperature and mood of the worker to be monitored are identified by performing feature transformation and aggregation on the temporal feature matrix and the intermediate spatiotemporal features respectively.

[0063] In another preferred embodiment, the training of the aforementioned preset physiological data analysis model includes: Acquire several facial video samples with a first real label; wherein, the first real label is used to represent the real physiological condition information of the worker corresponding to the facial video sample; the real physiological condition information includes: real heart rate, real blood oxygen, real blood pressure, real body temperature and real emotion; The above-mentioned face video samples are input into the physiological data analysis model to be trained for iterative training until the first loss function converges, and the above-mentioned preset physiological data analysis model is obtained. In each iteration of training, the current predicted physiological condition data is generated based on the current face video samples; the current first loss function is calculated based on the current predicted physiological condition data and the corresponding first true label, and it is determined whether the current first loss function has converged; if the current first loss function has converged, the current physiological data analysis model is used as the above-mentioned preset physiological data analysis model; otherwise, the model parameters in the current physiological data analysis model are adjusted and training continues.

[0064] Specifically, the aforementioned facial video samples were recorded when the workers' heart rate, blood oxygen, blood pressure, body temperature, and emotions were known at the time. Therefore, they can be used as training samples for the model to be trained iteratively.

[0065] Specifically, the formula for calculating the first loss function mentioned above is as follows: In the formula, Denotes the first loss function. The weight represents the physiological condition data i (excluding emotions), and its value ranges from [0.15, 0.95]. Represents the MSE loss function. This represents the actual physiological condition information corresponding to physiological condition data i (excluding emotions). This represents the predicted physiological condition data corresponding to physiological condition data i (excluding emotions). Indicates the weight of the sentiment item. Represents the standard cross-entropy loss function. This represents the true emotions expressed in physiological data. This refers to the predicted emotion in the data predicting physiological conditions.

[0066] In this preferred embodiment, the physiological data analysis model was iteratively trained using face video samples with first real labels to obtain a preset physiological data analysis model.

[0067] The aforementioned safe operation monitoring module is used to input the aforementioned physiological condition data, historical physical examination data, work area location, and identity ID into a preset safety risk assessment model, so that the preset safety risk assessment model can perform work risk assessment and safety qualification assessment on the aforementioned worker to be monitored, and obtain the safe operation monitoring results of the aforementioned worker to be monitored.

[0068] Specifically, the heart rate, blood oxygen, blood pressure, body temperature and mood output by the preset physiological data analysis model are used as physiological condition data. Combined with the worker's historical physical examination data, work area location and identity ID, the preset safety risk assessment model is used to determine whether the worker can carry out the work safety monitoring results.

[0069] Preferably, the present invention combines a preset physiological data analysis model and a preset safety risk assessment model, which realizes the collection of physiological data and risk assessment and screening of power grid workers before work. Compared with traditional health detection methods, it is more timely and accurate.

[0070] Preferably, this system of the present invention can be integrated into a smart badge that uses multiple positioning technologies including BeiDou, GPS, LBS, and Wi-Fi. This badge is worn by power grid workers to enable the system to record and transmit the wearer's location information with high accuracy in outdoor high-temperature or electromagnetic interference scenarios, preventing unauthorized personnel from entering the work area. Furthermore, the badge can be equipped with 4G CAT1 full network connectivity and voice broadcasting functions, enabling real-time communication for rapid decision-making and dispatching in emergency situations, maximizing processing capacity.

[0071] In a preferred embodiment, the aforementioned physiological condition data, historical medical examination data, work area location, and identity ID are input into a preset safety risk assessment model, so that the preset safety risk assessment model performs a work risk assessment and safety qualification assessment on the worker to be monitored, and obtains the work safety monitoring results of the worker to be monitored, including: The aforementioned physiological condition data, historical medical examination data, work area location, and identity ID are input into a preset safety risk assessment model. This model retrieves the risk information corresponding to the work area location from a preset area risk database based on the work area location. The model then obtains a work risk assessment result based on the risk information, physiological condition data, and historical medical examination data. This work risk assessment result indicates whether the worker being monitored is suitable for the job. Specifically, the aforementioned pre-defined risk database stores risk information for different work areas, such as whether the area is suitable for people with hypertension, acrophobia, or vertigo; whether the area is a high-temperature outdoor work environment; or whether the area is an electric field work environment. This risk information is then compared and analyzed with physiological condition data and historical medical examination data to obtain the work risk assessment result. Similarly, if any risk information is found to pose a risk to the worker's current physiological condition data and historical medical examination data during the comparison, the work risk assessment result is deemed unqualified.

[0072] Specifically, the preset safety risk assessment model can determine whether a worker is suitable to carry out the work under the current condition by conducting a work risk assessment on the worker to be monitored before the operation.

[0073] The aforementioned preset safety risk assessment model retrieves the validity period of the safety qualifications of the workers to be monitored from the preset database based on the aforementioned identity ID, and compares the validity period of the aforementioned safety qualifications with the preset safety validity period to obtain the safety qualification assessment result of the workers to be monitored. Specifically, the preset safety risk assessment model determines the safety qualification assessment result by judging whether the safety qualification of the worker to be monitored has expired. If it has expired, the safety qualification assessment result is unqualified, otherwise it is qualified.

[0074] If the above-mentioned work risk assessment result is unqualified or the safety qualification assessment result is unqualified, the safety operation monitoring result of the above-mentioned worker to be monitored shall be deemed unqualified; otherwise, the safety operation monitoring result of the above-mentioned worker to be monitored shall be deemed qualified.

[0075] Specifically, the preset safety risk assessment model will obtain a predicted risk value based on the input data to quantify the safety operation monitoring results of the worker to be monitored. At the same time, a risk threshold is set. If the predicted risk value exceeds the risk threshold, it means that either the operation risk assessment result or the safety qualification assessment result is unqualified, and the worker is not allowed to carry out the operation. Otherwise, the worker is allowed to carry out the operation.

[0076] In this preferred embodiment, a preset safety risk assessment model is used to conduct job risk assessment and safety qualification assessment of the workers to be monitored, thereby obtaining the safety monitoring results of the workers to be monitored.

[0077] In another preferred embodiment, the training of the aforementioned preset security risk assessment model includes: Acquire several safety risk assessment sample data with a second authenticity label; wherein, the safety risk assessment sample includes: historical physiological condition data of several workers, historical physical examination data corresponding to the historical physiological condition data, historical work area location and identity ID; the second authenticity label is used to represent the real work risk assessment result and the real safety qualification assessment result of the above safety risk assessment sample data; The above security risk assessment sample data is input into the security risk assessment model to be trained for iterative training until the second loss function converges, thereby generating the above security risk assessment model. In each iteration, the current predicted safety operation monitoring result is generated based on the current safety risk assessment sample data; the current second loss function is calculated based on the current predicted safety operation monitoring result and the corresponding second true label, and it is determined whether the current second loss function has converged; if the current second loss function has converged, the current safety risk assessment model is used as the above-mentioned preset safety risk assessment model; otherwise, the model parameters in the current safety risk assessment model are adjusted and training continues.

[0078] Specifically, the aforementioned safety risk assessment sample data includes historical physiological condition data of several workers, historical medical examination data corresponding to the historical physiological condition data, historical work area location and identity ID. Each set of historical physiological condition data, historical medical examination data, historical work area location and identity ID corresponds to a worker.

[0079] Specifically, the second loss function mentioned above is calculated using the following formula: In the formula, This represents the second loss function. Indicates the regression loss weights. Represents the regression loss function. Represents the sorting loss function. Indicates the ID recognition loss weight, This represents the ID recognition loss function, where N represents the total number of security risk assessment sample data. This represents the true risk value of sample data i' in the safety risk assessment. This represents the predicted risk value of sample data i' in the safety risk assessment. This represents the true risk value of another safety risk assessment sample data j, normalized to [0,1]. This represents the boundary margin of the sorting loss, initialized to 0.2. The feature vector representing the identity of security risk assessment sample data i' Let represent the predicted risk value of another safety risk assessment sample data j, and p represent the base of the logarithmic function.

[0080] In this preferred embodiment, a pre-trained security risk assessment model is obtained by iteratively training the pre-set security risk assessment model using security risk assessment sample data with a second real label.

[0081] In another preferred embodiment, it further includes: Physiological condition monitoring module; The aforementioned physiological condition monitoring module is used to acquire facial video data of the worker to be monitored at preset monitoring intervals after the worker to be monitored starts working. The above-mentioned facial video data is input into the above-mentioned preset physiological data analysis model to obtain the working heart rate, working blood oxygen, working blood pressure, working body temperature and working mood of the worker to be monitored during work. If any of the above-mentioned work heart rate, work blood oxygen, work blood pressure and work body temperature do not meet the preset safe work threshold, or if the work mood does not meet the preset safe work mood category, a safety risk warning will be issued.

[0082] Specifically, in order to obtain timely information about the physiological condition of workers during work, it is necessary to record facial video data during work. Then, through a preset physiological data analysis model, the health status of workers during work can be analyzed and identified. In the event of abnormalities in their heart rate, blood pressure, blood oxygen, body temperature, or emotions, timely safety risk warnings can be issued to improve the safety of workers in outdoor working environments.

[0083] Preferably, through a real-time monitoring and early warning mechanism, the present invention can issue an alarm in a timely manner when abnormal physiological parameters are detected, thereby improving the safety of power grid workers in high-temperature outdoor working environments.

[0084] It should be noted that the device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without creative effort. The above schematic diagrams are merely examples of a worker safety monitoring system and do not constitute a limitation on a worker safety monitoring system. It may include more or fewer components than illustrated, or combine certain components, or use different components.

[0085] Based on the above-described apparatus embodiments, the present invention provides corresponding method embodiments.

[0086] Indicative, such as Figure 3 As shown, another embodiment of the present invention provides a method for monitoring the safe operation of workers. The above method is applicable to the facial video analysis module in the above-described worker safety monitoring system in any of the above embodiments. The safe operation monitoring method includes: Step S101: Input the original facial video of the worker to be monitored before the operation into the preset physiological data analysis model, so that the input layer built into the preset physiological data analysis model can enhance the image quality of the original facial video to obtain the first facial video; wherein, the original facial video is obtained by the data acquisition module; Specifically, the aforementioned original facial video is recorded by the worker to be monitored before starting work. The video can be thirty seconds long, resulting in an original facial video with a preset recording duration.

[0087] Specifically, the aforementioned pre-defined physiological data analysis model adopts a hybrid neural network architecture with multi-task learning, combining temporal analysis and spatial feature extraction. The core modules include: a video preprocessing module, a spatiotemporal feature extraction backbone network (including 3D CNN and Transformer), a multi-branch physiological parameter prediction head, and a sentiment analysis branch.

[0088] Specifically, in the input layer of the preset physiological data analysis model, facial key points are detected for each video frame of the aforementioned original facial video. These facial key points are location points used to represent facial features (such as contours and facial features), and in this invention, they can be the locations of the facial features within a video frame. Illustratively, the model's built-in MTCNN (Multi-Task Convolutional Neural Network) can be used to detect facial key points.

[0089] Subsequently, to eliminate geometric differences in the video caused by factors such as head rotation, tilt, and distance, and to provide stable input for subsequent accurate feature extraction, the detected facial regions with key points need to be standardized into fixed-size (e.g., 112×112) images with frontalized facial poses through geometric transformation. In this invention, affine transformation is used to preprocess the faces in the video frames; this process is existing technology and will not be described in detail here.

[0090] Then, to eliminate image blur caused by camera shake and slight worker movement, and to correct uneven lighting (such as shadows and local overexposure), laying a high-quality image foundation for subsequent accurate extraction of skin color changes and subtle features such as micro-expressions, it is necessary to deblur the facial images. This can be achieved using the built-in, pre-trained DeblurGAN-v2 model. The DeblurGAN-v2 model is a generative adversarial network (GAN) whose generator is a U-Net-style encoder-decoder structure containing multiple downsampling and upsampling layers with skip connections. The network analyzes the image's blur pattern (usually motion blur) layer by layer and reconstructs lost high-frequency details (such as clear skin texture, eyebrow hairs, and pupil edges) in the feature space. The final output is a series of deblurred facial images that match the input.

[0091] Specifically, for each face image, its luminance channel image and color channel image are extracted, and the luminance channel image is divided into several image blocks; then the histogram of each image block is calculated, and the histogram is cropped and equalized according to a preset contrast limit to obtain the processed histogram; then the processed histograms are stitched together to obtain the processed luminance channel image; finally, the processed luminance channel image and the corresponding color channel image are merged to obtain the face image with enhanced image quality, and thus the first face video is obtained.

[0092] It should be noted that in this invention, the CLAHE algorithm can be used to achieve balanced illumination. Through the above process, this algorithm will output an image with uniform illumination, where shadows on the face are faded, overexposed areas are suppressed, and the overall contrast is moderate, resulting in a more consistent appearance of the skin area.

[0093] Preferably, by using a pre-set physiological data analysis model, workers' physiological data can be collected without contact before work begins, providing them with comprehensive health analysis and risk monitoring. This method reduces interference with workers' work while improving the convenience and real-time nature of data collection.

[0094] Step S102: Extract features from the first facial video according to the spatiotemporal feature extraction backbone network and encoder built into the preset physiological data analysis model to obtain the temporal feature matrix and mid-level spatiotemporal features; Specifically, the aforementioned spatiotemporal feature extraction backbone network mainly consists of four 3D ResNet blocks (3×3×3 convolutional kernels, 1×2×2 stride), capable of capturing subtle movements such as facial micro-expressions and vascular pulsations in the input first facial video. The encoder is a temporal Transformer encoder with six Transformer Encoder layers (8 heads, 256 hidden layers), which transforms the features extracted by the spatiotemporal feature extraction backbone network into temporal features capable of understanding the global context and long-range dependencies of the entire video sequence.

[0095] Specifically, in the four 3D ResNet layers of the spatiotemporal feature extraction backbone network, each layer performs convolution processing on the input. The aforementioned local spatiotemporal features are the final output of the entire spatiotemporal feature extraction backbone network, while the aforementioned intermediate spatiotemporal features are the output of the intermediate layers of the spatiotemporal feature extraction backbone network, which in this invention can be the output of the second layer. For the aforementioned spatiotemporal feature extraction backbone network, in this invention, the size of the video frame sequence of the first facial video as input can be 900 (time steps) × 112 (height) × 112 (width) × 3 (channels).

[0096] Specifically, the aforementioned local spatiotemporal features are input into the built-in encoder to add temporal position encoding to the spatial features within these local spatiotemporal features, resulting in the aforementioned temporal feature matrix. In the encoder, the spatial features are first flattened to obtain spatial vectors. Subsequently, since the Transformer itself lacks the ability to process sequence order, positional information needs to be explicitly injected, i.e., temporal position encoding is added to the spatial vectors, allowing the model to clearly know the position of each feature vector on the time axis of the video frame sequence. These spatial vectors are then weighted and concatenated through a multi-head self-attention layer in the encoder to obtain an abstract understanding of the overall first facial video input, thus yielding the aforementioned temporal feature matrix.

[0097] Step S103: Perform feature transformation and aggregation processing on the time-series feature matrix and the intermediate spatiotemporal features respectively to obtain the physiological status data of the worker to be monitored; Specifically, the physiological parameter prediction head obtains physiological condition data such as heart rate, blood oxygen, blood pressure, and body temperature, while the emotion analysis branch obtains emotion. Heart rate, blood oxygen, blood pressure, body temperature, and emotion together constitute physiological condition data.

[0098] Specifically, in the physiological parameter prediction head for heart rate and blood oxygen detection (composed of 1D convolution, LSTM network and PPG signal decoder), an attention mechanism is used to select the feature subset most relevant to the minute color fluctuations on the skin surface caused by changes in blood volume from the entire temporal feature matrix, thus obtaining a feature sequence for characterizing skin color changes.

[0099] Subsequently, a causal convolution operation is performed on the aforementioned feature sequence to generate a PPG waveform signal, which is the waveform signal of the photoplethysmography (PPG) wave. In this process, causal convolution is achieved through 1D convolution, effectively compressing the feature sequence and converting it into a scalar value, thereby obtaining the aforementioned PPG waveform signal.

[0100] Then, the PPG waveform signal is input into the built-in LSTM network for periodic analysis to obtain heart rate and blood oxygen saturation. Next, average pooling is performed on the temporal feature matrix along the time dimension to obtain the global feature vector. Subsequently, the vascular morphology features output from the preset shallow layer of the aforementioned spatiotemporal feature extraction backbone network are obtained. The global feature vector and vascular morphology features are then concatenated, nonlinearly transformed, and regressed to obtain blood pressure and body temperature.

[0101] Specifically, for emotions, spatial attention focusing is applied to the mid-level spatiotemporal features, followed by global average pooling in the spatial dimension to obtain a micro-expression feature sequence that represents facial micro-expressions. This sequence is then combined with a built-in GRU network to capture the facial emotion transformation process in the sequence, resulting in a global micro-expression feature vector. Finally, the global micro-expression feature vector is subjected to fully connected operations and normalization to obtain the emotion.

[0102] Step S104: The physiological condition data is sent to the safe operation monitoring module, so that the safe operation monitoring module inputs the physiological condition data, historical physical examination data, work area location, and identity ID into the preset safety risk assessment model, so that the preset safety risk assessment model performs work risk assessment and safety qualification assessment on the worker to be monitored, and obtains the safe operation monitoring result of the worker to be monitored; wherein, the historical physical examination data, work area location, and identity ID are obtained by the data acquisition module.

[0103] Specifically, the heart rate, blood oxygen, blood pressure, body temperature and mood output by the preset physiological data analysis model are used as physiological condition data. Combined with the worker's historical physical examination data, work area location and identity ID, the preset safety risk assessment model is used to determine whether the worker can carry out the work safety monitoring results.

[0104] Preferably, the present invention combines a preset physiological data analysis model and a preset safety risk assessment model, which realizes the collection of physiological data and risk assessment and screening of power grid workers before work. Compared with traditional health detection methods, it is more timely and accurate.

[0105] The above are preferred embodiments of the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A safety monitoring system for workers, characterized in that, include: Data acquisition module, facial video analysis module, and safe operation monitoring module; The data acquisition module is used to acquire the original facial video, historical physical examination data, work area location, and identity ID of the worker to be monitored before the operation. The facial video analysis module is used to input the original facial video into a preset physiological data analysis model to obtain the physiological condition data of the worker to be monitored; wherein, the original facial video is input into the input layer built into the preset physiological data analysis model to enhance the image quality of the original facial video to obtain a first facial video; Based on the built-in spatiotemporal feature extraction backbone network and encoder of the preset physiological data analysis model, feature extraction is performed on the first facial video to obtain the temporal feature matrix and mid-level spatiotemporal features; The time-series feature matrix and the intermediate-level spatiotemporal features are subjected to feature transformation and aggregation processes respectively to obtain the physiological status data of the worker to be monitored; The safety operation monitoring module is used to input the physiological condition data, historical physical examination data, work area location and identity ID into a preset safety risk assessment model, so that the preset safety risk assessment model can perform work risk assessment and safety qualification assessment on the worker to be monitored, and obtain the safety operation monitoring results of the worker to be monitored.

2. The safety monitoring system for workers according to claim 1, characterized in that, The process of inputting the original facial video into the input layer built into the preset physiological data analysis model to enhance the image quality of the original facial video and obtain a first facial video includes: Based on each video frame of the original facial video, the facial key points of each video frame are detected; Affine transformation alignment is performed based on the facial key points to obtain several facial images of a preset standard size; Each face image is deblurred to obtain several deblurred face images; For each face image, extract the luminance channel image and color channel image of the current face image, and divide the current luminance channel image into several image blocks; Calculate the histogram of each image block, and crop and equalize the histogram according to a preset contrast limit to obtain the processed histogram. The processed histograms are stitched together to obtain the processed luminance channel image; The processed luminance channel image and the corresponding color channel image are merged to obtain an enhanced face image, which in turn yields the first facial video.

3. The safety monitoring system for workers according to claim 2, characterized in that, The step of extracting features from the first facial video based on the built-in spatiotemporal feature extraction backbone network and encoder of the preset physiological data analysis model to obtain a temporal feature matrix and intermediate spatiotemporal features includes: The spatiotemporal feature extraction backbone network performs hierarchical convolution and feature transformation processing on the video frame sequence of the first facial video to obtain local spatiotemporal features of facial motion features and mid-level spatiotemporal features of facial deformation features; wherein, the local spatiotemporal features are obtained from the output of the last layer of the spatiotemporal feature extraction backbone network, and the mid-level spatiotemporal features are obtained from the output of the middle layer of the spatiotemporal feature extraction backbone network. The local spatiotemporal features are input into the built-in encoder to add time position encoding to the spatial features in the local spatiotemporal features, thereby obtaining the temporal feature matrix.

4. The safety monitoring system for workers according to claim 3, characterized in that, The step of inputting the local spatiotemporal features into a built-in encoder to add temporal location encoding to the spatial features in the local spatiotemporal features to obtain the temporal feature matrix includes: The encoder flattens the spatial features in the local spatiotemporal features to obtain a spatial vector; After adding temporal position encoding to the spatial vector, a long-term temporal dependency is modeled on the spatial vector with added temporal position encoding through a self-attention mechanism to obtain the temporal feature matrix used to characterize the global temporal context.

5. A safety monitoring system for workers according to claim 4, characterized in that, The physiological data include: heart rate, blood oxygen, blood pressure, body temperature, and mood; The process of performing feature transformation and aggregation on the temporal feature matrix and the intermediate spatiotemporal features respectively to obtain the physiological status data of the worker to be monitored includes: Based on the attention mechanism, a feature sequence for characterizing skin color changes is extracted from the temporal feature matrix; The feature sequence is subjected to causal convolution to generate a PPG waveform signal; The PPG waveform signal is input into the built-in LSTM network so that the LSTM network performs periodic analysis on the PPG waveform signal to obtain the heart rate and blood oxygen. The time-series feature matrix is ​​subjected to average pooling along the time dimension to obtain the global feature vector. Obtain the vascular morphology features output by the preset shallow layer of the spatiotemporal feature extraction backbone network; The global feature vector and blood vessel morphology features are concatenated and nonlinearly transformed to obtain a fused feature vector. A regression operation is performed on the fused feature vector to obtain the blood pressure and the body temperature; Spatial attention focusing processing is applied to the mid-level spatiotemporal features to obtain several weighted feature maps; Global average pooling is performed on all weighted feature maps in the spatial dimension to obtain a micro-expression feature sequence for representing facial micro-expressions; The micro-expression feature sequence is input into the built-in GRU network so that the GRU network can capture the facial emotion transformation process in the micro-expression feature sequence and obtain the micro-expression global feature vector. The global feature vectors of micro-expressions are sequentially subjected to fully connected operations and normalization to obtain the probability distributions corresponding to each emotion type, and the emotion type with the highest probability is taken as the emotion.

6. A safety monitoring system for workers according to claim 5, characterized in that, The physiological condition data, historical medical examination data, work area location, and identity ID are input into a preset safety risk assessment model so that the preset safety risk assessment model can perform work risk assessment and safety qualification assessment on the worker to be monitored, and obtain the work safety monitoring results of the worker to be monitored, including: The physiological condition data, historical medical examination data, work area location, and identity ID are input into a preset safety risk assessment model. This model retrieves risk information corresponding to the work area location from a preset area risk database based on the work area location. The model then obtains a work risk assessment result based on the risk information, physiological condition data, and historical medical examination data. The work risk assessment result indicates whether the worker being monitored is suitable for the job. The preset safety risk assessment model retrieves the validity period of the safety qualification of the worker to be monitored from the preset database based on the identity ID, and compares the validity period of the safety qualification with the preset safety validity period to obtain the safety qualification assessment result of the worker to be monitored. If the job risk assessment result is unqualified or the safety qualification assessment result is unqualified, the safety job monitoring result of the worker to be monitored is determined to be unqualified; otherwise, the safety job monitoring result of the worker to be monitored is determined to be qualified.

7. A safety monitoring system for workers according to claim 6, characterized in that, The training of the preset physiological data analysis model includes: Acquire several facial video samples with a first real label; wherein, the first real label is used to represent the real physiological condition information of the worker corresponding to the facial video sample; the real physiological condition information includes: real heart rate, real blood oxygen, real blood pressure, real body temperature and real emotion. The face video samples are input into the physiological data analysis model to be trained for iterative training until the first loss function converges, thus obtaining the preset physiological data analysis model. In each iteration of training, the current predicted physiological condition data is generated based on the current face video samples; the current first loss function is calculated based on the current predicted physiological condition data and the corresponding first true label, and it is determined whether the current first loss function has converged; if the current first loss function has converged, the current physiological data analysis model is used as the preset physiological data analysis model; otherwise, the model parameters in the current physiological data analysis model are adjusted, and training continues.

8. A safety monitoring system for workers according to claim 7, characterized in that, The training of the preset security risk assessment model includes: Acquire several safety risk assessment sample data with a second real label; wherein, the safety risk assessment sample includes: historical physiological condition data of several workers, historical physical examination data corresponding to the historical physiological condition data, historical work area location and identity ID; the second real label is used to represent the real work risk assessment result and the real safety qualification assessment result of the safety risk assessment sample data; The security risk assessment sample data is input into the security risk assessment model to be trained for iterative training until the second loss function converges, thereby generating the security risk assessment model. In each iteration, the current predicted safety operation monitoring result is generated based on the current safety risk assessment sample data; the current second loss function is calculated based on the current predicted safety operation monitoring result and the corresponding second true label, and it is determined whether the current second loss function has converged; if the current second loss function has converged, the current safety risk assessment model is used as the preset safety risk assessment model; otherwise, the model parameters in the current safety risk assessment model are adjusted, and training continues.

9. A safety monitoring system for workers according to claim 8, characterized in that, Also includes: Physiological condition monitoring module; The physiological condition monitoring module is used to acquire the facial video data of the worker to be monitored at preset monitoring intervals after the worker to be monitored starts working. The facial video data is input into the preset physiological data analysis model to obtain the worker's heart rate, blood oxygen, blood pressure, body temperature and mood during work. If any of the working heart rate, working blood oxygen, working blood pressure, and working body temperature does not meet the preset safe working threshold, or if the working mood does not meet the preset safe working mood category, a safety risk warning will be issued.

10. A method for monitoring the safe operation of workers, characterized in that, Applicable to the facial video analysis module in the safety operation monitoring system as described in claims 1-9; The safe operation monitoring method includes: The original facial video of the worker to be monitored before the operation is input into a preset physiological data analysis model, so that the input layer built into the preset physiological data analysis model can enhance the image quality of the original facial video to obtain a first facial video; wherein, the original facial video is obtained by the data acquisition module; Based on the built-in spatiotemporal feature extraction backbone network and encoder of the preset physiological data analysis model, feature extraction is performed on the first facial video to obtain the temporal feature matrix and mid-level spatiotemporal features; The time-series feature matrix and the intermediate-level spatiotemporal features are subjected to feature transformation and aggregation processes respectively to obtain the physiological status data of the worker to be monitored; The physiological condition data is sent to the safe operation monitoring module, which then inputs the physiological condition data, historical medical examination data, work area location, and identity ID into a preset safety risk assessment model. This allows the preset safety risk assessment model to perform an operation risk assessment and safety qualification assessment on the worker to be monitored, thereby obtaining the safe operation monitoring results for the worker. The historical medical examination data, work area location, and identity ID are obtained by the data acquisition module.