Intelligent construction site confined space operation monitoring and early warning method based on artificial intelligence

By applying artificial intelligence technology in a confined space operation environment, combining target detection, attitude estimation calculation method and risk prediction model, the problem of safety accidents in a confined space operation environment is solved, real-time monitoring and efficient early warning are achieved.

CN120126280APending Publication Date: 2025-06-10DATANG HUIZHOU THERMAL POWER CO LTD +1
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510275130.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In a confined space operation environment, it is difficult for the prior art to achieve real-time dynamic assessment of the operators and the environmental status, resulting in the occurrence of safety accidents.

Method used

Using an artificial intelligence-based method, combined with the DINOv2 object detection model, MotionBERT pose estimation algorithm and SARIMA model, operator status monitoring, behavior analysis and environmental risk prediction are carried out, comprehensive risk scores are calculated, and early warning parameters are dynamically adjusted.

Benefits of technology

Real-time monitoring and risk warning of the confined space operating environment is realized, the safety monitoring ability of the operator is improved, false alarms and missed reports are reduced, and the adaptability and accuracy of the early warning system is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126280A_ABST
    Figure CN120126280A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent construction site confined space operation monitoring and early warning method based on artificial intelligence, and the method comprises the following steps: S1, collecting source data in a confined space operation environment, and carrying out the preprocessing; s2, analyzing the environmental safety condition in the closed space; s3, using a DINOv2 target detection model to identify the existence state of the operator and the wearing condition of the protective equipment, and monitoring the behavior action of the operator in combination with a MotionBERT attitude estimation algorithm; s4, performing health state analysis on the operating personnel; s5, predicting the risk change trend of the closed space operation environment by using an SARIMA model; s6, all the results are fused to generate a final risk assessment result; s7, when a set early warning threshold value is reached, alarm information is generated; and S8, dynamically adjusting early warning parameters. According to the method, a DINOv2 target detection model, a MotionBERT attitude estimation algorithm and an SARIMA model are combined, monitoring and early warning of closed space operation are achieved, and the method has high-precision detection, intelligent identification and risk prediction capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent construction site safety monitoring, and particularly to a method for monitoring and warning of confined space operations in an intelligent construction site based on artificial intelligence. Background Art

[0002] In the confined space operation environment, such as in scenarios like mines, tunnels, underground pipe galleries, storage tanks, and chemical workshops, operators face great safety risks. These environments often have poor ventilation and complex gas compositions, making it easy to occur dangerous accidents such as oxygen deficiency, poisoning, and explosion. In addition, due to the limited nature of the operation space, it is difficult for manual inspections to comprehensively cover all areas. Traditional monitoring methods often rely on fixed sensor arrangements or manual observations, making it difficult to achieve real-time dynamic assessment of the operator and environmental conditions. Therefore, how to comprehensively monitor the confined space operation environment through intelligent means, analyze the operator's status in real time, and provide an efficient and low false-alarm safety warning mechanism is an important research direction in current intelligent construction site safety management.

[0003] Currently, the monitoring of confined space operations mainly relies on a single data source for environmental perception. For example, gas sensors are used to monitor the oxygen concentration and harmful gas concentration, and temperature and humidity sensors are used to determine whether the operation environment is suitable. These sensor technologies have good real-time performance in environmental parameter monitoring, but they cannot provide an effective assessment of the operator's status. In addition, due to the limited deployment positions of sensors, they are easily affected by environmental conditions. For example, poor ventilation may cause local accumulation of gas concentration, and the sensor installation positions may not cover key areas, resulting in deviation of monitoring data. At the same time, the data of a single sensor cannot effectively model the dynamic change trend of the operation environment, making it difficult to accurately predict future safety risks.

[0004] In terms of operator status monitoring, traditional methods mainly rely on manual inspections or video monitoring for observation. For example, cameras are used to monitor the operation site, and safety management personnel remotely view it. However, this method is limited by monitoring blind spots and cannot ensure continuous tracking of all operators. In addition, due to the strong subjectivity of manual judgment and the deviation of judgment criteria among different operators, false alarms or missed alarms may occur. Some studies introduce computer vision technology, such as using object detection algorithms to identify the presence status of operators and detect the wearing conditions of protective equipment such as safety helmets. However, the detection accuracy of traditional object detection methods (such as YOLO, Faster R-CNN) is limited in complex environments and difficult to handle problems such as light changes and occlusions. Especially in the low-light and highly reflective environment of confined spaces, the detection accuracy may drop significantly. At the same time, these methods can only identify the location and basic status of operators and cannot effectively analyze the behavior patterns of operators, making it difficult to determine whether the operator is in an abnormal state, such as falling or staying still for a long time.

[0005] In terms of the physiological state monitoring of operators, traditional methods usually adopt portable physiological monitoring devices, such as heart rate belts, smart watches, etc. These devices can provide data such as the heart rate, body temperature, and blood oxygen saturation of operators. However, the data collection of existing physiological monitoring devices relies on individual wearing, and may lead to incomplete or lost data due to problems such as improper wearing, device damage, or battery depletion. In addition, traditional physiological state monitoring methods usually rely on static threshold judgment. For example, an alarm is triggered when the heart rate exceeds the set range. However, due to large individual physiological differences, fixed thresholds may lead to false alarms or missed alarms, and lack the ability to analyze the long-term trends of operators' physiological states.

[0006] In terms of risk assessment and early warning, traditional methods mainly rely on historical experience or analysis based on fixed rules. For example, environmental parameters are alarmed by setting thresholds, or the safety state is judged based on simple logical rules. These methods lack the ability to deeply model data and cannot comprehensively analyze the working environment by combining multi-source data. In recent years, some risk assessment methods based on machine learning have begun to be applied in the field of industrial safety. For example, time series prediction models (such as LSTM, ARIMA) are used to analyze sensor data to predict the changing trends of environmental parameters. However, these methods usually only model based on environmental data and fail to fully combine the state information of operators, resulting in the lack of comprehensiveness of prediction results. At the same time, due to the high data dimension, existing methods have certain limitations in terms of real-time performance and computational efficiency, and it is difficult to meet the real-time monitoring requirements of high-risk working environments. Summary of the Invention

[0007] An object of the present invention is to propose an artificial intelligence-based monitoring and early warning method for confined space operations in intelligent construction sites. The present invention uses the DINOv2 object detection model to identify the presence state of operators and the wearing situation of protective equipment, and combines the MotionBERT pose estimation algorithm to analyze the behavior patterns of operators to judge whether the operators are in abnormal states. The SARIMA model is used to predict the changing trends of environmental risks. By combining the state data of operators and the results of physiological health assessment, the comprehensive risk score is calculated to achieve early warning of future safety risks, and the warning parameters are dynamically adjusted based on the risk level to improve the adaptability and accuracy of the early warning system.

[0008] An artificial intelligence-based monitoring and early warning method for confined space operations in intelligent construction sites according to an embodiment of the present invention includes the following steps:

[0009] S1. Collect the source data in the confined space operation environment and preprocess the collected source data;

[0010] S2. Analyze the environmental safety situation in the enclosed space based on the gas sensor data and temperature and humidity data in the source data, determine whether the oxygen concentration, harmful gas content, and temperature and humidity are within the safe range, extract the characteristics of the environmental change trend, and obtain the environmental safety status data;

[0011] S3. Use the DINOv2 object detection model to process the video data in the source data, identify the presence status of the operators and the wearing situation of the protective equipment, monitor the behavior of the operators in combination with the MotionBERT pose estimation algorithm, and determine whether there is an abnormal status to obtain the operator status characteristic data;

[0012] S4. Combine the physiological status data of the operators in the source data, analyze the health status of the operators, determine whether there are abnormal heart rates, abnormal body temperatures, or symptoms of hypoxia, and obtain the physiological health assessment results of the operators;

[0013] S5. Based on the environmental safety status data, operator status characteristic data, and physiological health assessment results of the operators, use the SARIMA model to predict the risk change trend of the enclosed space operation environment, construct a risk assessment model, and calculate the risk score to obtain the risk assessment result;

[0014] S6. Adopt a multimodal data fusion method to fuse the environmental safety status data, operator status characteristic data, physiological health assessment results of the operators, and risk assessment results to generate the final risk assessment result;

[0015] S7. When the final risk assessment result reaches the set warning threshold, trigger the warning mechanism, automatically generate an alarm message, and send a warning signal to the on-site operators and the remote management platform through the sound and light alarm device and the wireless communication network. At the same time, dynamically adjust the warning level according to the risk assessment result;

[0016] S8. After triggering the warning, dynamically adjust the warning parameters based on the operator feedback, environmental changes, and historical warning data to optimize the risk assessment model.

[0017] Optionally, the source data includes gas sensor data, temperature and humidity data, video data, and physiological status data of the operators.

[0018] Optionally, the specific content of S3 includes:

[0019] S31. Based on the video data, use the DINOv2 object detection model to perform object detection on the operation scene, detect the presence status of the operators, extract the position information of the operators, and generate a set of operator bounding boxes:

[0020]

[0021] Among them, B t represents the set of bounding boxes of workers at time t, represents the upper left coordinate of the bounding box of the i-th worker, represents the lower right coordinate of the bounding box of the i-th worker, and N represents the total number of workers detected currently;

[0022] S32. Based on the DINOv2 object detection results, perform keypoint detection on the worker area, and extract the skeletal keypoint information of the workers. The keypoints include the head, shoulders, elbows, wrists, knees, and ankles:

[0023]

[0024] Among them, represents the set of keypoints of the i-th worker at time t, represents the two-dimensional pixel coordinates of the j-th keypoint of the i-th worker, and K represents the total number of keypoints;

[0025] S33. Construct a keypoint temporal trajectory and define a time window:

[0026]

[0027] Among them, P i represents the keypoint temporal trajectory of the i-th worker within the time window [t - n, t], n represents the size of the time window, represents the set of keypoints of the i-th worker at time t - n, represents the set of keypoints of the i-th worker at time t;

[0028] S34. Input the keypoint temporal trajectory into the MotionBERT pose estimation algorithm for temporal modeling, and calculate the behavioral feature vector at time t. The MotionBERT pose estimation algorithm includes a bidirectional Transformer network, a graph convolutional network, and an attention mechanism;

[0029] S35. Based on the behavioral feature vector, calculate the behavioral category of the worker. The behavioral categories include standing, walking, bending, falling, and staying still for a long time:

[0030]

[0031] Among them, represents the behavioral category of the i-th worker at time t, MLP represents a multi-layer perceptron, represents the behavioral feature vector of the i-th worker at time t;

[0032] S36. Output the wearing situation of the protective equipment of the operator based on the object detection result, where the protective equipment includes a safety helmet, goggles and a protective suit;

[0033] S37. Based on the behavior category of the operator and the wearing situation of the protective equipment, determine whether the operator is in an abnormal state:

[0034]

[0035] Among them, represents the abnormal state of the operator, A abnormal represents the set of abnormal behavior categories, represents the behavior category of the operator, represents the wearing situation of the protective equipment of the operator, E missing represents the set of protective equipment categories;

[0036] S38. Generate the state characteristic data of the operator based on the abnormal state judgment result:

[0037]

[0038] Among them, represents the state characteristic data of the operator, represents the set of operator bounding boxes, represents the behavior category of the operator, represents the wearing situation of the protective equipment of the operator, represents the abnormal state of the operator.

[0039] Optionally, the S31 specifically includes:

[0040] S311. Extract video frames from the input video stream at the set frame rate f r Extract the video frame I t Normalize it to a fixed size of H×W and standardize it, where H in the fixed size represents the standardized height and W represents the standardized width;

[0041] S312. Use the Vision Transformer network of the DINOv2 object detection model as the backbone network to extract features from the standardized image, divide the standardized image into P×P small blocks, and map each small block to a feature vector:

[0042]

[0043] Among them, represents the set of small blocks at time t, SplitIntoPatches represents the image segmentation function, I t” represents the input image, and P represents the size of the patch;

[0044] The image segmentation function SplitIntoPatches includes:

[0045] Using a sliding window method, the input image I t ” is divided into non - overlapping patches of size P×P;

[0046] Using the patch division formula, assuming the input image I t ” has dimensions H×W, then the number of patches obtained after division is:

[0047]

[0048] where N patch represents the total number of patches, H represents the height of the patch, W represents the width of the patch, and P represents the size of the patch;

[0049] Perform patch embedding:

[0050]

[0051] where X t represents the patch feature vector at time t, W e represents the weight matrix for patch embedding, and b e represents the bias term for patch embedding;

[0052] S313. Calculate the positional encoding and add positional information to each patch:

[0053] P pos = sin(ωk)+cos(ωk);

[0054] X t ' = X t + P pos ;

[0055] where P pos represents the positional encoding matrix, ω represents the frequency parameter, k represents the Patch index, X t represents the patch feature vector at time t, and X t ' represents the patch feature vector after adding the positional encoding;

[0056] S314. Input the patch feature vector X t ' after adding the positional encoding into the Transformer encoder of the DINOv2 object detection model for global feature extraction, and calculate the attention weights of each image patch to obtain the global representation of each image patch:

[0057] ”'

[0058] Qt = X t W Q , K t = X t W K , V t = X t W V ;

[0059]

[0060] Among them, Z t represents the global feature representation at time t, A t represents the attention weight matrix at time t, V t represents the value matrix, softmax represents normalization, Q t represents the query matrix, d k represents the dimension of the key matrix, K t represents the key matrix, T represents the transpose operation, W Q , W K and W V respectively represent the weight matrices of the query matrix, key matrix and value matrix, X t ' represents the small block feature vector after adding the position encoding;

[0061] S315. Execute object detection based on the global feature Z extracted by the DINOv2 object detection model, calculate the bounding box of the object, calculate the class score of the operator through the classification head of DINOv2, and predict the bounding box of the object based on the regression head of DINOv2. The regression head represents a bounding box regression network, and the bounding box regression network includes a fully connected layer, a ReLU activation function and a normalization layer: t C

[0062] = sigmoid(MLP(Z t )); t ))

[0063]

[0064] Among them, B t represents the set of operator bounding boxes at time t, C t represents the class probability at time t, used to determine whether the detected object belongs to the operator, sigmoid represents the activation function, MLP represents the multi-layer perceptron, used to classify different object types, MLP bbox represents the bounding box regression network, used to predict the coordinates of the bounding box, represents the upper left corner coordinates of the bounding box of the i-th operator, denotes the coordinates of the lower right corner of the bounding box of the i-th operator, and N denotes the total number of operators currently detected;

[0065] S316. Calculate the overlap degree of the targets in the front and rear frames, remove duplicate targets, and when it is lower than the set threshold τ, eliminate this target:

[0066]

[0067]

[0068] where B t ′ represents the set of optimized target bounding boxes, denotes the matching degree of the i-th target at time t, IoU represents the intersection over union calculation function, B t represents the set of operator bounding boxes at time t, and B t represents the set of operator bounding boxes at time t - 1, and τ represents the set threshold.

[0069] Optionally, the S34 specifically includes:

[0070] S341. Perform spatial feature encoding on the temporal trajectory of the key points of the human body, use the graph convolutional network in the MotionBERT pose estimation algorithm to extract the topological relationship between different joints of the human body, and generate spatial features:

[0071]

[0072] where denotes the spatial feature of joint j at time t, N(j) denotes the set of adjacent joints of joint j, and α jk denotes the topological weight between joint j and k, and W s and b s respectively represent the spatial feature transformation matrix and the bias term, denotes the two-dimensional coordinates of joint k at time t, describing the temporal trajectory of the key points of the human body;

[0073] S342. Perform temporal feature encoding on the temporal trajectory of the key points of the human body, combine the temporal attention mechanism to calculate the dynamic correlation between time steps, and extract temporal information:

[0074]

[0075] where denotes the temporal feature of joint j at time t, T(t) denotes the time steps within the time window, and β tt′ denotes the attention weight between time steps t and t′, and W t and b trespectively represent the time feature transformation matrix and the bias term, represents the two-dimensional coordinates of joint j at time t′, describing the temporal trajectory of the key points of the human body;

[0076] S343. Integrate the spatial features and time features to calculate the spatio-temporal feature vector of the human body key points:

[0077]

[0078] where, represents the spatio-temporal feature vector of joint j at time t, and λ represents the adaptive integration coefficient, controlling the weights of the spatial features and time features;

[0079] S344. Input the spatio-temporal feature vector into the MotionBERT pose estimation algorithm, use the bidirectional Transformer network to calculate the human motion pattern, calculate the behavioral dependency of the key points based on the attention mechanism, and generate the spatio-temporal attention matrix:

[0080]

[0081] where, M t represents the spatio-temporal attention matrix at time t, softmax represents normalization, Y t represents the query matrix, d represents the dimension of the key matrix, U t represents the key matrix, and T represents the transpose operation;

[0082] S345. Calculate the human behavioral feature vector and extract the global behavioral information of the key points:

[0083]

[0084] where, represents the behavioral feature vector of joint j at time t, M t represents the spatio-temporal attention matrix at time t, G t represents the value matrix, represents the spatio-temporal feature vector of joint j at time t, W g represents the weight matrix of the value matrix;

[0085] S346. Calculate the global behavioral feature, calculate the mean value using the behavioral features of all joints, and calculate the change rate of the behavioral feature to analyze the human motion trend:

[0086]

[0087]

[0088] where, Denote the global behavior characteristics at time t, and J denote the total number of joints. Denote the behavior feature vector of joint j at time t. Denote the global behavior characteristics at time t-1. Denote the change trend of behavior characteristics.

[0089] Optionally, the S5 specifically includes:

[0090] S51. Collect environmental safety status data, operator status characteristic data, and operator physiological health assessment results, and construct an input feature set:

[0091] H t ={L t ,S t ,C t}, t = 1, 2,..., P;

[0092] Among them, H t denotes the input feature set, L t denotes the environmental safety status data at time t, S t denotes the operator status characteristic data, C t denotes the operator physiological health assessment result, and P denotes the total number of frames within the time window;

[0093] S52. Perform standardization processing on the input feature set H t and calculate the standardized value using the Z-score method;

[0094] S53. Use the SARIMA model to model the change trend of the job environment risk and construct a risk assessment model:

[0095]

[0096] Among them, y t denotes the risk feature value at time t, c denotes the constant term, φ i denotes the i-th order autoregressive coefficient, y t-i denotes the risk feature value at time t-i, θ j denotes the j-th order moving average coefficient, ε t-j denotes the white noise term at time t-i, ψ s denotes the s-th order seasonal coefficient, y t-s denotes the risk feature value at time t-s, ε t denotes the white noise term at time t, p denotes the number of autoregressive terms, q denotes the number of moving average terms, and S denotes the seasonal period;

[0097] S54. Calculate the job environment risk score according to the standardized input feature set.

[0098] R t = W r X′ t + b r ;

[0099] Wherein, R t represents the job environment risk score at time t, W r represents the weight matrix for risk score calculation, b r represents the bias term for job environment risk score calculation, and X′ t represents the set of standardized input features;

[0100] S55. Use the SARIMA model to predict the risk change trend in the next H steps and calculate the predicted value:

[0101]

[0102] Wherein, represents the risk eigenvalue at time t + H, y t+H-i represents the risk eigenvalue at time t + H - i, ε t+H-j represents the white noise term at time t + H - i, and y t+H-s represents the risk eigenvalue at time t + H - s;

[0103] S56. Combine the job environment risk score and the predicted value to calculate the final risk score:

[0104]

[0105] Wherein, S t represents the final risk score at time t, α 1 and α 2 respectively represent the adjustment coefficients of the job environment risk score and the predicted value;

[0106] S57. Based on the risk assessment threshold, determine whether there is a risk and generate a risk assessment result, where the risk assessment result includes the job environment risk score, the predicted value, the final risk score, and the risk determination result.

[0107] The beneficial effects of the present invention are:

[0108] First, the present invention adopts the DINOv2 object detection model, which makes the detection of operators more accurate and can identify the presence status of operators and the wearing situation of protective equipment in complex environments, providing reliable data support for subsequent behavior analysis and safety assessment. At the same time, the MotionBERT pose estimation algorithm combined with the time series modeling method can effectively extract the key point trajectories of operators, judge their behavior patterns, and identify abnormal states such as falls and long-term stillness, improving the safety monitoring ability of operators.

[0109] Secondly, in terms of risk assessment, the present invention uses the SARIMA model to conduct time series analysis and trend prediction on environmental safety status, operator status characteristics, and physiological health assessment data. Compared with traditional warning methods based on static thresholds, it can identify potential safety hazards in advance and improve the accuracy of risk prediction.

[0110] Finally, in terms of the warning mechanism, the present invention adopts a dynamic warning strategy. According to the change trend of the risk assessment results, it automatically adjusts the warning level to ensure that reasonable safety measures can be taken before the risk occurs. When the risk assessment result reaches the set threshold, the system automatically triggers an audible and visual alarm and sends warning information to operators and the remote management platform through the wireless communication network to ensure that relevant personnel can take corresponding measures in a timely manner. At the same time, the method also has the ability of self-learning and optimization, adjusting warning parameters based on historical data and optimizing the risk assessment model, so that the system can continuously improve the warning accuracy during long-term operation and reduce false alarms and missed alarms. BRIEF DESCRIPTION OF THE DRAWINGS

[0111] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0112] Figure 1 is a flowchart of a method for monitoring and warning confined space operations on a smart construction site based on artificial intelligence proposed by the present invention;

[0113] Figure 2 is a schematic diagram of the DINOv2 object detection model of a method for monitoring and warning confined space operations on a smart construction site based on artificial intelligence proposed by the present invention to identify operators and protective equipment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0114] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.

[0115] Reference Figure 1 and Figure 2, an artificial intelligence-based monitoring and early warning method for confined space operations on a smart construction site, including the following steps:

[0116] S1. Collect the source data in the confined space operation environment and preprocess the collected source data;

[0117] S2. Based on the gas sensor data and temperature and humidity data in the source data, analyze the environmental safety situation in the confined space, judge whether the oxygen concentration, harmful gas content, and temperature and humidity are within the safe range, and extract the environmental change trend characteristics to obtain the environmental safety status data;

[0118] S3. Use the DINOv2 object detection model to process the video data in the source data, identify the presence status of the operators and the wearing situation of the protective equipment, combine the MotionBERT pose estimation algorithm to monitor the behavior actions of the operators, and judge whether there is an abnormal status to obtain the operator status characteristic data;

[0119] S4. Combine the physiological status data of the operators in the source data, analyze the health status of the operators, judge whether there are abnormal heart rates, abnormal body temperatures, or hypoxia symptoms, and obtain the physiological health assessment results of the operators;

[0120] S5. Based on the environmental safety status data, operator status characteristic data, and physiological health assessment results of the operators, use the SARIMA model to predict the risk change trend of the confined space operation environment, construct a risk assessment model, and calculate the risk score to obtain the risk assessment result;

[0121] S6. Adopt a multi-modal data fusion method to fuse the environmental safety status data, operator status characteristic data, physiological health assessment results of the operators, and risk assessment results to generate the final risk assessment result;

[0122] S7. When the final risk assessment result reaches the set early warning threshold, trigger the early warning mechanism, automatically generate an alarm message, and send a warning signal to the on-site operators and the remote management platform through the sound and light alarm device and the wireless communication network, and at the same time dynamically adjust the early warning level according to the risk assessment result;

[0123] S8. After triggering the early warning, dynamically adjust the early warning parameters based on the operator feedback, environmental changes, and historical early warning data to optimize the risk assessment model.

[0124] In this embodiment, the source data includes gas sensor data, temperature and humidity data, video data, and physiological status data of the operators.

[0125] In this embodiment, the specific content of S3 is as follows:

[0126] S31. Based on the video data, use the DINOv2 object detection model to perform object detection on the operation scenario, detect the presence status of the operators, extract the location information of the operators, and generate a set of operator bounding boxes:

[0127]

[0128] Among them, B t represents the set of operator bounding boxes at time t, represents the upper left corner coordinates of the bounding box of the i-th operator, represents the lower right corner coordinates of the bounding box of the i-th operator, and N represents the total number of operators detected currently;

[0129] S32. Based on the DINOv2 object detection results, perform key point detection on the operator area, and extract the skeletal key point information of the operators. The key points include the head, shoulders, elbows, wrists, knees, and ankles:

[0130]

[0131] Among them, represents the set of key points of the i-th operator at time t, represents the two-dimensional pixel coordinates of the j-th key point of the i-th operator, and K represents the total number of key points;

[0132] S33. Construct the key point time series trajectory and define the time window:

[0133]

[0134] Among them, P i represents the key point time series trajectory of the i-th operator within the time window [t - n, t], n represents the size of the time window, represents the set of key points of the i-th operator at time t - n, represents the set of key points of the i-th operator at time t;

[0135] S34. Input the key point time series trajectory into the MotionBERT pose estimation algorithm for time series modeling, and calculate the behavior feature vector at time t. The MotionBERT pose estimation algorithm includes a bidirectional Transformer network, a graph convolutional network, and an attention mechanism;

[0136] S35. Based on the behavior feature vector, calculate the behavior category of the operator. The behavior categories include standing, walking, bending, falling, and staying stationary for a long time:

[0137]

[0138] Among them, represents the behavior category of the i-th operator at time t, and MLP represents a multi-layer perceptron. represents the behavior feature vector of the i-th operator at time t;

[0139] S36. Based on the object detection results, output the wearing situation of the operators' protective equipment, where the protective equipment includes safety helmets, goggles, and protective clothing;

[0140] S37. Based on the behavior category of the operator and the wearing situation of the protective equipment, determine whether the operator is in an abnormal state:

[0141]

[0142] Among them, represents the abnormal state of the operator, and A abnormal represents the set of abnormal behavior categories, represents the behavior category of the operator, represents the wearing situation of the operator's protective equipment, and E missing represents the set of protective equipment categories;

[0143] S38. Based on the abnormal state judgment result, generate the operator status feature data:

[0144]

[0145] Among them, represents the operator status feature data, represents the set of operator bounding boxes, represents the behavior category of the operator, represents the wearing situation of the operator's protective equipment, represents the abnormal state of the operator.

[0146] In this embodiment, the S31 specifically includes:

[0147] S311. Extract video frames from the input video stream at a set frame rate f r and normalize the extracted video frames I t to a fixed size of H×W and standardize them, where H in the fixed size represents the standardized height and W represents the standardized width;

[0148] S312. Use the Vision Transformer network of the DINOv2 object detection model as the backbone network to extract features from the standardized image, divide the standardized image into P×P small blocks, and map each small block to a feature vector:

[0149]

[0150] Among them, represents the set of patches at time t, SplitIntoPatches represents the image segmentation function, I t ” represents the input image, and P represents the size of the patches;

[0151] The image segmentation function SplitIntoPatches includes:

[0152] Using a sliding window method, the input image I t ” is divided into non-overlapping patches of size P×P;

[0153] Using the patch division formula, assuming the input image I t ” has dimensions H×W, then the number of patches obtained after division is:

[0154]

[0155] Among them, N patch represents the total number of patches, H represents the height of the patches, W represents the width of the patches, and P represents the size of the patches;

[0156] Perform patch embedding:

[0157]

[0158] Among them, X t represents the patch feature vector at time t, W e represents the weight matrix of patch embedding, and b e represents the bias term of patch embedding;

[0159] S313. Calculate the positional encoding and add positional information to each patch:

[0160] P pos = sin(ωk) + cos(ωk);

[0161] X′ t = X t + P pos ;

[0162] Among them, P pos represents the positional encoding matrix, ω represents the frequency parameter, k represents the Patch index, X t represents the patch feature vector at time t, and X t ′ represents the patch feature vector after adding the positional encoding;

[0163] S314. The patch feature vector X t'Input the Transformer encoder of the DINOv2 object detection model for global feature extraction, calculate the attention weights of each image patch, and obtain the global representation of each image patch:

[0164] ”'

[0165] Q t = X t W Q ,K t = X t W K ,V t = X t W V ;

[0166]

[0167] Among them, Z t represents the global feature representation at time t, A t represents the attention weight matrix at time t, V t represents the value matrix, softmax represents normalization, Q t represents the query matrix, d k represents the dimension of the key matrix, K t represents the key matrix, T represents the transpose operation, W Q 、W K and W V respectively represent the weight matrices of the query matrix, key matrix and value matrix, X t 'represents the small patch feature vector after adding position encoding;

[0168] S315. Based on the global feature Z t extracted by the DINOv2 object detection model, perform object detection, calculate the bounding box of the object, calculate the class score of the operator through the classification head of DINOv2, and predict the bounding box of the object based on the regression head of DINOv2. The regression head represents a bounding box regression network, and the bounding box regression network includes a fully connected layer, a ReLU activation function and a normalization layer:

[0169] C t = sigmoid(MLP(Z t ));

[0170]

[0171] Among them, B t represents the set of operator bounding boxes at time t, C tRepresents the class probability at time t, which is used to determine whether the detected target belongs to an operator. sigmoid represents the activation function, and MLP represents the multi-layer perceptron, which is used to classify different target types. MLP bbox Represents the bounding box regression network, which is used to predict the coordinates of the bounding box. Represents the upper left coordinate of the bounding box of the i-th operator. Represents the lower right coordinate of the bounding box of the i-th operator. N represents the total number of operators detected currently.

[0172] S316. Calculate the overlap degree of the targets in the front and back frames, and remove duplicate targets. When it is lower than the set threshold τ, eliminate this target:

[0173]

[0174]

[0175] Among them, B t ′ represents the optimized set of target bounding boxes. Represents the matching degree of the i-th target at time t. IoU represents the intersection over union calculation function. B t Represents the set of operator bounding boxes at time t. B t Represents the set of operator bounding boxes at time t - 1. τ represents the set threshold.

[0176] In this embodiment, the S34 specifically includes:

[0177] S341. Perform spatial feature encoding on the temporal trajectory of the key points of the human body, and use the graph convolutional network in the MotionBERT pose estimation algorithm to extract the topological relationship between different joints of the human body to generate spatial features:

[0178]

[0179] Among them, Represents the spatial feature of joint j at time t. N(j) represents the set of adjacent joints of joint j. α jk Represents the topological weight between joint j and k. W s and b s respectively represent the spatial feature transformation matrix and the bias term. Represents the two-dimensional coordinate of joint k at time t, describing the temporal trajectory of the key points of the human body.

[0180] S342. Perform temporal feature encoding on the temporal trajectory of the key points of the human body, and combine the temporal attention mechanism to calculate the dynamic correlation between time steps to extract temporal information:

[0181]

[0182] Among them, represents the time feature of joint j at time t, T(t) represents the time step within the time window, and β tt′ represents the attention weight between time steps t and t′, and W t and b t represent the time feature transformation matrix and the bias term respectively. represents the two-dimensional coordinate of joint j at time t′, describing the temporal trajectory of the key points of the human body;

[0183] S343. Fuse the spatial feature and the time feature, and calculate the spatio-temporal feature vector of the human body key points:

[0184]

[0185] Among them, represents the spatio-temporal feature vector of joint j at time t, λ represents the adaptive fusion coefficient, controlling the weights of the spatial feature and the time feature;

[0186] S344. Input the spatio-temporal feature vector into the MotionBERT pose estimation algorithm, use the bidirectional Transformer network to calculate the human motion pattern, calculate the key point behavior dependence relationship based on the attention mechanism, and generate the spatio-temporal attention matrix:

[0187]

[0188] Among them, M t represents the spatio-temporal attention matrix at time t, softmax represents normalization, Y t represents the query matrix, d represents the dimension of the key matrix, U t represents the key matrix, and T represents the transpose operation;

[0189] S345. Calculate the human behavior feature vector and extract the global behavior information of the key points:

[0190]

[0191] Among them, represents the behavior feature vector of joint j at time t, M t represents the spatio-temporal attention matrix at time t, G t represents the value matrix, represents the spatio-temporal feature vector of joint j at time t, and W g represents the weight matrix of the value matrix;

[0192] S346. Calculate the global behavior features, calculate the mean using the behavior features of all joints, calculate the change rate of the behavior features, and analyze the human motion trend:

[0193]

[0194]

[0195] Among them, represents the global behavior feature at time t, J represents the total number of joints, represents the behavior feature vector of joint j at time t, represents the global behavior feature at time t-1, represents the change trend of the behavior feature.

[0196] In this embodiment, the S5 specifically includes:

[0197] S51. Collect the environmental safety status data, the status feature data of the operator, and the physiological health assessment results of the operator, and construct an input feature set:

[0198] H t ={L t ,S t ,C t},t = 1,2,...,P;

[0199] Among them, H t represents the input feature set, L t represents the environmental safety status data at time t, S t represents the status feature data of the operator, C t represents the physiological health assessment results of the operator, and P represents the total number of frames within the time window;

[0200] S52. Perform standardization processing on the input feature set H t and calculate the standardized value using the Z-score method;

[0201] S53. Use the SARIMA model to model the change trend of the job environment risk and construct a risk assessment model:

[0202]

[0203] Among them, y t represents the risk feature value at time t, c represents the constant term, φ i represents the i-th order autoregressive coefficient, y t-i represents the risk feature value at time t-i, θ j represents the j-th order moving average coefficient, ε t-j represents the white noise term at time t-i, ψs represents the seasonal coefficient of the s-th order, y t-s represents the risk eigenvalue at time t - s, ε t represents the white noise term at time t, p represents the number of autoregressive terms, q represents the number of moving average terms, and S represents the seasonal period;

[0204] S54. Calculate the risk score of the operating environment based on the standardized input feature set:

[0205] R t = W r X′ t + b r ;

[0206] where, R t represents the risk score of the operating environment at time t, W r represents the weight matrix for calculating the risk score, b r represents the bias term for calculating the risk score of the operating environment, and X′ t represents the standardized input feature set;

[0207] S55. Use the SARIMA model to predict the risk change trend in the next H steps and calculate the predicted value:

[0208]

[0209] where, represents the risk eigenvalue at time t + H, y t+H-i represents the risk eigenvalue at time t + H - i, ε t+H-j represents the white noise term at time t + H - i, y t+H-s represents the risk eigenvalue at time t + H - s;

[0210] S56. Calculate the final risk score by combining the risk score of the operating environment and the predicted value:

[0211]

[0212] where, S t represents the final risk score at time t, α 1 and α 2 respectively represent the adjustment coefficients of the risk score of the operating environment and the predicted value;

[0213] S57. Judge whether there is a risk based on the risk assessment threshold and generate a risk assessment result, where the risk assessment result includes the risk score of the operating environment, the predicted value, the final risk score, and the risk judgment result.

[0214] Example 1:

[0215] To verify the feasibility of the present invention in implementation, the present invention is applied to an underground enclosed working space of a large chemical plant. The working space is mainly used for storing and handling hazardous chemicals. The internal environment is complex, with risks of high temperature, high humidity, low oxygen, and leakage of toxic gases. When working in this environment, operators need to strictly abide by safety regulations. However, due to the limitations of manual monitoring methods, accidents still occur from time to time. For example, operators do not wear safety protection equipment as required, remain stationary for a long time, or suddenly fall. Traditional monitoring systems are difficult to detect in time and issue effective warnings. Therefore, the intelligent monitoring and warning system of the present invention is deployed in this enclosed working space to achieve all-round safety protection.

[0216] In this scenario, the intelligent monitoring system first monitors the oxygen concentration, temperature and humidity, and harmful gases (such as carbon monoxide, hydrogen sulfide, benzene, etc.) in the working area in real time through an environmental sensor network. The collected environmental data is preliminarily processed by edge computing nodes and then uploaded to the central server. In addition, the DINOv2 object detection model is used for video data analysis to automatically detect the presence status of operators, the wearing situation of protective equipment, and the hazard sources in the working scenario, such as chemical leakage, equipment damage, etc. The behavior analysis of operators is completed by the MotionBERT pose estimation algorithm. Based on skeleton key point extraction and time series modeling technology, it can accurately identify behaviors such as standing, walking, bending, and falling, and trigger an alarm when an abnormal state occurs.

[0217] During the actual test process, the risk assessment function of the system was also verified. After being trained based on the environmental data of the past three months, the SARIMA model is used to predict the change trend of oxygen concentration within the next 24 hours to ensure early warning of low oxygen risk. In addition, the multi-modal data fusion module jointly analyzes environmental data, operator status data, and physiological health data, and adjusts the warning threshold according to different risk levels, enabling the system to provide accurate safety warnings before a real dangerous event occurs. For example, during a certain operation, the system detected that an operator was not wearing a protective mask, and at the same time, the physiological monitoring device reported that their blood oxygen saturation was lower than 90%. Combining the environmental data analysis, the oxygen concentration in this area was decreasing. The system immediately issued a secondary risk warning and sent an alarm to the operator and the remote safety management platform through the wireless communication network. Finally, the operator was safely evacuated within 3 minutes, avoiding a possible hypoxia accident.

[0218] To further verify the effectiveness of the present invention, the safety event data within 6 months before and after the system deployment was compared and analyzed, focusing on evaluating key indicators such as accident incidence rate, false alarm rate, and operation efficiency.

[0219] Table 1 Experimental comparison data table

[0220] Monitoring indicators Before system deployment (average value) After system deployment (average value) Change rate Monthly average number of safety accidents 6.8 cases 2.1 cases -69% Number of hypoxic accidents 3.2 cases 0.9 cases -72% Accidents of harmful gas poisoning 2.1 cases 0.7 cases -67% Equipment damage accidents 1.5 cases 0.5 cases -67% False alarm rate 12.5% 4.3% -66% Manual inspection time 15.6 hours / week 10.1 hours / week -35% Average warning response time 8.4 minutes 3.1 minutes -63% Accuracy of behavior recognition 81.3% 96.7% +19% Accuracy of target detection 85.5% 98.1% +15% Accuracy of risk prediction 78.2% 94.6% +21%

[0221] In terms of preventing safety accidents, after the system is deployed, the average monthly number of safety accidents in the confined space has decreased from 6.8 to 2.1, a reduction of 69%. Among them, the low-oxygen accidents and harmful gas poisoning accidents have decreased by 72% and 67% respectively, and the equipment damage accidents have decreased by 67%. This result shows that through the multi-source data fusion analysis and intelligent early warning mechanism of the present invention, the safety of the working environment is effectively improved, and the sudden accidents caused by environmental factors and abnormal conditions of operators are reduced.

[0222] In terms of false alarm rate and manual inspection, the intelligent monitoring method of the present invention has greatly reduced the unnecessary safety alarms and the need for manual intervention. After the system is deployed, the false alarm rate has decreased from 12.5% to 4.3%, a reduction of 66%, which proves the improvement of the accuracy of the present invention in intelligent data processing and behavior analysis. In addition, the manual inspection time has decreased from 15.6 hours per week to 10.1 hours per week, a reduction of 35%, indicating that the intelligent monitoring system has reduced the workload of manual safety inspections, improved the inspection efficiency, and enabled operators to devote more energy to core production tasks.

[0223] In terms of early warning response ability, the present invention has significantly shortened the response time after an accident occurs. Before the system is deployed, the average early warning response time for safety incidents is 8.4 minutes, while with the assistance of the intelligent early warning system, this time has been shortened to 3.1 minutes, a reduction of 63%, improving the emergency response speed. This improvement enables safety management personnel to take measures more quickly and effectively intervene before the accident expands, reducing the risk of personnel injury and equipment damage.

[0224] In terms of the accuracy of intelligent monitoring, the object detection, behavior recognition, and risk prediction capabilities of the present invention have all been improved. The accuracy of object detection has increased from 85.5% to 98.1%, the accuracy of behavior recognition has increased from 81.3% to 96.7%, and the accuracy of risk prediction has increased from 78.2% to 94.6%. These improvements show that the DINOv2 object detection model accurately identifies operators and their protective equipment wearing conditions, the MotionBERT pose estimation algorithm accurately judges the behavior state of operators, and the SARIMA prediction model effectively analyzes environmental data and makes accurate risk assessments.

[0225] The above is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. A method for monitoring and early warning of confined space operations in smart construction sites based on artificial intelligence, characterized in that: The steps include: S1. Collect source data in a confined space working environment and pre-process the collected source data; S2. Analyze the environmental safety situation in the confined space based on the gas sensor data and temperature and humidity data in the source data, determine whether the oxygen concentration, harmful gas content, temperature and humidity are within a safe range, and extract the environmental change trend characteristics to obtain environmental safety status data; S3. Use the DINOv2 target detection model to process the video data in the source data, identify the presence status of the operator and the wearing of protective equipment, and use the MotionBERT posture estimation algorithm to monitor the behavior of the operator and determine whether there is an abnormal state to obtain the operator status feature data; S4. Combine the physiological status data of the operators in the source data to analyze the health status of the operators, determine whether there are abnormal heart rate, abnormal body temperature or hypoxia symptoms, and obtain the physiological health assessment results of the operators; S5. Based on the environmental safety status data, the status characteristic data of the operators and the physiological health assessment results of the operators, the SARIMA model is used to predict the risk change trend of the confined space working environment, a risk assessment model is constructed, and the risk score is calculated to obtain the risk assessment results; S6. Use multimodal data fusion methods to fuse environmental safety status data, operator status characteristic data, operator physiological health assessment results and risk assessment results to generate the final risk assessment results; S7. When the final risk assessment result reaches the set warning threshold, the warning mechanism is triggered, alarm information is automatically generated, and warning signals are sent to on-site operators and remote management platforms through sound and light alarm devices and wireless communication networks. At the same time, the warning level is dynamically adjusted according to the risk assessment results; S8. After the warning is triggered, the warning parameters are dynamically adjusted and the risk assessment model is optimized based on the feedback from operators, environmental changes and historical warning data.

2. According to claim 1, a method for monitoring and early warning of confined space operations in a smart construction site based on artificial intelligence is characterized in that: The source data includes gas sensor data, temperature and humidity data, video data and operator physiological status data.

3. The method for monitoring and early warning of confined space operations in a smart construction site based on artificial intelligence according to claim 1 is characterized in that: The S3 specifically includes: S31. Based on the video data, the DINOv2 target detection model is used to detect the target of the operation scene, detect the existence status of the operator, extract the location information of the operator, and generate a set of operator bounding boxes: Among them, B t represents the set of bounding boxes of operators at time t, represents the coordinates of the upper left corner of the bounding box of the i-th operator, represents the coordinate of the lower right corner of the bounding box of the ith operator, and N represents the total number of operators currently detected; S32. Based on the DINOv2 target detection result, key point detection is performed on the operator area to extract the skeleton key point information of the operator, where the key points include the head, shoulders, elbows, wrists, knees and ankles: in, represents the key point set of the i-th operator at time t, represents the 2D pixel coordinates of the jth key point of the i-th operator, and K represents the total number of key points; S33. Construct the key point time series trajectory and define the time window: Among them, P i represents the key point time series trajectory of the i-th operator in the time window [tn, t], n represents the time window size, represents the key point set of the i-th operator at time tn, Represents the key point set of the i-th operator at time t; S34, inputting the key point time series trajectory into the MotionBERT posture estimation algorithm for time series modeling, and calculating the behavior feature vector at time t, wherein the MotionBERT posture estimation algorithm includes a bidirectional Transformer network, a graph convolutional network, and an attention mechanism; S35. Calculate the behavior category of the operator based on the behavior feature vector, where the behavior category includes standing, walking, bending, falling, and long-term stillness: in, represents the behavior category of the i-th operator at time t, MLP represents multi-layer perceptron, Represents the behavior feature vector of the i-th operator at time t; S36. Based on the target detection result, output the wearing status of the protective equipment of the operator, where the protective equipment includes a safety helmet, goggles and protective clothing; S37. Based on the operator's behavior type and protective equipment wearing situation, determine whether the operator is in an abnormal state: in, Indicates the abnormal status of the operator, A abnormal represents a set of abnormal behavior categories, Indicates the behavior category of the operator. Indicates the wearing of protective equipment by the operator, E missing Represents a collection of protective equipment categories; S38. Generate operator status characteristic data based on the abnormal status judgment result: in, Indicates the operator status characteristic data, represents the set of bounding boxes of workers, Indicates the behavior category of the operator. Indicates whether the workers are wearing protective equipment. Indicates the abnormal status of the operator.

4. The method for monitoring and early warning of confined space operations in a smart construction site based on artificial intelligence according to claim 3 is characterized in that: The S31 specifically includes: S311, from the input video stream according to the set frame rate f r Extract the video frame and convert the extracted video frame I t Normalize to a fixed size H×W and perform standardization on the image, where H in the fixed size represents the standardized height and W represents the standardized width; S312, using the Vision Transformer network of the DINOv2 target detection model as the backbone network, extract features from the standardized image, divide the standardized image into P×P small blocks, and map each small block to a feature vector: in, represents a set of small patches at time t, SplitIntoPatches represents the image segmentation function, I t " represents the input image, P represents the size of the small patch; The image segmentation function SplitIntoPatches includes: The input image I is transformed into t "Divided into non-overlapping blocks of size P×P; Using the small block partitioning formula, let the input image I t "The dimension is H×W, then the number of small blocks obtained after division is: Among them, N patch represents the total number of small blocks, H represents the height of a small block, W represents the width of a small block, and P represents the size of a small block; To embed small blocks: Among them, X t Represents a small feature vector at time t, W e represents the weight matrix of the small patch embedding, b e represents the bias term for small patch embedding; S313. Calculate the position code and add position information to each small block: P pos =sin(ωk)+cos(ωk); X t '=X t +P pos ; Among them, P pos represents the position encoding matrix, ω represents the frequency parameter, k represents the Patch index, X t Represents a small feature vector at time t, X t ' represents the small block feature vector after adding position encoding; S314, adding the small block feature vector X after position encoding t 'Input the Transformer encoder of the DINOv2 target detection model to extract global features and calculate the attention weight of each image block to obtain the global representation of each image block: Q t =X t 'W Q ,K t =X t 'W K ,V t =X t 'W V ; Among them, Z t represents the global feature representation at time t, A t Represents the attention weight matrix at time t, V t represents the value matrix, softmax represents normalization, Q t represents the query matrix, d k represents the dimension of the key matrix, K t represents the key matrix, T represents the transpose operation, W Q , W K and W V Denote the weight matrices of query matrix, key matrix and value matrix respectively, X t ′ represents the small block feature vector after adding position encoding; S315, global feature Z extracted based on DINOv2 target detection model t , perform target detection, calculate the target's bounding box, calculate the worker's category score through the DINOv2 classification head, and predict the target's bounding box based on the DINOv2 regression head. The regression head represents a bounding box regression network, which includes a fully connected layer, a ReLU activation function, and a normalization layer: C t =sigmoid(MLP(Z t )); Among them, B t represents the bounding box set of operators at time t, C t represents the category probability at time t, which is used to determine whether the detected target belongs to the operator. Sigmoid represents the activation function. MLP represents the multi-layer perceptron, which is used to classify different target types. bbox represents a bounding box regression network, used to predict the coordinates of the bounding box, represents the coordinates of the upper left corner of the bounding box of the i-th operator, represents the coordinate of the lower right corner of the bounding box of the ith operator, and N represents the total number of operators currently detected; S316, calculate the overlap degree of the previous and next frame targets, remove the duplicate targets, when When it is lower than the set threshold τ, the target is eliminated: Among them, B t ′ represents the optimized target bounding box set, represents the matching degree of the i-th target at time t, IoU represents the intersection-over-union calculation function, B t represents the bounding box set of operators at time t, B t represents the set of bounding boxes of the operators at time t-1, and τ represents the set threshold.

5. The method for monitoring and early warning of confined space operations in a smart construction site based on artificial intelligence according to claim 3 is characterized in that: The S34 specifically includes: S341. Encode the spatial features of the key points of the human body in time series, and use the graph convolutional network in the MotionBERT posture estimation algorithm to extract the topological relationship between different joints of the human body to generate spatial features: in, represents the spatial features of joint j at time t, N(j) represents the set of adjacent joints of joint j, α jk represents the topological weight between joints j and k, W s and b s Represent the spatial feature transformation matrix and bias term respectively, Represents the two-dimensional coordinates of joint k at time t, describing the time-series trajectory of key points of the human body; S342, encode the time feature of the key point time sequence trajectory of the human body, calculate the dynamic correlation between time steps in combination with the time attention mechanism, and extract the time sequence information: in, represents the temporal features of joint j at time t, T(t) represents the time step in the time window, and β tt′ represents the attention weight between time steps t and t′, W t and b t Represent the temporal feature transformation matrix and the bias term respectively, Represents the two-dimensional coordinates of joint j at time t′, describing the temporal trajectory of key points of the human body; S343, integrating spatial features and temporal features, and calculating the spatiotemporal feature vectors of key points of the human body: in, represents the spatiotemporal feature vector of joint j at time t, λ represents the adaptive fusion coefficient, which controls the weight of spatial and temporal features; S344. Input the spatiotemporal feature vector to the MotionBERT posture estimation algorithm, use the bidirectional Transformer network to calculate the human motion pattern, calculate the key point behavior dependency based on the attention mechanism, and generate the spatiotemporal attention matrix: Among them, M t represents the spatiotemporal attention matrix at time t, softmax represents normalization, Y t represents the query matrix, d represents the dimension of the key matrix, U t represents the key matrix, T represents the transpose operation; S345, calculate the human behavior feature vector and extract the global behavior information of the key points: in, represents the behavior feature vector of joint j at time t, M t represents the spatiotemporal attention matrix at time t, G t represents the value matrix, represents the spatiotemporal feature vector of joint j at time t, W g The weight matrix representing the value matrix; S346. Calculate the global behavior characteristics, use the behavior characteristics of all joints to calculate the mean, and calculate the behavior characteristic change rate to analyze the human body movement trend: in, represents the global behavior characteristics at time t, J represents the total number of joints, represents the behavior feature vector of joint j at time t, represents the global behavior characteristics at time t-1, Indicates the changing trend of behavioral characteristics.

6. The method for monitoring and early warning of confined space operations in a smart construction site based on artificial intelligence according to claim 1 is characterized in that: The S5 specifically includes: S51. Collect environmental safety status data, operator status feature data, and operator physiological health assessment results, and construct an input feature set: H t ={L t ,S t ,C t },t=1,2,...,P; Among them, H t represents the input feature set, L t Represents the environmental safety status data at time t, S t Indicates the operator status characteristic data, C t represents the physiological health assessment result of the operator, and P represents the total number of frames in the time window; S52, input feature set H t Standardization was performed and the standardized value was calculated using the Z-score method; S53. Use the SARIMA model to model the risk change trend of the working environment and build a risk assessment model: Among them, y t represents the risk characteristic value at time t, c represents the constant term, φ i represents the i-th order autoregressive coefficient, y t-i represents the risk characteristic value at time ti, θ j represents the j-th order moving average coefficient, ε t-j represents the white noise term at time ti, ψ s represents the s-th order seasonal coefficient, y t-s represents the risk characteristic value at time ts, ε t represents the white noise term at time t, p represents the number of autoregressive terms, q represents the number of moving average terms, and S represents the seasonal cycle; S54. Calculate the operating environment risk score based on the standardized input feature set: R t =W r X' t +b r ; Among them, R t represents the risk score of the working environment at time t, W r represents the weight matrix for risk score calculation, b r represents the bias term in the calculation of the operating environment risk score, X' t Represents the standardized input feature set; S55. Use the SARIMA model to predict the risk change trend in the next H steps and calculate the predicted value: in, represents the risk characteristic value at time t+H, y t+H-i represents the risk characteristic value at time t+Hi, ε t+H-j represents the white noise term at time t+Hi, y t+H-s Represents the risk characteristic value at time t+Hs; S56. Calculate the final risk score by combining the operating environment risk score and the predicted value: Among them, S t represents the final risk score at time t, α1 and α2 represent the adjustment coefficients of the operating environment risk score and the predicted value, respectively; S57. Determine whether there is a risk based on the risk assessment threshold and generate a risk assessment result, wherein the risk assessment result includes an operating environment risk score, a predicted value, a final risk score and a risk determination result.

Citation Information

Cited By

  • Comprehensive pipe gallery abnormal state early warning method and system based on Internet of Things

    CN120316691A

  • Method and system for detecting wearing of labor protection appliances for operating personnel in petroleum and natural gas station

    CN120388396A

  • Method and system for detecting the wearing of labor protection equipment by workers at oil and gas stations

    CN120388396B

  • Hierarchical intelligent emergency response method and system for high-risk operation

    CN121214629A

  • Intelligent safety helmet management method and equipment based on Internet of Things

    CN121279792A