Limited space operation behavior monitoring method and system
By preprocessing images and performing multi-model detection in confined space work environments, the accuracy problem of work behavior recognition under low light conditions was solved, and the reliability of the work was improved.
Patent Information
- Application Number
- CN202511096594.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-07
AI Technical Summary
Existing technologies struggle to accurately identify confined space operations in low-light conditions, reducing operational reliability.
Low-light enhancement models are used to preprocess images of safety equipment, supervisors, and workers. Combined with pre-trained equipment detection models, face target detection models, and motion detection models, equipment detection, job status detection, and dangerous motion detection are performed to generate work behavior monitoring results.
It improves image clarity in low-light environments, ensuring accurate and reliable identification of operations in confined spaces.
Smart Images

Figure CN120912988A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of limited space operation monitoring, and particularly relates to a limited space operation behavior monitoring method and system. BACKGROUND
[0002] Limited space operation behavior refers to operation activities in a closed or partially closed space such as a cable well, a deep foundation pit, the inside of a transformer and an accident oil pool. Such operation environment is usually complex, and there are dangerous factors including low illumination, poor ventilation and toxic and harmful gases, and the safety risk is high, which can easily threaten the life safety of operation personnel. Therefore, how to reliably monitor limited space operation behavior is crucial.
[0003] At present, the prior art mainly identifies the abnormal behavior of limited space operation through video playback or simple algorithms (such as movement detection), but in a low-illumination environment, the image is unclear, the features are confused, and it is difficult to accurately identify the abnormal operation behavior of the limited space, thereby reducing the reliability of the limited space operation. SUMMARY
[0004] The present application provides a limited space operation behavior monitoring method and system, which solves the technical problem that the prior art mainly identifies the abnormal behavior of limited space operation through video playback or simple algorithms (such as movement detection), but in a low-illumination environment, the image is unclear, the features are confused, and it is difficult to accurately identify the abnormal operation behavior of the limited space, thereby reducing the reliability of the limited space operation.
[0005] The present application provides a limited space operation behavior monitoring method and system, which solves the technical problem that the prior art mainly identifies the abnormal behavior of limited space operation through video playback or simple algorithms (such as movement detection), but in a low-illumination environment, the image is unclear, the features are confused, and it is difficult to accurately identify the abnormal operation behavior of the limited space, thereby reducing the reliability of the limited space operation.
[0006] The present application provides a limited space operation behavior monitoring method and system, which solves the technical problem that the prior art mainly identifies the abnormal behavior of limited space operation through video playback or simple algorithms (such as movement detection), but in a low-illumination environment, the image is unclear, the features are confused, and it is difficult to accurately identify the abnormal operation behavior of the limited space, thereby reducing the reliability of the limited space operation.
[0007] The present application provides a limited space operation behavior monitoring method and system, which solves the technical problem that the prior art mainly identifies the abnormal behavior of limited space operation through video playback or simple algorithms (such as movement detection), but in a low-illumination environment, the image is unclear, the features are confused, and it is difficult to accurately identify the abnormal operation behavior of the limited space, thereby reducing the reliability of the limited space operation.
[0008] The present application provides a limited space operation behavior monitoring method and system, which solves the technical problem that the prior art mainly identifies the abnormal behavior of limited space operation through video playback or simple algorithms (such as movement detection), but in a low-illumination environment, the image is unclear, the features are confused, and it is difficult to accurately identify the abnormal operation behavior of the limited space, thereby reducing the reliability of the limited space operation.
[0009] The present application provides a limited space operation behavior monitoring method and system, which solves the technical problem that the prior art mainly identifies the abnormal behavior of limited space operation through video playback or simple algorithms (such as movement detection), but in a low-illumination environment, the image is unclear, the features are confused, and it is difficult to accurately identify the abnormal operation behavior of the limited space, thereby reducing the reliability of the limited space operation.
[0010] Perform risk assessment according to the equipment list, the post state and the action detection result, and obtain a work behavior monitoring result corresponding to the limited space.
[0011] Optionally, the low-illumination enhancement model comprises a decomposition network, an enhancement network, a non-local mean filtering module and a channel-wise multiplication layer, and a processing process of the low-illumination enhancement model is specifically as follows:
[0012] Perform image decomposition operation on the input limited-space work image through the decomposition network to obtain a corresponding reflection image and an illumination image, wherein the decomposition network comprises five first convolutional layers and a first activation function layer connected in sequence;
[0013] Perform image enhancement processing on the illumination image through the enhancement network to obtain a corresponding target illumination image;
[0014] Perform denoising processing on the reflection image through the non-local mean filtering module to obtain a corresponding target reflection image;
[0015] Perform pixel-by-pixel multiplication processing on the target illumination image and the target reflection image through the channel-wise multiplication layer to obtain a target limited-space work image, wherein the target limited-space work image comprises a target safety equipment image, a target supervisor image or a target worker image.
[0016] Optionally, the equipment detection model comprises a backbone network, a three-way feature pyramid network and a detection head, and the step of performing equipment detection on the target safety equipment image through the pre-trained equipment detection model to obtain a corresponding equipment list comprises:
[0017] Perform feature extraction on the target safety equipment image through the backbone network to output a plurality of equipment feature maps layer by layer;
[0018] Perform feature fusion on the plurality of equipment feature maps through the three-way feature pyramid network to generate a plurality of equipment fusion feature maps;
[0019] Perform equipment recognition on each of the equipment fusion feature maps through the detection head to obtain a plurality of equipment prediction regions;
[0020] Select an equipment prediction region with the largest confidence from each of the equipment prediction regions as a reference region, and calculate an intersection-over-union between the reference region and each of the equipment prediction regions;
[0021] Determine whether each of the intersection-overs-union is less than or equal to a preset intersection-over-union threshold;
[0022] When the Jaccard index is less than or equal to the Jaccard index threshold, the equipment prediction area associated with the Jaccard index is determined as a target equipment prediction area;
[0023] The equipment data of the reference area and the equipment data of the target equipment prediction area are used to generate a corresponding equipment list.
[0024] Optionally, the step of detecting the post state of the target supervisor image by using the pre-trained face target detection model comprises:
[0025] The pre-trained face target detection model is used to extract features of the target supervisor image to obtain a corresponding identity feature vector;
[0026] Cosine similarities between the identity feature vector and each standard identity feature vector in a preset identity database are calculated respectively;
[0027] When each cosine similarity is less than or equal to a preset identity threshold, a preset authentication number is updated, and the step of obtaining the safety equipment image, the supervisor image and the worker image of the limited space is executed;
[0028] When the authentication number is greater than or equal to a preset frame number threshold, the post state is determined as an off-duty state;
[0029] When any cosine similarity is greater than the identity threshold, the post state is determined as an on-duty state.
[0030] Optionally, the action detection model comprises a high-resolution network, a three-dimensional convolutional neural network, a feature fusion layer and an action classification network, and the step of detecting the dangerous action of the target worker image by using the pre-trained action detection model to obtain a corresponding action detection result comprises:
[0031] The high-resolution network is used to extract two-dimensional posture features of the target worker image to obtain corresponding two-dimensional posture features;
[0032] The three-dimensional convolutional neural network is used to extract spatio-temporal convolution features of a worker image frame sequence associated with the target worker image to obtain corresponding spatio-temporal convolution features;
[0033] The feature fusion layer is used to perform a splicing operation on the two-dimensional posture features and the spatio-temporal convolution features to obtain a corresponding splicing vector;
[0034] The action classification network is used to detect the splicing vector to obtain a corresponding action detection result.
[0035] Optionally, the step of performing risk assessment according to the equipment list, the post state and the action detection result to obtain the work behavior monitoring result corresponding to the limited space comprises:
[0036] When the action detection result is abnormal action, the action detection result is taken as the work behavior monitoring result corresponding to the limited space;
[0037] When the action detection result is normal work, the equipment list is matched with a preset standard equipment list;
[0038] When the equipment list is not matched with the standard equipment list, equipment wearing violation is taken as the work behavior monitoring result corresponding to the limited space;
[0039] When the equipment list is matched with the standard equipment list, it is judged whether the post state is an on-duty state;
[0040] If the post state is the on-duty state, normal work is taken as the work behavior monitoring result corresponding to the limited space;
[0041] If the post state is an off-duty state, abnormal work is taken as the work behavior monitoring result corresponding to the limited space.
[0042] The second aspect of the present application provides a limited space work behavior monitoring system, comprising:
[0043] The acquisition module is used for acquiring safety equipment images, supervisor images and worker images of a limited space, and performing image preprocessing on the safety equipment images, the supervisor images and the worker images through a preset low-illumination enhancement model to obtain target safety equipment images, target supervisor images and target worker images;
[0044] The equipment detection module is used for performing equipment detection on the target safety equipment images through a pre-trained equipment detection model to obtain a corresponding equipment list;
[0045] The post state detection module is used for performing post state detection on the target supervisor images through a pre-trained face target detection model to obtain a corresponding post state;
[0046] The action detection module is used for performing dangerous action detection on the target worker images through a pre-trained action detection model to obtain a corresponding action detection result;
[0047] The evaluation module is used for performing risk assessment according to the equipment list, the post state and the action detection result to obtain the work behavior monitoring result corresponding to the limited space.
[0048] The third aspect of the present application provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer program, when executed by the processor, causes the processor to perform the steps of the limited space operation behavior monitoring method according to any one of the preceding aspects.
[0049] The fourth aspect of the present application provides a computer readable storage medium, which stores a computer program, and the computer program, when executed, implements the limited space operation behavior monitoring method according to any one of the preceding aspects.
[0050] The fifth aspect of the present application provides a computer program product, which comprises a computer program stored on a non-transitory computer readable storage medium, and the computer program comprises program instructions, wherein when the program instructions are executed by a computer, the computer performs the limited space operation behavior monitoring method according to any one of the preceding aspects.
[0051] From the above technical solutions, the present application has the following advantages:
[0052] The present application pre-processes the safety equipment image, the supervisor image and the worker image by the preset low-illumination enhancement model to obtain the target safety equipment image, the target supervisor image and the target worker image, detects the target safety equipment image by the pre-trained equipment detection model to obtain the corresponding equipment list, detects the post state of the target supervisor image by the pre-trained face target detection model to obtain the corresponding post state, detects the dangerous action of the target worker image by the pre-trained action detection model to obtain the corresponding action detection result, and evaluates the risk according to the equipment list, the post state and the action detection result to obtain the corresponding operation behavior monitoring result of the limited space. The present application overcomes the technical problem that the prior art mainly identifies the abnormal operation behavior of the limited space through video playback or simple algorithm, but it is difficult to accurately identify the abnormal operation behavior of the limited space in a low-illumination environment, and reduces the reliability of the operation of the limited space. Compared with the traditional limited space operation behavior monitoring method, the present application pre-processes the safety equipment image, the supervisor image and the worker image by the preset low-illumination enhancement model to obtain the target safety equipment image, the target supervisor image and the target worker image, which improves the cleaning degree of the collected images in a low-illumination environment, and at the same time, the equipment list, the post state and the action detection result are combined to evaluate the risk of the limited space operation, which ensures the accuracy of the identification of the abnormal operation behavior of the limited space and improves the reliability of the operation of the limited space. BRIEF DESCRIPTION OF DRAWINGS
[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings described below are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.
[0054] Figure 1 A step flow chart of a limited space operation behavior monitoring method provided for the first embodiment of the present application.
[0055] Figure 2 A step flow chart of a limited space operation behavior monitoring method provided for the second embodiment of the present application.
[0056] Figure 3 A structural schematic diagram of an equipment detection model provided for the second embodiment of the present application.
[0057] Figure 4 A structural schematic diagram of a neural network classifier provided for the second embodiment of the present application.
[0058] Figure 5 A structural block diagram of a limited space operation behavior monitoring system provided for the third embodiment of the present application.
[0059] Figure 6 A structural block diagram of an electronic device provided for the fourth embodiment of the present application. DETAILED DESCRIPTION
[0060] The embodiments of the present application provide a limited space operation behavior monitoring method and system, which are used to solve the technical problem that the prior art mainly identifies the abnormal behavior of the limited space operation through video playback or simple algorithm (such as movement detection), but the image is not clear in low-illumination environment, the features are confused, and it is difficult to accurately identify the abnormal operation behavior of the limited space, thereby reducing the reliability of the limited space operation.
[0061] In order to make the purposes, features and advantages of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the embodiments described below are only some of the embodiments of the present application, but not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.
[0062] Please refer to Figure 1 , Figure 1 A step flow chart of a limited space operation behavior monitoring method provided for the first embodiment of the present application.
[0063] The application provides a limited space operation behavior monitoring method, which comprises the following steps.
[0064] In step 101, the safety equipment image, the supervisor image and the operation personnel image of the limited space are acquired, the safety equipment image, the supervisor image and the operation personnel image are respectively subjected to image preprocessing through a preset low-illumination enhancement model, and the target safety equipment image, the target supervisor image and the target operation personnel image are obtained.
[0065] The safety equipment image refers to visual data of the power safety protection equipment worn by the operation personnel in the limited space operation scene.
[0066] The supervisor image refers to the facial feature data of the supervisor in the limited space operation scene.
[0067] The operation personnel image refers to visual data containing the human body posture and action time sequence information of the operation personnel in the limited space operation monitoring scene, which is acquired and processed through a video.
[0068] The target safety equipment image refers to the safety equipment image subjected to the preprocessing of the low-illumination enhancement model.
[0069] The target supervisor image refers to the supervisor image subjected to the preprocessing of the low-illumination enhancement model.
[0070] The target operation personnel image refers to the operation personnel image subjected to the preprocessing of the low-illumination enhancement model.
[0071] In the embodiment of the application, the safety equipment image, the supervisor image and the operation personnel image of the limited space are acquired in real time through a high-definition wide-angle camera arranged in the limited space, the safety equipment image, the supervisor image and the operation personnel image are respectively subjected to image preprocessing through a preset low-illumination enhancement model, and the target safety equipment image, the target supervisor image and the target operation personnel image are obtained.
[0072] In step 102, the equipment detection model is trained in advance, and the target safety equipment image is subjected to equipment detection through the pre-trained equipment detection model, so that the corresponding equipment list is obtained.
[0073] The equipment list refers to the number and type of the equipment worn by the operation personnel.
[0074] In the embodiment of the present application, the target safety equipment image is input into the pre-trained equipment detection model to obtain a plurality of equipment prediction regions. The equipment prediction region with the maximum confidence is selected as a reference region from each equipment prediction region, and the intersection over union between the reference region and each equipment prediction region is calculated respectively. It is judged whether each intersection over union is less than or equal to a preset intersection over union threshold. When the intersection over union is less than or equal to the intersection over union threshold, the equipment prediction region associated with the intersection over union is determined as the target equipment prediction region. The equipment data of the reference region and the equipment data of the target equipment prediction region are used to generate a corresponding equipment list.
[0075] Step 103, performing post state detection on the target supervisor image by using the pre-trained face target detection model to obtain a corresponding post state.
[0076] The post state refers to the working state of the supervisor.
[0077] In the embodiment of the present application, the pre-trained face target detection model is used to extract features from the target supervisor image to obtain a corresponding identity feature vector. The cosine similarity between the identity feature vector and each standard identity feature vector in the preset identity database is calculated respectively. When each cosine similarity is less than or equal to a preset identity threshold, the preset authentication times are updated, and step 101 is executed. When the authentication times are greater than the identity threshold, the on-duty state is determined as the corresponding post state. Greater than or equal to the preset frame number threshold, the off-duty state is determined as the corresponding post state.
[0078] Step 104, performing dangerous action detection on the target worker image by using the pre-trained action detection model to obtain a corresponding action detection result.
[0079] In the embodiment of the present application, the pre-trained action detection model is used to perform dangerous action detection on the target worker image to obtain a corresponding action detection result, wherein the action detection model includes a high-resolution network, a three-dimensional convolutional neural network, a feature fusion layer, and an action classification network.
[0080] Step 105, performing risk assessment according to the equipment list, the post state and the action detection result to obtain a corresponding work behavior monitoring result of the limited space.
[0081] In the embodiment of the present application, when the action detection result is abnormal action, the action detection result is taken as the operation behavior monitoring result corresponding to the limited space. When the action detection result is normal operation, the equipment list is matched with the preset standard equipment list. When the equipment list is not matched with the standard equipment list, the equipment wearing violation is taken as the operation behavior monitoring result corresponding to the limited space. When the equipment list is matched with the standard equipment list, it is judged whether the post state is on-duty state. If the post state is on-duty state, the operation is normal, which is taken as the operation behavior monitoring result corresponding to the limited space. If the post state is off-duty state, the operation is abnormal, which is taken as the operation behavior monitoring result corresponding to the limited space.
[0082] In the embodiment of the present application, the safety equipment image, the supervisor image and the operation personnel image are respectively preprocessed by the preset low-illumination enhancement model to obtain the target safety equipment image, the target supervisor image and the target operation personnel image. The equipment detection model is used for detecting the target safety equipment image to obtain the corresponding equipment list. The face target detection model is used for detecting the target supervisor image to obtain the corresponding post state. The action detection model is used for detecting the target operation personnel image to obtain the corresponding action detection result. The risk is evaluated according to the equipment list, the post state and the action detection result to obtain the operation behavior monitoring result corresponding to the limited space. The technical problem that the existing technology mainly identifies the abnormal operation behavior of the limited space through video playback or simple algorithm, but it is difficult to accurately identify the abnormal operation behavior of the limited space in the low-illumination environment, and the reliability of the operation of the limited space is reduced is overcome. Compared with the traditional limited space operation behavior monitoring method, the safety equipment image, the supervisor image and the operation personnel image are respectively preprocessed by the preset low-illumination enhancement model to obtain the target safety equipment image, the target supervisor image and the target operation personnel image, so that the cleaning degree of the collected image in the low-illumination environment is improved. At the same time, the risk of the limited space operation is evaluated in combination with the equipment list, the post state and the action detection result, so that the accuracy of the identification of the abnormal operation behavior of the limited space is ensured, and the reliability of the operation of the limited space is improved.
[0083] Please refer to Figure 2 , Figure 2 The step flow chart of a limited space operation behavior monitoring method provided in the second embodiment of the present application is shown in the figure.
[0084] The limited space operation behavior monitoring method provided by the present application comprises the following steps:
[0085] Step 201, acquire the safety equipment image, the supervisor image and the operation personnel image of the limited space, and perform image preprocessing on the safety equipment image, the supervisor image and the operation personnel image through a preset low-illumination enhancement model to obtain a target safety equipment image, a target supervisor image and a target operation personnel image.
[0086] In the embodiment of the application, the safety equipment image, the supervisor image and the operation personnel image corresponding to the current frame in the safety equipment video stream, the supervisor video stream and the operation personnel video stream of the limited space are acquired in real time by the high-definition wide-angle camera deployed in the limited space, and image preprocessing is performed on the safety equipment image, the supervisor image and the operation personnel image through a preset low-illumination enhancement model to obtain a target safety equipment image, a target supervisor image and a target operation personnel image.
[0087] It should be noted that the low-illumination enhancement model includes a decomposition network, an enhancement network, a non-local mean filtering module and a channel-wise multiplication layer, and the processing process of the low-illumination enhancement model is specifically as follows:
[0088] A1, performing image decomposition operation on the input limited space operation image through the decomposition network to obtain the corresponding reflection image and the illumination image, wherein the decomposition network includes five first convolution layers and a first activation function layer connected in sequence.
[0089] The limited space operation image refers to the safety equipment image, the supervisor image and the operation personnel image input into the low-illumination enhancement model.
[0090] The first convolution layer refers to a 3*3 convolution layer.
[0091] The first activation function layer refers to a LeakyReLU activation function layer (leaky linear rectifier function).
[0092] In the embodiment of the application, the image decomposition operation is performed on the input limited space operation image through the five 3*3 convolution layers and the LeakyReLU activation function layers connected in sequence to obtain the corresponding reflection image and the illumination image.
[0093] A2, performing image enhancement processing on the illumination image through the enhancement network to obtain the corresponding target illumination image.
[0094] The target illumination image refers to the illumination image after balancing the contrast of bright and dark areas.
[0095] In the embodiment of the application, the dynamic range compression and smooth constraint processing are performed on the illumination image through the enhancement network to obtain the corresponding target illumination image, wherein the enhancement network includes five convolution layers connected in sequence.
[0096] It should be noted that the convolution kernel of the first layer to the third layer of the convolution layer in the enhanced network is 3*3, and the activation function is LeakyReLU activation function. The convolution kernel of the convolution layer of the fourth layer is 1*1, and the activation function is Sigmoid activation function (i.e. logic function).
[0097] A3, the reflection image is denoised by the non-local mean filtering module to obtain the corresponding target reflection image.
[0098] The target reflection image refers to the reflection image after suppressing noise and retaining edges.
[0099] In the embodiment of the application, the non-local mean filtering module is used to filter the reflection image to obtain the corresponding target reflection image, wherein the target limited space operation image includes a target safety equipment image, a target supervisor image or a target operator image.
[0100] A4, the target illumination image and the target reflection image are multiplied by the pixel point by using the channel-by-channel multiplication layer to obtain the target limited space operation image.
[0101] The target limited space operation image refers to the limited space operation image after the low-illumination enhancement model image preprocessing, including the target safety equipment image, the target supervisor image and the target operator image.
[0102] In the embodiment of the application, the target illumination image and the target reflection image are multiplied by the pixel point by using the channel-by-channel multiplication layer to obtain the target limited space operation image, wherein the channel-by-channel multiplication layer includes a multiplication layer and a gamma correction layer. For example, the target illumination image and the target reflection image are multiplied by the pixel point by using the multiplication layer to obtain a point multiplication image, and then the point multiplication image is gamma corrected by using the gamma correction layer to obtain the target limited space operation image.
[0103] Step 202, the equipment detection model is pre-trained, and the target safety equipment image is detected to obtain the corresponding equipment list.
[0104] Further, the equipment detection model includes a backbone network, a three-way feature pyramid network and a detection head, and step 202 includes the following substeps:
[0105] S11, the target safety equipment image is extracted by using the backbone network to output a plurality of equipment feature maps layer by layer.
[0106] In the embodiment of the application, the target safety equipment image is extracted by using the backbone network to output a plurality of equipment feature maps layer by layer. Wherein, the backbone network is ResNet-50 (i.e. 50-layer residual network). For example, refer to Figure 3As shown, the backbone network is used to extract features of the target safety equipment image, and the feature maps of C2 level (i.e., the first equipment feature map), the feature maps of C3 level (i.e., the second equipment feature map), the feature maps of C4 level (i.e., the third equipment feature map), and the feature maps of C5 level (i.e., the fourth equipment feature map) are output layer by layer.
[0107] S12, feature fusion is performed on the plurality of equipment feature maps by the three-way feature pyramid network to generate a plurality of equipment fusion feature maps.
[0108] In the embodiment of the present application, the three-way feature pyramid network is used to perform feature fusion on the plurality of equipment feature maps to generate a plurality of equipment fusion feature maps. For example, referring to Figure 3 As shown, the three-way feature pyramid network includes a plurality of 1x1 convolutional layers, a plurality of feature fusion layers, and a plurality of 3x3 convolutional layers. The first equipment feature map is extracted by the 1x1 convolutional layer to obtain the first level feature map. The second equipment feature map is extracted by the 1x1 convolutional layer to obtain the second level feature map. The third equipment feature map is extracted by the 1x1 convolutional layer to obtain the third level feature map. The fourth equipment feature map is extracted by the 1x1 convolutional layer to obtain the fourth level feature map. The fourth level feature map and the third level feature map are fused by the feature fusion layer to obtain the first fusion equipment feature map. The first fusion equipment feature map is extracted by the 1x1 convolutional layer to obtain the fifth level feature map. The fifth level feature map and the second level feature map are fused to obtain the second fusion equipment feature map. The second fusion equipment feature map is extracted by the 1x1 convolutional layer to obtain the sixth level feature map. The first level feature map and the sixth level feature map are fused by the feature fusion layer to obtain the third fusion equipment feature map. The third fusion equipment feature map is extracted by the 3x3 convolutional layer to obtain the first equipment fusion feature map. The first equipment fusion feature map and the sixth level feature map are fused by the feature fusion layer to obtain the first intermediate feature map. The first intermediate feature map is extracted by the 3x3 convolutional layer to obtain the second equipment fusion feature map. The second equipment fusion feature map and the fifth level feature map are fused by the feature fusion layer to obtain the second intermediate feature map. The second intermediate feature map is extracted by the 3x3 convolutional layer to obtain the third equipment fusion feature map. The third equipment fusion feature and the fourth level feature map are fused by the feature fusion layer and the 3x3 convolutional layer in sequence to obtain the fourth equipment fusion feature map.
[0109] S13, each equipment fusion feature map is identified by the detection head to obtain a plurality of equipment prediction regions.
[0110] The equipment prediction region refers to the bounding box prediction in the multi-scale target feature map.
[0111] In the embodiments of the present application, referring to Figure 3 As shown, each equipment recognition is performed on each equipment fusion feature map by the detection head to obtain a plurality of multi-scale target feature maps. The bounding box prediction in each multi-scale target feature map is taken as an equipment prediction region.
[0112] S14, selecting the equipment prediction region with the maximum confidence from each equipment prediction region as a reference region, and calculating the intersection over union between the reference region and each equipment prediction region.
[0113] In the embodiments of the present application, the equipment prediction region with the maximum confidence is selected from each equipment prediction region as a reference region, and the intersection over union between the reference region and each equipment prediction region is calculated (i.e. the intersection over union is the ratio of the intersection area and the union area between the reference region and the equipment prediction region).
[0114] S15, judging whether each intersection over union is less than or equal to a preset intersection over union threshold.
[0115] The intersection over union threshold refers to a key parameter for eliminating redundant equipment prediction regions.
[0116] S16, when the intersection over union is less than or equal to the intersection over union threshold, the equipment prediction region associated with the intersection over union is determined as a target equipment prediction region.
[0117] In the embodiments of the present application, it is judged whether each intersection over union is less than or equal to a preset intersection over union threshold. When the intersection over union is less than or equal to the intersection over union threshold, it is considered that the equipment prediction region and the reference region do not point to the same target, and the equipment prediction region associated with the intersection over union is determined as a target equipment prediction region. When the intersection over union is greater than the intersection over union threshold, it is considered that the equipment prediction region and the reference region point to the same target, and is regarded as redundant and is eliminated.
[0118] S17, generating a corresponding equipment list by using the equipment data of the reference region and the equipment data of the target equipment prediction region.
[0119] In the embodiments of the present application, the equipment data of the reference region and the equipment data of the target equipment prediction region are used to construct a corresponding equipment list.
[0120] It should be noted that the training process of the equipment detection model is specifically:
[0121] B1, obtain a plurality of training equipment images, perform image preprocessing on the plurality of training equipment images (i.e., label each training equipment image), and obtain a training equipment set.B2, train the equipment detection model using the training equipment set, and obtain training equipment prediction region data.B3, based on a preset training loss function, calculate the training loss function value of the training equipment set through the training equipment prediction region data.B4, when the training loss function value is greater than or equal to a preset standard loss function value, adjust the model parameters of the equipment detection model, and jump to B2-B5, until the training loss function value is less than the standard loss function value.B5, when the training loss function value is less than the standard loss function value, a trained equipment detection model is generated.
[0122] The training loss function is specifically:
[0123]
[0124] Among them, is a classification loss value, is a bounding box regression loss, is a training loss function value, is a first weight coefficient, is a second weight coefficient, is a cth real class label, is a probability of model prediction class c, is a parameter of a prediction box associated with an ith coordinate point, is a parameter of a real box associated with an ith coordinate point, is a difference between the parameter of the prediction box and the parameter of the real box, x is the horizontal coordinate of the coordinate point, y is the vertical coordinate of the coordinate point, i is the index of the coordinate point, and c is the second index of the class.
[0125] Step 203, using a pre-trained face target detection model to perform feature extraction on the target supervisor image to obtain a corresponding identity feature vector.
[0126] In the embodiments of the present application, referring to Figure 3 , a pre-trained face target detection model is used to perform feature extraction on the target supervisor image to obtain a corresponding identity feature vector. The face target detection model has the same structure as the equipment detection model.
[0127] It should be noted that the training process of the face target detection model is specifically:
[0128] C1, obtain a plurality of training face images, perform image preprocessing on the plurality of training face images (i.e., label each training face image), and obtain a training face set.C2, train the face target detection model using the training face set, and obtain training identity feature vector data.C3, based on a preset training loss function, calculate the training loss function value of the training face set through the training identity feature vector data.C4, when the training loss function value is greater than or equal to a preset standard loss function value, adjust the model parameters of the face target detection model, and jump to execute C2-C5, until the training loss function value is less than the standard loss function value.C5, when the training loss function value is less than the standard loss function value, generate a trained face target detection model.
[0129] Step 204, respectively calculate the cosine similarity between the identity feature vector and each standard identity feature vector in the preset identity database.
[0130] The standard identity feature vector refers to the reference database stored feature for comparing the identity of the to-be-tested personnel.
[0131] In the embodiment of the present application, the cosine similarity between the identity feature vector and each standard identity feature vector in the preset identity database is calculated respectively through the preset cosine similarity function.
[0132] It should be noted that the cosine similarity function is specifically:
[0133] ;
[0134] Wherein, is the cosine similarity, is the identity feature vector, is the standard identity feature vector.
[0135] Step 205, when each cosine similarity is less than or equal to a preset identity threshold, update the preset authentication number, and jump to execute the step of obtaining the security equipment image, the supervisor image and the operator image of the limited space.
[0136] In the embodiment of the present application, when each cosine similarity is less than or equal to a preset identity threshold, it is determined that the current face identity authentication fails, the preset authentication number is updated, and the supervisor image at the current time is obtained to jump to execute steps 203-205.
[0137] It should be noted that when the first face identity authentication fails, the authentication number is updated from 0 to 1, and when the second personnel identity authentication fails, the authentication number is updated from 1 to 2 (i.e., the preset authentication number is set to 0, and is incremented by 1 for iteration update after each face identity authentication fails).
[0138] Step 206: When the number of authentication attempts is greater than or equal to a preset frame count threshold, the job status is determined to be off-duty.
[0139] The frame count threshold refers to the upper limit of authentication failures, with a value of 3.
[0140] In this embodiment of the invention, when the number of authentications is greater than or equal to 3, it is considered that the supervisor has left his post, and the off-duty status is determined as the corresponding post status.
[0141] Step 207: When any cosine similarity is greater than the identity threshold, the job status is determined to be on duty.
[0142] In this embodiment of the invention, when any cosine similarity is greater than the identity threshold, the supervisor is considered to be continuously on duty, and the on-duty status is determined as the corresponding job status.
[0143] Step 208: Detect dangerous actions in the target worker image using a pre-trained action detection model to obtain the corresponding action detection results.
[0144] Furthermore, the action detection model includes a high-resolution network, a three-dimensional convolutional neural network, a feature fusion layer, and an action classification network. Step 208 includes the following sub-steps:
[0145] S21. Extract two-dimensional pose features from the target worker image using a high-resolution network to obtain the corresponding two-dimensional pose features.
[0146] In this embodiment of the invention, two-dimensional pose features are extracted from the target worker image using an HRNet network (i.e., a high-resolution network) to obtain the corresponding two-dimensional pose features. For example, the target worker image (e.g., 512*512 pixels) is input into the high-resolution network to obtain the corresponding two-dimensional pose features, wherein the two-dimensional pose features are a set of coordinates of 17 human body key points + 21 hand key points.
[0147] It should be noted that before using a high-resolution network to extract 2D pose features from the target worker images, a top-down strategy can be employed. This involves using a human detector to process the target worker images to obtain bounding boxes for each worker. The image regions within these bounding boxes are then input into the high-resolution network to obtain a set of heatmaps corresponding to key points. The heatmaps are then decoded, and the coordinates of the key points are located by solving for the coordinates of the maximum response point. These coordinates are then used as 2D pose features. The calculation process is as follows:
[0148]
[0149] in, To return the maximum value of x and y, a heat map of the i-th key point, an abscissa of the i-th maximum response point, an ordinate of the i-th maximum response point, x is an abscissa of a coordinate point, y is an ordinate of the coordinate point, and i is an index of the coordinate point.
[0150] S22, spatio-temporal convolution feature extraction is performed on the target worker image associated worker image frame sequence through a three-dimensional convolutional neural network to obtain corresponding spatio-temporal convolution features.
[0151] It should be noted that the worker image frame sequence refers to a collection of worker images arranged in chronological order. Its core role is to provide time series visual data support for subsequent spatio-temporal convolution feature extraction and dangerous action prediction.
[0152] In the embodiment of the application, spatio-temporal convolution feature extraction is performed on the target worker image associated worker image frame sequence through a 3DCNN network (i.e., a three-dimensional convolutional neural network), the dynamic changes of the worker's body movements are captured, and corresponding spatio-temporal convolution features are obtained.
[0153] S23, a two-dimensional pose feature and a spatio-temporal convolution feature are spliced through a feature fusion layer to obtain a corresponding splicing vector.
[0154] In the embodiment of the application, as shown in Figure 4 a two-dimensional pose feature and a spatio-temporal convolution feature are spliced through a feature fusion layer to obtain a corresponding splicing vector.
[0155] S24, action detection is performed on the splicing vector through an action classification network to obtain a corresponding action detection result.
[0156] In the embodiment of the application, action detection is performed on the splicing vector through an action classification network to obtain a corresponding action detection result, wherein the action classification network includes a fully connected layer and a Softmax output layer (i.e., a flexible maximum value layer) connected in sequence. For example, as shown in Figure 4 the splicing vector is mapped to a dimension equal to the number of preset action categories through a fully connected layer to obtain original scores corresponding to each action. The Softmax output layer is used to process each original score to obtain a corresponding current action category probability (i.e., an action detection result).
[0157] It should be noted that the Softmax output layer calculation process is as follows:
[0158]
[0159] wherein, Probability that the current action is determined to be class c, z c Raw score output for class c, z j Raw score output for class j, K is the total number of hazard state classes, j is the action class first index, c is the class second index.
[0160] Step 209, risk assessment is performed according to the equipment list, the post state and the action detection result, and the working behavior monitoring result corresponding to the limited space is obtained.
[0161] Further, step 209 includes the following sub-steps:
[0162] S31, when the action detection result is action abnormality, the action detection result is taken as the working behavior monitoring result corresponding to the limited space.
[0163] In the embodiment of the present application, it is judged whether the action detection result is action abnormality. When the action detection result is action abnormality, it indicates that the worker has dangerous action, and the action detection result is taken as the working behavior monitoring result corresponding to the limited space.
[0164] S32, when the action detection result is normal working, the equipment list is matched with the preset standard equipment list.
[0165] The standard equipment list refers to the equipment list pre-configured according to the current working specification.
[0166] In the embodiment of the present application, when the action detection result is normal working, the equipment list is matched with the preset standard equipment list to judge whether the equipment list is adapted to the preset standard equipment list.
[0167] S33, when the equipment list is not adapted to the standard equipment list, the equipment wearing violation is taken as the working behavior monitoring result corresponding to the limited space.
[0168] In the embodiment of the present application, when the types and quantities of the equipment list and the standard equipment list are inconsistent, it indicates that the current safety equipment wearing condition does not conform to the regulation, and the safety equipment name of the equipment wearing violation, the missing and the incorrect wearing is taken as the working behavior monitoring result corresponding to the limited space.
[0169] S34, when the equipment list is adapted to the standard equipment list, it is judged whether the post state is on-duty state.
[0170] In the embodiment of the present application, when the types and quantities of the equipment list and the standard equipment list are consistent, it indicates that the current safety equipment wearing condition conforms to the regulation, and it is judged whether the post state is on-duty state.
[0171] S35, if the post state is an on-duty state, the operation normal is taken as the operation behavior monitoring result corresponding to the limited space.
[0172] In the embodiment of the application, when the post state is an on-duty state, the operation normal is taken as the operation behavior monitoring result corresponding to the limited space.
[0173] S36, if the post state is an off-duty state, the operation abnormal is taken as the operation behavior monitoring result corresponding to the limited space.
[0174] In the embodiment of the application, when the post state is an off-duty state, it indicates that the supervisor is off-duty, and the operation abnormal is taken as the operation behavior monitoring result corresponding to the limited space.
[0175] It is worth mentioning that when the monitoring result is not operation normal, the monitoring result can be input into a preset emergency situation accident case library, a historical case corresponding to the monitoring result is matched, and environmental parameters (the environmental parameters include oxygen concentration, toxic gas concentration, flammable gas concentration, temperature, humidity, light intensity, ventilation state, limited space entrance / exit state, obstacle distribution, operation area, and operation duration) of the limited space are obtained. A case evaluation method based on an entropy weight method and an approximation ideal solution sorting method is used to analyze the importance weight of the monitoring result and the environmental parameters in this decision, and the closeness degree score of each historical case and the current output monitoring result is calculated, and the historical case with the highest closeness degree score is automatically selected as the best case. The related emergency auxiliary disposal procedure is extracted from the best case as the emergency auxiliary decision of the limited space, and the emergency auxiliary disposal procedure includes specific response steps, contact objects, and matters needing attention.
[0176] It is worth mentioning that the video picture in the limited space and the operation behavior monitoring result corresponding to the limited space can be played in real time through a preset comprehensive information display platform of the supervisor. The target safety equipment image, the target supervisor image, and the target operation personnel image are projected on the main display area of the monitoring interface of the comprehensive information display platform, thereby providing the supervisor with basic visual perception of the limited space operation site. At the same time, the graphic information associated with the monitoring result can be superimposed and rendered on the target safety equipment image, the target supervisor image, and the target operation personnel image, and the corresponding emergency auxiliary decision is presented in the form of structured text, and the graphic information includes a highlight box indicating missing safety equipment, a red line box reminding the supervisor to be off-duty, and a motion trajectory line indicating the falling of the operation personnel.
[0177] In the embodiment of the present application, the safety equipment image, the supervisor image and the worker image are respectively preprocessed by the preset low-illumination enhancement model to obtain a target safety equipment image, a target supervisor image and a target worker image, the equipment detection model is used for detecting the equipment in the target safety equipment image to obtain a corresponding equipment list, the face target detection model is used for detecting the post state of the target supervisor image to obtain a corresponding post state, the action detection model is used for detecting the dangerous action of the target worker image to obtain a corresponding action detection result, and the risk assessment is performed according to the equipment list, the post state and the action detection result to obtain the operation behavior monitoring result of the limited space. The technical problem that the prior art mainly identifies the abnormal operation behavior in the limited space through video playback or simple algorithm, but it is difficult to accurately identify the abnormal operation behavior in the limited space in a low-illumination environment, and the reliability of the operation in the limited space is reduced. Compared with the traditional limited space operation behavior monitoring method, the safety equipment image, the supervisor image and the worker image are respectively preprocessed by the preset low-illumination enhancement model to obtain a target safety equipment image, a target supervisor image and a target worker image, so that the cleaning degree of the collected image in the low-illumination environment is improved, and the risk assessment of the limited space operation is performed in combination with the equipment list, the post state and the action detection result, so that the accuracy of the identification of the abnormal operation behavior in the limited space is ensured, and the reliability of the operation in the limited space is improved.
[0178] Please refer to Figure 5 , Figure 5 The structure block diagram of a limited space operation behavior monitoring system provided in the third embodiment of the present application is shown in FIG. 3.
[0179] The limited space operation behavior monitoring system provided in the present application comprises:
[0180] The acquisition module 301 is configured to acquire the safety equipment image, the supervisor image and the worker image of the limited space, and pre-process the safety equipment image, the supervisor image and the worker image by the preset low-illumination enhancement model to obtain a target safety equipment image, a target supervisor image and a target worker image.
[0181] The equipment detection module 302 is configured to detect the equipment in the target safety equipment image by the pre-trained equipment detection model to obtain a corresponding equipment list.
[0182] The post state detection module 303 is configured to detect the post state of the target supervisor image by the pre-trained face target detection model to obtain a corresponding post state.
[0183] The action detection module 304 is configured to perform dangerous action detection on the target worker image by using a pre-trained action detection model, and obtain an action detection result.
[0184] The evaluation module 305 is configured to perform risk evaluation according to the equipment list, the post state and the action detection result, and obtain an operation behavior monitoring result corresponding to the limited space.
[0185] Further, the low-illumination enhancement model comprises a decomposition network, an enhancement network, a non-local mean filtering module and a channel-wise multiplication layer, and a processing process of the low-illumination enhancement model is specifically as follows:
[0186] The decomposition network is configured to perform image decomposition on the input limited-space operation image, and obtain a corresponding reflection image and an illumination image, wherein the decomposition network comprises five first convolutional layers and first activation function layers connected in sequence.
[0187] The enhancement network is configured to perform image enhancement on the illumination image, and obtain a corresponding target illumination image.
[0188] The non-local mean filtering module is configured to perform denoising on the reflection image, and obtain a corresponding target reflection image.
[0189] The channel-wise multiplication layer is configured to perform pixel-by-pixel multiplication on the target illumination image and the target reflection image, and obtain a target limited-space operation image, wherein the target limited-space operation image comprises a target safety equipment image, a target supervisor image or a target worker image.
[0190] Further, the equipment detection model comprises a backbone network, a three-way feature pyramid network and a detection head, and the equipment detection module 302 comprises:
[0191] The feature extraction submodule is configured to perform feature extraction on the target safety equipment image by using the backbone network, and output a plurality of equipment feature maps layer by layer.
[0192] The feature fusion submodule is configured to perform feature fusion on the plurality of equipment feature maps by using the three-way feature pyramid network, and generate a plurality of equipment fusion feature maps.
[0193] The equipment recognition submodule is configured to perform equipment recognition on each equipment fusion feature map by using the detection head, and obtain a plurality of equipment prediction regions.
[0194] The first analysis submodule is configured to select an equipment prediction region with the maximum confidence as a reference region from the plurality of equipment prediction regions, and calculate an intersection over union between the reference region and each equipment prediction region.
[0195] The second analysis submodule is configured to determine whether each intersection over union is less than or equal to a preset intersection over union threshold.
[0196] When the intersection over union is less than or equal to the intersection over union threshold, the equipment prediction region associated with the intersection over union is determined as the target equipment prediction region;
[0197] The equipment data of the reference region and the equipment data of the target equipment prediction region are used to generate a corresponding equipment list.
[0198] Further, the post state detection module 303 comprises:
[0199] The third analysis submodule is configured to use a pre-trained face target detection model to perform feature extraction on the target supervisor image, to obtain a corresponding identity feature vector;
[0200] The fourth analysis submodule is configured to calculate the cosine similarity between the identity feature vector and each standard identity feature vector in the preset identity database, respectively;
[0201] When each cosine similarity is less than or equal to the preset identity threshold, the preset authentication number is updated, and the step of obtaining the safety equipment image, the supervisor image and the worker image of the limited space is executed;
[0202] When the authentication number is greater than or equal to the preset frame number threshold, the off-duty state is determined as the corresponding post state;
[0203] When any cosine similarity is greater than the identity threshold, the on-duty state is determined as the corresponding post state.
[0204] Further, the action detection model comprises a high-resolution network, a three-dimensional convolutional neural network, a feature fusion layer and an action classification network, and the action detection module 304 comprises:
[0205] The two-dimensional pose feature submodule is configured to perform two-dimensional pose feature extraction on the target worker image through the high-resolution network, to obtain a corresponding two-dimensional pose feature;
[0206] The spatio-temporal convolution feature submodule is configured to perform spatio-temporal convolution feature extraction on the worker image frame sequence associated with the target worker image through the three-dimensional convolutional neural network, to obtain a corresponding spatio-temporal convolution feature;
[0207] The splicing submodule is configured to perform splicing operation on the two-dimensional pose feature and the spatio-temporal convolution feature through the feature fusion layer, to obtain a corresponding splicing vector;
[0208] The action detection submodule is configured to perform action detection on the splicing vector through the action classification network, to obtain a corresponding action detection result.
[0209] Further, the evaluation module 305 comprises:
[0210] a fifth analysis submodule, configured to, when the action detection result is an abnormal action, take the action detection result as the working behavior monitoring result corresponding to the limited space;
[0211] when the action detection result is a normal action, match the equipment list with a preset standard equipment list;
[0212] a sixth analysis submodule, configured to, when the equipment list is not matched with the standard equipment list, take equipment wearing violation as the working behavior monitoring result corresponding to the limited space;
[0213] a seventh analysis submodule, configured to, when the equipment list is matched with the standard equipment list, judge whether the post state is an on-duty state;
[0214] if the post state is the on-duty state, take normal working as the working behavior monitoring result corresponding to the limited space;
[0215] if the post state is an off-duty state, take abnormal working as the working behavior monitoring result corresponding to the limited space.
[0216] Please refer to Figure 6 , Figure 6 a structural block diagram of an electronic device provided in the fourth embodiment of the present application.
[0217] The electronic device in the embodiment of the present application comprises a memory 401 and a processor 402, the memory 401 stores a computer program, and the computer program is executed by the processor 402 to make the processor 402 execute the limited space working behavior monitoring method in any of the above embodiments.
[0218] The memory 401 can be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk, or a ROM. The memory 401 has a storage space 403 for program codes 413 for performing any of the method steps in the above-described methods. For example, the storage space 403 for program codes can include individual program codes 413 for implementing various steps in the above-described methods, respectively. These program codes can be read from or written to one or more computer program products. These computer program products include program code carriers such as a hard disk, a compact disc (CD), a memory card, or a floppy disk. The program codes can be compressed in a suitable form, for example. These codes, when run by a computing processing device, cause the computing processing device to perform the individual steps in the above-described methods. These program codes can be read from or written to one or more computer program products. These computer program products include program code carriers such as a hard disk, a compact disc (CD), a memory card, or a floppy disk. The program codes can be compressed in a suitable form, for example. These codes, when run by a computing processing device, cause the computing processing device to perform the individual steps in the above-described limited space work behavior monitoring method.
[0219] The embodiment five of the present application further provides a computer readable storage medium, having stored thereon a computer program, which, when executed by a processor, implements the limited space work behavior monitoring method according to any one of the above-described embodiments.
[0220] The embodiment six of the present application further provides a computer program product, which includes a computer program stored on a non-transitory computer readable storage medium, and the computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer performs the limited space work behavior monitoring method according to any one of the above-described embodiments.
[0221] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described herein.
[0222] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the division of the units is only a logical function division, and there can be another division manner for the actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different units, can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or in other forms.
[0223] The units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments of the present application.
[0224] In addition, each functional unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically as a separate unit, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware, or in the form of a software functional unit.
[0225] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solutions of the present application essentially, or the part that contributes to the prior art, or all or a part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, and various media that can store program codes.
[0226] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method of monitoring behavior in a confined space operation, characterized by, The method comprises the following steps: obtaining a safety equipment image, a supervisor image and a worker image of a limited space, and performing image preprocessing on the safety equipment image, the supervisor image and the worker image respectively through a preset low-illumination enhancement model to obtain a target safety equipment image, a target supervisor image and a target worker image; performing equipment detection on the target safety equipment image through a pre-trained equipment detection model to obtain a corresponding equipment list; performing post state detection on the target supervisor image through a pre-trained face target detection model to obtain a corresponding post state; performing dangerous action detection on the target worker image through a pre-trained action detection model to obtain a corresponding action detection result; performing risk assessment according to the equipment list, the post state and the action detection result to obtain a corresponding work behavior monitoring result of the limited space.
2. The method of claim 1, wherein, The low-illumination enhancement model comprises a decomposition network, an enhancement network, a non-local mean filtering module and a channel-wise multiplication layer, and the processing process of the low-illumination enhancement model is specifically as follows: performing image decomposition operation on the input limited space work image through the decomposition network to obtain a corresponding reflection image and an illumination image, wherein the decomposition network comprises five first convolution layers and a first activation function layer connected in sequence; performing image enhancement processing on the illumination image through the enhancement network to obtain a corresponding target illumination image; performing denoising processing on the reflection image through the non-local mean filtering module to obtain a corresponding target reflection image; performing pixel-by-pixel multiplication processing on the target illumination image and the target reflection image through the channel-wise multiplication layer to obtain a target limited space work image, wherein the target limited space work image comprises a target safety equipment image, a target supervisor image or a target worker image.
3. The method of claim 1, wherein, The equipment detection model comprises a backbone network, a three-way feature pyramid network and a detection head, and the step of performing equipment detection on the target safety equipment image through the pre-trained equipment detection model to obtain a corresponding equipment list comprises: extracting features of the target safety equipment image through the backbone network to output a plurality of equipment feature maps layer by layer; performing feature fusion on the plurality of equipment feature maps through the three-way feature pyramid network to generate a plurality of equipment fusion feature maps; performing equipment recognition on each equipment fusion feature map through the detection head to obtain a plurality of equipment prediction regions; selecting the equipment prediction region with the maximum confidence as a reference region, and calculating the intersection over union between the reference region and each equipment prediction region; determining whether each intersection over union is less than or equal to a preset intersection over union threshold; when the intersection over union is less than or equal to the intersection over union threshold, determining the equipment prediction region associated with the intersection over union as a target equipment prediction region; generating a corresponding equipment list by using the equipment data of the reference region and the equipment data of the target equipment prediction region.
4. The method of claim 1, wherein, The step of detecting the post state of the target supervisor image by the pre-trained face target detection model comprises: The pre-trained face target detection model is used to extract features of the target supervisor image to obtain a corresponding identity feature vector; The cosine similarity between the identity feature vector and each standard identity feature vector in the preset identity database is calculated respectively; When each cosine similarity is less than or equal to a preset identity threshold, the preset authentication number is updated, and the step of acquiring the safety equipment image, the supervisor image and the worker image of the limited space is executed; When the authentication number is greater than or equal to a preset frame number threshold, it is determined that the post state is an off-duty state; When any cosine similarity is greater than the identity threshold, it is determined that the post state is an on-duty state.
5. The method of claim 1, wherein, The action detection model comprises a high-resolution network, a three-dimensional convolutional neural network, a feature fusion layer and an action classification network. The step of detecting the dangerous action of the target worker image by the pre-trained action detection model to obtain a corresponding action detection result comprises: The high-resolution network is used to extract two-dimensional posture features of the target worker image to obtain corresponding two-dimensional posture features; The three-dimensional convolutional neural network is used to extract spatio-temporal convolution features of a worker image frame sequence associated with the target worker image to obtain corresponding spatio-temporal convolution features; The feature fusion layer is used to perform a splicing operation on the two-dimensional posture features and the spatio-temporal convolution features to obtain a splicing vector; The action classification network is used to detect the splicing vector to obtain a corresponding action detection result.
6. The method of confined space work behavior monitoring according to any one of claims 1-5, wherein, The step of performing risk assessment according to the equipment list, the post state and the action detection result to obtain a corresponding work behavior monitoring result of the limited space comprises: When the action detection result is an action anomaly, the action detection result is taken as the work behavior monitoring result of the limited space; When the action detection result is normal work, the equipment list is matched with a preset standard equipment list; When the equipment list is not adapted to the standard equipment list, equipment wearing violation is taken as the work behavior monitoring result of the limited space; When the equipment list is adapted to the standard equipment list, it is determined whether the post state is an on-duty state; If the post state is an on-duty state, normal work is taken as the work behavior monitoring result of the limited space; If the post state is an off-duty state, abnormal work is taken as the work behavior monitoring result of the limited space.
7. A confined space work behavior monitoring system characterized by, The system comprises: A collection module is configured to acquire safety equipment images, supervisor images and worker images of a limited space, and perform image preprocessing on the safety equipment images, the supervisor images and the worker images by a preset low-illumination enhancement model to obtain target safety equipment images, target supervisor images and target worker images. The equipment detection module is configured to perform equipment detection on the target safety equipment image by using a pre-trained equipment detection model to obtain a corresponding equipment list. The post state detection module is configured to perform post state detection on the target supervisor image by using a pre-trained face target detection model to obtain a corresponding post state. The action detection module is configured to perform dangerous action detection on the target worker image by using a pre-trained action detection model to obtain a corresponding action detection result. The evaluation module is configured to perform risk evaluation according to the equipment list, the post state, and the action detection result to obtain a corresponding work behavior monitoring result of the limited space.
8. An electronic device, comprising: The computer program is executed by the processor to cause the processor to perform the steps of the limited space work behavior monitoring method according to any one of claims 1-6.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed to implement the limited space work behavior monitoring method according to any one of claims 1-6.
10. A computer program product, characterised in that, The computer program product comprises a computer program stored on a non-transitory computer-readable storage medium, and the computer program comprises program instructions, wherein when the program instructions are executed by a computer, the computer performs the limited space work behavior monitoring method according to any one of claims 1-6.