Bucket arm vehicle personnel dangerous posture recognition method based on multi-source sensor fusion, electronic equipment and storage medium

By using multi-source sensor fusion technology, combining RGB video, depth images, and IMU signals, the system can intelligently identify and provide real-time warnings of dangerous postures of workers operating at heights in bucket trucks. This solves the problems of insufficient manual supervision and delayed response in traditional methods, and improves the safety assurance capabilities of high-altitude power operations.

CN120976837AActive Publication Date: 2025-11-18STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO

Patent Information

Application Number
CN202511491975.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2025-11-18
Estimated Expiration
2045-10-20

AI Technical Summary

Technical Problem

Existing methods for safety supervision of high-altitude power operations rely on manual supervision, which is subject to strong subjectivity, inability to quantify safe distances, and delayed response, making it difficult to achieve intelligent identification and real-time early warning of dangerous behaviors of personnel working at height on bucket trucks.

Method used

By employing a multi-source sensor fusion method, combining RGB video, depth images, IMU signals, and human posture key point sequences, and through temporal alignment, spatial alignment, and feature extraction, and utilizing Cross-Attention interaction and adaptive weight fusion technology, high-precision identification and real-time early warning of dangerous postures of personnel in bucket trucks are achieved.

Benefits of technology

It significantly improves the accuracy and real-time performance of hazardous behavior identification, reduces false alarm and false alarm rates, maintains stable identification performance in complex environments, and enables dynamic perception and early warning for high-altitude operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976837A_ABST
    Figure CN120976837A_ABST
Patent Text Reader

Abstract

The invention discloses a bucket arm vehicle personnel dangerous posture recognition method based on multi-source sensor fusion, electronic equipment and a storage medium. The method comprises the following steps: acquiring multi-modal data, including RGB (Red, Green, Blue) videos, depth images and IMU (Inertial Measurement Unit) signals acquired in the working process of personnel on the bucket arm vehicle, and a human body posture key point sequence extracted from the RGB videos; preprocessing and feature extraction are carried out on the multi-modal data to obtain multi-modal features, and preprocessing comprises time sequence alignment, space alignment and data denoising; rGB video features are used as a main mode, Cross-Attention interaction is carried out on the RGB video features and features of the depth image, the IMU signal and the human body posture key point sequence, and the RGB video features are enhanced; calculating the self-adaptive weight of each modal feature, and fusing the features of the four modals; and based on the fused features, identifying the posture category of the personnel on the bucket arm vehicle. According to the invention, high-precision identification and real-time early warning of dangerous behaviors of the boom personnel can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power distribution network non-power operation safety monitoring and personnel posture intelligent identification, and particularly relates to a dangerous posture identification method for personnel on a boom truck based on multi-source sensor fusion, an electronic device and a storage medium. BACKGROUND

[0002] In the daily maintenance of the power industry, especially in the maintenance operation of transmission lines, the operating personnel generally need to perform high-altitude operation on the conductor or the tower. At present, the climbing tool that can be used for high-altitude live maintenance operation is mainly an insulating boom truck.

[0003] The high-altitude operation of the power grid, especially the live operation based on the insulating boom truck, is one of the work with the highest technical content and the largest risk coefficient in the power operation and maintenance. The traditional safety supervision highly depends on the naked-eye patrol and supervision of safety officers, and has problems such as strong subjectivity, inability to quantify the safety distance, and reaction lag. How to intelligently identify and warn the dangerous behavior of the high-altitude operating personnel on the boom truck has become the focus of the academic and industrial circles.

[0004] The statements herein only provide background technology related to the present application, and do not necessarily constitute the prior art. SUMMARY

[0005] The purpose of the present application is to provide a dangerous posture identification method for personnel on a boom truck based on multi-source sensor fusion, an electronic device and a storage medium, to realize high-precision identification and real-time warning of the dangerous behavior of the high-altitude operating personnel on the boom truck, and thus effectively improve the safety guarantee capability in the high-altitude operation scene of the power grid.

[0006] In order to achieve the above purpose, the present application realizes the following technical scheme: One aspect of the present application provides a dangerous posture identification method for personnel on a boom truck based on multi-source sensor fusion, comprising: acquiring multi-modal data, including RGB video, depth image, IMU signal collected in the working process of the personnel on the boom truck, and a human posture key point sequence extracted from the RGB video; respectively pre-processing and feature extracting the multi-modal data to obtain multi-modal features, the pre-processing including time alignment, space alignment and data denoising; taking the RGB video feature as the main mode, respectively performing Cross-Attention interaction with the features of the depth image, the IMU signal and the human posture key point sequence to enhance the RGB video feature; calculating the adaptive weight of each modal feature, and fusing the features of the four modes; based on the fused features, identifying the posture category of the personnel on the boom truck.

[0007] Optionally, the multi-modal data is time-aligned in the following manner: Taking the modality data with the highest sampling frequency as the reference, time-interpolation is used to supplement the missing timestamp values for the modality data with lower sampling frequency; The four kinds of multi-modal data are synchronously sliced in the same time period through a sliding window.

[0008] Optionally, the multi-modal data is spatially aligned in the following manner: For each modality data, the data in different coordinate systems is converted to a unified reference system through the corresponding homogeneous transformation matrix; All modality data is aligned to the body reference system with the target pose key point as the origin through pose estimation.

[0009] Optionally, the multi-modal data is respectively feature-extracted, including: A video spatio-temporal modeling network model is used to extract features from the RGB video; A three-dimensional convolution expansion network model is used to extract features from the depth image; A bidirectional recurrent network model is used to extract features from the IMU signal; A spatio-temporal graph convolution network model is used to extract features from the human pose key point sequence.

[0010] Optionally, before enhancing the RGB video features, it further includes: A fully connected layer is used to uniformly map the features of all modalities to the same dimension.

[0011] Optionally, the adaptive weight of each modality feature is calculated, including: The adaptive weight of each modality feature is calculated through a fully connected layer, and the adaptive weight value is in the range of [0, 1] through a Sigmoid activation function.

[0012] Optionally, the features of the four modalities are fused, including: The following fusion formula is used to calculate the fused features: wherein, is the fused feature, and is the element-wise product, is the feature of the enhanced RGB image, is the feature of the depth image, is the feature of the IMU signal, is the feature of the human pose key point sequence, , is the adaptive weight tensor of the corresponding modality feature.

[0013] Optionally, the posture category of the person on the arm truck is identified based on the fused features, including: The fused features are flattened or pooled in sequence dimensions; The flattened features are input into a fully connected network to determine the probability of the feature mapping to each posture category, thereby identifying the posture category of the person on the arm truck, the fully connected network including two fully connected layers, the activation function of the first layer being ReLU, and the activation function of the second layer being Softmax.

[0014] In a second aspect, the present application also provides an electronic device, including a processor and a memory, the memory storing a computer program, the computer program being executed by the processor to implement the method as described above.

[0015] In a third aspect, the present application also provides a readable storage medium, the readable storage medium storing a computer program, the computer program being executed by a processor to implement the method as described above.

[0016] The present application has at least the following technical effects: In summary, the method for recognizing dangerous posture of person on arm truck based on multi-source sensor fusion provided by the present application can significantly improve the accuracy and real-time performance of dangerous behavior recognition by combining RGB video, depth image, IMU signal and human posture key point sequence, solve the problem of spatio-temporal heterogeneity of multi-source vision and sensor data through cross-modal spatio-temporal alignment processing, realize effective fusion and joint modeling of cross-modal features, realize multi-modal complementation through Cross-Attention interaction method, maintain stable recognition performance relying on other modalities in the case of video occlusion, IMU signal loss or noise interference, thereby effectively reducing the false positive rate and the false negative rate, and through the adaptive weight feature fusion method, the dynamic change trend of dangerous behavior can be captured to realize dynamic perception and early warning of high-altitude operation dangerous behavior. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed to be used in the description. Obviously, the drawings in the following description are one embodiment of the present application, and other drawings can also be obtained by those skilled in the art without creative labor: Figure 1 The flowchart of the method for recognizing dangerous posture of person on arm truck based on multi-source sensor fusion provided by an embodiment of the present application is shown in the figure; Figure 2 The overall architecture of the present application is shown in the figure. DETAILED DESCRIPTION

[0018] The scheme proposed by the present application is further described in detail below in combination with the drawings and specific embodiments. The advantages and features of the present application will be clearer according to the following description. It should be noted that the drawings are greatly simplified and all use non-precise proportions, only for the purpose of facilitating and clarifying the purpose of assisting in the description of the embodiments of the present application. In order to make the purpose, features and advantages of the present application more obvious and easy to understand, please refer to the drawings. It should be noted that the structure, proportion, size, etc. shown in the drawings attached to the present specification are only used to cooperate with the content disclosed in the specification, so that those skilled in the art can understand and read, and are not used to limit the conditions for implementing the present application, therefore, any modification of structure, change of proportion relationship or adjustment of size, which does not affect the effect and purpose that can be achieved by the present application, should still fall within the scope of the technical content disclosed by the present application.

[0019] In the prior art, the dangerous behavior recognition method for the operation of the boom truck in the aerial operation of the power grid mainly includes the following types: a method based on traditional machine learning, a method based on single computer vision, a method based on deep learning, and a method based on multi-source sensor fusion. However, the existing recognition methods have the following problems: 1) the multi-source vision and sensor data have space-time heterogeneity, and the sampling rate, delay and coordinate system difference are not effectively aligned, which limits the joint modeling capability of cross-modal features; 2) dangerous posture recognition mainly depends on single-frame images or local time series fragments, and lacks modeling of continuous evolution of actions, resulting in insufficient early warning capability; 3) in a complex operation environment, factors such as light changes, shaking, electromagnetic interference, etc. easily cause noise or loss of video and sensor signals, and the existing methods lack robust data compensation and redundancy mechanisms.

[0020] In view of the above problems existing in the prior art, the present application provides a dangerous posture recognition method for personnel of a boom truck based on multi-source sensor fusion, which provides efficient and accurate dangerous posture detection by combining RGB video, depth image, inertial measurement unit (IMU) signal and posture key point sequence.

[0021] In combination with the method shown in Figure 1 , Figure 2 , the present embodiment provides a dangerous posture recognition method for personnel of a boom truck based on multi-source sensor fusion, which includes the following steps: Step S1, acquiring multi-modal data, the multi-modal data including: RGB video, depth image, IMU signal collected during the working process of personnel on the boom truck, and human posture key point sequence extracted from the RGB video.

[0022] Generally, a camera and a depth sensor are installed on the fence of a boom truck, and a high-altitude worker needs to wear an IMU (Inertial Measurement Unit) device on key parts of the body before work. During high-altitude work, the camera captures color video of the human body and the environment (RGB video), usually collecting 30 frames of image data per second, and the image data is stored in RGB format. At the same time, a depth sensor (such as Kinect) measures the distance by infrared light beams to generate a depth image, providing three-dimensional spatial information. Each pixel in the depth image represents the depth value of a certain point in the scene, with a unit of meters, and is collected synchronously with the RGB video. The IMU device measures acceleration and angular velocity in real time through an accelerometer and a gyroscope, with a sampling frequency of usually 100 Hz or higher, capturing X, Y, and Z axis acceleration and angular velocity information.

[0023] In order to ensure the original accuracy of the skeleton data, the two-dimensional or three-dimensional coordinates of each joint of the human body are extracted from the RGB video before subsequent processing, obtaining a sequence of human posture key points. The OpenPose algorithm can be used for extraction, collecting 30 frames of data per second, thereby providing skeleton model information of the human body.

[0024] The above four kinds of multi-modal data can provide comprehensive support for subsequent dangerous posture recognition, improving recognition accuracy and system robustness.

[0025] In step S2, multi-modal data is pre-processed and feature extracted to obtain multi-modal features. Pre-processing includes time alignment, spatial alignment, and data denoising.

[0026] It can be understood that the data collected by multi-source sensors often has significant differences in time and space, and this data heterogeneity brings great challenges to subsequent feature modeling and system recognition. Therefore, in order to solve the heterogeneity problem between multi-source data, the embodiment performs time and spatial alignment operations when pre-processing multi-source data, as follows: First, the following method is used for time alignment of multi-modal data: taking the modality data with the highest sampling frequency as the reference, and for the modality data with low sampling frequency, time interpolation is used to supplement the missing timestamp value; through the sliding window method, the four kinds of multi-modal data are synchronously sliced in the same time period.

[0027] Suppose the sampling points of the modality are , where is not necessarily continuous, and the goal of time interpolation is to predict the value at the new time point from the existing data points. Linear interpolation can be used, which predicts the intermediate time point by the rate of change of the previous and next data points.

[0028] For time series data, the four modalities of data are sliced by sliding window to ensure that the data in each window can be processed as a time series. By setting the window size as w and the step size as s , the data set D of each time period can be represented as: , where is the start time of the window, is all the data in the i th window. The four modalities of data contained in each window will be sliced synchronously in the same time period to facilitate subsequent feature extraction and fusion. By sliding window slicing, long time series data can be effectively segmented, so that the data in each time period can be processed independently, thereby improving recognition accuracy and system efficiency.

[0029] Then, the multi-modal data is spatially aligned in the following way: for each modality of data, the data in different coordinate systems is converted to a unified reference system by the corresponding homogeneous transformation matrix; all modalities of data are aligned to the body reference system with the target pose key point as the origin by pose estimation.

[0030] Let the data point of a certain modality be in the original coordinate system, and the new coordinate point after conversion to the target coordinate system can be converted by the following homogeneous transformation matrix : , where T is a 4x4 homogeneous transformation matrix containing translation, rotation, scaling and other transformation information, which can map data from the original coordinate system to the unified reference system. The specific value of the homogeneous transformation matrix can be calculated by the calibration process or the known rotation and translation relationship between coordinate systems.

[0031] All data is aligned to the body reference system with a specific pose key point (e.g. pelvis) as the origin by pose estimation, ensuring spatial consistency between different modalities. This facilitates accurate spatial comparison and feature fusion between different modalities. Although the pelvis is usually used as the reference point, in specific applications, key points such as shoulders, spine or head can also be selected.

[0032] In addition, data preprocessing also includes denoising the data. It can be understood that during the sensor acquisition process, data is often affected by noise, interference and other problems, especially video images may be affected by occlusion, blur and other problems, IMU signals may be affected by vibration or other external factors, and depth images may also produce random noise and holes. In order to improve the quality of the data, the above three kinds of data are denoised and enhanced.

[0033] Specifically, for RGB video, blurring and denoising methods can be used to remove noise and details from the image and smooth the video data. Gaussian filtering or mean filtering is typically used for image blurring.

[0034] For depth image denoising, bilateral filtering and interpolation inpainting methods can be used to address random noise, holes, and edge artifacts that may appear in the depth image. This method can remove noise while preserving object edge features and filling in missing areas, thereby improving the integrity and reliability of the depth data.

[0035] For IMU signal denoising, low-pass filtering and signal normalization methods can be used to process the IMU signal. Low-pass filtering can effectively remove high-frequency noise while retaining the useful information in the IMU signal.

[0036] After preprocessing the multimodal data, feature extraction is performed. In this embodiment, different deep learning networks are used to independently extract effective features of the multimodal data in order to give full play to the advantages of each model on different data types.

[0037] Specifically, RGB video is an important modality in dangerous pose recognition, as its spatiotemporal characteristics contain dynamic information about human movement and scene changes. To efficiently extract spatiotemporal features from RGB video, this embodiment employs a video spatiotemporal modeling network model (such as the Video Swin Transformer or TimeSformer model) to extract RGB video features. Taking the Video Swin Transformer model as an example, the input data format is (batch_size, time_steps, height, width, channels), and the input data is a sequence of video frames from an RGB video. During feature extraction, video frames are input into the Video Swin Transformer model and divided into local image windows. Spatial features are extracted through local window self-attention, and inter-frame temporal relationships are modeled by cross-window self-attention. Finally, the output data format is (batch_size, time_steps, 768), i.e. This is output as an RGB video feature.

[0038] The depth image provides spatial geometric information complementary to the RGB image, which can effectively assist in identifying dangerous postures. In order to extract the spatial geometric features of the depth image, the embodiment adopts a three-dimensional convolution expansion network model (such as I3D or Depth-Transformer model) to extract features from the depth image. Taking the I3D model as an example, the input data format is (batch_size, time_steps, height, width, channels), and the input data is a video frame sequence of the depth image . During the feature extraction process, the video frame is input into the I3D model, and the spatial geometric information is modeled by three-dimensional convolution. The I3D network can capture the dynamic changes in the depth image and extract spatial structure features related to human posture. Finally, the output data format is (batch_size, time_steps, 1024), that is , as the depth image feature output.

[0039] The IMU signal contains time series data such as acceleration and angular velocity of human motion, which can provide dynamic information for posture recognition. In order to extract the time series features of the IMU signal, the embodiment adopts a bidirectional recurrent network model (such as BiLSTM or TCN) to extract features from the IMU signal. Taking the BiLSTM model as an example, the input data format is (batch_size, time_steps, num_features), and the input data is an IMU signal sequence (acceleration, angular velocity, etc.). During the feature extraction process, after the IMU signal is input into the BiLSTM network, the BiLSTM can capture the dynamic change rule according to the information before and after the time series, and extract the time series features related to the dangerous posture. Finally, the output data format is (batch_size, time_steps, 64), that is , as the IMU signal output.

[0040] The human posture key point sequence provides two-dimensional coordinate information of each joint position of the human body, which is an important structural feature in dangerous posture recognition. In order to effectively capture the spatial structure features of the human skeleton, the embodiment adopts a spatio-temporal graph convolution network model (such as ST-GCN model) to extract features from the human posture key point sequence. Taking the ST-GCN model as an example, the input data format is (batch_size, time_steps, num_keypoints * 2), and the input data is a human posture key point sequence. During the feature extraction process, the ST-GCN extracts the structure features of the human skeleton and models the spatio-temporal relationship between the key points. Finally, the output data format is (batch_size, time_steps, 256), that is , as the key point feature output, which is expressed by the formula: wherein, is the feature obtained by spatial convolution (spatial feature at time t), and is the weight matrix of the time convolution operation, used to process the features in the time dimension, is the final spatio-temporal feature, combining time and space information.

[0041] Step S3, taking the RGB video feature as the main modality, respectively interacts with the features of the depth image, IMU signal and human pose key point sequence through Cross-Attention to enhance the RGB video feature.

[0042] It should be noted that, since the features extracted from different modal data have different dimensions, the feature dimensions need to be aligned first to ensure that the subsequent fusion operation can be performed in the same dimension. Specifically, all modal features can be uniformly mapped to the same dimension (e.g. 728 dimensions) using a fully connected layer (MLP), as shown in the following formula: wherein, , , respectively represent the RGB video feature, the depth image feature, the IMU signal feature and the human pose key point sequence feature mapped to the same dimension.

[0043] The Cross-Attention mechanism enhances the interaction information between modalities to ensure that the data of different modalities can be efficiently fused in the feature dimension. To reduce the complexity of multi-modal interaction, the embodiment adopts a main modality driven interaction mechanism. Since the RGB video contains the most abundant spatio-temporal information among the four modalities, and is often used as the time reference for multi-modal acquisition, the RGB video feature is selected as the main modality, and respectively interacts with the features of the depth image, IMU signal and human pose key point sequence through Cross-Attention to enhance the RGB video feature.

[0044] Specifically, it can be realized by the following formula: wherein, is a learnable projection matrix, respectively used to generate a query vector , key vector and value vector , is the vector dimension for normalizing dot product results. represents the interaction results of RGB video and modalities , which include depth image , IMU signal and human pose keypoint sequence , , by fusing the interaction results of RGB video and the remaining three modalities (such as splicing or weighted sum), the final is the enhanced RGB feature. It should be noted that the four modality features involved in the formula here refer to the aforementioned features mapped to the same dimension.

[0045] Step S4, calculate the adaptive weight of each modality feature, and fuse the features of the four modalities.

[0046] This embodiment can make full use of the respective advantages of the four modality features (enhanced RGB video, depth image, IMU signal and pose keypoint) by fusing them: RGB video provides dynamic information of human motion and scene changes, depth image supplements spatial geometric features, IMU signal provides acceleration and angular velocity data about motion, and pose keypoint accurately describes the positions of human joints. Traditional single model cannot effectively combine these information.

[0047] This embodiment first calculates the adaptive weight of each modality feature, so as to dynamically adjust the importance of each modality feature in the fusion process, thereby flexibly controlling the contribution of different modalities to the final fused feature, and significantly improving the accuracy and stability of recognition. Specifically, the adaptive weight of each modality feature is calculated through a fully connected layer, and the adaptive weight value is in the range of [0, 1] through a Sigmoid activation function, and the specific formula is as follows: Among them, , are the adaptive weight tensors of RGB video, depth image, IMU signal and human pose keypoint sequence respectively, and the output dimension is consistent with the corresponding modality feature to ensure that point-by-point weighting operation can be performed at the element level.

[0048] Then the four features are fused by adaptive weights, and the fusion formula is: wherein, is the element-wise product, is the fused feature, and the dimension is (batch_size, time_steps, 728), , is the adaptive weight tensor corresponding to the modal feature.

[0049] Step S5, based on the fused feature, the posture category of the personnel on the arm truck is identified.

[0050] In this embodiment, the fused feature is classified, and finally the posture category of the personnel is output. The posture category can include three categories of "normal", "danger warning" and "failure". Among them, the danger warning indicates that the posture has a certain danger, but has not yet occurred failure, and the failure indicates that the posture has entered a dangerous area and needs to be handled urgently. This classification method directly corresponds to the actual application of the dangerous posture recognition system, which can judge whether the posture is abnormal or dangerous in real time, and provide a warning function for the system to improve safety.

[0051] In the identification, the fused feature is first flattened or pooled according to the sequence dimension. Specifically, the time dimension is first pooled to compress the time sequence information to obtain (batch_size, 728), and then mapped by full connection to reduce the dimension to (batch_size, 128) as the input for subsequent classification.

[0052] Then, the flattened feature is input into a full connection network to determine the probability of feature mapping to each posture category, so as to identify the posture category of the personnel on the arm truck. The full connection network includes two full connection layers, the activation function of the first layer is ReLU, and the activation function of the second layer is Softmax.

[0053] Specifically, the flattened feature is input into the first full connection layer FC1, the first layer has 256 neurons, and the activation function adopts ReLU, which can increase the nonlinear ability of the model. The output of the second full connection layer FC2 is mapped to the number of categories C (such as normal, dangerous, failure, etc.), and the activation function adopts Softmax to convert the result into a category probability distribution.

[0054] In summary, the personnel dangerous posture recognition method based on multi-source sensor fusion provided by the application can significantly improve the accuracy and real-time performance of dangerous behavior recognition by combining RGB video, depth image, IMU signal and human body posture key point sequence, solve the spatio-temporal heterogeneity problem of multi-source vision and sensor data through cross-modal spatio-temporal alignment processing, realize effective fusion and joint modeling of cross-modal features, realize multi-modal complementation through the Cross-Attention interaction method, and can maintain stable recognition performance in the case of video occlusion, IMU signal loss or noise interference, thereby effectively reducing the false positive rate and the false negative rate, and the adaptive weight feature fusion method can capture the dynamic change trend of dangerous behavior, and realize dynamic perception and early warning of high-altitude operation dangerous behavior.

[0055] The training and optimization of the model in the embodiment are described below.

[0056] The data source of the data in the application relates to actual multi-source sensor actual capture data and public multi-source sensor related dangerous recognition data set. The data content includes: RGB video collected by an RGB camera (frame rate 30fps), depth image collected by a Kinect depth sensor (frame rate 30fps, synchronized with the RGB camera), IMU signal (acceleration / angular velocity) collected by a wearable IMU device (sampling frequency 100Hz), human body posture key point sequence (frame rate 30fps) extracted from the RGB video. The data label is divided into three categories: “normal”, “abnormal warning” (instability, about to fall, etc.), “fault” (abnormal illegal posture), which is labeled by experts combined with operation records. The sample distribution satisfies: about 50,000 segments; the proportion of “abnormal warning” and “fault” is about 15-25%; at least 5000 segments of each type to reduce extreme imbalance.

[0057] The loss function adopts a multi-class cross-entropy loss, as follows: Wherein, N is the total number of samples, C is the number of categories, is the true label, is the predicted probability.

[0058] The Adam optimizer is used for model training, and the initial learning rate is set to 0.001. Every 10 epochs, the learning rate is decayed to 0.95 times of the original.

[0059] The data set is divided into a training set, a validation set and a test set in a ratio of 7:1:2. The training set is used for model training, the validation set is used for performance evaluation and parameter adjustment during training, and the test set is used for final model evaluation after training. Training is stopped when the validation set loss no longer decreases, even if the training accuracy continues to improve. This method can ensure that the model does not overfit the training set.

[0060] The present application adopts a plurality of mainstream classification performance indicators for evaluation, including accuracy (Accuracy), precision (Precision), recall (Recall), and F1 score (F1-score). Specifically as follows: Recall (Recall) is used to measure how many real samples of each category are successfully identified by the model, reflecting the model's ability to identify minority categories such as anomalies, and the formula is as follows: F1 score (F1-score) is used to comprehensively reflect the accuracy and recall ability of the model, and is suitable for class-imbalance data sets, and the formula is as follows: The method of the present application is compared with other methods, and the comparison results are shown in Table 1, wherein RGB-only is to use only the Video Swin Transformer model to recognize dangerous postures based on RGB video, Depth-only represents to use only the I3D model to recognize dangerous postures based on depth image, IMU-only represents to use only the BiLSTM model to recognize dangerous postures based on IMU signal, Pose-only represents to use only the ST-GCN model to recognize dangerous postures based on human posture key point sequence, Early-Concat+MLP represents to align four kinds of modal data, then average, and then classify through mlp layer, and Late-AvgLogits represents to equally average the logits of four single-mode classifiers.

[0061] Table 1 As can be seen from the above table, the accuracy, precision, recall and F1 value of the method of the present application are significantly better than those of the prior art.

[0062] In some other embodiments, the present application also provides an electronic device comprising a processor and a memory, wherein the memory has a computer program stored thereon, and the computer program is executed by the processor to implement the method as described above.

[0063] In some other embodiments, the present application also provides a readable storage medium, wherein a computer program is stored in the readable storage medium, and the computer program is executed by a processor to implement the method described above. All algorithms in the method of the present embodiment can be completed in the software of the electronic device or the readable storage medium, and are adapted to different application environments.

[0064] In summary, compared with the prior art, the present application has the following technical effects: Traditional multi-modal fusion methods often face the problems of time alignment and spatial reference unification. Especially in the process of multi-source data acquisition, the sampling rate, coordinate system difference and time synchronization of different modalities lead to difficulties in fusion and low recognition accuracy. The present application proposes a time alignment method based on interpolation and sliding window, and combines the homogeneous transformation matrix and the skeleton key points as anchor points to realize spatial normalization, so as to ensure that the RGB, depth, IMU and attitude data are processed under a unified space-time reference. Compared with the traditional alignment method, the present application has significant improvement in accuracy and robustness.

[0065] Existing dangerous posture recognition often relies on single time point state detection, which is difficult to capture the dynamic evolution process of dangerous behavior. The present application introduces Video Swin Transformer, BiLSTM and ST-GCN and other time series modeling methods to realize continuous modeling of multi-modal sequences, combines sliding window and dynamic weight mechanism to perceive the evolution trend of dangerous actions, and sets the "danger warning" intermediate state at the output end, so as to realize forward-looking identification and warning before the danger occurs completely.

[0066] Existing dangerous posture recognition systems often have some missing or disturbed modal data (such as video occlusion, IMU signal loss or noise interference) when processing multi-modal data. The present application takes RGB video features as the main modality, introduces depth image, IMU signal and attitude key point features through Cross-Attention to enhance and correct the RGB features, and combines an adaptive weight mechanism to dynamically adjust the contribution of each modality. Even in the case of video occlusion, the system can still maintain stable recognition performance, significantly improving robustness and reliability.

[0067] It is to be appreciated that the term "comprises", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0068] While the application has been described in detail by reference to preferred embodiments thereof, it is to be understood that the detailed description is not intended to limit the application to the embodiments described herein. Many modifications and variations of the present application will be apparent to those of ordinary skill in the art upon reading this description. It is therefore contemplated to be within the scope of the present application to set forth and claim such modifications and variations of the application that come within the scope of the appended claims.

Claims

1. A method for identifying dangerous postures of personnel in a boom truck based on multi-source sensor fusion, characterized in that, include: Acquire multimodal data, including: RGB video, depth images, IMU signals collected during the work of personnel on the boom truck, and human pose key point sequences extracted from RGB video; Multimodal data are preprocessed and feature extracted separately to obtain multimodal features. The preprocessing includes temporal alignment, spatial alignment and data denoising. Using RGB video features as the primary modality, cross-attention interactions are performed with features from depth images, IMU signals, and human pose keypoint sequences to enhance the RGB video features. Calculate the adaptive weights for each modality feature and fuse the features of the four modalities; Based on the fused features, the posture categories of the personnel on the boom truck are identified.

2. The method for identifying dangerous postures of personnel in a bucket truck based on multi-source sensor fusion as described in claim 1, characterized in that, The following method is used to perform time-series alignment of multimodal data: Using the modal data with the highest sampling frequency as a benchmark, missing timestamp values ​​are supplemented for modal data with low sampling frequency through time interpolation; Four types of multimodal data are synchronously sliced ​​within the same time period using a sliding window method.

3. The method for identifying dangerous postures of personnel in a bucket truck based on multi-source sensor fusion as described in claim 1, characterized in that, Spatial alignment of multimodal data is performed using the following method: For each modal data, the data from different coordinate systems are transformed to a unified reference system using the corresponding homogeneous transformation matrix; All modal data are aligned to a body reference frame with the target pose keypoints as the origin using pose estimation.

4. The method for identifying dangerous postures of personnel in a bucket truck based on multi-source sensor fusion as described in claim 1, characterized in that, Feature extraction is performed on the multimodal data, including: Feature extraction from RGB video is performed using a video spatiotemporal modeling network model; A three-dimensional convolutional extended network model is used to extract features from depth images; A bidirectional recurrent network model is used to extract features from IMU signals; Feature extraction of human pose keypoint sequences is performed using a spatiotemporal graph convolutional network model.

5. The method for identifying dangerous postures of personnel in a bucket truck based on multi-source sensor fusion as described in claim 1, characterized in that, Before enhancing the RGB video features, the following is also included: Use a fully connected layer to uniformly map the features of all modalities to the same dimension.

6. The method for identifying dangerous postures of personnel in a bucket truck based on multi-source sensor fusion as described in claim 1, characterized in that, The calculation of the adaptive weights for each modal feature includes: For each modality feature, adaptive weights are calculated through a fully connected layer, and the Sigmoid activation function is used to ensure that the adaptive weight values ​​are within the range of [0, 1].

7. The method for identifying dangerous postures of personnel in a bucket truck based on multi-source sensor fusion as described in claim 1, characterized in that, The fusion of features from the four modalities includes: The fused features are calculated using the following fusion formula: in, The features are those after fusion, and ⊙ represents element-wise product. To enhance the features of the RGB image, Features of depth images Characteristics of IMU signals, Features of the sequence of key points in human posture. , This is the adaptive weight tensor for the corresponding modal features.

8. The method for identifying dangerous postures of personnel in a bucket truck based on multi-source sensor fusion as described in claim 1, characterized in that, The identification of the posture category of the personnel on the boom truck based on the fused features includes: Flatten or pool the fused features along the sequence dimension; The flattened features are input into a fully connected network to determine the probability of the features being mapped to each posture category, thereby identifying the posture category of the person on the boom truck. The fully connected network includes two fully connected layers, with the activation function of the first layer being ReLU and the activation function of the second layer being Softmax.

9. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, which, when executed by the processor, implements the method of any one of claims 1 to 8.

10. A readable storage medium, characterized in that, The readable storage medium stores a computer program, which, when executed by a processor, implements the method of any one of claims 1 to 8.

Citation Information

Patent Citations

  • Three-dimensional human body posture estimation method fusing color image and depth image

    CN117132651A

  • Method and system for realizing three-dimensional human body posture estimation by fusing image and sparse IMU (Inertial Measurement Unit)

    CN117351564A

  • Assistant decision-making platform for water conservancy project operation and maintenance based on AI unmanned aerial vehicle

    CN119990627A

  • Human shape posture recognition method and system based on image analysis

    CN120431639A

  • Method of processing multimodal tasks, and an apparatus for the same

    US20230259779A1

Cited By

  • Dexterous hand action recognition method and device based on deep learning and medium

    CN121963321A