Behavior classification method and device, computer device, storage medium and program product
By training a behavior classification network using a multi-head cross-attention mechanism, and utilizing key point feature data and behavior query data from power monitoring images for multi-head cross-attention learning, the problem of low accuracy in classifying the behavior of power workers in existing technologies is solved, and more efficient anomaly detection is achieved.
Patent Information
- Application Number
- CN202310708684.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-14
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2043-06-14
AI Technical Summary
Existing methods for classifying the behavior of power workers have low accuracy and cannot effectively identify abnormal behaviors at power work sites.
A behavior classification network trained based on a multi-head cross-attention mechanism is adopted. By acquiring key point feature data and behavior query data in power monitoring images, multi-head cross-attention learning is performed, focusing on the spatial feature relationships and correlations between key points to improve classification accuracy.
It improves the accuracy of behavior classification, enabling better identification of abnormal behaviors at power operation sites and enhancing the safety of power equipment and personnel.
Smart Images

Figure CN116704610B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electric power, and in particular to a behavior classification method and device, computer equipment, storage medium and program product. BACKGROUND
[0002] In the electric power industry, in order to ensure the safety of operating personnel and electric power equipment, the behavior of operating personnel needs to be monitored in real time.
[0003] At present, the real-time monitoring method of the behavior of electric power operating personnel is to obtain a target image including the human body of the operating personnel through a monitoring camera or other monitoring equipment, to perform human key point detection on the target image, to perform classification recognition on the detected key point data, and to obtain a classification result, wherein the classification result is used to represent the behavior category of the operating personnel in the target image, such as making a phone call, falling down, and climbing, etc.
[0004] However, the above behavior classification method has low accuracy. SUMMARY
[0005] Therefore, it is necessary to provide a behavior classification method, device, computer equipment, storage medium and program product capable of improving classification accuracy in view of the above technical problems.
[0006] In a first aspect, the present application provides a behavior classification method. The method comprises:
[0007] obtaining an electric power monitoring image, wherein the electric power monitoring image includes a target image;
[0008] obtaining key point feature data corresponding to the target image according to the electric power monitoring image;
[0009] inputting the key point feature data and at least one behavior query data into a behavior classification network to obtain a behavior classification result of the target image output by the behavior classification network, wherein the behavior classification network is trained based on a multi-head cross-attention mechanism, the at least one behavior query data is determined according to a behavior classification requirement, and each behavior query data corresponds to a different behavior category.
[0010] In one embodiment, the behavior classification network comprises a multi-head self-attention network and a multi-head cross-attention prediction network connected to each other; the key point feature data and the at least one behavior query data are input into the behavior classification network to obtain the behavior classification result of the target image output by the behavior classification network, which comprises:
[0011] inputting each behavior query data into the multi-head self-attention network to obtain self-attention query data corresponding to each behavior query data, wherein a spatial embedding tensor of the multi-head self-attention network is obtained by position encoding according to the dimension of the key point feature data;
[0012] According to the respective attention query data, the key point feature data and the multi-head cross attention prediction network, the behavior classification result is obtained.
[0013] In one of the embodiments, the multi-head cross attention prediction network comprises a multi-head cross attention network and a prediction network connected in sequence with the multi-head self-attention network; and according to the respective attention query data, the key point feature data and the multi-head cross attention prediction network, the behavior classification result is obtained, comprising:
[0014] The respective attention query data and the key point feature data are input into the multi-head cross attention network to obtain cross query data corresponding to the respective attention query data output by the multi-head cross attention network;
[0015] The cross query data are input into the prediction network to obtain the behavior classification result output by the prediction network.
[0016] In one of the embodiments, the prediction network comprises a feedforward sub-network and a pooling sub-network connected in sequence; and the cross query data are input into the prediction network to obtain the behavior classification result output by the prediction network, comprising:
[0017] The cross query data are input into the feedforward sub-network to obtain cross enhancement data corresponding to the cross query data output by the feedforward sub-network;
[0018] According to the cross enhancement data and the pooling sub-network, the behavior classification result is obtained.
[0019] In one of the embodiments, the prediction network further comprises an affine transformation sub-network connected between the feedforward sub-network and the pooling sub-network; and according to the cross enhancement data and the pooling sub-network, the behavior classification result is obtained, comprising:
[0020] The cross enhancement data are input into the affine transformation sub-network to obtain at least one cross transformation data corresponding to the cross enhancement data output by the affine transformation sub-network;
[0021] The cross transformation data are input into the pooling sub-network to obtain the behavior classification result.
[0022] In one of the embodiments, the key point feature data corresponding to the target portrait is obtained according to the power monitoring image, comprising:
[0023] The power monitoring image is input into the portrait key point feature extraction network to obtain the key point feature data output by the portrait key point feature extraction network, and the portrait key point feature extraction network is trained based on a multi-resolution parallel mechanism.
[0024] In a second aspect, the application further provides a behavior classification device. The device comprises:
[0025] an image acquisition module, configured to acquire a power monitoring image, the power monitoring image comprising a target portrait;
[0026] a feature acquisition module, configured to acquire key point feature data corresponding to the target portrait according to the power monitoring image;
[0027] a behavior classification module, configured to input the key point feature data and at least one behavior query data into a behavior classification network to obtain a behavior classification result corresponding to the target portrait output by the behavior classification network, the behavior classification network being trained based on a multi-head cross attention mechanism, the at least one behavior query data being determined according to a behavior classification requirement, each behavior query data corresponding to a different behavior category.
[0028] In a third aspect, the present application also provides a computer device. The computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the steps of the method in the first aspect when executing the computer program.
[0029] In a fourth aspect, the present application also provides a computer readable storage medium. The computer readable storage medium stores a computer program, and the computer program implements the steps of the method in the first aspect when executed by a processor.
[0030] In a fifth aspect, the present application also provides a computer program product. The computer program product comprises a computer program, and the computer program implements the steps of the method in the first aspect when executed by a processor.
[0031] The aforementioned behavior classification method, apparatus, computer equipment, storage medium, and program product acquire power monitoring images, including target human images; based on the power monitoring images, acquire key point feature data corresponding to the target human images; input the key point feature data and at least one behavior query data into a behavior classification network to obtain the behavior classification result corresponding to the target human images output by the behavior classification network. The behavior classification network is trained based on a multi-head cross-attention mechanism, and at least one behavior query data is determined according to behavior classification requirements, with each behavior query data corresponding to a different behavior category; using the behavior classification network trained based on the multi-head cross-attention mechanism, multi-head cross-attention learning is performed on the behavior query data determined according to behavior classification requirements and the key point feature data corresponding to the target human images, utilizing multi-head cross-attention... The force mechanism focuses on important information and ignores unimportant information when processing data, enabling each behavior query data to learn the spatial feature relationships between key points in the key point feature data. Simultaneously, the behavior classification network, based on a multi-head cross-attention mechanism, can focus on the correlations between different key points in the target portrait, thus improving classification accuracy. This avoids the problem of low classification accuracy caused by traditional behavior classification methods using support vector machines, which can only analyze the positional relationships of key points in the key point feature data and cannot cross-learn the spatial feature relationships corresponding to the behavior categories determined according to the behavior classification requirements. The embodiments of this application utilize a behavior classification network to perform multi-head cross-attention learning on key point feature data and each behavior query data, resulting in highly accurate behavior classification results. Attached Figure Description
[0032] Figure 1 This is a diagram illustrating the application environment of the behavior classification method in one embodiment;
[0033] Figure 2 This is a flowchart illustrating a behavior classification method in one embodiment;
[0034] Figure 3 This is a schematic diagram of the behavior classification network structure in one embodiment;
[0035] Figure 4 This is a flowchart illustrating the process of obtaining behavior classification results in one embodiment;
[0036] Figure 5 This is a schematic diagram of the behavior classification network structure in another embodiment;
[0037] Figure 6 This is a flowchart illustrating the process of obtaining behavior classification results in another embodiment;
[0038] Figure 7 This is a flowchart illustrating the behavior classification method in another embodiment;
[0039] Figure 8 A structural block diagram of a behavior classification device in an embodiment;
[0040] Figure 9 An internal structural diagram of a computer device in an embodiment. DETAILED DESCRIPTION
[0041] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.
[0042] The behavior classification method provided by the embodiments of the present application can be applied to an application environment as shown in Figure 1 . The terminal 102 communicates with the server 104 through a network. The data storage system can store data required to be processed by the server 104. The data storage system can be integrated on the server 104, or placed on a cloud or other network server. The terminal 102 can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things device can be a smart speaker, a smart television, a smart air conditioner, a smart vehicle device, etc. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0043] The terminal 102 acquires a power monitoring image, and the power monitoring image includes a target portrait. According to the power monitoring image, key point feature data corresponding to the target portrait is acquired. The key point feature data and at least one behavior query data are input into a behavior classification network to obtain a behavior classification result of the target portrait output by the behavior classification network. The behavior classification network is trained based on a multi-head cross attention mechanism. The at least one behavior query data is determined according to a behavior classification requirement, and each behavior query data corresponds to a different behavior category.
[0044] In an embodiment, as shown in Figure 2 , a behavior classification method is provided. Taking the terminal 102 in Figure 1 as an example, the method includes the following steps:
[0045] Step 202, acquiring a power monitoring image, and the power monitoring image includes a target portrait.
[0046] The power monitoring image refers to an image including a target person image obtained by a monitoring system in a power operation site. The behavior type of the target person image in the power monitoring image is recognized to determine whether the staff in the power operation site has abnormal behavior. The behavior type recognition can also be referred to as human posture estimation. The target person image in the power monitoring image can be one person image or multiple person images.
[0047] The power monitoring image refers to an image including a target person image obtained by a monitoring system in a power operation site. The behavior type of the target person image in the power monitoring image is recognized to determine whether the staff in the power operation site has abnormal behavior. The behavior type recognition can also be referred to as human posture estimation. The target person image in the power monitoring image can be one person image or multiple person images.
[0048] For example, the power inspection can be performed by using a drone, and the power monitoring image can be obtained by using a high-definition camera, infrared thermal imaging, etc. mounted on the drone. The power monitoring image can also be obtained by using a visual inspection tool, such as a steel wire indoor inspection robot, an insulation high-speed inspection device, etc. The power monitoring image can also be obtained by using a monitoring device in the power operation site.
[0049] In step 204, key point feature data corresponding to the target person image is obtained according to the power monitoring image.
[0050] The key point feature data refers to position information of each key point in the target person image in the power monitoring image and a position relationship between the key points. For example, the key point feature data is a key point heat map including 17 dimensions, one dimension corresponds to one key point, and the key points include left ear, right ear, left eye, right eye, nose, right shoulder, head, left shoulder, right hand, right elbow, left hand, left elbow, right waist, left waist, right knee, left knee, right foot, and left foot. In the field of human posture estimation technology, the key point feature data can be referred to as a human skeleton feature map.
[0051] If the power monitoring image includes multiple target person images, the key point feature data corresponding to each target person image is obtained. For example, the power monitoring image includes a target person image 1 and a target person image 2, and the key point feature data corresponding to the target person image 1 and the key point feature data corresponding to the target person image 2 are obtained.
[0052] The key point feature data corresponding to the target portrait refers to performing key point feature extraction on the power monitoring image by using a pose estimation algorithm to obtain corresponding key point feature data. The pose estimation algorithm used in this embodiment is not specifically limited. Exemplarily, the OPENPOSE network can be used to perform key point detection on the power monitoring image to obtain key point feature data output by the OPENPOSE network. Exemplarily, the CPM (Convolutional Pose Machines) network or the LCR-Net (Localization-Classification-Regression Net) or the like can be used to perform key point detection on the power monitoring image to obtain key point feature data.
[0053] In step 206, the key point feature data and the at least one behavior query data are input into the behavior classification network to obtain a behavior classification result of the target portrait output by the behavior classification network.
[0054] The behavior classification network is trained based on a multi-head cross attention mechanism, the at least one behavior query data is determined according to a behavior classification requirement, and each behavior query data corresponds to a different behavior category. One behavior query data can correspond to one behavior category or multiple different behavior categories.
[0055] The behavior classification method adopted in this embodiment can be regarded as a two-stage target detection network architecture. The first stage is a feature extraction stage, and the second stage is a target detection stage based on the key point feature data.
[0056] The behavior query data can be regarded as a feature matrix describing a behavior category. Exemplarily, the operation process of the multi-head cross attention mechanism can be represented by the following formula:
[0057]
[0058]
[0059] wherein, N is the number of cross attention heads in the multi-head cross attention mechanism, Q is an input query vector of the i-th cross attention head, which is transformed from the corresponding behavior query data by a query conversion matrix, K is a key vector input into the i-th cross attention head, which is transformed from the key point feature data by a key conversion matrix, V is a value vector input into the i-th cross attention head, which is transformed from the key point feature data by a value conversion matrix, and C is a cross attention output vector of the i-th cross attention head. h i i i h is the dimension of each cross-attention head to implement scaled dot-product attention, W O is a learnable conversion matrix.
[0060] The behavior classification network cross-learns the key point feature data and the at least one behavior query data, inputs the cross-learned data into a normalization layer and a residual connection layer for processing, and then performs feedforward prediction to obtain a prediction probability of the target portrait corresponding to each behavior query data. The behavior classification result corresponding to the target portrait is determined as a behavior category corresponding to the behavior query data with the maximum prediction probability, that is, the behavior category corresponding to the target portrait is obtained.
[0061] Because human postures have strong spatial correlation, the features at different positions in the power monitoring image have different importance degrees on the behavior classification result, and the non-important information of the background needs to be suppressed when classifying the behavior of the target portrait. The spatial relationship of each key point in the human posture is different in different behavior categories. In this embodiment, the behavior classification network trained based on the multi-head cross-attention mechanism cross-learns the key point feature data and each behavior query data, so that each behavior query data learns the spatial feature relationship between each key point in the key point feature data, thereby well utilizing the useful features and achieving more accurate behavior classification recognition.
[0062] In the above behavior classification method, an electric power monitoring image is obtained, and the electric power monitoring image includes a target portrait; key point feature data corresponding to the target portrait is obtained according to the electric power monitoring image; the key point feature data and at least one behavior query data are input into a behavior classification network to obtain a behavior classification result corresponding to the target portrait output by the behavior classification network, the behavior classification network is trained based on a multi-head cross attention mechanism, the at least one behavior query data is determined according to a behavior classification requirement, and each behavior query data corresponds to a different behavior category; the behavior classification network trained based on the multi-head cross attention mechanism is used for multi-head cross attention learning of the behavior query data determined according to the behavior classification requirement and the key point feature data corresponding to the target portrait, the multi-head cross attention mechanism is used to focus on important information and ignore unimportant information when processing data, so that each behavior query data learns the spatial feature relationship between each key point in the key point feature data, and meanwhile, the behavior classification network based on the multi-head cross attention mechanism can focus on the correlation between different key points in the target portrait, so as to improve the classification accuracy; the behavior classification method using a support vector machine in the traditional technology can only analyze the positional relationship of each key point in the key point feature data, and cannot cross-learn the spatial feature relationship between the key point feature data and the behavior category corresponding to the behavior classification requirement, which leads to low classification accuracy; the behavior classification result obtained by the behavior classification network through multi-head cross attention learning of the key point feature data and each behavior query data has high accuracy.
[0063] In one embodiment, based on Figure 2 As shown in the embodiment shown in FIG. 6, the behavior classification network provided in this embodiment includes a multi-head self-attention network and a multi-head cross-attention prediction network connected to each other. Figure 3 As shown in FIG. 6, the process includes steps 402 and 404. Figure 4 The embodiment relates to how to input the key point feature data and the at least one behavior query data into the behavior classification network to obtain the behavior classification result corresponding to the target portrait output by the behavior classification network. Figure 4 As shown in FIG. 6, the process includes steps 402 and 404.
[0064] In step 402, each behavior query data is input into the multi-head self-attention network to obtain self-attention query data corresponding to each behavior query data. The spatial embedding tensor E of the multi-head self-attention network is obtained by position coding according to the dimension of the key point feature data.
[0065] In order to add the spatial position information of the key point feature data to the multi-head self-attention network, the spatial embedding tensor E of the multi-head self-attention network is obtained by position coding according to the dimension of the key point feature data.
[0066] Exemplarily, the spatial embedding tensor E in the embodiment can be calculated using the sine-cosine position encoding method. The position encoding (PE) process can be represented by the following formula:
[0067] PE(i,2j)=sin(i / 10000 2j / d )
[0068] PE(i,2j+1)=cos(i / 10000 2j / d ),
[0069] wherein d is the total dimension of the spatial embedding tensor E (which is the same as the dimension of the key feature data), i is the position index, and j is the index of the encoding dimension.
[0070] Exemplarily, the multi-head self-attention network includes a multi-head self-attention subnetwork and a normalization and residual connection subnetwork, wherein the operation process of the multi-head self-attention subnetwork can be represented by the following formula:
[0071]
[0072]
[0073] wherein N h is the number of self-attention heads in the multi-head self-attention network, Q i is the input query vector of the i-th self-attention head, which is transformed by the query transformation matrix from the behavior query data and the spatial embedding tensor E, K i is the input key vector of the i-th self-attention head, which is transformed by the key transformation matrix from the behavior query data and the spatial embedding tensor E, V i is the input value vector of the i-th self-attention head, which is transformed by the value transformation matrix from the behavior query data and the spatial embedding tensor E, C h is the dimension of each self-attention head to realize the scaled dot-product attention, W O is a learnable transformation matrix.
[0074] The output of the multi-head self-attention subnetwork and the input data of the multi-head self-attention subnetwork are jointly input into the normalization and residual connection subnetwork to obtain the self-attention query data.
[0075] Step 404, obtaining the behavior classification result according to the respective self-attention query data, the key point feature data, and the multi-head cross-attention prediction network.
[0076] In the embodiment, the position encoding of the key point feature data in the dimension is used to obtain the spatial embedding tensor E of the multi-head self-attention network, so that the position information of the key point feature data is added to the multi-head self-attention network, and then each behavior query data is subjected to multi-head self-attention learning through the multi-head self-attention network, so that the behavior classification network is associated with the position information of the key point feature data. Thus, the behavior classification network does not need to be designed according to the manner of obtaining the key point feature data, that is, the behavior classification network provided in the embodiment has universality.
[0077] In one embodiment, based on Figure 4 As shown in the embodiment, the multi-head cross-attention prediction network in the behavior classification network provided in the embodiment includes a multi-head cross-attention network and a prediction network connected in sequence with the multi-head self-attention network, as shown in Figure 5 As shown in Figure 6 The embodiment relates to a process of obtaining a behavior classification result according to the respective attention query data, the key point feature data and the multi-head cross-attention prediction network. As shown in Figure 6 The process includes steps 602 and 604.
[0078] In step 602, the respective attention query data and the key point feature data are input into the multi-head cross-attention network to obtain cross query data corresponding to the respective attention query data output by the multi-head cross-attention network.
[0079] In the multi-head cross-attention network, the input Q vector is derived from the self-attention query data, and the K vector and the V vector are derived from the key point feature data.
[0080] Through the multi-head cross-attention network, the respective attention query data learns the similarity between the respective attention query data and the key point feature data to obtain the corresponding cross query data.
[0081] In step 604, the respective cross query data is input into the prediction network to obtain a behavior classification result output by the prediction network.
[0082] In one possible implementation, the prediction network includes a feedforward subnetwork and a pooling subnetwork connected in sequence. In this implementation, step 604 inputs the respective cross query data into the prediction network to obtain a behavior classification result output by the prediction network, including:
[0083] In step A1, the respective cross query data is input into the feedforward subnetwork to obtain cross-enhanced data corresponding to the respective cross query data output by the feedforward subnetwork.
[0084] In step A2, the behavior classification result is obtained according to the respective cross-enhanced data and the pooling subnetwork.
[0085] In a possible implementation, the step A2 directly inputs each cross-enhanced data into the pooling subnetwork to obtain a behavior classification result output by the pooling subnetwork.
[0086] In another possible implementation, the prediction network further includes an affine transformation subnetwork connected between the feedforward subnetwork and the pooling subnetwork, and in this possible implementation, the step A2 obtains the behavior classification result according to each cross-enhanced data and the pooling subnetwork, including:
[0087] The step B1 inputs each cross-enhanced data into the affine transformation subnetwork to obtain at least one cross-transformed data corresponding to each cross-enhanced data output by the affine transformation subnetwork.
[0088] The step B2 inputs each cross-transformed data into the pooling subnetwork to obtain the behavior classification result.
[0089] Generally, one behavior query data corresponds to one behavior category. In some extreme behavior classification scenarios, the number of behavior categories is large, which means that if more behavior categories are to be parsed, the number of response behavior query data needs to be increased, but this will bring heavy parameter quantity and calculation quantity. With the increase of the number of behavior categories, the behavior classification network will consume a lot of computing resources and may reduce the classification performance. Among them, the multi-head self-attention network and the multi-head cross-attention network will present a quadratic complexity O(n 2 ) with the increase of the sequence length of the input data:
[0090] O(n 2 )=4nh 2 d 2 +2n 2 hd;
[0091] Wherein, n is the sequence length of the input data; d is the dimension of each head; h is the number of heads.
[0092] In order to solve the problem that the calculation complexity of the behavior classification network increases quadratically with the increase of the behavior classification, the grouping decoding method is adopted in the embodiment to reduce the parameter quantity and the calculation quantity. By adding an affine transformation layer between the feedforward subnetwork and the pooling subnetwork, one behavior query data corresponds to multiple behavior categories through affine transformation and pooling operation. In this way, the behavior classification network and the behavior query data present a linear complexity, and the embodiment has great advantages in some application scenarios that need to be actually landed (such as edge terminal). In terms of calculation complexity, the calculation quantity of the affine transformation subnetwork is equal to that of the full connection layer, and the calculation quantity of the pooling subnetwork is NxD times of multiplication (N is the number of behavior query data, and D is the dimension of key point feature data), so the calculation quantities of the affine transformation subnetwork and the pooling subnetwork are both in linear relationship with the number N of input behavior query data.
[0093] For example, first define the grouping factor g = K / N, where N is the number of behavioral query data and K is the actual number of behavioral categories to be queried; the output of the affine transformation subnetwork is the actual number of behavioral categories to be classified, L. i :
[0094] L i =(W k ·Q k ) j n = i divg, j = imodg;
[0095] In the formula: The query data corresponds to the nth row; It is the nth learnable transformation matrix.
[0096] When g=1, the behavior classification network in this embodiment can also be called a fully decoded behavior classification network, where one behavior query data corresponds to one behavior category. When g≠1, the behavior classification network in this embodiment can also be called a grouped decoded behavior classification network, which means that each behavior query data corresponds to several behavior categories. The behavior classification network in this embodiment randomly divides the behavior categories into several groups.
[0097] In this embodiment, an efficient grouping factor g can be set according to different task requirements. Typically, g is greater than 1. The relationship between behavioral query data and grouping is established through an affine transformation subnetwork, and then the final classification result is obtained by pooling along the embedding dimension of the token.
[0098] The behavior classification network provided in this embodiment achieves lightweight design while maintaining accuracy and stability. At the same time, the grouping method can be easily extended to any category, making it highly universal.
[0099] In one embodiment, based on Figure 2 The illustrated embodiment describes the process of obtaining key point feature data corresponding to a target human image based on a power monitoring image. This process includes: inputting the power monitoring image into a human key point feature extraction network to obtain key point feature data output by the network. The human key point feature extraction network is trained using a multi-resolution parallel mechanism.
[0100] For example, a human keypoint feature extraction network includes a stem, a backbone, and a head. The backbone in the human keypoint feature extraction network uses a multi-resolution branch parallel computing method and cross-resolution connections in different branches to exchange features, while maintaining strong semantic and positional information, making the predicted heatmap more spatially accurate.
[0101] The base is used to reduce the input power monitoring image to 1 / 4 of the original image size to obtain a reduced image to reduce the subsequent calculation amount. Then the reduced image is put into the main body divided into stage layer and transition layer to extract features; the stage layer in the main body mainly plays a role of image feature extraction and multi-feature fusion; the feature extraction module uses a residual connection structure and is repeated 4 times to more fully extract features; multi-feature fusion is realized by nearest neighbor upsampling and convolution downsampling. The transition layer in the main body increases the downsampling branch on the basis of the original branch, performs convolution downsampling operation on the existing feature map, and further obtains a feature map with more rich semantic information. Finally, the regression head fuses each feature obtained from the last stage layer to predict the key point heat map corresponding to the target portrait, i.e. the key point feature data corresponding to a target portrait containing 17 key points.
[0102] The loss function of the portrait key point feature extraction network in the training process in the embodiment is as follows: similar to other deep learning algorithms, the loss function is mean square error loss. Since the heat map is predicted, if a single real labeled point (ground truth, GT) in the data set is directly used for calculation, too many negative samples will cause the training to be difficult to converge, so it is necessary to expand the GT into a GT heat map with two-dimensional Gaussian distribution. Since different key points have different prediction difficulties, corresponding weights need to be calculated for different key points when calculating the total loss.
[0103] Illustratively, the portrait key point feature extraction network can be constructed based on the HRNet network.
[0104] In one embodiment, referring to Figure 7 , a behavior classification method is provided, comprising:
[0105] Step 702, obtaining a power monitoring image, the power monitoring image including a target portrait.
[0106] Step 704, inputting the power monitoring image into a portrait key point feature extraction network to obtain key point feature data output by the portrait key point feature extraction network, the portrait key point feature extraction network being trained based on a multi-resolution parallel mechanism.
[0107] Step 706, inputting at least one behavior query data into a multi-head self-attention network to obtain self-attention query data corresponding to each behavior query data, the spatial embedding tensor of the multi-head self-attention network being obtained by position encoding according to the dimension of the key point feature data. Among them, at least one behavior query data is determined according to the behavior classification requirement, and each behavior query data corresponds to a different behavior category.
[0108] Step 708, input the respective attention query data and key point feature data into the multi-head cross attention network to obtain cross query data corresponding to the respective attention query data output by the multi-head cross attention network.
[0109] Step 710, input each cross query data into a feedforward subnetwork to obtain cross enhancement data corresponding to each cross query data output by the feedforward subnetwork.
[0110] Step 712, input each cross enhancement data into an affine transformation subnetwork to obtain at least one cross transformation data corresponding to each cross enhancement data output by the affine transformation subnetwork.
[0111] Step 714, input each cross transformation data into a pooling subnetwork to obtain a behavior classification result corresponding to the target portrait.
[0112] The process of the above behavior classification method can correspond to a two-stage network structure of a feature encoder and a classification decoder. The applicant provides a way of training the two-stage network structure corresponding to the behavior classification method provided in the embodiment. The exemplary training process is as follows:
[0113] The two-stage network structure is trained as a whole by using the pre-training fine-tuning method. The feature encoder part uses the Microsoft COCO2017 public dataset as the training dataset. The COCO2017 dataset is the mainstream dataset for pose estimation, which includes rich images with single / multiple people, large / medium / small targets, and related research uses it as the training and test dataset. The classification decoder part is trained using a private dataset. The private dataset is mainly based on the substation power scene, and there are 12510 sample pictures, including various indoor and outdoor scenes, various personnel working postures (climbing, crossing, making a phone call, falling down, lifting), different number of people distribution (single person, multiple people), and other situations to simulate various scenes in the actual field. The sample quantity of the private dataset is shown in Table 1.
[0114] Table 1 Sample quantity of private dataset
[0115] Behavior class Sample number Climbing 2278 Crossing 2470 Calling 2721 Falling 2486 Carrying 2555
[0116] The hardware platform conditions adopted by the training process provided by the embodiment include: Intel Xeon Gold 6242R processor (memory 754 GB), single NVIDIA A100 (display memory 80 GB), Ubuntu 18.04 operating system and PyTorch deep learning framework. In the pre-training stage, the feature encoder is first trained, and the classification decoder is not added to the training at this time; the training input picture size is 384 pixels x 288 pixels, the batch size is 24, the Adam optimizer is used, the initial learning rate is set to 1e-3, the learning rate adjustment method is MultiStep, the drop points are 170 and 200, the drop rate is set to 0.1, and a total of 210 cycles are trained. In the fine-tuning stage, the feature encoder is fixed, only the classification decoder is trained, the learning rate is set to 1e-5, the learning rate adjustment method is cosine annealing, the batch size is 128, the weight decay is 1e-8, and a total of 30 cycles are trained. The rest of the hyperparameters are the same as in the pre-training stage.
[0117] In the training process, new training samples can also be created by data augmentation to expand the size of the existing data set, to improve the accuracy of the feature encoder and the classification decoder, effectively prevent overfitting, and allow the model to better generalize, so that it is more accurate and reliable in actual application. For example, the data augmentation method shown in Table 2 can be used.
[0118] Table 2 Data augmentation method for training samples
[0119] BoF (Bag of Freebies) f sp / Hz]]> Random flip Probability changed to 0.5 Brightness change Probability 0.5, from -10 to +10 Contrast change Probability 0.5, from 75% to 125% Saturation change Probability 0.5, from 90% to 110%
[0120] It should be understood that although each step in the flowchart involved in each embodiment as described above is displayed in sequence according to the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, as described above, at least part of the steps in the flowchart involved in each embodiment can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or steps or stages in other steps.
[0121] Based on the same inventive concept, the embodiments of the present application also provide a behavior classification device for implementing the above-mentioned behavior classification method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, and therefore the specific limitations in one or more behavior classification device embodiments provided below can refer to the limitations of the behavior classification method described above, which will not be described here again.
[0122] In one embodiment, as shown in Figure 8 a behavior classification device is provided, comprising: an image acquisition module 802, a feature acquisition module 804, and a behavior classification module 806, wherein:
[0123] The image acquisition module 802 is configured to acquire a power monitoring image, wherein the target portrait is included in the power monitoring image.
[0124] The feature acquisition module 804 is configured to acquire key point feature data corresponding to the target portrait according to the power monitoring image.
[0125] The behavior classification module 806 is configured to input the key point feature data and at least one behavior query data into a behavior classification network to obtain a behavior classification result corresponding to the target portrait output by the behavior classification network, wherein the behavior classification network is trained based on a multi-head cross-attention mechanism, and each behavior query data corresponds to a different behavior category.
[0126] In one embodiment, the behavior classification network comprises a multi-head self-attention network and a multi-head cross-attention prediction network connected to each other; the behavior classification module 806 is configured to input each behavior query data into the multi-head self-attention network to obtain self-attention query data corresponding to each behavior query data, and the spatial embedding tensor of the multi-head self-attention network is obtained by position encoding according to the dimension of the key point feature data; and configured to acquire the behavior classification result according to each self-attention query data, the key point feature data, and the multi-head cross-attention prediction network.
[0127] In one embodiment, the multi-head cross-attention prediction network comprises a multi-head cross-attention network and a prediction network connected to each other in sequence; the behavior classification module 806 is configured to input each self-attention query data and the key point feature data into the multi-head cross-attention network to obtain cross-query data corresponding to each self-attention query data output by the multi-head cross-attention network; and configured to input each cross-query data into the prediction network to obtain the behavior classification result output by the prediction network.
[0128] In an embodiment, the prediction network comprises a feedforward subnetwork and a pooling subnetwork connected in sequence; the behavior classification module 806 is configured to input each of the cross query data into the feedforward subnetwork to obtain cross enhanced data corresponding to each of the cross query data output by the feedforward subnetwork; and configured to obtain the behavior classification result according to each of the cross enhanced data and the pooling subnetwork.
[0129] In an embodiment, the prediction network further comprises an affine transformation subnetwork connected between the feedforward subnetwork and the pooling subnetwork; the behavior classification module 806 is configured to input each of the cross enhanced data into the affine transformation subnetwork to obtain at least one cross transformed data corresponding to each of the cross enhanced data output by the affine transformation subnetwork; and configured to input each of the cross transformed data into the pooling subnetwork to obtain the behavior classification result.
[0130] In an embodiment, the feature acquisition module 804 is configured to input the power monitoring image into a portrait key point feature extraction network to obtain the key point feature data output by the portrait key point feature extraction network, wherein the portrait key point feature extraction network is trained based on a multi-resolution parallel mechanism.
[0131] Each of the above behavior classification apparatuses can be implemented by software, hardware, and combinations thereof, in whole or in part. Each of the above modules can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a memory in a computer device in software form, so as to be called and executed by a processor to perform the operations corresponding to each of the above modules.
[0132] In an embodiment, a computer device is provided, which can be a server, and an internal structure diagram thereof can be as shown in Figure 9 The computer device comprises a processor, a memory, an input / output interface (I / O), and a communication interface. The processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The database of the computer device is configured to store power monitoring images. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to communicate with external terminals through network connection. The computer program is executed by the processor to implement a behavior classification method.
[0133] Those skilled in the art can understand that Figure 9 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0134] In one embodiment, a computer device is provided, including a memory and a processor, the memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0135] In one embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the steps in the above method embodiments.
[0136] In one embodiment, a computer program product is provided, which includes a computer program, and the computer program is executed by a processor to implement the steps in the above method embodiments.
[0137] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant national and regional laws, regulations and standards.
[0138] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (Read-Only Memory, ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (Magnetoresistive Random Access Memory, MRAM), ferroelectric memory (Ferroelectric Random Access Memory, FRAM), phase change memory (Phase Change Memory, PCM), graphene memory, etc. Volatile memory can include random access memory (Random Access Memory, RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (Static Random Access Memory, SRAM) or dynamic random access memory (Dynamic Random Access Memory, DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0139] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.
[0140] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A behavior classification method, characterized in that, The method includes: Acquire power monitoring images, wherein the power monitoring images include the image of the target person; Based on the power monitoring image, key point feature data corresponding to the target portrait is obtained. The key point feature data refers to the position information of each key point in the target portrait in the power monitoring image and the position information between each key point. At least one behavioral query data is input into a multi-head self-attention network in the classification network to obtain self-attention query data corresponding to each behavioral query data. The spatial embedding tensor of the multi-head self-attention network is obtained by position encoding according to the dimension of the key point feature data. The at least one behavioral query data is determined according to the behavioral classification requirements, and each behavioral query data corresponds to a different behavioral category. Based on the self-attention query data, the key point feature data, and the multi-head cross-attention prediction network in the classification network, the behavior classification result corresponding to the target portrait is obtained.
2. The method according to claim 1, characterized in that, The multi-head cross-attention prediction network includes a multi-head cross-attention network and a prediction network sequentially connected to the multi-head self-attention network; obtaining the behavior classification result based on each of the self-attention query data, the key point feature data, and the multi-head cross-attention prediction network includes: Each self-attention query data and the key point feature data are input into the multi-head cross-attention network to obtain the cross-query data corresponding to each self-attention query data output by the multi-head cross-attention network. The cross-query data are input into the prediction network to obtain the behavior classification result output by the prediction network.
3. The method according to claim 2, characterized in that, The prediction network includes a feedforward subnetwork and a pooling subnetwork connected in sequence; the step of inputting each of the cross-query data into the prediction network to obtain the behavior classification result output by the prediction network includes: Each of the cross-query data is input into the feedforward sub-network to obtain the cross-enhancement data corresponding to each of the cross-query data output by the feedforward sub-network; The behavior classification result is obtained based on the cross-enhanced data and the pooling sub-network.
4. The method according to claim 3, characterized in that, The prediction network further includes an affine transformation subnetwork connected between the feedforward subnetwork and the pooling subnetwork; obtaining the behavior classification result based on each of the cross-enhancement data and the pooling subnetwork includes: Each of the aforementioned cross-enhancement data is input into the affine transform sub-network to obtain at least one cross-transform data corresponding to each of the aforementioned cross-enhancement data output by the affine transform sub-network; The cross-transformation data are input into the pooling sub-network to obtain the behavior classification result.
5. The method according to claim 1, characterized in that, The step of obtaining key point feature data corresponding to the target human image based on the power monitoring image includes: The power monitoring image is input into the human key point feature extraction network to obtain the key point feature data output by the human key point feature extraction network. The human key point feature extraction network is trained based on a multi-resolution parallel mechanism.
6. A behavior classification device, characterized in that, The device includes: An image acquisition module is used to acquire power monitoring images, wherein the power monitoring images include a target human image; The feature acquisition module is used to acquire key point feature data corresponding to the target portrait based on the power monitoring image. The key point feature data refers to the position information of each key point in the target portrait in the power monitoring image and the position information between each key point. The behavior classification module is used to input at least one behavior query data into a multi-head self-attention network in the classification network to obtain self-attention query data corresponding to each behavior query data. The spatial embedding tensor of the multi-head self-attention network is obtained by position encoding according to the dimension of the key point feature data. The module is also used to obtain the behavior classification result corresponding to the target image based on each self-attention query data, the key point feature data, and the multi-head cross-attention prediction network in the classification network. The at least one behavior query data is determined according to the behavior classification requirements, and each behavior query data corresponds to a different behavior category.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Human interaction behavior detection method based on self-attention mechanism
CN114782995A