Pose Recognition Method and Device Based on Skeleton Separation, Unification and Attention Mechanism

By using graph convolution technology of frame separation and unity and attention mechanism in human posture recognition, multi-scale structural features and long-term dependencies are extracted, and attention mechanism processing is added at important limb nodes, the accuracy and real-time problems of posture recognition in the existing technology are solved, and efficient worker posture recognition is achieved.

CN113989849BActive Publication Date: 2025-06-20HANGZHOU LIGHT ELEPHANT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111299036.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-04
Publication Date
2025-06-20
Estimated Expiration
2041-11-04

AI Technical Summary

Technical Problem

The existing human posture recognition methods are difficult to accurately detect human position in complex environments, and the real-time and accuracy are low, limited to local joint connectivity, and fail to effectively extract multi-scale structural features and long-term dependencies.

Method used

The graph convolution technology based on skeleton separation and unity and attention mechanism is adopted. By obtaining skeleton data, multi-scale learning graph convolution processing is performed, the first skeleton features are extracted, and attention mechanism processing is added to important limb nodes, the weighted feature map is obtained as the second skeleton feature, and finally input into the Softmax classifier for pose recognition.

Benefits of technology

It realizes accurate recognition of body movements and postures of workers in factory workshops, improves the efficiency of skeleton recognition, and can accurately detect workers' positions, with high real-time and high accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113989849B_ABST
    Figure CN113989849B_ABST
Patent Text Reader

Abstract

This application relates to a pose recognition method and device based on skeleton separation and unification and attention mechanism. The method includes: obtaining skeleton data; selecting a graph sequence from the skeleton data, and then performing graph convolution processing of multi-scale learning on the graph sequence by a unified spatio-temporal operator based on a time window to obtain a first skeleton feature; performing attention mechanism processing on the first skeleton feature and completing recalibration of the first skeleton feature to obtain a weighted feature map as a second skeleton feature; performing global average pooling processing on the second skeleton feature, and inputting the result of the global average pooling processing into a Softmax classifier; the Softmax classifier recognizes and outputs the pose type. This application processes skeleton data, extracts multi-scale structural features and long-term dependencies, and then adds attention mechanism processing at important limb joint points to obtain enhanced skeleton features, thereby realizing accurate recognition of the limb movements and poses of workers in the factory workshop during production line work and improving the skeleton recognition efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of object detection and pose recognition, and specifically to a pose recognition method and device based on skeleton separation, unification, and attention mechanism. The method includes: obtaining skeleton data; selecting a graph sequence from the skeleton data, and then performing graph convolution processing of multi-scale learning on the graph sequence using a unified spatio-temporal operator based on a time window to obtain a first skeleton feature; performing attention mechanism processing on the first skeleton feature and completing recalibration of the first skeleton feature to obtain a weighted feature map as a second skeleton feature; performing global average pooling processing on the second skeleton feature, and inputting the result of the global average pooling processing into a Softmax classifier; and the Softmax classifier identifying and outputting the pose type. This application processes skeleton data, extracts multi-scale structural features and long-term dependencies, and then adds attention mechanism processing at important limb joint points to obtain enhanced skeleton features, thereby realizing accurate recognition of the limb movements and poses of workers in the factory workshop during production line work and improving the skeleton recognition efficiency. Background Art

[0002] With the development of deep learning, optical image object detection technology has penetrated into various fields of industry, and its application in factory workshops has gradually become widespread. Deep learning-based methods can detect the working status of workers on the production line in the factory workshop to achieve machine replacement of manual monitoring and thus improve office efficiency.

[0003] However, most current human pose estimation methods are difficult to accurately detect the position of a person due to the relatively complex environmental background. In addition, current human pose recognition methods are limited to local joint connectivity, regarding human joints as a set of independent features, and there are problems such as low real-time performance and low accuracy. Summary of the Invention

[0004] Based on the above technical problems, the present invention aims to adopt a graph convolution technical solution based on the combination of skeleton separation, unification, and attention mechanism to recognize the poses of workers on the production line in the factory workshop. When the production line in the workshop is operating, the sitting postures of workers are relatively fixed, and the upper limb movements on the production line are relatively single. By focusing on the joint point features of the workers' arms, we process the skeleton data, transcend local joint connectivity, extract multi-scale structural features and long-term dependencies, and then add attention mechanism processing at important limb joint points, thereby realizing the recognition of limb movements of factory workers during production line work.

[0005] The embodiments of the present application provide a pose recognition method, apparatus, and computer-readable storage medium based on skeleton separation, unification, and attention mechanism. To provide a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary part is not a general review, nor is it intended to identify key / important constituent elements or delineate the protection scope of these embodiments. Its sole purpose is to present some concepts in a simple form as a prelude to the detailed description that follows.

[0006] The first aspect of the present invention provides a pose recognition method based on skeleton separation, unification, and attention mechanism, including:

[0007] Obtain skeleton data;

[0008] Select a graph sequence from the skeleton data, and then perform graph convolution processing of multi-scale learning on the graph sequence using a unified spatio-temporal operator based on a time window to obtain a first skeleton feature, where the graph sequence includes multiple frames of spatio-temporal subgraphs;

[0009] Perform attention mechanism processing on the first skeleton feature and complete recalibration of the first skeleton feature to obtain a weighted feature map as the second skeleton feature;

[0010] Perform global average pooling processing on the second skeleton feature, and input the result of the global average pooling processing into a Softmax classifier;

[0011] The Softmax classifier recognizes and outputs the pose type.

[0012] Specifically, the step of selecting a graph sequence from the skeleton data, and then performing graph convolution processing of multi-scale learning on the graph sequence using a unified spatio-temporal operator based on a time window to obtain a first skeleton feature, where the graph sequence includes multiple frames of spatio-temporal subgraphs, includes:

[0013] Select a graph sequence from the skeleton data, where the graph sequence includes multiple frames of spatio-temporal subgraphs;

[0014] For any current frame of the multiple frames of spatio-temporal subgraphs, the adjacent matrix nodes are used to extrapolate the spatial connectivity in the frame direction to adjacent frames of the multiple frames of spatio-temporal subgraphs in the time domain to obtain a unified spatio-temporal operator of the time window;

[0015] Use the unified spatio-temporal operator of the time window to combine with the learned weight matrix to transform the graph sequence into a third skeleton feature;

[0016] Perform convolution operations on the third skeleton feature using spatio-temporal graph convolution blocks expanded at different expansion rates to obtain the first skeleton feature.

[0017] More specifically, the unified spatio-temporal operator of the time window is:

[0018]

[0019] Among them, \(t\) represents the current moment, and \(\tau\) represents the sliding time window. represents the adjacency matrix is the diagonal matrix of represents the activation function represents the learnable weight matrix.

[0020] Furthermore, the spatio-temporal graph convolution block expanded with different expansion rates performs a convolution operation on the third skeleton feature to obtain the first skeleton feature. The operation method is as follows:

[0021]

[0022] Among them, \(V\) represents the expansion rate, \(F\) represents the feature before expansion, and \(H\) represents the first skeleton feature.

[0023] Further preferably, the process of performing attention mechanism processing on the first skeleton feature and completing the recalibration of the first skeleton feature to obtain a weighted feature map as the second skeleton feature includes:

[0024] S31. Obtain skeleton training samples;

[0025] S32. Perform attention mechanism processing on the skeleton training samples and complete the spatio-temporal weight conversion of the skeleton training samples to obtain attention weights;

[0026] S33. Iteratively execute S32 to obtain optimized attention weights;

[0027] S34. Backtrack and process the first skeleton feature based on the optimized attention weights to obtain a weighted feature map as the second skeleton feature.

[0028] In the second aspect of the present invention, a skeleton neural network model is provided. The skeleton neural network model includes an input module, a multi-scale feature extraction module, an attention mechanism module, a pooling module, a classification module, and an output module. The multi-scale feature extraction module performs the step of performing multi-scale learning graph convolution processing on the graph sequence by the unified spatio-temporal operator based on the time window to obtain the first skeleton feature; the attention mechanism module performs the step of performing attention mechanism processing on the first skeleton feature and completing the recalibration of the first skeleton feature to obtain a weighted feature map as the second skeleton feature.

[0029] Preferably, the input module is used to input a graph sequence; the pooling module is used to perform a pooling operation on the processing result of the attention mechanism module; the classification module is used to classify and identify the pooling result output by the pooling module; and the output module is used to output the posture category identified by the classification module.

[0030] The third aspect of the present invention provides a posture recognition method based on a skeleton neural network model. The method applies the skeleton neural network model proposed in the second aspect of the present invention. The posture recognition method based on the skeleton neural network model includes:

[0031] Obtain skeleton data;

[0032] Input the skeleton data into the trained skeleton neural network model for recognition;

[0033] Output the posture type recognized by the skeleton neural network model.

[0034] The fourth aspect of the present invention provides a posture recognition device based on skeleton separation and unification and attention mechanism. The device includes:

[0035] An acquisition module, configured to acquire skeleton data;

[0036] A multi-scale module, configured to select a graph sequence from the skeleton data, and then perform graph convolution processing of multi-scale learning on the graph sequence based on a unified spatio-temporal operator of a time window to obtain a first skeleton feature, where the graph sequence includes multiple frames of spatio-temporal sub-graphs;

[0037] An attention module, configured to perform attention mechanism processing on the first skeleton feature and complete recalibration of the first skeleton feature to obtain a weighted feature map as a second skeleton feature;

[0038] A classification module, configured to perform global average pooling processing on the second skeleton feature, and input the global average pooling processing result into a Softmax classifier for recognition;

[0039] An output module, configured to output the posture type recognized by the Softmax classifier.

[0040] The fifth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0041] Obtain skeleton data;

[0042] Select a graph sequence from the skeleton data, and then perform graph convolution processing of multi-scale learning on the graph sequence based on a unified spatio-temporal operator of a time window to obtain a first skeleton feature, where the graph sequence includes multiple frames of spatio-temporal sub-graphs;

[0043] Perform attention mechanism processing on the first skeleton feature and complete recalibration of the first skeleton feature to obtain a weighted feature map as the second skeleton feature;

[0044] Perform global average pooling on the second skeleton feature and input the result of the global average pooling into a Softmax classifier;

[0045] The Softmax classifier identifies and outputs the pose type.

[0046] A sixth aspect of the present invention provides a computer program product, including a computer program, which when executed by a processor implements the following steps:

[0047] Obtain skeleton data;

[0048] Select a graph sequence from the skeleton data, and then perform graph convolution processing of multi-scale learning on the graph sequence based on a unified spatio-temporal operator of a time window to obtain a first skeleton feature, where the graph sequence includes multiple frames of spatio-temporal subgraphs;

[0049] Perform attention mechanism processing on the first skeleton feature and complete recalibration of the first skeleton feature to obtain a weighted feature map as the second skeleton feature;

[0050] Perform global average pooling on the second skeleton feature and input the result of the global average pooling into a Softmax classifier;

[0051] The Softmax classifier identifies and outputs the pose type.

[0052] The beneficial effects of the present application are as follows: The present application processes skeleton data, transcends the local joint connectivity in the skeleton data, extracts multi-scale structural features and long-term dependencies, and then adds attention mechanism processing at important limb joint points, and backtracks the skeleton features based on the optimized attention weights to obtain enhanced skeleton features, thereby realizing the recognition of the limb movements and postures of factory workshop workers during production line work, being able to accurately detect the positions of workers, with high real-time performance and high accuracy. In addition, the adjacency matrix used in the spatio-temporal operator is a multi-scale adjacency matrix, and through graph convolution processing of multi-scale learning, the present application can better extract feature information at different node distances and improve the efficiency of skeleton recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The drawings forming a part of the specification depict embodiments of the present application and, together with the description, are used to explain the principles of the present application.

[0054] Referring to the drawings, the present application can be more clearly understood from the following detailed description, where:

[0055] Figure 1 The schematic flow chart of the method according to an exemplary embodiment of the present application is shown;

[0056] Figure 2 The schematic diagram showing the process of obtaining the second skeleton feature through attention mapping in an exemplary embodiment of the present application is shown;

[0057] Figure 3 The schematic diagram of the device structure according to an exemplary embodiment of the present application is shown;

[0058] Figure 4 The schematic diagram of the structure of an electronic device provided by an exemplary embodiment of the present application is shown;

[0059] Figure 5 The schematic diagram of a storage medium provided by an exemplary embodiment of the present application is shown. Detailed implementation manners

[0060] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present application. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application. It is obvious to those skilled in the art that the present application can be implemented without one or more of these details. In other examples, some well-known technical features in the art are not described to avoid confusing the present application.

[0061] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "include" and / or "comprise" are used in this specification, they indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or combinations thereof.

[0062] Now, exemplary embodiments according to the present application will be described in more detail with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many different forms and should not be construed as being limited only to the embodiments set forth herein. The drawings are not drawn to scale, where some details may be enlarged for the purpose of clear expression and some details may be omitted. The shapes of various regions and layers shown in the drawings and their relative sizes and positional relationships are only exemplary, and may deviate in practice due to manufacturing tolerances or technical limitations, and those skilled in the art can design regions / layers with different shapes, sizes, and relative positions according to actual needs.

[0063] The following will describe several embodiments in conjunction with the accompanying drawings Figures 1-5 to describe the exemplary embodiments according to the present application. It should be noted that the following application scenarios are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard. On the contrary, the embodiments of the present application can be applied to any applicable scenario.

[0064] Embodiment 1:

[0065] This embodiment implements a pose recognition method based on skeleton separation and unification and attention mechanism, as Figure 1 shown, including:

[0066] S1. Obtain skeleton data;

[0067] S2. Select a graph sequence from the skeleton data, and then perform graph convolution processing of multi-scale learning on the graph sequence based on the unified spatio-temporal operator of the time window to obtain a first skeleton feature, where the graph sequence includes multiple frames of spatio-temporal subgraphs;

[0068] S3. Perform attention mechanism processing on the first skeleton feature and complete recalibration of the first skeleton feature to obtain a weighted feature map as the second skeleton feature;

[0069] S4. Perform global average pooling processing on the second skeleton feature, and input the result of the global average pooling processing into a Softmax classifier;

[0070] S5. The Softmax classifier recognizes and outputs the pose type.

[0071] Specifically, the step of selecting a graph sequence from the skeleton data, and then performing graph convolution processing of multi-scale learning on the graph sequence based on the unified spatio-temporal operator of the time window to obtain a first skeleton feature, where the graph sequence includes multiple frames of spatio-temporal subgraphs, includes:

[0072] Select a graph sequence from the skeleton data, where the graph sequence includes multiple frames of spatio-temporal subgraphs;

[0073] The adjacency matrix nodes of any current frame in the multiple frames of spatio-temporal subgraphs obtain the unified spatio-temporal operator of the time window by extrapolating the spatial connectivity in the frame direction to adjacent frames in the multiple frames of spatio-temporal subgraphs in the time domain;

[0074] Use the unified spatio-temporal operator of the time window to combine with the learned weight matrix to transform the graph sequence into a third skeleton feature;

[0075] Perform convolution operations on the third skeleton feature using spatio-temporal graph convolution blocks expanded at different expansion rates to obtain the first skeleton feature.

[0076] More specifically, the unified spatio-temporal operator of the time window is:

[0077]

[0078] where Y represents the spatio-temporal operator, but here it is for iterative processing, t represents the current moment, τ represents the sliding time window, represents the adjacency matrix, is the diagonal matrix of represents the activation function, represents the learnable weight matrix.

[0079] Further, the spatio-temporal graph convolution block expanded with different expansion rates is used to perform a convolution operation on the third skeleton feature to obtain the first skeleton feature, and the operation method is:

[0080]

[0081] where V represents the expansion rate, F represents the feature before expansion, and H represents the first skeleton feature.

[0082] Further preferably, the process of performing an attention mechanism processing on the first skeleton feature and completing the recalibration of the first skeleton feature to obtain a weighted feature map as the second skeleton feature includes:

[0083] S31. Obtain skeleton training samples;

[0084] S32. Perform an attention mechanism processing on the skeleton training samples and complete the spatio-temporal weight conversion of the skeleton training samples to obtain attention weights;

[0085] S33. Iteratively execute S32 to obtain optimized attention weights;

[0086] S34. Based on the optimized attention weights, perform a backtracking process on the first skeleton feature to obtain a weighted feature map as the second skeleton feature.

[0087] This application processes skeleton data, transcends the local joint connectivity in the skeleton data, extracts multi-scale structural features and long-term dependencies, then adds an attention mechanism processing at important limb joint points, and performs a backtracking process on the skeleton features based on the optimized attention weights to obtain data-enhanced skeleton features, thereby realizing the recognition of the limb movements and postures of workers in the factory workshop during production line work, being able to accurately detect the positions of workers, having high real-time performance and high accuracy, and improving the skeleton recognition efficiency.

[0088] Example 2:

[0089] This embodiment implements a posture recognition method based on skeleton separation and unification and an attention mechanism, and the steps are described in detail as follows.

[0090] Step 1: Obtain skeleton data.

[0091] Specifically, obtaining skeleton data includes obtaining the skeleton data of the worker's sitting posture in an actual factory workshop scenario. It should be noted that this application adopts a graph convolutional factory workshop production line worker posture recognition method based on skeleton separation, unification, and an attention mechanism module. Considering that when the workshop production line is operating, the worker's sitting posture is relatively fixed, and the upper limb movements on the production line are relatively single. We process the skeleton data by focusing on the characteristics of the worker's arm joint points. Therefore, the first step is to obtain the skeleton data of the worker's sitting posture on the factory workshop production line. In a specific implementation, obtaining skeleton data can be achieved by using a camera to shoot a video and then extracting the skeleton data through the Kinect platform. The Kinect platform is a platform specifically for extracting bone points and can directly obtain the spatial coordinates of the bones. Preferably, all the skeleton data is transformed into a five-dimensional array.

[0092] Step 2: Select a graph sequence from the skeleton data, and then perform graph convolutional processing of multi-scale learning on the graph sequence based on the unified spatio-temporal operator of the time window to obtain the first skeleton feature, where the graph sequence includes multiple frames of spatio-temporal subgraphs.

[0093] Preferably, the length of each video is 300 frames. If it is less than 300 frames, zeros are padded at the end. The selected graph sequence should include at least 300 frames of images, that is, 300 frames of spatio-temporal subgraphs. Additionally, optionally, the maximum and minimum values in the XYZ three directions can be found respectively in all the skeleton data and normalized.

[0094] To obtain the skeleton data, a graph sequence needs to be selected from it, where the graph sequence includes multiple frames of spatio-temporal subgraphs. Then, graph convolutional processing of multi-scale learning is performed on the graph sequence based on the unified spatio-temporal operator of the time window to obtain the first skeleton feature. The principle is that for spatio-temporal subgraphs, it is necessary to go beyond local joint connectivity and extract multi-scale structural features and long-term dependencies because structurally separated joints can also have strong correlations. An attention mechanism is added at the important joint points of the limbs, and then through subsequent processing, the recognition of the limb movements of factory workers during production line work is achieved. It should be noted here that the graph convolutional processing of multi-scale learning uses spatial convolutions with different dilation rates to obtain a larger receptive field without increasing the size of the convolutional kernel. Specifically, the formula for the multi-scale adjacency matrix is:

[0095]

[0096] where V i and V j represent two bone points, is V i and Vj The shortest distance between skeletal points. By setting different K values, adjacency matrices of different scales are obtained, eliminating the redundant dependence of distant neighborhoods on the weights of closer neighborhoods (i.e., the weights of nodes closer to the node are higher), and solving the problem of biased weights. The adjacency matrix used in the spatio-temporal operator is exactly this multi-scale adjacency matrix. Through graph convolution processing of multi-scale learning, the present application can better extract feature information of different node distances.

[0097] Specifically, select a graph sequence from the skeleton data, and then perform graph convolution processing of multi-scale learning on the graph sequence based on the unified spatio-temporal operator of the time window to obtain the first skeleton feature. Among them, the graph sequence includes multiple frames of spatio-temporal subgraphs, including: selecting a graph sequence from the skeleton data, where the graph sequence includes multiple frames of spatio-temporal subgraphs; the adjacency matrix nodes of any current frame in the multiple frames of spatio-temporal subgraphs obtain the unified spatio-temporal operator of the time window by extrapolating the spatial connectivity in the frame direction to adjacent frames in the time domain of the multiple frames of spatio-temporal subgraphs; use the unified spatio-temporal operator of the time window combined with the learned weight matrix to transform the graph sequence into the third skeleton feature; perform convolution operations on the third skeleton feature using spatio-temporal graph convolution blocks expanded at different expansion rates to obtain the first skeleton feature.

[0098] In a possible implementation manner, for example, we first consider a sliding time window of size τ on the input graph sequence, which obtains a frame of spatio-temporal subgraph at each step :

[0099]

[0100] Among them, is the union of all node sets across τ frames in the window, and the initial edge set is defined by tiling into the block adjacency matrix to define, is expressed as:

[0101]

[0102] Intuitively, each node in each submatrix is connected to itself and its adjacent frames in the current frame by extrapolating the spatial connectivity in the frame direction to the time domain. Therefore, in all τ frames, each node inside is densely connected to itself and adjacent frames, thus obtaining the unified spatio-temporal graph convolution operator of the time window:

[0103]

[0104] Among them, Y represents the spatio-temporal operator, but here it is processed iteratively, t represents the current moment, τ represents the sliding time window, represents the adjacency matrix, is The diagonal matrix, represents the activation function, represents the learnable weight matrix, which represents a learnable weight matrix in the l-th layer of the network.

[0105] Furthermore, the spatio-temporal graph convolution block expanded with different expansion rates is used to perform a convolution operation on the third skeleton feature to obtain the first skeleton feature. The operation method is as follows:

[0106]

[0107] Among them, V represents the expansion rate, V t1 、V t2 、V t3 represent using different expansion rates for different frames. F represents the feature before expansion, H represents the first skeleton feature, H t1 、H t2 、H t3 represent the features after expansion of different frames, and these combined together constitute the first skeleton feature. Preferably, different expansion rates are obtained according to the size of the multi-scale convolution kernel, and the expansion rate ranges from 64 frames per second to 100 frames per second. V t1 、V t2 、V t3 can respectively adopt different expansion rates such as 64, 67, and 73. H is composed of H t1 、H t2 、H t3 combined, but it can be understood that it does not necessarily only contain three frames. Here, it only describes the process of performing a convolution operation on the third skeleton feature using the spatio-temporal graph convolution block expanded with different expansion rates.

[0108] In the third step, the first skeleton feature is processed by an attention mechanism and recalibrated for the first skeleton feature to obtain a weighted feature map as the second skeleton feature.

[0109] Furthermore, as shown in Figure 2 , the process of processing the first skeleton feature by an attention mechanism and recalibrating the first skeleton feature to obtain a weighted feature map as the second skeleton feature includes: S31, obtaining skeleton training samples; S32, processing the skeleton training samples by an attention mechanism and completing the spatio-temporal weight conversion of the skeleton training samples to obtain attention weights; S33, iteratively executing S32 to obtain optimized attention weights; S34, backtracking and processing the first skeleton feature based on the optimized attention weights to obtain a weighted feature map as the second skeleton feature. Among them, the so-called backtracking processing means that on the basis of training the attention weights, the first skeleton feature to be processed currently is processed with them.

[0110] In a possible specific implementation, assume that the first skeleton feature is X. When processing the first skeleton feature X based on the optimized attention weight backtracking, X + X * M will be obtained, where M is the attention weight. This attention mechanism is introduced at the place after feature extraction of the skeleton points to obtain a weighted feature map, so as to strengthen the key skeleton points. The attention mechanism network belongs to a type of convolutional neural network structure. For example, given an input Y with the number of channels C', after a series of general transformations such as convolution, a feature map with the number of feature channels C is obtained. This feature map is sent into the attention mechanism network for processing. It will perform spatio-temporal transformation, change each two-dimensional feature channel into a real number, and the output dimension matches the number of input feature channels. Then, weights are generated for each channel, and finally, each channel is weighted by multiplication to the previous feature. For example, when currently processing the first skeleton feature X, recalibration of the first skeleton feature X is completed to obtain a weighted feature map.

[0111] In the fourth step, perform global average pooling on the second skeleton feature, and input the result of the global average pooling into the Softmax classifier.

[0112] In the fifth step, the Softmax classifier identifies and outputs the pose type.

[0113] This application processes the skeleton data, transcends the local joint connectivity in the skeleton data, extracts multi-scale structural features and long-term dependencies, and then adds an attention mechanism at the important joint points of the limbs. Based on the optimized attention weight backtracking to process the skeleton features, enhanced skeleton features are obtained, thereby realizing the recognition of the limb movements and postures of workers in the factory workshop during production line work, being able to accurately detect the positions of workers, with high real-time performance and high accuracy. In addition, the adjacency matrix used in the spatio-temporal operator is a multi-scale adjacency matrix, and through graph convolutional processing of multi-scale learning, this application can better extract feature information at different node distances and improve the skeleton recognition efficiency.

[0114] Embodiment 3:

[0115] This embodiment provides a skeleton neural network model. The skeleton neural network model includes an input module, a multi-scale feature extraction module, an attention mechanism module, a pooling module, a classification module, and an output module. The multi-scale feature extraction module performs the step of performing graph convolutional processing of multi-scale learning on the graph sequence using the unified spatio-temporal operator based on the time window to obtain the first skeleton feature; the attention mechanism module performs the step of performing attention mechanism processing on the first skeleton feature and completing the recalibration of the first skeleton feature to obtain a weighted feature map as the second skeleton feature.

[0116] Preferably, the input module is used to input a graph sequence; the pooling module is used to perform a pooling operation on the processing result of the attention mechanism module; the classification module is used to classify and identify the pooling result output by the pooling module; and the output module is used to output the pose category identified by the classification module.

[0117] Embodiment 4:

[0118] This embodiment provides a pose recognition method based on a skeleton neural network model. The method applies the skeleton neural network model in Embodiment 3. The pose recognition method based on the skeleton neural network model includes:

[0119] Obtain skeleton data;

[0120] Input the skeleton data into the trained skeleton neural network model for recognition;

[0121] Output the pose type recognized by the skeleton neural network model.

[0122] It should be noted that the trained skeleton neural network model needs to be trained before being trained. It is iteratively trained to a certain number of times, and the loss function needs to be continuously adjusted. The specific training process and the like are not specifically limited here.

[0123] Embodiment 5:

[0124] This embodiment implements a pose recognition device based on skeleton separation, unification, and attention mechanism, as Figure 3 shown. The device includes:

[0125] An acquisition module 701, configured to acquire skeleton data;

[0126] A multi-scale module 702, configured to select a graph sequence from the skeleton data, and then perform graph convolution processing of multi-scale learning on the graph sequence based on a unified spatio-temporal operator of a time window to obtain a first skeleton feature, where the graph sequence includes multiple frames of spatio-temporal subgraphs;

[0127] An attention module 703, configured to perform attention mechanism processing on the first skeleton feature and complete recalibration of the first skeleton feature to obtain a weighted feature map as a second skeleton feature;

[0128] A classification module 704, configured to perform global average pooling processing on the second skeleton feature, and input the global average pooling processing result into a Softmax classifier for recognition;

[0129] An output module 705, configured to output the pose type recognized by the Softmax classifier.

[0130] Next, please refer to Figure 4, which shows a schematic diagram of an electronic device provided by some embodiments of the present application. As Figure 4 shown, the electronic device 2 includes: a processor 200, a memory 201, a bus 202, and a communication interface 203. The processor 200, the communication interface 203, and the memory 201 are connected through the bus 202; a computer program that can run on the processor 200 is stored in the memory 201, and when the processor 200 runs the computer program, it executes the posture recognition method based on skeleton separation and unification and attention mechanism provided by any of the foregoing embodiments of the present application.

[0131] Among them, the memory 201 may include a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 203 (which can be wired or wireless), a communication connection is realized between the system network element and at least one other network element, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.

[0132] The bus 202 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 201 is used to store a program. After receiving an execution instruction, the processor 200 executes the program. The posture recognition method based on skeleton separation and unification and attention mechanism disclosed in any of the foregoing embodiments of the present application can be applied to the processor 200 or implemented by the processor 200.

[0133] The processor 200 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 200 or the instructions in the form of software. The above-mentioned processor 200 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 201, and the processor 200 reads the information in the memory 201 and combines its hardware to complete the steps of the above method.

[0134] The electronic device provided in the embodiments of the present application and the posture recognition method based on skeleton separation and unification and attention mechanism provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by them.

[0135] The embodiments of the present application also provide a computer-readable storage medium corresponding to the posture recognition method based on skeleton separation and unification and attention mechanism provided in the foregoing embodiments. Please refer to Figure 5 , Figure 5 The computer-readable storage medium shown is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will execute the posture recognition method based on skeleton separation and unification and attention mechanism provided in any of the foregoing embodiments.

[0136] In addition, examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated here one by one.

[0137] The computer-readable storage medium provided by the above embodiments of the present application and the method for allocating quantum key distribution channels in the space-division multiplexing optical network provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.

[0138] It should be noted that: The algorithms and displays provided herein are not inherently related to any particular computer, virtual device, or other equipment. Various general-purpose devices can also be used in conjunction with the teachings herein. Based on the above description, the structure required to construct such devices is obvious. In addition, the present application is not directed to any specific programming language. It should be understood that the content of the present application described herein can be implemented using various programming languages, and the description of the specific language above is to disclose the best implementation manner of the present application.

[0139] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and technologies have not been shown in detail so as not to obscure the understanding of this specification.

[0140] Similarly, it should be understood that, in order to streamline the present application and assist in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting the intention that the claimed subject matter of the present application requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate embodiment of the present application.

[0141] Those skilled in the art can understand that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature providing the same, equivalent, or similar purpose.

[0142] Each component embodiment of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the virtual machine creation device according to the embodiments of the present application. The present application can also be implemented as a device or device program (including a computer program product) for executing part or all of the methods described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or in any other form.

[0143] As described above, only the preferred specific embodiments of the present application are provided, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A posture recognition method based on skeleton separation, unification and attention mechanism, characterized in that Including: Obtain skeleton data; Select a graph sequence from the skeleton data, and then perform graph convolution processing of multi-scale learning on the graph sequence based on the unified spatio-temporal operator of the time window to obtain the first skeleton feature, where the graph sequence includes multiple frames of spatio-temporal subgraphs; Perform attention mechanism processing on the first skeleton feature and complete recalibration of the first skeleton feature to obtain a weighted feature map as the second skeleton feature; Perform global average pooling processing on the second skeleton feature, and input the result of the global average pooling processing into a Softmax classifier; The Softmax classifier identifies and outputs the pose type; The step of selecting a graph sequence from the skeleton data, and then performing graph convolution processing of multi-scale learning on the graph sequence based on the unified spatio-temporal operator of the time window to obtain the first skeleton feature, where the graph sequence includes multiple frames of spatio-temporal subgraphs, includes: Select a graph sequence from the skeleton data, where the graph sequence includes multiple frames of spatio-temporal subgraphs; the adjacency matrix nodes of any current frame in the multiple frames of spatio-temporal subgraphs obtain the unified spatio-temporal operator of the time window by extrapolating the spatial connectivity in the frame direction to adjacent frames of the multiple frames of spatio-temporal subgraphs in the time domain; Use the unified spatio-temporal operator of the time window to combine the learned weight matrix to transform the graph sequence into a third skeleton feature; Perform convolution operations on the third skeleton feature using spatio-temporal graph convolution blocks expanded with different expansion rates to obtain the first skeleton feature.

2. The posture recognition method based on skeleton separation, unification and attention mechanism according to claim 1, characterized in that The unified spatio-temporal operator of the time window is: where \(t\) represents the current moment and \(\tau\) represents the sliding time window. represents the adjacency matrix. is the diagonal matrix of (l) and \(\sigma\) represents the activation function, \(\theta\) 3. The posture recognition method based on skeleton separation, unification and attention mechanism according to claim 1, characterized in that The step of performing attention mechanism processing on the first skeleton feature and complete recalibration of the first skeleton feature to obtain a weighted feature map as the second skeleton feature, includes: S31. Obtain skeleton training samples; S32. Perform attention mechanism processing on the skeleton training samples and complete spatio-temporal weight conversion of the skeleton training samples to obtain attention weights; S33. Iteratively execute S32 to obtain optimized attention weights; S34. Backtrack and process the first skeleton feature based on the optimized attention weights to obtain a weighted feature map as the second skeleton feature.

4. A posture recognition device based on skeleton separation, unification and attention mechanism, characterized in that The device includes: An acquisition module for acquiring skeleton data; A multi-scale module for selecting a graph sequence from the skeleton data, and then performing graph convolution processing of multi-scale learning on the graph sequence based on the unified spatio-temporal operator of the time window to obtain the first skeleton feature, where the graph sequence includes multiple frames of spatio-temporal subgraphs; An attention module for performing attention mechanism processing on the first skeleton feature and complete recalibration of the first skeleton feature to obtain a weighted feature map as the second skeleton feature; a classification module for performing global average pooling processing on the second skeleton feature and inputting the result of the global average pooling processing into a Softmax classifier for identification; An output module for outputting the pose type identified by the Softmax classifier; The step of selecting a graph sequence from the skeleton data, and then performing graph convolution processing of multi-scale learning on the graph sequence based on the unified spatio-temporal operator of the time window to obtain the first skeleton feature, where the graph sequence includes multiple frames of spatio-temporal subgraphs, includes: Select a graph sequence from the skeleton data, where the graph sequence includes multiple frames of spatio-temporal subgraphs; the adjacency matrix nodes of any current frame in the multiple frames of spatio-temporal subgraphs obtain a unified spatio-temporal operator for the time window by extrapolating the spatial connectivity in the frame direction to adjacent frames of the spatio-temporal subgraphs in the time domain. Use the unified spatio-temporal operator of the time window to combine with the learned weight matrix to transform the graph sequence into the third skeleton feature. Perform a convolution operation on the third skeleton feature using spatio-temporal graph convolutional blocks expanded with different expansion rates to obtain the first skeleton feature.

5. A computer-readable storage medium, on which a computer program is stored, characterized in that When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-3.

6. A computer program product, including a computer program, characterized in that When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-3.

Citation Information

Patent Citations

  • Action structure self-attention graph convolutional network for action recognition

    CN112543936A