Bone sequence behavior recognition method, device and equipment and storage medium

CN116824689BActive Publication Date: 2026-09-29SHANTOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310446265.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-23
Publication Date
2026-09-29
Estimated Expiration
2043-04-23

AI Technical Summary

Technical Problem

然而,由于传统的基于图像或视频的动作识别方法受到了光照、噪声等因素的影响,因此在实际应用中存在一定的局限性

Benefits of technology

[0042]本发明的有益效果:基于胶囊网络模型对骨骼序列图像的动作类别进行识别,使用结合了针对时空特征的注意力机制的残差时间卷积模块从骨骼数据提取时空特征,通过区分出携带目标时空特征的骨骼点来确定骨骼点的重要程度,在主胶囊子层构造主胶囊来表示骨骼点,通过动作胶囊子层迭代调整主胶囊与预设的动作胶囊之间的分配方式,使用调整后的动作胶囊提取主胶囊的动作特征,每个主胶囊可通过动态路由将输出送至动作胶囊,迭代过程将剔除导致错误分类的骨骼点信息并强调有效骨骼点,使胶囊网络模型能有效地通过关注特定骨骼点以识别对应的动作类别,能更好地捕捉骨骼点之间的关系和组件的方向,使相似动作的区分得到较大的改善,提高骨骼动作识别的准确性和可用性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116824689B_ABST
    Figure CN116824689B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of image recognition, and discloses a skeleton sequence behavior recognition method, device and equipment and a storage medium. The method comprises the following steps: acquiring skeleton data of a skeleton sequence image; inputting the skeleton data into a preset capsule network model; extracting the space-time features of skeleton points, distinguishing the skeleton points carrying target space-time features according to an attention mechanism for the space-time features, and obtaining a distinguishing result; constructing a main capsule according to the distinguishing result and the space-time features of the skeleton points; iteratively adjusting the distribution mode between the main capsule and preset action capsules, extracting the space-time features of the main capsule by using the adjusted action capsules, obtaining action prediction information, and each action capsule corresponding to an action category; and aggregating the action prediction information output by the action capsule sublayers at different stages in a soft voting manner to obtain a recognition result. The application can improve the accuracy and availability of skeleton action recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image recognition technology, and in particular to a method, apparatus, device, and storage medium for skeletal sequence behavior recognition. Background Technology

[0002] In recent years, skeletal motion recognition has become a hot research area in computer vision. In many practical scenarios, such as sports competitions, medical rehabilitation, and security, the accurate recognition and classification of human movements is of great value. However, traditional image- or video-based motion recognition methods are affected by factors such as lighting and noise, thus having certain limitations in practical applications.

[0003] To address the aforementioned technical challenges, researchers have employed neural network-based action recognition methods to identify input skeletal joint data, thereby achieving skeletal action recognition. These methods exhibit strong robustness and accuracy. However, most existing neural network-based action recognition methods utilize traditional deep neural network models, such as convolutional neural networks and recurrent neural networks. These models have limitations in their ability to model complex human movement patterns and exhibit relatively low accuracy. Summary of the Invention

[0004] The purpose of this invention is to provide a method, apparatus, device, and storage medium for skeletal sequence behavior recognition, aiming to improve the accuracy and usability of skeletal action recognition.

[0005] Firstly, a method for recognizing skeletal sequence behavior is provided, including:

[0006] Acquire skeletal data from a skeletal sequence image, wherein the skeletal data consists of multiple skeletal points;

[0007] Skeletal data is input into a preset capsule network model, which includes residual temporal convolutional layers and capsule network layers. Each capsule network layer includes a main capsule sub-layer, an action capsule sub-layer, and a classification sub-layer.

[0008] Based on the residual temporal convolutional layer, the spatiotemporal features of the skeletal points are extracted. The skeletal points carrying the target spatiotemporal features are distinguished according to the attention mechanism for the spatiotemporal features, and the distinction results are obtained.

[0009] Based on the master capsule sublayer, a master capsule is constructed according to the differentiation results and the spatiotemporal characteristics of the skeletal points. The master capsule is a set of neurons that conforms to the dynamic routing mechanism.

[0010] Based on the action capsule sub-layer, the allocation method between the main capsule and the preset action capsule is iteratively adjusted. The spatiotemporal features of the main capsule are extracted using the adjusted action capsule to obtain action prediction information. Each action capsule corresponds to an action category.

[0011] Based on the classification sub-layer, the action prediction information output by the action capsule sub-layer at different stages is aggregated through soft voting to obtain the recognition result.

[0012] In some embodiments, the capsule network model further includes a preprocessing layer, and the skeletal sequence behavior recognition method further includes:

[0013] Each skeletal point is transformed into the same coordinate system for representation, resulting in coordinate transformation data.

[0014] The coordinate transformation data is normalized to obtain normalized transformation data;

[0015] Action features are extracted from the coordinate vectors of each normalized transformed data. The extracted action feature vectors are then concatenated to obtain the preprocessed data used as input to the residual temporal convolutional layer.

[0016] In some embodiments, the extraction of spatiotemporal features of skeletal points based on residual temporal convolutional layers, and the differentiation of skeletal points carrying target spatiotemporal features based on an attention mechanism targeting these features, to obtain a differentiation result, includes:

[0017] The skeletal data is mapped to a high-dimensional space, and the spatiotemporal features of the skeletal data mapped to the high-dimensional space are extracted through convolution operations to obtain the spatiotemporal feature vector.

[0018] Pooling is performed on the spatiotemporal feature vectors at the frame and joint levels to obtain pooled features.

[0019] The importance of the corresponding skeleton point is determined by the attention score of the pooling feature in the attention mechanism for spatiotemporal features. The spatiotemporal features of the skeleton point are then assigned corresponding weights based on the importance of the skeleton point, which serves as the distinction result.

[0020] In some embodiments, the attention mechanism targeting spatiotemporal features is specifically represented as follows:

[0021]

[0022] Among them, f in and f out Represents the input and output feature mapping. Indicates a connection operation. ⊙ and ⊙ represent the channel outer product and element-wise product, respectively. pool t (·) and pool v (·) represent the average pooling operations at the frame and joint levels, respectively, and σ(·) and θ(·) represent the Sigmoid and HardSwish activation functions, respectively. These are trainable parameters.

[0023] In some embodiments, the iterative adjustment of the allocation method between the master capsule and the preset action capsule, and the extraction of the spatiotemporal features of the master capsule using the adjusted action capsule to obtain action prediction information, includes:

[0024] Based on the dynamic routing mechanism, the spatiotemporal features of the main capsule are transferred to the action capsule. The action capsule learns the feature representation of the corresponding action of the main capsule, calculates the spatial relationship and direction between each action capsule, and obtains the pose matrix.

[0025] The distance and relative direction between each action capsule are calculated based on the attitude matrix. The weight allocation vector and the weighting coefficient of the weight allocation vector are iteratively updated by the reverse dynamic routing mechanism to obtain the adjusted weight allocation vector and weighting coefficient.

[0026] Based on the adjusted weight allocation vector and weighting coefficients, the action capsule is used to extract the action features of the main capsule to obtain action prediction information.

[0027] In some embodiments, the aggregation of action prediction information output by the action capsule sublayer at different stages based on the classification sublayer using soft voting to obtain the recognition result includes:

[0028] The distance between the action prediction information output by the action capsule and the preset prediction information in the training set is measured to obtain the prediction gap value;

[0029] The predicted gap value is mapped to a specific action category, the probability value of the predicted gap value relative to each action category is calculated to obtain the action probability value, and the action probability value is normalized to obtain the normalized probability value.

[0030] The normalized probability values ​​of the action category obtained at each stage are summed, and the action category with the largest summation result is taken as the recognition result.

[0031] In some embodiments, the skeletal sequence behavior recognition method further includes:

[0032] By quantifying the relationship between the three dimensions, namely configuring the width, depth and resolution of the model according to a certain scaling factor, multiple networks of different sizes are obtained. Then, by comparing the performance of networks of different sizes, a suitable model is selected to obtain the capsule network model.

[0033] Secondly, a skeletal sequence behavior recognition device is provided, the device comprising:

[0034] The first module is used to acquire bone data of a bone sequence image, wherein the bone data consists of multiple bone points;

[0035] The second module is used to input skeletal data into a preset capsule network model. The capsule network model includes residual temporal convolutional layers and capsule network layers. The capsule network layers include a main capsule sub-layer, an action capsule sub-layer, and a classification sub-layer.

[0036] The third module is used to extract the spatiotemporal features of skeletal points based on residual temporal convolutional layers, and to distinguish skeletal points carrying target spatiotemporal features based on the attention mechanism for spatiotemporal features, so as to obtain the distinction results.

[0037] The fourth module is used to construct a master capsule based on the master capsule sublayer, according to the differentiation results and the spatiotemporal characteristics of the skeletal points. The master capsule is a set of neurons that conforms to the dynamic routing mechanism.

[0038] The fifth module is used to iteratively adjust the allocation method between the main capsule and the preset action capsules based on the action capsule sub-layer. The adjusted action capsules are used to extract the spatiotemporal features of the main capsule to obtain action prediction information. Each action capsule corresponds to an action category.

[0039] The sixth module is used to aggregate the action prediction information output by the action capsule sublayer at different stages based on the classification sublayer, through soft voting, to obtain the recognition result.

[0040] Thirdly, an electronic device is provided, the electronic device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the skeletal sequence behavior recognition method described in the first aspect.

[0041] Fourthly, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program, which, when executed by a processor, implements the skeletal sequence behavior recognition method described in the first aspect.

[0042] The beneficial effects of this invention are as follows: It identifies action categories in skeletal sequence images based on a capsule network model. A residual temporal convolution module incorporating an attention mechanism targeting spatiotemporal features extracts spatiotemporal features from the skeletal data. The importance of skeletal points is determined by distinguishing those carrying target spatiotemporal features. A master capsule is constructed in the master capsule sub-layer to represent skeletal points. The allocation between the master capsule and preset action capsules is iteratively adjusted through the action capsule sub-layer. The adjusted action capsules are used to extract action features from the master capsule. Each master capsule can dynamically route its output to the action capsule. The iterative process removes skeletal point information that leads to misclassification and emphasizes effective skeletal points. This allows the capsule network model to effectively identify corresponding action categories by focusing on specific skeletal points, better capture the relationships between skeletal points and the orientation of components, significantly improves the differentiation of similar actions, and enhances the accuracy and usability of skeletal action recognition. Attached Figure Description

[0043] Figure 1 This is a flowchart of the skeletal sequence behavior recognition method provided in the first embodiment of this application.

[0044] Figure 2 This is a flowchart of the skeletal sequence behavior recognition method provided in the second embodiment of this application.

[0045] Figure 3 yes Figure 1 The flowchart for step S103.

[0046] Figure 4 yes Figure 1 The flowchart for step S105 in the process.

[0047] Figure 5 yes Figure 1 The flowchart for step S106.

[0048] Figure 6 This is a schematic diagram of the skeletal sequence behavior recognition device provided in the embodiments of this application.

[0049] Figure 7 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0051] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0053] First, let's analyze some of the terms used in this application:

[0054] Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.

[0055] Capsule Networks (CapsNet) are vector-based neural networks that preserve the pose information of the object being processed (such as precise target position, rotation, thickness, tilt, size, etc.) rather than recovering it after loss. Capsule Networks not only contain the probability of feature occurrences but also retain pose information such as position, tilt, and size that is lost in classic convolutional neural networks due to operations like pooling. Capsule Networks are composed of "capsules," where a "capsule" is a group of neurons, not a single neuron. Classic neural networks are composed of neurons, each with scalar inputs and outputs. Each capsule's inputs and outputs are vectors, and each capsule represents an attribute of the object being processed; the capsule's vector value corresponds to the value of that attribute.

[0056] In related technologies, most motion recognition methods based on skeletal joint data adopt traditional deep neural network models, such as convolutional neural networks (GCN) and recurrent neural networks (RNN). These models have certain limitations in their ability to model complex human motion patterns.

[0057] Recurrent Neural Networks (RNNs) have inherent limitations, lacking the ability to effectively learn the spatial relationships between skeletal joints. Furthermore, RNN-based methods require automatic adjustment of the observation point, increasing the complexity of addressing dynamic viewpoint changes in action recognition. Additionally, processing spatiotemporal dynamics necessitates deeper neural networks to ensure complete feature extraction; however, excessively deep RNNs suffer from unstable weights, leading to gradient explosion and vanishing gradient problems, resulting in suboptimal skeletal motion feature capture.

[0058] Convolutional Neural Networks (CNNs) require a large amount of labeled data for training, making it a very time-consuming and resource-intensive task when dealing with temporal and spatial dynamics. Furthermore, due to structural limitations, CNNs cannot preserve the spatial relationships between joints involved in an action; for example, spatial information is discarded in pooling layers, and each layer of a CNN only generates local features, leading to a loss of information about joint relationships. Finally, CNNs are typically used for image-based tasks, while action recognition tasks based on skeleton sequences are highly time-dependent, requiring a balanced and more efficient use of both spatial and temporal information. To meet the input requirements of CNNs, 3D skeleton sequence data is usually converted from vector sequences into pseudo-images. However, obtaining representations with spatiotemporal information is not easy. Solutions that encode skeleton joints into multiple 2D pseudo-images and then input them into the CNN to learn features only consider adjacent joints and cannot account for the co-occurrence of all joints in later layers. Most methods based on Graph Convolutional Neural Networks (GCNs) employ static graphs and fixed adjacency matrices to capture the spatial correlations of node feature representations. However, every node in the human body is dynamically changing during movement, and graph convolutions based on static graphs cannot dynamically capture the features of nodes. Furthermore, graph convolutional neural networks struggle to extract the underlying information behind 3D skeleton sequence data. Since skeleton data itself is spatiotemporally coupled, the connections between joints and bones should also be spatiotemporally coupled when converting skeleton data into a graph. Effectively extracting the dynamics of skeleton data in both time and space dimensions remains challenging.

[0059] Based on this, embodiments of this application provide a method, apparatus, device, and storage medium for skeletal sequence behavior recognition, aiming to improve the accuracy and usability of skeletal action recognition.

[0060] The skeletal sequence behavior recognition method, device, electronic device, and storage medium provided in this application are specifically described through the following embodiments. First, the skeletal sequence behavior recognition method in this application embodiment is described.

[0061] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0062] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0063] The skeletal sequence behavior recognition method provided in this application relates to the field of artificial intelligence technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the skeletal sequence behavior recognition method, but is not limited to the above forms.

[0064] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0065] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards of the relevant countries and regions. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirects to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data for the normal operation of the embodiments of this application obtained.

[0066] Figure 1 This is the first optional flowchart of the skeletal sequence behavior recognition method provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S101 to S106.

[0067] Step S101: Obtain the skeletal data of the skeletal sequence image.

[0068] The skeletal data consists of multiple skeletal points.

[0069] In step S101, the skeletal data of the skeletal sequence image is obtained. Specifically, the skeletal points (neck, shoulder, elbow, etc.) of the skeletal sequence image are marked from images of different frames. Through standardization processing, the three-dimensional skeletal data of each human body movement has the same number and format of skeletal points. Then, the skeletal points are connected to form a skeleton, and finally the skeletal sequence data is obtained.

[0070] Step S102: Input the skeletal data into the preset capsule network model.

[0071] The capsule network model includes residual temporal convolutional layers and capsule network layers. The capsule network layers include main capsule sub-layers, action capsule sub-layers, and classification sub-layers.

[0072] Step S103: Based on the residual temporal convolutional layer, extract the spatiotemporal features of the skeletal points, and distinguish the skeletal points carrying the target spatiotemporal features according to the attention mechanism for the spatiotemporal features to obtain the distinction result.

[0073] Among them, spatiotemporal features include the temporal and spatial features of each feature point in the skeletal data.

[0074] In step S103, the residual temporal convolutional layer is used to encode the features of each bone point according to the spatiotemporal features of each bone data, and an attention mechanism targeting the spatiotemporal features is used to distinguish the bone points with more information, that is, the bone points carrying the target spatiotemporal features, and assign different weights to each bone point as the distinction result.

[0075] Residual temporal convolutional layers are used to extract local features from the input data. They extract spatiotemporal features from the skeletal data, determine the importance of these features, and distinguish skeletal points carrying the target spatiotemporal features. Specifically, this layer uses a residual temporal convolutional network containing multiple convolutional layers and activation functions. Each convolutional layer incorporates an attention mechanism targeting the spatiotemporal features. The convolutional operations map the input skeleton data into a high-dimensional space and extract the feature vector for each skeletal point.

[0076] Understandably, the attention mechanism targeting spatiotemporal features is used to find the most critical skeletal points from the entire skeleton sequence, enhancing the model's ability to extract discriminative features. At the same time, this attention module can eliminate the step of manually dividing parts in the skeleton diagram when processing data, thereby avoiding the problem of designing appropriate pooling rules on the nodes of each part.

[0077] Step S104: Based on the main capsule sub-layer, construct the main capsule according to the differentiation results and the spatiotemporal characteristics of the skeletal points.

[0078] The main capsule is a collection of neurons that conforms to the dynamic routing mechanism.

[0079] In step S104, each main capsule in the main capsule sub-layer corresponds to a skeletal point. The input of the main capsule sub-layer is the feature vector of the corresponding skeletal point output by the residual temporal convolution module. The feature vector output by the residual temporal convolution layer is transformed into a capsule representation by a dynamic routing algorithm.

[0080] In the main capsule sub-layer, the main capsule is constructed using the feature vectors of the corresponding skeletal points. Specifically, the feature vector of a skeletal point is represented as u. i The weight matrix for information transfer between main capsule i and action capsule j is represented as W. ij Using the coupling coefficient c ij Describe the contribution strength of the output of main capsule i to action capsule j, and initialize it to obtain the log prior matrix b. init During the construction of the master capsule, the weight matrix is ​​multiplied by the input, as shown below:

[0081]

[0082] At the same time, the coupling coefficients are initialized using the softmax function:

[0083] c ij ←softmax(b ij ).

[0084] Step S105: Based on the action capsule sub-layer, iteratively adjust the allocation method between the main capsule and the preset action capsule, and use the adjusted action capsule to extract the spatiotemporal features of the main capsule to obtain action prediction information.

[0085] Each exercise capsule corresponds to an exercise category, which can be pull-ups, bent-arm hangs, sit-ups, push-ups, etc.

[0086] In step S105, multiple preset action capsules are used, each representing a specific action category. When the action capsule sub-layer receives the feature vector output by the main capsule sub-layer, each action capsule calculates the spatial relationship and direction with other action capsules and represents this information as a pose matrix and vector. The magnitude of the vector output by the action capsule represents the probability of the action in the sequence. Then, the action capsule sub-layer calculates the distance and relative direction between action capsule units through the pose matrix and learns the pattern and process of the entire action.

[0087] Preferably, the main capsule sublayer and the action capsule sublayer transmit information through a dynamic routing algorithm. This algorithm aggregates information from the main capsule and transmits it to the action capsule in an iterative mechanism. During the iteration process, by changing the coupling coefficient, it emphasizes the information of important nodes while diluting or eliminating erroneous joint information.

[0088] Step S106: Based on the classification sub-layer, the action prediction information output by the action capsule sub-layer at different stages is aggregated through soft voting to obtain the recognition result.

[0089] In step S106, the action category of the skeletal sequence image is determined based on the action prediction information. Specifically, the dynamic temporal rule algorithm is used to measure the distance between the action prediction information output by the action capsule sub-layer and the capsule output vector of each category in the training set. The action prediction information output by the action capsule sub-layer is mapped to the specific action category through a fully connected layer and the Softmax function, and the probability value of each action category is calculated.

[0090] Steps S101 to S106 of this embodiment identify the action categories of skeletal sequence images based on a capsule network model. A residual temporal convolution module incorporating an attention mechanism targeting spatiotemporal features is used to extract spatiotemporal features from the skeletal data. The importance of skeletal points is determined by distinguishing those carrying target spatiotemporal features. A master capsule is constructed in the master capsule sub-layer to represent the skeletal points. The allocation between the master capsule and preset action capsules is iteratively adjusted through the action capsule sub-layer. The adjusted action capsules are used to extract the action features of the master capsule. Each master capsule can dynamically route its output to the action capsule. The iterative process removes skeletal point information that leads to misclassification and emphasizes effective skeletal points. This allows the capsule network model to effectively identify the corresponding action category by focusing on specific skeletal points, better capture the relationships between skeletal points and the orientation of components, significantly improve the differentiation of similar actions, and enhance the accuracy and usability of skeletal action recognition.

[0091] Figure 2 This is a second optional flowchart of the skeletal sequence behavior recognition method provided in the embodiments of this application. Figure 1 Based on the embodiment, the following steps are executed after step S102. Figure 2 The method may include, but is not limited to, steps S201 to S203.

[0092] The capsule network model also includes a preprocessing layer, which consists of a coordinate system transformation sublayer, a normalization sublayer, and a convolution sublayer.

[0093] Step S201: Transform each skeletal point into the same coordinate system for representation to obtain coordinate transformation data.

[0094] Step S202: Normalize the coordinate transformation data to obtain normalized transformation data.

[0095] Step S203: Extract action features from the coordinate vectors of each normalized transformed data, and concatenate the extracted action feature vectors to obtain preprocessed data for inputting the residual temporal convolutional layer.

[0096] In step S201, based on the coordinate system transformation sublayer, a transformation method of rotation followed by translation is adopted to unify the skeletal data into the same coordinate system. Specifically, the coordinates of the first skeletal point are used as the origin of the coordinate system. The vector between the second and first skeletal points is rotated along the Y-axis to lie on the XY plane. Then, the skeletal data is translated along the X and Y axes to lie within a standard rectangular area. The joints in the skeletal sequence are represented as fixed-length vectors, calculated using the following formula:

[0097]

[0098] Among them, v i This indicates the position of joint point i in the x, y, and z directions. Let a represent the velocities of skeletal point i in the x, y, and z directions, respectively. i This represents other attributes of joint i, such as angle and acceleration, while j represents the number of bone points.

[0099] Operations such as translation, rotation, and scaling on a skeletal sequence are defined as follows:

[0100] Translation: P(v) i ) = v i +t,

[0101] Rotation: R(v) i )=R*v i ,

[0102] Scaling: S(v) i )=s*v i ,

[0103] Wherein, P(v i ) indicates the relationship between the skeletal point v i The result after the translation operation, R(v) i ) indicates the relationship between the skeletal point v i The result after the rotation operation, S(v) i ) indicates the relationship between the skeletal point v i The result after scaling is given, where t represents the translation vector, R represents the rotation matrix, and s represents the scaling factor.

[0104] In step S202, based on the normalization sublayer, the skeletal data after coordinate system transformation is normalized. Specifically, all skeletal data is divided by height and scaled to the range of [-1, 1]. The advantage of doing this is that it can eliminate the differences in body size between different people and make the model more robust.

[0105] In step S203, based on the convolutional sub-layer, the convolutional kernel is used to extract features from the normalized skeletal data and encode the skeletal sequence. Specifically, a bidirectional time series network is used as the convolutional sub-layer to convolve the skeletal data in both forward and backward directions. The time series network extracts features from the coordinates of each joint, and then concatenates these feature vectors into a fixed-length vector, which is then input into the subsequent capsule network for further processing.

[0106] The concatenation of all the vectors at the key points into a fixed-length vector can be represented as:

[0107] V = [v1, v2, ..., v] j ]

[0108] Where V represents the encoded vector, and j represents the number of joints in the skeleton sequence.

[0109] In summary, the three data processing operations are all for preprocessing the skeletal data to facilitate subsequent network training and recognition. Among them, coordinate system transformation and normalization operations can eliminate differences between different acquisition devices and different people, making the model more robust.

[0110] Please see Figure 3 In some embodiments, step S103 may include, but is not limited to, steps S301 to S303.

[0111] Step S301: Map the skeletal data to a high-dimensional space, and extract the spatiotemporal features of the skeletal data mapped to the high-dimensional space through convolution operations to obtain the spatiotemporal feature vector.

[0112] Step S302: Perform pooling processing on the spatiotemporal feature vectors at the frame level and joint level to obtain pooled features.

[0113] Step S303: Determine the importance of the corresponding skeleton point based on the attention score of the pooling feature in the attention mechanism for spatiotemporal features, and assign corresponding weights to the spatiotemporal features of the skeleton point according to the importance of the skeleton point as the distinction result.

[0114] In this embodiment, a residual temporal convolutional module is used to extract action features. This module consists of three consecutive temporal convolutional layers and a max-pooling layer. Residual connections exist between the temporal convolutional layers and the max-pooling layer. Following this is an attention mechanism targeting spatiotemporal features, which assigns corresponding weights to the encoded feature vectors of different key points based on their importance. The attention module identifies the most critical bone points from the entire skeletal sequence, enhancing the model's ability to extract discriminative features. Furthermore, the attention module eliminates the need for manually dividing the skeleton into parts during preprocessing, thus avoiding the need to design appropriate pooling rules for each node.

[0115] In this embodiment, the attention mechanism targeting spatiotemporal features is specifically represented as follows:

[0116]

[0117] Among them, f in and f out Represents the input and output feature mapping. Indicates a connection operation. ⊙ and ⊙ represent the channel outer product and element-wise product, respectively. pool t (·) and pool v(·) represent the average pooling operations at the frame and joint levels, respectively, and σ(·) and θ(·) represent the Sigmoid and HardSwish activation functions, respectively. These are trainable parameters.

[0118] Please see Figure 4 In some embodiments, step S105 may include, but is not limited to, steps S401 to S403.

[0119] Step S401: Based on the dynamic routing mechanism, the spatiotemporal features of the main capsule are transmitted to the action capsule. The action capsule learns the feature representation of the corresponding action of the main capsule, calculates the spatial relationship and direction between each action capsule, and obtains the pose matrix.

[0120] Step S402: Calculate the distance and relative direction between each action capsule based on the attitude matrix, and iteratively update the weight allocation vector and weighting coefficient of each action capsule passed upward through the reverse dynamic routing mechanism to obtain the adjusted weight allocation vector and weighting coefficient.

[0121] Step S403: Based on the adjusted weight allocation vector and weighting coefficients, use the action capsule to extract the action features of the main capsule to obtain action prediction information.

[0122] In step S401, the master capsule sublayer, constructed using the motion features of the target skeletal points, is passed to the action capsule. The representation of the motion features corresponding to each action capsule is learned using the calculated output vector and feature matrix. Specifically, each action capsule corresponds to a specific action class, and the vector output by action capsule j is represented as s. j For each output vector, the squash function is used to compress the magnitude of the vector output by the action capsule to between 0 and 1. The magnitude of the compressed output vector then represents the probability of the action existing. The squash function's compressed vector representation is as follows:

[0123]

[0124] The pose matrix is ​​used to capture pose differences between different actions, and the calculation formula is as follows:

[0125]

[0126] Among them, squash(s j ) represents the output vector of the j-th action capsule, s j Let u represent the input vector of the j-th action capsule. i and u j Let M represent the prediction vectors for the i-th action capsule and the j-th action capsule, respectively. i,jLet represent the pose matrix from the j-th action capsule to the ith action capsule.

[0127] Understandably, the pose matrix is ​​generated from each action capsule in the action capsule sub-layer. The pose matrix can represent the rotation, translation, and scaling information of each action capsule, which varies between different actions. The pose matrix can be viewed as a distance metric between representations of different actions, helping the model to better distinguish between them.

[0128] The information transfer between the master capsule and the action capsule is performed using a dynamic routing algorithm, which is described in detail below:

[0129]

[0130]

[0131] In step S402, the distance and relative direction between each action capsule are calculated based on the pose matrix. Specifically, this is done by using the rotation, translation, and scaling information of the action capsules represented by the pose matrix to calculate the distance and relative direction between each action capsule. The weight allocation vector and the weighting coefficient of the weight allocation vector are iteratively updated through the reverse dynamic routing mechanism. Specifically, after the output vector of the bottom action capsule is calculated, the pose matrix of the top action capsule and the weight allocation vector of the current action capsule are calculated. Then, the weight allocation and routing between capsules are iteratively updated repeatedly to obtain a suitable weight allocation vector. This can help the model learn the complex features of the input data better. In the reverse dynamic routing algorithm, the capsule network layer learns a weight coefficient matrix and dynamically adjusts the distribution of input data between capsules to achieve better input feature learning.

[0132] Please see Figure 5 In some embodiments, step S106 may include, but is not limited to, steps S501 to S503.

[0133] Step S501: Measure the distance between the action prediction information output by the action capsule and the preset prediction information in the training set to obtain the prediction gap value.

[0134] Step S502: Map the predicted gap value to a specific action category, calculate the probability value of the predicted gap value relative to each action category to obtain the action probability value, and normalize the action probability value to obtain the normalized probability value.

[0135] Step S503: Sum the normalized probability values ​​of the action categories obtained at each stage, and take the action category with the largest sum of probabilities as the recognition result.

[0136] In this embodiment, the classifier layer includes two sub-classifiers: a classifier based on dynamic time warping and a classifier based on a fully connected layer and a Softmax function.

[0137] In step S501, the distance matrix between the action prediction information and the preset prediction information is first calculated, where the matrix elements represent the distance between the output vector of an action capsule in the action prediction information and the output vector of an action capsule in the training sequence. Then, the dynamic programming algorithm is used to calculate the minimum distance path from the starting position of the current action capsule sub-layer to any position in the training sequence. Finally, the dynamic time warping algorithm distance between the action capsule sub-layer and the training sequence is calculated by adding up the distances of all paths to obtain the prediction gap value.

[0138] Step S502 is implemented based on a classifier using a fully connected layer and a Softmax function. Its goal is to map the input to the probability distribution of the output label. The fully connected layer and the Softmax function are used as a feature matching-based classifier to map the input action feature vector (action prediction information) to the probability distribution of the output label, thereby classifying the input sequence into the correct action category.

[0139] In step S502, the input vector of the fully connected layer is the output vector (action prediction information) of the action capsule sub-layer. The weight matrix and bias vector are the model parameters to be trained. The output vector of the fully connected layer is used as the input of the classifier layer to calculate the probability value of each category. The Softmax function converts the output vector of the fully connected layer into the probability value of each category, and finally outputs the prediction result. More specifically, the output vector of the action capsule sub-layer is mapped to a specific action category through the fully connected layer and the Softmax function, and the probability value of each category is calculated. The calculation formula of the classifier layer is as follows:

[0140] z = Wv + by = Softmax(z);

[0141] Where v represents the output vector of the capsule network layer, W and b represent the weight matrix and bias vector, respectively, z represents the output vector of the fully connected layer, and y represents the probability value of each category;

[0142] The formula for calculating the Softmax function is as follows:

[0143]

[0144] Where z represents the output vector of the fully connected layer, K represents the number of classes, and Softmax(z) = 1 / K. i ) represents the probability value of the i-th category, which is used as the normalized probability value.

[0145] In step S503, the actions corresponding to the skeletal data are classified according to the magnitude of the normalized probability value, thereby determining the action category of the skeletal sequence image.

[0146] In some embodiments, the skeletal sequence behavior recognition method further includes: quantifying the relationship between the three dimensions, i.e., configuring the width, depth and resolution of the model according to a certain scaling factor, to obtain multiple networks of different sizes, and then selecting a suitable model by comparing the performance of networks of different sizes to obtain a capsule network model.

[0147] Specifically, an optimal set of parameters is obtained based on neural architecture search technology as composite coefficients φ to quantify the relationship between the network's depth, width, and resolution, and to uniformly scale the network's depth, width, and resolution. This results in multiple networks of different sizes. By comparing and validating the performance of networks of different sizes, the best-performing network size can be selected, or the model can be adapted to different scales and types of data, significantly improving the model's accuracy and efficiency.

[0148] In this embodiment, the composite parameters of the strategy are defined as follows:

[0149] depth:d=α φ ;

[0150] width:w = β φ

[0151] resolution: r = γ φ ;

[0152] stα·β 2 ·γ 2 ≈2;

[0153] α≥1, β≥1, γ≥1;

[0154] Here, α, β, and γ are a set of parameters that need to be solved. This is a constrained optimal parameter solution, which measures the weight of depth, width, and resolution, respectively. The constraint includes a square because if the width or resolution is doubled, the computational cost increases fourfold; however, if the depth is doubled, the computational cost only increases twofold.

[0155] Please see Figure 6 This application also provides a skeletal sequence behavior recognition device, which can implement the above-described skeletal sequence behavior recognition method. The device includes:

[0156] The first module 601 is used to acquire bone data of a bone sequence image, wherein the bone data consists of multiple bone points;

[0157] The second module 602 is used to input skeletal data into a preset capsule network model. The capsule network model includes a residual temporal convolutional layer and a capsule network layer. The capsule network layer includes a main capsule sub-layer, an action capsule sub-layer, and a classification sub-layer.

[0158] The third module 603 is used to extract the spatiotemporal features of skeletal points based on residual temporal convolutional layers, and to distinguish skeletal points carrying target spatiotemporal features based on the attention mechanism for spatiotemporal features, so as to obtain the distinction result.

[0159] The fourth module 604 is used to construct a master capsule based on the master capsule sublayer, according to the differentiation results and the spatiotemporal characteristics of the skeletal points. The master capsule is a set of neurons that conforms to the dynamic routing mechanism.

[0160] The fifth module 605 is used to iteratively adjust the allocation method between the main capsule and the preset action capsules based on the action capsule sub-layer, and use the adjusted action capsules to extract the spatiotemporal features of the main capsule to obtain action prediction information. Each action capsule corresponds to an action category.

[0161] The sixth module, 606, is used to aggregate the action prediction information output by the action capsule sublayer at different stages based on the classification sublayer through soft voting to obtain the recognition result.

[0162] The specific implementation of the skeletal sequence behavior recognition device is basically the same as the specific implementation of the skeletal sequence behavior recognition method described above, and will not be repeated here.

[0163] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described skeletal sequence behavior recognition method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0164] Please see Figure 7 , Figure 7 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0165] The processor 701 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0166] The memory 702 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 702 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 702 and is called and executed by the processor 701 using the skeletal sequence behavior recognition method of the embodiments of this application.

[0167] The input / output interface 703 is used to implement information input and output;

[0168] The communication interface 704 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0169] Bus 705 transmits information between various components of the device (e.g., processor 701, memory 702, input / output interface 703, and communication interface 704);

[0170] The processor 701, memory 702, input / output interface 703 and communication interface 704 are connected to each other within the device via bus 705.

[0171] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described skeletal sequence behavior recognition method.

[0172] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0173] The skeletal sequence behavior recognition method, apparatus, device, and storage medium provided in this application embodiment recognize the action categories of skeletal sequence images based on a capsule network model. It uses a residual temporal convolution module incorporating an attention mechanism targeting spatiotemporal features to extract spatiotemporal features from skeletal data. The importance of skeletal points is determined by distinguishing those carrying target spatiotemporal features. A master capsule is constructed in the master capsule sub-layer to represent skeletal points. The allocation between the master capsule and preset action capsules is iteratively adjusted through the action capsule sub-layer. The adjusted action capsules are used to extract action features from the master capsule. Each master capsule can dynamically route its output to the action capsule. The iterative process removes skeletal point information that leads to misclassification and emphasizes effective skeletal points. This allows the capsule network model to effectively identify corresponding action categories by focusing on specific skeletal points, better capture the relationships between skeletal points and the orientation of components, significantly improve the differentiation of similar actions, and enhance the accuracy and usability of skeletal action recognition.

[0174] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0175] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0176] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0177] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0178] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0179] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0180] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0181] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0182] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0183] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0184] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method for recognizing behaviors based on skeletal sequences, characterized in that, include: Acquire skeletal data from a skeletal sequence image, wherein the skeletal data consists of multiple skeletal points; Skeletal data is input into a preset capsule network model, which includes residual temporal convolutional layers and capsule network layers. Each capsule network layer includes a main capsule sub-layer, an action capsule sub-layer, and a classification sub-layer. Based on the residual temporal convolutional layer, the spatiotemporal features of the skeletal points are extracted. The skeletal points carrying the target spatiotemporal features are distinguished according to the attention mechanism for the spatiotemporal features, and the distinction results are obtained. Based on the master capsule sublayer, a master capsule is constructed according to the differentiation results and the spatiotemporal characteristics of the skeletal points. The master capsule is a set of neurons that conforms to the dynamic routing mechanism. Based on the action capsule sub-layer, the spatiotemporal features of the main capsule are transferred to the action capsule based on the dynamic routing mechanism. The action capsule learns the feature representation of the corresponding action of the main capsule, calculates the spatial relationship and direction between each action capsule, and obtains the pose matrix. The pose matrix is ​​used to capture pose differences between different actions, and the calculation formula is as follows: in, Indicates the first The output vector of each action capsule Indicates the first The input vector of each action capsule, and They represent the first The first action capsule and the first The prediction vector of each action capsule, Indicates from the first The action capsule to the first The posture matrix of each action capsule; The distance and relative direction between each action capsule are calculated based on the attitude matrix. The rotation, translation and scaling information of the action capsules represented by the attitude matrix are used to calculate the distance and relative direction between each action capsule. The weight allocation vector and weighting coefficient of each action capsule are iteratively updated by the reverse dynamic routing mechanism. After the output vector of the bottom action capsule is calculated, the pose matrix of the top action capsule and the weight allocation vector of the current action capsule are calculated. Then, the weight allocation and routing between capsules are iteratively updated repeatedly to finally obtain a suitable weight allocation vector. In the reverse dynamic routing algorithm, the capsule network layer will learn a weight coefficient matrix to dynamically adjust the distribution of input data between capsules. Based on the updated weight allocation vector and weighting coefficients, action capsules are used to extract action features from the main capsule to obtain action prediction information. Each action capsule corresponds to an action category. Based on the classification sub-layer, the distance matrix between the action prediction information and the preset prediction information is first calculated, where the matrix elements represent the distance between the output vector of an action capsule in the action prediction information and the output vector of an action capsule in the training sequence. Then, the dynamic programming algorithm is used to calculate the minimum distance path from the starting position of the current action capsule sub-layer to any position in the training sequence. Finally, the dynamic time warping algorithm distance between the action capsule sub-layer and the training sequence is calculated by adding up the distances of all paths, and the prediction gap value is obtained. The classifier layer contains two sub-classifiers: a classifier based on dynamic time warping and a classifier based on a fully connected layer and a softmax function; The predicted gap value is mapped to a specific action category, and the probability value of the predicted gap value relative to each action category is calculated to obtain the action probability value. The action probability value is then normalized to obtain the normalized probability value. A classifier based on fully connected layers and the Softmax function is implemented, and its goal is to map the input to the probability distribution of the output label. The fully connected layers and the Softmax function are used as a feature matching-based classifier to map the input action prediction information to the probability distribution of the output label, thereby classifying the input sequence into the correct action category. The normalized probability values ​​of the action category obtained at each stage are summed, and the action category with the largest summation result is taken as the recognition result.

2. The skeletal sequence behavior recognition method according to claim 1, characterized in that, The capsule network model further includes a preprocessing layer, and the skeletal sequence behavior recognition method further includes: Each skeletal point is transformed into the same coordinate system for representation, resulting in coordinate transformation data. The coordinate transformation data is normalized to obtain normalized transformation data; Action features are extracted from the coordinate vectors of each normalized transformed data. The extracted action feature vectors are then concatenated to obtain the preprocessed data used as input to the residual temporal convolutional layer.

3. The skeletal sequence behavior recognition method according to claim 1, characterized in that, The residual temporal convolutional layer extracts the spatiotemporal features of skeletal points, and distinguishes skeletal points carrying target spatiotemporal features based on an attention mechanism targeting these features, yielding the distinction results, including: The skeletal data is mapped to a high-dimensional space, and the spatiotemporal features of the skeletal data mapped to the high-dimensional space are extracted through convolution operations to obtain the spatiotemporal feature vector. Pooling is performed on the spatiotemporal feature vectors at the frame and joint levels to obtain pooled features. The importance of the corresponding skeleton point is determined by the attention score of the pooling feature in the attention mechanism for spatiotemporal features. The spatiotemporal features of the skeleton point are then assigned corresponding weights based on the importance of the skeleton point, which serves as the distinction result.

4. The skeletal sequence behavior recognition method according to claim 3, characterized in that, The attention mechanism targeting spatiotemporal features is specifically represented as follows: , , in, and This represents the input and output feature mapping, and ⊕ represents the join operation. and These represent the channel outer product and the element-wise product, respectively. and These are the frame-level and joint-level average pooling operations, respectively. and These represent the Sigmoid and HardSwish activation functions, respectively. These are trainable parameters.

5. The skeletal sequence behavior recognition method according to claim 1, characterized in that, The skeletal sequence behavior recognition method further includes: By quantifying the relationship between the three dimensions, namely configuring the width, depth and resolution of the model according to a certain scaling factor, multiple networks of different sizes are obtained. Then, by comparing the performance of networks of different sizes, a suitable model is selected to obtain the capsule network model.

6. A skeletal sequence behavior recognition device, characterized in that, The device includes: The first module is used to acquire bone data of a bone sequence image, wherein the bone data consists of multiple bone points; The second module is used to input skeletal data into a preset capsule network model. The capsule network model includes residual temporal convolutional layers and capsule network layers. The capsule network layers include a main capsule sub-layer, an action capsule sub-layer, and a classification sub-layer. The third module is used to extract the spatiotemporal features of skeletal points based on residual temporal convolutional layers, and to distinguish skeletal points carrying target spatiotemporal features based on the attention mechanism for spatiotemporal features, so as to obtain the distinction results. The fourth module is used to construct a master capsule based on the master capsule sublayer, according to the differentiation results and the spatiotemporal characteristics of the skeletal points. The master capsule is a set of neurons that conforms to the dynamic routing mechanism. The fifth module is used for: Based on the action capsule sub-layer, the spatiotemporal features of the main capsule are transferred to the action capsule based on the dynamic routing mechanism. The action capsule learns the feature representation of the corresponding action of the main capsule, calculates the spatial relationship and direction between each action capsule, and obtains the pose matrix. The pose matrix is ​​used to capture pose differences between different actions, and the calculation formula is as follows: in, Indicates the first The output vector of each action capsule Indicates the first The input vector of each action capsule, and They represent the first The first action capsule and the first The prediction vector of each action capsule, Indicates from the first The action capsule to the first The posture matrix of each action capsule; The distance and relative direction between each action capsule are calculated based on the attitude matrix. The rotation, translation and scaling information of the action capsules represented by the attitude matrix are used to calculate the distance and relative direction between each action capsule. The weight allocation vector and weighting coefficient of each action capsule are iteratively updated by the reverse dynamic routing mechanism. After the output vector of the bottom action capsule is calculated, the pose matrix of the top action capsule and the weight allocation vector of the current action capsule are calculated. Then, the weight allocation and routing between capsules are iteratively updated repeatedly to finally obtain a suitable weight allocation vector. In the reverse dynamic routing algorithm, the capsule network layer will learn a weight coefficient matrix to dynamically adjust the distribution of input data between capsules. Based on the updated weight allocation vector and weighting coefficients, action capsules are used to extract action features from the main capsule to obtain action prediction information. Each action capsule corresponds to an action category. Module 6 is used for: Based on the classification sub-layer, the distance matrix between the action prediction information and the preset prediction information is first calculated, where the matrix elements represent the distance between the output vector of an action capsule in the action prediction information and the output vector of an action capsule in the training sequence. Then, the dynamic programming algorithm is used to calculate the minimum distance path from the starting position of the current action capsule sub-layer to any position in the training sequence. Finally, the dynamic time warping algorithm distance between the action capsule sub-layer and the training sequence is calculated by adding up the distances of all paths, and the prediction gap value is obtained. The classifier layer contains two sub-classifiers: a classifier based on dynamic time warping and a classifier based on a fully connected layer and a softmax function; The predicted gap value is mapped to a specific action category, and the probability value of the predicted gap value relative to each action category is calculated to obtain the action probability value. The action probability value is then normalized to obtain the normalized probability value. A classifier based on fully connected layers and the Softmax function is implemented, and its goal is to map the input to the probability distribution of the output label. The fully connected layers and the Softmax function are used as a feature matching-based classifier to map the input action prediction information to the probability distribution of the output label, thereby classifying the input sequence into the correct action category. The normalized probability values ​​of the action category obtained at each stage are summed, and the action category with the largest summation result is taken as the recognition result.

7. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the skeletal sequence behavior recognition method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the skeletal sequence behavior recognition method according to any one of claims 1 to 5.