Motion recognition system and method thereof
The high-dimensional implicit features of yoga movements are extracted through deep learning technology, which solves the problem that existing systems are difficult to identify beginners’ irregular yoga movements, provides accurate real-time feedback and guidance, and improves the effectiveness of yoga practice.
Patent Information
- Application Number
- CN202510553664.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing yoga action recognition system is difficult to accurately identify and evaluate beginners’ irregular yoga movements, which makes the practice effect counterproductive.
Using deep learning-based image processing technology, yoga standard action images are collected, and high-dimensional implicit features are extracted using convolutional neural networks of different scales, and combined with Gaussian density maps and block structure feature extraction modules to perform yoga action matching scores.
Accurate detection and analysis of yoga movements is realized, real-time feedback and guidance are provided to users, and the accuracy and reliability of recognition are improved.
Smart Images

Figure CN120340135A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent recognition technology, and more specifically, to an action recognition system and method thereof. Background Art
[0002] Yoga is a series of self-weight postures and movement processes. Yoga postures use ancient and easy-to-master techniques, which have the ability to help people release stress, enhance self-confidence, promote blood circulation, enhance endurance and body flexibility, shape the body and lose weight. It is a way of exercise that achieves the harmonious unity of body, mind and spirit. Yoga has developed to the present day and is highly regarded because of its obvious effects on psychological stress reduction and physical health care. It has become a widely spread physical and mental exercise method in the world.
[0003] Due to the modern lifestyle, most people practice yoga at home. However, for beginners, incorrect yoga postures may have the opposite effect on the practice effect. An action recognition system is a system that uses computer vision and machine learning technologies to automatically detect and analyze yoga actions. It can monitor and analyze the user's yoga postures in real time, and provide accurate feedback and guidance to ensure correct yoga practice. However, due to the large number of yoga action categories and uncommon postures, some existing action recognition systems are difficult to effectively recognize and evaluate.
[0004] Therefore, there is an expectation for an action recognition system and method thereof. Summary of the Invention
[0005] In order to solve the above technical problems, this application is proposed. Embodiments of this application provide an action recognition system and method thereof. First, it collects multiple yoga standard action images, uses deep learning-based image processing technology to extract high-dimensional implicit features from the multiple yoga standard action images and the yoga action images to be recognized, and uses the features of the yoga action images to be recognized as query features to query the matching degree between the yoga action to be recognized and the standard action from the yoga standard action feature library, and then scores the yoga action. In this way, it can accurately detect and analyze yoga actions, provide real-time feedback and guidance for users, and has high accuracy and reliability.
[0006] Correspondingly, according to one aspect of this application, an action recognition system is provided, which includes: An image acquisition module, configured to collect yoga standard action images to establish a yoga standard action set, and obtain yoga action images to be recognized through a camera; A standard action feature extraction module, configured to respectively pass each yoga standard action image in the yoga standard action set through a first convolutional neural network model including a first convolutional layer and a second convolutional layer to obtain a plurality of standard action feature vectors, wherein the first convolutional layer uses a two-dimensional convolutional kernel with a first scale, the second convolutional layer uses a two-dimensional convolutional kernel with a second scale, and the first scale is different from the second scale; A Gaussian fusion module, configured to fuse the plurality of standard action feature vectors based on a Gaussian density map to obtain a standard action feature matrix; A feature enhancement module, configured to pass the standard action feature matrix through a second convolutional neural network model including a block structure feature extraction module to obtain an enhanced reference data feature matrix; A to-be-recognized action feature extraction module, configured to pass the to-be-recognized yoga action image through the first convolutional neural network model including the first convolutional layer and the second convolutional layer to obtain a query feature vector; A query module, configured to multiply the query feature vector by the enhanced reference data feature matrix to obtain a decoded feature vector; A latent space mapping adjustment module, configured to perform latent space mapping adjustment based on principal component reconstruction on the decoded feature vector to obtain an optimized decoded feature vector; An identification result generation module, configured to decode and regress the optimized decoded feature vector through a decoder to obtain a decoded value, and the decoded value is used to represent the score of the yoga action.
[0007] In the above action recognition system, the standard action feature extraction module includes: a first-scale standard image encoding unit, configured to respectively perform, in the forward pass of each layer of the first convolutional layer of the first convolutional neural network model, the following operations on the input data: performing convolutional processing, global average pooling processing, and non-linear activation processing on the input data based on the two-dimensional convolutional kernel with the first scale to obtain a first-scale standard action feature vector; a second-scale standard image encoding unit, configured to respectively perform, in the forward pass of each layer of the second convolutional layer of the first convolutional neural network model, the following operations on the input data: performing convolutional processing, global average pooling processing, and non-linear activation processing on the input data based on the two-dimensional convolutional kernel with the second scale to obtain a second-scale standard action feature vector; a standard image feature fusion unit, configured to fuse the first-scale standard action feature vector and the second-scale standard action feature vector to obtain the standard action feature vector.
[0008] In the above action recognition system, the Gaussian fusion module includes: a fusion Gaussian density map construction unit, configured to use a Gaussian density map to fuse the plurality of standard action feature vectors according to the following fusion formula to obtain a fusion Gaussian density map; wherein the fusion formula is: ; Wherein, represents the position-wise mean vector among the multiple standard action feature vectors, and the value of each position of represents the variance among the feature values of each position in the multiple standard action feature vectors, represents the Gaussian density probability function, represents the variable of the Gaussian density map; a Gaussian discretization unit for discretizing the Gaussian distribution of each position of the fused Gaussian density map to obtain the standard action feature matrix.
[0009] In the above action recognition system, the feature enhancement module is used for: using each layer of the second convolutional neural network model including the block structure feature extraction module to respectively perform on the input data during the forward pass of the layer: performing convolution processing on the input data based on the convolution kernel to obtain a convolution feature map; performing global average pooling on each feature matrix along the channel dimension of the convolution feature map to obtain a channel feature vector; passing the channel feature vector through the Softmax function to obtain a normalized channel feature vector; using the feature values of each position in the normalized channel feature vector as weights to weight the feature matrix along the channel dimension of the convolution feature map to obtain a channel attention map; respectively performing average pooling and max pooling along the channel dimension on the channel attention map to obtain an average feature matrix and a max feature matrix; concatenating and channel adjusting the average feature matrix and the max feature matrix to obtain a channel feature matrix; using the convolutional layer of the spatial attention module to perform convolutional encoding on the channel feature matrix to obtain a convolutional feature matrix; passing the convolutional feature matrix through the Softmax function to obtain a spatial attention score matrix; performing position-wise multiplication of the spatial attention score matrix and each feature matrix along the channel dimension of the channel attention map to obtain the channel-spatial attention map; performing pooling along the channel dimension on the channel-spatial attention map to obtain a pooled feature map; performing non-linear activation on the pooled feature map to obtain an activated feature map; wherein, the input of the first layer of the second convolutional neural network model including the block structure feature extraction module is the standard action feature matrix, and the output of the last layer of the second convolutional neural network model including the block structure feature extraction module is the enhanced reference data feature matrix.
[0010] In the above action recognition system, the action feature extraction module to be recognized includes: a first-scale action feature extraction unit to be recognized, configured to use each layer of the first convolutional layer of the first convolutional neural network model to respectively perform the following operations on the input data during the forward pass of the layer: performing convolutional processing, global average pooling processing, and non-linear activation processing on the input data based on the two-dimensional convolutional kernel with the first scale to obtain a first-scale action feature vector to be recognized; a second-scale action feature extraction unit to be recognized, configured to use each layer of the second convolutional layer of the first convolutional neural network model to respectively perform the following operations on the input data during the forward pass of the layer: performing convolutional processing, global average pooling processing, and non-linear activation processing on the input data based on the two-dimensional convolutional kernel with the second scale to obtain a second-scale action feature vector to be recognized; an action feature extraction fusion unit to be recognized, configured to fuse the first-scale action feature vector to be recognized and the second-scale action feature vector to be recognized to obtain the query feature vector.
[0011] In the above action recognition system, the latent space mapping and adjustment module is configured to: construct a pixel-level semantic association matrix of the decoded feature vector; perform kernel space feature extraction on the pixel-level semantic association matrix based on a convolutional layer to obtain a decoded feature latent space non-linear response matrix; perform covariance matrix decomposition on the pixel-level semantic association matrix to obtain a set of decoded feature principal component coding vectors; input each decoded feature principal component coding vector in the set of decoded feature principal component coding vectors into a feature importance calibration module based on a self-attention mechanism to obtain a set of decoded feature enhanced principal component coding vectors; map each decoded feature enhanced principal component coding vector in the set of decoded feature enhanced principal component coding vectors to the decoded feature latent space non-linear response matrix to obtain a set of decoded feature principal component latent space mask coding vectors; fuse the set of decoded feature principal component latent space mask coding vectors to obtain the optimized decoded feature vector.
[0012] In the above action recognition system, the recognition result generation module is configured to: use the decoder to perform decoding regression on the optimized decoded feature vector according to the following decoding formula to obtain the decoded value; where the decoding formula is: ; where is the optimized decoded feature vector, is the decoded value, is the weight matrix, represents matrix multiplication.
[0013] According to another aspect of the present application, there is provided an action recognition method, which includes: Collect yoga standard action images to establish a yoga standard action set, and obtain the yoga action images to be recognized through a camera; Respectively pass each yoga standard action image in the yoga standard action set through a first convolutional neural network model including a first convolutional layer and a second convolutional layer to obtain a plurality of standard action feature vectors, wherein the first convolutional layer uses a two-dimensional convolutional kernel with a first scale, the second convolutional layer uses a two-dimensional convolutional kernel with a second scale, and the first scale is different from the second scale; Fuse the plurality of standard action feature vectors based on a Gaussian density map to obtain a standard action feature matrix; Pass the standard action feature matrix through a second convolutional neural network model including a block structure feature extraction module to obtain an enhanced reference data feature matrix; Pass the yoga action image to be recognized through the first convolutional neural network model including the first convolutional layer and the second convolutional layer to obtain a query feature vector; Multiply the query feature vector by the enhanced reference data feature matrix to obtain a decoded feature vector; Perform hidden space mapping adjustment based on principal component reconstruction on the decoded feature vector to obtain an optimized decoded feature vector; Pass the optimized decoded feature vector through a decoder for decoded regression to obtain a decoded value, and the decoded value is used to represent the score of the yoga action.
[0014] In the above action recognition method, respectively passing each yoga standard action image in the yoga standard action set through a first convolutional neural network model including a first convolutional layer and a second convolutional layer to obtain a plurality of standard action feature vectors includes: using each layer of the first convolutional layer of the first convolutional neural network model to respectively perform the following operations on the input data during the forward pass of the layer: performing convolutional processing, global average pooling processing, and non-linear activation processing on the input data based on the two-dimensional convolutional kernel with the first scale to obtain a first-scale standard action feature vector; using each layer of the second convolutional layer of the first convolutional neural network model to respectively perform the following operations on the input data during the forward pass of the layer: performing convolutional processing, global average pooling processing, and non-linear activation processing on the input data based on the two-dimensional convolutional kernel with the second scale to obtain a second-scale standard action feature vector; fusing the first-scale standard action feature vector and the second-scale standard action feature vector to obtain the standard action feature vector.
[0015] In the above action recognition method, fusing the plurality of standard action feature vectors based on a Gaussian density map to obtain a standard action feature matrix includes: using a Gaussian density map to fuse the plurality of standard action feature vectors with the following fusion formula to obtain a fused Gaussian density map; wherein, the fusion formula is: ; Among them, represents the position-wise mean vector among the multiple standard action feature vectors, and the value at each position of represents the variance among the feature values at each position in the multiple standard action feature vectors, represents the Gaussian density probability function,
[0016] Compared with the prior art, the action recognition system and method provided by this application first collect multiple yoga standard action images, adopt image processing technology based on deep learning to extract high-dimensional implicit features from the multiple yoga standard action images and the yoga action images to be recognized, and use the features of the yoga action images to be recognized as query features to query the matching degree between the yoga action to be recognized and the standard actions from the yoga standard action feature library, and then score the yoga action. In this way, yoga actions can be accurately detected and analyzed, providing real-time feedback and guidance for users, with high accuracy and reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present application will become more obvious. The accompanying drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification, and are used to explain the present application together with the embodiments of the present application, and do not constitute a limitation to the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0018] Figure 1 is a block diagram of an action recognition system according to an embodiment of the present application.
[0019] Figure 2 is a schematic structural diagram of an action recognition system according to an embodiment of the present application.
[0020] Figure 3 is a block diagram of a standard action feature extraction module in an action recognition system according to an embodiment of the present application.
[0021] Figure 4 is a block diagram of an action feature extraction module to be recognized in an action recognition system according to an embodiment of the present application.
[0022] Figure 5 is a flowchart of an action recognition method according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] Hereinafter, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.
[0024] Figure 1 It is a block diagram of an action recognition system according to an embodiment of the present application. Figure 2 It is a schematic diagram of the architecture of an action recognition system according to an embodiment of the present application. As Figure 1 and Figure 2 shown, the action recognition system 100 according to an embodiment of the present application includes: an image acquisition module 110, configured to collect yoga standard action images to establish a yoga standard action set, and acquire yoga action images to be recognized through a camera; a standard action feature extraction module 120, configured to respectively pass each yoga standard action image in the yoga standard action set through a first convolutional neural network model including a first convolutional layer and a second convolutional layer to obtain a plurality of standard action feature vectors, wherein the first convolutional layer uses a two-dimensional convolutional kernel with a first scale, the second convolutional layer uses a two-dimensional convolutional kernel with a second scale, and the first scale is different from the second scale; a Gaussian fusion module 130, configured to fuse the plurality of standard action feature vectors based on a Gaussian density map to obtain a standard action feature matrix; a feature enhancement module 140, configured to pass the standard action feature matrix through a second convolutional neural network model including a block structure feature extraction module to obtain an enhanced reference data feature matrix; a to-be-recognized action feature extraction module 150, configured to pass the yoga action image to be recognized through the first convolutional neural network model including the first convolutional layer and the second convolutional layer to obtain a query feature vector; a query module 160, configured to multiply the query feature vector by the enhanced reference data feature matrix to obtain a decoded feature vector; a hidden space mapping adjustment module 170, configured to perform hidden space mapping adjustment based on principal component reconstruction on the decoded feature vector to obtain an optimized decoded feature vector; and an identification result generation module 180, configured to decode and regress the optimized decoded feature vector through a decoder to obtain a decoded value, and the decoded value is used to represent the score of the yoga action.
[0025] In the above action recognition system 100, the image acquisition module 110 is configured to collect yoga standard action images to establish a yoga standard action set, and acquire yoga action images to be recognized through a camera. As mentioned in the above background art, for beginners, the non-standard yoga actions may counteract the practice effect. Therefore, it is desired to be able to monitor and analyze the user's yoga postures in real time, and provide accurate feedback and guidance to ensure correct yoga practice. However, due to the large number of yoga action categories and uncommon postures, some existing action recognition systems are difficult to effectively recognize and evaluate.
[0026] Correspondingly, considering that when performing intelligent recognition of yoga poses, the recognition and evaluation should be based on the comparison between the yoga poses to be recognized and the standard yoga poses. Therefore, in the technical solution of this application, first, a variety of standard yoga pose images are collected, and image processing technology based on deep learning is used to extract high-dimensional implicit features from the variety of standard yoga pose images and the yoga pose images to be recognized. Then, taking the features of the yoga pose images to be recognized as query features, the matching degree between the yoga pose to be recognized and the standard pose is queried from the standard yoga pose feature library, and then the yoga pose is scored. In this way, yoga poses can be accurately detected and analyzed, providing real-time feedback and guidance for users, with high accuracy and reliability. Specifically, in the technical solution of this application, first, standard yoga pose images are collected to establish a standard yoga pose set, and, yoga pose images to be recognized are collected through a camera.
[0027] In the above action recognition system 100, the standard pose feature extraction module 120 is used to respectively pass each standard yoga pose image in the standard yoga pose set through a first convolutional neural network model including a first convolutional layer and a second convolutional layer to obtain a plurality of standard pose feature vectors. Among them, the first convolutional layer uses a two-dimensional convolutional kernel with a first scale, and the second convolutional layer uses a two-dimensional convolutional kernel with a second scale, and the first scale is different from the second scale. It should be understood that the convolutional neural network model has excellent performance in image feature extraction. In the technical solution of this application, the first convolutional layer and the second convolutional layer with convolutional kernels of different scales are used to perform convolutional operations on the image to capture action features at different scales. The first convolutional layer uses a two-dimensional convolutional kernel with a first scale to capture coarser action features at a lower level. That is, a convolutional kernel with a larger scale can better capture the features of the overall action, such as large movements of the whole body. And the second convolutional layer uses a two-dimensional convolutional kernel with a second scale to capture more detailed action features at a higher level. That is, a smaller convolutional kernel can better capture detailed information, such as local subtle movements. In this way, by using convolutional kernels of different scales, the first convolutional neural network model can extract the features of yoga poses from different abstraction levels, enabling subsequent feature fusion and classification tasks to more accurately recognize and evaluate yoga poses.
[0028] Figure 3 Block diagram of the standard pose feature extraction module in the action recognition system according to an embodiment of the present application. As Figure 3As shown, the standard action feature extraction module 120 includes: a first-scale standard image encoding unit 121, which is configured to use each layer of the first convolutional layer of the first convolutional neural network model to perform the following operations on the input data during the forward pass of the layer: performing convolutional processing, global average pooling processing, and non-linear activation processing on the input data based on the two-dimensional convolutional kernel with the first scale to obtain a first-scale standard action feature vector; a second-scale standard image encoding unit 122, which is configured to use each layer of the second convolutional layer of the first convolutional neural network model to perform the following operations on the input data during the forward pass of the layer: performing convolutional processing, global average pooling processing, and non-linear activation processing on the input data based on the two-dimensional convolutional kernel with the second scale to obtain a second-scale standard action feature vector; a standard image feature fusion unit 123, which is configured to fuse the first-scale standard action feature vector and the second-scale standard action feature vector to obtain the standard action feature vector.
[0029] In the above action recognition system 100, the Gaussian fusion module 130 is configured to fuse the multiple standard action feature vectors based on the Gaussian density map to obtain a standard action feature matrix. It should be understood that by performing probability density estimation on the multiple standard action feature vectors and constructing a Gaussian density map based on their probability distribution in the feature space, where the regions with high density represent the important regions of the feature space. In this way, the multiple standard action feature vectors are fused into the same feature space, and the relative importance of different features is reflected through the Gaussian density map to highlight important feature information and suppress interference noise, further enhancing the expression ability of the features. Then, the Gaussian density map is discretized to obtain the standard action feature matrix.
[0030] Correspondingly, in a specific example, the Gaussian fusion module 130 includes: a fused Gaussian density map construction unit, which is configured to use the Gaussian density map to fuse the multiple standard action feature vectors according to the following fusion formula to obtain a fused Gaussian density map; where the fusion formula is: ; where, represents the position-wise mean vector between the multiple standard action feature vectors, and the value of each position of represents the variance between the feature values of each position in the multiple standard action feature vectors, represents the Gaussian density probability function,
[0031] In the above action recognition system 100, the feature enhancement module 140 is configured to obtain an enhanced reference data feature matrix by passing the standard action feature matrix through a second convolutional neural network model including a block structure feature extraction module. It should be understood that the second convolutional neural network model including the block structure feature extraction module is composed of multiple convolutional layers and pooling layers, and is used to learn and represent more complex action features. By performing convolutional encoding and attention application on the standard action feature matrix, associated feature extraction and feature enhancement expression are performed on the standard action feature matrix, so as to capture higher-level action features, such as action combinations, part relationships, and overall structures, thereby improving the model's understanding and expression ability of yoga actions.
[0032] Correspondingly, in a specific example, the feature enhancement module 140 is configured to: use each layer of the second convolutional neural network model of the block structure feature extraction module to respectively perform the following operations on the input data during the forward pass of the layer: perform convolutional processing on the input data based on a convolutional kernel to obtain a convolutional feature map; perform global average pooling on each feature matrix along the channel dimension of the convolutional feature map to obtain a channel feature vector; pass the channel feature vector through a Softmax function to obtain a normalized channel feature vector; use the feature values at each position in the normalized channel feature vector as weights to weight the feature matrices along the channel dimension of the convolutional feature map to obtain a channel attention map; perform average pooling and maximum pooling along the channel dimension on the channel attention map respectively to obtain an average feature matrix and a maximum feature matrix; perform concatenation and channel adjustment on the average feature matrix and the maximum feature matrix to obtain a channel feature matrix; use the convolutional layer of the spatial attention module to perform convolutional encoding on the channel feature matrix to obtain a convolutional feature matrix; pass the convolutional feature matrix through a Softmax function to obtain a spatial attention score matrix; perform element-wise multiplication of the spatial attention score matrix and each feature matrix along the channel dimension of the channel attention map to obtain the channel-spatial attention map; perform pooling along the channel dimension on the channel-spatial attention map to obtain a pooled feature map; perform non-linear activation on the pooled feature map to obtain an activated feature map; wherein, the input of the first layer of the second convolutional neural network model of the block structure feature extraction module is the standard action feature matrix, and the output of the last layer of the second convolutional neural network model of the block structure feature extraction module is the enhanced reference data feature matrix.
[0033] In the above-mentioned action recognition system 100, the action feature extraction module 150 to be recognized is used to obtain a query feature vector by passing the yoga action image to be recognized through the first convolutional neural network model including the first convolutional layer and the second convolutional layer. It should be understood that the first convolutional neural network model is also used to extract action features from the yoga action image to be recognized to obtain a query feature vector, so as to match and recognize it with the standard action.
[0034] Figure 4 It is a block diagram of the action feature extraction module to be recognized in the action recognition system according to the embodiment of the present application. As Figure 4 shown, the action feature extraction module 150 to be recognized includes: a first-scale action feature extraction unit 151 to be recognized, which is used to use each layer of the first convolutional layer of the first convolutional neural network model to perform the following operations on the input data respectively during the forward pass of the layer: performing convolutional processing, global average pooling processing, and non-linear activation processing on the input data based on the two-dimensional convolutional kernel with the first scale to obtain a first-scale action feature vector to be recognized; a second-scale action feature extraction unit 152 to be recognized, which is used to use each layer of the second convolutional layer of the first convolutional neural network model to perform the following operations on the input data respectively during the forward pass of the layer: performing convolutional processing, global average pooling processing, and non-linear activation processing on the input data based on the two-dimensional convolutional kernel with the second scale to obtain a second-scale action feature vector to be recognized; an action feature extraction fusion unit 153 to be recognized, which is used to fuse the first-scale action feature vector to be recognized and the second-scale action feature vector to be recognized to obtain the query feature vector.
[0035] In the above-mentioned action recognition system 100, the query module 160 is used to multiply the query feature vector by the enhanced reference data feature matrix to obtain a decoded feature vector. It should be understood that through the vector multiplication operation, the standard action reference features in the enhanced reference data feature matrix are mapped into the query feature vector, so that the decoded feature vector contains the query feature and the reference feature. That is to say, the matching degree feature between the yoga action to be recognized and the standard action is queried from the standard action reference data feature library.
[0036] In the above action recognition system 100, the latent space mapping adjustment module 170 is configured to perform latent space mapping adjustment based on principal component reconstruction on the decoded feature vector to obtain an optimized decoded feature vector. Since there may be redundant information or noise interference in the high-dimensional space in the decoded feature vector, and the diversity of yoga action postures leads to a complex feature distribution, directly using the original feature vector is vulnerable to non-critical factors and it is difficult to accurately match the core feature patterns of standard actions. Based on this, in the technical solution of this application, latent space mapping adjustment based on principal component reconstruction is performed on the decoded feature vector.
[0037] Specifically, the latent space mapping adjustment module 170 is configured to: First, construct a pixel-level semantic association matrix of the decoded feature vector, which is expressed by the formula: ; Wherein, represents the decoded feature vector, and respectively represent the feature values at the th and th positions of the decoded feature vector, represents calculating the Euclidean distance, represents the feature value at the position of the pixel-level semantic association matrix.
[0038] That is, convert the decoded feature vector from discrete units into an explicit relational graph representation. By quantifying the association strength or similarity of the feature information at each pixel position, encode more complex dependency structures within the features, reveal potential patterns, groupings, and underlying organizational structures, enable the decoded feature vector to explicitly reflect pixel-level semantic associations, enhance the feature's ability to describe action posture details, thereby more accurately capture the collaborative relationships or response patterns between action features, and improve the accuracy of subsequent action recognition and evaluation.
[0039] Secondly, perform kernel space feature extraction on the pixel-level semantic association matrix based on a convolutional layer to obtain a decoded feature latent space non-linear response matrix, which is expressed by the formula: ; Wherein, represents the pixel-level semantic association matrix, represents the convolutional layer, represents the decoded feature latent space non-linear response matrix.
[0040] That is, a convolution layer is used to construct a kernel domain beyond simple second-order statistics, enabling the model to perceive complex internal correlations of features in a non-linear manner and providing a context representation containing high-level structural information for subsequent processing. Specifically, through data-driven convolutional encoding, a non-linear response matrix of the decoded feature latent space is obtained, which encodes high-level and non-linear structural information extracted from the original correlation information, can capture complex relationships such as multi-level correlations and asymmetric dependencies between feature positions, provides richer semantic information support for accurately measuring the feature matching degree between the action to be recognized and the standard action, and further enhances the model's semantic representation ability for complex yoga postures.
[0041] Then, covariance matrix decomposition is performed on the pixel-level semantic correlation matrix to obtain a set of decoded feature principal component encoding vectors, which is expressed by the formula: ; Among them, represents the transpose of the vector, represents a diagonal matrix, and respectively represent the first and the th eigenvalues on the diagonal of the diagonal matrix, represents a set of decoded feature principal component encoding vectors, , , respectively represent the first, second and the th decoded feature principal component encoding vectors.
[0042] That is, through covariance matrix decomposition, the core structural main axis of the pixel-level semantic correlation matrix is mined, and the high-dimensional correlation information is mapped to a low-dimensional space spanned by orthogonal basis vectors, thereby realizing the dimensionality reduction and structured refinement of the original correlation information. The set of decoded feature principal component encoding vectors obtained in this way constitutes the basic atomic pattern representation of the correlation structure, retains the key correlation features while removing redundant information, enables the model to focus on the core correlation patterns that dominate the action posture differences, provides a more efficient and discriminative feature input for subsequent latent space mapping adjustment and decoding regression, and further improves the accuracy and computational efficiency of complex action recognition.
[0043] Next, each decoded feature principal component encoding vector in the set of decoded feature principal component encoding vectors is input into a feature importance calibration module based on the self-attention mechanism to obtain a set of decoded feature enhanced principal component encoding vectors, which is expressed by the formula: ; Among them, represents a sequence model based on the self-attention mechanism, represents a set of decoded feature enhanced principal component encoding vectors, , , respectively represent the first, second, and the th decoded feature enhanced principal component coding vectors.
[0044] That is, the self-attention mechanism is used to dynamically evaluate the interdependence between each decoded feature principal component coding vector, adaptively adjust the attention of different principal components according to the global feature structure context, so as to screen out the more "critical" structural patterns for the current action recognition task, suppress redundant or secondary information components, make the set representation of the generated decoded feature enhanced principal component coding vectors more conform to the structural characteristics of the current yoga action, provide a more discriminative feature input for subsequent latent space mapping adjustment and decoding regression, and then improve the model's ability to capture the core correlation patterns in complex postures and the adaptive adjustment ability of the recognition process.
[0045] Then, each decoded feature enhanced principal component coding vector in the set of the decoded feature enhanced principal component coding vectors is mapped to the decoded feature latent space non-linear response matrix to obtain a set of decoded feature principal component latent space mask coding vectors, which is expressed by the formula: ; where, represents matrix multiplication, represents the feature scale of the decoded feature latent space non-linear response matrix, represents the th decoded feature enhanced principal component coding vector, represents the length of the decoded feature enhanced principal component coding vector, represents the th decoded feature principal component latent space mask coding vector.
[0046] That is, through the generalized mapping operation, the interaction and fusion of the globally decoupled principal component structure information and the locally data-driven non-linear correlation patterns are realized, so that the decoded feature enhanced principal component coding vector can perceive and absorb the complex correlation context encoded in the latent space, thereby generating an intermediate representation with both global structural skeletons and local fine correlation features. The set of decoded feature principal component latent space mask coding vectors obtained in this way not only retains the key structural change directions refined by principal component analysis, but also incorporates non-linear local patterns such as multi-level correlation and asymmetric dependence learned by the convolution kernel, forming a feature representation with richer levels, and providing composite semantic information that takes into account both global structure and local details for the subsequent recognition process.
[0047] Finally, the set of the decoded feature principal component latent space mask coding vectors is fused to obtain the optimized decoded feature vector, which is expressed by the formula: ; Among them, represents a cascaded function, , , respectively represent the first, second, and th decoded feature principal component hidden space mask encoding vectors, represents the optimized decoded feature vector.
[0048] That is, the feature information generated in different processing links with different structural perspectives is effectively aggregated to form a unified feature representation that combines global structural stability and local detail richness, providing more comprehensive semantic support for subsequent decoding regression. Specifically, the optimized decoded feature vector obtained through the fusion operation can integrate the complementary structural information in each mask encoding vector, enabling the model to comprehensively consider global structural consistency and local detail accuracy when evaluating the matching degree between the action posture and the standard action, thereby improving the overall semantic modeling ability of complex yoga actions and the accuracy of recognition scores.
[0049] In the above action recognition system 100, the recognition result generation module 180 is used to perform decoding regression on the optimized decoded feature vector through a decoder to obtain a decoded value, and the decoded value is used to represent the score of the yoga action. The decoder can map the abstract feature vector back to the original data space, convert it into an interpretable and comparable numerical value, and use it to evaluate and represent the quality or accuracy of the yoga action, so as to provide accurate feedback and guidance for users to ensure correct yoga practice.
[0050] Correspondingly, in a specific example, the recognition result generation module 180 is used to: use the decoder to perform decoding regression on the optimized decoded feature vector according to the following decoding formula to obtain the decoded value; where the decoding formula is: ; where is the optimized decoded feature vector, is the decoded value, is the weight matrix, represents matrix multiplication.
[0051] In summary, an action recognition system according to an embodiment of the present application is described. First, it collects various yoga standard action images, and uses deep learning-based image processing technology to extract high-dimensional implicit features from the various yoga standard action images and the yoga action images to be recognized. Then, taking the features of the yoga action images to be recognized as query features, it queries the matching degree between the yoga action to be recognized and the standard actions from the yoga standard action feature library, and then scores the yoga action. In this way, it can accurately detect and analyze yoga actions, provide real-time feedback and guidance for users, and has high accuracy and reliability.
[0052] Figure 5 FIG. is a flowchart of an action recognition method according to an embodiment of the present application. As Figure 5 shown, the action recognition method according to an embodiment of the present application includes the steps of: S110, collecting yoga standard action images to establish a yoga standard action set, and obtaining yoga action images to be recognized through a camera; S120, respectively passing each yoga standard action image in the yoga standard action set through a first convolutional neural network model including a first convolutional layer and a second convolutional layer to obtain a plurality of standard action feature vectors, wherein the first convolutional layer uses a two-dimensional convolutional kernel with a first scale, the second convolutional layer uses a two-dimensional convolutional kernel with a second scale, and the first scale is different from the second scale; S130, fusing the plurality of standard action feature vectors based on a Gaussian density map to obtain a standard action feature matrix; S140, passing the standard action feature matrix through a second convolutional neural network model including a block structure feature extraction module to obtain an enhanced reference data feature matrix; S150, passing the yoga action images to be recognized through the first convolutional neural network model including the first convolutional layer and the second convolutional layer to obtain a query feature vector; S160, multiplying the query feature vector by the enhanced reference data feature matrix to obtain a decoded feature vector; S170, performing a hidden space mapping adjustment based on principal component reconstruction on the decoded feature vector to obtain an optimized decoded feature vector; S180, passing the optimized decoded feature vector through a decoder for decoding regression to obtain a decoded value, and the decoded value is used to represent the score of the yoga action.
[0053] Here, those skilled in the art can understand that the specific operations of each step in the above action recognition method have been described in detail in the description of the above Figures 1 to 4 action recognition system, and therefore, the repeated description thereof will be omitted.
Claims
1. An action recognition system, characterized in that, Including: An image acquisition module, configured to collect yoga standard action images to establish a yoga standard action set, and to obtain the yoga action image to be recognized through a camera; A standard action feature extraction module, configured to respectively pass each yoga standard action image in the yoga standard action set through a first convolutional neural network model including a first convolutional layer and a second convolutional layer to obtain a plurality of standard action feature vectors, wherein the first convolutional layer uses a two-dimensional convolutional kernel with a first scale, the second convolutional layer uses a two-dimensional convolutional kernel with a second scale, and the first scale is different from the second scale; A Gaussian fusion module, configured to fuse the plurality of standard action feature vectors based on a Gaussian density map to obtain a standard action feature matrix; A feature enhancement module, configured to pass the standard action feature matrix through a second convolutional neural network model including a block structure feature extraction module to obtain an enhanced reference data feature matrix; A to-be-recognized action feature extraction module, configured to pass the yoga action image to be recognized through the first convolutional neural network model including the first convolutional layer and the second convolutional layer to obtain a query feature vector; A query module, configured to multiply the query feature vector by the enhanced reference data feature matrix to obtain a decoded feature vector; A latent space mapping adjustment module, configured to perform latent space mapping adjustment based on principal component reconstruction on the decoded feature vector to obtain an optimized decoded feature vector; An identification result generation module, configured to decode and regress the optimized decoded feature vector through a decoder to obtain a decoded value, and the decoded value is used to represent the score of the yoga action.
2. The action recognition system according to claim 1, characterized in that The standard action feature extraction module includes: A first-scale standard image encoding unit, configured to respectively perform, in the forward pass of each layer of the first convolutional layer of the first convolutional neural network model, the following operations on the input data: performing convolutional processing, global average pooling processing, and non-linear activation processing on the input data based on the two-dimensional convolutional kernel with the first scale to obtain a first-scale standard action feature vector; A second-scale standard image encoding unit, configured to respectively perform, in the forward pass of each layer of the second convolutional layer of the first convolutional neural network model, the following operations on the input data: performing convolutional processing, global average pooling processing, and non-linear activation processing on the input data based on the two-dimensional convolutional kernel with the second scale to obtain a second-scale standard action feature vector; A standard image feature fusion unit, configured to fuse the first-scale standard action feature vector and the second-scale standard action feature vector to obtain the standard action feature vector.
3. The action recognition system according to claim 2, wherein The Gaussian fusion module includes: A fused Gaussian density map construction unit, configured to use a Gaussian density map to fuse the plurality of standard action feature vectors according to the following fusion formula to obtain a fused Gaussian density map; wherein the fusion formula is: ; Among them, represents the position-wise mean vector among the multiple standard action feature vectors, and the value of each position of represents the variance between the eigenvalue of each position in the multiple standard action feature vectors, represents the variable of the Gaussian density map; A Gaussian discretization unit, configured to perform discretization processing on the Gaussian distribution at each position of the fused Gaussian density map to obtain the standard action feature matrix.
4. The action recognition system according to claim 3, characterized in that, The feature enhancement module is used for: using each layer of the second convolutional neural network model including the block structure feature extraction module to respectively perform the following operations on the input data during the forward pass of the layer: Performing convolution processing on the input data based on a convolution kernel to obtain a convolution feature map; Performing global average pooling on each feature matrix along the channel dimension of the convolution feature map to obtain a channel feature vector; Passing the channel feature vector through a Softmax function to obtain a normalized channel feature vector; Using the feature values at each position in the normalized channel feature vector as weights to weight the feature matrices along the channel dimension of the convolution feature map to obtain a channel attention map; Performing average pooling and max pooling along the channel dimension on the channel attention map respectively to obtain an average feature matrix and a max feature matrix; Cascading and channel adjusting the average feature matrix and the max feature matrix to obtain a channel feature matrix; Using the convolutional layer of the spatial attention module to perform convolutional encoding on the channel feature matrix to obtain a convolutional feature matrix; Passing the convolutional feature matrix through a Softmax function to obtain a spatial attention score matrix; Performing element-wise multiplication at each position along the channel dimension between the spatial attention score matrix and each feature matrix of the channel attention map to obtain the channel-spatial attention map; Performing pooling along the channel dimension on the channel-spatial attention map to obtain a pooled feature map; Performing non-linear activation on the pooled feature map to obtain an activated feature map; Wherein, the input of the first layer of the second convolutional neural network model of the block structure feature extraction module is the standard action feature matrix, and the output of the last layer of the second convolutional neural network model of the block structure feature extraction module is the enhanced reference data feature matrix.
5. The action recognition system according to claim 4, characterized in that The action feature to be recognized extraction module includes: The first-scale action feature to be recognized extraction unit is used for using each layer of the first convolutional layer of the first convolutional neural network model to respectively perform the following operations on the input data during the forward pass of the layer: performing convolution processing, global average pooling processing and non-linear activation processing on the input data based on the two-dimensional convolution kernel with the first scale to obtain a first-scale action feature vector to be recognized; The second-scale action feature to be recognized extraction unit is used for using each layer of the second convolutional layer of the first convolutional neural network model to respectively perform the following operations on the input data during the forward pass of the layer: performing convolution processing, global average pooling processing and non-linear activation processing on the input data based on the two-dimensional convolution kernel with the second scale to obtain a second-scale action feature vector to be recognized; The action feature to be recognized extraction fusion unit is used for fusing the first-scale action feature vector to be recognized and the second-scale action feature vector to be recognized to obtain the query feature vector.
6. The action recognition system according to claim 5, wherein The latent space mapping adjustment module is used for: Constructing a pixel-level semantic association matrix of the decoded feature vector; Performing kernel space feature extraction on the pixel-level semantic association matrix based on a convolutional layer to obtain a decoded feature latent space non-linear response matrix; Perform covariance matrix decomposition on the pixel-level semantic association matrix to obtain a set of decoded feature principal component encoding vectors; Input each decoded feature principal component encoding vector in the set of decoded feature principal component encoding vectors into a feature importance calibration module based on the self-attention mechanism to obtain a set of decoded feature enhanced principal component encoding vectors; Map each decoded feature enhanced principal component encoding vector in the set of decoded feature enhanced principal component encoding vectors to the decoded feature latent space non-linear response matrix to obtain a set of decoded feature principal component latent space mask encoding vectors; Fuse the set of decoded feature principal component latent space mask encoding vectors to obtain the optimized decoded feature vector.
7. The action recognition system according to claim 6, wherein The recognition result generation module is used to: use the decoder to perform decoding regression on the optimized decoded feature vector according to the following decoding formula to obtain the decoded value; where the decoding formula is: ; wherein is the optimized decoding feature vector, is the decoding value, is the weight matrix, represents matrix multiplication.
8. A method for action recognition, characterized in that, including: Collect yoga standard action images to establish a yoga standard action set, and obtain the yoga action images to be recognized through a camera; Pass each yoga standard action image in the yoga standard action set through a first convolutional neural network model including a first convolutional layer and a second convolutional layer to obtain a plurality of standard action feature vectors, where the first convolutional layer uses a two-dimensional convolutional kernel with a first scale, the second convolutional layer uses a two-dimensional convolutional kernel with a second scale, and the first scale is different from the second scale; Fuse the plurality of standard action feature vectors based on the Gaussian density map to obtain a standard action feature matrix; Pass the standard action feature matrix through a second convolutional neural network model including a block structure feature extraction module to obtain an enhanced reference data feature matrix; Pass the yoga action image to be recognized through the first convolutional neural network model including the first convolutional layer and the second convolutional layer to obtain a query feature vector; Multiply the query feature vector by the enhanced reference data feature matrix to obtain a decoded feature vector; Perform latent space mapping adjustment based on principal component reconstruction on the decoded feature vector to obtain an optimized decoded feature vector; Perform decoding regression on the optimized decoded feature vector through a decoder to obtain a decoded value, and the decoded value is used to represent the score of the yoga action.
9. The action recognition method according to claim 8, wherein Pass each yoga standard action image in the yoga standard action set through a first convolutional neural network model including a first convolutional layer and a second convolutional layer to obtain a plurality of standard action feature vectors, including: Use each layer of the first convolutional layer of the first convolutional neural network model to perform, respectively, in the forward pass of the layer on the input data: convolutional processing, global average pooling processing, and non-linear activation processing on the input data based on the two-dimensional convolutional kernel with the first scale to obtain a first-scale standard action feature vector; Use each layer of the second convolutional layer of the first convolutional neural network model to perform, respectively, in the forward pass of the layer on the input data: convolutional processing, global average pooling processing, and non-linear activation processing on the input data based on the two-dimensional convolutional kernel with the second scale to obtain a second-scale standard action feature vector; Fuse the first-scale standard action feature vector and the second-scale standard action feature vector to obtain the standard action feature vector.
10. The action recognition method according to claim 9, wherein Based on the Gaussian density map, fuse the multiple standard action feature vectors to obtain a standard action feature matrix, including: Use the Gaussian density map to fuse the multiple standard action feature vectors with the following fusion formula to obtain a fused Gaussian density map; wherein, the fusion formula is: ; Among them, represents the position-wise mean vector among the multiple standard action feature vectors, and the value at each position represents the variance among the feature values at each position in the multiple standard action feature vectors, represents the Gaussian density probability function, represents the variable of the Gaussian density map; Perform discretization processing on the Gaussian distribution at each position of the fused Gaussian density map to obtain the standard action feature matrix.