Repetitive motion posture counting method, device and equipment and storage medium

By dividing human joints as main sub-map in the video sequence and establishing mapping relationships using the graph attention mechanism and classification head, the problem of insufficient mapping relationship between features and action classes in the prior art is solved, and a more efficient repetitive motion posture counting is achieved.

CN120260135AInactive Publication Date: 2025-07-04QUANZHOU INST OF EQUIP MFG
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510712226.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, video-level methods and pose-level methods lack a good mapping relationship between features and action classes, resulting in low model performance.

Method used

By obtaining image frames of video sequences, the joints are divided into subject subgraphs based on the natural topology of the human body, the graph attention mechanism is used to establish local and global dependencies, and the repetitive motion posture counting is performed by combining the classification head and action trigger.

Benefits of technology

The mapping relationship between the model between features and action classes is improved, and the accuracy and efficiency of repeated motion pose counting is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260135A_ABST
    Figure CN120260135A_ABST
Patent Text Reader

Abstract

The invention is suitable for the field of computer vision, and provides a repetitive motion posture counting method, device and equipment and a storage medium. The method comprises the following steps: acquiring a video sequence, and dividing human joints of an image frame into a plurality of main sub-images based on a natural topological structure of a human body; establishing a dependency relationship in the main body sub-graphs locally and establishing a dependency relationship between the main body sub-graphs globally through graph attention to obtain representative features about the image frames; establishing a mapping relationship between the representative features and action categories by using a classification head to obtain a score of the action category of each image frame in the video data; the repetitive actions are counted using an action trigger. According to the method, the mutual relation of the sub-graphs in the components and among the components is explored through graph attention, the local significant features in the components are captured, and the global dependency relation among the components is captured, so that the model performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and particularly relates to a method, device, equipment and storage medium for counting repeated motion postures. Background Art

[0002] In the field of computer vision, the counting of repeated human actions refers to the technology of analyzing the motion trajectories of human key points in a video or image sequence, identifying the periodic changes of specific actions, and counting their repetition times. This technology is widely used in fields such as fitness, medical rehabilitation, and industrial production. For example, it can count the number of push-ups and skipping ropes, or monitor the standardized operation frequency of workers on the production line.

[0003] In the prior art, video-level methods and pose-level methods are two different data processing paradigms in computer vision. Video-level methods assign a single label (such as "running" or "swimming") to the entire video without distinguishing the specific action details within the video. For example, a basketball game video may only be labeled as "basketball movement". Compared with traditional video-level methods, pose-level methods focus on the changes in human joint coordinates and significantly improve performance. However, although pose-level methods associate poses with specific action labels, there is a lack of a good mapping relationship between features and action classes, resulting in low model performance. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide a method for counting repeated motion postures, aiming to solve the problem of the lack of a good mapping relationship between features and action classes and the low model performance.

[0005] The embodiments of the present invention are implemented as follows. A method for counting repeated motion postures, the method for counting repeated motion postures includes: Obtain a video sequence, extract a series of image frames from the video sequence, convert the extraction of the image frames into joint coordinates, and divide the human joints of the image frames into several main subgraphs based on the natural topological structure of the human body; Establish the dependency relationship within the main subgraphs locally and the dependency relationship between the main subgraphs globally through graph attention to obtain representative features of the image frames; Use a classification head to establish a mapping relationship between the representative features and action categories to obtain the scores of the action categories of each image frame in the video data; Use an action trigger to count repeated actions.

[0006] Further, convert each frame in the video sequence into joint coordinates through the Blazepose algorithm. The process is as follows: ; Wherein, V represents joint coordinates. represents each of the image frames in the video sequence, C represents the number of channels, H represents the height, W represents the width, and T represents the number of frames. represents each significant part feature point in each of the image frames, D×K×T represents the feature points of the significant parts in the image frame, D represents the dimension of each pose key point, and K represents the number of significant part feature points.

[0007] Further, based on the natural topological structure of the human body, the human joints of the image frame are divided into several body subgraphs, and the same feature extraction operation is performed on each of the body subgraphs, and the operation is as follows: ; wherein, represents that the human topological structure is different body subgraphs, represents the number of human pose key points in the j-th body subgraph, the CBR layer consists of Convolution, BatchNorm, and ReLU, and D' represents the D dimension of after being transformed by CBR, represents

[0008] Further, the graph attention is used to establish the internal dependencies of the body subgraphs locally, and the specific steps are as follows: Use the graph attention mechanism to expand the feature representation and extract features: ; wherein, represents the linear transformation matrix for obtaining the Query in the attention mechanism, respectively represent the linear transformation matrices for Key and Value in the attention mechanism, represents the Query, Key, and Value of the j-th subgraph.

[0009] Perform graph attention calculation, perform max pooling operation on the result of graph attention, repeat the feature mapping U times along the point dimension, splice the mapped features along the channel dimension, and perform average pooling along the channel dimension, and repeat the feature mapping M times on the channel dimension to obtain the output feature : ; wherein, represents the graph attention operation, Maxpooling(·) represents the max pooling operation, RP represents the feature mapping, U represents the number of times of feature mapping, represents the feature of the j-th body subgraph, Denote the features obtained by mapping and splicing all the subject subgraphs. M is the number of feature mappings, Avgpooling represents the average pooling operation, and Concat(·) represents splicing in the same dimension.

[0010] Further, establishing the dependencies between the subject subgraphs globally specifically includes: Embed the entire node into a new M-dimensional feature space through two feature embedding layers in the channel dimension and reshape the dimension to ensure that the dimensions are equal; connect the J subject subgraphs, arrange the dimensions, and then input the embedded features into the self-attention mechanism to obtain the global dependencies between the subject subgraphs; use the splicing method to and fuse them to obtain the representative features. The specific formula is as follows: ; ; where represents the node, the number of nodes is j, CBR(·) represents feature embedding, M represents the dimension, represents the global dependency, Reshape(·) represents reshaping the dimension, Contact(·) represents the combination operation, represents the representative feature.

[0011] Further, the method of using the classification head for classification is specifically: ; where the representative feature is input into a feed-forward neural network, and the feed-forward neural network consists of three cascaded and three MLPs with different parameter settings are used in the classification head. The pose mapping is completed through the linear layer to obtain the score of a single frame, where O represents the number of output channels, Flatten(·) means flattening all dimensions except the batch dimension into one dimension, and Linear(·) performs a linear transformation on the flattened features to map the features to the action categories to obtain the classification score.

[0012] Further, the repeated motion pose counting method further includes loss calculation. The elements for calculating the loss include the triple margin loss and the binary cross-entropy loss; the calculation of the triple margin loss is as follows: ; where It represents the triple boundary loss. Here, a, p, and n represent the anchor point, positive sample, and negative sample respectively. CS represents the cosine similarity, which is used to measure the similarity between features; The calculation of the binary cross-entropy loss is as follows: ; Among them, It represents the binary cross-entropy loss. N represents the batch size, where each frame constitutes a batch. y represents the true label, p is the predicted result, and c represents the number of action categories; The total loss function is as follows: ; Among them, It represents the total loss function, and α is the weighting factor that controls the two losses.

[0013] Another object of the embodiments of the present invention is a repetitive motion posture counting device, and the repetitive motion posture counting device includes: An information extraction and segmentation module, which is used to obtain a video sequence, extract a series of image frames from the video sequence, convert the extracted image frames into joint coordinates, and divide the human joints of the image frames into several main subgraphs based on the natural topological structure of the human body; A local and global feature learning module, which is used to establish the dependency relationship inside the main subgraphs locally and the dependency relationship between the main subgraphs globally through graph attention to obtain representative features of the image frames; A classification head, which is used to establish the mapping relationship between the representative features and the action categories by using the classification head to obtain the scores of the action categories of each image frame in the video data; A repetitive counting module, which is used to count repetitive actions by using an action trigger.

[0014] Another object of the embodiments of the present invention is a computer device, including a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor is caused to execute the steps of the repetitive motion posture counting method.

[0015] Another object of the embodiments of the present invention is a computer-readable storage medium. When the computer program stored on the computer-readable storage medium is executed by a processor, the processor is caused to execute the steps of the repetitive motion posture counting method.

[0016] A method for counting repeated motion postures provided by an embodiment of the present invention has the beneficial effect that: according to the natural human topology, key points are segmented into several parts, and the mutual relationships of subgraphs within and between components are explored through graph attention to capture local significant features within components and capture global dependencies between components, so as to solve the problem of counting repeated motion postures in video sequences, thereby improving the model performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is an application environment diagram of the repeated motion posture counting method provided by an embodiment of the present invention; Figure 2 It is a flowchart of the repeated motion posture counting method provided by an embodiment of the present invention; Figure 3 It is an algorithm network structure diagram of the repeated motion posture counting method provided by an embodiment of the present invention; Figure 4 It is a main body subgraph between and within components of the natural human topology segmentation provided by an embodiment of the present invention; Figure 5 It is a legend of the action trigger module provided by an embodiment of the present invention; Figure 6 It is an internal structure block diagram of a computer device in one embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0019] Figure 1 An application environment diagram of a repeated motion posture counting method provided by an embodiment of the present invention is as Figure 1 shown. In this application environment, it includes an image acquisition device and a computer device.

[0020] The image acquisition device refers to a hardware tool for capturing, recording or generating digital image sequences or videos, and can be a camera, scanner, digital camera or industrial camera, etc.

[0021] The computer device can be a smart phone, tablet computer, notebook computer or desktop computer, or can also be a cloud server providing basic cloud computing services such as cloud servers, cloud databases, cloud storage and CDN. The image acquisition device and the computer device can be connected through a network, and the present invention does not limit this here.

[0022] Figure 2 It is a flowchart of a repeated motion posture counting method, Figure 3It is the algorithm network structure diagram of this method. In one embodiment, a method for counting repeated motion postures is proposed. This embodiment mainly takes the application of this method to the computer device in the above Figure 1 as an example for illustration. A method for counting repeated motion postures may specifically include the following steps: Step S202, obtain a video sequence, extract a series of image frames from the video sequence, convert the extraction of the image frames into joint coordinates, and divide the human joints of the image frames into several main sub - graphs based on the natural topological structure of the human body.

[0023] In this embodiment, the video sequence can select an existing video file, extract image frames from it, and process each image frame. As Figure 4 shown, the natural topological structure of the human body usually refers to the connections between joint points, such as parts like the limbs and torso. The number of main sub - graphs can be determined according to specific circumstances; in this embodiment, the number of main sub - graphs is selected as 5. Step S202 converts each frame in the video sequence into joint coordinates through the Blazepose algorithm. The process is as follows: ; where V represents joint coordinates, represents each of the image frames in the video sequence, C represents the number of channels, H represents the height, W represents the width, T represents the number of frames, represents each significant part feature point in each of the image frames, D×K×T represents the feature points of the significant parts in the image frame, D represents the dimension of each pose key point, and K represents the number of significant part feature points.

[0024] Step S202 divides the human joints of the image frames into several main sub - graphs based on the natural topological structure of the human body, and performs the same feature extraction operation on each of the main sub - graphs. The operation is as follows: ; where, represents that the human topological structure is different main sub - graphs, represents the number of human pose key points in the j - th main sub - graph. The CBR layer consists of Convolution, BatchNorm, and ReLU. D' represents the D dimension of after CBR transformation, represents the new representation symbol after

[0025] Step S204, establish the dependence relationship within the main sub - graphs locally and the dependence relationship between the main sub - graphs globally through graph attention, and obtain the representative features of the image frames.

[0026] In this embodiment, step S204 includes two main steps. One is to learn the significant features within components, which is implemented by the Figure 3 significant feature learning module within components (SIFL-Module) in Figure 3 , and the other is to learn the global features between components, which is implemented by the

[0027] global feature learning module between components (GIFL-Module) in

[0028] . By exploring the mutual relationships of subgraphs within and between components through graph attention, local significant features within components are captured, and global dependencies between components are captured, solving the problem of counting repeated motion postures in video sequences, thereby improving the model performance. ; Among them, in order to predict the final classification scores of all action categories in the current frame the representative features are input into a feed-forward neural network, and the feed-forward neural network consists of three cascaded as shown in Figure 3 . In order to further refine the characteristics of the output signal, reduce the number of parameters of the model, and reduce the number of parameters in the network, three MLPs with different parameter settings are used in the classification head, and the output dimensions of each MLP are gradually reduced to I⁄2, I⁄4, and I⁄8 respectively. Finally, the posture mapping is completed through a linear layer to obtain the score of a single frame , where O represents the number of output channels. Flatten(·) means flattening all dimensions except the batch dimension into one dimension, and Linear(·) performs a linear transformation on the flattened features to map the features to the action categories to obtain the classification scores.

[0029] Step S208, using an action trigger to count repeated actions.

[0030] In this embodiment, in order to obtain the final output Y of repeated motion posture counting and keep the model lightweight, the method of this embodiment uses a lightweight Actiontrigger module. As shown in Figure 5 , first, the method of this embodiment scans all frames of the input video sequence and obtains the action scores from each frame. Subsequently, two thresholds are set respectively to judge the start and end of the action. Finally, when these two thresholds are triggered continuously, the count will be incremented by 1.

[0031] In one embodiment, as Figure 3 shown, in step S204, the dependency relationships within the main subgraph are established locally through graph attention, specifically as follows: For some fine-grained operations, the changes in the operations mainly occur within the main subgraph. Therefore, this method uses the self-attention mechanism to expand the feature representation, enabling the model to focus on more important information and learn the dependency relationships between the pose key points in the internal main subgraph. Specifically, as Figure 3 shown, three linear projection layers are used to map each to the learnable and to extract the feature representations and . The formula for using the graph attention mechanism to expand the feature representation and extract features is as follows: ; where, where, represents the linear transformation matrix for obtaining the Query in the attention mechanism, represent the linear transformation matrices for Key and Value in the attention mechanism respectively, represents the Query, Key, and Value of the j-th subgraph.

[0032] To highlight the important information in the key points from the attention map of the key point coordinates and extract the most significant features in the local area, this method performs a max-pooling operation on the attention results in the point dimension. Since the feature dimension changes after pooling, the feature mapping needs to be repeated U times along the point dimension. Subsequently, to maintain the overall trend of the internal point coordinates and be able to capture a larger range of context information, all are concatenated along the channel dimension and an average pooling is performed along the channel dimension. To maintain consistency, the feature mapping also needs to be repeated M times along the channel dimension. The output feature of this process is calculated as follows: ; where, represents the graph attention operation, Maxpooling(·) represents the max-pooling operation, RP represents the feature mapping, U represents the number of times of feature mapping, represents the feature of the j-th main subgraph, represents the feature obtained by mapping and concatenating all the main subgraphs, M is the number of times of feature mapping, Avgpooling represents the average pooling operation, and Concat(·) represents concatenation in the same dimension.

[0033] In step S204, the dependency relationships between the subject subgraphs are established globally to obtain representative features of the image frames, specifically as follows: Since the changes during the movement are the result of the work of the subject sub Figure 1 graphs, it is necessary to make full use of the mutual relationships between different subject subgraphs. Therefore, this method also needs to model the relationships between the subject subgraphs at the global level. As Figure 3 shown, the entire node is embedded into a new M-dimensional feature space through two feature embedding layers in the channel dimension to provide more significant high-level features for subsequent global relationship capture. Through dimension reshaping to ensure that the dimension is equal to the dimension; the J subject subgraphs are connected, the dimensions are arranged, and then the embedded features are input into the self-attention mechanism to obtain the global dependency relationships between the subject subgraphs. Finally, the splicing method is used to and are fused so that the model can not only obtain the local context information inside the components but also capture the global features between the components, thereby fusing multi-source information, improving the feature representation, and obtaining the representative features. The specific formulas for the above process are as follows: ; ; where, represents the node, the number of nodes is j, CBR(·) represents feature embedding, M represents the dimension, represents the global dependency relationship, Reshape(·) represents reshaping the dimension, Contact(·) represents the concatenation operation, represents the representative feature.

[0034] In one embodiment, a method for calculating the loss is proposed. The repeated motion pose counting method further includes loss calculation, and the elements for calculating the loss include the triplet margin loss and the binary cross-entropy loss; the calculation of the triplet margin loss is as follows: ; where, represents the triplet margin loss, a, p, and n respectively represent the anchor, positive sample, and negative sample, and CS represents the cosine similarity, which is used to measure the similarity between features; The calculation of the binary cross-entropy loss is as follows: ; where, It represents the binary cross-entropy loss. N represents the batch size, where each frame constitutes a batch. y represents the true label, p is the predicted result, and c represents the number of action categories; The total loss function is as follows: ; Among them, represents the total loss function, and α is the weighting factor that controls the weighting of the two losses.

[0035] In one embodiment, a specific experiment of the above-mentioned repeated motion pose counting method is given.

[0036] During the training process, the method of this embodiment uses the PyTorch-Lightning framework to train the model. This framework performs a training step before officially starting the training, monitors the change of loss in the batch processing to automatically select the initial optimal learning rate. In addition, after each epoch is completed, a validation is performed. If the validation loss does not decrease for 6 consecutive epochs, the learning rate will be automatically adjusted. In addition, the method of this embodiment sets the optimizer to Adam and uses Triplet Margin Loss and BCELoss to train the overall architecture on the NVIDIA PCle A100 GPU.

[0037] Compared with the traditional video-level method, the pose-level method focuses on the change of human joint coordinates and significantly improves the performance. However, the pose-level method ignores the natural topological structure and dependencies existing between human body parts during the movement. Therefore, the method of this embodiment will further study the change relationship of the natural topological structure of the human body during the movement on the basis of the pose-level method. As shown in Table 1, the method BIGC-Net proposed in this embodiment is compared with some state-of-the-art methods on the RepCount-pose dataset. The best results are MAE of 0.215 and OBO of 0.579. Compared with the state-of-the-art video-level method TransRAC, the method of this embodiment reduces the MAE by 22.8% and increases the OBO by 28.8%. In addition, compared with the latest method PoseRAC at the pose level, the method of this embodiment reduces the MAE by 2.1% and increases the OBO by 1.9%. The experimental results show that the model effectively learns the relationship between human body parts and establishes a good mapping relationship between features and action categories, improving the performance of the model.

[0038] Table 1: Comparison of two key objective indicators between the BIGC-Net algorithm and existing algorithms on the RepCount-pose dataset

[0039] In one embodiment, a repetitive motion posture counting device is proposed. The repetitive motion posture counting device includes: An information extraction and segmentation module, configured to obtain a video sequence, extract a series of image frames from the video sequence, convert the extraction of the image frames into joint coordinates, and divide the human joints of the image frames into several main sub - graphs based on the natural topological structure of the human body; A local and global feature learning module, configured to establish the dependency relationships within the main sub - graphs locally and the dependency relationships between the main sub - graphs globally through graph attention, to obtain representative features of the image frames; A classification head, configured to establish a mapping relationship between the representative features and action categories by using the classification head, to obtain the scores of the action categories of each image frame in the video data; A repetition counting module, configured to count repetitive actions by using an action trigger.

[0040] In this embodiment, the repetitive motion posture counting device can execute the steps of the repetitive motion posture counting method in the above - mentioned embodiment. Each module of the repetitive motion posture counting device can be divided into more detailed sub - modules according to actual needs. For example, the local and global feature learning module can be divided into a significant feature learning module within components and a global feature learning module between components. The significant feature learning module within components is configured to establish the dependency relationships within the main sub - graphs, and the global feature learning module between components is configured to establish the dependency relationships between the main sub - graphs.

[0041] Figure 6 The internal structure diagram of a computer device in one embodiment is shown. As Figure 6 shown, the computer device includes a processor, a memory, a network interface, and an input device connected through a system bus. Among them, the memory includes a non - volatile storage medium and an internal memory. The non - volatile storage medium of the computer device stores an operating system and can also store a computer program. When the computer program is executed by the processor, the processor can implement a repetitive motion posture counting method. The internal memory can also store a computer program. When the computer program is executed by the processor, the processor can execute a repetitive motion posture counting method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the outer shell of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0042] Those skilled in the art can understand that Figure 6The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0043] In one embodiment, a computer device is provided. The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented: Obtain a video sequence, extract a series of image frames from the video sequence, convert the extraction of the image frames into joint coordinates, and divide the human joints of the image frames into several body sub - graphs based on the natural topological structure of the human body; Establish the dependency relationships within the body sub - graphs locally and the dependency relationships between the body sub - graphs globally through graph attention to obtain representative features of the image frames; Use a classification head to establish a mapping relationship between the representative features and action categories to obtain the scores of the action categories of each image frame in the video data; Use an action trigger to count repeated actions.

[0044] In one embodiment, a computer - readable storage medium is provided. A computer program is stored on the computer - readable storage medium. When the computer program is executed by a processor, the processor is caused to execute the following steps: Obtain a video sequence, extract a series of image frames from the video sequence, convert the extraction of the image frames into joint coordinates, and divide the human joints of the image frames into several body sub - graphs based on the natural topological structure of the human body; Establish the dependency relationships within the body sub - graphs locally and the dependency relationships between the body sub - graphs globally through graph attention to obtain representative features of the image frames; Use a classification head to establish a mapping relationship between the representative features and action categories to obtain the scores of the action categories of each image frame in the video data; Use an action trigger to count repeated actions.

[0045] It should be understood that although the steps in the flowcharts of the embodiments of the present invention are shown in sequence according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0046] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0047] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0048] The above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for counting repeated motion postures, characterized in that, The repeated motion posture counting method includes: Obtain a video sequence, extract a series of image frames from the video sequence, convert the extraction of the image frames into joint coordinates, and divide the human joints of the image frames into several main sub - graphs based on the natural topological structure of the human body; Establish the dependency relationships within the main sub - graphs locally and between the main sub - graphs globally through graph attention to obtain representative features of the image frames; Use a classification head to establish a mapping relationship between the representative features and action categories, and obtain the scores of the action categories of each image frame in the video data; Use an action trigger to count repeated actions.

2. The repetitive motion posture counting method according to claim 1, characterized in that Convert each frame in the video sequence into joint coordinates through the Blazepose algorithm. The process is as follows: ; Among them, V represents joint coordinates, represents each of the image frames in the video sequence, C represents the number of channels, H represents the height, W represents the width, and T represents the number of frames. represents each significant part feature point in each of the image frames, D×K×T represents the feature points of the significant parts in the image frame, D represents the dimension of each pose key point, and K represents the number of significant part feature points.

3. The repeated motion posture counting method according to claim 2, wherein Divide the human joints of the image frames into several main sub - graphs based on the natural topological structure of the human body, and perform the same feature extraction operation on each main sub - graph. The operation is as follows: ; Among them, represents that the human body topology is different from the subject sub-graphs, represents the number of human body pose key points in the j-th subject sub-graph. The CBR layer consists of Convolution, BatchNorm, and ReLU. D' represents the D-dimension of which has undergone a CBR transformation, represents the new representation symbol after the CBR transformation.

4. The repetitive motion gesture counting method according to claim 1, characterized in that The steps of establishing the dependency relationships within the main sub - graphs locally through graph attention specifically include the following steps: Use the graph attention mechanism to expand the feature representation and extract features: ; Among them, represents the linear transformation matrix for obtaining Query in the attention mechanism, respectively represent the linear transformation matrices for Key and Value in the attention mechanism, represent the Query, Key, and Value of the j-th sub-graph; Perform graph attention calculation, perform max pooling operation on the result of graph attention, repeat the feature mapping U times along the point dimension, concatenate the mapped features along the channel dimension, and perform average pooling along the channel dimension. Repeat the feature mapping M times on the channel dimension to obtain the output feature : ; Among them, represents the graph attention operation, Maxpooling(·) represents the max pooling operation, RP represents the feature mapping, U represents the number of feature mapping times, represents the feature of the j-th said subject subgraph, represents the feature obtained after all subject subgraphs are mapped and concatenated. M is the number of feature mapping times, Avgpooling represents the average pooling operation, and Concat(·) represents concatenation in the same dimension.

5. The repetitive motion gesture counting method according to claim 4, characterized in that, The steps of establishing the dependency relationships between the main sub - graphs globally specifically include: Embed the entire node into a new M-dimensional feature space through two feature embedding layers in the channel dimension, and ensure the dimensionality is equal to that of through dimension reshaping; connect J of the said main subgraphs, arrange the dimensions, and then input the embedded features into the self-attention mechanism to obtain the global dependencies between the said main subgraphs; use the splicing method to fuse and to obtain the said representative features, and the specific formula is as follows: Embed into a new M-dimensional feature space through dimension reshaping Ensure The dimensionality is equal to that of Connect J of the said main subgraphs, arrange the dimensions, and then input the embedded features Into the self-attention mechanism to obtain the global dependencies between the said main subgraphs; use the splicing method to And Fuse to obtain the said representative features, and the specific formula is as follows: ; ; Among them, represents a node, the number of nodes is j, CBR(·) represents feature embedding, M represents the dimension, represents the global dependency, Reshape(·) represents reshaping the dimension, Contact(·) represents the operation from where, represents the representative feature.

6. The repeated motion gesture counting method according to claim 5, wherein The method of using a classification head for classification is specifically: ; Among them, the representative feature is input into a feed-forward neural network, which is composed of three cascaded . Three MLPs with different parameter settings are used in the classification head, and the pose mapping is completed through a linear layer to obtain the score of a single frame , where O represents the number of output channels; Flatten(·) means flattening all dimensions except the batch dimension into one dimension, and Linear(·) performs a linear transformation on the Flattened features to map the features to the action categories to obtain the classification score 7. The repetitive motion posture counting method according to claim 1, characterized in that, The repeated motion posture counting method further includes loss calculation. The elements for calculating the loss include triple - margin loss and binary cross - entropy loss. The calculation of the triple - margin loss is as follows: ; Among them, represents the triple boundary loss, where a, p, and n represent the anchor, positive sample, and negative sample respectively, and CS represents the cosine similarity, which is used to measure the similarity between features; The calculation of the binary cross - entropy loss is as follows: ; Among them, represents the binary cross-entropy loss, N represents the batch size, where each frame constitutes a batch, y represents the true label, p is the predicted result, and c represents the number of action categories; The total loss function is as follows: ; Among them, represents the total loss function, and α is the weighting factor that controls the weighting of the two losses.

8. A repetitive motion posture counting device, characterized in that, The repeated motion posture counting device executes the repeated motion posture counting method according to any one of claims 1 - 7. The repeated motion posture counting device includes: An information extraction and segmentation module, configured to obtain a video sequence, extract a series of image frames from the video sequence, convert the extraction of the image frames into joint coordinates, and divide the human joints of the image frames into several main sub - graphs based on the natural topological structure of the human body; A local and global feature learning module, configured to establish the dependency relationships within the main sub - graphs locally and between the main sub - graphs globally through graph attention to obtain representative features of the image frames; A classification head, configured to use the classification head to establish a mapping relationship between the representative features and action categories, and obtain the scores of the action categories of each image frame in the video data; A repeated counting module, configured to use an action trigger to count repeated actions.

9. A computer device, characterized in that, It includes a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the processor executes the steps of the repeated motion posture counting method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer - readable storage medium. When the computer program is executed by the processor, the processor executes the steps of the repeated motion posture counting method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • User identity recognition method and system in combination with user gait information

    CN112101176A

  • Skipping rope counting method and system based on space-time diagram convolutional network

    CN115346149A

  • Repeated action counting method and device based on multi-structure information sensing network

    CN118781663A

  • Methods and apparatus for human pose estimation from images using dynamic multi-headed convolutional attention

    WO2023219901A1