Action recognition method and device, computer equipment, storage medium and computer program product

By extracting visual content features from video frames, calculating inter-frame similarity and periodic motion patterns over time, and combining localization recognition and enhanced feature processing, the problems of inconsistent duration and noise interference in periodic motion recognition are solved, thus improving the accuracy of motion recognition.

CN121921831APending Publication Date: 2026-04-24TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2024-10-22
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies suffer from inconsistent action durations and noise interference when recognizing periodic actions in real-world scenarios, leading to reduced accuracy in action recognition.

Method used

By extracting visual content features from video frames, calculating inter-frame similarity, extracting periodic motion pattern features at different time scales, and fusing them, combined with localization recognition and enhanced feature processing, the number of periodic actions can be identified.

Benefits of technology

It effectively reduces the inconsistency in the duration of cyclic movements and the noise interference introduced by rest, thereby improving the accuracy of cyclic movement quantity recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121921831A_ABST
    Figure CN121921831A_ABST
Patent Text Reader

Abstract

The invention relates to an action recognition method and device, computer equipment, a storage medium and a computer program product. The method comprises the following steps: acquiring a to-be-identified video, and extracting visual content of each video frame in the to-be-identified video to obtain content features of each video frame; calculating the similarity between the content features of the video frames to obtain similar features, extracting periodic action mode features corresponding to the time scales based on the similar features, and fusing the periodic action mode features to obtain target periodic action mode features; based on the content feature of each video frame, positioning and identifying a periodic action to obtain a periodic action positioning feature; fusing the target periodic action mode feature and the periodic action positioning feature to obtain a periodic action enhancement feature; and identifying the periodic action number of the target object in the to-be-identified video based on the periodic action enhancement feature to obtain the periodic action number of the target object in the to-be-identified video. By adopting the method, the accuracy of action recognition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video recognition technology, and in particular to an action recognition method, apparatus, computer equipment, storage medium, and computer program product. Background Technology

[0002] With the development of video recognition technology, it is possible to identify the number of periodic actions performed by objects in a video. For example, it is possible to identify the number of periodic actions of robots in a video, the number of periodic actions of humans in a video, and so on.

[0003] Currently, in real-world scenarios involving the recognition of periodic actions, the number of periodic actions is typically counted by calculating the similarity of video frames. However, due to the inconsistent duration of periodic actions in real-world videos and the presence of noise interference such as breaks between them, counting periodic actions by calculating the similarity of video frames can easily lead to a decrease in the accuracy of action recognition. Summary of the Invention

[0004] Therefore, it is necessary to provide a method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve the accuracy of action recognition in response to the above-mentioned technical problems.

[0005] Firstly, this application provides an action recognition method. The method includes:

[0006] The video to be identified is acquired, and the visual content of each video frame in the video to be identified is extracted to obtain the content features of each video frame.

[0007] The similarity between the content features of each video frame is obtained, and each similar feature is obtained. Based on each similar feature, periodic motion patterns at different time scales are extracted to obtain the periodic motion pattern features corresponding to each time scale.

[0008] The periodic motion pattern features corresponding to each time scale are fused to obtain the target periodic motion pattern features.

[0009] Based on the content features of each video frame, the location and recognition of periodic actions are performed to obtain periodic action location features;

[0010] The target periodic action pattern features and periodic action localization features are fused to obtain the periodic action enhancement features corresponding to the video to be identified.

[0011] Based on the enhanced features of periodic actions, the number of periodic actions of the target object in the video to be identified is determined, thus obtaining the number of periodic actions of the target object in the video to be identified.

[0012] Secondly, this application also provides an action recognition device. The device includes:

[0013] The feature extraction module is used to acquire the video to be identified and extract the visual content of each video frame in the video to be identified, so as to obtain the content features of each video frame.

[0014] The pattern feature extraction module is used to obtain the similarity between the content features of each video frame, obtain each similar feature, and extract periodic motion patterns at different time scales based on each similar feature, so as to obtain the periodic motion pattern features corresponding to each time scale.

[0015] The pattern feature fusion module is used to fuse the periodic action pattern features corresponding to each time scale to obtain the target periodic action pattern features.

[0016] The positioning and recognition module is used to locate and recognize periodic actions based on the content features of each video frame, and to obtain periodic action positioning features.

[0017] The enhanced feature acquisition module is used to fuse the target periodic action pattern features and periodic action localization features to obtain the periodic action enhanced features corresponding to the video to be identified.

[0018] The quantity recognition module is used to identify the number of periodic actions of the target object in the video to be recognized based on the enhanced periodic action features, and to obtain the number of periodic actions of the target object in the video to be recognized.

[0019] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:

[0020] The video to be identified is acquired, and the visual content of each video frame in the video to be identified is extracted to obtain the content features of each video frame.

[0021] The similarity between the content features of each video frame is obtained, and each similar feature is obtained. Based on each similar feature, periodic motion patterns at different time scales are extracted to obtain the periodic motion pattern features corresponding to each time scale.

[0022] The periodic motion pattern features corresponding to each time scale are fused to obtain the target periodic motion pattern features.

[0023] Based on the content features of each video frame, the location and recognition of periodic actions are performed to obtain periodic action location features;

[0024] The target periodic action pattern features and periodic action localization features are fused to obtain the periodic action enhancement features corresponding to the video to be identified.

[0025] Based on the enhanced features of periodic actions, the number of periodic actions of the target object in the video to be identified is determined, thus obtaining the number of periodic actions of the target object in the video to be identified.

[0026] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:

[0027] The video to be identified is obtained, and the embedded representations of the video frames in the video to be identified are extracted to obtain the content features of each video frame;

[0028] The similarity between the content features of each video frame is obtained, and each similar feature is obtained. Based on each similar feature, periodic motion patterns at different time scales are extracted to obtain the periodic motion pattern features corresponding to each time scale.

[0029] The periodic motion pattern features corresponding to each time scale are fused to obtain the target periodic motion pattern features.

[0030] Based on the content features of each video frame, the location and recognition of periodic actions are performed to obtain periodic action location features;

[0031] The target periodic action pattern features and periodic action localization features are fused to obtain the periodic action enhancement features corresponding to the video to be identified.

[0032] Based on the enhanced features of periodic actions, the number of periodic actions of the target object in the video to be identified is determined, thus obtaining the number of periodic actions of the target object in the video to be identified.

[0033] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:

[0034] The video to be identified is acquired, and the visual content of each video frame in the video to be identified is extracted to obtain the content features of each video frame.

[0035] The similarity between the content features of each video frame is obtained, and each similar feature is obtained. Based on each similar feature, periodic motion patterns at different time scales are extracted to obtain the periodic motion pattern features corresponding to each time scale.

[0036] The periodic motion pattern features corresponding to each time scale are fused to obtain the target periodic motion pattern features.

[0037] Based on the content features of each video frame, the location and recognition of periodic actions are performed to obtain periodic action location features;

[0038] The target periodic action pattern features and periodic action localization features are fused to obtain the periodic action enhancement features corresponding to the video to be identified.

[0039] The number of periodic actions of the target object in the video to be identified is determined by using periodic action enhancement features.

[0040] The aforementioned action recognition method, apparatus, computer equipment, storage medium, and computer program product acquire a video to be recognized and extract the visual content of each video frame to obtain the content features of each video frame. The similarity between the content features of each video frame is obtained to obtain similarity features, and periodic action patterns at different time scales are extracted based on these similarity features to obtain periodic action pattern features corresponding to each time scale. The periodic action pattern features corresponding to each time scale are fused to obtain target periodic action pattern features. Periodic action localization and recognition are performed based on the content features of each video frame to obtain periodic action localization features. The target periodic action pattern features and periodic action localization features are fused to obtain enhanced periodic action features corresponding to the video to be recognized. The number of periodic actions of the target object in the video to be recognized is identified based on the enhanced periodic action features to obtain the number of periodic actions of the target object in the video to be recognized. In essence, by extracting the periodic action pattern features corresponding to each time scale, the target periodic action pattern features are obtained, enabling the identification of periodic actions at different time scales. This reduces noise interference introduced by inconsistent periodic action durations. Then, the periodic action localization features are obtained through localization recognition, enabling the localization of periodic actions and reducing noise interference introduced by breaks in the periodic actions. Finally, the periodic action enhancement features are fused together, and the number of periodic actions is identified using the periodic action enhancement features. This effectively overcomes the noise interference introduced by inconsistent periodic action durations and breaks in the periodic actions, thereby improving the accuracy of the obtained number of periodic actions, and thus improving the accuracy of action recognition. Attached Figure Description

[0041] Figure 1 This is a diagram illustrating the application environment of the action recognition method in one embodiment;

[0042] Figure 2 This is a flowchart illustrating an action recognition method in one embodiment;

[0043] Figure 3 This is a schematic diagram of the architecture for obtaining the target periodic action pattern characteristics in a specific embodiment;

[0044] Figure 4 This is a schematic diagram of the architecture for obtaining the periodic action localization features in a specific embodiment;

[0045] Figure 5This is a schematic diagram of the architecture for obtaining the periodic motion density map in a specific embodiment;

[0046] Figure 6 This is a schematic diagram of the backbone of an action recognition model in a specific embodiment;

[0047] Figure 7 This is a schematic diagram illustrating the training process of an action recognition model in one embodiment.

[0048] Figure 8 This is a schematic diagram illustrating the principle of action recognition model training in a specific embodiment;

[0049] Figure 9 This is a schematic diagram of the overall architecture of an action recognition model in a specific embodiment;

[0050] Figure 10 This is a flowchart illustrating an action recognition method in a specific embodiment;

[0051] Figure 11 This is a schematic diagram of a squatting exercise video in a specific embodiment;

[0052] Figure 12 This is a schematic diagram of a leg-raising exercise video in a specific embodiment;

[0053] Figure 13 This is a schematic diagram of the action recognition results in a specific embodiment;

[0054] Figure 14 This is a structural block diagram of an action recognition device in one embodiment;

[0055] Figure 15 This is an internal structural diagram of a computer device in one embodiment;

[0056] Figure 16 This is a diagram of the internal structure of a computer device in another embodiment. Detailed Implementation

[0057] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0058] The action recognition method provided in this application embodiment can be applied to, for example, Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed on the cloud or other servers. Server 104 can acquire the video to be recognized from terminal 102 and extract the visual content of each video frame to obtain the content features of each video frame; server 104 obtains the similarity between the content features of each video frame to obtain similar features, and extracts periodic action patterns at different time scales based on these similar features to obtain periodic action pattern features corresponding to each time scale; server 104 fuses the periodic action pattern features corresponding to each time scale to obtain the target periodic action pattern features; based on the content features of each video frame, it performs periodic action localization and recognition to obtain periodic action localization features; server 104 fuses the target periodic action pattern features and the periodic action localization features to obtain the periodic action enhancement features corresponding to the video to be recognized; server 104 identifies the number of periodic actions of the target object in the video to be recognized based on the periodic action enhancement features to obtain the number of periodic actions of the target object in the video to be recognized. Server 104 can return the number of periodic actions of the target object in the video to be identified to terminal 102 for display. Terminal 102 can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein.

[0059] In one embodiment, such as Figure 2 As shown, an action recognition method is provided, which can be applied to... Figure 1 Taking a server as an example, it can be understood that this method can also be applied to a terminal, and also to a system that includes both a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0060] S202, acquire the video to be identified, and extract the visual content of each video frame in the video to be identified to obtain the content features of each video frame.

[0061] The video to be identified refers to a video in which the number of periodic actions of a target object needs to be identified. This target object can be a real object or a virtual object. A real object can be an object that exists in the real world, such as a person, animal, robot, or robotic animal. A virtual object can be an object that does not exist in the real world, such as a virtual character or animal. Periodic actions refer to actions that can be repeated periodically. For example, periodic actions can be human fitness movements, such as squats, leg raises, push-ups, and rope skipping. The time required to complete one periodic action can be the same or different; that is, the action cycles corresponding to the same periodic action can be the same or different. For example, the time required to perform a squat at the beginning is usually less than the time required to perform a squat after a period of time. Content features are feature vectors used to represent the visual content in a video frame, and can be extracted through embedding representation. Embedding representation is the process of mapping a video frame to a low-dimensional space. Through embedding representation, a video frame can be converted into a vector representation, capturing the feature information in the video frame. Video frame features refer to the embedding representation vector of the video frame, used to represent the feature information of the video frame.

[0062] Specifically, the server can obtain the video to be recognized uploaded by the terminal, retrieve the video from a database, obtain the video from the Internet, or acquire the video through a video capture device. It can also obtain the video to be recognized sent by a service provider offering action recognition services. The server then extracts the visual content of each video frame from the video to be recognized. This extraction can be done using a neural network, such as a convolutional neural network or a recurrent neural network, employing deep learning techniques. The server can extract the embedding representations of each video frame in parallel or sequentially to obtain the content features of each video frame.

[0063] In some embodiments, the server may sample the video to be recognized to obtain video frames, and then encode each video frame using an encoder to obtain the content features of each video frame. The encoder may be a pre-trained neural network for feature encoding, which can convert the input sequence into a fixed-length context vector. The encoder may be built using a recurrent neural network.

[0064] S204, obtain the similarity between the content features of each video frame, obtain each similar feature, and extract the periodic motion pattern at different time scales based on each similar feature, and obtain the periodic motion pattern feature corresponding to each time scale.

[0065] Among them, similarity features are feature vectors used to characterize the similarity between features of video frames. This feature vector can be a one-dimensional vector or a matrix vector. Periodic motion patterns are used to characterize the feature information of periodic motions in the video to be identified. Periodic motion patterns at different time scales refer to the feature information of periodic motions at different time scales in the video to be identified. Periodic motion pattern features are feature vectors used to characterize the periodic motion information at the corresponding time scale.

[0066] In this embodiment, the server can obtain the similarity between features of two video frames in the content features of each video frame. This similarity can be calculated using similarity algorithms, such as distance similarity algorithms or cosine similarity algorithms, or using pre-trained similarity weights. The server obtains various similar features by calculating the similarity between each video frame and other video frame features. Then, the server uses these similar features to extract periodic motion patterns at different time scales. This can be done by using a pre-set neural network to extract periodic motion patterns at different time scales, resulting in output periodic motion pattern features for each time scale. These time scales are pre-set and can be obtained through testing. Each similar feature is used to extract the periodic motion pattern at its corresponding time scale, resulting in periodic motion pattern features corresponding to each similar feature.

[0067] S206, fuse the periodic motion pattern features corresponding to each time scale to obtain the target periodic motion pattern features.

[0068] Among them, the target periodic motion pattern feature is a feature vector used to characterize the periodic motion information in the video to be identified, which is obtained by fusing periodic motion pattern features at different time scales.

[0069] Specifically, the server sequentially fuses the periodic action pattern features corresponding to each time scale. This can be achieved by concatenating the periodic action pattern features at each time scale sequentially in chronological order to obtain the target periodic action pattern feature. Alternatively, feature vector operations can be performed on the periodic action pattern features across all time scales to obtain the target periodic action pattern feature. For example, this could be done by summing the periodic action pattern features across all time scales, or by multiplying the periodic action pattern features across all time scales.

[0070] S208, Based on the content features of each video frame, perform periodic action localization and recognition to obtain periodic action localization features.

[0071] Localization recognition refers to identifying the positional information of periodic actions within a video to be recognized. This positional information can be any video frame containing the periodic action. In other words, it can be used to classify and identify the periodic action as the foreground and non-periodic actions as the background. Periodic action localization features are feature vectors used to characterize the positional information of periodic actions within the video to be recognized.

[0072] Specifically, the server can identify periodic actions in the video at the instance level. This means the server can locate and identify periodic actions using the content features of each video frame. Specifically, it can use the content features of each video frame to locate the position of the periodic action in the video to be identified, thus obtaining periodic action location features. The server can use a trained neural network to locate and identify periodic actions using the content features of each video frame, and this trained neural network can be a convolutional neural network, a recurrent neural network, a feedforward neural network, etc.

[0073] S210, the target periodic action pattern features and periodic action localization features are fused to obtain the periodic action enhancement features corresponding to the video to be identified.

[0074] Among them, periodic action enhancement features refer to enhanced periodic action features, which are feature vectors used to characterize periodic actions in the video to be identified. Periodic action enhancement features are a more fine-grained representation of periodic actions corresponding to the video to be identified.

[0075] Specifically, the server can fuse the target periodic action pattern features with the periodic action localization features. This can be achieved by concatenating the target periodic action pattern features and the periodic action localization features end-to-end; for example, the target periodic action pattern features can be used as the beginning and the periodic action localization features as the end, or vice versa. Alternatively, feature vector operations can be performed on the target periodic action pattern features and the periodic action localization features to obtain enhanced periodic action features. For example, this could involve calculating the sum of the feature vectors of the target periodic action pattern features and the periodic action localization features, or calculating the product of the feature vectors of the target periodic action pattern features and the periodic action localization features, etc.

[0076] In some embodiments, the server may also fuse the target periodic action pattern features with periodic action positioning features through a neural network, which is a pre-trained neural network for fusing features.

[0077] In some embodiments, the server may also use an aggregation function to fuse the target periodic action pattern features with the periodic action location features. For example, the average value of the target periodic action pattern features and the periodic action location features may be calculated to obtain the periodic action enhancement features.

[0078] In some embodiments, the server may also fuse the target periodic action pattern features with periodic action localization features through convolution operations to obtain periodic action enhancement features.

[0079] In some embodiments, the server can also obtain pre-set weights, weight the target periodic action pattern features and periodic action positioning features respectively, and then fuse them to obtain periodic action enhancement features.

[0080] S212, Based on the periodic action enhancement features, identify the number of periodic actions of the target object in the video to be identified, and obtain the number of periodic actions of the target object in the video to be identified.

[0081] The number of periodic actions of the target object refers to the statistical count of repeated periodic actions performed by the target object in the video to be identified. In this embodiment, the number of periodic actions is identified in a video to be identified that includes one target object.

[0082] Specifically, the server can use periodic motion enhancement features to identify the number of periodic motions of the target object in the video to be identified. This can be achieved by using periodic motion enhancement features to perform mapping calculations of the periodic motion density map to obtain the density map of the video to be identified, and then using the density map to statistically calculate the number of periodic motions. Alternatively, the number of periodic motions can be identified by using a neural network for periodic motion counting to obtain the number of periodic motions of the target object in the video to be identified.

[0083] In some embodiments, when the target object in the video to be identified includes at least two objects, the server can segment the video to be identified according to the target objects to obtain the video corresponding to each target object, and then use the video corresponding to each target object as the video to be identified for action recognition.

[0084] The aforementioned action recognition method acquires the video to be recognized and extracts the visual content of each video frame to obtain the content features of each video frame. It then obtains the similarity between the content features of each video frame to obtain similarity features, and extracts periodic action patterns at different time scales based on these similarity features to obtain periodic action pattern features corresponding to each time scale. The periodic action pattern features corresponding to each time scale are fused to obtain the target periodic action pattern features. Based on the content features of each video frame, periodic actions are located and recognized to obtain periodic action location features. The target periodic action pattern features and the periodic action location features are fused to obtain the enhanced periodic action features corresponding to the video to be recognized. Based on the enhanced periodic action features, the number of periodic actions of the target object in the video to be recognized is identified to obtain the number of periodic actions of the target object in the video to be recognized. In essence, by extracting the periodic action pattern features corresponding to each time scale, the target periodic action pattern features are obtained, enabling the identification of periodic actions at different time scales. This reduces noise interference introduced by inconsistent periodic action durations. Then, the periodic action localization features are obtained through localization recognition, enabling the localization of periodic actions and reducing noise interference introduced by breaks in the periodic actions. Finally, the periodic action enhancement features are fused together, and the number of periodic actions is identified using the periodic action enhancement features. This effectively overcomes the noise interference introduced by inconsistent periodic action durations and breaks in the periodic actions, thereby improving the accuracy of the obtained number of periodic actions, and thus improving the accuracy of action recognition.

[0085] In some embodiments, S204, namely, extracting periodic action patterns at different time scales based on each similar feature to obtain periodic action pattern features corresponding to each time scale, includes the following steps:

[0086] Each similar feature is pooled at its corresponding time scale to extract perceptual domain information, resulting in a pooled feature. Each pooled feature is then interacted with for temporal information through self-attention, resulting in its own attention feature. Finally, each attention feature is subjected to a fully connected operation to obtain the periodic action pattern feature corresponding to each time scale.

[0087] Pooling features refer to the features obtained after performing pooling operations. Pooling operations allow the pre-obtained pooling features to contain information from a specific receptive domain. Self-attention features refer to the features extracted through a self-attention mechanism. The self-attention mechanism can extract the temporal correlation information between different elements in the pooling features. The self-attention mechanism calculates the representation of the input by comparing each element of the input with other elements in the input.

[0088] Specifically, the server extracts receptive domain information for each similar feature through pooling operations at the corresponding time scale, obtaining pooled features for each similar feature. The time scale pooling operations for each similar feature are pre-set. Then, the server interacts with each pooled feature using a self-attention mechanism to obtain self-attention features. For example, the server can use self-attention parameters for self-attention feature extraction; these parameters can be pre-set or obtained from the service provider offering the competition. Alternatively, the server can directly input each pooled feature into the corresponding self-attention neural network for self-attention feature extraction; this self-attention neural network can be pre-trained. To ensure consistency in feature dimensions, the server can perform fully connected operations on each self-attention feature to obtain periodic action pattern features. This can be done using a fully connected neural network, which can be pre-trained. Alternatively, the server can obtain fully connected operation parameters and use these parameters to perform fully connected operations on the self-attention features; these parameters can be pre-set or obtained from the parameter provider. Finally, the server can fuse the periodic action pattern features corresponding to each time scale to obtain the target periodic action pattern features.

[0089] In the above embodiments, perceptual domain information is extracted by pooling each similar feature at its corresponding time scale to obtain pooled features. Then, each pooled feature is interacted with for temporal information through self-attention to obtain its own attention features. Finally, each attention feature is subjected to a fully connected operation to obtain the periodic action pattern features corresponding to each time scale. That is, by using pooling and self-attention at different temporal levels for feature extraction, the obtained periodic action pattern features can contain periodic action information at the corresponding time scale, thereby improving the accuracy of the obtained periodic action pattern features.

[0090] In some embodiments, each similar feature is subjected to pooling at its corresponding time scale to extract perceptual domain information, resulting in each pooled feature, including the following steps:

[0091] The target similarity features in each similarity feature are pooled according to the corresponding pooling window to obtain the pooled features corresponding to the target similarity features.

[0092] In this embodiment, the pooling window refers to the kernel size used for pooling operations. This pooling window is pre-set and represents the corresponding time scale. There are multiple pooling windows, and different pooling windows are used for different similar features during pooling operations. Different pooling windows represent pooling at different time scales. The target similar feature refers to the similar feature to be pooled. The server can perform max pooling on the target similar feature according to the corresponding pooling window to obtain the pooled feature. This max pooling operation can select the maximum value among the feature elements corresponding to the pooling window. The server can also perform average pooling on the target similar feature according to the corresponding pooling window to obtain the pooled feature. This max pooling operation can calculate the average value of all feature elements corresponding to the pooling window. The server can simultaneously perform pooling operations on each similar feature according to its corresponding pooling window to obtain the corresponding pooled feature. The server can also sequentially perform pooling operations on each similar feature according to its corresponding pooling window in chronological order to obtain the corresponding pooled feature.

[0093] In the above embodiments, by performing pooling operations on the target similar features in each similar feature according to the corresponding pooling window, the pooled features corresponding to the target similar features are obtained. Through the pooling operation of the corresponding pooling window, the size of the feature map can be reduced, while retaining the feature information of the periodic actions of the corresponding time scale, thereby improving the accuracy of the obtained pooled features.

[0094] In some embodiments, each pooling feature is subjected to temporal information interaction through self-attention to obtain its own attention feature, including the following steps:

[0095] Extract the query feature and key feature of the target pooling feature from each pooling feature; calculate the self-attention weight based on the query feature and key feature to obtain the self-attention weight; weight the target pooling feature according to the self-attention weight to obtain the self-attention feature corresponding to the target pooling feature.

[0096] In this embodiment, the target pooling feature is the pooling feature determined from all pooling features that requires self-attention feature extraction. The server can simultaneously use each pooling feature as a target pooling feature for parallel self-attention feature extraction, or it can sequentially use each pooling feature as a target pooling feature for serial self-attention feature extraction. During self-attention feature extraction, the server obtains pre-set self-attention parameters, which may include linear transformation parameters. The pooling features are then linearly transformed using these parameters to obtain query features and key features, where the linear transformation parameters for the query and key features are different. Self-attention weights are then calculated using the query and key features, i.e., the similarity between the query and key features is calculated. This similarity can be obtained using dot product or scaled dot product operations. Normalization is then performed based on the similarity calculation result to obtain the self-attention weights. Finally, the server weights the target pooling feature according to the self-attention weights, i.e., it calculates the product of the self-attention weights and the target pooling feature to obtain the self-attention feature corresponding to the target pooling feature.

[0097] In some embodiments, the server can also obtain the parameters for linear transformation to perform a linear transformation on the pooling features, resulting in value features. Then, the value features are weighted using self-attention weights, i.e., the product of the self-attention weights and the value features is calculated to obtain the self-attention features.

[0098] In some embodiments, the server may also use convolutional layers and dilated convolutions with ReLU activation to interact with temporal information to obtain self-attention features.

[0099] In the above embodiments, by extracting self-attention weights and weighting the target pooling features according to the self-attention weights, the self-attention features corresponding to the target pooling features are obtained. In this way, the time information corresponding to the periodic actions can be extracted through the self-attention mechanism, which improves the accuracy of the obtained self-attention features and thus improves the accuracy of action recognition.

[0100] In a specific embodiment, such as Figure 3The diagram illustrates an architecture for obtaining target periodic action pattern features. Each similarity matrix undergoes max pooling with pooling layers of different kernel sizes to extract information with a specific receptive domain. Each kernel-sized pooling layer operates at different time scales, from time scale 1 to time scale 10. One-dimensional max pooling is performed at each time scale to obtain the output pooled features. Each pooled feature is then passed through a corresponding self-attention network (SA) for self-attention feature extraction, yielding its own attention features. Finally, each self-attention feature is processed through a corresponding fully connected layer to obtain the periodic action pattern features corresponding to each time scale. These periodic action pattern features are then merged to obtain the target periodic action pattern features. In other words, by applying a scale-specific attention mechanism to the temporal self-similarity representation, various periodic action distributions can be matched, thereby extracting finer-grained periodic action representation features.

[0101] In some embodiments, S208, namely, performing periodic action localization and recognition based on the content features of each video frame to obtain periodic action localization features, includes the following steps:

[0102] Based on the content features of each video frame, contextual semantic information is extracted through temporal convolution to obtain initial periodic action localization features; fully connected operations are then performed on the initial periodic action localization features to obtain periodic action localization features.

[0103] Among them, the initial periodic action localization features refer to the initial features used to characterize the positional information of periodic actions in the video to be identified. Temporal convolution is a pre-set convolution used to extract semantic information from temporal data. Temporal convolution operations can be performed to extract contextual semantic information.

[0104] In this embodiment, the server can use the content features of each video frame to extract contextual semantic information through temporal convolution to obtain initial periodic action localization features. The server can obtain temporal convolution parameters and use these parameters to perform temporal convolution operations on the content features of each video frame to extract the contextual semantic information of the video to be identified. These temporal convolution parameters can be pre-set, obtained from a database, or obtained from a service provider offering parameter services. Alternatively, the server can input the content features of each video frame into a temporal convolutional neural network for contextual semantic information extraction to obtain the output initial periodic action localization features. The temporal convolutional neural network can be pre-trained. To ensure consistency in feature dimensions, the server can further perform fully connected operations on the initial periodic action localization features to obtain periodic action localization features. This fully connected operation unifies the feature dimensions, ensuring feature accuracy and facilitating subsequent calculations.

[0105] In some embodiments, the server can also use initial periodic motion localization features to perform classification and recognition through a classification network, which may be built using a fully connected network. This involves using the initial periodic motion localization features to identify the foreground category of periodic motions through the classification network, where the foreground is the periodic motion portion of the video to be identified, and the background is the portion other than the periodic motion. Inputting the initial periodic motion localization features into the classification network yields the probability of the periodic motion category for each video frame in the video to be identified. Based on the periodic motion category probability, the video frames in the video to be identified that belong to periodic motions are determined; for example, video frames exceeding a threshold can be considered as video frames of periodic motions.

[0106] In a specific embodiment, such as Figure 4 The diagram illustrates the architecture for obtaining periodic motion localization features. Initial periodic motion localization features can be extracted using a temporal convolutional network (TCN), such as a TCN. This TCN combines the parallel processing capabilities of convolutional neural networks (CNNs) with the long-term dependency modeling capabilities of recurrent neural networks (RNNs), enabling the modeling of time-series data. The TCN consists of dilated convolutions and causal convolutions. The server inputs the content features of each video frame into this TCN for computation, extracting contextual information from the input feature sequence and locating motion cycles from noisy information, thus obtaining the initial periodic motion localization features. Finally, a fully connected layer is used to perform fully connected operations on the initial periodic motion localization features to obtain the final periodic motion localization features.

[0107] In the above embodiments, the initial periodic action localization features are obtained by extracting contextual semantic information through temporal convolution based on the content features of each video frame; the initial periodic action localization features are then subjected to a fully connected operation to obtain periodic action localization features. This allows the obtained periodic action localization features to contain the positional information of periodic actions, thereby improving the accuracy of the obtained periodic action localization features and thus improving the accuracy of action recognition.

[0108] In some embodiments, S204, namely, obtaining the similarity between the content features of each video frame to obtain each similar feature, includes the following steps:

[0109] The first video frame content features and the second video frame features are obtained from the content features of each video frame, and the similarity weights corresponding to the first video frame content features and the second video frame content features are obtained. The transpose of the first video frame content features is calculated to obtain the first transpose feature. Based on the first transpose feature, the similarity weights and the second video frame content features, the similarity features between the first video frame content features and the second video frame content features are calculated.

[0110] In this embodiment, the first video frame content feature and the second video frame content feature refer to the two video frame features whose similarity features are to be calculated from the content features of each video frame. The first transposed feature is obtained by transposing the first video frame content feature. The similarity weight refers to the weight used when calculating the similarity feature, which is pre-set, for example, it can be obtained and set through training. That is, the server can sequentially select the first video frame content feature and the second video frame content feature from the content features of each video frame to perform pairwise similarity calculation, that is, calculate the transpose of the first video frame content feature to obtain the first transposed feature. Then, the first transposed feature is weighted using the similarity weight to obtain the weighted feature. Then, the product of the weighted feature and the second video frame content feature is calculated to obtain the similarity feature between the first video frame content feature and the second video frame content feature. Then, the server traverses the content features of each video frame and calculates the similarity feature between each video frame content feature and other video frame content features to obtain each similar feature. This similarity feature is a feature vector used to characterize the degree of similarity between two video frame features, or it can be a feature matrix.

[0111] In a specific embodiment, similar features can be calculated using the formula (1) shown below.

[0112] Formula (1)

[0113] in, It refers to the feature of the i-th video frame. This refers to the feature of the j-th video frame. W represents the similarity weight. T represents the transpose. This represents the similarity features between the features of the i-th video frame and the features of the j-th video frame. This similarity feature can be a similarity matrix.

[0114] In the above embodiments, similarity calculation is performed by using the first transpose feature, similarity weight, and second video frame content feature to obtain similar features between the first video frame content feature and the second video frame content feature. That is, similarity calculation is performed by using pre-set similarity weight, which improves the accuracy of the obtained similar features.

[0115] In some embodiments, S212, namely, identifying the number of periodic actions of the target object in the video to be identified based on the periodic action enhancement features, to obtain the number of periodic actions of the target object in the video to be identified, includes the following steps:

[0116] The video enhancement features are subjected to attention transformation to obtain target transformation features. Fully connected operations are performed on the target transformation features to obtain a periodic action density map. The number of periodic actions of the target object in the video to be identified is calculated based on the periodic action density map to obtain the number of periodic actions of the target object in the video to be identified.

[0117] The target transformation feature refers to the feature obtained through transformation by a transformer. This transformer can be a transformer built using an attention mechanism, such as a Transformer (a sequence network based on an attention mechanism). The Transformer has a global perspective and can extract richer semantic information features. This Transformer adopts an encoder-decoder architecture, with both the encoder and decoder consisting of multi-head attention layers and fully connected feedforward networks. The periodic action density map is used to represent the distribution information of periodic actions in video frames and is the predicted distribution information. During training, the sum of the distributions of periodic actions in the video can be preset to 1 to generate the labels of the periodic action density map. This distribution information can be normal distribution information, Gaussian distribution information, etc. For example, if a periodic action corresponds to 5 video frames, then the normal distribution values ​​of the corresponding 5 video frames can be set, and the sum of the normal distribution values ​​of the 5 video frames is 1.

[0118] Specifically, the server can obtain a pre-set transformer and then input the video enhancement features into the transformer for attention transformation to obtain the target transformation features. The server can also obtain attention transformation parameters and apply these parameters to the video enhancement features to obtain the target transformation features. These attention transformation parameters can be pre-trained or obtained from a parameter provider. Then, a fully connected operation is performed on the target transformation features to obtain a periodic action density map. This can be done using pre-set fully connected parameters or a pre-trained fully connected network that maps the density map. The result is the periodic action density map. Finally, the server calculates the sum of all density distribution values ​​in the periodic action density map to obtain the number of periodic actions of the target object in the video to be identified. That is, the values ​​in the periodic action density map are between 0 and 1. The sum of all elements in the periodic action density map is the number of periodic actions in the video to be identified.

[0119] In a specific embodiment, such as Figure 5The diagram illustrates the architecture for obtaining the periodic motion density map. Video enhancement features are input into a counting network for density map mapping. This counter is built using a Transformer and two fully connected layers. The output periodic motion density map is obtained by transforming the video enhancement features and performing fully connected operations. This periodic motion density map can be a one-dimensional feature vector. The sum of all feature elements in this one-dimensional feature vector is then calculated to obtain the number of periodic actions.

[0120] In the above embodiments, by performing attention transformation on the video enhancement features, target transformation features are obtained. A fully connected operation is then performed on these target transformation features to obtain a periodic action density map. The number of periodic actions of the target object in the video to be identified is calculated based on the periodic action density map. In other words, by extracting the periodic action density map and using it to calculate the number of periodic actions, the accuracy of the obtained number of periodic actions can be guaranteed.

[0121] In some embodiments, S202, the video to be identified is acquired, and the visual content of each video frame in the video to be identified is extracted to obtain the content features of each video frame, including the following steps:

[0122] The video to be identified is acquired, and the video to be identified is uniformly sampled according to the number of target frames to obtain each video frame; the visual content of each video frame is extracted to obtain the content features of each video frame.

[0123] In this embodiment, the target frame number refers to the pre-set number of video frames to be uniformly sampled. Uniform sampling means sampling video frames from the video at equal intervals, that is, selecting video frames from the video at equal intervals to ensure that the video frames are evenly distributed on the timeline. This interval can be determined based on the total number of video frames in the video to be identified and the target frame number. For example, the video to be identified can be downsampled to a fixed 64 frames, that is, uniformly sampled from the video to be identified to obtain 64 video frames.

[0124] In some embodiments, the server can uniformly sample the video to be identified. When the number of video frames in the video to be identified is less than the target number of frames, the target number of video frames can be obtained by repeatedly sampling video frames. For example, if the target number of frames is 64, and the length of the video to be identified is less than 64 frames, 64 frames can be obtained by repeatedly sampling video frames, for example, by repeatedly sampling the last frame of the video to be identified. Then, the server extracts the embedding representation from each uniformly sampled video frame to obtain the embedding representation corresponding to each video frame, thus obtaining each video frame.

[0125] In the above embodiments, each video frame is obtained by uniformly sampling the video to be identified according to the number of target frames; the embedding representation of each video frame is extracted to obtain the content features of each video frame. That is, by extracting the visual content of the uniformly sampled video frames, the accuracy of obtaining the content features of each video frame can be improved.

[0126] In some embodiments, the action recognition method further includes the step of:

[0127] The video to be identified is input into the action recognition model, and the visual content of each video frame in the video to be identified is extracted by the action recognition model to obtain the content features of each video frame.

[0128] The similarity between the content features of each video frame is obtained by the pattern feature extraction network in the action recognition model. Each similar feature is obtained, and periodic action patterns at different time scales are extracted based on each similar feature to obtain the periodic action pattern features corresponding to each time scale. The periodic action pattern features corresponding to each time scale are then fused to obtain the target periodic action pattern features.

[0129] The localization feature extraction network in the action recognition model is used to locate and identify periodic actions in the content features of each video frame, and the periodic action localization features are obtained.

[0130] By fusing the target periodic action pattern features and periodic action localization features through an action recognition model, the periodic action enhancement features corresponding to the video to be recognized are obtained.

[0131] The number of periodic actions of the target object in the video to be identified is obtained by using the action recognition model to identify the number of periodic actions of the target object in the video according to the periodic action enhancement features.

[0132] The action recognition model refers to a pre-trained neural network model that identifies the number of periodic actions in a video. For example, it can be trained using a training video and labels representing the number of periodic actions within that video. This action recognition model includes a pattern feature extraction network and a localization feature extraction network, used to extract key information about the repetition of actions and enhance the embedded representation of the video. The pattern feature extraction network is a neural network used to extract pattern features of periodic actions, capable of extracting features from the video at multiple different time scales. The localization feature extraction network is a neural network used to extract localization features of periodic actions, enabling it to locate the action cycle in the video and understand contextual information from the feature sequence at the instance level. The pattern feature extraction network and the localization feature extraction network are two different branches of the action recognition model.

[0133] In this embodiment, the server can train an initial action recognition model using a training video and periodic action quantity labels within that training video. Once training is complete, an action recognition model is obtained and can be deployed and used. When action recognition is needed, the server invokes the deployed action recognition model, inputting the video to be recognized into it. The action recognition model can then execute the action recognition steps described in any of the above embodiments to obtain the output number of periodic actions. For example, when the action recognition model receives the input video to be recognized, it can extract the visual content of each video frame, obtaining the content features of each video frame. These content features are then simultaneously input into a pattern feature extraction network and a localization feature extraction network. The pattern feature extraction network calculates the similarity between the content features of each video frame, obtaining similar features. Based on these similar features, periodic action patterns at different time scales are extracted, yielding periodic action pattern features corresponding to each time scale. These periodic action pattern features are then fused to obtain the output target periodic action pattern features. Finally, the localization feature extraction network performs localization recognition of the periodic actions based on the content features of each video frame, obtaining the output periodic action localization features. Then, the action recognition model fuses the target's periodic action pattern features and periodic action localization features to obtain enhanced periodic action features. Finally, the action recognition model identifies the number of periodic actions of the target object in the video to be recognized based on the enhanced periodic action features, thus obtaining the number of periodic actions of the target object in the video to be recognized.

[0134] In a specific embodiment, such as Figure 6The diagram illustrates the backbone of an action recognition model. The model extracts an embedding representation of the video to be recognized, which can be a 64*768 dimensional feature vector, obtained by concatenating features from each video frame. This embedding representation is then input into a pattern feature extraction network and a localization feature extraction network for feature extraction, yielding target periodic action pattern features and periodic action localization features. These are then concatenated and fed into a fully connected layer for fully connected operations, resulting in enhanced output features—a 64*112 dimensional periodic action enhanced feature. This enhanced feature is then used for density map mapping, achieved through a Transformer and two fully connected layers, to obtain a periodic action density map, which can be a 64*1 dimensional feature vector. Finally, the sum of all feature values ​​in the corresponding 64*1 dimensional feature vector is calculated, and the sum is rounded. Specifically, if the first decimal place in the sum is 4 or less, that first decimal place and all subsequent decimal places are discarded. If the first digit after the decimal point in the sum is 5 or greater, that first digit and all subsequent digits are discarded, and the integer digits are incremented by 1 to obtain the number of periodic actions. The pattern feature extraction network can be as follows: Figure 3 The network architecture shown is established. A localization feature extraction network can be like this... Figure 4 The network architecture shown is established.

[0135] In the above embodiments, by inputting the video to be identified into the action recognition model and performing action recognition through the action recognition model, the number of periodic actions of the target object in the video to be identified can be obtained. That is, by performing action recognition through a pre-trained action recognition model, the accuracy and efficiency of action recognition can be improved.

[0136] In some embodiments, such as Figure 7 As shown, training the action recognition model includes the following steps:

[0137] S702, obtain training video, training quantity label, and training density map label.

[0138] S704: Input the training video into the initial action recognition model to identify the number of periodic actions, and obtain the training periodic action density map and the number of training periodic actions.

[0139] S706: Obtain the loss between the number of actions in the training cycle and the training quantity label to obtain quantity loss information; and obtain the loss between the action density map in the training cycle and the training density map label to obtain density loss information.

[0140] S708 trains the initial action recognition model based on quantity loss information and density loss information. When the training completion condition is met, the trained action recognition model is obtained.

[0141] In this context, "training video" refers to the pre-annotated video used to train the action recognition model. "Training quantity label" refers to the label of the number of periodic actions in the pre-annotated training video. "Training density map label" refers to the label of the periodic action distribution information corresponding to the pre-annotated training video. "Initial action recognition model" refers to the action recognition model with initialized model parameters, which requires model parameter training. "Training periodic action density map" refers to the periodic action density map obtained by using the initial action recognition model to be trained during action recognition. "Training periodic action quantity" refers to the number of periodic actions obtained by using the initial action recognition model to be trained during action recognition; this number is obtained through statistical calculation using the training periodic action density map. Quantity loss information is used to characterize the error between the training periodic action quantity and the training quantity label. Density loss information is used to characterize the error between the training periodic action density map and the training density map label.

[0142] Specifically, the server can retrieve training videos, training quantity labels, and training density map labels from a database. The server can also obtain these from a data service provider. Furthermore, the server can retrieve training videos, training quantity labels, and training density map labels uploaded by the terminal. The server can also retrieve training videos, annotate them to obtain training quantity labels and training density map labels. Specifically, the training quantity labels can be annotated based on the number of periodic actions in the training video, and then the corresponding training density map labels can be generated using a normal or Gaussian distribution.

[0143] At this point, the server trains the initial action recognition model using training videos, training quantity labels, and training density map labels. During training, the server inputs the training videos into the initial action recognition model, which then performs action recognition on the training videos to obtain the output number of periodic actions. Specifically, the initial action recognition model can identify initial periodic action pattern features and initial periodic action localization features. These features are then fused to obtain enhanced periodic action features. Finally, these enhanced features are used to identify the number of later actions of the target object in the training video, yielding the training periodic actions and the training number of periodic actions.

[0144] The server then uses pre-set loss functions to calculate the loss information. Specifically, the mean absolute error loss function can be used to calculate the loss between the number of actions in a training epoch and the training quantity labels, yielding the quantity loss information. Then, the mean squared error loss function is used to calculate the loss between the action density map in a training epoch and the training density map labels, yielding the density loss information. Finally, the server determines whether training completion conditions have been met, such as whether the loss information has reached a pre-set loss threshold, whether the number of training iterations has reached the maximum number of iterations, and whether the model parameters have stopped changing, etc.

[0145] When the training completion condition is not met, the server uses quantity loss information and density loss information to iterate on the initial action recognition model. That is, the gradient descent algorithm is used to update the model parameters in the initial action recognition model in reverse using quantity loss information and density loss information to obtain an updated action recognition model. The updated action recognition model is used as the initial action recognition model, and the steps of obtaining training videos, training quantity labels and training density map labels are executed again until the training completion condition is met, and the trained action recognition model is obtained.

[0146] In a specific embodiment, the sum of the quantity loss information and the density loss information can be calculated using the formula (2) shown below to obtain the overall count loss information.

[0147] Formula (2)

[0148] Where c represents the number of actions in a training cycle. The training quantity label refers to the actual number of cycles. This represents the action density during the training cycle in frame j. This represents the training density label in frame j, which is the true density. B represents the batch size, and N represents the video length. As a hyperparameter, this The value is set to 0.2. This represents the quantity loss information calculated using the mean absolute error loss function. This represents the density loss information calculated using the mean square error loss function. This indicates the overall count loss information.

[0149] In the above embodiments, the initial action recognition model is trained using quantity loss information and density loss information. When the training completion condition is met, a trained action recognition model is obtained, which can supervise the density map and periodic action count, thereby ensuring the accuracy of action recognition by the trained action recognition model.

[0150] In some embodiments, S708, the initial action recognition model is trained based on quantity loss information and density loss information. When the training completion condition is met, a trained action recognition model is obtained, including the following steps:

[0151] The anchor point pattern features, foreground pattern features, and background pattern features of the anchor point video, foreground pattern features, and background pattern features are obtained. The loss between the anchor point pattern features, foreground pattern features, and background pattern features is obtained to obtain triplet loss information. The initial pattern feature extraction network in the initial action recognition model is trained based on the triplet loss information, and the initial action recognition model is trained based on the quantity loss information and density loss information. When the training completion condition is met, the first action recognition model that has been trained is obtained.

[0152] In this context, the anchor video refers to the reference video used for periodic action recognition; this anchor video is the anchor sample. Anchor pattern features are the target periodic action pattern features extracted from the anchor video using the initial pattern feature extraction network. The foreground video is a video belonging to the same category as the anchor video; this foreground video is a positive sample. Foreground pattern features are the target periodic action pattern features extracted from the foreground video using the initial pattern feature extraction network. The background video belongs to a different category than the anchor video; this background video is a negative sample. Background pattern features are the target periodic action pattern features extracted from the background video using the initial pattern feature extraction network. Triplet loss information, also known as triplet loss information, is used to characterize the loss between anchor pattern features, foreground pattern features, and background pattern features.

[0153] In this embodiment, the server can use anchor video, foreground video, and background video as triple samples to train the initial action recognition model. Specifically, the anchor video, foreground video, and background video are input into the initial action recognition model, and features are extracted through the initial pattern feature extraction network to obtain the anchor pattern features corresponding to the anchor video, the foreground pattern features corresponding to the foreground video, and the background pattern features corresponding to the background video. Then, action recognition continues, obtaining the periodic action density map and the number of periodic actions corresponding to the anchor video output by the initial action recognition model. The quantity loss information and density loss information corresponding to the anchor video are then calculated. Similarly, the quantity loss information and density loss information corresponding to the foreground video and the background video are calculated. Simultaneously, the loss between the anchor pattern features, foreground pattern features, and background pattern features is calculated using a triple loss function to obtain the triple loss information. Finally, the parameters of the initial pattern feature extraction network are updated and iterated using the triple loss information. The server can then update and train the model parameters of the initial action recognition model using both quantity loss and density loss information separately. Alternatively, it can calculate the average quantity loss of all quantity loss information and the average density loss of all density loss information, and then use these averages to update and train the model parameters of the initial action recognition model. In other words, the server uses triplet loss to train only the initial pattern feature extraction network, and uses quantity and density loss information to train the initial action recognition model as a whole. If the training completion condition is not met, new triplet samples are acquired for retraining until the training completion condition is met. The initial action recognition model that meets the training completion condition is then used as the first action recognition model that has been successfully trained. Finally, the server can deploy and use this first action recognition model for action recognition.

[0154] In some specific embodiments, the triplet loss information can be calculated using the formula (3) shown below.

[0155] Formula (3)

[0156] in, This refers to the loss information of the triplet. This refers to the anchor point pattern feature. This refers to foreground pattern characteristics. This refers to background pattern features. It means and The L2 pairwise distance between them. It means and The L2 pairwise distance between them. This refers to the loss information of the triplet. It is a hyperparameter, a pre-set threshold used to control the difference between foreground video samples and background video samples. Typically, the distance between the anchor point recognition sample and the background video sample is at least greater than the distance between the anchor point recognition sample and the foreground video sample.

[0157] In the above embodiments, triplet loss information is obtained, and then the initial pattern feature extraction network in the initial action recognition model is trained using the triplet loss information. The initial action recognition model is also trained based on quantity loss information and density loss information. When the training completion condition is met, the first action recognition model that has been trained is obtained. That is, the foreground and background are distinguished by triplet loss; specifically, the feature representation is enhanced by clustering the embeddings of repetitive action cycles and deducing the embeddings of background cycles, thereby improving the accuracy of the obtained periodic action pattern features and thus improving the accuracy of action recognition.

[0158] In some embodiments, the initial action recognition model includes an initial video frame classification network. The initial action recognition model is trained based on quantity loss information and density loss information. When the training completion condition is met, a trained action recognition model is obtained, including:

[0159] The periodic action localization features are input into the initial video frame classification network for periodic action classification and recognition, and the classification and recognition results corresponding to the video frames in the video to be recognized are obtained. The training category labels corresponding to the training videos are obtained, and the loss between the training category labels and the classification and recognition results is obtained to obtain category loss information. Based on the category loss information, the initial localization feature extraction network and the initial video frame classification network in the initial action recognition model are trained, and the initial action recognition model is trained based on the quantity loss information and density loss information. When the training completion condition is met, the second action recognition model that has been trained is obtained.

[0160] The initial video frame classification network refers to the video frame classification network whose network parameters are initialized. This network classifies video frames into foreground and background frames. Foreground frames are those performing periodic actions, while background frames are those not performing periodic actions. The classification result refers to the classification result of each video frame in the training video, including the classification result of video frames that are foreground frames and those that are background frames. The training class label refers to the class label of each video frame in the training video, where label 1 can be used to represent foreground frames and label 0 can be used to represent background frames.

[0161] In this embodiment, the initial action recognition model also includes an initial video frame classification network, which can be built using a fully connected neural network. When classification training is required, the server can input periodic action localization features into the initial video frame classification network to classify and recognize periodic actions, obtaining the classification results corresponding to each video frame in the output video to be recognized. The server can then obtain the training category labels corresponding to the training videos and calculate the loss between the training category labels and the classification results using a classification loss function, which can be a binary cross-entropy loss function. Simultaneously, the server can obtain the quantity loss information and density loss information corresponding to the training videos through training the initial action recognition model. At this point, the server can use the category loss information to iteratively update the initial localization feature extraction network and the initial video frame classification network in the initial action recognition model, and then use the quantity loss information and density loss information to iteratively update the initial action recognition model. That is, the server can use a gradient descent algorithm to iteratively update the network parameters of the initial localization feature extraction network and the initial video frame classification network using the category loss information, and after the update is complete, iteratively update the model parameters of the initial action recognition model using the quantity loss information and density loss information. The server can also calculate the sum of category loss information, quantity loss information, and density loss information, and use the sum to iteratively update the initial action recognition model. Then, when the training completion condition is met, a trained second action recognition model is obtained based on the initial action recognition model at the time of training completion. This can be achieved by deleting the initial video frame classification network from the initial action recognition model, and then deploying and using the trained second action recognition model. Alternatively, the initial action recognition model at the time of training completion can be directly used as the trained second action recognition model, meaning that the second action recognition model can be used to simultaneously perform video frame classification on the input video.

[0162] In a specific embodiment, the category loss information can be calculated using the formula (4) shown below.

[0163] Formula (4)

[0164] in, This refers to category loss information. It refers to the classification and recognition result of a video frame at time t, which can be the predicted probability of the video frame being a foreground frame or a background frame. B refers to the true category label of the video frame at time t, and B represents the batch size of the video frames.

[0165] In the above embodiment, periodic action localization features are input into an initial video frame classification network for periodic action classification and recognition, obtaining the classification and recognition results corresponding to the video frames in the video to be recognized, and then calculating the category loss information. Finally, the initial localization feature extraction network and the initial video frame classification network in the initial action recognition model are trained using the category loss information, and the initial action recognition model is trained using quantity loss information and density loss information. When the training completion condition is met, a second action recognition model that has been trained is obtained. That is, updating and iterating the initial localization feature extraction network through category loss information can enhance the corresponding feature representation, thereby improving the accuracy of the obtained periodic action localization features, and thus improving the accuracy of action recognition.

[0166] In some embodiments, the initial video frame classification network in the initial action recognition model is trained based on category loss information, and the initial action recognition model is trained based on category loss information and density loss information. When the training completion condition is met, a first action recognition model that has been trained is obtained, including the following steps:

[0167] The loss between the anchor point pattern features corresponding to the anchor point video, the foreground pattern features corresponding to the foreground video, and the background pattern features corresponding to the background video is calculated to obtain triplet loss information. Based on the triplet loss information, the initial pattern feature extraction network in the initial action recognition model is trained. Based on the category loss information, the initial localization feature extraction network and the initial video frame classification network are trained. Based on the quantity loss information and density loss information, the initial action recognition model is trained. When the training completion condition is met, the trained target action recognition model is obtained.

[0168] In this embodiment, when training the initial action recognition model, triplet loss information, category loss information, quantity loss information, and density loss information can be calculated. The triplet loss information and category loss information are used as auxiliary supervision loss information during training, allowing the resulting similarity feature matrix to directly obtain higher similarity scores during action execution cycles and lower similarity scores between irrelevant cycles, thereby reducing the impact of noisy cycles. Specifically, the server uses triplet loss information to iteratively update the initial pattern feature extraction network in the initial action recognition model, uses category loss information to iteratively update the initial localization feature extraction network and the initial video frame classification network, and simultaneously uses quantity loss information and density loss information to iteratively update the initial action recognition model. In other words, the initial pattern feature extraction network is trained using triplet loss information, quantity loss information, and density loss information, and the initial localization feature extraction network is trained using category loss information, quantity loss information, and density loss information. Then, when the training completion condition is met, the trained target action recognition model is obtained. Finally, the server can deploy and use the target action recognition model.

[0169] In the above embodiments, by training the initial pattern feature extraction network using triplet loss information, training the initial localization feature extraction network and the initial video frame classification network using category loss information, and training the initial action recognition model using quantity loss information and density loss information, the pattern feature extraction network and localization feature extraction network in the trained target action recognition model can extract more robust features. This improves the accuracy and stability of the target action recognition model in recognizing periodic actions in real-world scenes, and enhances the generalizability and practicality of action recognition.

[0170] In a specific embodiment, such as Figure 8 The diagram illustrates the principle of training an action recognition model. Specifically, the server inputs a training video into an initial action recognition model to extract video embeddings. These embeddings are then fed into an initial localization feature extraction network and an initial pattern feature extraction network to obtain initial periodic action pattern features and initial periodic action localization features. These features are then fused to obtain enhanced periodic action features. These enhanced features are used to identify the number of periodic actions of the target object in the video, resulting in a periodic action density map. Finally, the periodic action density map is used for statistical calculations to determine the total number of periodic actions.

[0171] At this point, the server calculates the category loss information. That is, the server inputs the initial periodic motion localization features into the classification network of the video frames for classification and recognition, obtaining the classification and recognition results of the training video. Then, it uses the classification and recognition results and the training category labels of the training video to calculate the category loss information. The server can then use the category loss information to iteratively update the parameters of the video frame classification network and the initial periodic motion localization feature network, obtaining the updated periodic motion localization feature network and the updated classification network. Finally, the updated periodic motion localization feature network and the updated classification network are used as the initial periodic motion localization feature network and the initial classification network for the next iterative update.

[0172] The server then calculates the triplet loss information. When the training video is a positive sample, the server can obtain the pattern features of the plotted video and the background video. It then uses the initial periodic motion pattern features of the training video, the pattern features of the plotted video, and the pattern features of the background video to calculate the triplet loss information. This triplet loss information is then used to back-update the initial localization feature extraction network, resulting in an updated localization feature extraction network. This updated localization feature extraction network is then used as the initial localization feature extraction network for the next iteration.

[0173] The server then calculates density and quantity loss information. Specifically, it uses the periodic action density map and training density labels to calculate density loss information, and simultaneously uses the number of periodic actions and training quantity labels to calculate quantity loss information. The server then calculates the sum of these two losses and uses this sum to iteratively update the model parameters in the initial action recognition model, resulting in an updated action recognition model. This updated model is then used as the initial action recognition model for the next iteration. When the server determines that training is complete, it uses the trained initial action recognition model as the final action recognition model.

[0174] Finally, the server deploys the trained action recognition model and uses it to identify actions in the video to obtain the number of periodic actions of the target object. In other words, by using triplet loss information and category loss information as auxiliary supervision loss information, and density loss information and quantity loss information as quantity recognition supervision loss information to train the initial action recognition model, the trained action recognition model can effectively overcome the inconsistency in the duration of periodic actions and the noise interference introduced by the rest periods between periodic actions, thereby improving the accuracy of the obtained number of periodic actions, and thus improving the accuracy of action recognition.

[0175] In a specific embodiment, such as Figure 9The diagram illustrates the overall architecture of an action recognition model. This model includes a Multi-Scale Periodic Aware Representation (MPR) branch (pattern feature extraction network) and a Repeating Foreground Localization (RFL) branch (localization feature extraction network). Specifically, by introducing periodic priors through dual-branch multi-scale representation (DMR), time-related video enhancement representations can be extracted, thereby improving the accuracy of counting periodic actions in real-world scenes. In detail: the server embeds each video frame of video v into an embedding representation using an encoder, obtaining the video embedding representation X, which represents the content features of each video frame. Then, the video embedding representation X is input into different branch networks.

[0176] Branch 1 is constructed using hierarchical structures with different time scales, such as... Figure 3 The architecture is shown. For each timescale, there are two components: a distribution-specific similarity matrix and a scale-specific attention component. This multi-scale periodic awareness representation can be multiplied by different weight matrices on the similarity matrix to match various action distributions. Specifically, branch 1 obtains the DSM similarity matrix by calculating pairwise similarities between video frame features. Then, for each timescale, the similarity matrix is ​​passed through max-pooling layers with different kernel sizes to extract information with specific receptive domains, resulting in pooled features. For example, pooling can be performed using 1D max-pooling, where k represents the corresponding timescale, i.e., the pooling window size. These pooled features are then subjected to self-attention (SA) feature extraction to interact with event information, thereby obtaining periodic action pattern features for each timescale. Finally, the periodic action pattern features from each timescale are concatenated to obtain the target periodic action pattern features.

[0177] Branch 2 is a repetitive foreground localization branch based on a Temporal Convolutional Network (TCN). The goal of this branch is to locate motion cycles from noisy information and understand contextual information from feature sequences at the instance level. Branch 2 also includes a fully connected layer for foreground and background classification, providing classification results for calculating foreground and background localization losses. Initial periodic motion localization features are obtained by inputting the content features of each video frame into the TCN for feature extraction. These initial features are then fully connected to obtain the output periodic motion localization features of Branch 2. Finally, the target periodic motion pattern features and the periodic motion localization features are concatenated and dimensionality reduced using a fully connected operation to obtain the enhanced embedding representation X' of the video, i.e., the enhanced periodic motion features.

[0178] At this point, the enhanced periodic motion features are decoded by a decoder to obtain the output periodic motion density map D. Finally, the number of periodic motions of the target object in the input video is calculated using the periodic motion density map. That is, using this motion recognition model can effectively overcome the inconsistency in the duration of periodic motions and the noise interference introduced by the rest periods between periodic motions, thereby improving the accuracy of the obtained number of periodic motions, and thus improving the accuracy of motion recognition.

[0179] In a specific embodiment, such as Figure 10 The diagram illustrates a flowchart of an action recognition method, executed by a computer device, which can be a server or a terminal, preferably a server. The method specifically includes the following steps:

[0180] S1002, acquire the video to be recognized, input the video to be recognized into the action recognition model, extract the visual content of each video frame in the video to be recognized through the action recognition model, obtain the content features of each video frame, and input the content features of each video frame into the pattern feature extraction network and the localization feature extraction network respectively.

[0181] S1004, the first video frame content features and the second video frame content features are obtained from the content features of each video frame through the pattern feature extraction network, and the similarity weights corresponding to the first video frame content features and the second video frame content features are obtained. The transpose of the first video frame content features is calculated to obtain the first transposed feature.

[0182] S1006, the pattern feature extraction network uses the first transposed feature, similarity weight, and second video frame content feature to calculate the similarity feature between the first video frame content feature and the second video frame content feature, and then iterates through the content features of each video frame to obtain each similar feature.

[0183] S1008: The pattern feature extraction network performs pooling operations on each similar feature according to the corresponding pooling window to obtain the pooled features corresponding to each similar feature.

[0184] S1010: The query feature and key feature corresponding to each pooling feature are extracted using a pattern feature extraction network. Self-attention weights are calculated based on the query feature and key feature to obtain the self-attention weights. The corresponding pooling features are then weighted according to these self-attention weights to obtain the self-attention features corresponding to each pooling feature.

[0185] S1012, through the pattern feature extraction network, performs fully connected operations on the attention features of each time scale to obtain the periodic action pattern features corresponding to each time scale, and fuses the periodic action pattern features corresponding to each time scale to obtain the output target periodic action pattern features.

[0186] S1014: The localization feature extraction network extracts contextual semantic information from the content features of each video frame through temporal convolution to obtain initial periodic action localization features. The initial periodic action localization features are then subjected to fully connected operations to obtain the output periodic action localization features.

[0187] S1016, the target periodic action pattern features and periodic action localization features are fused by the action recognition model to obtain the periodic action enhancement features corresponding to the video to be recognized.

[0188] S1018: The video enhancement features are subjected to attention transformation using the action recognition model to obtain target transformation features. Fully connected operations are then performed on these target transformation features to obtain a periodic action density map. The number of periodic actions of the target object in the video to be recognized is calculated based on the periodic action density map.

[0189] In the above embodiments, by using an action recognition model that includes a pattern feature extraction network and a location feature extraction network for action recognition, the inconsistency in the duration of periodic actions and the noise interference introduced by the rest between periodic actions can be effectively overcome, thereby improving the accuracy of the number of periodic actions obtained, that is, improving the accuracy of action recognition.

[0190] In one specific embodiment, this action recognition method is applied to an intelligent sports platform. Specifically, the intelligent sports platform can acquire real-time videos of users exercising and then detect these videos. For example, it can identify the number of periodic movements in the user's exercise video, i.e., perform Relative Athleticism Score (RAC). This Relative Athleticism Score refers to the task of calculating the number of periodic human movements in a video. Figure 11The image shows a schematic diagram of a user performing squat exercises in a video. However, the video exhibits inconsistencies in the squatting motions; for example, squat 1 lasts 1.1 seconds, squat 2 lasts 2.0 seconds, and squat 3 lasts 2.9 seconds. Clearly, the timescales of squats 1, 2, and 3 are inconsistent. The intelligent sports platform then inputs the user's squatting video into a target motion recognition model, which is pre-trained and deployed within the platform. The target motion recognition model extracts the embedding representation of each video frame from the squatting video, obtaining the content features of each frame. These content features are then input into a pattern feature extraction network for feature extraction, yielding the output target periodic motion pattern features. Simultaneously, the content features are input into a localization feature extraction network for feature extraction, yielding the output periodic motion localization features. The target motion recognition model then fuses the target periodic motion pattern features and the periodic motion localization features to obtain enhanced periodic motion features. These enhanced features are then used to identify the number of squatting movements performed by the user in the squatting video. Figure 13 The diagram illustrates the motion recognition results. A user uploads a squat video to the server via their terminal. The server uses a target motion recognition model to count the number of squat movements in the uploaded video. Finally, the server returns this result to the user's terminal for display. The user's terminal shows the uploaded squat video and the recognized number of squat movements: 20. The intelligent sports platform can also recognize the cyclical movements of other sports, such as... Figure 12 The image shows a schematic diagram of a user performing leg raises, with rest periods between the leg raises. The foreground represents the leg raise portion of the video, while the background represents the rest portion. The intelligent fitness platform then inputs this leg raise video into a target motion recognition model. Through pattern feature extraction and localization feature extraction networks, the model performs motion recognition, resulting in the number of leg raises performed by the user in the video. In other words, by using the target motion recognition model, the number of periodic movements performed by the user can be accurately monitored in complex fitness scenarios, improving the reliability and robustness of periodic motion recognition.

[0191] In one specific embodiment, this action recognition method is applied to a robot testing platform. This platform can test data on robots performing periodic actions, such as identifying the number of periodic actions a robot can perform within a fixed time period. Specifically, when testing a robot, the testing platform can collect video footage of the robot's periodic motion over a fixed time period and then perform action recognition on this video. For example, the video can be input into a target action recognition model, where pattern feature extraction and localization feature extraction networks are used for action recognition. This yields the number of periodic actions performed by the robot in the video output by the target action recognition model. Using this target action recognition model improves the accuracy of counting the robot's periodic actions.

[0192] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0193] Based on the same inventive concept, this application also provides an action recognition device for implementing the action recognition method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more action recognition device embodiments provided below can be found in the limitations of the action recognition method described above, and will not be repeated here.

[0194] In one embodiment, such as Figure 14 As shown, an action recognition device 1400 is provided, including: a representation feature extraction module 1402, a pattern feature extraction module 1404, a pattern feature fusion module 1406, a positioning recognition module 1408, an enhanced feature acquisition module 1410, and a quantity recognition module 1412, wherein:

[0195] The feature extraction module 1402 is used to acquire the video to be identified and extract the visual content of each video frame in the video to be identified, so as to obtain the content features of each video frame.

[0196] The pattern feature extraction module 1404 is used to obtain the similarity between the content features of each video frame, obtain each similar feature, and extract periodic action patterns at different time scales based on each similar feature, so as to obtain the periodic action pattern features corresponding to each time scale.

[0197] The pattern feature fusion module 1406 is used to fuse the periodic action pattern features corresponding to each time scale to obtain the target periodic action pattern features.

[0198] The positioning and recognition module 1408 is used to perform positioning and recognition of periodic actions based on the content features of each video frame, and obtain the positioning features of periodic actions.

[0199] The enhanced feature module 1410 is used to fuse the target periodic action pattern features and periodic action localization features to obtain the periodic action enhanced features corresponding to the video to be identified.

[0200] The quantity recognition module 1412 is used to recognize the number of periodic actions of the target object in the video to be recognized based on the periodic action enhancement features, so as to obtain the number of periodic actions of the target object in the video to be recognized.

[0201] In one embodiment, the pattern feature extraction module 1404 is further configured to extract perceptual domain information from each similar feature by pooling at the corresponding time scale to obtain each pooling feature; to perform time information interaction on each pooling feature by self-attention to obtain each attention feature; and to perform fully connected operations on each attention feature to obtain the periodic action pattern features corresponding to each time scale.

[0202] In one embodiment, the pattern feature extraction module 1404 is further configured to perform pooling operations on the target similar features among the similar features according to the corresponding pooling window to obtain the pooled features corresponding to the target similar features.

[0203] In one embodiment, the pattern feature extraction module 1404 is further configured to extract the query features and key features of the target pooling feature from each pooling feature; calculate the self-attention weights based on the query features and key features to obtain the self-attention weights; and weight the target pooling feature according to the self-attention weights to obtain the self-attention features corresponding to the target pooling feature.

[0204] In one embodiment, the positioning and recognition module 1408 is further configured to extract contextual semantic information based on the content features of each video frame through temporal convolution to obtain initial periodic action positioning features; and to perform fully connected operations on the initial periodic action positioning features to obtain periodic action positioning features.

[0205] In one embodiment, the pattern feature extraction module 1404 is further configured to obtain the first video frame content feature and the second video frame content feature from the content features of each video frame, and obtain the similarity weights corresponding to the first video frame content feature and the second video frame content feature; calculate the transpose of the first video frame content feature to obtain the first transposed feature, and calculate the similarity feature between the first video frame content feature and the second video frame content feature based on the first transposed feature, the similarity weights, and the second video frame content feature.

[0206] In one embodiment, the quantity recognition module 1412 is further configured to perform attention transformation on the video enhancement features to obtain target transformation features, perform fully connected operation based on the target transformation features to obtain a periodic action density map, and calculate the number of periodic actions of the target object in the video to be recognized based on the periodic action density map to obtain the number of periodic actions of the target object in the video to be recognized.

[0207] In one embodiment, the feature extraction module 1402 is further configured to acquire the video to be identified, uniformly sample the video to be identified according to the number of target frames to obtain each video frame, and extract the visual content of each video frame to obtain the content features of each video frame.

[0208] In one embodiment, the motion recognition device 1400 further includes:

[0209] The model recognition module is used to input the video to be recognized into the action recognition model. The action recognition model extracts the visual content of each video frame, obtaining the content features of each frame. The pattern feature extraction network in the action recognition model obtains the similarity between the content features of each video frame, obtaining similar features. Based on these similar features, periodic action patterns at different time scales are extracted, obtaining periodic action pattern features corresponding to each time scale. These periodic action pattern features are then fused to obtain the target periodic action pattern features. The localization feature extraction network in the action recognition model performs localization and recognition of periodic actions based on the content features of each video frame, obtaining periodic action localization features. The action recognition model then fuses the target periodic action pattern features and the periodic action localization features to obtain the periodic action enhancement features corresponding to the video to be recognized. Finally, the action recognition model identifies the number of periodic actions of the target object in the video to be recognized according to the periodic action enhancement features, obtaining the number of periodic actions of the target object in the video to be recognized.

[0210] In one embodiment, the motion recognition device 1400 further includes:

[0211] The model training module is used to acquire training videos, training quantity labels, and training density map labels; input the training videos into the initial action recognition model to perform periodic action quantity recognition, and obtain the training periodic action density map and the training periodic action quantity; obtain the loss between the training periodic action quantity and the training quantity labels to obtain quantity loss information, and obtain the loss between the training periodic action density map and the training density map labels to obtain density loss information; train the initial action recognition model based on the quantity loss information and density loss information, and obtain the trained action recognition model when the training completion condition is met.

[0212] In one embodiment, the model training module is further configured to acquire anchor pattern features corresponding to the anchor video, foreground pattern features corresponding to the foreground video, and background pattern features corresponding to the background video; acquire the loss between the anchor pattern features, foreground pattern features, and background pattern features to obtain triplet loss information; train the initial pattern feature extraction network in the initial action recognition model based on the triplet loss information, and train the initial action recognition model based on the quantity loss information and density loss information; and obtain the first action recognition model after training is completed when the training completion condition is met.

[0213] In one embodiment, the initial action recognition model includes an initial video frame classification network. The model training module is further configured to input periodic action localization features into the initial video frame classification network to classify and recognize periodic actions, thereby obtaining the classification and recognition results corresponding to the video frames in the video to be recognized; obtain the training category labels corresponding to the training videos, and obtain the loss between the training category labels and the classification and recognition results to obtain category loss information; train the initial localization feature extraction network and the initial video frame classification network in the initial action recognition model based on the category loss information, and train the initial action recognition model based on the quantity loss information and density loss information. When the training completion condition is met, a second action recognition model that has been trained is obtained.

[0214] In one embodiment, the model training module is further used to obtain the loss between the anchor point pattern features corresponding to the anchor point video, the foreground pattern features corresponding to the foreground video, and the background pattern features corresponding to the background video, to obtain triplet loss information; based on the triplet loss information, the initial pattern feature extraction network in the initial action recognition model is trained; based on the category loss information, the initial localization feature extraction network and the initial video frame classification network are trained; and based on the quantity loss information and density loss information, the initial action recognition model is trained. When the training completion condition is met, the trained target action recognition model is obtained.

[0215] Each module in the aforementioned motion recognition device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0216] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 15 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores data such as videos to be recognized, training videos, training labels, and action recognition models. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements an action recognition method.

[0217] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 16As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements an action recognition method. The display unit of the computer device is used to form a visually visible image. It can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0218] Those skilled in the art will understand that Figure 14 or Figure 15 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0219] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0220] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0221] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0222] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0223] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0224] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0225] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An action recognition method, characterized in that, The method includes: The video to be identified is acquired, and the visual content of each video frame in the video to be identified is extracted to obtain the content features of each video frame. The similarity between the content features of each video frame is obtained to obtain each similar feature, and the periodic motion patterns at different time scales are extracted based on the similar features to obtain the periodic motion pattern features corresponding to each time scale. The periodic action pattern features corresponding to each time scale are fused to obtain the target periodic action pattern features; Based on the content features of each video frame, the location and recognition of periodic actions are performed to obtain periodic action location features; The target periodic motion pattern features and the periodic motion localization features are fused to obtain the periodic motion enhancement features corresponding to the video to be identified; Based on the enhanced periodic motion features, the number of periodic motions of the target object in the video to be identified is determined, thereby obtaining the number of periodic motions of the target object in the video to be identified.

2. The method according to claim 1, characterized in that, The step of extracting periodic action patterns at different time scales based on the aforementioned similar features to obtain periodic action pattern features corresponding to each time scale includes: The receptive domain information of each similar feature is extracted by pooling at the corresponding time scale to obtain each pooling feature; Each pooling feature is then subjected to temporal information interaction through self-attention to obtain its respective attention feature; The attention features are each subjected to a fully connected operation to obtain the periodic action pattern features corresponding to each time scale.

3. The method according to claim 2, characterized in that, The step of extracting perceptual domain information from each of the similar features through pooling at the corresponding time scale to obtain each pooled feature includes: The target similar features among the aforementioned similar features are pooled according to the corresponding pooling window to obtain the pooled features corresponding to the target similar features.

4. The method according to claim 2, characterized in that, The step of interacting with each pooling feature through self-attention to obtain its respective attention feature includes: Extract the query feature and key feature of the target pooling feature from each pooling feature; Self-attention weights are calculated based on the query features and the key features to obtain the self-attention weights. The target pooling feature is weighted according to the self-attention weights to obtain the self-attention feature corresponding to the target pooling feature.

5. The method according to claim 1, characterized in that, The process of locating and identifying periodic actions based on the content features of each video frame to obtain periodic action location features includes: Based on the content features of each video frame, contextual semantic information is extracted through temporal convolution to obtain initial periodic action localization features; The initial periodic action localization features are subjected to a fully connected operation to obtain the periodic action localization features.

6. The method according to claim 1, characterized in that, The step of obtaining the similarity between the content features of each video frame to obtain each similar feature includes: The first video frame content features and the second video frame content features are obtained from the content features of each video frame, and the similarity weights corresponding to the first video frame content features and the second video frame content features are obtained. Obtain the transpose of the content features of the first video frame to obtain the first transpose feature, and calculate the similarity features between the content features of the first video frame and the content features of the second video frame based on the first transpose feature, the similarity weight and the content features of the second video frame.

7. The method according to claim 1, characterized in that, The step of identifying the number of periodic actions of a target object in the video to be identified based on the periodic action enhancement features, to obtain the number of periodic actions of the target object in the video to be identified, includes: The video enhancement features are subjected to attention transformation to obtain target transformation features. Based on the target transformation features, a fully connected operation is performed to obtain a periodic action density map. The number of periodic actions of the target object in the video to be identified is calculated based on the periodic action density map.

8. The method according to claim 1, characterized in that, The process of acquiring the video to be identified and extracting the visual content of each video frame from the video to obtain the content features of each video frame includes: The video to be identified is obtained, and the video to be identified is sampled uniformly according to the number of target frames to obtain each video frame; The visual content of each video frame is extracted to obtain the content features of each video frame.

9. The method according to claim 1, characterized in that, The method further includes: The video to be identified is input into the action recognition model, and the visual content of each video frame in the video to be identified is extracted by the action recognition model to obtain the content features of each video frame; The similarity between the content features of each video frame is obtained by the pattern feature extraction network in the action recognition model, and each similar feature is obtained. Based on the similar features, periodic action patterns at different time scales are extracted to obtain periodic action pattern features corresponding to each time scale. The periodic action pattern features corresponding to each time scale are then fused to obtain target periodic action pattern features. The localization feature extraction network in the action recognition model is used to locate and identify periodic actions in the content features of each video frame, thereby obtaining periodic action localization features. The motion recognition model fuses the target periodic motion pattern features and the periodic motion localization features to obtain the periodic motion enhancement features corresponding to the video to be recognized. The motion recognition model identifies the number of periodic actions of the target object in the video to be identified according to the periodic motion enhancement features, thereby obtaining the number of periodic actions of the target object in the video to be identified.

10. The method according to claim 9, characterized in that, The training of the action recognition model includes the following steps: Obtain training video, training quantity labels, and training density map labels; The training video is input into the initial action recognition model to identify the number of periodic actions, thereby obtaining the training periodic action density map and the number of training periodic actions. Obtain the loss between the number of actions in the training cycle and the training quantity label to obtain quantity loss information, and obtain the loss between the action density map in the training cycle and the training density map label to obtain density loss information; The initial action recognition model is trained based on the quantity loss information and the density loss information. When the training completion condition is met, the trained action recognition model is obtained.

11. The method according to claim 10, characterized in that, The process of training the initial action recognition model based on the quantity loss information and the density loss information, and obtaining the trained action recognition model when the training completion condition is met, includes: Obtain the anchor point pattern features corresponding to the anchor point video, the foreground pattern features corresponding to the foreground video, and the background pattern features corresponding to the background video. The loss among the anchor point pattern features, the foreground pattern features, and the background pattern features is obtained to obtain triplet loss information; The initial pattern feature extraction network in the initial action recognition model is trained based on the triplet loss information, and the initial action recognition model is trained based on the quantity loss information and the density loss information. When the training completion condition is met, the first action recognition model that has been trained is obtained.

12. The method according to claim 10, characterized in that, The initial action recognition model includes an initial video frame classification network. The initial action recognition model is trained based on the quantity loss information and the density loss information. When the training completion condition is met, a trained action recognition model is obtained, including: The periodic action localization features are input into the initial video frame classification network to classify and identify periodic actions, thereby obtaining the classification and identification results corresponding to the video frames in the video to be identified. Obtain the training category label corresponding to the training video, and obtain the loss between the training category label and the classification recognition result to obtain category loss information; The initial localization feature extraction network and the initial video frame classification network in the initial action recognition model are trained based on the category loss information, and the initial action recognition model is trained based on the quantity loss information and the density loss information. When the training completion condition is met, the trained second action recognition model is obtained.

13. The method according to claim 12, characterized in that, The method further includes: The loss between the anchor point pattern features corresponding to the anchor point video, the foreground pattern features corresponding to the foreground video, and the background pattern features corresponding to the background video is obtained to get the triplet loss information. The initial pattern feature extraction network in the initial action recognition model is trained based on the triplet loss information, the initial localization feature extraction network and the initial video frame classification network are trained based on the category loss information, and the initial action recognition model is trained based on the quantity loss information and the density loss information. When the training completion condition is met, the trained target action recognition model is obtained.

14. A motion recognition device, characterized in that, The device includes: The feature extraction module is used to acquire the video to be identified and extract the visual content of each video frame in the video to be identified, so as to obtain the content features of each video frame. The pattern feature extraction module is used to obtain the similarity between the content features of each video frame, obtain each similar feature, and extract periodic action patterns at different time scales based on each similar feature, so as to obtain the periodic action pattern features corresponding to each time scale. The pattern feature fusion module is used to fuse the periodic action pattern features corresponding to each time scale to obtain the target periodic action pattern features. The positioning and recognition module is used to perform positioning and recognition of periodic actions based on the content features of each video frame, and obtain periodic action positioning features; The enhanced feature acquisition module is used to fuse the target periodic action pattern features and the periodic action localization features to obtain the periodic action enhanced features corresponding to the video to be identified; The quantity recognition module is used to identify the number of periodic actions of the target object in the video to be recognized based on the periodic action enhancement features, so as to obtain the number of periodic actions of the target object in the video to be recognized.

15. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 13.

16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 13.

17. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 13.