Action recognition methods, devices, terminal equipment, and storage media
By using frame-by-frame feature map feature transfer and action recognition, the problem of insufficient action recognition speed in existing technologies is solved, and fast action recognition is achieved on lightweight embedded development boards.
Patent Information
- Application Number
- CN202310955722.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-31
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2043-07-31
AI Technical Summary
In existing action recognition technologies, 3D convolutional networks are slow and difficult to use in engineering practice, while 2D networks are faster but still not fast enough, especially for action recognition on lightweight embedded development boards.
By extracting several frames of feature maps from the video to be identified, feature transfer and action recognition are performed sequentially. Using preset feature transfer and recognition rules, action differences are judged frame by frame, and finally the action recognition result is obtained through analysis.
This reduces the processing load on the motion recognition system and improves the motion recognition speed, enabling rapid motion recognition on lightweight embedded development boards.
Smart Images

Figure CN117058753B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of action recognition, and more particularly to an action recognition method, apparatus, terminal device, and storage medium. Background Technology
[0002] Recognition technology is now widely used in various industries, with a variety of methods and all showing remarkable performance.
[0003] Action recognition can be viewed as a data structure composed of frames arranged in chronological order over a period of time. Existing solutions include 3D convolution-based action recognition, such as C3D, Res3D / 3D ResNet, LTC, and I3D. 3D convolution can learn the temporal relationships between video frames. However, 3D convolutional networks are relatively slow, with slow training and inference speeds, making them difficult to use in engineering practice. While 2D networks are faster, they are still not fast enough, especially for some lightweight embedded development boards. For example, current 2D networks need to infer 8 or 16 frames of images simultaneously, then exchange frame information based on the temporal relationship of the frame rate before classification. Summary of the Invention
[0004] The main objective of this application is to provide an action recognition method, device, terminal equipment, and storage medium, with the aim of improving the speed of action recognition.
[0005] To achieve the above objectives, this application provides an action recognition method, the action recognition method comprising:
[0006] Obtain the video to be recognized;
[0007] Feature extraction is performed on the video to be identified to obtain several frames of feature maps to be identified;
[0008] According to the preset feature transfer and recognition rules, the feature transfer and action recognition are performed on the several frames of feature maps to be recognized in sequence to obtain several action recognition results;
[0009] The results of the action recognition are analyzed to obtain the final action recognition result.
[0010] Optionally, the step of sequentially performing feature transfer and action recognition on the plurality of frames of feature maps to be identified according to preset feature transfer and recognition rules includes:
[0011] Determine whether the current feature map to be identified is the same as the feature map in the first frame;
[0012] If the current feature map to be identified is the first frame feature map to be identified, then action recognition is performed on the current feature map to be identified;
[0013] Based on the preset learning ratio parameter, features are extracted from the current feature map to be identified according to the ratio to obtain the transferred features of the current feature map to be identified;
[0014] If the current feature map to be identified is not the first frame feature map to be identified, then according to the learning ratio parameter, the current feature map to be identified is extracted proportionally to obtain the current feature map after feature extraction and the transferred features of the current feature map to be identified.
[0015] The current feature map after feature extraction is fused with the transferred features extracted from the feature map to be identified in the previous frame to obtain the fused feature map to be identified.
[0016] Action recognition is performed on the fused feature map to be identified.
[0017] Optionally, the step of performing action recognition on the plurality of frames of feature maps to be identified includes:
[0018] The preset action recognition neural network model is used to perform action recognition on the several frames of feature maps to be recognized.
[0019] Optionally, before the step of performing action recognition on the plurality of frames of feature maps to be recognized using a preset action recognition neural network model, the method further includes:
[0020] The steps for creating the action recognition neural network model include:
[0021] Acquire sample action videos and corresponding sample action information;
[0022] The motion recognition neural network model is obtained by training the sample motion video and the sample motion information based on the focal loss function.
[0023] Optionally, the step of training the action recognition neural network model based on the focal loss function on the sample action video and the sample action information includes:
[0024] Based on preset image frame extraction rules, several frame sample feature maps are extracted from the sample action video.
[0025] The feature maps of the several frames are classified to obtain several sample classification values;
[0026] The total sample classification value is obtained by summing the classification values of the given samples.
[0027] The action recognition neural network model is obtained by training the total sample classification values and the sample action information using the focal loss function.
[0028] Optionally, before the step of training the action recognition neural network model based on the focal loss function on the sample action video and the sample action information, the method further includes:
[0029] Obtain the parameters of the pre-trained model;
[0030] Based on the pre-trained model parameters and the focal loss function, the sample action video and the sample action information are trained to obtain the action recognition neural network model.
[0031] Optionally, before the step of training the sample action video and the sample action information based on the pre-trained model parameters and the focal loss function to obtain the action recognition neural network model, the method further includes:
[0032] Freeze the batch normalization layer in the neural network.
[0033] This application also proposes an action recognition device, the action recognition device comprising:
[0034] The video acquisition module is used to acquire the video to be recognized;
[0035] The feature extraction module is used to extract features from the video to be identified, and obtain several frames of feature maps to be identified;
[0036] The behavior recognition module is used to sequentially perform feature transfer and action recognition on the several frames of feature maps to be recognized according to preset feature transfer and recognition rules, and obtain several action recognition results.
[0037] The comprehensive judgment module is used to analyze the several action recognition results to obtain the final action recognition result.
[0038] This application also proposes a terminal device, which includes a memory, a processor, and an action recognition program stored in the memory and executable on the processor. When the action recognition program is executed by the processor, it implements the steps of the action recognition method described above.
[0039] This application also proposes a computer-readable storage medium storing an action recognition program, which, when executed by a processor, implements the steps of the action recognition method described above.
[0040] The action recognition method, apparatus, terminal device, and storage medium proposed in this application involve: acquiring a video to be recognized; extracting features from the video to obtain several frames of feature maps to be recognized; sequentially performing feature transfer and action recognition on the several frames of feature maps according to preset feature transfer and recognition rules to obtain several action recognition results; and analyzing the several action recognition results to obtain the final action recognition result. By extracting several frames of feature maps from the video to be recognized, and then sequentially transferring some features from each frame of feature maps to the next frame of feature maps according to preset feature transfer and recognition rules before performing action recognition, it is possible to determine the feature differences between the previous and current frames when recognizing the current feature map, thereby determining the difference in the user's actions between the current and previous frame of feature maps, and thus obtaining the action recognition result for each frame of feature maps. Then, analyzing the several action recognition results to obtain the final action recognition result, it is understood that the above process performs action recognition frame by frame on the feature maps to be recognized, which reduces the processing pressure on the action recognition system and improves the speed of action recognition. Attached Figure Description
[0041] Figure 1 This is a schematic diagram of the functional modules of the terminal device to which the motion recognition device of this application belongs;
[0042] Figure 2 This is a flowchart illustrating a first exemplary embodiment of the action recognition method of this application;
[0043] Figure 3 This is a flowchart illustrating a second exemplary embodiment of the action recognition method of this application;
[0044] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0045] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0046] The main solution of this application embodiment is as follows: First, acquire the video to be identified; second, extract features from the video to obtain several frames of feature maps to be identified; third, according to preset feature transfer and recognition rules, sequentially transfer features and perform action recognition on the several frames of feature maps to be identified, obtaining several action recognition results; fourth, analyze the several action recognition results to obtain the final action recognition result. Extracting several frames of feature maps to be identified from the video to be identified, and then sequentially transferring some features from each frame of feature maps to the next frame of feature maps to be identified according to preset feature transfer and recognition rules before performing action recognition, allows for the determination of feature differences between the previous and current frames when recognizing the current feature map, thereby determining the difference in user actions between the current and previous frames of feature maps to be identified, and thus obtaining the action recognition result for each frame of feature maps to be identified. Then, analyzing the several action recognition results yields the final action recognition result. It can be understood that the above process performs action recognition frame by frame on the feature maps to be identified, which reduces the processing pressure on the action recognition system and improves the speed of action recognition.
[0047] This application takes into account that existing solutions in related technologies include action recognition based on 3D convolution, such as C3D, Res3D / 3D ResNet, LTC, and I3D. 3D convolution can learn the temporal relationships between video frames. However, 3D convolutional networks are relatively slow, with slow training and inference speeds, making them difficult to use in engineering practice. While 2D networks are faster, they are still not fast enough, especially for some lightweight embedded development boards. For example, current 2D networks need to infer 8 or 16 frames of images simultaneously, then exchange frame information based on the temporal relationship of the frame rate before classification.
[0048] Based on this, this application proposes a solution that extracts several frames of feature maps to be identified from the video to be identified. Then, according to preset feature transfer and recognition rules, partial features of each frame of feature maps are sequentially transferred to the next frame of feature maps for action recognition. This allows for the identification of feature differences between the previous and current frames when recognizing the current feature map, thereby determining the difference in user actions between the current and previous frame of feature maps and obtaining the action recognition result for each frame of feature maps. The results of these action recognitions are then analyzed to obtain the final action recognition result. It is understood that the above process performs action recognition frame by frame on the feature maps to be identified, which reduces the processing load on the action recognition system and thus improves the speed of action recognition.
[0049] Specifically, refer to Figure 1 , Figure 1This is a schematic diagram of the functional modules of the terminal device to which the motion recognition device of this application belongs. The motion recognition device can be a data processing device independent of the terminal device, or it can be carried on the terminal device in the form of hardware or software.
[0050] In this embodiment, the terminal device to which the motion recognition device belongs includes at least an output module 110, a processor 120, a memory 130, and a communication module 140.
[0051] The memory 130 stores the operating system and motion recognition program to acquire the video to be recognized; it extracts features from the video to be recognized to obtain several frames of feature maps to be recognized; according to preset feature transfer and recognition rules, it sequentially performs feature transfer and motion recognition on the several frames of feature maps to be recognized to obtain several motion recognition results; it analyzes the several motion recognition results to obtain the final motion recognition result, which is stored in the memory 130; the output module 110 can be a display screen, speaker, etc. The communication module 140 can include a WIFI module, a mobile communication module, and a Bluetooth module, etc., and communicates with external devices or servers through the communication module 140.
[0052] When the action recognition program in memory 130 is executed by the processor, it performs the following steps:
[0053] Obtain the video to be recognized;
[0054] Feature extraction is performed on the video to be identified to obtain several frames of feature maps to be identified;
[0055] According to the preset feature transfer and recognition rules, the feature transfer and action recognition are performed on the several frames of feature maps to be recognized in sequence to obtain several action recognition results;
[0056] The results of the action recognition are analyzed to obtain the final action recognition result.
[0057] Furthermore, when the action recognition program in memory 130 is executed by the processor, it also performs the following steps:
[0058] If the current feature map to be identified is the first frame feature map to be identified, then action recognition is performed on the current feature map to be identified;
[0059] Based on the preset learning ratio parameter, features are extracted from the current feature map to be identified according to the ratio to obtain the transferred features of the current feature map to be identified;
[0060] If the current feature map to be identified is not the first frame feature map to be identified, then according to the learning ratio parameter, the current feature map to be identified is extracted proportionally to obtain the current feature map after feature extraction and the transferred features of the current feature map to be identified.
[0061] The current feature map after feature extraction is fused with the transferred features extracted from the feature map to be identified in the previous frame to obtain the fused feature map to be identified.
[0062] Action recognition is performed on the fused feature map to be identified.
[0063] Furthermore, when the action recognition program in memory 130 is executed by the processor, it also performs the following steps:
[0064] Action recognition is performed on the several frames of feature maps to be recognized using a pre-set action recognition neural network model.
[0065] Furthermore, when the action recognition program in memory 130 is executed by the processor, it also performs the following steps:
[0066] The steps for creating the action recognition neural network model include: acquiring sample action videos and corresponding sample action information;
[0067] The motion recognition neural network model is obtained by training the sample motion video and the sample motion information based on the focal loss function.
[0068] Furthermore, when the action recognition program in memory 130 is executed by the processor, it also performs the following steps:
[0069] Based on preset image frame extraction rules, several frame sample feature maps are extracted from the sample action video.
[0070] The feature maps of the several frames are classified to obtain several sample classification values;
[0071] The total sample classification value is obtained by summing the classification values of the given samples.
[0072] The action recognition neural network model is obtained by training the total sample classification values and the sample action information using the focal loss function.
[0073] Furthermore, when the action recognition program in memory 130 is executed by the processor, it also performs the following steps:
[0074] Obtain the parameters of the pre-trained model;
[0075] Based on the pre-trained model parameters and the focal loss function, the sample action video and the sample action information are trained to obtain the action recognition neural network model.
[0076] Furthermore, when the action recognition program in memory 130 is executed by the processor, it also performs the following steps:
[0077] Freeze the batch normalization layer in the neural network.
[0078] This embodiment, through the above-described scheme, acquires a video to be identified; extracts features from the video to obtain several frames of feature maps to be identified; according to preset feature transfer and recognition rules, performs feature transfer and action recognition sequentially on the several frames of feature maps to be identified, obtaining several action recognition results; and analyzes the several action recognition results to obtain the final action recognition result. Extracting several frames of feature maps to be identified from the video to be identified, and then, according to preset feature transfer and recognition rules, sequentially transferring some features from each frame of feature maps to the next frame of feature maps to be identified before action recognition, allows for the determination of feature differences between the previous and current frames when recognizing the current feature map, thereby determining the difference in user actions between the current and previous frames of feature maps to be identified, thus obtaining the action recognition result for each frame of feature maps to be identified. Then, analyzing the several action recognition results yields the final action recognition result. It can be understood that the above process performs action recognition frame by frame on the feature maps to be identified, which reduces the processing pressure on the action recognition system and improves the speed of action recognition.
[0079] Based on, but not limited to, the terminal device architecture described above, this application proposes method embodiments.
[0080] Reference Figure 2 , Figure 2 This is a flowchart illustrating a first exemplary embodiment of the action recognition method of this application.
[0081] An embodiment of the present invention provides an action recognition method, the method comprising:
[0082] Step S10: Obtain the video to be recognized;
[0083] Existing action recognition solutions include 3D convolutional action recognition, which can learn the temporal relationships between video frames. Examples include C3D, Res3D / 3D ResNet, LTC, and I3D. However, 3D convolutional networks are relatively slow, with slow training and inference speeds, making them difficult to use in engineering practice. While 2D networks are faster, they are still not fast enough, especially for some lightweight embedded development boards. For example, current 2D networks need to infer 8 or 16 frames of images simultaneously, then exchange frame information based on the temporal relationship of the frame rate before classification.
[0084] In existing technologies, at least eight frames of images to be recognized are generally required. Specifically, a complete action recognition is achieved by analyzing the feature changes of these eight frames. However, inputting eight feature maps into the action recognition system at once increases the processing pressure on the system and slows down the process.
[0085] Therefore, this embodiment proposes to extract several frames of feature maps to be identified from the video to be identified, and then, according to preset feature transfer and recognition rules, sequentially transfer some features of each frame of feature maps to the next frame of feature maps before performing action recognition. This allows for the identification of feature differences between the previous and current frames each time the current feature map is identified, thereby determining the difference in user actions between the current and previous frame of feature maps, and thus obtaining the action recognition result for each frame of feature maps. Then, the several action recognition results are analyzed to obtain the final action recognition result. It can be understood that the above process performs action recognition frame by frame on the feature maps to be identified, which reduces the processing pressure on the action recognition system and thus improves the speed of action recognition.
[0086] Specifically, the video to be identified needs to be acquired first, which can be done by using a camera or other video capture devices to capture live video. This is suitable for applications that require real-time motion recognition or processing of live streaming video. For example, capturing and recognizing real-time human movements using a webcam or mobile device camera.
[0087] Step S20: Extract features from the video to be identified to obtain several frames of feature maps to be identified;
[0088] To better perform action recognition on videos, it is necessary to extract effective features from the video and remove noise outside of action recognition. Action features can be extracted using an image extraction model to obtain several frames of feature maps to be recognized.
[0089] Specifically, image features can be extracted based on frames: traditional feature extraction methods in computer vision, such as SIFT (Scale Invariant Feature Transform) and HOG (Histogram of Oriented Gradients), can be used to extract image features for each frame.
[0090] Specifically, image feature extraction can be achieved through deep learning-based methods: pre-trained deep learning models, such as convolutional neural network (CNN) models (e.g., VGG, ResNet, Inception, etc.), can be used to extract feature maps for each frame. These models are typically trained on large-scale image datasets and can capture high-level semantic features in images.
[0091] Specifically, feature extraction can be achieved through optical flow: optical flow calculates the motion information of pixels by analyzing the changes in pixel intensity between consecutive frames. Optical flow algorithms (such as the Lucas-Kanade algorithm and the Farneback algorithm) can be used to extract motion features from the video to be identified.
[0092] Step S30: According to the preset feature transfer and recognition rules, the feature transfer and action recognition are performed on the several frames of feature maps to be recognized in sequence to obtain several action recognition results;
[0093] To avoid putting excessive pressure on the action system program by inputting multiple feature maps at once, which would slow down action recognition, feature transfer and action recognition can be performed on multiple frames of feature maps sequentially through preset feature transfer and recognition rules.
[0094] Specifically, partial features of each frame of the feature map to be identified are sequentially passed to the next frame of the feature map to be identified. Then, the feature map to be identified in the next frame is fused with the features passed from the previous frame to form a new feature map to be identified, and then action recognition is performed.
[0095] This ensures that each fused current feature map (except for the first frame feature map) contains features from both the previous frame feature map and the current feature map. Therefore, when the action recognition program performs action recognition on the fused current feature map, it can analyze the action changes between the current feature map and the previous frame feature map, thus enabling action recognition to be performed from the first frame feature map and sequentially up to the last frame feature map.
[0096] Multiple experiments have shown that the total action recognition time for sequentially recognizing multiple frames of feature maps is less than that for recognizing multiple frames of feature maps at once, thus enabling a faster action recognition process.
[0097] Understandably, each time an action is performed on the current feature map to be recognized, the action recognition result for each frame of the feature map to be recognized is output, thus yielding several action recognition results. Specifically, action recognition can be performed in the following ways:
[0098] Temporal classification models: Recurrent Neural Networks (RNNs) or their variants (such as Long Short-Term Memory Networks (LSTM) and Gated Recurrent Units (GRUs)) can be used to process the temporal information of feature maps. RNN models can accept time-series input and update the hidden state at each time step, then pass the final hidden state to a classifier for action recognition.
[0099] Convolutional Neural Networks (CNNs): Feature maps can typically be viewed as two-dimensional images, and two-dimensional CNN models can be used for action recognition. Convolution and pooling operations can be applied to the feature maps, followed by classification through fully connected layers.
[0100] Spatiotemporal 3D Convolutional Neural Networks (3D CNNs): 3D CNN models can be used to simultaneously consider temporal and spatial information. 3D CNNs apply convolutional operations in the temporal dimension to learn temporal features in video sequences and perform action classification.
[0101] Optical flow: Optical flow can capture motion information between consecutive frames. It can calculate the optical flow of pixels in a feature map and then use the changing patterns of optical flow to identify different actions.
[0102] Joint-coordinate-based methods: If the feature map is represented by the coordinates of joint points, joint-coordinate-based methods can be used for action recognition. These methods typically utilize the spatiotemporal patterns of joint motion for classification, such as using Hidden Markov Models (HMMs) or methods based on Dynamic Time Warping (DTW).
[0103] Step S40: Analyze the several action recognition results to obtain the final action recognition result.
[0104] Specifically, action recognition can also be completed by classifying several frames of feature maps to be recognized. Understandably, the action recognition result of each frame is a classification value. The final action recognition result can be obtained by analyzing each classification value.
[0105] Specifically, each category value can be analyzed through simple statistical analysis, such as calculating the frequency or proportion of each category. This can help understand the relative importance or distribution of each action category in the video.
[0106] Specifically, each category value can be analyzed by setting a threshold, and the judgment can be made based on the threshold set for the category value. Different thresholds can be set according to actual needs, such as confidence thresholds or occurrence frequency thresholds. Category values exceeding the threshold are considered valid action categories and can be used for the final action recognition result.
[0107] In addition to conventional classification result analysis, anomaly detection methods can also be applied. For example, they can be used to detect the existence of unknown actions or anomalous action categories that may not be accurately identified by the classification model.
[0108] Understandably, the final action recognition result can be either action recognition passed or action recognition failed. Each action recognition result can be a classification value, which can correspond to whether the action recognition passed or failed.
[0109] The action recognition method proposed in this application involves: acquiring a video to be recognized; extracting features from the video to obtain several frames of feature maps to be recognized; sequentially performing feature transfer and action recognition on the several frames of feature maps according to preset feature transfer and recognition rules to obtain several action recognition results; and analyzing the several action recognition results to obtain the final action recognition result. Extracting several frames of feature maps from the video to be recognized, and then sequentially transferring some features from each frame of feature maps to the next frame of feature maps according to preset feature transfer and recognition rules before performing action recognition, allows for the identification of feature differences between the previous and current frames when recognizing the current feature map, thereby determining the difference in user actions between the current and previous frame of feature maps, and thus obtaining the action recognition result for each frame of feature maps. Analyzing the several action recognition results to obtain the final action recognition result, it is understood that the above process performs action recognition frame by frame on the feature maps to be recognized, which reduces the processing pressure on the action recognition system and improves the speed of action recognition.
[0110] Reference Figure 3 , Figure 3 This is a flowchart illustrating a second exemplary embodiment of the action recognition method of this application.
[0111] Based on the first embodiment, a second embodiment of this application is proposed. The difference between the second embodiment and the first embodiment is as follows:
[0112] Step S30, which involves sequentially performing feature transfer and action recognition on the several frames of feature maps to be recognized according to preset feature transfer and recognition rules to obtain several action recognition results, is further refined. This step may include:
[0113] In this embodiment, step S30, which involves sequentially performing feature transfer and action recognition on the several frames of feature maps to be recognized according to preset feature transfer and recognition rules to obtain several action recognition results, includes:
[0114] Step S301: Determine whether the current feature map to be identified is the first frame feature map to be identified;
[0115] Step S31: If the current feature map to be identified is the first frame feature map to be identified, then perform action recognition on the current feature map to be identified;
[0116] Step S32: Extract features from the current feature map to be identified according to the preset learning ratio parameter to obtain the transferred features of the current feature map to be identified.
[0117] To more accurately determine the amount of features to be transmitted to the feature map to be identified in the next frame, a pre-set learning ratio parameter can be used. Understandably, the learning ratio parameter is a number less than 1 and greater than 0.
[0118] This learning ratio parameter can be adjusted according to the specific action recognition task to achieve optimal performance. Different action recognition tasks may require different learning ratio parameters.
[0119] Understandably, the transmitted features are first cached in memory, then fused with the feature map to be recognized in the next frame before action recognition is performed. The number of transmitted features will affect resource consumption and the amount of transmitted features.
[0120] Understandably, the learning scale parameter can affect the amount of features passed. A larger learning scale parameter may result in more features being passed to the next frame, thus preserving more detail and motion information. This may be more suitable for tasks that require capturing rapid changes and subtle movements.
[0121] A smaller learning scale parameter may limit feature propagation, ensuring that only key features are passed to the next frame. This can be used to filter noise or ignore unimportant details, making it suitable for action recognition tasks that require greater stability and robustness.
[0122] Therefore, the optimal learning scaling parameter can be determined through cross-validation or experiments based on the validation set. Different parameter values can be tried, their performance on the validation set evaluated, and the parameter that performs best can be selected as the final learning scaling parameter.
[0123] Understandably, different learning parameter ratios are suitable for different action recognition methods. The learning ratio parameters need to be adjusted based on experience or experiments according to the specific task requirements in order to obtain the best action recognition results.
[0124] If the current feature map to be identified is the first frame of feature maps to be identified, then action recognition is performed on the current feature map. Understandably, since the first frame of feature maps does not contain the previous frame's feature map, a feature fusion process is not required. Action recognition is then performed on the first frame of feature maps to be identified, yielding the action recognition result for the first frame of feature maps.
[0125] Step S33: If the current feature map to be identified is not the first frame feature map to be identified, then according to the learning ratio parameter, extract features from the current feature map to be identified according to the ratio, and obtain the current feature map after feature extraction and the transferred features of the current feature map to be identified.
[0126] Step S34: The current feature map after feature extraction and the transferred features extracted from the feature map to be identified in the previous frame are fused to obtain the fused feature map to be identified.
[0127] Step S35: Perform action recognition on the fused feature map to be identified.
[0128] Then, based on the preset learning ratio parameter, feature extraction is performed on the first frame of the feature map to be identified. The learning ratio parameter can be 1 / 10, in which case 1 / 10 of the total number of features of the first frame of the feature map to be identified is extracted as the transferred features of the first frame, and then the transferred features of the first frame are cached in memory.
[0129] When recognizing the second frame feature map, firstly, 1 / 10 of the total features of the second frame feature map is extracted as the transferred features of the second frame according to the learning ratio parameter. The remaining 9 / 10 of the second frame feature map is the current feature map after feature extraction. Then, the remaining 9 / 10 of the second frame feature map is fused with the transferred features of the first frame feature map, which is 1 / 10 of the total features of the first frame feature map, to obtain the fused second frame feature map. Then, the fused second frame feature map is recognized.
[0130] Understandably, the fused second frame of the feature map to be identified contains 1 / 10 of the total number of features in the first frame of the feature map to be identified and 9 / 10 of the total number of features in the second frame of the feature map to be identified. That is, it contains the features of the first and second frames. The action recognition program can analyze the fused second frame of the feature map to be identified to obtain the action difference between the second frame of the feature map to be identified and the first frame of the feature map to be identified, thereby completing the action recognition of the second frame of the feature map to be identified.
[0131] Understandably, the action recognition process for the third and fourth frame feature maps can be deduced similarly.
[0132] Furthermore, in this embodiment, when performing action recognition on each frame of the feature map to be recognized, 11 input items and 11 output items are used. That is, 11 input items are input into the action recognition program, and then 11 output items are input. Among them, input item 1 for action recognition of the first frame of the feature map to be recognized is the first frame of the feature map to be recognized, and input items 2-11 and other input items are zero.
[0133] The input for action recognition, except for the first frame of the image to be recognized, is:
[0134] Input item 1: The current feature map to be identified, represented as an image tensor.
[0135] Input item 2-11: Tensor of the transferred features extracted from the previous frame of the current feature map to be identified.
[0136] The output is similar to the input, and they are:
[0137] Output item 1: Action recognition result of the current feature map to be recognized.
[0138] Outputs 2-11: The remaining 10 outputs are tensors used to cache the extracted features from the second feature map to be identified. Understandably, these outputs may contain feature representations and contextual information from previous frames to enhance the model's utilization of historical information and improve accuracy.
[0139] In addition, the input image dimensions are 1×3×128×128, and other input dimensions are inferred from the designed network.
[0140] Understandably, the feature map to be identified can be divided into channels, and a certain number of channels can be extracted according to the learning ratio parameter, so that the features expressed by the channels in the previous frame are fused with the features expressed by the current number of channels, thereby increasing the information exchange between channels.
[0141] Understandably, the fused feature map can better reflect the feature differences between two consecutive frames of feature maps, enabling action recognition to be started in a single frame, thus speeding up the overall action recognition process.
[0142] The action recognition method proposed in this application involves: if the current feature map to be recognized is the first frame feature map to be recognized, then performing action recognition on the current feature map to be recognized; extracting features from the current feature map to be recognized proportionally according to a pre-set learning ratio parameter to obtain the transferred features of the current feature map to be recognized; if the current feature map to be recognized is not the first frame feature map to be recognized, then extracting features from the current feature map to be recognized proportionally according to the learning ratio parameter to obtain the current feature map after feature extraction and the transferred features of the current feature map to be recognized; fusing the current feature map after feature extraction and the transferred features extracted from the previous frame feature map to obtain the fused feature map to be recognized; and performing action recognition on the fused feature map to be recognized. The fused feature map to be recognized better reflects the feature differences between two consecutive frames of feature maps to be recognized, enabling action recognition to be started in a single frame, thus accelerating the overall action recognition process.
[0143] Based on the first embodiment, a third embodiment of this application is proposed. The difference between the third embodiment and the first embodiment is as follows:
[0144] Step S30, the step of performing action recognition on the plurality of frames of feature maps to be identified, is further refined, wherein the refined steps may include:
[0145] In this embodiment, step S30, the step of performing action recognition on the plurality of frames of feature maps to be identified, includes:
[0146] Step S36: Perform action recognition on the several frames of feature maps to be recognized using a preset action recognition neural network model.
[0147] Action recognition can be performed on several frames of feature maps to be recognized using an action recognition neural network model.
[0148] Specifically, for each frame of feature map to be identified, preprocessing steps and a pre-defined neural network model can be used to extract features. This may involve resizing the feature map to the size of the model input, normalization, or other necessary preprocessing operations.
[0149] Furthermore, the extracted feature maps are organized into the input format required by the model. This typically involves stacking multiple feature maps together to form a multi-channel input to meet the model's input requirements.
[0150] Then, the prepared feature map is input into the preset action recognition neural network model. The model will perform forward propagation on the input and output the corresponding action category prediction result.
[0151] Then, based on the model's output, some post-processing steps can be used, such as using a softmax function to transform the predictions into a probability distribution and to find the most likely action category.
[0152] The action recognition method proposed in this application improves the action recognition effect by performing action recognition on the several frames of feature maps to be recognized through a preset action recognition neural network model.
[0153] Based on the third embodiment, a fourth embodiment of this application is proposed. The difference between the fourth embodiment and the third embodiment is as follows:
[0154] Before step S36, which involves performing action recognition on the several frames of feature maps to be recognized using a preset action recognition neural network model, supplementary steps may include:
[0155] In this embodiment, before step S36, which involves performing action recognition on the plurality of frames of feature maps to be recognized using a preset action recognition neural network model, the following steps are also included:
[0156] Step S00: Create the action recognition neural network model;
[0157] Step S01: Obtain the sample action video and the corresponding sample action information;
[0158] Step S02: Based on the focal loss function, train the sample action video and the sample action information to obtain the action recognition neural network model.
[0159] Weighting each class is done to address the problem of imbalanced samples. In practical action recognition tasks, different action categories may have different numbers of samples; some categories may have fewer samples, while others may have more.
[0160] If imbalanced samples are not considered and training is performed using a standard loss function, the model may be biased towards learning the class with a larger number of samples, while performing poorly on classes with fewer samples. This can lead to poor performance in identifying classes with fewer samples in tests or real-world applications.
[0161] To address this issue, a weighted loss function, such as Focal Loss, can be used. Focal Loss assigns higher weights to classes with fewer samples, causing the model to focus more on these classes during training, thus balancing the influence of different classes. By using a weighted loss function, the impact of imbalanced samples on training can be mitigated, improving the model's performance on classes with fewer samples.
[0162] By assigning weights to each class, the model focuses more on learning and recognizing classes with fewer samples, thereby improving the overall accuracy and balance of action recognition. This ensures that the model achieves good performance across all classes and addresses the problem of imbalanced samples.
[0163] Therefore, training the model on sample action videos and corresponding sample action information using the focal loss function can ensure that the model can achieve good performance across all categories.
[0164] The action recognition method proposed in this application, by creating the action recognition neural network model, includes the following steps: acquiring sample action videos and corresponding sample action information; training the sample action videos and sample action information based on the focal loss function to obtain the action recognition neural network model, which can ensure that the model can achieve good performance in various categories.
[0165] Based on the fourth embodiment, a fifth embodiment of this application is proposed. The difference between the fifth embodiment and the fourth embodiment is as follows:
[0166] Before step S02, which involves training the sample action video and the sample action information based on the focal loss function to obtain the action recognition neural network model, a refinement process is performed. This refinement process may include:
[0167] In this embodiment, step S02, training the sample action video and the sample action information based on the focal loss function to obtain the action recognition neural network model, includes:
[0168] Step S021: Based on preset image frame extraction rules, extract several frame sample feature maps from the sample action video.
[0169] Step S022: Classify the feature maps of the several frames of samples to obtain several sample classification values;
[0170] Step S023: Sum the classification values of the plurality of samples to obtain the total classification value of the samples;
[0171] Step S024: The total sample classification values and the sample action information are trained using the focal loss function to obtain the action recognition neural network model.
[0172] For each input frame, the neural network generates a corresponding output classification value. Instead of considering the classification result of each frame individually, this method sums the classification values of each frame. By summing the classification values of the feature maps of all frames, the classification result of the entire video sequence can be obtained.
[0173] Next, the classification results of the entire video sequence are used to calculate the loss function. The focal loss function measures the difference between the model's predictions and the actual labels. By summing the output values of all frames and calculating the overall loss, the information of the video sequence can be utilized more comprehensively, thereby improving the training accuracy.
[0174] The advantage of this method lies in its consideration of the contextual information of the entire video sequence, rather than focusing solely on the classification results of a single frame. By summing the classification values of each frame, the relationships between frames can be modeled, and action patterns and temporal features in the video sequence can be extracted. This allows for more accurate action recognition and classification tasks.
[0175] In summary, by summing the output classification values of each frame and calculating the overall loss, the contextual information of the video sequence can be utilized to improve training accuracy. This method can better capture action patterns and temporal features in video sequences, thereby improving the performance of action recognition tasks.
[0176] The action recognition method proposed in this application extracts several frames of sample feature maps from the sample action video based on preset image frame extraction rules; classifies the several frames of sample feature maps to obtain several sample classification values; sums the several sample classification values to obtain a total sample classification value; and trains the total sample classification value and the sample action information using a focal loss function to obtain the action recognition neural network model, which can perform action recognition and classification tasks more accurately.
[0177] Based on the fourth embodiment, a sixth embodiment of this application is proposed. The difference between the sixth embodiment and the fourth embodiment is as follows:
[0178] Step S02, which involves training the sample action video and the sample action information based on the focal loss function to obtain the action recognition neural network model, is further refined. The refined steps may include:
[0179] In this embodiment, step S02, training the sample action video and the sample action information based on the focal loss function to obtain the action recognition neural network model, includes:
[0180] Step S025: Obtain the parameters of the pre-trained model;
[0181] Step S027: Based on the pre-trained model parameters and the focal loss function, train the sample action video and the sample action information to obtain the action recognition neural network model.
[0182] Considering that pre-training can greatly reduce the workload of training the model, we can obtain the pre-trained model parameters through ImageNet. Based on the trained model parameters, we can train the sample action video and sample action information through the focal loss function to obtain the action recognition neural network model.
[0183] In the field of computer vision, ImageNet is a widely used large-scale image classification dataset that contains a large number of images of different categories. By pre-training on ImageNet, models can learn rich image features and general visual representations.
[0184] The use of pre-trained models helps accelerate and improve the training process for new tasks. Since pre-trained models on ImageNet have already learned some general features and patterns, they can serve as a starting point for new tasks, providing better initial parameter and weight initialization. This reduces the need for training on large datasets and allows for good performance even on relatively small datasets.
[0185] Therefore, if the network architecture and task of this invention are similar to the image classification task on ImageNet, then using parameters pre-trained on ImageNet is reasonable and can bring many benefits, including accelerating training, improving convergence, and providing better initial weights.
[0186] It is important to note that if the network architecture and task of this invention differ significantly from the tasks on ImageNet, fine-tuning or transfer learning may be necessary to adapt to the requirements of the new task. Fine-tuning can be performed by further training on top of the pre-trained model to better adapt the model to the specific features and data distribution of the new task.
[0187] The action recognition method proposed in this application obtains pre-trained model parameters; based on the pre-trained model parameters and the focal loss function, it trains the sample action video and the sample action information to obtain the action recognition neural network model. This can reduce the training requirements for large-scale datasets and still achieve good performance on relatively small datasets.
[0188] Based on the sixth embodiment, a seventh embodiment of this application is proposed. The difference between the seventh embodiment and the sixth embodiment is as follows:
[0189] Before step S027, which involves training the sample action video and the sample action information based on the pre-trained model parameters and the focal loss function to obtain the action recognition neural network model, the supplementary steps may include:
[0190] In this embodiment, before step S027, which involves training the sample action video and the sample action information based on the pre-trained model parameters and the focal loss function to obtain the action recognition neural network model, the following steps are included:
[0191] Step S26: Freeze the batch normalization layer in the neural network.
[0192] Batch Normalization is a commonly used technique in deep learning models to normalize and standardize inputs, thereby accelerating training and improving model accuracy. It normalizes the inputs for each mini-batch, ensuring that the input features have zero mean and unit variance.
[0193] Consider a neural network action recognition model based on a neural network. Since the number of frames required for action recognition is relatively small, the batch normalization layer in the neural network action may not be necessary.
[0194] Calculating the mean and variance of a small batch size may not be stable or meaningful. In such cases, the Batch Normalization layer can be frozen. This can reduce resource overhead and improve model training speed.
[0195] The table below lists the relevant metrics for action recognition models built using different methods.
[0196]
[0197] The action recognition method proposed in this application reduces resource overhead and improves model training speed by freezing the batch normalization layer in the neural network.
[0198] Furthermore, this application also proposes an action recognition device, which includes:
[0199] The video acquisition module is used to acquire the video to be recognized;
[0200] The feature extraction module is used to extract features from the video to be identified, and obtain several frames of feature maps to be identified;
[0201] The behavior recognition module is used to sequentially perform feature transfer and action recognition on the several frames of feature maps to be recognized according to preset feature transfer and recognition rules, and obtain several action recognition results.
[0202] The comprehensive judgment module is used to analyze the several action recognition results to obtain the final action recognition result.
[0203] The principle and implementation process of action recognition in this embodiment are explained in the above embodiments and will not be repeated here.
[0204] Furthermore, this application also proposes a terminal device, which includes a memory, a processor, and an action recognition program stored in the memory and executable on the processor. When the action recognition program is executed by the processor, it implements the steps of the action recognition method described above.
[0205] Since this action recognition program employs all the technical solutions of all the aforementioned embodiments when executed by the processor, it has at least all the beneficial effects brought about by all the technical solutions of all the aforementioned embodiments, which will not be elaborated here.
[0206] Furthermore, embodiments of this application also propose a computer-readable storage medium storing an action recognition program, which, when executed by a processor, implements the steps of the action recognition method described above.
[0207] Since this action recognition program employs all the technical solutions of all the aforementioned embodiments when executed by the processor, it has at least all the beneficial effects brought about by all the technical solutions of all the aforementioned embodiments, which will not be elaborated here.
[0208] Compared to existing technologies, the action recognition method, apparatus, terminal device, and storage medium proposed in this application acquire a video to be recognized; extract features from the video to obtain several frames of feature maps to be recognized; sequentially perform feature transfer and action recognition on the several frames of feature maps to be recognized according to preset feature transfer and recognition rules, obtaining several action recognition results; and analyze the several action recognition results to obtain the final action recognition result. Extracting several frames of feature maps to be recognized from the video to be recognized, and then sequentially transferring some features from each frame of feature maps to the next frame of feature maps to be recognized according to preset feature transfer and recognition rules before performing action recognition, allows for the determination of feature differences between the previous and current frames when recognizing the current feature map, thereby determining the difference in user actions between the current and previous frames of feature maps to be recognized, thus obtaining the action recognition result for each frame of feature maps to be recognized. Then, analyzing the several action recognition results to obtain the final action recognition result, it can be understood that the above process performs action recognition frame by frame on the feature maps to be recognized, which reduces the processing pressure on the action recognition system and improves the speed of action recognition.
[0209] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0210] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0211] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, controlled terminal, or network device, etc.) to execute the methods of each embodiment of this application.
[0212] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. An action recognition method, characterized in that, The action recognition method includes: Obtain the video to be recognized; Feature extraction is performed on the video to be identified to obtain several frames of feature maps to be identified; According to preset feature transfer and recognition rules, feature transfer and action recognition are performed sequentially on the several frames of feature maps to be recognized, resulting in several action recognition results, including: Determine whether the current feature map to be identified is the first frame feature map to be identified; if the current feature map to be identified is the first frame feature map to be identified, then perform action recognition on the current feature map to be identified; according to a preset learning ratio parameter, extract a preset portion of the total feature amount from the current feature map to be identified as the transferred feature of the current feature map to be identified, wherein the ratio is the learning ratio parameter, and the optimal learning ratio parameter is determined through cross-validation or experiments based on the validation set; if the current feature map to be identified is not the first frame feature map to be identified, then according to the learning ratio parameter, perform action recognition on the current feature map to be identified. After extracting features from the feature map proportionally, the process yields the current feature map with extracted features and the transferred features of the current feature map to be identified. This includes: extracting a preset portion of the total feature amount from the current feature map to be identified proportionally, based on the learning ratio parameter, as the transferred features of the current feature map to be identified; using the remaining portion of the current feature map to be identified as the current feature map with extracted features; fusing the current feature map with extracted features and the transferred features extracted from the previous frame's feature map to be identified, to obtain a fused feature map to be identified; and performing action recognition on the fused feature map to be identified. The results of the action recognition are analyzed to obtain the final action recognition result.
2. The action recognition method according to claim 1, characterized in that, The step of performing action recognition on the plurality of frame feature maps to be identified includes: The preset action recognition neural network model is used to perform action recognition on the several frames of feature maps to be recognized.
3. The action recognition method according to claim 2, characterized in that, Before the step of performing action recognition on the several frames of feature maps to be recognized using a preset action recognition neural network model, the method further includes: The steps for creating the action recognition neural network model include: Acquire sample action videos and corresponding sample action information; The action recognition neural network model is obtained by training the sample action video and the sample action information based on the focal loss function.
4. The action recognition method according to claim 3, characterized in that, The step of training the action recognition neural network model based on the focal loss function to obtain the sample action video and the sample action information includes: Based on preset image frame extraction rules, several frame sample feature maps are extracted from the sample action video. The feature maps of the several frames are classified to obtain several sample classification values; The total sample classification value is obtained by summing the classification values of the given samples. The action recognition neural network model is obtained by training the total sample classification values and the sample action information using the focal loss function.
5. The action recognition method according to claim 3, characterized in that, Before the step of training the action recognition neural network model based on the focal loss function on the sample action video and the sample action information, the method further includes: Obtain the parameters of the pre-trained model; Based on the pre-trained model parameters and the focal loss function, the sample action video and the sample action information are trained to obtain the action recognition neural network model.
6. The action recognition method according to claim 5, characterized in that, Before the step of training the sample action video and the sample action information based on the pre-trained model parameters and the focal loss function to obtain the action recognition neural network model, the method further includes: Freeze the batch normalization layer in the neural network.
7. A motion recognition device, characterized in that, The motion recognition device includes: The video acquisition module is used to acquire the video to be recognized. The feature extraction module is used to extract features from the video to be identified, and obtain several frames of feature maps to be identified; The behavior recognition module is used to sequentially perform feature transfer and action recognition on several frames of feature maps to be recognized according to preset feature transfer and recognition rules, and obtain several action recognition results, including: determining whether the current feature map to be recognized is the first frame of feature maps to be recognized; if the current feature map to be recognized is the first frame of feature maps to be recognized, then performing action recognition on the current feature map to be recognized; according to a preset learning ratio parameter, extracting a preset portion of the total feature amount from the current feature map to be recognized as the transferred feature of the current feature map, wherein the ratio is the learning ratio parameter, and the optimal learning ratio parameter is determined through cross-validation or experiments based on a validation set; if the current feature map to be recognized... If the image is not the first frame of the feature map to be identified, then according to the learning ratio parameter, features are extracted from the current feature map to be identified proportionally to obtain the current feature map after feature extraction and the transferred features of the current feature map to be identified. This includes: extracting a preset portion of the total number of features from the current feature map to be identified proportionally according to the learning ratio parameter as the transferred features of the current feature map to be identified, where the ratio is the learning ratio parameter; using the remaining portion of the current feature map to be identified as the current feature map after feature extraction; fusing the current feature map after feature extraction and the transferred features extracted from the previous frame of the feature map to be identified to obtain a fused feature map to be identified; and performing action recognition on the fused feature map to be identified. The comprehensive judgment module is used to analyze the several action recognition results to obtain the final action recognition result.
8. A terminal device, characterized in that, The terminal device includes a memory, a processor, and an action recognition program stored in the memory and executable on the processor. When the action recognition program is executed by the processor, it implements the steps of the action recognition method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an action recognition program, which, when executed by a processor, implements the steps of the action recognition method as described in any one of claims 1-6.
Citation Information
Patent Citations
Gesture recognition method based on deep residual network
CN111444764A
Facial recognition method and device, computer equipment and storage medium
CN114333021A
Behavior posture recognition method for shielded part of human body
CN114582013A
Badminton action recognition method and device, electronic equipment and storage medium
CN115205961A