Action recognition method, model training method, device and electronic device
By using the self-attention model to calculate the similarity probability between video and action categories, the problem of CNN action recognition method consumes a lot of computing resources in long videos is solved, and more efficient and accurate action recognition is achieved.
Patent Information
- Application Number
- CN202210072157.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-21
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-01-21
AI Technical Summary
The existing CNN-based action recognition method requires a large number of convolution operations when the video is long, resulting in large consumption of computing resources and affecting device performance.
The self-attention model is used to calculate the similarity probability distribution between the video to be identified and multiple action categories, and the target action category of the video to be identified is directly determined, avoiding convolutional operations.
It saves computing resources, improves device performance, and improves the accuracy of action recognition through learning of image feature sequences in time and space dimensions.
Smart Images

Figure CN114429675B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method for action recognition, a method for model training, a device, and an electronic device. Background Art
[0002] In scenarios such as human-computer interaction, video understanding, and security, action recognition methods based on convolutional neural networks (CNNs) are usually used to recognize various behavioral actions in videos. Specifically, an electronic device uses a CNN to detect images in a video to obtain a detection result of human target points in the image and a preliminary action recognition result, and trains an action recognition neural network according to the detection result of human target points and the action recognition result. Further, the electronic device recognizes the behavioral actions in the above images according to the trained action recognition neural network.
[0003] However, in the detection process of the above action recognition method, a large number of convolutional operations need to be performed based on the CNN. Especially when the above video is long, the CNN convolutional operation requires a large amount of computing resources, which affects the device performance. Summary of the Invention
[0004] On the one hand, there is provided a method for action recognition, including: obtaining a plurality of image frames of a video to be recognized; determining a probability distribution of the video to be recognized being similar to multiple action categories according to the plurality of image frames and a pre-trained self-attention model; the self-attention model is used to calculate the similarity of the probability of the image feature sequence being similar to multiple action categories through a self-attention mechanism; the image feature sequence is obtained based on the plurality of image frames in the time dimension or the space dimension; the probability distribution includes the probability of the video to be recognized being similar to each action category among the multiple action categories; determining a target action category corresponding to the video to be recognized based on the probability distribution of the video to be recognized being similar to multiple action categories; the probability of the video to be recognized being similar to the target action category is greater than or equal to a preset threshold.
[0005] The technical solution provided by the embodiments of the present disclosure calculates the probability distribution of the video to be recognized being similar to multiple action categories based on the self-attention model, and can directly determine the target action category where the video to be recognized is located from multiple action categories. Compared with the prior art, there is no need to set up a CNN, avoiding a large amount of calculations brought by using convolutional operations, and finally saving the computing resources of the device.
[0006] Meanwhile, since the image feature sequence is obtained based on multiple image frames in the time dimension or the space dimension, the image feature sequence can represent the time sequence of multiple image frames or the time sequence and space sequence of multiple image frames, and can learn the similarity between the video to be recognized and each action category from the time dimension and space dimension of multiple image frames to a certain extent, which can make the subsequent obtained probability distribution more accurate.
[0007] In some embodiments, the self-attention model includes a self-attention encoding layer and a classification layer. The self-attention encoding layer is used to calculate the similarity features of a sequence composed of multiple image features with respect to multiple action categories of different action categories, and the classification layer is used to calculate the probability distribution corresponding to the similarity features. Determining the probability distribution corresponding to the similarity between the video to be recognized and multiple action categories of different action categories according to multiple image frames and the pre-trained self-attention model includes: determining the target similarity features of the video to be recognized with respect to multiple action categories of different action categories according to multiple image frames and the self-attention encoding layer; the target similarity features are used to represent the similarity between the video to be recognized and each action category of different action categories; inputting the target similarity features into the classification layer to obtain the probability distribution corresponding to the similarity between the video to be recognized and multiple action categories of different action categories.
[0008] The above technical solution provided by the embodiments of the present disclosure can, by using the preset self-attention encoding layer, determine the similarity features between multiple image frames and multiple action categories based on the self-attention mechanism, and perform classification processing on the similarity features based on the classification layer to obtain the probability distribution of the video to be recognized belonging to multiple action categories, and can provide an implementation method for determining the probability distribution of the video to be recognized belonging to multiple action categories without using CNN, saving the computing resources consumed by the convolution operation.
[0009] In some embodiments, after determining the target similarity features of the video to be recognized with respect to multiple action categories according to multiple image frames and the self-attention encoding layer, the following is included: segmenting each image frame of the multiple image frames to obtain multiple subsampled images; in this case, determining the target similarity features of the video to be recognized with respect to multiple action categories according to multiple image frames and the self-attention encoding layer includes: determining the sequence features of the video to be recognized according to the multiple subsampled images and the self-attention encoding layer, and determining the target similarity features according to the sequence features of the video to be recognized; the sequence features include time sequence features, or time sequence features and space sequence features; the time sequence features are used to represent the similarity between the video to be recognized and multiple action categories in the time dimension, and the space sequence features are used to represent the similarity between the video to be recognized and multiple action categories in the space dimension.
[0010] The above technical solution provided by the embodiments of the present disclosure divides each image frame into multiple sub-sampled images of a preset size, and determines the time series features from the time dimension and the spatial series features from the spatial dimension based on the multiple sub-sampled images. The determined target similarity features can reflect the time features and spatial features of the video to be recognized, which can make the determined target action category more accurate in the subsequent process.
[0011] In some embodiments, the above determination of the time series features of the video to be recognized includes: determining at least one time sampling sequence from the multiple sub-sampled images; each time sampling sequence includes the sub-sampled images located at the same position in each image frame; determining the sub-time series features of each time sampling sequence according to each time sampling sequence and the self-attention encoding layer; the sub-time series features are used to characterize the similarity between each time sampling sequence and multiple action categories; determining the time series features of the video to be recognized according to the sub-time series features of at least one time sampling sequence.
[0012] The above technical solution provided by the embodiments of the present disclosure at least brings the following: dividing the multiple sub-sampled images into at least one time sampling sequence, determining the sub-time series features of each time sampling sequence, and determining the time series features of the video to be recognized according to the multiple sub-time series features. Since the positions of the sub-sampled images in each time sampling sequence are the same in different image frames, the determined time series features are more comprehensive and accurate based on this.
[0013] In some embodiments, the above determination of the sub-time series features of each time sampling sequence according to each time sampling sequence and the self-attention encoding layer includes: determining multiple first image input features and category input features; each first image input feature is obtained by performing position encoding and merging on the image features of the sub-sampled images included in the first time sampling sequence, and the first time sampling sequence is any one of the at least one time sampling sequence; the category input feature is obtained by performing position encoding and merging on the category feature, and the category feature is used to characterize multiple action categories; inputting the multiple first image input features and the category input feature into the self-attention encoding layer, and determining the output feature corresponding to the category input feature output by the self-attention encoding layer as the sub-time series feature of the first time sampling sequence.
[0014] The above technical solution provided by the embodiments of the present disclosure can determine the sub-time series features of each time sampling sequence with respect to multiple action categories by using the self-attention encoding layer, and can save the corresponding computing resources without using convolution operations compared with the prior art.
[0015] In some embodiments, determining the spatial sequence features of the video to be recognized includes: determining at least one spatial sampling sequence from multiple subsampled images; each spatial sampling sequence includes subsampled images in an image frame; determining the subspace sequence features of each spatial sampling sequence according to each spatial sampling sequence and the self-attention encoding layer; the subspace sequence features are used to characterize the similarity between each spatial sampling sequence and multiple action categories; determining the spatial sequence features of the video to be recognized according to the subspace sequence features of at least one spatial sampling sequence.
[0016] In the above technical solution provided by the embodiments of the present disclosure, in the process of determining each spatial sampling sequence, a preset number of target subsampled images at preset positions can be used to generate at least one spatial sequence feature. In this way, without affecting the spatial sequence features, the number of subsampled images in each spatial sampling sequence can be reduced, and the computational consumption of the subsequent self-attention encoding layer can be reduced.
[0017] In some embodiments, determining at least one spatial sampling sequence from multiple subsampled images includes: for the first image frame, determining a preset number of target subsampled images located at preset positions from the subsampled images included in the first image frame, and determining the target subsampled images as the spatial sampling sequence corresponding to the first image frame; the first image frame is any one of the multiple image frames.
[0018] In the above technical solution provided by the embodiments of the present disclosure, multiple subsampled images are divided into at least one spatial sampling sequence, the subspace sequence features of each spatial sampling sequence are determined, and the spatial sequence features of the video to be recognized are determined according to the multiple subspace sequence features. In this way, the spatial sequence features determined based on this are more comprehensive and accurate.
[0019] In some embodiments, determining the subspace sequence features of each spatial sampling sequence according to each spatial sampling sequence and the self-attention encoding layer includes: determining multiple second image input features and class input features; each second image input feature is obtained by performing position encoding and merging on the image features of the subsampled images included in the first spatial sampling sequence, and the first spatial sampling sequence is any one of the at least one spatial sampling sequence; the class input feature is obtained by performing position encoding and merging on the class feature, and the class feature is used to characterize multiple action categories; inputting the multiple second image input features and the class input features into the self-attention encoding layer, and determining the output feature corresponding to the class input feature output by the self-attention encoding layer as the subspace sequence feature of the first spatial sampling sequence.
[0020] The above technical solution provided by the embodiments of the present disclosure can, by using the self-attention encoding layer, determine the subspace sequence features of each spatial sampling sequence with respect to multiple action categories, and can avoid consuming computing resources caused by using convolution operations.
[0021] In some embodiments, the above multiple image frames are obtained based on image preprocessing, and the image preprocessing includes at least one operation of cropping, image enhancement, and scaling.
[0022] On the other hand, a model training method is provided, including: obtaining multiple sample image frames of a sample video and the sample action category to which the sample video belongs; performing self-attention training based on the multiple sample image frames and the sample action category to obtain a trained self-attention model; the self-attention model is used to calculate the similarity of the probability that the sample image feature sequence is similar to multiple action categories; the sample image feature sequence is obtained based on the multiple sample image frames in the time dimension or the spatial dimension.
[0023] The above technical solution provided by the embodiments of the present disclosure performs self-attention training on the initial self-attention model based on the multiple sample image frames of the sample video and the sample action category to which the sample video belongs, and trains to obtain a self-attention model. Since only the sample similarity features of the multiple sample image frames similar to different sample categories need to be determined based on the self-attention mechanism during the training process, compared with the prior art, there is no need to perform convolution operations based on CNN, avoiding a large amount of calculations caused by using convolution operations, and finally saving the computing resources of the device.
[0024] On the other hand, an action recognition device is provided, including an acquisition unit and a determination unit; the acquisition unit is used to acquire multiple image frames of a video to be recognized; the determination unit is used to, after the acquisition unit acquires the multiple image frames, determine the probability distribution of the similarity between the video to be recognized and multiple action categories according to the multiple image frames and the pre-trained self-attention model; the self-attention model is used to calculate the similarity of the image feature sequence similar to multiple action categories through the self-attention mechanism; the image feature sequence is obtained based on the multiple image frames in the time dimension or the spatial dimension; the probability distribution includes the probability that the video to be recognized is similar to each action category among the multiple action categories; the determination unit is further used to determine the target action category corresponding to the video to be recognized based on the probability distribution of the similarity between the video to be recognized and different action categories; the probability that the video to be recognized is similar to the target action category is greater than or equal to a preset threshold.
[0025] In some embodiments, the self-attention model described above includes a self-attention encoding layer and a classification layer. The self-attention encoding layer is used to calculate the similarity features of the image feature sequence with respect to multiple action categories, and the classification layer is used to calculate the probabilities corresponding to the similarity features. The determining unit is specifically configured to: determine the target similarity features of the video to be recognized with respect to multiple action categories according to multiple image frames and the self-attention encoding layer; the target similarity features are used to characterize the similarity between the video to be recognized and multiple action categories; input the target similarity features into the classification layer to obtain the probabilities of similarity between the video to be recognized and multiple action categories.
[0026] In some embodiments, the above-mentioned action recognition device further includes a processing unit; the processing unit is configured to segment each of the multiple image frames to obtain multiple subsampled images before the determining unit determines the target similarity features of the video to be recognized with respect to multiple action categories according to the multiple image frames and the self-attention encoding layer; the determining unit is specifically configured to: determine the sequence features of the video to be recognized according to the multiple subsampled images and the self-attention encoding layer, and determine the target similarity features according to the sequence features of the video to be recognized; the sequence features include temporal sequence features, or temporal sequence features and spatial sequence features; the temporal sequence features are used to characterize the similarity between the video to be recognized and multiple action categories in the temporal dimension, and the spatial sequence features are used to characterize the similarity between the video to be recognized and multiple action categories in the spatial dimension.
[0027] In some embodiments, the determining unit is specifically configured to: determine at least one temporal sampling sequence from the multiple subsampled images; each temporal sampling sequence includes the subsampled images at the same position in each image frame; determine the sub-temporal sequence features of each temporal sampling sequence according to each temporal sampling sequence and the self-attention encoding layer; the sub-temporal sequence features are used to characterize the similarity between each temporal sampling sequence and multiple action categories; determine the temporal sequence features of the video to be recognized according to the sub-temporal sequence features of at least one temporal sampling sequence.
[0028] In some embodiments, the determining unit is specifically configured to: determine multiple first image input features and a category input feature; each first image input feature is obtained by performing position encoding merging on the image features of the subsampled images included in the first temporal sampling sequence, and the first temporal sampling sequence is any one of the at least one temporal sampling sequence; the category input feature is obtained by performing position encoding merging on the category features, and the category features are used to characterize multiple action categories; input the multiple first image input features and the category input feature into the self-attention encoding layer, and determine the output feature corresponding to the category input feature output by the self-attention encoding layer as the sub-temporal sequence features of the first temporal sampling sequence.
[0029] In some embodiments, the above-mentioned determination unit is specifically configured to: determine at least one spatial sampling sequence from multiple subsampled images; each spatial sampling sequence includes subsampled images in an image frame; determine subspace sequence features of each spatial sampling sequence according to each spatial sampling sequence and the self-attention encoding layer; the subspace sequence features are used to characterize the similarity between each spatial sampling sequence and multiple action categories; determine the spatial sequence features of the video to be recognized according to the subspace sequence features of at least one spatial sampling sequence.
[0030] In some embodiments, the above-mentioned determination unit is specifically configured to: for the first image frame, determine a preset number of target subsampled images located at preset positions from the subsampled images included in the first image frame, and determine the target subsampled images as the spatial sampling sequence corresponding to the first image frame; the first image frame is any one of multiple image frames.
[0031] In some embodiments, the above-mentioned determination unit is specifically configured to: determine multiple second image input features and category input features; each second image input feature is obtained by performing position encoding and merging on the image features of the subsampled images included in the first spatial sampling sequence, and the first spatial sampling sequence is any one of at least one spatial sampling sequence; the category input feature is obtained by performing position encoding and merging on the category feature, and the category feature is used to characterize multiple action categories; input the multiple second image input features and the category input features into the self-attention encoding layer, and determine the output feature corresponding to the category input feature output by the self-attention encoding layer as the subspace sequence feature of the first spatial sampling sequence.
[0032] In some embodiments, the above-mentioned multiple image frames are obtained based on image preprocessing, and the image preprocessing includes at least one operation of cropping, image enhancement, and scaling.
[0033] On the other hand, a model training device is provided, including an acquisition unit and a training unit; the acquisition unit is configured to acquire multiple sample image frames of a sample video and the sample action category where the sample video is located; the training unit is configured to, after the acquisition unit acquires multiple sample image frames and the sample action category, perform self-attention training according to the multiple sample image frames and the sample action category to obtain a trained self-attention model; the self-attention model is used to calculate the similarity between the sample image feature sequence and multiple action categories; the sample image feature sequence is obtained based on the multiple sample image frames in the time dimension or the spatial dimension.
[0034] In another aspect, there is provided an electronic device, including: a processor and a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the action recognition method provided in the first aspect and any possible design manner thereof, or the model training method provided in the second aspect and any possible design manner thereof.
[0035] In another aspect, there is provided a computer-readable storage medium storing computer program instructions, which, when running on a computer (such as an electronic device, an action recognition device or a model training device), cause the computer to execute the action recognition method or the model training method of any of the above embodiments.
[0036] In another aspect, there is provided a computer program product. The computer program product includes computer program instructions, which, when executed on a computer (such as an electronic device, an action recognition device or a model training device), cause the computer to execute the action recognition method or the model training method of any of the above embodiments.
[0037] In another aspect, there is provided a computer program. When the computer program is executed on a computer (such as an electronic device, an action recognition device or a model training device), the computer program causes the computer to execute the action recognition method or the model training method of any of the above embodiments. Description of the Drawings
[0038] To more clearly illustrate the technical solutions in the present disclosure, the following will briefly introduce the drawings required for use in some embodiments of the present disclosure. Obviously, the drawings in the following description are only the drawings of some embodiments of the present disclosure, and those of ordinary skill in the art can also obtain other drawings based on these drawings. In addition, the drawings in the following description can be regarded as schematic diagrams and do not limit the actual sizes of the products, the actual processes of the methods, the actual timings of the signals, etc. involved in the embodiments of the present disclosure.
[0039] Figure 1 is a structural diagram of an action recognition system shown according to some embodiments;
[0040] Figure 2 is one of the flowcharts of an action recognition method shown according to some embodiments;
[0041] Figure 3 is a flowchart of a custom sampling shown according to some embodiments;
[0042] Figure 4 is the second flowchart of an action recognition method shown according to some embodiments;
[0043] Figure 5 One of the timing diagrams of an action recognition method shown according to some embodiments;
[0044] Figure 6 The third flowchart of an action recognition method shown according to some embodiments;
[0045] Figure 7 A schematic diagram of an image segmentation process shown according to some embodiments;
[0046] Figure 8 The fourth flowchart of an action recognition method shown according to some embodiments;
[0047] Figure 9 A schematic diagram of a time sampling sequence shown according to some embodiments;
[0048] Figure 10 The fifth flowchart of an action recognition method shown according to some embodiments;
[0049] Figure 11 A timing diagram of determining time series features shown according to some embodiments;
[0050] Figure 12 The sixth flowchart of an action recognition method shown according to some embodiments;
[0051] Figure 13 A schematic diagram of a spatial sampling sequence shown according to some embodiments;
[0052] Figure 14 The seventh flowchart of an action recognition method shown according to some embodiments;
[0053] Figure 15 The second timing diagram of an action recognition method shown according to some embodiments;
[0054] Figure 16 The flowchart of a model training method shown according to some embodiments;
[0055] Figure 17 The structural diagram of an action recognition device shown according to some embodiments;
[0056] Figure 18 The structural diagram of a model training device shown according to some embodiments;
[0057] Figure 19 The structural diagram of an electronic device shown according to some embodiments. Detailed implementation manners
[0058] Next, in conjunction with the accompanying drawings, the technical solutions in some embodiments of the present disclosure will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the embodiments provided by the present disclosure, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present disclosure.
[0059] Unless the context requires otherwise, throughout the specification and claims, the term "comprise" and its other forms, such as the third-person singular form "comprises" and the present participle form "comprising", are interpreted as open and inclusive, that is, "including, but not limited to". In the description of the specification, the terms "one embodiment", "some embodiments", "exemplary embodiments", "example", "specific example", or "some examples", etc. are intended to indicate that the specific features, structures, materials, or characteristics related to the embodiment or example are included in at least one embodiment or example of the present disclosure. The schematic representations of the above terms do not necessarily refer to the same embodiment or example. In addition, the described specific features, structures, materials, or characteristics may be included in any one or more embodiments or examples in any appropriate manner.
[0060] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present disclosure, unless otherwise stated, the meaning of "a plurality" is two or more.
[0061] "At least one of A, B, and C" has the same meaning as "at least one of A, B, or C", and both include the following combinations of A, B, and C: only A, only B, only C, the combination of A and B, the combination of A and C, the combination of B and C, and the combination of A, B, and C.
[0062] "A and / or B" includes the following three combinations: only A, only B, and the combination of A and B.
[0063] The use of "suitable for" or "configured to" herein means open and inclusive language, which does not exclude a device that is suitable for or configured to perform additional tasks or steps.
[0064] In addition, the use of "based on" implies openness and inclusiveness because a process, step, calculation, or other action "based on" one or more conditions or values can, in practice, be based on additional conditions or values beyond those.
[0065] The following introduces the inventive principles of the action recognition method and the model training method provided by the embodiments of the present disclosure:
[0066] In related technologies, when an electronic device recognizes behavioral actions in a video, it usually pre-trains an action recognition model based on a CNN. During the training process of this action recognition model, the electronic device can perform frame extraction on the sample video to obtain multiple sample image frames, and train a preset convolutional neural network with the multiple sample image frames and the labels of the action categories where the sample videos are located to train this action recognition model. Subsequently, during the use process of this action recognition model, the electronic device performs frame extraction on the video to be recognized to obtain multiple image frames, and inputs the image features of the multiple image frames into the action recognition model. Correspondingly, the action recognition model outputs the action category to which the video to be recognized belongs.
[0067] Since in the training and use processes of the above-mentioned action recognition model, a large number of convolutional operations need to be performed based on the CNN to learn the features of the input image frames, it will consume a large amount of device computing resources.
[0068] In some embodiments adopting an action recognition model, in related technologies, the action recognition model is also combined with the optical flow method to analyze the action categories of the video. Among them, it is also necessary to load optical flow images in the CNN, and there will also be a large number of convolutional operations, resulting in the consumption of a large amount of computing resources.
[0069] The embodiments of the present disclosure consider that the convolutional operations of the CNN consume a large amount of computing. The self-attention model is used to calculate the similarity between the video to be recognized and multiple action categories, and based on the determined similarity, the probability that the video to be recognized is similar to multiple action categories is determined. Furthermore, the action category where the video to be recognized is located can also be determined. Since only the encoder in the self-attention model is required and no convolutional operations are needed, a large amount of computing resources can be saved.
[0070] An action recognition method provided by the embodiments of the present disclosure can be applied to an action recognition system. Figure 1 A schematic structural diagram of this action recognition system is shown. As Figure 1As shown, the action recognition system 10 is used to solve the problem in the related art that the calculation resources consumed for action recognition of videos are large. The action recognition system 10 includes an action recognition device 11 and an electronic device 12. The action recognition device 11 is connected to the electronic device 12. The action recognition device 11 and the electronic device 12 can be connected by a wired method or a wireless method, and the embodiments of the present disclosure do not limit this.
[0071] The action recognition device 11 can be used to interact with the electronic device 12. For example, the action recognition device 11 can obtain the video to be recognized and the sample video from the electronic device 12.
[0072] At the same time, the action recognition device 11 can also execute the model training method provided by the embodiments of the present disclosure. For example, the action recognition device 11 uses the sample video as a sample to train an action recognition model based on the self-attention mechanism, and obtains a trained self-attention model.
[0073] It should be noted that in some embodiments, when the action recognition device is used to train the self-attention model, the action recognition device can also be referred to as a model training device.
[0074] On the other hand, the action recognition device 11 can also execute the action recognition method provided by the embodiments of the present disclosure. For example, the action recognition device 11 can also process the video to be recognized or input the video to be recognized into the self-attention model for the video to be recognized, so as to determine the target action category corresponding to the video to be recognized.
[0075] It should be noted that the video to be recognized or the sample video involved in the embodiments of the present disclosure can be a video captured by a shooting device in the electronic device, or a video received by the electronic device from other similar devices. The multiple action categories involved in the present disclosure can specifically include categories such as falling, climbing, and chasing. The action recognition system involved in the present disclosure can specifically be applied to public monitoring places such as nursing homes, stations, hospitals, and shopping malls, and can also be used in scenarios such as smart homes, augmented reality (AR) / virtual reality technology (VR), and video analysis and understanding.
[0076] The action recognition device 11 and the electronic device 12 can be independent devices or integrated into the same device, and the present disclosure does not specifically limit this.
[0077] When the action recognition device 11 and the electronic device 12 are integrated into the same device, the communication method between the action recognition device 11 and the electronic device 12 is the communication between internal modules of this device. In this case, the communication process between the two is the same as "the communication process between the action recognition device 11 and the electronic device 12 when they are independent of each other".
[0078] In the following embodiments provided by the present disclosure, the present disclosure takes the action recognition device 11 and the electronic device 12 being independently arranged as an example for illustration.
[0079] In practical applications, the action recognition method provided by the embodiments of the present disclosure can be applied to the action recognition device or the electronic device. Below, in combination with the accompanying drawings, taking the action recognition method being applied to the electronic device as an example, the action recognition method provided by the embodiments of the present disclosure will be described.
[0080] As shown in Figure 2 the action recognition method provided by the embodiments of the present disclosure includes the following S201 - S203.
[0081] S201. The electronic device acquires multiple image frames of the video to be recognized.
[0082] As a possible implementation manner, the electronic device acquires the video to be recognized, decodes and extracts frames from the video to be recognized, and uses the multiple sampled frames obtained by the decoding and frame extraction processing as the multiple image frames.
[0083] As another possible implementation manner, after the electronic device acquires the video to be recognized, after decoding and extracting frames from the video to be recognized, it pre - processes the multiple sampled frames obtained by the frame extraction to obtain multiple image frames.
[0084] Among them, the image pre - processing includes at least one operation of cropping, image enhancement, and scaling.
[0085] As a third possible implementation manner, after the electronic device acquires the video to be recognized, it can decode the video to be recognized to obtain multiple decoded frames, and perform the above - mentioned pre - processing on the multiple decoded frames to obtain the pre - processed decoded frames. Further, the electronic device performs frame extraction and random sampling on the pre - processed decoded frames to obtain multiple image frames.
[0086] It should be noted that the above - mentioned process of frame extraction and random sampling can use the method of adding random noise and blur processing to expand the samples of the pre - processed decoded frames. Exemplarily, the above - mentioned random noise can be Gaussian noise.
[0087] During the above sampling process, custom sampling in the time dimension can also be adopted, or custom sampling in the space dimension can be adopted, or a combination of custom sampling in the time dimension and the space dimension can be adopted. Exemplarily, Figure 3 shows a sampling method based on custom sampling in the time dimension, such as Figure 3 shown. After the video to be recognized is decoded, multiple decoded frames are obtained. The electronic device can perform frame extraction sampling on the multiple decoded frames based on a custom sampling method in the time dimension to obtain multiple image frames.
[0088] It can be understood that sampling the image frames using multiple different sampling methods can extract as many feature information of the video to be recognized as possible to ensure the accuracy of determining the target action category of the video to be recognized subsequently.
[0089] In some embodiments, the electronic device may also be pre-set with a preset sampling rate. During the above random sampling process, sampling can be performed on the multiple decoded frames obtained by decoding or on the preprocessed decoded frames based on the preset sampling rate. For example, in the case of using the preset sampling rate, the number of multiple image frames can be 96 frames. In some embodiments, the preset sampling rate can be set to be greater than the sampling rate when using a CNN convolutional neural network.
[0090] Since the video to be processed may have image distortion and prominent edge parts, when the electronic device obtains multiple sampled frames and performs image preprocessing, each sampled frame can be cropped based on a preset cropping size.
[0091] It should be noted that the above cropping can be performed in a central cropping manner to crop off the severely distorted parts around the sampled frame. Exemplarily, if the size of the sampled frame before cropping is 640*480 and the preset cropping size is 256*256, then after cropping each sampled frame, the size of the obtained image frame is 256*256.
[0092] It can be understood that adopting the central cropping method can, to a certain extent, reduce the influence brought by image distortion, and at the same time, it can remove the invalid feature information around the sampled frame, enabling the subsequent self-attention model to converge more easily, recognize more accurately, and faster.
[0093] Since the shooting conditions of the video to be recognized are different, from different environments and different lighting conditions. Therefore, when the electronic device performs image preprocessing, image enhancement processing can be performed on the multiple sampled frames.
[0094] It should be noted that the above image enhancement operation includes brightness enhancement. When performing image enhancement processing, a pre-packaged image enhancement function can be called to process each sampled frame.
[0095] It is understandable that image enhancement processing can adjust features such as brightness, color, and contrast of each sampled frame, and can correspondingly improve the generalization ability of each sampled frame.
[0096] In some cases, since the self-attention model involved in the embodiments of the present disclosure is pre-trained, when it inputs the image features of the input image frame, it has certain constraints on the pixel size of the image frame. In this case, if the pixel sizes of the multiple sampled image frames are different from the pixel size of the image frame constrained by the self-attention model, the electronic device needs to scale the obtained sampled frames to adjust to the pixel size that the self-attention model can adapt to. For example, if the pixel size of the sample image frame used in the training process of the self-attention model is 256*256, then during action recognition, the pixel sizes of the multiple image frames obtained after scaling can be 256*256.
[0097] S202. The electronic device determines the probability distribution of the video to be recognized being similar to multiple action categories according to the multiple image frames and the pre-trained self-attention model.
[0098] Among them, the self-attention model is used to calculate the similarity between the image feature sequence and multiple action categories through the self-attention mechanism. The image feature sequence is obtained based on the multiple image frames in the time dimension or the space dimension. The probability distribution includes the probabilities of the video to be recognized being similar to each action category among the multiple action categories.
[0099] It should be noted that the self-attention model includes a self-attention encoding layer and a classification layer. The self-attention encoding layer is used to calculate the similarity of the input feature sequence based on the self-attention mechanism to calculate the similarity features of each feature in the feature sequence relative to other features respectively. The classification layer is used to calculate the probabilities of the similarities of the similarity features of each input feature relative to other features to output the probability distribution of the similarity of each feature to other features.
[0100] As a possible implementation, the electronic device converts the multiple image frames into multiple image features respectively, and determines the sequence features of the video to be recognized based on the multiple converted image features and the self-attention encoding layer. Further, the electronic device inputs the sequence features of the video to be recognized into the classification layer, and then determines the multiple probability distributions output by the classification layer as the probability distribution of the video to be recognized being similar to multiple action categories.
[0101] In this case, the image feature sequence is generated based on the image features of each image frame in the multiple image frames in the time dimension or the space dimension.
[0102] Among them, the sequence features of the video to be recognized are used to characterize the similarity between the video to be recognized and multiple action categories.
[0103] As another possible implementation, the electronic device separately performs segmentation processing on each of the multiple image frames, divides each image frame into subsampled images of a preset size, and determines the sequence features of the video to be recognized based on the subsampled images included in the multiple image frames. Further, the electronic device inputs the sequence features of the video to be recognized into the classification layer, and then determines the multiple probabilities output by the classification layer as the probability distribution of the video to be recognized being similar to multiple action categories.
[0104] In this case, the image feature sequence is generated in the time dimension or the spatial dimension based on the image features of the subsampled images obtained by segmenting each of the multiple image frames.
[0105] For the specific implementation of this step, reference can be made to the subsequent descriptions of the embodiments of the present disclosure, and details will not be elaborated here.
[0106] S203. The electronic device determines the target action category corresponding to the video to be recognized based on the probability distribution of the video to be recognized being similar to multiple action categories.
[0107] Among them, the probability that the video to be recognized is similar to the target action category is greater than or equal to a preset threshold.
[0108] As a possible implementation, the electronic device determines the action category with the highest probability from the probability distribution of the video to be recognized being similar to multiple action categories as the target action category.
[0109] In this case, the preset threshold can be the maximum value among all the probabilities in the determined probability distribution.
[0110] As another possible implementation, the electronic device determines the action categories greater than the preset threshold from the probability distribution of the video to be recognized being similar to multiple action categories as the target action categories.
[0111] The above technical solution provided by the embodiments of the present disclosure calculates the probability distribution of the video to be recognized being similar to multiple action categories based on the self-attention model, and can directly determine the target action category where the video to be recognized is located from multiple action categories. Compared with the prior art, there is no need to set up a CNN, avoiding a large amount of calculations brought by convolution operations, and ultimately saving the computing resources of the device.
[0112] At the same time, since the image feature sequence is obtained based on multiple image frames in the time dimension or the spatial dimension, the image feature sequence can represent the time sequence of multiple image frames or the time sequence and spatial sequence of multiple image frames, and can learn the similarity between the video to be recognized and each action category from the time dimension and spatial dimension of multiple image frames to a certain extent, making the subsequent obtained probability distribution more accurate.
[0113] In one design, in order to be able to determine the probability distribution of the video to be recognized being similar to multiple action categories, the self-attention model provided by the embodiments of the present disclosure includes a self-attention encoding layer and a classification layer. The self-attention encoding layer is used to calculate the similarity features of the sequence composed of multiple image features with respect to multiple action categories, and the classification layer is used to calculate the probability distribution corresponding to the similarity features.
[0114] Meanwhile, as Figure 4 shown, S202 provided by the embodiments of the present disclosure specifically includes the following S2021-S2022.
[0115] S2021. The electronic device determines the target similarity features of the video to be recognized with respect to multiple action categories according to multiple image frames and the self-attention encoding layer.
[0116] Wherein, the target similarity features are used to characterize the similarity between the video to be recognized and multiple action categories.
[0117] As a possible implementation, the electronic device performs feature extraction processing on multiple image frames to obtain the image features of multiple image frames. Exemplarily, the image features of each image frame can be represented in the form of a vector, for example, it can be a vector with a length of 512 dimensions.
[0118] Furthermore, the electronic device merges the image features of each image frame with the corresponding position encoding features respectively to obtain multiple image input features of the self-attention encoding layer.
[0119] It can be understood that in this case, the sequence composed of multiple image input features is the above-mentioned image feature sequence.
[0120] It should be noted that each image feature corresponds to a position encoding feature, and the position encoding feature is used to identify the relative position of the corresponding image feature in the input sequence. The position encoding feature can be pre-generated by the electronic device according to the image features of multiple image frames. Exemplarily, the position encoding feature corresponding to the image feature can be a vector with 512 dimensions. In this case, the image input feature obtained by merging the image feature with the corresponding position encoding feature is also a vector with a length of 512 dimensions.
[0121] As an example, Figure 5 shows the timing diagram of the action recognition method provided by some embodiments. As Figure 4 shown, the number of multiple image frames is 9, and the electronic device converts 9 image frames into 9 image features (corresponding to Figure 5 0-9 in Figure 5In the *) above. Further, the electronic device merges the 9 image features with the 9 position encoding features respectively to obtain an image feature sequence composed of 9 image input features (in this case, the image feature sequence is obtained in the time dimension according to multiple image frames). In Figure 5 In the example of, the shape of the image feature sequence composed of 9 image input features is (b, 9, 512). Among them, b represents the image input feature, 9 represents the number of image input features, and 512 represents the length of the image input feature.
[0122] In some cases, the position encoding features corresponding to the image frames can be determined by the electronic device based on network braking learning, or can be determined based on the preset sin-cos rule. For the specific determination method of the position encoding features here, reference can be made to the description in the prior art, and details will not be elaborated here.
[0123] At the same time, the electronic device obtains a learnable category feature (denoted as the feature 0 in Figure 5 ), and merges the category feature with the corresponding position encoding feature to obtain a category input feature.
[0124] Among them, the category feature is used to represent the features of multiple action categories. The category feature can be pre-set in the self-attention encoding layer. Exemplarily, the category feature can be a vector with a length of N dimensions, where N can be the number of multiple action categories.
[0125] It can be understood that the category input feature is obtained by merging the category feature with the corresponding position encoding feature, and the category input feature is the feature used to input the self-attention encoding layer.
[0126] Taking Figure 5 as an example, after the electronic device determines the category feature, it merges the category feature with the corresponding position encoding feature to obtain a category input feature.
[0127] Subsequently, the electronic device takes the category input feature and the image feature sequence composed of multiple image input features as a sequence and inputs it into the self-attention encoding layer, and uses the sequence feature corresponding to the category input feature output by the self-attention encoding layer as the target similarity feature of the video to be recognized.
[0128] It can be understood that the sequence feature or the target similarity feature of the video to be recognized represents the similarity between the image features of multiple image frames and multiple action categories.
[0129] Taking Figure 5For example, the electronic device takes the category input feature as the 0th feature and the 9 image input features as the 1st - 9th features (image sequence features), and forms an input sequence to input into the self - attention encoding layer. In this case, the shape of the formed input sequence is (b, 10, 512). Among them, 10 is the number of features in the input sequence.
[0130] It should be noted that for the input and output of the self - attention encoding layer, taking Figure 5 as an example, if the input sequence of the self - attention encoding layer is (b, 10, 512), then its output sequence is also (b, 10, 512). At the same time, the input sequence input to the self - attention encoding layer includes 10 input features, and its output sequence also includes 10 output features. The 10 input features correspond one - to - one with the 10 output features. Each output feature reflects the weighted sum of the similarity features of the corresponding input feature relative to other input features.
[0131] For the specific implementation of this step, reference can be made to the subsequent similar specific descriptions in this disclosure, and details will not be elaborated here.
[0132] In some embodiments, the position of the category input feature in the input sequence can be the 0th, or any other position. The difference lies in the determined position encoding features being different.
[0133] The above - mentioned embodiments describe the implementation manner of directly taking multiple video frames as the input features of the self - attention encoding layer. As another possible implementation manner, the electronic device can also perform segmentation processing on each image frame, and determine the target similarity features of the video to be recognized based on the subsampled images obtained from the segmentation processing and the self - attention encoding layer.
[0134] For the specific implementation of this step, reference can be made to the subsequent descriptions in the embodiments of this disclosure, and details will not be elaborated here.
[0135] S2022. The electronic device inputs the target similarity features into the classification layer to obtain the probability distribution of the video to be recognized being similar to multiple action categories.
[0136] As a possible implementation manner, the electronic device inputs the target similarity features of the video to be recognized into the classification layer of the self - attention model to obtain the probability distribution of the similarity between the video to be recognized output by the classification layer and multiple action categories.
[0137] Exemplarily, the classification layer can be a multi - layer perceptron (MLP) connected to the self - attention encoding layer, which includes at least one fully - connected layer and a logistic regression softmax layer, and is used to classify the input target similarity features and calculate the probability distribution of each classification.
[0138] For the specific implementation of this step, reference can be made to the description in the prior art, and details will not be elaborated here.
[0139] As Figure 5 shown, the electronic device inputs the target similarity feature into the classification layer. The classification layer calculates through two fully connected layers and performs normalization operations through softmax, calculates and outputs the probabilities corresponding to each action category.
[0140] The following separately introduces the specific implementation of the self-attention encoder involved in the embodiments of the present disclosure:
[0141] After inputting the category input feature and multiple image input features into the self-attention encoder, the self-attention encoder calculates the input features based on the self-attention mechanism, and respectively obtains the output results corresponding to each input feature. Among them, the output feature corresponding to the category input feature satisfies the following formula under the constraint of the self-attention mechanism:
[0142]
[0143] Among them, S is the output feature corresponding to the category input feature, Q is the query transformation vector of the category input feature, K T is the transpose of the key transformation vector of the category input feature, V is the value transformation vector of the category input feature, and d is the dimension of the category input feature. Exemplarily, d can be 512.
[0144] In practical applications, the above self-attention encoding layer can adopt the multi-headed self-attention mechanism Multi-headedSelf-attention, or can also be processed by the single-headed self-attention mechanism.
[0145] It can be understood that the above QK T can be understood as the self-attention score in the self-attention encoding layer, and Softmax is the normalization process, that is, converting the reduced self-attention score into a probability distribution. Further, multiplying the probability distribution by V can be understood as performing weighted summation of the probability distribution and V.
[0146] It should be noted that the self-attention mechanism can process the input categorical input features, determine the feature weights of the categorical input features and multiple image input features, and transform the input categorical input features based on the categorical input features and the feature weights of each image input feature to obtain the output features corresponding to the categorical input features. After the categorical input features are processed by the self-attention mechanism, the corresponding output features will introduce the encoded information of multiple image input features through the self-attention mechanism. The process of the electronic device performing query transformation, key transformation, and value transformation on different input features based on the self-attention mechanism can specifically refer to the prior art and will not be elaborated here.
[0147] The above technical solution provided by the embodiments of the present disclosure can, by using the preset self-attention encoding layer, determine the similarity features between multiple image frames and multiple action categories based on the self-attention mechanism, and perform classification processing on the similarity features based on the classification layer to obtain the probability distribution of the video to be recognized belonging to multiple action categories, and can provide an implementation method for determining the probability distribution of the video to be recognized belonging to multiple action categories without using CNN, saving the computing resources consumed by the convolution operation.
[0148] In one design, in order to be able to learn more detailed features in multiple image frames of the video to be recognized, such as Figure 6 As shown, the action recognition method provided by the embodiments of the present disclosure further includes the following S204 before S2021:
[0149] S204. The electronic device divides each image frame of the multiple image frames to obtain multiple subsampled images.
[0150] As a possible implementation, the electronic device can divide each image frame according to the preset segmentation pixel size to obtain multiple subsampled images.
[0151] Among them, the segmentation pixel size can be pre-set by the operation and maintenance personnel of the action recognition system in the electronic device in advance.
[0152] Exemplarily, when the size of each image frame is 256*256 and the segmentation pixel size is 32*32, each image frame can be divided into 64 subsampled images. If there are 10 multiple image frames of the video to be recognized, 640 subsampled images can be obtained after all the image frames are divided.
[0153] Figure 7 Shows a schematic diagram of image segmentation processing. As shown in Figure 7 For each image frame in multiple image frames, each image frame can be divided into multiple subsampled images based on the size of each image frame and the segmentation pixel size.
[0154] In this case, as Figure 6 shown, the above S2021 provided by the embodiments of the present disclosure may specifically further include the following S301 - S302.
[0155] S301. The electronic device determines the sequence features of the video to be recognized according to a plurality of sub - sampled images and the self - attention encoding layer.
[0156] Among them, the sequence features include temporal sequence features, or temporal sequence features and spatial sequence features. The temporal sequence features are used to characterize the similarity of the video to be recognized with multiple action categories in the time dimension, and the spatial sequence features are used to characterize the similarity of the video to be recognized with multiple action categories in the space dimension.
[0157] As a possible implementation manner, the electronic device divides the plurality of sub - sampled images into a plurality of temporal sampling sequences according to the time sequence, and determines the sub - temporal sequence features of each temporal sampling sequence according to each temporal sampling sequence and the self - attention encoding layer. Further, the electronic device determines the temporal sequence features of the video to be recognized according to the determined plurality of sub - temporal sequence features.
[0158] For the specific implementation manner of this step, reference may be made to the subsequent description of the embodiments of the present disclosure, and details are not described herein again.
[0159] Meanwhile, in the case where the sequence features include temporal sequence features and spatial sequence features, the electronic device also divides the plurality of sub - sampled images into a plurality of spatial sampling sequences according to the spatial sequence of the image frames. Further, the electronic device determines the sub - spatial sequence features of each spatial sampling sequence according to each spatial sampling sequence and the self - attention encoding layer. Finally, the electronic device determines the spatial sequence features of the video to be recognized according to the plurality of sub - spatial sequence features.
[0160] For the specific implementation manner of this step, reference may be made to the subsequent description of the embodiments of the present disclosure, and details are not described herein again.
[0161] S302. The electronic device determines the target similarity features according to the sequence features of the video to be recognized.
[0162] In the case where the sequence features include temporal sequence features, the electronic device determines the temporal sequence features of the video to be recognized as the target similarity features of the video to be recognized.
[0163] In the case where the sequence features include temporal sequence features and spatial sequence features, the electronic device merges the determined temporal sequence features and spatial sequence features, and determines the merged features obtained by merging as the target similarity features of the video to be recognized.
[0164] It should be noted that the above combined features can also be obtained by fusing time series features and spatial series features based on other fusion methods, and the embodiments of the present disclosure do not limit this.
[0165] The above technical solution provided by the embodiments of the present disclosure divides each image frame into multiple sub-sampled images of a preset size, and determines time series features from the time dimension and spatial series features from the spatial dimension according to the multiple sub-sampled images. The determined target similarity features can reflect the time features and spatial features of the video to be recognized, which can make the subsequent determined target action categories more accurate.
[0166] In one design, in order to determine the time series features of the video to be recognized, as Figure 8 shown, S301 provided by the embodiments of the present disclosure specifically includes the following S3011 - S3013.
[0167] S3011. The electronic device determines at least one time sampling sequence from the multiple sub-sampled images.
[0168] Wherein, each time sampling sequence includes sub-sampled images located at the same position in each image frame.
[0169] As a possible implementation, the electronic device divides the multiple sub-sampled images into at least one time sampling sequence based on the time series.
[0170] It should be noted that the number of time sampling sequences is the number of sub-sampled images obtained by dividing each image frame.
[0171] Figure 9 Fig. shows a schematic diagram of a time sampling sequence. As Figure 9 shown, the multiple image frames include Image Frame 1, Image Frame 2, and Image Frame 3, and each image frame includes 9 sub-sampled images. Among the 3 image frames, the sub-sampled images at the upper left corner position of each image frame can form the first time sampling sequence. For another example, the sub-sampled images at the middle right position of each image frame can form the second time sampling sequence.
[0172] S3012. The electronic device determines the sub-time series features of each time sampling sequence according to each time sampling sequence and the self-attention encoding layer.
[0173] Wherein, the sub-time series features are used to characterize the similarity between each time sampling sequence and multiple action categories.
[0174] As a possible implementation, for each time sampling sequence, the electronic device performs position encoding merging based on the image features of each sub-sampled image to obtain a first image input feature (in combination with the above embodiments, the sequence composed of multiple first image input features corresponding to each time sampling sequence feature corresponds to the above image feature sequence. In this case, the image feature sequence is obtained based on multiple image frames in the time dimension). At the same time, the electronic device also performs position encoding merging according to the category features to obtain a category input feature. Further, the electronic device inputs the sequence composed of the category input feature and all the first image input features (image feature sequence) into the self-attention encoding layer, and uses the feature corresponding to the category input feature output by the self-attention encoding layer as the sub-time sequence feature of this time sampling sequence.
[0175] For the specific implementation of this step, reference can be made to the subsequent description of the embodiments of the present disclosure, and details will not be elaborated here.
[0176] S3013. The electronic device determines the time sequence feature of the video to be recognized according to the sub-time sequence features of at least one time sampling sequence.
[0177] As a possible implementation, the electronic device merges the sub-time sequence features of at least one time sampling sequence, and determines the merged feature obtained by the merging as the time sequence feature of the video to be recognized.
[0178] It should be noted that the above merged feature can also be obtained by fusing multiple sub-time sequence features based on other fusion methods, and the embodiments of the present disclosure do not limit this.
[0179] The above technical solutions provided by the embodiments of the present disclosure at least bring the following: dividing multiple sub-sampled images into at least one time sampling sequence, determining the sub-time sequence features of each time sampling sequence, and determining the time sequence feature of the video to be recognized according to multiple sub-time sequence features. Since the positions of the sub-sampled images in different image frames in each time sampling sequence are the same, the time sequence features determined based on this are more comprehensive and accurate.
[0180] In one design, in order to be able to determine the sub-time sequence features of each time sampling sequence, as Figure 10 shown, S3012 provided by the embodiments of the present disclosure specifically includes the following S401 - S403.
[0181] S401. The electronic device determines multiple first image input features and a category input feature.
[0182] Among them, each first image input feature is obtained by performing positional encoding and merging on the image features of the subsampled images included in the first time sampling sequence, and the first time sampling sequence is any one of at least one time sampling sequence. The category input feature is obtained by performing positional encoding and merging on the category feature, and the category feature is used to represent multiple action categories.
[0183] As a possible implementation, the electronic device determines the image features of each subsampled image in the first time sampling sequence. Further, the electronic device merges the image features of each subsampled image with the corresponding positional encoding features to obtain the first image input feature corresponding to the image feature of each subsampled image.
[0184] Meanwhile, the electronic device also obtains the category features corresponding to multiple action categories, and merges the category features with the corresponding positional encoding features to obtain the category input feature.
[0185] For the specific implementation of this step, reference may be made to the specific description of S2021 in the embodiments of the present disclosure above, and details will not be elaborated here.
[0186] S402. The electronic device inputs multiple first image input features and the category input feature into the self-attention encoding layer to obtain the output features of the self-attention encoding layer.
[0187] For the specific implementation of this step, reference may be made to the specific description of S2021 in the embodiments of the present disclosure above, and details will not be elaborated here.
[0188] S403. The electronic device determines the output feature corresponding to the category input feature output by the self-attention encoding layer as the sub-time sequence feature of the first time sampling sequence.
[0189] As a possible implementation, the electronic device determines the output feature corresponding to the category input feature as the sub-time sequence feature of the first time sampling sequence.
[0190] Figure 11 Shows a timing diagram for determining the time sequence feature in the above embodiment. As Figure 11 shown, for the first time sampling sequence, the electronic device converts the subsampled images included in the first time sampling sequence into 9 image features (0-9), and merges them with the corresponding positional encoding features to obtain 9 image input features corresponding to the first time sampling sequence (corresponding to the image sequence features obtained from the above multiple image frames based on the time dimension). Further, the electronic device inputs the category input feature and the 9 image input features into the self-attention encoding layer to obtain the output feature corresponding to the category input feature, and determines this output feature as the sub-time sequence feature of the first time sampling sequence.
[0191] The above technical solution provided by the embodiments of the present disclosure can, by using the self-attention encoding layer, determine the sub-time series features of each time sampling sequence with respect to multiple action categories, and compared with the prior art, convolution operations can be avoided, thereby saving corresponding computing resources.
[0192] In one design, when the sequence features of the video to be recognized include time series features and spatial sequence features, in order to determine the spatial sequence features of the video to be recognized, as Figure 12 shown, the above S301 provided by the embodiments of the present disclosure specifically further includes the following S3014-S3016.
[0193] S3014. The electronic device determines at least one spatial sampling sequence from multiple subsampled images.
[0194] Wherein, each spatial sampling sequence includes the subsampled images in one image frame.
[0195] As a possible implementation, the electronic device divides the multiple subsampled images into at least one spatial sampling sequence based on the spatial sequence.
[0196] Exemplarily, the subsampled images included in one image frame can be determined as one spatial sampling sequence. In this case, the number of at least one spatial sampling sequence is the same as the number of multiple image frames. For example, in Figure 7 , all the subsampled images included in each image frame can be used as one spatial sampling sequence.
[0197] As another possible implementation, for the first image frame among the multiple image frames, the electronic device can also determine a preset number of target subsampled images located at preset positions from the subsampled images included in the first image frame, and determine the target subsampled images as the spatial sampling sequence corresponding to the first image frame.
[0198] Wherein, the first image frame is any one of the multiple image frames.
[0199] Exemplarily, the target subsampled images in the first image frame can be any adjacent M sampled value images.
[0200] Figure 13 shows a schematic diagram of a spatial sampling sequence. As Figure 13 shown, in image frame 1, the subsampled images at adjacent preset positions can be set as the first spatial sampling sequence. For example, the first spatial sampling sequence can be 4 subsampled images in the upper left part of image frame 1, and the first spatial sampling sequence can also be 4 subsampled images in the lower right part of image 1.
[0201] In the above technical solution provided by the embodiments of the present disclosure, when determining each spatial sampling sequence, at least one spatial sequence feature can be generated by using a preset number of target sub-sampled images at preset positions. In this way, without affecting the spatial sequence features, the number of sub-sampled images in each spatial sampling sequence can be reduced, and the computational consumption of the subsequent self-attention encoding layer can be reduced.
[0202] S3015. The electronic device determines the subspace sequence features of each spatial sampling sequence according to each spatial sampling sequence and the self-attention encoding layer.
[0203] Among them, the subspace sequence features are used to characterize the similarity between each spatial sampling sequence and multiple action categories.
[0204] For the specific implementation manner of this step, reference can be made to the specific description of S3012 above. The difference is that the processing objects are different, so it will not be elaborated here.
[0205] S3016. The electronic device determines the spatial sequence features of the video to be recognized according to the subspace sequence features of at least one spatial sampling sequence.
[0206] As a possible implementation manner, the electronic device merges the subspace sequence features of at least one spatial sampling sequence, and determines the merged feature obtained by the merging as the spatial sequence features of the video to be recognized.
[0207] It should be noted that the above merged feature can also be obtained by fusing multiple subspace sequence features based on other fusion methods. The embodiments of the present disclosure do not limit this.
[0208] In the above technical solution provided by the embodiments of the present disclosure, multiple sub-sampled images are divided into at least one spatial sampling sequence, the subspace sequence features of each spatial sampling sequence are determined, and the spatial sequence features of the video to be recognized are determined according to the multiple subspace sequence features. In this way, the spatial sequence features determined based on this are more comprehensive and accurate.
[0209] In one design, in order to be able to determine the subspace sequence features of each temporal sampling sequence, as Figure 14 shown, S3015 provided by the embodiments of the present disclosure specifically includes the following S501-S503.
[0210] S501. The electronic device determines multiple second image input features and category input features.
[0211] Among them, each second image input feature is obtained by performing positional encoding and merging on the image features of the subsampled images included in the first spatial sampling sequence, and the first spatial sampling sequence is any one of at least one spatial sampling sequence. The category input feature is obtained by performing positional encoding and merging on the category feature, and the category feature is used to represent multiple action categories.
[0212] As a possible implementation, the electronic device determines the image features of each subsampled image in the first spatial sampling sequence. Further, the electronic device merges the image features of each subsampled image with the corresponding positional encoding features to obtain the second image input feature corresponding to the image feature of each subsampled image (in some embodiments, multiple second image input features correspond to the image sequence features in the above embodiments. In this case, the image feature sequence is obtained in the spatial dimension according to multiple image frames).
[0213] At the same time, the electronic device also obtains the category features corresponding to multiple action categories, and merges the category features with the corresponding positional encoding features to obtain the category input feature.
[0214] For the specific implementation of this step, reference may be made to the specific description of S2021 in the embodiments of the present disclosure, and details are not described herein again.
[0215] S502. The electronic device inputs multiple second image input features and the category input feature into the self-attention encoding layer to obtain the output features of the self-attention encoding layer.
[0216] For the specific implementation of this step, reference may be made to the specific description of S2021 in the embodiments of the present disclosure, and details are not described herein again.
[0217] S503. The electronic device determines the output feature corresponding to the category input feature output by the self-attention encoding layer as the subspace sequence feature of the first spatial sampling sequence.
[0218] As a possible implementation, the electronic device determines the output feature corresponding to the category input feature as the sub-time sequence feature of the first time sampling sequence.
[0219] The above technical solution provided by the embodiments of the present disclosure can determine the subspace sequence feature of each spatial sampling sequence relative to multiple action categories by using the self-attention encoding layer, and can avoid consuming computing resources caused by using convolution operations.
[0220] Figure 15 Shows a timing diagram of an action recognition method, as Figure 15As shown, after the electronic device performs segmentation processing on each of the multiple image frames, multiple subsampled images are obtained. Further, the electronic device determines at least one temporal sampling sequence and at least one spatial sampling sequence among the multiple subsampled images, and respectively determines the respective sub-temporal sequence features of each temporal sampling sequence according to each temporal sampling sequence and the self-attention encoding layer, and determines the respective sub-spatial sequence features of each spatial sampling sequence according to each spatial sampling sequence and the self-attention encoding layer. Subsequently, the electronic device combines the determined multiple sub-temporal sequence features to obtain the temporal sequence features of the video to be recognized, and the electronic device combines the determined sub-spatial sequence features to obtain the spatial sequence features of the video to be recognized. Further, the electronic device combines the temporal sequence features and the spatial sequence features of the video to be recognized to obtain the target similarity features of the video to be recognized, and inputs the target similarity features into the classification layer to determine the probability distribution of the similarity between the video to be recognized and multiple action categories.
[0221] In one design, in order to be able to train the above-mentioned self-attention model provided by the embodiments of the present disclosure, the embodiments of the present disclosure further provide a model training method, and this model training method can also be applied to the above-mentioned action recognition system.
[0222] In practical applications, the model training method provided by the embodiments of the present disclosure can be applied to a model training device or an electronic device. Below, with reference to the accompanying drawings, taking the application of the model training method to an electronic device as an example, the action recognition method provided by the embodiments of the present disclosure will be described.
[0223] As Figure 16 shown, the model training method provided by the embodiments of the present disclosure includes the following S601-S602.
[0224] S601. The electronic device obtains multiple sample image frames of a sample video and the sample action category where the sample video is located.
[0225] As a possible implementation manner, the electronic device obtains a sample video, decodes and extracts frames from the sample video, and uses the multiple sampled frames obtained by the decoding and frame extraction as multiple sample image frames.
[0226] As another possible implementation manner, after the electronic device obtains the sample video, it decodes and extracts frames from the sample video, and preprocesses the multiple sampled frames obtained by the frame extraction to obtain multiple sample image frames.
[0227] Among them, the image preprocessing includes at least one operation of cropping, image enhancement, and scaling.
[0228] As a third possible implementation, after obtaining the sample video, the electronic device can decode the sample video to obtain multiple decoded frames, and perform the above-mentioned preprocessing on the multiple decoded frames to obtain the preprocessed decoded frames. Further, the electronic device performs frame extraction and random sampling on the preprocessed decoded frames to obtain multiple sample image frames.
[0229] For the specific implementation of this step, reference can be made to the specific description in S201 provided in the embodiments of the present disclosure. The difference lies in the different processing objects, and details will not be elaborated here.
[0230] S602. The electronic device performs self-attention training based on multiple sample image frames and the sample action category to obtain a trained self-attention model.
[0231] Among them, the self-attention model is used to calculate the similarity between the sample image feature sequence and multiple action categories, and the sample image feature sequence is obtained based on multiple sample image frames. The sample image feature sequence is obtained based on multiple sample image frames in the time dimension or the spatial dimension.
[0232] As a possible implementation, the electronic device determines the sample similarity feature of the sample video based on multiple sample image frames and the self-attention encoding layer, uses the sample similarity feature as the sample feature, uses the sample action category as the label, trains a preset neural network to obtain a trained classification layer, and finally trains the self-attention model.
[0233] Among them, the sample similarity feature is used to represent the similarity between the sample video and multiple action categories.
[0234] In this case, the initial self-attention model includes the above-mentioned self-attention encoding layer and the preset neural network.
[0235] As another possible implementation, the electronic device can also perform self-attention training on the initial self-attention model as a whole, use the image features of multiple sample image frames as the sample features, use the sample action category of the sample video as the label, and perform supervised training on the input and output of the entire initial self-attention model until a trained self-attention model is obtained.
[0236] As a third possible implementation, the electronic device can also perform self-attention training on the initial self-attention model as a whole, divide each sample image frame in multiple sample image frames to obtain multiple sub-sampled sample images, and perform supervised training on the initial self-attention model based on the multiple sub-sampled sample images to obtain a trained self-attention model.
[0237] In the above process of training the initial self-attention model as a whole, the gradient parameters to be adjusted in the initial self-attention model include parameters such as query, key, and value in the self-attention encoding layer and weight parameters in the classification layer.
[0238] In the above process of adjusting the parameters of query, key, value, etc. in the self-attention encoding layer and weight parameters in the classification layer, it can specifically refer to the prior art and will not be elaborated here.
[0239] It should be noted that in the above process of iterative training of the neural network, the cross-entropy loss function celoss can be specifically used for training.
[0240] The specific steps for the electronic device to determine the sample similarity features of the sample video based on multiple sample image frames and the self-attention encoding layer in this step can refer to the specific description of S2021 in the above embodiments of the present disclosure. The difference is that the processing objects are different and will not be elaborated here.
[0241] The above technical solution provided by the embodiments of the present disclosure performs self-attention training on the initial self-attention model based on multiple sample image frames of the sample video and the sample action category where the sample video is located, and trains to obtain the self-attention model. Since only the sample similarity features similar to different sample categories of multiple sample image frames need to be determined based on the self-attention mechanism during the training process, compared with the prior art, there is no need to perform convolution operations based on CNN, avoiding a large amount of calculations brought by using convolution operations, and finally saving the computing resources of the device.
[0242] In one design, in order to be able to determine the sample similarity features of the sample video based on multiple sample image frames and the self-attention encoding layer, the model training method provided by the embodiments of the present disclosure further includes the following S603.
[0243] S603: The electronic device divides each sample image frame of the multiple sample image frames to obtain multiple sub-sampled sample images.
[0244] The specific implementation manner of this step can refer to the specific description of S204 in the above of the present disclosure and will not be elaborated here.
[0245] In this case, the above S602 provided by the embodiments of the present disclosure specifically includes the following S6021-S6022.
[0246] S6021: The electronic device determines the sample sequence features of the sample video according to the multiple sub-sampled sample images and the self-attention encoding layer.
[0247] Among them, the sample sequence features include sample time series features, or sample time series features and sample spatial sequence features. The sample time series features are used to characterize the similarity between the sample video and multiple action categories in the time dimension, and the sample spatial sequence features are used to characterize the similarity between the sample video and multiple action categories in the spatial dimension.
[0248] For the specific implementation of this step, reference may be made to the specific description of S301 above in this disclosure, and details will not be elaborated here.
[0249] S6022. The electronic device determines sample similarity features according to the sample sequence features of the sample video.
[0250] For the specific implementation of this step, reference may be made to the specific description of S302 above in this disclosure, and details will not be elaborated here.
[0251] In the above technical solution provided by the embodiments of this disclosure, each sample image frame is segmented into multiple sub-sampled sample images of a preset size, and the sample time series features are determined from the time dimension according to the multiple sub-sampled sample images, and the sample spatial sequence features are determined from the spatial dimension. The sample similarity features determined in this way can reflect the time features and spatial features of the sample video, and can make the subsequent self-attention model obtained by training more accurate.
[0252] In one design, in order to be able to determine the time sample sequence features of the sample video, S6021 provided by the embodiments of this disclosure specifically includes the following S701-S703.
[0253] S701. The electronic device determines at least one sample time sampling sequence from the multiple sub-sampled sample images.
[0254] Among them, each sample time sampling sequence includes the sub-sampled sample images at the same position in each sample image frame.
[0255] For the specific implementation of this step, reference may be made to the specific description of S3011 above in this disclosure. The difference lies in the processing object, and details will not be elaborated here.
[0256] S702. The electronic device determines the sample sub-time series features of each sample time sampling sequence according to each sample time sampling sequence and the self-attention encoding layer.
[0257] Among them, the sample sub-time series features are used to characterize the similarity between each sample time sampling sequence and multiple action categories.
[0258] For the specific implementation of this step, reference may be made to the specific description of S3012 above in this disclosure. The difference lies in the processing object, and details will not be elaborated here.
[0259] S703. The electronic device determines the sample time series feature of the sample video according to the sample sub-time series features of at least one sample time sampling series.
[0260] For the specific implementation of this step, reference may be made to the specific description of S3013 above in this disclosure. The difference lies in the different processing objects, so it will not be elaborated here.
[0261] In the above technical solution provided by the embodiments of this disclosure, multiple sub-sampled sample images are divided into at least one sample time sampling series, the sample sub-time series features of each sample time sampling series are determined, and the sample time series feature of the sample video is determined according to the multiple sample sub-time series features. Since the positions of the sub-sampled sample images in each sample time sampling series are the same in different sample image frames, the sample time series features determined based on this are more comprehensive and accurate.
[0262] In one design, in order to be able to determine the sub-time series features of each time sampling series, S702 provided by the embodiments of this disclosure specifically includes the following S7021 - S7023.
[0263] S7021. The electronic device determines multiple first sample image input features and category input features.
[0264] Among them, each first sample image input feature is obtained by performing position encoding and merging on the image features of the sub-sampled sample images included in the first sample time sampling series, and the first sample time sampling series is any one of at least one sample time sampling series. The category input feature is obtained by performing position encoding and merging on the category features, and the category features are used to represent multiple action categories.
[0265] Combined with the above embodiments, the sequence composed of multiple first sample image input features corresponds to the above sample image feature sequence. In this case, the sample image feature sequence is obtained based on multiple sample image frames in the time dimension.
[0266] For the specific implementation of this step, reference may be made to the specific description of S401 above in this disclosure. The difference lies in the different processing objects, so it will not be elaborated here.
[0267] S7022. The electronic device inputs the multiple first sample image input features and the category input features into the self-attention encoding layer to obtain the output features of the self-attention encoding layer.
[0268] For the specific implementation of this step, reference may be made to the specific description of S402 above in this disclosure. The difference lies in the different processing objects, so it will not be elaborated here.
[0269] S7023. The electronic device determines the output feature corresponding to the class input feature output by the self-attention encoding layer as the sample sub-time series feature of the first sample time sampling sequence.
[0270] For the specific implementation of this step, reference can be made to the specific description of S403 above in this disclosure. The difference lies in the different processing objects, so it will not be elaborated here.
[0271] With the above technical solution provided by the embodiments of this disclosure, by using the self-attention encoding layer, it is possible to determine the sample sub-time series features of each sample time sampling sequence relative to multiple action categories, which can avoid consuming computing resources caused by using convolution operations.
[0272] In one design, when the sample sequence feature of the sample video includes the sample time series feature and the sample space series feature, in order to determine the sample space series feature of the sample video, S6022 provided by the embodiments of this disclosure specifically further includes the following S704 - S706.
[0273] S704. The electronic device determines at least one sample space sampling sequence from multiple sample sub-sampled images.
[0274] Each sample space sampling sequence includes sub-sampled sample images in a sample image frame.
[0275] For the specific implementation of this step, reference can be made to the specific description of S3014 above in this disclosure. The difference lies in the different processing objects, so it will not be elaborated here.
[0276] S705. The electronic device determines the sample sub-space sequence feature of each sample space sampling sequence according to each sample space sampling sequence and the self-attention encoding layer.
[0277] The sample sub-space sequence feature is used to characterize the similarity between each sample space sampling sequence and multiple action categories.
[0278] For the specific implementation of this step, reference can be made to the specific description of S3015 above in this disclosure. The difference lies in the different processing objects, so it will not be elaborated here.
[0279] S706. The electronic device determines the sample space sequence feature of the sample video according to the sample sub-space sequence features of at least one sample space sampling sequence.
[0280] For the specific implementation of this step, reference can be made to the specific description of S3016 above in this disclosure. The difference lies in the different processing objects, so it will not be elaborated here.
[0281] The above technical solution provided by the embodiments of the present disclosure divides multiple subsampled sample images into at least one sample space sampling sequence, determines the sample subspace sequence features of each sample space sampling sequence, and determines the sample space sequence features of the sample video based on the multiple sample subspace sequence features. In this way, the determined sample space sequence features are more comprehensive and accurate.
[0282] In one design, in order to be able to determine the sample sub-time sequence features of each sample time sampling sequence, S705 provided by the embodiments of the present disclosure specifically includes the following S7051-S7053.
[0283] S7051. The electronic device determines multiple second sample image input features and class input features.
[0284] Wherein, each second sample image input feature is obtained by performing position encoding and merging on the image features of the subsampled sample images included in the first sample space sampling sequence, and the first sample space sampling sequence is any one of the at least one sample space sampling sequence. The class input feature is obtained by performing position encoding and merging on the class feature, and the class feature is used to represent multiple action classes.
[0285] Combined with the above embodiments, the sequence composed of multiple second sample image input features corresponds to the above sample image feature sequence. In this case, the sample image feature sequence is obtained in the spatial dimension according to multiple sample image frames.
[0286] For the specific implementation manner of this step, reference may be made to the specific description of S501 above in the present disclosure. The difference lies in the processing object, and details will not be elaborated here.
[0287] S7052. The electronic device inputs the multiple second sample image input features and the class input features into the self-attention encoding layer to obtain the output features of the self-attention encoding layer.
[0288] For the specific implementation manner of this step, reference may be made to the specific description of S502 above in the present disclosure. The difference lies in the processing object, and details will not be elaborated here.
[0289] S7053. The electronic device determines the output features corresponding to the class input features output by the self-attention encoding layer as the sample subspace sequence features of the first sample space sampling sequence.
[0290] For the specific implementation manner of this step, reference may be made to the specific description of S503 above in the present disclosure. The difference lies in the processing object, and details will not be elaborated here.
[0291] Figure 17 It is a schematic structural diagram of an action recognition device shown according to an exemplary embodiment. Refer toFigure 17 As shown in Figure 17 , the action recognition device 80 provided by an embodiment of the present disclosure can be applied to an electronic device and is used to execute the action recognition method provided by the above embodiment. The action recognition device 80 includes an acquisition unit 801 and a determination unit 802.
[0292] The acquisition unit 801 is used to acquire multiple image frames of the video to be recognized.
[0293] The determination unit 802 is used to determine the probability that the video to be recognized is similar to multiple action categories according to the multiple image frames and a pre-trained self-attention model after the acquisition unit 801 acquires the multiple image frames. The self-attention model is used to calculate the similarity between the image feature sequence and multiple action categories through the self-attention mechanism. The image feature sequence is obtained based on the multiple image frames in the time dimension or the space dimension. The probability distribution includes the probability that the video to be recognized is similar to each action category among the multiple action categories.
[0294] The determination unit 802 is further used to determine the target action category corresponding to the video to be recognized based on the probability distribution that the video to be recognized is similar to multiple action categories. The probability that the video to be recognized is similar to the target action category is greater than or equal to a preset threshold.
[0295] Optionally, as Figure 17 shown in Figure 17 , the self-attention model provided by an embodiment of the present disclosure includes a self-attention encoding layer and a classification layer. The self-attention encoding layer is used to calculate the similarity features of the image feature sequence with respect to multiple action categories, and the classification layer is used to calculate the probability distribution corresponding to the similarity features. The above determination unit 802 is specifically used for:
[0296] Determine the target similarity features of the video to be recognized with respect to multiple action categories according to the multiple image frames and the self-attention encoding layer. The target similarity features are used to characterize the similarity between the video to be recognized and multiple action categories.
[0297] Input the target similarity features into the classification layer to obtain the probability distribution that the video to be recognized is similar to multiple action categories.
[0298] Optionally, as Figure 17 shown in Figure 17 , the action recognition device 80 provided by an embodiment of the present disclosure further includes a processing unit 803.
[0299] The processing unit 803 is used to segment each image frame of the multiple image frames to obtain multiple subsampled images before the determination unit 802 determines the target similarity features of the video to be recognized with respect to multiple action categories according to the multiple image frames and the self-attention encoding layer.
[0300] The determination unit 802 is specifically configured to determine the sequence features of the video to be recognized according to a plurality of subsampled images and the self-attention encoding layer, and determine the target similarity features according to the sequence features of the video to be recognized. The sequence features include temporal sequence features, or temporal sequence features and spatial sequence features. The temporal sequence features are used to characterize the similarity between the video to be recognized and multiple action categories in the temporal dimension, and the spatial sequence features are used to characterize the similarity between the video to be recognized and multiple action categories in the spatial dimension.
[0301] Optionally, as Figure 17 shown, the determination unit 802 provided in the embodiment of the present disclosure is specifically configured to:
[0302] Determine at least one temporal sampling sequence from the multiple subsampled images. Each temporal sampling sequence includes the subsampled images at the same position in each image frame.
[0303] According to each temporal sampling sequence and the self-attention encoding layer, determine the sub-temporal sequence features of each temporal sampling sequence. The sub-temporal sequence features are used to characterize the similarity between each temporal sampling sequence and multiple action categories.
[0304] According to the sub-temporal sequence features of at least one temporal sampling sequence, determine the temporal sequence features of the video to be recognized.
[0305] Optionally, as Figure 17 shown, the determination unit 802 provided in the embodiment of the present disclosure is specifically configured to:
[0306] Determine a plurality of first image input features and category input features. Each first image input feature is obtained by performing position encoding merging on the image features of the subsampled images included in the first temporal sampling sequence, and the first temporal sampling sequence is any one of the at least one temporal sampling sequence. The category input feature is obtained by performing position encoding merging on the category features, and the category features are used to characterize multiple action categories.
[0307] Input the plurality of first image input features and the category input features into the self-attention encoding layer, and determine the output features corresponding to the category input features output by the self-attention encoding layer as the sub-temporal sequence features of the first temporal sampling sequence.
[0308] Optionally, as Figure 17 shown, the determination unit 802 provided in the embodiment of the present disclosure is specifically configured to:
[0309] Determine at least one spatial sampling sequence from the multiple subsampled images. Each spatial sampling sequence includes the subsampled images in one image frame.
[0310] Determine the subspace sequence features of each spatial sampling sequence according to each spatial sampling sequence and the self-attention encoding layer. The subspace sequence features are used to characterize the similarity between each spatial sampling sequence and multiple action categories.
[0311] Determine the spatial sequence features of the video to be recognized according to the subspace sequence features of at least one spatial sampling sequence.
[0312] Optionally, as Figure 17 shown, the determination unit 802 provided in the embodiment of the present disclosure is specifically configured to:
[0313] For the first image frame, determine a preset number of target sub-sampled images located at preset positions from the sub-sampled images included in the first image frame, and determine the target sub-sampled images as the spatial sampling sequence corresponding to the first image frame. The first image frame is any one of multiple image frames.
[0314] Optionally, as Figure 17 shown, the determination unit 802 provided in the embodiment of the present disclosure is specifically configured to:
[0315] Determine multiple second image input features and class input features. Each second image input feature is obtained by performing position encoding merging on the image features of the sub-sampled images included in the first spatial sampling sequence, and the first spatial sampling sequence is any one of at least one spatial sampling sequence. The class input feature is obtained by performing position encoding merging on the class feature, and the class feature is used to characterize multiple action categories.
[0316] Input the multiple second image input features and the class input features into the self-attention encoding layer, and determine the output feature corresponding to the class input feature output by the self-attention encoding layer as the subspace sequence feature of the first spatial sampling sequence.
[0317] Optionally, as Figure 17 shown, the multiple image frames provided in the embodiment of the present disclosure are obtained based on image preprocessing, and the image preprocessing includes at least one operation of cropping, image enhancement, and scaling.
[0318] Figure 18 is a schematic structural diagram of a model training device shown according to an exemplary embodiment. Refer to Figure 18 shown, the model training device 90 provided in the embodiment of the present disclosure can be used for the above-mentioned electronic device, and is specifically configured to execute the model training method provided in the above embodiment. The model training device 90 includes an acquisition unit 901 and a training unit 901.
[0319] The acquisition unit 901 is configured to acquire multiple sample image frames of the sample video and the sample action category where the sample video is located.
[0320] A training unit 901, which is configured to, after an acquisition unit 901 acquires a plurality of sample image frames and sample action categories, perform self-attention training based on the plurality of sample image frames and the sample action categories to obtain a trained self-attention model. The self-attention model is used to calculate the similarity between a sample image feature sequence and a plurality of action categories, and the sample image feature sequence is obtained based on the plurality of sample image frames in the time dimension or the space dimension.
[0321] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0322] Figure 19 It is a schematic structural diagram of an electronic device provided by the present disclosure. As Figure 19 , the electronic device 100 may include at least one processor 1001 and a memory 1003 for storing processor-executable instructions. Among them, the processor 1001 is configured to execute the instructions in the memory 1003 to implement the action recognition method in the above embodiments.
[0323] In addition, the electronic device 100 may further include a communication bus 1002 and at least one communication interface 1004.
[0324] The processor 1001 may be a central processing unit (CPU), a microprocessing unit, an ASIC, or one or more integrated circuits for controlling the execution of the program of the present disclosure solution.
[0325] The communication bus 1002 may include a path for transmitting information between the above components.
[0326] The communication interface 1004 uses any transceiver-like device for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.
[0327] The memory 1003 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory can exist independently and be connected to the processing unit through a bus. The memory can also be integrated with the processing unit.
[0328] Among them, the memory 1003 is used to store the instructions for executing the solution of the present disclosure and is controlled by the processor 1001 to execute. The processor 1001 is used to execute the instructions stored in the memory 1003, thereby implementing the functions in the method of the present disclosure.
[0329] As an example, in combination with Figure 17 , the functions implemented by the acquisition unit 801, the determination unit 802, and the processing unit 803 in the action recognition device are the same as those of Figure 19 the processor 1001.
[0330] In a specific implementation, as an embodiment, the processor 1001 may include one or more CPUs, such as Figure 19 the CPU0 and CPU1 in
[0331] In a specific implementation, as an embodiment, the electronic device 100 may include multiple processors, such as Figure 19 the processor 1001 and the processor 1007 in . Each of these processors can be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0332] In a specific implementation, as an example, the electronic device 100 may further include an output device 1005 and an input device 1006. The output device 1005 communicates with the processor 1001 and can display information in various ways. For example, the output device 1005 may be a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector, etc. The input device 1006 communicates with the processor 1001 and can accept input of an account in various ways. For example, the input device 1006 may be a mouse, a keyboard, a touch screen device, or a sensing device, etc.
[0333] Those skilled in the art can understand that Figure 19 the structure shown in does not constitute a limitation on the electronic device 100, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt different component arrangements.
[0334] Meanwhile, another schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present disclosure may also refer to the description of the electronic device in the above Figure 19 wherein the difference is that the processor included in the electronic device is used to execute the steps in the prediction training method executed by the model training device in the above embodiments.
[0335] Some embodiments of the present disclosure provide a computer-readable storage medium (for example, a non-transitory computer-readable storage medium), in which computer program instructions are stored. When the computer program instructions run on a computer (for example, an electronic device), the computer is caused to execute the action recognition method or the model training method in any one of the above embodiments.
[0336] Exemplarily, the above computer-readable storage medium may include, but is not limited to: magnetic storage devices (such as hard disks, floppy disks, or magnetic tapes, etc.), optical discs (such as CDs (Compact Disks), DVDs (Digital Versatile Disks), etc.), smart cards, and flash memory devices (such as EPROMs (Erasable Programmable Read-Only Memories), cards, sticks, or key drives, etc.). The various computer-readable storage media described in the present disclosure may represent one or more devices and / or other machine-readable storage media for storing information. The term "machine-readable storage medium" may include, but is not limited to, wireless channels and various other media capable of storing, containing, and / or carrying instructions and / or data.
[0337] Some embodiments of the present disclosure also provide a computer program product. For example, the computer program product is stored on a non-transitory computer-readable storage medium. The computer program product includes computer program instructions. When the computer program instructions are executed on a computer (e.g., an electronic device), the computer program instructions cause the computer to execute the action recognition method or the model training method of any one of the above embodiments.
[0338] Some embodiments of the present disclosure also provide a computer program. When the computer program is executed on a computer (e.g., an electronic device), the computer program causes the computer to execute the action recognition method or the model training method of any one of the above embodiments.
[0339] The beneficial effects of the above computer-readable storage medium, computer program product, and computer program are the same as those of the action recognition method or the model training method of any one of the above embodiments, and will not be elaborated herein.
[0340] The above are only the specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure who thinks of changes or substitutions should be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. An action recognition method, characterized in that, comprising: Obtaining multiple image frames of the video to be recognized; Segmenting each of the multiple image frames to obtain multiple subsampled images; Determining the sequence features of the video to be recognized according to the multiple subsampled images and the self-attention encoding layer of the pre-trained self-attention model, and determining the target similarity features for multiple action categories according to the sequence features of the video to be recognized; the sequence features include time series features, or the time series features and spatial sequence features; The time series features are used to characterize the similarity between the video to be recognized and the multiple action categories in the time dimension, and the spatial sequence features are used to characterize the similarity between the video to be recognized and the multiple action categories in the spatial dimension; The target similarity features are used to characterize the similarity between the video to be recognized and each action category; Inputting the target similarity features into the classification layer of the pre-trained self-attention model to obtain the probability distribution of the similarity between the video to be recognized and the multiple action categories; The probability distribution includes the probability of similarity between the video to be recognized and each action category in the multiple action categories; Based on the probability distribution of the similarity between the video to be recognized and the multiple action categories, determining the target action category corresponding to the video to be recognized; the probability of similarity between the video to be recognized and the target action category is greater than or equal to a preset threshold.
2. The action recognition method according to claim 1, characterized in that, Determining the time series features of the video to be recognized includes: Determining at least one time sampling sequence from the multiple subsampled images; each time sampling sequence includes the subsampled images at the same position in each image frame; Determining the sub-time series features of each time sampling sequence according to each time sampling sequence and the self-attention encoding layer; the sub-time series features are used to characterize the similarity between each time sampling sequence and the multiple action categories; Determining the time series features of the video to be recognized according to the sub-time series features of the at least one time sampling sequence.
3. The action recognition method according to claim 2, characterized in that, The determining the sub-time series features of each time sampling sequence according to each time sampling sequence and the self-attention encoding layer includes: Determining a plurality of first image input features and class input features; each first image input feature is obtained by position encoding and merging of the image features of the subsampled images included in the first time sampling sequence, and the first time sampling sequence is any one of the at least one time sampling sequence; the class input feature is obtained by position encoding and merging of the class features, and the class features are used to characterize the multiple action categories; Inputting the plurality of first image input features and the class input features into the self-attention encoding layer, and determining the output feature corresponding to the class input feature output by the self-attention encoding layer as the sub-time series feature of the first time sampling sequence.
4. The action recognition method according to any one of claims 1-3, wherein, determining the spatial sequence feature of the video to be recognized includes: determining at least one spatial sampling sequence from the multiple subsampled images; each spatial sampling sequence includes the subsampled images in one image frame; determining the subspace sequence feature of each spatial sampling sequence according to each spatial sampling sequence and the self-attention encoding layer; the subspace sequence feature is used to characterize the similarity between each spatial sampling sequence and the multiple action categories; determining the spatial sequence feature of the video to be recognized according to the subspace sequence features of the at least one spatial sampling sequence.
5. The action recognition method according to claim 4, wherein, the determining at least one spatial sampling sequence from the multiple subsampled images includes: for the first image frame, determining a preset number of target subsampled images located at preset positions from the subsampled images included in the first image frame, and determining the target subsampled images as the spatial sampling sequence corresponding to the first image frame; the first image frame is any one of the multiple image frames.
6. The action recognition method according to claim 4, wherein, the determining the subspace sequence feature of each spatial sampling sequence according to each spatial sampling sequence and the self-attention encoding layer includes: determining a plurality of second image input features and a class input feature; each second image input feature is obtained by position encoding and merging the image features of the subsampled images included in the first spatial sampling sequence, and the first spatial sampling sequence is any one of the at least one spatial sampling sequence; the class input feature is obtained by position encoding and merging the class features, and the class features are used to characterize the multiple action categories; inputting the plurality of second image input features and the class input feature into the self-attention encoding layer, and determining the output feature corresponding to the class input feature output by the self-attention encoding layer as the subspace sequence feature of the first spatial sampling sequence.
7. The action recognition method according to claim 4, wherein, the multiple image frames are obtained based on image preprocessing, and the image preprocessing includes at least one operation of cropping, image enhancement, and scaling.
8. A model training method, wherein, comprising: acquiring a plurality of sample image frames of a sample video and the sample action category where the sample video is located; determining the sample similarity feature of the sample video based on the plurality of sample image frames and the self-attention encoding layer of an initial self-attention model; the self-attention encoding layer is used to determine the sequence feature of the sample video and, according to the sequence feature of the sample video, determine the sample similarity feature for the sample action category, and the sequence feature includes a time sequence feature, or the time sequence feature and a spatial sequence feature; The time series features are used to characterize the similarity between the sample video and the sample action category in the time dimension, and the spatial series features are used to characterize the similarity between the sample video and the sample action category in the spatial dimension; The sample similarity features are used to characterize the similarity between the sample video and the sample action category; Using the sample similarity features as sample features and the sample action category as a label to train a preset neural network to obtain a trained classification layer; the trained classification layer is used to obtain the probability of similarity between the sample video and the sample action category.
9. An action recognition device, characterized in that, it includes an acquisition unit, a processing unit, and a determination unit; The acquisition unit is used to acquire multiple image frames of the video to be recognized; The processing unit is used to segment each of the multiple image frames to obtain multiple sub-sampled images; The determination unit is used to determine the sequence features of the video to be recognized according to the multiple sub-sampled images and the self-attention encoding layer of the pre-trained self-attention model, and determine the target similarity features for multiple action categories according to the sequence features of the video to be recognized; the sequence features include time series features, or the time series features and spatial series features; The time series features are used to characterize the similarity between the video to be recognized and the multiple action categories in the time dimension, and the spatial series features are used to characterize the similarity between the video to be recognized and the multiple action categories in the spatial dimension; The target similarity features are used to characterize the similarity between the video to be recognized and each action category; The determination unit is further used to input the target similarity features into the classification layer of the pre-trained self-attention model to obtain a probability distribution of similarity between the video to be recognized and the multiple action categories; the probability distribution includes the probability of similarity between the video to be recognized and each action category among the multiple action categories; The determination unit is further used to determine the target action category corresponding to the video to be recognized based on the probability distribution of similarity between the video to be recognized and the multiple action categories; the probability of similarity between the video to be recognized and the target action category is greater than or equal to a preset threshold.
10. The action recognition device according to claim 9, characterized in that, The determination unit is specifically used for: Determine at least one time sampling sequence from the multiple sub-sampled images; each time sampling sequence includes the sub-sampled images at the same position in each image frame; Determine the sub-time series features of each time sampling sequence according to each time sampling sequence and the self-attention encoding layer; the sub-time series features are used to characterize the similarity between each time sampling sequence and the multiple action categories; Determine the time series features of the video to be recognized according to the sub-time series features of the at least one time sampling sequence.
11. The action recognition device according to claim 9 or 10, characterized in that, The determination unit is specifically used for: Determine at least one spatial sampling sequence from the plurality of subsampled images; each spatial sampling sequence includes the subsampled images in one image frame; Determine the subspace sequence features of each spatial sampling sequence according to each spatial sampling sequence and the self-attention encoding layer; the subspace sequence features are used to characterize the similarity between each spatial sampling sequence and the plurality of action categories; Determine the spatial sequence features of the video to be recognized according to the subspace sequence features of the at least one spatial sampling sequence.
12. The action recognition device according to claim 11, wherein, The determining unit is specifically configured to: For the first image frame, determine a preset number of target subsampled images located at preset positions from the subsampled images included in the first image frame, and determine the target subsampled images as the spatial sampling sequence corresponding to the first image frame; The first image frame is any one of the plurality of image frames.
13. A model training device, wherein, It includes an acquisition unit, a processing unit and a training unit; The acquisition unit is configured to acquire a plurality of sample image frames of a sample video and the sample action category to which the sample video belongs; The processing unit is configured to segment each sample image frame of the plurality of sample image frames to obtain a plurality of sample subsampled images; The training unit is configured to determine the sample similarity features of the sample video based on the plurality of sample image frames and the self-attention encoding layer of the initial self-attention model; The self-attention encoding layer is configured to determine the sequence features of the sample video, and determine the sample similarity features for the sample action category according to the sequence features of the sample video, where the sequence features include time sequence features, or the time sequence features and spatial sequence features; The time sequence features are used to characterize the similarity between the sample video and the sample action category in the time dimension, and the spatial sequence features are used to characterize the similarity between the sample video and the sample action category in the spatial dimension; The sample similarity features are used to characterize the similarity between the sample video and the sample action category; The training unit is further configured to use the sample similarity features as sample features and the sample action category as a label to train a preset neural network to obtain a trained classification layer; the trained classification layer is used to obtain the probability that the sample video is similar to the sample action category.
14. An electronic device, wherein, It includes: A processor and a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions so that the electronic device implements the action recognition method according to any one of claims 1-7, or the model training method according to claim 8.
15. A computer-readable storage medium, wherein, When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the action recognition method according to any one of claims 1-7, or the model training method according to claim 8.
Citation Information
Patent Citations
Human motion recognition method and device
CN108985259A
Method for correcting training action in real time based on deep learning and related equipment thereof
CN113743362A