Action recognition method and system, intelligent terminal and storage medium
By performing image encoding and feature decoupling of video frames, separating feature information and redundant information, the problem of low accuracy in action recognition in complex scenarios is solved, and higher accuracy in action recognition is achieved.
Patent Information
- Application Number
- CN202510103427.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is prone to information interference when recognizing actions in complex scenarios, resulting in low accuracy of action recognition.
By image encoding of the video frame to be processed, the first image feature is obtained, and feature decoupling processing is performed to separate feature information and redundant information, and the second image feature is obtained. Then, based on the second image feature and the text feature vector, the target action category is determined from the action category label.
By filtering out redundant information, reduce information interference and improve the accuracy of action recognition, especially in complex scenarios.
Smart Images

Figure CN119992657A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to an action recognition method, system, intelligent terminal and storage medium. Background Art
[0002] Video action recognition refers to the automatic recognition and understanding of human actions and behaviors from videos through computer vision technology. Action recognition technology plays an important role in many practical applications, especially in security, health, sports, smart driving and other fields.
[0003] In the prior art, when performing action recognition on a video, the image features obtained by extracting features from the video are usually directly used to match the text features corresponding to the action category, so as to determine the corresponding action category. The problem with the prior art is that the image features obtained by extracting features from the video contain a large amount of information, especially in complex scenes, which may contain a large amount of scene information. When the extracted image features are directly used for matching to identify the action category, it is easy to be interfered by information and is not conducive to improving the accuracy of action recognition.
[0004] Therefore, relevant technologies still need to be improved and developed. Summary of the invention
[0005] The main purpose of the present application is to provide an action recognition method, system, intelligent terminal and storage medium, aiming to solve the technical problem in the related art that the image features obtained by extracting features from a video contain a large amount of information, and when the extracted image features are directly used for matching to identify the action category, it is easy to be interfered by the information and is not conducive to improving the accuracy of action recognition.
[0006] In order to achieve the above-mentioned object, the first aspect of the present application provides a motion recognition method, wherein the motion recognition method comprises:
[0007] Obtain a video frame to be processed and at least one action category label;
[0008] Performing image encoding on the video frame to be processed and obtaining a first image feature;
[0009] Performing feature decoupling processing on the first image feature to separate feature information and redundant information in the first image feature, and obtaining a second image feature corresponding to the feature information in the first image feature, wherein the feature information includes information related to the action category, and the redundant information includes information unrelated to the action category;
[0010] Perform text encoding on the above action category labels and obtain text feature vectors;
[0011] According to the second image feature and the text feature vector, a target action category corresponding to the to-be-processed video frame is determined from the action category label.
[0012] Optionally, the obtaining of the video frame to be processed includes:
[0013] Get the video to be processed;
[0014] According to a preset video frame extraction rule, at least one video frame to be processed is extracted from the video to be processed.
[0015] Optionally, the performing image encoding on the video frame to be processed and obtaining the first image feature includes:
[0016] Inputting the above-mentioned video frame to be processed into a preset image encoder, wherein the above-mentioned image encoder is provided with a feature extraction module and a first feature decoupling module;
[0017] Extracting image features from the video frame to be processed by the feature extraction module to obtain candidate image features;
[0018] The first feature decoupling module is used to perform feature decoupling processing on the candidate image features to separate feature information and redundant information in the candidate image features, and obtain the first image features corresponding to the feature information in the candidate image features.
[0019] Optionally, the first feature decoupling module includes an instance normalization submodule and a channel mask submodule;
[0020] The above-mentioned first feature decoupling module performs feature decoupling processing on the above-mentioned candidate image features to separate feature information and redundant information in the above-mentioned candidate image features, and obtains the first image features corresponding to the feature information in the above-mentioned candidate image features, including:
[0021] Performing instance normalization on the candidate image features through the instance normalization submodule to obtain normalized features;
[0022] Obtaining residual features according to the candidate image features and the normalized features;
[0023] Performing a mask operation on the residual feature through the channel mask submodule to separate and obtain feature information and redundant information corresponding to the residual feature;
[0024] The feature information corresponding to the residual feature and the normalized feature are fused to obtain a first image feature.
[0025] Optionally, the performing feature decoupling processing on the first image feature to separate feature information and redundant information in the first image feature and obtaining a second image feature corresponding to the feature information in the first image feature includes:
[0026] Inputting the first image feature into a second feature decoupling module, so as to perform feature decoupling processing on the first image feature through the second feature decoupling module to obtain a second image feature;
[0027] Among them, the above-mentioned second characteristic decoupling module has the same structure as the above-mentioned first characteristic decoupling module.
[0028] Optionally, the step of determining the target action category corresponding to the to-be-processed video frame from the action category label according to the second image feature and the text feature vector includes:
[0029] Performing global pooling processing on the second image feature to obtain a global feature vector;
[0030] Input the above global feature vector into a preset anchor self-attention module, so as to learn the context dependency between distant pixels and enhance the central features through the above anchor self-attention module, and obtain a processed target global feature vector;
[0031] According to the target global feature vector and the text feature vector, a target action category corresponding to the to-be-processed video frame is determined from the action category labels.
[0032] Optionally, determining the target action category corresponding to the to-be-processed video frame from the action category label according to the target global feature vector and the text feature vector includes:
[0033] Calculate the similarity between the target global feature vector and the text feature vector;
[0034] The target action category corresponding to the above-mentioned video frame to be processed is determined according to the similarity calculation result.
[0035] A second aspect of the present application provides a motion recognition system, wherein the motion recognition system comprises:
[0036] A data acquisition module, used to acquire a video frame to be processed and at least one action category label;
[0037] An image encoding module, used to perform image encoding on the above-mentioned video frame to be processed and obtain a first image feature;
[0038] a feature decoupling module, configured to perform feature decoupling processing on the first image feature to separate feature information and redundant information in the first image feature, and obtain a second image feature corresponding to the feature information in the first image feature, wherein the feature information includes information related to the action category, and the redundant information includes information unrelated to the action category;
[0039] A text encoding module, used for performing text encoding on the above action category labels and obtaining a text feature vector;
[0040] The action recognition module is used to determine the target action category corresponding to the video frame to be processed from the action category label according to the second image feature and the text feature vector.
[0041] The third aspect of the present application provides a smart terminal, which includes a memory, a processor, and a motion recognition program stored in the memory and executable on the processor. When the motion recognition program is executed by the processor, any step of the motion recognition method is implemented.
[0042] A fourth aspect of the present application provides a computer-readable storage medium, on which a motion recognition program is stored. When the motion recognition program is executed by a processor, any step of the motion recognition method is implemented.
[0043] As can be seen from the above, in the scheme of the present application, a video frame to be processed and at least one action category label are obtained; image encoding is performed on the above-mentioned video frame to be processed and a first image feature is obtained; feature decoupling processing is performed on the above-mentioned first image feature to separate the feature information and redundant information in the above-mentioned first image feature, and a second image feature corresponding to the feature information in the above-mentioned first image feature is obtained, wherein the above-mentioned feature information includes information related to the action category, and the above-mentioned redundant information includes information unrelated to the action category; text encoding is performed on the above-mentioned action category label and a text feature vector is obtained; based on the above-mentioned second image feature and the above-mentioned text feature vector, the target action category corresponding to the above-mentioned video frame to be processed is determined from the above-mentioned action category label.
[0044] Compared with the prior art, in the scheme corresponding to the action recognition method provided by the present application, after the first image feature is obtained by image encoding the video frame to be processed, the first image feature is subjected to feature decoupling processing to separate the feature information and redundant information therein, and then the second image feature corresponding to the feature information is obtained, and the second image feature corresponding to the feature information is matched to determine the corresponding target action category. In this way, the interference of redundant information on the action recognition process is filtered out, which is conducive to improving the accuracy of action recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0046] Figure 1 It is an action recognition heat map based on the prior art provided in the embodiment of the present application;
[0047] Figure 2 It is a flowchart of an action recognition method provided by an embodiment of the present application;
[0048] Figure 3 This is a schematic diagram of a specific process of action recognition provided by an embodiment of the present application;
[0049] Figure 4 It is a schematic diagram of a specific data processing flow corresponding to a feature decoupling module provided in an embodiment of the present application;
[0050] Figure 5 It is a structural diagram of an anchor point self-attention module provided in an embodiment of the present application;
[0051] Figure 6 It is a schematic diagram of the components of a motion recognition system provided by an embodiment of the present application;
[0052] Figure 7 It is a block diagram of the internal structure principle of a smart terminal provided in an embodiment of the present application. DETAILED DESCRIPTION
[0053] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0054] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0055] It should also be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, the singular forms "a", "an" and "the" are intended to include plural forms unless the context clearly indicates otherwise.
[0056] It should be further understood that the term “and / or” used in the specification and appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0057] As used in this specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to being classified into," depending on the context. Similarly, the phrase "if it is determined" or "if classified into [described condition or event]" may be interpreted as meaning "upon determination" or "in response to determining" or "upon classification into [described condition or event]" or "in response to being classified into [described condition or event]," depending on the context.
[0058] The following is a clear and complete description of the technical solutions in the embodiments of the present application in conjunction with the drawings of the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0059] In the following description, many specific details are set forth to facilitate a full understanding of the present application, but the present application may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present application. Therefore, the present application is not limited to the specific embodiments disclosed below.
[0060] Action recognition is an important research direction in the field of computer vision, which aims to identify and understand human actions and behaviors from video data. Unlike static image recognition, video action recognition not only needs to analyze the spatial features in the image, but also needs to process the changes in the time dimension to capture the temporal dynamics of the action. With the continuous development of technologies such as deep learning, convolutional neural networks, recurrent neural networks, spatiotemporal convolutional networks, and self-attention mechanisms, the performance and application areas of video action recognition have been significantly improved. With the rapid rise of multimodal and large model technologies, traditional single-modal video understanding tasks have also begun to shift to large models and multimodal. Large models and multimodal models have good transferability and generalization. In an application scenario, the video action recognition model uses the Contrastive Language-Image Pre-Training (CLIP) model as the backbone network for training. CLIP is a multimodal model that aims to provide a model that can understand images and texts at the same time by jointly training visual and language information. It is trained through large-scale image-text pairs, enabling the model to perform image understanding tasks without special supervision.
[0061] At present, there are few studies on video tasks with complex scenes. Most of them are optimized and improved in the target detection task, and also optimize the image field. The existing target detection task is also only for the recognition task of small targets in large scenes in the case of pictures.
[0062] In one application scenario, based on the video action recognition method ActionCLIP, the CLIP model is extended to video, combining the multimodal learning capabilities of images, videos, and text. This method breaks through the traditional video action recognition method and can recognize actions in videos through natural language descriptions. In another application scenario, CLIP can also be applied to the field of video action recognition, using a teacher model to optimize CLIP recognition tasks in the video field. In another application scenario, data augmentation can also be used to cut and splice single images of the data set into a complete image to enhance the recognition of small targets in the image.
[0063] However, there is little research on multi-scale issues in video tasks in related technologies, such as style changes (lighting, color contrast, etc.), small subject targets, and crowded crowds. Existing video action recognition tasks have relatively good performance in a few relatively simple public datasets, but their accuracy is greatly reduced in complex scenes, which are the scenes that can be encountered in actual applications.
[0064] Figure 1This is an action recognition heat map based on the prior art provided in the embodiment of the present application. Figure 1 The corresponding prompt words used are: two men is fighting. It should be noted that the prompt words used in the embodiment of the present application are in English, but are not specifically limited. Figure 1 As shown in the figure, when performing action recognition in complex scenes based on the existing technology, the model pays more attention to the areas that are not related to the prompt words, that is, the existing video action recognition model has poor recognition effect on complex scenes. Specifically, when directly using the extracted image features for matching to identify the action category in the case of a large amount of scene information, it is easy to be interfered by information and is not conducive to improving the accuracy of action recognition.
[0065] In order to solve at least one of the above-mentioned technical problems, in the solution of the present application, a video frame to be processed and at least one action category label are obtained; image encoding is performed on the above-mentioned video frame to be processed and a first image feature is obtained; feature decoupling processing is performed on the above-mentioned first image feature to separate feature information and redundant information in the above-mentioned first image feature, and a second image feature corresponding to the feature information in the above-mentioned first image feature is obtained, wherein the above-mentioned feature information includes information related to the action category, and the above-mentioned redundant information includes information unrelated to the action category; text encoding is performed on the above-mentioned action category label and a text feature vector is obtained; based on the above-mentioned second image feature and the above-mentioned text feature vector, the target action category corresponding to the above-mentioned video frame to be processed is determined from the above-mentioned action category label.
[0066] Compared with the prior art, in the scheme corresponding to the action recognition method provided by the present application, after the first image feature is obtained by image encoding the video frame to be processed, the first image feature is subjected to feature decoupling processing to separate the feature information and redundant information therein, and then the second image feature corresponding to the feature information is obtained, and the second image feature corresponding to the feature information is matched to determine the corresponding target action category. In this way, the interference of redundant information on the action recognition process is filtered out, which is conducive to improving the accuracy of action recognition.
[0067] Specifically, in order to solve the problem that action recognition is difficult to handle complex scenes in practical applications, such as when pedestrians account for a very small part of the entire screen under a panoramic camera, there are many people in the screen, insufficient light at night, etc. In the embodiment of the present application, a feature decoupling method is used to enhance the model's ability to recognize specific actions, so as to enhance its generalization performance and attention to specific areas. Furthermore, a contextual anchor self-attention method is added during the training process to establish contextual interdependencies between pixels to enhance the network's attention to specific areas. In this way, a decoupling and enhanced attention to key areas method is proposed to enhance the model's recognition of specific actions in complex scenes, which can increase the accuracy of action recognition in real complex scenes.
[0068] like Figure 2 As shown, the embodiment of the present application provides a method for action recognition. Specifically, the method includes the following steps:
[0069] Step S100, obtaining a video frame to be processed and at least one action category label.
[0070] Among them, the video frame to be processed is a video frame that needs to be action recognized, and the above-mentioned action category label is a label corresponding to the action category used when performing action recognition. In the embodiment of the present application, action recognition is performed on the video frame to be processed to determine which action category label matches the action of the person in the video frame to be processed, thereby determining the action corresponding to the person. It should be noted that the specific action category label used can be set and adjusted according to actual needs. For example, it can include fighting, hugging, raising the head, etc., which are not specifically limited here.
[0071] Specifically, the above-mentioned obtaining of the video frame to be processed includes:
[0072] Get the video to be processed;
[0073] According to a preset video frame extraction rule, at least one video frame to be processed is extracted from the video to be processed.
[0074] Among them, the above-mentioned video to be processed is a video that needs to perform action recognition. In the embodiment of the present application, some video frames to be processed are extracted from the video to be processed for action recognition. The above-mentioned video extraction rules can be set and adjusted according to actual needs. For example, some video frames to be processed can be extracted from the video to be processed based on a preset sampling period, or random sampling can be performed to determine the video frames to be processed. Other methods can also be used to determine the video frames to be processed, and the number of video frames to be processed can be determined and adjusted according to actual needs, which is not specifically limited here.
[0075] In an application scenario, the steps corresponding to the action recognition method provided in the embodiment of the present application are performed based on a pre-trained action recognition model, and the video frames to be processed can be extracted through the model. The complete video to be processed is input into the model, and the model randomly extracts the corresponding frames to be processed based on the preset video frame extraction rules. For example, if it is preset that 16 frames need to be extracted, the video to be processed is cut into 16 segments, and one video frame is extracted from each segment as the video frame to be processed.
[0076] Step S200: performing image encoding on the video frame to be processed and obtaining a first image feature.
[0077] Specifically, a preset image encoder may be used to perform image encoding on the above-mentioned video frames to be processed. The image encoder may be set and trained according to actual needs. For example, the image encoder in the CLIP model may be used for processing, but this is not a specific limitation.
[0078] In the embodiment of the present application, the above-mentioned image encoding of the above-mentioned video frame to be processed and obtaining the first image feature includes:
[0079] Inputting the above-mentioned video frame to be processed into a preset image encoder, wherein the above-mentioned image encoder is provided with a feature extraction module and a first feature decoupling module;
[0080] Extracting image features from the video frame to be processed by the feature extraction module to obtain candidate image features;
[0081] The first feature decoupling module is used to perform feature decoupling processing on the candidate image features to separate feature information and redundant information in the candidate image features, and obtain the first image features corresponding to the feature information in the candidate image features.
[0082] Specifically, in the embodiment of the present application, the image encoder of the CLIP model is fine-tuned, and a first feature decoupling module is added to the image encoder, so that the candidate image features are subjected to feature decoupling processing based on the first feature decoupling module, and feature information related to the action category and redundant information not related to the action category are separated, and the part corresponding to the feature information is used as the first image feature. The redundant information may include background information, lighting information, etc.
[0083] It should be noted that the above-mentioned candidate features and the first image features may be in the form of feature maps or feature vectors, which is not specifically limited here.
[0084] Furthermore, the first feature decoupling module includes an instance normalization submodule and a channel mask submodule;
[0085] The above-mentioned first feature decoupling module performs feature decoupling processing on the above-mentioned candidate image features to separate feature information and redundant information in the above-mentioned candidate image features, and obtains the first image features corresponding to the feature information in the above-mentioned candidate image features, including:
[0086] Performing instance normalization on the candidate image features through the instance normalization submodule to obtain normalized features;
[0087] Obtaining residual features according to the candidate image features and the normalized features;
[0088] Performing a mask operation on the residual feature through the channel mask submodule to separate and obtain feature information and redundant information corresponding to the residual feature;
[0089] The feature information corresponding to the residual feature and the normalized feature are fused to obtain a first image feature.
[0090] In the embodiment of the present application, the normalized feature is subtracted from the candidate image feature to obtain the corresponding residual feature. It should be noted that the structural setting of the first feature decoupling module is only for illustration and is not a specific limitation. In actual use, other division methods can also be used.
[0091] Step S300, performing feature decoupling processing on the above-mentioned first image feature to separate feature information and redundant information in the above-mentioned first image feature, and obtaining a second image feature corresponding to the feature information in the above-mentioned first image feature, wherein the above-mentioned feature information includes information related to the action category, and the above-mentioned redundant information includes information unrelated to the action category.
[0092] Specifically, the above-mentioned feature decoupling processing is performed on the above-mentioned first image feature to separate the feature information and redundant information in the above-mentioned first image feature, and obtain the second image feature corresponding to the feature information in the above-mentioned first image feature, including:
[0093] Inputting the first image feature into a second feature decoupling module, so as to perform feature decoupling processing on the first image feature through the second feature decoupling module to obtain a second image feature;
[0094] Among them, the above-mentioned second characteristic decoupling module has the same structure as the above-mentioned first characteristic decoupling module.
[0095] In the embodiment of the present application, the first feature decoupling module and the second feature decoupling module have the same structure and function. In the embodiment of the present application, feature decoupling is first performed based on the first feature decoupling module during the image encoding process to separate feature information from redundant information. Then, the first image feature is further processed based on the second feature decoupling module to further enhance the feature information, thereby improving the accuracy of the subsequent action recognition process.
[0096] Step S400: text encoding the action category label and obtaining a text feature vector.
[0097] In the embodiments of the present application, text encoding is performed using a text encoder based on the CLIP model. In actual use, text encoding may also be performed using other methods, which are not specifically limited here.
[0098] Step S500: determining the target action category corresponding to the to-be-processed video frame from the action category label according to the second image feature and the text feature vector.
[0099] Specifically, determining the target action category corresponding to the to-be-processed video frame from the action category label according to the second image feature and the text feature vector includes:
[0100] Performing global pooling processing on the second image feature to obtain a global feature vector;
[0101] Input the above global feature vector into a preset anchor self-attention module, so as to learn the context dependency between distant pixels and enhance the central features through the above anchor self-attention module, and obtain a processed target global feature vector;
[0102] According to the target global feature vector and the text feature vector, a target action category corresponding to the to-be-processed video frame is determined from the action category labels.
[0103] In this way, the attention to specific areas is enhanced through the anchor self-attention module, further improving the accuracy of action recognition.
[0104] In the embodiment of the present application, the target action category corresponding to the to-be-processed video frame is determined from the action category label according to the target global feature vector and the text feature vector, including:
[0105] Calculate the similarity between the target global feature vector and the text feature vector;
[0106] The target action category corresponding to the above-mentioned video frame to be processed is determined according to the similarity calculation result.
[0107] In an application scenario, for each action category label, its corresponding text feature vector is determined respectively, and then the cosine similarity between each text feature vector and the target global feature vector is calculated respectively, and the corresponding text feature vector with the highest cosine similarity is determined, and the action category label corresponding to the text feature vector with the highest cosine similarity is used as the target action category.
[0108] As can be seen from the above, in the scheme corresponding to the action recognition method provided by the embodiment of the present application, after the image encoding of the video frame to be processed is performed to obtain the first image feature, the first image feature is subjected to feature decoupling processing to separate the feature information and redundant information therein, and then the second image feature corresponding to the feature information is obtained, and matching is performed based on the second image feature corresponding to the feature information to determine the corresponding target action category. In this way, filtering out the interference of redundant information on the action recognition process is conducive to improving the accuracy of action recognition.
[0109] In the embodiment of the present application, the above-mentioned action recognition method is also described in detail based on a specific application scenario. Specifically, in the embodiment of the present application, a feature decoupling module is added in the CLIP image encoder and after the last output layer of the CLIP image encoder to remove some redundant information, such as useless background, lighting and other information, and then restore the information related to the feature, effectively separating the features related to and irrelevant to the specific category. After pooling, an anchor self-attention module is added to enhance the focus on specific areas.
[0110] It should be noted that in the embodiments of the present application, a specific description is made from the perspective of model training, that is, a specific description is given of the data processing process during the model training process. When the model is used to actually perform the task of action recognition, the corresponding data processing process can also refer to the data processing process during the training process.
[0111] Figure 3 is a specific flow chart of an action recognition provided by an embodiment of the present application, such as Figure 3As shown, in an embodiment of the present application, the video frame to be processed is input into the image encoder, the image encoder is used to process the frame to obtain the first image feature, and then the first image feature is processed by the feature decoupling module to obtain the second image feature. It should be noted that a feature decoupling module is also provided in the image encoder. Through the processing of the two feature decoupling modules, the feature decoupling effect is further improved, thereby improving the accuracy of subsequent action recognition. Furthermore, the second image feature is subjected to time pooling processing to obtain the pooled global feature, and is further processed by the anchor attention module to obtain the processed target global feature. On the other hand, during the training process, the text part corresponding to the model is jointly constituted by the learnable context vector (i.e., the prompt word) and the label category, and the prompt word is trained and learned during the training process. For example, the entire prompt word is aphoto of is{class}, where class represents the corresponding action category label, and the rest is the learnable prompt word part. It should be noted that, Figure 3 Ctx in is used to represent the corresponding prompt word. It should be further explained that the learnable prompt word part is fixed when the training is completed. When using the trained model to perform action recognition tasks, you only need to enter the corresponding action category label. The text part is encoded by the text encoder to obtain the corresponding text features, and finally the cosine similarity between the image and text features is calculated to achieve action recognition.
[0112] In the embodiment of the present application, action recognition in complex scenes is performed based on the CLIP model. Specifically, CLIP is selected as the backbone network, its image encoder and text encoder are fine-tuned, and a feature decoupling module (FD, Feature Decoupling) is added to the CLIP image encoder to separate relevant and irrelevant features and extract image features. The text encoder mainly learns text context prompts and extracts text features.
[0113] Specifically, given a video frame x to be processed with T frames and each frame has a spatial dimension of H×W i ∈ Input it to the CLIP image encoder (i.e. visual encoder) f v In the example, the image encoder provided with the first feature decoupling module (FD) is used for processing to obtain the first image feature, which is expressed as follows:
[0114]
[0115] in, represents the first image feature, FD(·) represents the processing operation corresponding to the first feature decoupling module in the image encoder, and f v (·) represents the processing operation of the image encoder.
[0116] Let the text action embedding (i.e. action category label) be t j , using the text encoder f t Get text features, the formula is expressed as:
[0117]
[0118] in, represents the text feature vector, f t (·) represents the processing of the text encoder, R D Represents the length of the text feature vector.
[0119] It should be noted that in the embodiment of the present application, a first feature decoupling module and a second feature decoupling module are used, but the structures and specific functions of the two feature decoupling modules are the same.
[0120] Figure 4 This is a schematic diagram of a specific data processing flow corresponding to a feature decoupling module provided in an embodiment of the present application. It should be noted that: Figure 4 In , if the input data used is the candidate image features obtained by processing the middle layer of the image encoder (i.e., the feature extraction module in the image encoder), then Figure 4 This corresponds to the processing of the first feature decoupling module, and finally the first image feature is obtained by adding the category-related feature and the normalized feature. If the input data used is the first image feature, then Figure 4 This corresponds to the processing process of the second feature decoupling module, and finally the second image feature is obtained by adding the category-related feature and the normalized feature.
[0121] In the embodiment of the present application, the specific processing process of the first feature decoupling module is taken as an example for specific description. Specifically, the first feature decoupling module includes an instance normalization submodule (IN) and a channel mask submodule M. Assume that a frame of an image passes through the middle layer of the image encoder (i.e., the feature extraction module in the image encoder) and outputs a feature map (i.e., candidate image feature) f∈R h×w×c , where h and w are the length and width of the feature map, respectively, and c is the number of channels of the feature map. The candidate image feature is used as the input of the first feature decoupling module, as shown in Figure 4 As shown, it is first processed by an instance normalization submodule to reduce the difference in input features. The implementation formula is as follows:
[0122]
[0123] in, represents the normalized features obtained after the instance normalization (IN) operation, u and σ represent the mean and standard deviation of each channel of each instance feature calculated in the spatial dimension, τ and ∈ are the learnable scaling parameters and translation parameters, respectively.
[0124] From the calculation method of IN, we can know that IN can maintain the independence between the features of each instance and reduce the mutual influence between different modes. Instance normalization (IN) reduces style differences and improves generalization ability, but after it is instantiated, some features that are important for category discrimination may be lost, so the lost information needs to be recovered. This can be done by Extracted from the network to restore features related to the label category to the network, residual features The calculation formula is defined as:
[0125]
[0126] Furthermore, through the channel attention vector Perform mask operation and transform Decoupling into feature information and redundant information The details are shown in the following formulas (5) to (7):
[0127]
[0128] in, It means that Perform a global average pooling operation, δ and g represent the ReLU activation function and the Sigmoid activation function, respectively, to convert linear features into nonlinearities, and m is the channel attention vector, which is obtained by the global average pooling layer and two fully connected layers. Represents the weight matrix corresponding to the Sigmoid activation function, which is a learned parameter and is usually optimized during training. Its function is to linearly transform the input features. Represents the weight matrix corresponding to the ReLU activation function, which is usually optimized during the training process and is used to linearly transform the input features. The values of both can be automatically adjusted during the training process or manually set according to actual needs.
[0129] Then the obtained feature information and redundant information and Perform addition fusion to obtain features related to the category (i.e., the first image feature) Information not related to category Finally, a double recovery loss constraint is imposed to promote the effective separation of features, making the category-related features more obvious and the category-irrelevant information smoother and less prominent.
[0130] The implementation formula is as follows:
[0131]
[0132] For category-related features Calculate the corresponding cross entropy loss, the formula is as follows:
[0133]
[0134] Among them, L classification Represents the loss function of category-related features, that is, classification loss, N represents the number of samples in the data set, ω j represents the weight vector of the jth class, ω i represents the weight vector of the i-th category, and g represents the total number of categories.
[0135] Category-irrelevant information Adding regularization loss encourages redundant information to have a smaller norm, and its formula is as follows:
[0136]
[0137] Among them, L redundancy Represents the class-independent feature loss function. Represents the category-independent information corresponding to the i-th sample.
[0138] The total loss of the FD module is:
[0139] L total =L classification +λL redundancy (12);
[0140] Among them, λ is a hyperparameter used to balance the two losses, and its value can be set and adjusted according to actual needs.
[0141] After the above processing, the output first image features are obtained, and the first image features are input into the second feature decoupling module, and then processed by a feature decoupling module (FD) to further separate the relevant and irrelevant features to obtain the processed second image features. After the second image features after the FD module are output, a temporal pooling operation is performed on the frame features of this batch to obtain the second image features after temporal pooling. The processing process is shown in the following formula (13) and formula (14):
[0142]
[0143] Among them, F(b,t,h,w,c) represents a batch of frames, b is the batch-size, which refers to the number of samples used in this training, t is the current time frame number, h is the feature map height, w is the feature map width, c is the number of channels, α t Represents the weight of the tth frame, and its value can be set and adjusted according to actual needs. T represents the total number of frames. represents the second image feature obtained after being processed by the second feature decoupling module, represents the second image feature after time pooling, F pool (·) represents the temporal pooling operation.
[0144] Furthermore, the second image features after the time pooling process are further processed. Specifically, the visual features of multiple frames are globally pooled, and their temporal information is aggregated into a feature vector containing the whole world. The pooled information is sent to the anchor self-attention module (CAF, Context Anchor Focus) to learn the contextual dependency between distant pixels and enhance the central features.
[0145] Figure 5 is a schematic diagram of the structure of an anchor self-attention module provided in an embodiment of the present application, such as Figure 5 As shown in the figure, the anchor self-attention module includes two 1×1 convolutions and two depth-separable convolutions. The depth-separable convolution mainly extracts features in the horizontal and vertical directions and can capture local information. The activation function Sigmoid is used to generate weights and limit their range to [0, 1] in order to perform weighted operations on features. After CAF, the video-level features (i.e., the target global feature vector) are obtained. As shown in the following formula:
[0146]
[0147] Furthermore, the target global feature vector v i With text feature vector Perform similarity calculation, specifically, perform cosine similarity calculation to determine the category of the correct action, as shown in the following formula:
[0148]
[0149] in, Represents the calculation of the cosine similarity loss function during training, Represents the calculated similarity. In this application, cosine similarity is used. τ is a parameter used to control the smoothness of the similarity distribution. The specific value can be determined and adjusted according to actual needs. A larger τ will make the distribution smoother; a smaller τ will amplify the contribution of high similarity samples, making the model pay more attention to high similarity samples.
[0150] It should be noted that Represents the calculation of the cosine similarity loss function during the training process. It is mainly used to measure the differences between sample feature representations, maximize the similarity between samples of the same type and minimize the similarity between samples of different types. This loss is used to optimize model parameters in back propagation, thereby improving model performance. In the actual application stage, it is no longer necessary to calculate
[0151] Thus, in the embodiment of the present application, the visual language model network is combined to perform the video action recognition task, which can solve the problem of inaccurate model recognition of action categories in complex scenes. The feature decoupling module is mainly used to eliminate redundant information irrelevant to the feature, maximize the information related to the feature, and the decoupled features have higher discriminability than the previous features. The possibility of category features after decoupling is clearer than before decoupling, which can reduce the ambiguity of the sample. Furthermore, the anchor self-attention module pays attention to different parts of the input feature map, enhancing the network's attention to specific areas, thereby further improving the accuracy of action recognition.
[0152] Using the visual language model CLIP, we can inherit its zero-sample performance. Adding a feature decoupling module can increase recognition in complex scenes. At the same time, the instance normalization of feature decoupling removes some irrelevant redundant information. This operation also increases the generalization of the model and enhances its migration ability. At the same time, adding an anchor attention module further optimizes its attention to feature areas in complex environments.
[0153] Furthermore, the applicability and feasibility of the embodiment of the present application in complex scenarios were verified by testing in an actual surveillance camera environment. The experimental results show that the embodiment of the present application can not only operate stably under challenging conditions such as dynamic and illumination changes, but also effectively improve task performance. Compared with the simple use of the CLIP model, the embodiment of the present application has achieved significant improvements in accuracy, robustness, and model adaptability, providing an efficient and reliable solution for target detection or recognition tasks in surveillance scenarios.
[0154] like Figure 6 As shown in , corresponding to the above-mentioned action recognition method, the embodiment of the present application further provides an action recognition system, and the above-mentioned action recognition system includes:
[0155] The data acquisition module 610 is used to acquire a video frame to be processed and at least one action category label;
[0156] An image encoding module 620 is used to perform image encoding on the above-mentioned video frame to be processed and obtain a first image feature;
[0157] A feature decoupling module 630 is used to perform feature decoupling processing on the first image feature to separate feature information and redundant information in the first image feature, and obtain a second image feature corresponding to the feature information in the first image feature, wherein the feature information includes information related to the action category, and the redundant information includes information unrelated to the action category;
[0158] A text encoding module 640, used to perform text encoding on the action category label and obtain a text feature vector;
[0159] The action recognition module 650 is used to determine the target action category corresponding to the video frame to be processed from the action category label according to the second image feature and the text feature vector.
[0160] In this way, after the processed video frame is encoded to obtain the first image feature, the first image feature is subjected to feature decoupling processing to separate the feature information and redundant information therein, thereby obtaining the second image feature corresponding to the feature information, and matching is performed based on the second image feature corresponding to the feature information to determine the corresponding target action category. In this way, the interference of redundant information on the action recognition process is filtered out, which is conducive to improving the accuracy of action recognition.
[0161] It should be noted that the specific structure and implementation of the above-mentioned action recognition system and its various modules or units can refer to the corresponding description in the above-mentioned method embodiment, and will not be repeated here.
[0162] It should be noted that the division method of the various modules of the above-mentioned action recognition system is not unique and is not specifically limited here.
[0163] Based on the above embodiments, the present application also provides a smart terminal, whose principle block diagram can be as follows: Figure 7As shown. The above-mentioned intelligent terminal includes a processor, a memory, a network interface and a display screen connected through a system bus. Among them, the processor of the intelligent terminal is used to provide computing and control capabilities. The memory of the intelligent terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and an action recognition program. The internal memory provides an environment for the operation of the operating system and the action recognition program in the non-volatile storage medium. The network interface of the intelligent terminal is used to communicate with an external terminal through a network connection. When the action recognition program is executed by the processor, the steps of any one of the above-mentioned action recognition methods are implemented. The display screen of the intelligent terminal can be a liquid crystal display screen or an electronic ink display screen.
[0164] Those skilled in the art will understand that Figure 7 The principle block diagram shown in the figure is only a block diagram of a partial structure related to the solution of the present application, and does not constitute a limitation on the smart terminal to which the solution of the present application is applied. The specific smart terminal may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0165] In one embodiment, a smart terminal is provided, which includes a memory, a processor, and a motion recognition program stored in the memory and executable on the processor. When the motion recognition program is executed by the processor, the steps of any one of the motion recognition methods provided in the embodiments of the present application are implemented.
[0166] An embodiment of the present application also provides a computer-readable storage medium, on which a motion recognition program is stored. When the motion recognition program is executed by a processor, the steps of any one of the motion recognition methods provided in the embodiment of the present application are implemented.
[0167] It should be understood that the serial numbers of the steps in the above embodiments do not imply a sequence of execution. The execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0168] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the above-mentioned device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0169] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0170] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0171] In the embodiments provided in the present application, it should be understood that the disclosed system / terminal device and method can be implemented in other ways. For example, the system / terminal device embodiments described above are only schematic, for example, the division of the above modules or units is only a logical function division, and in actual implementation, other division methods can be used, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0172] If the above-mentioned integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The above-mentioned computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the above-mentioned computer program includes computer program code, and the above-mentioned computer program code can be in source code form, object code form, executable file or some intermediate form. The above-mentioned computer-readable medium may include: any entity or device capable of carrying the above-mentioned computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal and software distribution medium, etc. It should be noted that the content contained in the above-mentioned computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction.
[0173] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. A method for motion recognition, characterized in that: The method comprises: Obtain a video frame to be processed and at least one action category label; Performing image encoding on the video frame to be processed and obtaining a first image feature; Performing feature decoupling processing on the first image feature to separate feature information and redundant information in the first image feature, and obtaining a second image feature corresponding to the feature information in the first image feature, wherein the feature information includes information related to the action category, and the redundant information includes information unrelated to the action category; Performing text encoding on the action category label and obtaining a text feature vector; According to the second image feature and the text feature vector, a target action category corresponding to the to-be-processed video frame is determined from the action category label.
2. The action recognition method according to claim 1, characterized in that: The step of obtaining a video frame to be processed includes: Get the video to be processed; At least one video frame to be processed is extracted from the video to be processed according to a preset video frame extraction rule.
3. The action recognition method according to claim 1, characterized in that: The performing image encoding on the video frame to be processed and obtaining the first image feature comprises: Inputting the to-be-processed video frame into a preset image encoder, wherein the image encoder is provided with a feature extraction module and a first feature decoupling module; Extracting image features from the video frame to be processed by the feature extraction module to obtain candidate image features; The first feature decoupling module performs feature decoupling processing on the candidate image features to separate feature information and redundant information in the candidate image features, and obtains a first image feature corresponding to the feature information in the candidate image features.
4. The action recognition method according to claim 3, characterized in that: The first feature decoupling module includes an instance normalization submodule and a channel mask submodule; The step of performing feature decoupling processing on the candidate image feature by the first feature decoupling module to separate feature information and redundant information in the candidate image feature and obtaining a first image feature corresponding to the feature information in the candidate image feature includes: Performing instance normalization on the candidate image features through the instance normalization submodule to obtain normalized features; Obtaining residual features according to the candidate image features and the normalized features; Performing a mask operation on the residual feature through the channel mask submodule to separate and obtain feature information and redundant information corresponding to the residual feature; The feature information corresponding to the residual feature and the normalized feature are fused to obtain a first image feature.
5. The action recognition method according to claim 4, characterized in that: The performing feature decoupling processing on the first image feature to separate feature information and redundant information in the first image feature and obtaining a second image feature corresponding to the feature information in the first image feature includes: Inputting the first image feature into a second feature decoupling module, so as to perform feature decoupling processing on the first image feature through the second feature decoupling module to obtain a second image feature; The second feature decoupling module has the same structure as the first feature decoupling module.
6. The motion recognition method according to any one of claims 1 to 5, characterized in that: The step of determining the target action category corresponding to the to-be-processed video frame from the action category label according to the second image feature and the text feature vector comprises: Performing global pooling processing on the second image feature to obtain a global feature vector; Inputting the global feature vector into a preset anchor point self-attention module, so as to learn the context dependency between distant pixels and enhance the central features through the anchor point self-attention module, and obtain a processed target global feature vector; According to the target global feature vector and the text feature vector, a target action category corresponding to the to-be-processed video frame is determined from the action category label.
7. The action recognition method according to claim 6, characterized in that: The step of determining the target action category corresponding to the to-be-processed video frame from the action category label according to the target global feature vector and the text feature vector comprises: Calculating similarity between the target global feature vector and the text feature vector; The target action category corresponding to the to-be-processed video frame is determined according to the similarity calculation result.
8. A motion recognition system, characterized in that: The system comprises: A data acquisition module, used to acquire a video frame to be processed and at least one action category label; An image encoding module, used for performing image encoding on the video frame to be processed and obtaining a first image feature; a feature decoupling module, configured to perform feature decoupling processing on the first image feature to separate feature information and redundant information in the first image feature, and obtain a second image feature corresponding to the feature information in the first image feature, wherein the feature information includes information related to the action category, and the redundant information includes information unrelated to the action category; A text encoding module, used for performing text encoding on the action category label and obtaining a text feature vector; An action recognition module is used to determine the target action category corresponding to the video frame to be processed from the action category label according to the second image feature and the text feature vector.
9. An intelligent terminal, characterized in that: The intelligent terminal includes a memory, a processor, and a motion recognition program stored in the memory and executable on the processor. When the motion recognition program is executed by the processor, the steps of the motion recognition method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores an action recognition program, and when the action recognition program is executed by the processor, the steps of the action recognition method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Action recognition method and device, electronic equipment and storage medium
CN116994188A
Visible light and infrared pedestrian re-identification method based on feature decoupling
CN117496552A
Action recognition method and device, electronic equipment and readable storage medium
CN117953581A