Open Behavior Recognition Method, Device and Storage Medium for Jointly Generating Discriminative Features
By adopting an open behavior recognition method that combines the generation of discriminant features in video behavior recognition, end-to-end in-depth evidence learning is performed using video feature encoder, generative model and classification network, the problem of difficulty in identifying unknown categories in the existing technology is solved, and more efficient video behavior recognition performance is achieved.
Patent Information
- Application Number
- CN202310170158.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-02-27
AI Technical Summary
Existing video behavior recognition methods are difficult to identify unknown categories that have not appeared during the training process, which may cause security problems in real application scenarios.
The open behavior recognition method of joint generation of discriminant features is adopted, and end-to-end deep evidence learning is carried out through video feature encoder, generative model and classification network to generate joint generation of discriminant features to identify unknown categories.
This method can not only identify known categories that appear during the training process, but also effectively identify unknown categories, solving open problems in video behavior recognition tasks and improving recognition performance.
Smart Images

Figure CN116152717B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to an open behavior recognition method, device, and storage medium for jointly generating discriminant features. Background Art
[0002] Video behavior recognition is an important and key task in video understanding. Most existing video behavior recognition methods are proposed under the closed-set assumption, that is, they can only recognize data from known classes. Among them, the known classes represent the classes that appear during the training process. However, in real application scenarios, video behavior recognition systems often encounter classes that do not appear during the training process, and these unseen classes are collectively referred to as unknown classes. Existing video behavior systems will recognize the data of these unknown classes as a certain known class, causing security problems in actual applications. Therefore, this requires that video behavior recognition methods should not only be able to accurately classify the known classes that appear during the training process, but also be able to recognize the unknown classes that do not appear during the training process.
[0003] Currently, the relevant video behavior methods mainly rely on deep evidence learning to solve the open problem of video behavior recognition. By using deep evidence learning, the open problem of video behavior recognition is transformed into an uncertainty estimation problem. Deep evidence learning uses a deep neural network to predict the Dirichlet distribution of class probabilities, which can be regarded as an evidence collection process. The learned evidence helps to quantify the prediction uncertainty of various human behaviors, so that behaviors from unknown classes generate high uncertainty. In this way, each behavior data will have an uncertainty value. When the uncertainty value is higher than the threshold, it is recognized as an unknown class, and when it is lower than the threshold, it is recognized as a certain known class. The threshold is theoretically the maximum uncertainty value in the training data. Summary of the Invention
[0004] The purpose of the embodiments of this specification is to provide an open behavior recognition method, device, and storage medium for jointly generating discriminant features.
[0005] To solve the above technical problems, the embodiments of the present application are implemented in the following ways:
[0006] In a first aspect, the present application provides an open behavior recognition method for jointly generating discriminant features, and the method includes:
[0007] Obtain the video to be recognized;
[0008] Use the open behavior recognition model to calculate the uncertainty score and classification score of the video to be recognized;
[0009] If the uncertainty score is greater than the threshold, the video to be recognized is classified as an unknown class. If the uncertainty score is less than or equal to the threshold, determine the classification label of the video to be recognized according to the classification score.
[0010] In one embodiment, the open behavior recognition model includes a video feature encoder, a generative model, and a classification network;
[0011] Calculating the uncertainty score of the video to be recognized using the open behavior recognition model includes:
[0012] Extracting the feature vector of the video to be recognized using the video feature encoder;
[0013] Generating a generated feature vector corresponding to the feature vector using the generative model;
[0014] Concatenating the feature vector and the corresponding generated feature vector to generate a joint generated discriminant feature;
[0015] Inputting the joint generated discriminant feature into the classification network to obtain the uncertainty score and classification score of the video to be recognized.
[0016] In one embodiment, when training the open behavior recognition model, end-to-end deep evidence learning is performed on the video feature encoder, the generative model, and the classification network.
[0017] In one embodiment, the loss function of the joint generative model and the loss function of the classification network Train the loss functions of the video feature encoder, the generative model, and the classification network in an end-to-end manner:
[0018]
[0019] where N is the number of videos in each iteration of training.
[0020] In one embodiment, the loss function of the classification network is:
[0021]
[0022] where t ij represents the 0-1 binary form of the class label y i of the video x i and α ij = e ij + 1, e ij is the non-negative output of the classification network C, and K is the number of known classes in the training set.
[0023] In one embodiment, the classification score of each video x i is: α ij / S i and the uncertainty score of each video x i is u i = K / Si 。
[0024] In one embodiment, the threshold is determined according to the uncertainty score during training.
[0025] In one embodiment, the threshold τ is:
[0026]
[0027] where u i is the uncertainty score of each video x i during training, X represents the entire training set, and s is a free parameter used to provide margin relaxation.
[0028] In a second aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the open behavior recognition method for jointly generating discriminant features as in the first aspect.
[0029] In a third aspect, the present application provides a readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the open behavior recognition method for jointly generating discriminant features as in the first aspect.
[0030] As can be seen from the technical solutions provided in the embodiments of the present specification above, this solution: can solve the open problem in the video behavior recognition task, and can not only recognize the categories that appear during the training process, but also recognize the data of unknown categories that have not appeared as unknown categories. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0032] Figure 1 is a schematic flow chart of the open behavior recognition method for jointly generating discriminant features provided by the present application;
[0033] Figure 2 is another schematic flow chart of the open behavior recognition method for jointly generating discriminant features provided by the present application;
[0034] Figure 3 is a schematic structural diagram of the electronic device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] To enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.
[0036] In the following description, specific details such as specific system structures and technologies are proposed for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of this application. However, those skilled in the art should clearly understand that this application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of this application.
[0037] Without departing from the scope or spirit of this application, various improvements and changes can be made to the specific implementation manners of this application's specification, which are obvious to those skilled in the art. Other implementation manners obtained from this application's specification are obvious to those skilled in the art. This application's specification and embodiments are only exemplary.
[0038] Regarding the use of "including", "comprising", "having", "containing", etc. in this article, they are all open-ended terms, that is, they are intended to include but not be limited to.
[0039] Existing technology deep evidence learning is based on discriminative features learning of known classes, and the discriminative features of known classes do not contain information of unknown classes. Therefore, it is not suitable to directly use this feature to solve the open problems in video behavior recognition tasks.
[0040] Based on the above defects, this application provides an open behavior recognition method for jointly generating discriminative features. The jointly generated discriminative features of this method are generated through the input and output features of a generative model, and deep evidence learning is performed based on this feature. This feature contains more knowledge of unknown classes and is more suitable for solving the open problems in video behavior recognition tasks.
[0041] The following further details the present invention in conjunction with the drawings and embodiments.
[0042] Refer to Figure 1 and Figure 2 , which shows a schematic flowchart of an open behavior recognition method for jointly generating discriminative features applicable to the embodiments of this application.
[0043] As Figure 1 shown, the open behavior recognition method for jointly generating discriminative features may include:
[0044] S110: Obtain a video to be identified.
[0045] Specifically, the video to be identified may be a video captured in real time, or a video obtained from a known data set, which is not limited here.
[0046] S120: Calculate the uncertainty score and classification score of the video to be identified using an open behavior recognition model.
[0047] like Figure 2 As shown, the open action recognition model includes a video feature encoder, a generation model, and a classification network.
[0048] S120 may specifically include:
[0049] A video feature encoder is used to extract a feature vector of the video to be identified;
[0050] A generated feature vector corresponding to the feature vector generated by the generative model is generated;
[0051] Concatenate the feature vector and the corresponding generated feature vector in series to generate a jointly generated discriminant feature;
[0052] The discriminant features are jointly generated and input into the classification network to obtain the uncertainty score and classification score of the video to be identified.
[0053] The video feature encoder E is used to project the original video into the feature space, that is, from the original video space to the feature space, which can reduce the complexity of classification and generation. It can be understood that when performing behavior recognition on a video, the original video is the video to be recognized, when training an open behavior recognition model, the original video is the video in the training set, and when testing the open behavior recognition model, the original video is the video in the test set.
[0054] The video feature encoder E in this application selects the currently popular video behavior recognition deep learning methods, including but not limited to I3D, TSM, SlowFast, TPN and ViT. The video feature encoder is used to i Extract a 768-dimensional feature vector z with high-level semantic information i =E(x i ), where x i represents the i-th input video, and E represents the video feature encoder.
[0055] The generative model G is used to generate the feature space of known classes. In this application, the generative model selects the currently popular generative network model, including but not limited to autoencoder (AutoEncoder), variational autoencoder (VAE) and flow model (Flow). For each input feature vector zi Generate a corresponding generated feature vector of the same dimensionality size where G represents the generation model.
[0056] To introduce the difference between the input and output of the generation model, this application concatenates the input feature z of the generation model i and the output feature to obtain a joint generation discriminant feature with a dimensionality size of 1536.
[0057] The classification network C is used to perform K-class classification on the input joint generation discriminant feature, where K is the number of known classes in the training set, and is used to classify the known class samples that have appeared during the training process.
[0058] The classification network C of this application selects a three-layer linear fully connected network, with the number of layers being 1536, 768, and K respectively, and the output uses the activation function Relu. The structure of the classification network includes but is not limited to this structure. The input is the joint generation discriminant feature and the output is the classification score and uncertainty score of the video to be recognized.
[0059] In one embodiment, when training the open behavior recognition model, end-to-end deep evidence learning is performed on the video feature encoder, the generation model, and the classification network.
[0060] Specifically, the loss function of the joint generation model and the loss function of the classification network are used to train the loss functions of the video feature encoder, the generation model, and the classification network in an end-to-end manner:
[0061]
[0062] where N is the number of videos in each iteration of training.
[0063] Taking the autoencoder as an example, the loss function of the generation model G is:
[0064]
[0065] Calculate the loss function of the classification network C based on deep evidence learning is:
[0066]
[0067] where t ij represents the 0-1 binary form of the class label y i of the video x i , and α ij = e ij+1, e ij is the non - negative output of the classification network C
[0068] For each video x i the classification score is α ij / S i and the uncertainty score is u i = K / S i .
[0069] In this embodiment, the difference between the input and output features of the generative model is used to help identify unknown classes. The generative model here is trained on known - class features and can be any deep - generative method. When an unknown - class input is fed into the generative model, since the generative model is trained on known classes, its output should be more like known classes, which results in a difference between the input and output of the generative model. And this difference can be used as an important piece of knowledge to help identify unknown classes.
[0070] Moreover, this application trains the generative model based on the feature space rather than the original video data, which can greatly reduce the computational complexity.
[0071] In addition, this application trains the entire open - behavior recognition model in an end - to - end manner, which can further improve the final open - recognition performance.
[0072] S130: If the uncertainty score is greater than the threshold, the video to be recognized is classified as an unknown class. If the uncertainty score is less than or equal to the threshold, the classification label of the video to be recognized is determined according to the classification score.
[0073] Among them, the threshold τ is determined according to the uncertainty scores during training and is used to distinguish between known classes and unknown classes. Theoretically, the size of the threshold is selected as the maximum value of the uncertainty scores of the training data. Specifically:
[0074]
[0075] where u i is the uncertainty score of each video x i during training, X represents the entire training set, and s is a free parameter used to provide margin relaxation.
[0076] According to the above - determined threshold, it is identified whether the video to be recognized belongs to an unknown class. If the uncertainty score is greater than the threshold, the video to be recognized is classified as the (K + 1) - th class, that is, the unknown class. If the uncertainty score is less than or equal to the threshold, an appropriate known - class label will be assigned to the video to be recognized from the classification model. The formula is as follows:
[0077]
[0078] In the formula, pred(xi ) represents the final classification label of video x i .
[0079] The open behavior recognition method for jointly generating discriminant features provided by this application can solve the open problems in video behavior recognition tasks. It can not only recognize the categories that appear during the training process, but also identify the data of unknown categories that have not appeared as unknown categories.
[0080] The open behavior recognition method for jointly generating discriminant features provided by this application has been experimentally demonstrated to improve by 3%-4% compared with the existing technology methods in two open behavior recognition scenarios, UCF101-HMDB51 and UCF101-MiTV2.
[0081] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 3 shown, it shows a schematic structural diagram of an electronic device 300 suitable for implementing the embodiments of this application.
[0082] As Figure 3 shown, the electronic device 300 includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 302 or the program loaded from the storage part 308 into the random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the device 300 are also stored. The CPU 301, ROM 302, and RAM 303 are connected to each other through a bus 304. The input / output (I / O) interface 305 is also connected to the bus 304.
[0083] The following components are connected to the I / O interface 305: an input part 306 including a keyboard, a mouse, etc.; an output part 307 including such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage part 308 including a hard disk, etc.; and a communication part 309 including a network interface card such as a LAN card, a modem, etc. The communication part 309 performs communication processing via a network such as the Internet. The drive 310 is also connected to the I / O interface 306 as required. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as required, so that the computer program read from it can be installed into the storage part 308 as required.
[0084] Specifically, according to the embodiments of the present disclosure, as referred to above Figure 1The described process can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product that includes a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing the above-described open behavior recognition method for jointly generating discriminative features. In such an embodiment, the computer program can be downloaded and installed from a network via a communication section 309, and / or installed from a removable medium 311.
[0085] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, and the foregoing module, segment of a program, or part of code includes one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0086] The units or modules involved in the embodiments described in this application can be implemented in software or in hardware. The described units or modules can also be provided in a processor. The names of these units or modules do not, in some cases, constitute a limitation on the units or modules themselves.
[0087] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by a computer chip or an entity, or by a product with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a mobile phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0088] As another aspect, the present application also provides a storage medium. The storage medium can be the storage medium included in the foregoing device in the above embodiments; or it can exist separately and be not assembled into the device. The storage medium stores one or more programs, and the foregoing programs are used by one or more processors to execute the open behavior recognition method for jointly generating discriminative features described in this application.
[0089] A storage medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0090] It should be noted that the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "including a..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0091] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the related content.
Claims
1. An open behavior recognition method for jointly generating discriminative features, characterized in that, the method includes: Obtain the video to be recognized; Use an open behavior recognition model to calculate the uncertainty score and classification score of the video to be recognized. The open behavior recognition model includes a video feature encoder, a generation model, and a classification network; If the uncertainty score is greater than the threshold, the video to be recognized is classified as an unknown class. If the uncertainty score is less than or equal to the threshold, determine the classification label of the video to be recognized according to the classification score; Combined with the loss function of the generative model and the loss function of the classification network The loss function L for training the video feature encoder, the generative model, and the classification network in an end-to-end manner is as follows: Among them, is the number of videos in each iterative training; The loss function of the classification network is as follows: Among them, represents the class label of the video in its 01 binary form, , is the non - negative output of the classification network , , K is the number of known classes in the training set.
2. The method according to claim 1, characterized in that, The step of using an open behavior recognition model to calculate the uncertainty score of the video to be recognized includes: Use a video feature encoder to extract the feature vector of the video to be recognized; Use a generation model to generate the generated feature vector corresponding to the feature vector; Concatenate the feature vector and the corresponding generated feature vector to generate a jointly generated discriminative feature; The jointly generated discriminative feature is input into the classification network to obtain the uncertainty score and classification score of the video to be recognized.
3. The method according to claim 2, characterized in that, When training the open behavior recognition model, perform end-to-end deep evidence learning on the video feature encoder, the generation model, and the classification network.
4. The method according to claim 1, characterized in that, Each video has a classification score of: , and each video has an uncertainty score of .
5. The method according to claim 1, characterized in that, The threshold is determined according to the uncertainty score during training.
6. The method according to claim 5, characterized in that, The threshold value is: Among them, is the uncertainty score for each video during training , represents the entire training set is a free parameter used to provide margin relaxation 7. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the open behavior recognition method for jointly generating discriminative features as described in any one of claims 1-6.
8. A readable storage medium, on which a computer program is stored, characterized in that, When the program is executed by the processor, it implements the open behavior recognition method for jointly generating discriminative features as described in any one of claims 1-6.
Citation Information
Patent Citations
Image recognition method and device based on hybrid model, and medium
CN113554127A
Picture recognition method and apparatus, computer device and computer- readable medium
US20180260621A1