A video recognition method
Patent Information
- Application Number
- CN202311560317.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-21
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-11-21
AI Technical Summary
因此,解决当前视频识别算法高精度而低效率的问题十分迫切
[0065]本公开实施例中,视频识别模型可以获取目标视频的目标视频帧,并且获取目标视频帧的目标图像区域,进而依据目标图像区域的局部特征图获取到目标视频的识别结果;其中,目标视频帧包含的信息量大于非目标视频帧包含的信息量,目标图像区域包含的信息量大于非目标图像区域包含的信息量。如此,相较于根据目标视频的各个视频帧进行视频识别,本公开实施例仅仅根据目标视频帧进行视频识别,从时间维度上降低了视频数据的冗余性;相较于根据视频帧的整张图像区域进行视频识别,本公开实施例仅仅根据目标视频帧的目标图像区域进行视频识别,从空间维度上降低了视频数据的冗余性。此外,视频识别模型是基于条件退出策略训练得到的,条件退出策略可以用于动态控制视频识别模型的训练过程中使用样本数据的数量,因此,从样本维度上降低了视频数据的冗余性。如此,本公开实施例通过压缩时间冗余信息、空间冗余信息、样本冗余信息,高效地完成了视频识别。
Smart Images

Figure CN117523454B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a video recognition method. Background Technology
[0002] In recent years, various online video platforms have developed rapidly, and the amount of online video data has surged. Thanks to advancements in deep neural network technology, accurate video recognition algorithms have been successfully applied in fields such as recommendation, surveillance, and content search.
[0003] Current technologies focus on designing larger, deeper, and more complex neural networks to improve the accuracy of video recognition algorithms, but they neglect the enormous computational resource overhead this incurs. In practical applications, the computational load of algorithms is directly related to inference latency, energy consumption, and carbon emissions; from economic, environmental, and security perspectives, computational load is a significant factor that cannot be ignored. Furthermore, in the widespread application scenarios of video recognition technology, such as motion capture and security systems, deep neural network-based algorithms often need to be deployed on edge devices with limited computing resources. In this situation, the bottleneck limiting the algorithm is no longer its accuracy, but its computational efficiency. Therefore, solving the problem of high accuracy but low efficiency in current video recognition algorithms is extremely urgent. Summary of the Invention
[0004] In view of the above problems, this disclosure provides a video recognition method to overcome or at least partially solve the above problems.
[0005] A first aspect of this disclosure provides a video recognition method applied to a video recognition model, the video recognition model comprising: a global feature extraction network, a local feature extraction network, a policy network, and a classifier; the video recognition model is trained based on a conditional exit policy, the conditional exit policy being used to: dynamically control the number of sample data used during the training process of the video recognition model; the method includes:
[0006] The target video is input into the video recognition model to obtain the global feature map of each video frame of the target video output by the global feature extraction network;
[0007] The global feature maps of each video frame are input into the policy network to obtain multiple target video frames; wherein, the target video frames contain more information than the non-target video frames.
[0008] The global feature map of each target video frame is input into the policy network to obtain the target image region of each target video frame; wherein the target image region contains more information than the non-target image region.
[0009] The target image region of each target video frame is input into the local feature extraction network to obtain the local feature map of each target video frame;
[0010] The local feature map of each target video frame is input into the classifier to obtain the recognition result of the target video.
[0011] Optionally, the training steps of the video recognition model include at least:
[0012] The video sample is input into the initial video recognition model to obtain global feature map samples of each video frame of the video sample output by the initial global feature extraction network; the video sample carries a category label; the initial video recognition model includes: the initial global feature extraction network, the initial local feature extraction network, the initial policy network, and the initial classifier;
[0013] The global feature map samples of each video frame sample are input into the initial policy network to obtain multiple target video frame samples; wherein, the number of target video frame samples is determined by the conditional exit policy;
[0014] The global feature map sample of each target video frame sample is input into the initial classifier to obtain the first recognition result of each target video frame;
[0015] Each target video frame sample is input into the initial policy network to obtain a target image region sample for each target video frame sample;
[0016] The target image region sample of each target video frame sample is input into the initial local feature extraction network to obtain the local feature map sample of each target video frame sample;
[0017] The local feature map sample of each target video frame sample is input into the initial classifier to obtain the second recognition result of each target video frame;
[0018] The global feature map samples of the video frame samples and the local feature map samples of each target video frame sample are input into the initial classifier to obtain the third recognition result;
[0019] Based on the first recognition result, the second recognition result, the third recognition result, and the category label, the initial video recognition model is trained to obtain the trained video recognition model.
[0020] Optionally, the step of inputting video samples into the initial video recognition model to obtain global feature map samples of each video frame of the video samples output by the initial global feature extraction network includes:
[0021] The video sample is input into the initial video recognition model to obtain the initial video frame samples that make up the video sample;
[0022] The initial video frame samples are uniformly sampled to obtain the individual video frame samples of the video sample.
[0023] Each video frame sample is input into the initial global feature extraction network to obtain global feature map samples of each video frame sample.
[0024] Optionally, the step of inputting the various video frame samples into the initial policy network to obtain multiple target video frame samples includes:
[0025] Each video frame sample is input into the initial policy network to obtain multiple initial target video frame samples;
[0026] According to the order of the multiple initial target video frame samples in the video sample, for each initial target video frame sample, it is determined whether the initial target video frame sample meets the exit condition; wherein, the exit condition is: the accuracy of the recognition result of the video sample predicted based on the initial target video frame sample and each video frame sample preceding the initial target video frame sample is greater than the exit threshold.
[0027] If any of the initial target video frame samples satisfies the exit condition, the initial target video frame sample and each of the initial target video frame samples preceding the initial target video frame sample are determined as the target video frame sample.
[0028] Optionally, determining whether the initial target video frame sample meets the exit condition includes:
[0029] The initial target video frame sample and each initial target video frame sample preceding the initial target video frame sample are input into the initial local feature extraction network to obtain multiple initial local feature map samples.
[0030] The global feature map samples of the initial target video frame sample, the global feature map samples of each video frame sample preceding the initial target video frame sample, and the multiple initial local feature maps are input into the initial classifier to obtain the conditional recognition result;
[0031] Based on the conditional recognition results and the category labels, determine the accuracy of the predicted recognition results for the video samples;
[0032] Determine whether the accuracy rate is greater than the exit threshold.
[0033] Optionally, the step of inputting each target video frame sample into the initial policy network to obtain a target image region sample for each target video frame sample includes:
[0034] Each target video frame sample is input into the initial policy network to obtain a quadruple corresponding to each target video frame. The quadruple includes: center coordinates, height, and width.
[0035] The target video frame samples are cropped based on the quadruple to obtain the target image region sample for each target video frame sample.
[0036] Optionally, training the initial video recognition model based on the first recognition result, the second recognition result, the third recognition result, and the category label to obtain the trained video recognition model includes:
[0037] Based on the first identification result and the category label, determine the time loss function;
[0038] Based on the second identification result and the category label, determine the spatial loss function;
[0039] Based on the third identification result and the category label, determine the category loss function;
[0040] The initial video recognition model is trained based on the time loss function, the spatial loss function, and the category loss function to obtain a trained video recognition model.
[0041] Optionally, determining the spatial loss function based on the second identification result and the category label includes:
[0042] Based on the second recognition result and the category label of each target video frame sample, a cross-entropy loss function is constructed;
[0043] Obtain the height and width of each target video frame sample, and obtain the height and width of the target image region sample of each target video frame sample;
[0044] Based on the height and width of each target video frame sample, and the height and width of the target image region sample of each target video frame sample, determine the height difference and width difference corresponding to each target video frame sample;
[0045] The spatial loss function is determined based on the cross-entropy loss function and the height and width differences corresponding to each target video frame sample.
[0046] Optionally, determining the category loss function based on the third identification result and the category label includes:
[0047] Obtain the category loss function corresponding to the global features, and obtain the category loss function corresponding to the local features;
[0048] Based on the third identification result and the category label, determine the classification category loss function;
[0049] The category loss function is determined by summing the category loss function corresponding to the global feature, the category loss function corresponding to the local feature, and the classification category loss function.
[0050] Optionally, obtaining the category loss function corresponding to the global features includes:
[0051] The global feature map samples of each of the video frame samples are input into the initial classifier to obtain the fourth recognition result;
[0052] Based on the fourth identification result and the category label, the category loss function corresponding to the global feature is obtained;
[0053] The category loss function for obtaining local features includes:
[0054] The local feature map samples of each of the target video frame samples are input into the initial classifier to obtain the fifth recognition result;
[0055] Based on the fifth identification result and the category label, the category loss function corresponding to the local feature is obtained.
[0056] A second aspect of this disclosure provides a video recognition apparatus applied to a video recognition model, the video recognition model comprising: a global feature extraction network, a local feature extraction network, a policy network, and a classifier; the video recognition model is trained based on a conditional exit policy, the conditional exit policy being used to: dynamically control the number of sample data used during the training process of the video recognition model; the apparatus includes:
[0057] The global feature extraction module is used to input the target video into the video recognition model and obtain the global feature map of each video frame of the target video output by the global feature extraction network.
[0058] The video frame determination module is used to input the global feature maps of each video frame into the policy network to obtain multiple target video frames; wherein, the target video frames contain more information than the non-target video frames.
[0059] The region determination module is used to input the global feature map of each target video frame into the policy network to obtain the target image region of each target video frame; wherein the target image region contains more information than the non-target image region.
[0060] The local feature extraction module is used to input the target image region of each target video frame into the local feature extraction network to obtain the local feature map of each target video frame;
[0061] The classification module is used to input the local feature map of each target video frame into the classifier to obtain the recognition result of the target video.
[0062] A third aspect of this disclosure provides an electronic device, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute instructions to implement the video recognition method of the first aspect.
[0063] A fourth aspect of this disclosure provides a computer-readable storage medium that, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to perform the video recognition method of the first aspect.
[0064] The embodiments disclosed herein have the following advantages:
[0065] In this embodiment, the video recognition model can acquire the target video frame of the target video and the target image region of the target video frame, and then obtain the recognition result of the target video based on the local feature map of the target image region. The target video frame contains more information than non-target video frames, and the target image region contains more information than non-target image regions. Thus, compared to performing video recognition based on each video frame of the target video, this embodiment performs video recognition only based on the target video frame, reducing video data redundancy in the temporal dimension; compared to performing video recognition based on the entire image region of a video frame, this embodiment performs video recognition only based on the target image region of the target video frame, reducing video data redundancy in the spatial dimension. Furthermore, the video recognition model is trained based on a conditional exit strategy, which can be used to dynamically control the number of sample data used during the training process of the video recognition model, thus reducing video data redundancy in the sample dimension. Therefore, this embodiment efficiently completes video recognition by compressing temporal redundancy, spatial redundancy, and sample redundancy. Attached Figure Description
[0066] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments of this disclosure will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0067] Figure 1 This is a flowchart of the steps of a video recognition method according to an embodiment of this disclosure;
[0068] Figure 2 This is a schematic diagram of the framework of a video recognition method according to an embodiment of this disclosure. Detailed Implementation
[0069] To make the above-mentioned objectives, features and advantages of this disclosure more apparent and understandable, the disclosure will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0070] Video recognition refers to identifying and classifying video content. Redundancy in input data is a significant reason for the inefficiency of related video recognition algorithms. Highly efficient video recognition methods often focus only on data redundancy in the temporal dimension, neglecting data redundancy in the spatial and sample dimensions, resulting in poor performance.
[0071] Some studies have proposed a coarse-to-fine resource-efficient video recognition framework (LiteEval). This method utilizes two Long Short-Term Memory (LSTM) networks to maintain coarse-grained and fine-grained features, and uses gating units to dynamically determine whether fine-grained computation is needed for each frame, reducing data redundancy and computational overhead in the time dimension. However, this method does not consider data redundancy in the spatial dimension and sample dimension. In addition, the LSTM network itself has a considerable computational cost, and the sequential frame selection method is not conducive to improving computation speed, thus resulting in low computational efficiency.
[0072] Some studies have proposed a conditionally early-retrieval-based efficient video recognition network (FrameExit). This method employs a temporal gating module to dynamically determine whether to stop computation at each frame, thereby reducing temporal redundancy and improving computational efficiency. However, this method also fails to consider spatial and sample dimensions and requires the features of the final selected frame to be temporally continuous. This continuous frame selection strategy makes it difficult to achieve optimal efficiency on some video datasets. Therefore, there is still significant room for improvement in the computational efficiency of this method.
[0073] This disclosure proposes a method for joint dynamic computation across the temporal, spatial, and sample dimensions, significantly reducing the redundancy of input video data in these dimensions. The disclosure improves the dynamic computation strategy in the temporal dimension by replacing recurrent neural networks (such as LiteEval's Long Short-Term Memory network) with more efficient convolutional neural networks. It models temporal dynamic frame selection as a multinomial-distributed, non-repeating, multiple-sampling process, which not only significantly reduces computational overhead but also flexibly allocates computational resources to the most important frames. Furthermore, the dynamic region selection based on bilinear interpolation in the spatial and sample dimensions, along with an adaptive conditional exit strategy based on different focused frame numbers, greatly improves the computational efficiency of the video recognition model.
[0074] Reference Figure 1 The diagram illustrates a flowchart of a video recognition method according to an embodiment of this disclosure. This video recognition method is applied to a video recognition model, which includes a global feature extraction network, a local feature extraction network, a policy network, and a classifier. The video recognition model is trained based on a conditional exit strategy, which dynamically controls the amount of sample data used during the training process of the video recognition model. Figure 1 As shown, the video recognition method may specifically include steps S11 to S15.
[0075] Step S11: Input the target video into the video recognition model to obtain the global feature map of each video frame of the target video output by the global feature extraction network.
[0076] The target video can be any video to be identified. Inputting the target video into a video recognition model yields individual video frames. Inputting each video frame into a global feature extraction network yields global feature maps for each frame. The network architecture of the global feature extraction network can refer to relevant technologies, and this disclosure does not impose any limitations on it.
[0077] Optionally, considering that similar video frames describe similar video content, uniform sampling can be performed on each video frame of the target video to obtain sampled video frames. These sampled video frames are then input into a global feature extraction network to obtain global feature maps for each video frame. By uniformly sampling the video frames and then performing subsequent processing based on them, computational resources can be significantly saved.
[0078] Alternatively, each video frame of the target video can be obtained first, and then each video frame of the target video can be input into the video recognition model to obtain the global feature map of each video frame output by the global feature extraction network.
[0079] Step S12: Input the global feature map of each video frame into the policy network to obtain multiple target video frames.
[0080] The target video frame contains more information than the non-target video frame.
[0081] The video recognition model is a pre-trained model, and its policy network is also a pre-trained policy network. Therefore, by inputting the global feature maps of each video frame into the policy network, the policy network can obtain multiple target video frames containing a large amount of information.
[0082] Step S13: Input the global feature map of each target video frame into the policy network to obtain the target image region of each target video frame.
[0083] The target image region contains more information than the non-target image region.
[0084] By inputting the global feature map of each target video frame into the policy network, the target image region of each target video frame can be obtained. The target image region of the target video frame is rectangular and can be determined by a four-tuple, which includes: center coordinates, height, and width. The policy network outputs a four-tuple, and the video recognition model crops the target video frame based on the four-tuple to obtain the target image region of the target video frame.
[0085] Step S14: Input the target image region of each target video frame into the local feature extraction network to obtain the local feature map of each target video frame.
[0086] By inputting the target image region of each target video frame into the local feature extraction network, the feature map of the target image region of each target video frame can be obtained, that is, the local feature map of the target video frame can be obtained.
[0087] Step S15: Input the local feature map of each target video frame into the classifier to obtain the recognition result of the target video.
[0088] In video recognition, only local feature maps of the target video frame are used for identification to reduce computational load. The local feature maps of the target video frame are input into a classifier to obtain the recognition result of the target video. The recognition result of the target video can characterize the category of the target video.
[0089] Experiments revealed that, compared to related technologies that use all image regions of all video frames, the accuracy of video recognition results obtained using the technical solution of this disclosure is improved. This is because, during video recognition, a portion of the video frames and image regions are discarded; these are frames and image regions containing less information, thus having a smaller impact on the video recognition results. Furthermore, discarded video frames and image regions might actually interfere with the video recognition results; therefore, discarding these video frames and image regions actually improves the accuracy of the video recognition results.
[0090] The technical solution of this disclosure, compared to video recognition based on each video frame of the target video, performs video recognition only based on the target video frame, reducing video data redundancy in the temporal dimension; compared to video recognition based on the entire image region of a video frame, this disclosure performs video recognition only based on the target image region of the target video frame, reducing video data redundancy in the spatial dimension. Furthermore, the video recognition model is trained based on a conditional exit strategy, which can be used to dynamically control the number of sample data used during the training process of the video recognition model, thus reducing video data redundancy in the sample dimension. In this way, this disclosure efficiently completes video recognition by compressing temporal redundancy, spatial redundancy, and sample redundancy.
[0091] The training method for the video recognition model is described below. The video recognition model is obtained by training an initial video recognition model, which is the video recognition model to be trained. The initial video model includes: an initial global feature extraction network, an initial local feature extraction network, an initial policy network, and an initial classifier. The initial global feature extraction network is the global feature extraction network to be trained, the initial local feature extraction network is the local feature extraction network to be trained, the initial policy network is the policy network to be trained, and the initial classifier is the classifier to be trained.
[0092] The initial global feature extraction network is represented as f G The initial local feature extraction network is represented as f L The initial policy network is represented as π, and the initial classifier is represented as f. C The parameters of these networks are randomly initialized. Among them, the initial local feature extraction network has the largest number of parameters and computational cost, but the strongest inference ability and is the main source of computational overhead.
[0093] The training steps for the video recognition model include at least steps S21 to S28.
[0094] Step S21: Input the video sample into the initial video recognition model to obtain the global feature map samples of each video frame of the video sample output by the initial global feature extraction network.
[0095] The video samples carry category labels.
[0096] The video sample is described using symbols. The length of the video sample is T0, and the resolution of the video sample is H×W. The frame sequence of the video sample can be represented as follows: The video sample carries the category label y.
[0097] The video sample can be any video sample in the training sample set. The initial video recognition model will be trained using multiple video samples from the training sample set. In order to avoid going into too much detail, this disclosure will only describe it using a single video sample.
[0098] By inputting video samples into the initial video recognition model, we can obtain individual video frame samples. By inputting individual video frame samples into the initial global feature extraction network, we can obtain global feature map samples for each video frame.
[0099] Optionally, inputting video samples into an initial video recognition model can yield initial video frame samples that make up the video samples; uniformly sampling the initial video frame samples yields individual video frame samples of the video samples; inputting the individual video frame samples into the initial global feature extraction network yields global feature map samples of the individual video frame samples.
[0100] Each initial video frame sample is a video frame represented by a frame sequence of video samples. The frame sequence V is uniformly sampled T. G A frame can be used to obtain the global frame sequence N. G N G Individual video frame samples that represent video samples.
[0101] N G Each frame is input into the initial global feature network, which yields global feature map samples for each video frame. Among them, v t The video frame sample represents the video sample at time t. The global feature map sample is a coarse-grained feature map.
[0102] Alternatively, one can first obtain the video frame samples of the video sample, and then input the video frame samples of the video sample into the initial video recognition model to obtain the global feature map samples of the video frame samples output by the initial global feature extraction network.
[0103] Step S22: Input the global feature map samples of each video frame sample into the initial policy network to obtain multiple target video frame samples.
[0104] The number of target video frame samples is determined by the conditional exit strategy.
[0105] Step S23: Input the global feature map sample of each target video frame sample into the initial classifier to obtain the first recognition result of each target video frame.
[0106] The method for determining the number of target video frame samples through a conditional exit strategy will be described in detail later. Here, we will only introduce the method for determining the target video frame.
[0107] The accuracy of the target video frame samples identified by the initial policy network to be trained is poor. By training the initial video recognition model, the initial policy network will also be trained, thereby improving the accuracy of the initial policy network in sampling target video frame samples.
[0108] This disclosure models the problem of selecting target video frame samples as a temporal sampling problem to achieve dynamic temporal computation. To reduce temporal redundancy of video data, this disclosure models the target video frame selection problem most relevant to the recognition task as a non-repeating sampling problem based on a temporal probability distribution. T video frames are selected from all T video frame samples according to their relevance to the task. L For each target video frame sample, the sequence number of the selected target video frame sample is denoted as... Specifically, each target video frame sample is collected sequentially in the following manner:
[0109]
[0110]
[0111] Where, n i+1 The sample represents the (i+1)th video frame; WeightedSample represents the weighted sampling; j represents the number of samples; w′ j The actual weighted sampling weights for the j-th weighted sampling without replacement; {n1,...,n i} represents the sequence of target video frame samples; w j "otherwise" represents the weighted sampling weight in the j-th weighted sampling without replacement; "otherwise" represents other characters; the remaining characters can be found in the previous text.
[0112] The weighted sampling follows a multinomial distribution. satisfy i represents the i-th video frame sample, and π is the input global feature map sample to the policy network. Output this multinomial distribution. Select video frame samples with high weight parameters as target video frame samples. This disclosure requires that the selected T... LThe target video frame samples are compared with the remaining T0-T L Each video frame sample contains more information, therefore the following optimization objective exists:
[0113]
[0114] Where, minimize represents minimization; Characterizing the time loss function; Characterizing expectation, Characterization using video frame samples The recognition loss function is calculated from the global feature map samples.
[0115] Sample each target video frame The global feature map samples are input into the initial classifier to obtain the first recognition result for each target video frame sample; based on the difference between the first recognition result and the class label for each target video frame sample, the recognition loss function can be determined. The mean of multiple recognition loss functions is used as the time loss function.
[0116] There are two layers of expectation here. One layer is the expectation of sampling different target video frame samples in a video sample, and the other layer is the expectation according to different videos in the training sample set.
[0117] To achieve the aforementioned optimization objectives without significantly increasing computational overhead, this disclosure samples T from frame T0 during training. L The frame is approximately from T G In-frame sampling Frame. At this point, the new weights can be obtained by downsampling. get, This can be obtained using existing auxiliary classifiers. Finally, the Monte Carlo method is used to calculate... According to different local frame sampling methods N L The expected value can be minimized by updating the policy network π using gradient descent.
[0118] To achieve deterministic reasoning, this disclosure includes, during testing... The cumulative distribution function formed is chosen by its T L A uniformly distributed quantile point is used to select the frame at the corresponding horizontal coordinate position on the time axis as a local frame.
[0119] Step S24: Input each of the target video frame samples into the initial policy network to obtain the target image region sample of each target video frame sample.
[0120] The initial policy network can also select target image region samples of the target video frame samples to obtain local feature maps of the target video frame samples, thereby achieving spatial dynamic computation.
[0121] To reduce spatial redundancy in video data, this disclosure applies a specific method to each selected target video frame sample v. t Cropping is performed to obtain a rectangular region with an adaptive shape, which is the target image region sample with the richest information. Then, it is input into the initial local feature extraction network to obtain fine-grained local feature map samples.
[0122] Each target video frame sample is input into the initial policy network to obtain a quadruple corresponding to each target video frame. The quadruple includes: center coordinates, height, and width. The target video frame sample is cropped according to the quadruple to obtain the target image region sample of each target video frame sample.
[0123] The shape adaptation here is reflected in the target image region sample. The aspect ratio and size are dynamically calculated; specifically, the initial policy network π is applied to each frame v. t Input global feature map samples Output a quadruple H represents the center coordinates of the sample in the target image region. t W represents the height of the sample in the labeled image region. t This represents the width of the target image region sample. The original target video frame sample is cropped based on the quadruple and resized to a square P×P to obtain the target image region sample.
[0124] Step S25: Input the target image region sample of each target video frame sample into the initial local feature extraction network to obtain the local feature map sample of each target video frame sample.
[0125] Step S26: Input the local feature map sample of each target video frame sample into the initial classifier to obtain the second recognition result of each target video frame.
[0126] To train the policy network π, this disclosure proposes the following optimization objective:
[0127]
[0128] The original target video frame sample has a size of H×W; the global feature map sample of the target video frame has a size of H. G ×W G The initial input size of the local feature extraction network is P×P; Characterizing the spatial loss function, L CECross-entropy loss function is represented by SoftMax, and soft maximization is represented by FC. Aux Characterize the auxiliary classifier, The feature vector representing a region of the global feature map sample obtained by interpolation using the quadruples of parameters output by the initial policy network π; α represents the coefficients; Interpolation represents interpolation; the other characters can be found in the previous text.
[0129] This disclosure guides the cropping of the pixel space of the original target video frame samples through proportional interpolation of the feature space, and the depth features are derived from global feature map samples. This will not significantly increase computational overhead. The cropping rectangle for the original target video frame sample is H. t ×W t The second regularization term promotes a larger rectangular region output by the policy network π. During training, gradient descent can be used to optimize the policy network π to achieve dynamic calculation of spatial dimensions. During testing, there is no need to calculate the above optimization objective; pruning can be performed directly.
[0130] Step S27: Input the global feature map samples of the video frame samples and the local feature map samples of each target video frame sample into the initial classifier to obtain the third recognition result.
[0131] For coarse-grained global feature map samples and fine-grained local feature map samples The feature vector is obtained by performing average pooling. By reusing global feature map samples from the initial global feature extraction network This can further improve accuracy and increase computational efficiency.
[0132] Input the two sets of feature vectors mentioned above into the initial classifier f. C Classification is performed to obtain the third identification result.
[0133] Step S28: Based on the first recognition result, the second recognition result, the third recognition result, and the category label, train the initial video recognition model to obtain the trained video recognition model.
[0134] Based on the first recognition result and the category label, a temporal loss function is determined; based on the second recognition result and the category label, a spatial loss function is determined; based on the third recognition result and the category label, a category loss function is determined; based on the temporal loss function, the spatial loss function, and the category loss function, the initial video recognition model is trained to obtain a trained video recognition model.
[0135] Based on the category labels and the first recognition results corresponding to each target video frame sample, a cross-entropy loss function can be established, and the mean of each cross-entropy loss function can be used as the time loss function.
[0136] Based on the second recognition result and the category label of each target video frame sample, a cross-entropy loss function is constructed; the height and width of each target video frame sample, and the height and width of the target image region sample of each target video frame sample are obtained; based on the height and width of each target video frame sample, and the height and width of the target image region sample of each target video frame sample, the height difference and width difference corresponding to each target video frame sample are determined; based on the cross-entropy loss function and the height difference and width difference corresponding to each target video frame sample, the spatial loss function is determined.
[0137] Based on the third identification result and the category label, determining the category loss function may include: obtaining the category loss function corresponding to the global features and obtaining the category loss function corresponding to the local features; determining the classification category loss function based on the third identification result and the category label; and determining the category loss function as the sum of the category loss function corresponding to the global features, the category loss function corresponding to the local features, and the classification category loss function.
[0138] The global feature map samples of each video frame are input into the initial classifier to obtain the fourth recognition result; based on the fourth recognition result and the category label, the category loss function corresponding to the global feature is obtained.
[0139] The local feature map samples of each of the target video frame samples are input into the initial classifier to obtain the fifth recognition result; based on the fifth recognition result and the category label, the category loss function corresponding to the local feature is obtained.
[0140] The category loss function can be determined using the following formula.
[0141]
[0142] Wherein, the first term L in the formula CE (p,y) represents the classification loss function determined based on the third recognition result and the category label; the second term of the formula represents the category loss function corresponding to the global features; the third term of the formula represents the category loss function corresponding to the local features. Characterizing target video frame samples n i The local feature map sample; the meaning of the other characters can be found in the previous text.
[0143] Understandably, the initial classifier can predict a first recognition result based on the global feature map samples of each target video frame sample, a second recognition result based on the local feature map samples of each target video frame sample, a third recognition result based on the global feature map samples of all video frame samples and the local feature map samples of all target video frames, a fourth recognition result based on the global feature map samples of all video frame samples, and a fifth recognition result based on the local feature map samples of all target video frames.
[0144] The total loss function can be determined using the following formula.
[0145]
[0146] To minimize To achieve this, training the initial video recognition model can improve its accuracy, thereby obtaining a well-trained video recognition model that can achieve efficient video recognition.
[0147] The following describes a method for determining the number of target video frame samples corresponding to each video sample using a conditional exit strategy.
[0148] Each video frame sample is input into the initial policy network to obtain multiple initial target video frame samples. According to the order of the multiple initial target video frame samples in the video samples, for each initial target video frame sample, it is determined whether the initial target video frame sample meets the exit condition. The exit condition is: the accuracy of the recognition result of the video sample predicted based on the initial target video frame sample and each video frame sample preceding it is greater than an exit threshold. If any initial target video frame sample meets the exit condition, the initial target video frame sample and each initial target video frame sample preceding it are determined as the target video frame sample.
[0149] The step of determining whether the initial target video frame sample meets the exit condition may include: inputting the initial target video frame sample and each of the initial target video frame samples preceding the initial target video frame sample into the initial local feature extraction network to obtain multiple initial local feature map samples; inputting the global feature map samples of the initial target video frame sample, the global feature map samples of each of the video frame samples preceding the initial target video frame sample, and the multiple initial local feature maps into the initial classifier to obtain a conditional recognition result; determining the accuracy of the predicted recognition result of the video sample based on the conditional recognition result and the category label; and determining whether the accuracy is greater than the exit threshold.
[0150] Local feature extraction networks (LFIs) are the components with the largest number of parameters and the highest computational cost, yet also the most powerful inference capabilities, in video recognition models, making them the primary source of computational overhead. LFIs are used to extract local features from target video frames; therefore, reducing the number of target video frames can effectively save computational resources.
[0151] The initial policy network can determine multiple initial target video frame samples, but during training, the initial target video frame samples input to the local feature extraction network need to be dynamically calculated.
[0152] Each time an initial target video frame sample is input into the initial local feature extraction network, an initial local feature map sample of that initial target video frame sample is obtained. Because the initial local feature map samples of the initial target video frame samples are extracted sequentially, the initial local feature map samples of each initial target video frame sample preceding that initial target video frame sample have already been obtained before the initial local feature map sample of that initial target video frame sample is obtained.
[0153] By inputting the obtained initial local feature map samples, the global feature map samples of the initial target video frame sample, and the global feature map samples of all video frames preceding the initial target video frame sample into the initial classifier, a conditional recognition result can be obtained. The conditional recognition result can characterize the confidence level of the video sample belonging to each category. Through the conditional recognition result and the category label, the accuracy of the predicted recognition result of the video sample can be determined. If the accuracy is greater than the exit threshold, it proves that there is no need to extract the local feature map samples of the next initial target video frame sample. Instead, the initial target video frame samples for which initial local feature map samples have already been extracted are identified as target video frame samples. In this way, the number of target video frame samples is effectively determined.
[0154] By employing a conditional exit strategy, dynamic calculation of samples can be achieved, thereby allocating more computational resources to more difficult video samples. This disclosure adopts a conditional exit strategy, allowing the initial local feature extraction network to process a variable number of target video frame samples, i.e., dynamically determining T. L In the specific implementation, the decision on whether there is sufficient confidence to exit is made by comparing the cross-entropy loss of the prediction results of the previous t frames with a predefined exit threshold.
[0155]
[0156]
[0157] Where, p t The representation is based on the conditional recognition results obtained from video frame samples at time t and before time t; p i The representation is based on the conditional recognition result obtained from the i-th video frame sample and the video samples before it.
[0158] When the above conditions are met, the calculation can be completed up to frame t to exit and obtain the prediction result. Therefore, for simple samples, a relatively small number of frames are required to meet the exit condition; for difficult samples, a sufficient number of frames need to be calculated to exit, thus realizing dynamic calculation according to the difficulty of sample prediction and achieving efficient video recognition on the validation training sample set.
[0159] Set different candidate exit thresholds; under a given computational cost, obtain the accuracy of the recognition results of the predicted video samples corresponding to each candidate exit threshold; and determine the candidate exit threshold corresponding to the highest accuracy as the exit threshold.
[0160] For the training sample set D val This disclosure determines the calculation exit threshold η by optimizing the following conditional problem. t :
[0161]
[0162] Where, maximize represents maximization, η1,η2,... represent multiple candidate exit thresholds, Acc represents accuracy, FLOPs represents the computational cost during network inference, and B represents the given computational cost, which is the expected upper limit of the network's computational cost.
[0163] Let q be a constant, representing the probability that each frame of a video sample meets the classification confidence requirement and exits the classification process after passing through the initial local feature extraction network. Then, the probability that the target video frame sample at time t achieves the conditional exit is q. t =z(1-q) t-1 q and z are normalization terms that guarantee ∑ t q t=1. Therefore, the constraint condition of the optimization problem can be transformed into |D val |∑ t q t η t ≤B, and then use a greedy strategy to find D. val The highest accuracy rate.
[0164] Figure 2 This is a schematic diagram of the framework of a video recognition method according to an embodiment of this disclosure. In this embodiment, the global feature extraction network, local feature extraction network, policy network, and classifier can be applied to efficient and fast video recognition tasks. The global feature extraction network and local feature extraction network can be deployed as arbitrary deep neural networks according to the computational resource requirements of the task, allowing for flexible deployment of the algorithm framework of this disclosure.
[0165] The technical solution of this disclosure achieves the following effective effects:
[0166] (1) Spatiotemporal Adaptive High-Efficiency Computation: A global feature extraction network and a local feature extraction network are designed. The global feature extraction network has low computational cost. It takes a complete video frame as input, obtains a coarse-grained global feature map of the global frame, and uses it for spatiotemporal adaptive selection. The local feature extraction network has high computational cost. It takes a portion of key regions of a temporally sampled target video frame as input, obtains a fine-grained local feature map of the target video frame, and uses it for prediction results. This saves computational overhead and improves computational efficiency without sacrificing accuracy.
[0167] (2) Dynamic calculation in the time dimension: Select the set of target video frames with the most information in the time dimension, and model the temporal sampling as multiple non-repeating multinomial distribution sampling. This probability distribution is dynamically calculated by the policy network. End-to-end training and optimization are achieved by calculating the expected cross-entropy loss of the prediction results of the feature combination of the target video frames using the Monte Carlo method.
[0168] (3) Dynamic calculation of spatial dimension: In each frame, the local rectangular region with the most information is selected in the spatial dimension. The position, shape and size of the rectangle are adaptive and are dynamically calculated by the policy network. The cross-entropy loss of the prediction results of the local feature rectangle region of the global feature map is calculated by using the same position, shape and size parameters to complete the gradient backpropagation. The combination of feature space guides the selection of frame information, realizing end-to-end training and optimization;
[0169] (4) Dynamic calculation of sample dimensions: The computational overhead is allocated according to the difficulty of video samples in the sample dimension, that is, the number of input target video frame samples. Dynamic calculation of sample dimensions solves the video recognition problem for a total dataset under limited resources, and effectively improves the inference efficiency and speed of video recognition.
[0170] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of this disclosure are not limited to the described order of actions, because according to the embodiments of this disclosure, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of this disclosure.
[0171] This disclosure also provides a schematic diagram of a video recognition device. The device is applied to a video recognition model, which includes a global feature extraction network, a local feature extraction network, a policy network, and a classifier. The video recognition model is trained based on a conditional exit strategy, which dynamically controls the amount of sample data used during the training process of the video recognition model. The device includes a global feature extraction module, a video frame determination module, a region determination module, a local feature extraction module, and a classification module, wherein:
[0172] The global feature extraction module is used to input the target video into the video recognition model and obtain the global feature map of each video frame of the target video output by the global feature extraction network.
[0173] The video frame determination module is used to input the global feature maps of each video frame into the policy network to obtain multiple target video frames; wherein, the target video frames contain more information than the non-target video frames.
[0174] The region determination module is used to input the global feature map of each target video frame into the policy network to obtain the target image region of each target video frame; wherein the target image region contains more information than the non-target image region.
[0175] The local feature extraction module is used to input the target image region of each target video frame into the local feature extraction network to obtain the local feature map of each target video frame;
[0176] The classification module is used to input the local feature map of each target video frame into the classifier to obtain the recognition result of the target video.
[0177] Optionally, the training steps of the video recognition model include at least:
[0178] The video sample is input into the initial video recognition model to obtain global feature map samples of each video frame of the video sample output by the initial global feature extraction network; the video sample carries a category label; the initial video recognition model includes: the initial global feature extraction network, the initial local feature extraction network, the initial policy network, and the initial classifier;
[0179] The global feature map samples of each video frame sample are input into the initial policy network to obtain multiple target video frame samples; wherein, the number of target video frame samples is determined by the conditional exit policy;
[0180] The global feature map sample of each target video frame sample is input into the initial classifier to obtain the first recognition result of each target video frame;
[0181] Each target video frame sample is input into the initial policy network to obtain a target image region sample for each target video frame sample;
[0182] The target image region sample of each target video frame sample is input into the initial local feature extraction network to obtain the local feature map sample of each target video frame sample;
[0183] The local feature map sample of each target video frame sample is input into the initial classifier to obtain the second recognition result of each target video frame;
[0184] The global feature map samples of the video frame samples and the local feature map samples of each target video frame sample are input into the initial classifier to obtain the third recognition result;
[0185] Based on the first recognition result, the second recognition result, the third recognition result, and the category label, the initial video recognition model is trained to obtain the trained video recognition model.
[0186] Optionally, the step of inputting video samples into the initial video recognition model to obtain global feature map samples of each video frame of the video samples output by the initial global feature extraction network includes:
[0187] The video sample is input into the initial video recognition model to obtain the initial video frame samples that make up the video sample;
[0188] The initial video frame samples are uniformly sampled to obtain the individual video frame samples of the video sample.
[0189] Each video frame sample is input into the initial global feature extraction network to obtain global feature map samples of each video frame sample.
[0190] Optionally, the step of inputting the various video frame samples into the initial policy network to obtain multiple target video frame samples includes:
[0191] Each video frame sample is input into the initial policy network to obtain multiple initial target video frame samples;
[0192] According to the order of the multiple initial target video frame samples in the video sample, for each initial target video frame sample, it is determined whether the initial target video frame sample meets the exit condition; wherein, the exit condition is: the accuracy of the recognition result of the video sample predicted based on the initial target video frame sample and each video frame sample preceding the initial target video frame sample is greater than the exit threshold.
[0193] If any of the initial target video frame samples satisfies the exit condition, the initial target video frame sample and each of the initial target video frame samples preceding the initial target video frame sample are determined as the target video frame sample.
[0194] Optionally, determining whether the initial target video frame sample meets the exit condition includes:
[0195] The initial target video frame sample and each initial target video frame sample preceding the initial target video frame sample are input into the initial local feature extraction network to obtain multiple initial local feature map samples.
[0196] The global feature map samples of the initial target video frame sample, the global feature map samples of each video frame sample preceding the initial target video frame sample, and the multiple initial local feature maps are input into the initial classifier to obtain the conditional recognition result;
[0197] Based on the conditional recognition results and the category labels, determine the accuracy of the predicted recognition results for the video samples;
[0198] Determine whether the accuracy rate is greater than the exit threshold.
[0199] Optionally, the step of inputting each target video frame sample into the initial policy network to obtain a target image region sample for each target video frame sample includes:
[0200] Each target video frame sample is input into the initial policy network to obtain a quadruple corresponding to each target video frame. The quadruple includes: center coordinates, height, and width.
[0201] The target video frame samples are cropped based on the quadruple to obtain the target image region sample for each target video frame sample.
[0202] Optionally, training the initial video recognition model based on the first recognition result, the second recognition result, the third recognition result, and the category label to obtain the trained video recognition model includes:
[0203] Based on the first identification result and the category label, determine the time loss function;
[0204] Based on the second identification result and the category label, determine the spatial loss function;
[0205] Based on the third identification result and the category label, determine the category loss function;
[0206] The initial video recognition model is trained based on the time loss function, the spatial loss function, and the category loss function to obtain a trained video recognition model.
[0207] Optionally, determining the spatial loss function based on the second identification result and the category label includes:
[0208] Based on the second recognition result and the category label of each target video frame sample, a cross-entropy loss function is constructed;
[0209] Obtain the height and width of each target video frame sample, and obtain the height and width of the target image region sample of each target video frame sample;
[0210] Based on the height and width of each target video frame sample, and the height and width of the target image region sample of each target video frame sample, determine the height difference and width difference corresponding to each target video frame sample;
[0211] The spatial loss function is determined based on the cross-entropy loss function and the height and width differences corresponding to each target video frame sample.
[0212] Optionally, determining the category loss function based on the third identification result and the category label includes:
[0213] Obtain the category loss function corresponding to the global features, and obtain the category loss function corresponding to the local features;
[0214] Based on the third identification result and the category label, determine the classification category loss function;
[0215] The category loss function is determined by summing the category loss function corresponding to the global feature, the category loss function corresponding to the local feature, and the classification category loss function.
[0216] Optionally, obtaining the category loss function corresponding to the global features includes:
[0217] The global feature map samples of each of the video frame samples are input into the initial classifier to obtain the fourth recognition result;
[0218] Based on the fourth identification result and the category label, the category loss function corresponding to the global feature is obtained;
[0219] The category loss function for obtaining local features includes:
[0220] The local feature map samples of each of the target video frame samples are input into the initial classifier to obtain the fifth recognition result;
[0221] Based on the fifth identification result and the category label, the category loss function corresponding to the local feature is obtained.
[0222] It should be noted that the device embodiments are similar to the method embodiments, so the description is relatively simple. For relevant details, please refer to the method embodiments.
[0223] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0224] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, apparatus, or computer program products. Therefore, embodiments of this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, embodiments of this disclosure can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0225] This disclosure describes embodiments of methods, apparatus, electronic devices, and computer program products according to embodiments of this disclosure with reference to flowchart illustrations and / or block diagrams. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0226] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0227] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0228] While preferred embodiments of the present disclosure have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the present disclosure.
[0229] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes the element.
[0230] The video recognition method provided by this disclosure has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this disclosure. The description of the above embodiments is only for the purpose of helping to understand the method and its core ideas. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this disclosure. Therefore, the content of this specification should not be construed as a limitation of this disclosure.
Claims
1. A video recognition method, characterized in that, The method is applied to a video recognition model, which includes a global feature extraction network, a local feature extraction network, a policy network, and a classifier. The video recognition model is trained based on a conditional exit strategy, which dynamically controls the amount of sample data used during the training process of the video recognition model. The target video is input into the video recognition model to obtain the global feature map of each video frame of the target video output by the global feature extraction network; The global feature maps of each video frame are input into the policy network to obtain multiple target video frames; wherein, the target video frames contain more information than the non-target video frames. The global feature map of each target video frame is input into the policy network to obtain the target image region of each target video frame; wherein the target image region contains more information than the non-target image region. The target image region of each target video frame is input into the local feature extraction network to obtain the local feature map of each target video frame; The local feature map of each target video frame is input into the classifier to obtain the recognition result of the target video; The training steps of the video recognition model include at least the following: The video sample is input into the initial video recognition model to obtain global feature map samples of each video frame of the video sample output by the initial global feature extraction network; the video sample carries a category label; the initial video recognition model includes: the initial global feature extraction network, the initial local feature extraction network, the initial policy network, and the initial classifier; The global feature map samples of each video frame sample are input into the initial policy network to obtain multiple target video frame samples; wherein, the number of target video frame samples is determined by the conditional exit policy; The global feature map sample of each target video frame sample is input into the initial classifier to obtain the first recognition result of each target video frame; Each target video frame sample is input into the initial policy network to obtain a target image region sample for each target video frame sample; The target image region sample of each target video frame sample is input into the initial local feature extraction network to obtain the local feature map sample of each target video frame sample; The local feature map sample of each target video frame sample is input into the initial classifier to obtain the second recognition result of each target video frame; The global feature map samples of the video frame samples and the local feature map samples of each target video frame sample are input into the initial classifier to obtain the third recognition result; Based on the first recognition result, the second recognition result, the third recognition result, and the category label, the initial video recognition model is trained to obtain the trained video recognition model; The step of inputting video samples into the initial video recognition model to obtain global feature map samples of each video frame of the video samples output by the initial global feature extraction network includes: The video sample is input into the initial video recognition model to obtain the initial video frame samples that make up the video sample; The initial video frame samples are uniformly sampled to obtain the individual video frame samples of the video sample. Each video frame sample is input into the initial global feature extraction network to obtain global feature map samples of each video frame sample; The step of inputting the global feature map samples of each video frame sample into the initial policy network to obtain multiple target video frame samples includes: Each video frame sample is input into the initial policy network to obtain multiple initial target video frame samples; According to the order of the multiple initial target video frame samples in the video sample, for each initial target video frame sample, it is determined whether the initial target video frame sample meets the exit condition; wherein, the exit condition is: the accuracy of the recognition result of the video sample predicted based on the initial target video frame sample and each video frame sample preceding the initial target video frame sample is greater than the exit threshold. If any of the initial target video frame samples satisfies the exit condition, the initial target video frame sample and each of the initial target video frame samples preceding the initial target video frame sample are determined as the target video frame sample.
2. The method according to claim 1, characterized in that, The step of determining whether the initial target video frame sample meets the exit condition includes: The initial target video frame sample and each initial target video frame sample preceding the initial target video frame sample are input into the initial local feature extraction network to obtain multiple initial local feature map samples. The global feature map samples of the initial target video frame sample, the global feature map samples of each video frame sample preceding the initial target video frame sample, and the multiple initial local feature maps are input into the initial classifier to obtain the conditional recognition result; Based on the conditional recognition results and the category labels, determine the accuracy of the predicted recognition results for the video samples; Determine whether the accuracy rate is greater than the exit threshold.
3. The method according to claim 1, characterized in that, The step of inputting each target video frame sample into the initial policy network to obtain a target image region sample for each target video frame sample includes: Each target video frame sample is input into the initial policy network to obtain a quadruple corresponding to each target video frame. The quadruple includes: center coordinates, height, and width. The target video frame samples are cropped based on the quadruple to obtain the target image region sample for each target video frame sample.
4. The method according to claim 1, characterized in that, The step of training the initial video recognition model based on the first recognition result, the second recognition result, the third recognition result, and the category label to obtain the trained video recognition model includes: Based on the first identification result and the category label, determine the time loss function; Based on the second identification result and the category label, determine the spatial loss function; Based on the third identification result and the category label, determine the category loss function; The initial video recognition model is trained based on the time loss function, the spatial loss function, and the category loss function to obtain a trained video recognition model.
5. The method according to claim 4, characterized in that, The step of determining the spatial loss function based on the second identification result and the category label includes: Based on the second recognition result and the category label of each target video frame sample, a cross-entropy loss function is constructed; Obtain the height and width of each target video frame sample, and obtain the height and width of the target image region sample of each target video frame sample; Based on the height and width of each target video frame sample, and the height and width of the target image region sample of each target video frame sample, determine the height difference and width difference corresponding to each target video frame sample; The spatial loss function is determined based on the cross-entropy loss function and the height and width differences corresponding to each target video frame sample.
6. The method according to claim 4, characterized in that, The step of determining the category loss function based on the third identification result and the category label includes: Obtain the category loss function corresponding to the global features, and obtain the category loss function corresponding to the local features; Based on the third identification result and the category label, determine the classification category loss function; The category loss function is determined by summing the category loss function corresponding to the global feature, the category loss function corresponding to the local feature, and the classification category loss function.
7. The method according to claim 6, characterized in that, The category loss function for obtaining the global features includes: The global feature map samples of each of the video frame samples are input into the initial classifier to obtain the fourth recognition result; Based on the fourth identification result and the category label, the category loss function corresponding to the global feature is obtained; The category loss function for obtaining local features includes: The local feature map samples of each of the target video frame samples are input into the initial classifier to obtain the fifth recognition result; Based on the fifth identification result and the category label, the category loss function corresponding to the local feature is obtained.
Citation Information
Patent Citations
Target tracking method based on deep transfer learning
CN109544603A
Video classification method and device, equipment and medium
CN112650885A