A video dynamic frame extraction method for experimental operation specification evaluation
Through the video dynamic frame extraction method of deep neural network, the problem of missing keyframes for experimental operation videos in the prior art is solved, and efficient and fair evaluation of experimental operation specifications is achieved, compressing the video duration and retaining the keyframes.
Patent Information
- Application Number
- CN202210952429.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-09
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-08-09
AI Technical Summary
Existing video summary technology cannot accurately describe experimental operations without specific sequences, resulting in the loss of key hand movements, affecting the efficiency and fairness of the evaluation of experimental operation specifications.
Using a video dynamic frame extraction method based on deep neural networks, pseudo-labels are generated through the K-Means algorithm and the YCbCr Gaussian skin color algorithm, a single-input single-output and single-input multi-output neural network is constructed, and two-stage training is carried out. Feature similarity classification and hand feature restoration modules are used to intelligently identify and retain keyframes, skip irrelevant frames, and compress video duration.
It realizes intelligent fast-forward compression of experimental operation videos, retains key interactions, improves evaluation efficiency and fairness, and has a compression rate of 70%. The fluency of key operation clips is consistent with the original video.
Smart Images

Figure CN115457425B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a method for dynamically extracting frames from a video for evaluating experimental operation specifications. Background Art
[0002] In recent years, the Ministry of Education has improved the "new high school entrance examination and new college entrance examination" to improve students' practical ability and proposed to strengthen the study and examination of students' physical, chemical and biological experimental operation abilities. With the development of video technology, the original method of organizing a large number of teachers to invigilate and score on-site has been changed to recording videos of the operation examination process for centralized scoring. Due to the large number of examination videos, the different lengths of the operation examination videos, and the fatigue caused by the fact that it is not conducive to the scoring teachers to watch the videos throughout the process for scoring, it is difficult to ensure the fairness of scoring. In the above background, video acceleration technology has emerged to reduce the viewing time of scoring teachers. However, the existing video acceleration all uses the technology of compressing the frame rate and does not consider the instantaneousness of key experimental operation actions, and cannot filter out the invalid duration existing in video recording. The existing video summary technology based on deep learning mainly focuses on the description of the plot or scene and cannot well describe the experimental operations without a specific order, resulting in the loss of key hand actions. Summary of the Invention
[0003] In view of the above problems, the present invention proposes a method for dynamically extracting frames from a video for evaluating experimental operation specifications, aiming to solve the problem that the existing video summary technology cannot accurately describe the experimental operations without a specific order, resulting in the loss of key hand actions.
[0004] To solve the above technical problems, the technical solution of the present invention is as follows:
[0005] A method for dynamically extracting frames from a video for evaluating experimental operation specifications includes the following steps:
[0006] S1. Extract the key frames of different experimental videos and integrate them into a first training set, and cluster each picture in the first training set into different categories as a pseudo-label T c , and extract a binary map containing only the hand from each picture in the first training set through the YCbCr Gaussian skin color algorithm as the pseudo-label T b ;
[0007] S2. Construct a first neural network with a single input and a single output for the following training;
[0008] S3. In the first-stage training, input an open-source second training set into the first neural network and update the parameters of the first neural network until the first preset condition is met, and then stop the first-stage training;
[0009] S4. Based on the training results of the first-stage training, replace the class prediction layer of the first neural network with a fully connected layer of 512 dimensions as the feature output layer, and splice two branches, one of which is a hand feature restoration module and the other is a feature similarity classification module, to form a single-input multi-output neural network, defined as the second neural network;
[0010] S5. In the second-stage training, transfer the pictures of the first training set to the second neural network for training. The pseudo-label T c and the pseudo-label T b are used as supervision signals to calculate the loss of the feature similarity classification module, defined as the classification accuracy loss function, and calculate the Euclidean distance between the hand picture P b predicted by the hand feature restoration module and the pseudo-label T b . Define it as the hand region loss function, set a preset weight for the hand region loss function, calculate the sum of the losses of the classification accuracy loss function and the hand region loss function, and update the parameters of the second neural network until the second preset condition is met to stop the second-stage training;
[0011] S6. In the usage stage, remove the hand feature restoration module and the feature similarity classification module, use the feature output layer as the only feature of the image, take the first frame of the video as the target frame, sequentially traverse the other frames of the video as the frames to be detected, calculate the cosine distance between the features of the frames to be detected and the target frame to obtain the inter-frame similarity. When the inter-frame similarity is less than the preset weight, extract and save the frame to be detected, and replace the frame to be detected with a new target frame to participate in the next calculation of the cosine distance.
[0012] The beneficial effects of the present invention are as follows: Based on the deep neural network to discriminate the similarity of video frames, when processing the experimental operation specification evaluation video, the operation video containing key interactions does not skip frames, and for the video frames of irrelevant operations, they are directly skipped, and the video duration is intelligently fast-forwarded and compressed. Description of the Drawings
[0013] Figure 1 It is a schematic flowchart of the video dynamic frame extraction method for experimental operation specification evaluation disclosed in the embodiment of the present invention;
[0014] Figure 2 It is a network architecture diagram of the second neural network disclosed in the embodiment of the present invention;
[0015] Figure 3 It is an architecture diagram of the hand feature restoration module disclosed in the embodiment of the present invention. Detailed Embodiments
[0016] To make the objectives, technical solutions and advantages of the present invention more clear and definite, the content of the present invention will be further described in detail below with reference to the drawings and specific embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that for the convenience of description, only the parts related to the present invention are shown in the drawings, rather than all the content.
[0017] This embodiment proposes a video dynamic frame extraction method for experimental operation specification evaluation. Based on a deep neural network to discriminate the similarity of video frames, when processing the experimental operation specification evaluation video, the operation video containing key interactions does not skip frames, while the video frames of irrelevant operations are directly skipped, and the video duration is intelligently fast-forwarded and compressed.
[0018] As Figure 1 shown, it includes the following steps:
[0019] S1. Extract the key frames of different experimental videos and integrate them into the first training set. Cluster each image in the first training set into different categories by the K-Means algorithm as the pseudo-label T c , and extract a binary image containing only the hand from each image in the first training set by the YCbCr Gaussian skin color algorithm as the pseudo-label T b . By generating training labels in this way, the manual annotation cost can be reduced, and the algorithm's fast adaptation ability for different experimental scenarios can be improved.
[0020] In step S1, the coordinate points of the YCbCr Gaussian skin color algorithm are respectively set as Cr = [138, 243], Cb = [77, 127].
[0021] S2. Construct a first neural network with a single input and a single output for the following training;
[0022] In step S2, the constructed first neural network can adopt the Resnet50 feature extraction backbone network.
[0023] S3. In the first-stage training, input the open-source second training set into the first neural network and update the parameters of the first neural network until the first preset condition is met, and then stop the first-stage training. The above steps can first let the neural network learn a large number of images without a specific environment, improve the algorithm's basic cognitive ability for texture, boundary and color, and also prevent the overfitting or underfitting problems caused by direct training when the training quantity is not large, and improve the robustness;
[0024] In step S3, the open-source ILSVRC2012 dataset with 1.2 million image data in 1000 categories can be downloaded from the network as the second training set.
[0025] It also includes step S301. After inputting the second training set into the first neural network, calculate the cross-entropy loss function value of each picture according to the predicted category and the true category:
[0026]
[0027] where L is the cross-entropy loss function value, N is the number of pictures in each batch of training of the second training set, G i is the true category of the i-th picture, and p i is the predicted category of the i-th picture.
[0028] Step S302, perform backpropagation on the cross-entropy loss function value L, and then update the parameters of the first neural network.
[0029] The first preset condition includes repeating the steps of S301 for a preset number of times, or stopping the first-stage training when the cross-entropy loss function value L is less than the first threshold.
[0030] S4. Based on the training results of the first stage, replace the category prediction layer of the first neural network with a 512-dimensional fully connected layer as the feature output layer, and splice two branches. One branch is the hand feature restoration module, and the other branch is the feature similarity classification module, forming a single-input multi-output neural network, defined as the second neural network;
[0031] In step S4, the hand feature restoration module is composed of multiple bilinear interpolations for upsampling, aiming to gradually restore the feature map to the input size of the network image, that is, the original picture input provided by the first training set. The specific network structure is as Figure 2 and 3 shown.
[0032] S5. In the second-stage training, transfer the pictures of the first training set to the second neural network for training. The pseudo-labels T c and the pseudo-labels T b are used as supervision signals to calculate the loss of the feature similarity classification module, defined as the classification accuracy loss function. Here, the classification loss function selects arcface loss, which can better improve the aggregation degree of the same category and the difference degree of different categories, and strengthen the feature uniqueness. And calculate the Euclidean distance between the predicted hand picture P b of the hand feature restoration module and the pseudo-label T b , defined as the hand region loss function, because the pseudo-label T bIf it is a binary image with data only in the hand region and 0 in other regions, the algorithm can be guided by this function to only consider the hand region. However, excessive attention should not be paid to the hand. Therefore, a preset weight is set for the loss function of the hand region, and the sum of the classification accuracy loss function and the hand region loss function is calculated. The parameters of the second neural network are updated until the second preset condition is met, and the second stage of training stops;
[0033] In step S5, it further includes step S501. According to the sum of losses L total Perform backpropagation and then update the parameters of the second neural network.
[0034] The classification accuracy loss function is:
[0035]
[0036] where L arcface is the value of the classification accuracy loss function, N is the number of images in each batch of the first training set, i is the i-th image in the first training set, K is the total number of categories in the first training set, k is the k-th category in the first training set, T ik is the k-th annotation result of the i-th image with the true annotation, and P ik is the k-th category result of the i-th image predicted by the network.
[0037] The hand region loss function is:
[0038]
[0039] where L b is the value of the hand region loss function, N is the number of images in each batch of the first training set, i is the i-th image in each batch of training, m is the total number of pixels in the image, j is the j-th pixel corresponding to the image, is the prediction result of the j-th pixel of the image predicted by the neural network, is the prediction result of the j-th pixel of the pseudo-label. The hand region loss function is used to enhance the attention to the hand.
[0040] The sum of losses L total is:
[0041] L total = L arcface + 0.5·L b (4)
[0042] where the final loss is the sum of the losses of the two outputs (formulas (2) and (3)). However, the hand is only used to specify that the model focuses on the hand region. Therefore, the preset weight is 0.5.
[0043] The second preset condition includes repeating the steps of S5 and S501 for a preset number of times, or when the sum of losses Ltotal When it is less than the second threshold, stop the second-stage training.
[0044] S6. In the usage stage, remove the hand feature restoration module and the feature similarity classification module, use the feature output layer as the only feature of the image, take the first frame of the video as the target frame, sequentially traverse the other frames of the video as the frames to be detected, calculate the cosine distance between the features of the frames to be detected and the target frame to obtain the inter-frame similarity. When the inter-frame similarity is less than the preset weight, extract and save the frame to be detected, and replace the frame to be detected with a new target frame to participate in the next calculation of the cosine distance. If it is greater than the preset weight, it is defaulted to be a similar frame and directly skipped.
[0045] In step S6, it is tested with 1000 segments of 10-minute videos of real examination rooms. After processing, the average duration is 2 minutes and 40 seconds, the compression rate is 70%, and the fluency of the key operation segments is the same as that of the original video, greatly improving the evaluation efficiency.
[0046] In this embodiment, first extract the key frames of different experimental videos, and extract a binary image containing only the hand through the YCbCr Gaussian skin color algorithm; construct a single-input multi-output neural network for two-stage training; in the first stage, learn the color and edge information of images without a specific environment through an open-source image classification training set, then replace the class prediction layer with a 512-dimensional fully connected layer as the feature output layer, and splice the hand feature restoration module and the feature similarity classification module. On the basis of image feature classification, assist in improving the algorithm's attention to the hand part and reducing the influence of changes in other lights or irrelevant objects by restoring the hand area of the experimental pictures.
[0047] The above embodiments are only for illustrating the technical concept and features of the present invention, and the purpose is to enable those of ordinary skill in the art to understand the content of the present invention and implement it accordingly. It is not intended to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the essence of the content of the present invention should be covered within the protection scope of the present invention.
Claims
1. A video dynamic frame extraction method for experimental operation specification evaluation, characterized in that It includes the following steps: S1. Extract the key frames of different experimental videos and integrate them into the first training set. Cluster each image in the first training set into different categories by the K-Means algorithm as the pseudo-label T c , and extract a binary image containing only the hand from each image in the first training set by the YCbCr Gaussian skin color algorithm as the pseudo-label T b ; S2. Construct a first neural network with single input and single output for the following training; S3. In the first-stage training, input the open-source second training set into the first neural network, and update the parameters of the first neural network until the first preset condition is met to stop the first-stage training; S4. Based on the training result of the first-stage training, replace the class prediction layer of the first neural network with a 512-dimensional fully connected layer as the feature output layer, and splice two branches, one branch is the hand feature restoration module, and the other branch is the feature similarity classification module, to form a neural network with single input and multiple outputs, defined as the second neural network; S5. In the second-stage training, transfer the pictures of the first training set to the second neural network for training, and use the pseudo-label T c and the pseudo-label T b as supervision signals, calculate the loss of the feature similarity classification module, defined as the classification accuracy loss function, and calculate the Euclidean distance between the hand picture P b predicted by the hand feature restoration module and the pseudo-label T b , defined as the hand region loss function. Set a preset weight for the hand region loss function, calculate the sum of the losses of the classification accuracy loss function and the hand region loss function, and update the parameters of the second neural network until the second preset condition is met to stop the second-stage training; S6. In the using stage, remove the hand feature restoration module and the feature similarity classification module, use the feature output layer as the only feature of the image, take the first frame of the video as the target frame, sequentially traverse the other frames of the video as the frames to be detected, calculate the cosine distance between the features of the frames to be detected and the target frame to obtain the inter-frame similarity. When the inter-frame similarity is less than the preset weight, extract and save the frame to be detected, and replace the frame to be detected with a new target frame to participate in the next calculation of the cosine distance.
2. The video dynamic frame extraction method for experimental operation specification evaluation according to claim 1, wherein, The coordinate points of the YCbCr Gaussian skin color algorithm are respectively set as Cr = [138, 243], Cb = [77, 127].
3. The video dynamic frame extraction method for experimental operation specification evaluation according to claim 1, wherein It also includes step S301. After inputting the second training set into the first neural network, calculate the cross-entropy loss function value of each picture according to the predicted class and the true class: Among them, L is the value of the cross-entropy loss function, N is the number of images in each batch of training of the second training set, G i is the true category of the i-th image, and p i is the predicted category of the i-th image; Step S302. Perform backpropagation on the cross-entropy loss function value, and then update the parameters of the first neural network.
4. The video dynamic frame extraction method for experimental operation specification evaluation according to claim 3, wherein, The first preset condition includes repeating the steps of S301 for a preset number of times, or stopping the first-stage training when the cross-entropy loss function value is less than the first threshold.
5. The video dynamic frame extraction method for experimental operation specification evaluation according to claim 1, characterized in that The hand feature restoration module is composed of multiple bilinear interpolations for upsampling.
6. The video dynamic frame extraction method for experimental operation specification evaluation according to claim 1, characterized in that The classification accuracy loss function is: Among them, L arcface is the value of the classification accuracy loss function, N is the number of images in each batch of training of the first training set, i is the i-th image in each batch of training, K is the total number of categories in the first training set, k is the k-th category in the first training set, T ik is the k-th annotation result of the i-th image with the true annotation, P ik is the k-th category result of the i-th image predicted by the network.
7. The video dynamic frame extraction method for experimental operation specification evaluation according to claim 6, wherein The hand region loss function is: Among them, L b is the loss function value of the hand region, N is the number of pictures in each batch of training of the first training set, i is the i-th picture in each batch of training, m is the total number of pixels of the picture, and j is the j-th pixel corresponding to the picture. is the prediction result of the j-th pixel of the picture predicted by the neural network. is the prediction result of the j-th pixel of the pseudo-label.
8. The video dynamic frame extraction method for experimental operation specification evaluation according to claim 7, wherein The loss sum L total is as follows: L total = L arcface + 0.5·L b Among them, the preset weight is 0.
5.
9. The video dynamic frame extraction method for experimental operation specification evaluation according to claim 8, wherein It further includes step S501, according to the loss sum L total perform backpropagation, and then update the parameters of the second neural network.
10. The video dynamic frame extraction method for experimental operation specification evaluation according to claim 9, wherein, The second preset condition includes repeating the steps of S5 and S501 for a preset number of times, or stopping the second-stage training when the loss sum L total is less than a second threshold value.
Citation Information
Patent Citations
Video crowd counting system and method
CN111860162A
Method for determining attribute characteristics of target video, storage medium and electronic device
CN112270231A