Classroom experiment teaching video abstract extraction method, equipment and medium

Through technical means such as kernel time period segmentation algorithm and deep information extraction model, the problem of low video abstract extraction accuracy in the existing technology is solved, and a more efficient classroom experimental teaching video abstract extraction is achieved.

CN120107861APending Publication Date: 2025-06-06NORTHWEST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510252197.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When processing classroom experimental teaching videos, existing video digest extraction algorithms are difficult to remove similar components of the video to the greatest extent, resulting in low accuracy of the extracted video digest.

Method used

The kernel time period segmentation algorithm is used to divide the video into time period sequences, and the deep information extraction model (including ResNet-50 and Bi-LSTM networks) and the MLP model are used, combined with the K-Means clustering algorithm and fixed threshold method to select keyframes from the video frames to generate a video summary.

Benefits of technology

This method significantly improves the accuracy of the video abstract extraction in classroom experimental teaching, and can capture important content of the video more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107861A_ABST
    Figure CN120107861A_ABST
Patent Text Reader

Abstract

The invention discloses a classroom experiment teaching video abstract extraction method and device and a medium, and relates to the technical field of video abstract extraction, and the method comprises the steps: determining the depth information of a current video frame sequence through employing a kernel time period segmentation algorithm and a depth information extraction model; inputting the depth information into the trained MLP model to obtain a prediction importance score of each video frame in the current video frame sequence; according to the predicted importance score of each video frame in the current video frame sequence, using a K-Means clustering algorithm to select multiple video frames from the current video frames as to-be-determined video key frames; and selecting a plurality of key frames from the to-be-determined video key frames by using a fixed threshold method to obtain a video abstract of the current video segment. Based on the training depth information extraction model and the MLP model, the classroom experiment teaching video abstract extraction precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of video abstract extraction, and in particular to a method, device and medium for extracting abstracts from classroom experimental teaching videos. Background Art

[0002] With the popularization and in-depth application of the Internet, online education has become an important platform for people to learn knowledge, which makes up for the limitations of conducting offline training courses. Whether it is online education or physical education, classroom experimental teaching videos are a commonly used teaching method. A large number of videos increases the difficulty of users' browsing. Video summarization technology allows users to browse the content of videos more effectively, and has received widespread attention in recent years.

[0003] At present, the achievements in video summarization are mainly based on the content analysis of key frames, expanding the key frames with surrounding video clips and linking them together, thus forming a simpler video browsing algorithm. In the dynamic video summarization part, the existing algorithms mainly focus on the similarity analysis at the key frame level. Since this algorithm relies heavily on the selection of key frames. When two similar shots are long in duration and contain large lens motion information, the extracted key frames cannot be guaranteed to be similar enough, but the video sequences represented by these key frames are likely to be very similar. Therefore, only performing redundancy analysis at the video key frame level cannot remove the similar components of the video to the greatest extent, which greatly reduces the accuracy of video summary extraction. Summary of the invention

[0004] The purpose of this application is to provide a method, device and medium for extracting classroom experimental teaching video summaries, which can improve the accuracy of classroom experimental teaching video summaries.

[0005] To achieve the above objectives, this application provides the following solutions:

[0006] In a first aspect, the present application provides a method for extracting a summary of a classroom experimental teaching video, comprising:

[0007] Get the classroom experimental teaching video to be extracted;

[0008] The kernel time segmentation algorithm is used to divide the time series of the classroom experimental teaching video to be extracted into time series;

[0009] Divide the classroom experimental teaching video to be extracted based on the time period sequence to obtain a video segment corresponding to each time period in the time period sequence;

[0010] Determine any video segment as the current video segment;

[0011] Perform frame processing on the current video segment to obtain a current video frame sequence;

[0012] Inputting the current video frame sequence into a trained depth information extraction model to obtain depth information of the current video frame sequence; the depth information extraction model includes a ResNet-50 model and a Bi-LSTM network connected in sequence; the Bi-LSTM network combines an input attention mechanism and a local attention mechanism;

[0013] Inputting the depth information into the trained MLP model to obtain a predicted importance score for each video frame in the current video frame sequence;

[0014] According to the predicted importance score of each video frame in the current video frame sequence, a K-Means clustering algorithm is used to select multiple video frames from the current video frames as the pending video key frames;

[0015] Determine the full-channel first-order color moment and full-channel second-order color moment of each frame of the pending video key frame respectively;

[0016] According to the full-channel first-order color moment and full-channel second-order color moment of each pending video key frame, a fixed threshold method is used to select multiple key frames from the pending video key frames to obtain the video summary of the current video segment.

[0017] Optionally, a kernel time segmentation algorithm is used to divide the time series of the classroom experimental teaching video to be extracted into a time series, including:

[0018] The time series of classroom experimental teaching videos to be extracted are mapped to a high-dimensional feature space, and the similarities of all subsequences in the time series are calculated through the kernel function to obtain multiple similarity matrices;

[0019] Clustering the similarity matrix using a density-based clustering algorithm to obtain multiple first cluster centers;

[0020] Assign each time series point in the time series of the extracted classroom experimental teaching video to the cluster where the nearest first cluster center is located, and obtain the cluster assignment vector corresponding to each first cluster center;

[0021] The time series of the classroom experimental teaching videos to be extracted are divided based on a plurality of the cluster allocation vectors to obtain a time period series.

[0022] Optionally, inputting the current video frame sequence into a trained depth information extraction model to obtain depth information of the current video frame sequence includes:

[0023] Inputting the current video frame sequence into the trained ResNet-50 model, wherein the average pooling layer in the trained ResNet-50 model outputs the deep spatial features of each video frame in the current video frame sequence;

[0024] Using the trained Bi-LSTM network, determine the first attention weight of each video frame in the current video frame sequence;

[0025] The deep spatial features are weighted summed according to the first attention weight of each video frame in the current video frame sequence to obtain a motion feature vector;

[0026] Inputting the current video frame sequence into the trained Bi-LSTM network to obtain an output feature sequence;

[0027] According to the output feature sequence, a context relevance score between each output feature in the output feature sequence and the output feature sequence is calculated;

[0028] Determine a second attention weight corresponding to each output feature based on multiple context relevance scores;

[0029] Use the local attention mechanism to update the second attention weight to obtain the third attention weight corresponding to each output feature;

[0030] Based on the third attention weight corresponding to each output feature, the output feature sequence is weighted summed to obtain the context vector;

[0031] The context vector is summed with the motion feature vector to obtain the depth information.

[0032] Optionally, the first attention weight is:

[0033]

[0034] Among them, α i,t is the first attention weight of the i-th video frame at time t; e i,t is the correlation score of the spatial features of the i-th video frame at time t, is the deep spatial feature x of the i-th video frame i The transpose of W e is the weight matrix parameter, h t-1 is the hidden state at time t-1 output by the trained Bi-LSTM network; e j,t is the correlation score of the spatial features of the j-th video frame at time t; N is the number of video frames in the current video frame sequence;

[0035] The contextual relevance score is:

[0036]

[0037] Among them, l i,t is the context relevance score between the i-th output feature and the output feature sequence at time t; W a and Ua are weight parameters to be learned, h i is the i-th output feature; h t Output feature sequence of the trained Bi-LSTM network at time t;

[0038] The second attention weight is:

[0039]

[0040] Among them, β i,t is the second attention weight corresponding to the i-th output feature at time t; l j,t is the context relevance score between the j-th output feature and the output feature sequence at time t;

[0041] The third attention weight is:

[0042]

[0043] β′ i,t is the third attention weight corresponding to the i-th output feature at time t; σ is the standard deviation of the attention weight; p t The local attention mechanism at time t reduces the intermediate amount of computational complexity, and W p are the learned parameters used by the model to predict position.

[0044] Optionally, according to the predicted importance score of each video frame in the current video frame sequence, a K-Means clustering algorithm is used to select multiple video frames from the current video frame as the pending video key frames, including:

[0045] Arrange all video frames in the current video frame sequence in descending order according to the predicted importance scores;

[0046] Acquiring a preset number of video frames as a first pending video key frame set;

[0047] Using the K-Means clustering algorithm to cluster all video frames in the current video frame sequence, a plurality of second clustering centers are obtained;

[0048] Determine the video frames corresponding to all the second cluster centers as the second to-be-determined video key frame set;

[0049] A union of the first pending video key frame set and the second pending video key frame set is determined as the pending video key frame.

[0050] Optionally, the full-channel first-order color moment includes: a red channel first-order color moment, a green channel first-order color moment, and a blue channel first-order color moment; the full-channel second-order color moment includes: a red channel second-order color moment, a green channel second-order color moment, and a blue channel second-order color moment;

[0051] The first-order color moment of the red channel is:

[0052]

[0053] Among them, μ R is the first-order color moment of the red channel of the key frame of the video to be determined; N 1 is the number of pixels in the key frame of the pending video; r i is the red pixel value of the i-th pixel in the key frame of the video to be determined;

[0054] The first-order color moment of the green channel is:

[0055]

[0056] Among them, μ G is the first-order color moment of the green channel of the key frame of the video to be determined; g i is the green pixel value of the i-th pixel in the key frame of the video to be determined;

[0057] The second-order color moment of the blue channel is:

[0058]

[0059] Among them, μ B is the first-order color moment of the blue channel of the key frame of the video to be determined; b i is the blue pixel value of the i-th pixel in the key frame of the video to be determined;

[0060] The second-order color moment of the red channel is:

[0061]

[0062] in, The second-order color moment of the red channel of the key frame of the video to be determined;

[0063] The second-order color moment of the green channel is:

[0064]

[0065] in, is the second-order color moment of the green channel of the key frame of the video to be determined;

[0066] The second-order color moment of the blue channel is:

[0067]

[0068] in, It is the second-order color moment of the blue channel of the key frame of the video to be determined.

[0069] Optionally, according to the full-channel first-order color moment and the full-channel second-order color moment of each frame of the pending video key frame, a fixed threshold method is used to select multiple key frames from the pending video key frames to obtain a video summary of the current video segment, including:

[0070] Determine the empty set as the key frame set;

[0071] Let the key frame number of the pending video i=1;

[0072] Determine the first Euclidean distance according to the full-channel first-order color moment of the (i+1)th frame of the pending video key frame and the full-channel first-order color moment of the (i)th frame of the pending video key frame;

[0073] Determine the second Euclidean distance according to the full-channel second-order color moment of the (i+1)th frame of the pending video key frame and the full-channel second-order color moment of the (i)th frame of the pending video key frame;

[0074] When a key frame condition is met, determining the i-th frame of the pending video key frame and the i+1-th frame of the pending video key frame as the current key frame; the key frame condition is that the first Euclidean distance is greater than the first threshold, and the second Euclidean distance is greater than the second threshold;

[0075] Add the current key frame as the last element to the key frame set;

[0076] When the key frame condition is not met, the i-th frame of the pending video key frame is determined to be a non-key frame;

[0077] Increase the value of the sequence number i of the pending video key frame by 1 and return to the step "determine the first Euclidean distance according to the full-channel first-order color moment of the i+1th frame of the pending video key frame and the full-channel first-order color moment of the i-th frame of the pending video key frame" until all the pending video key frames are traversed and the key frame set is determined to be the video summary of the current video segment.

[0078] Optionally, the first Euclidean distance is:

[0079]

[0080] Among them, |d i1 | is the first Euclidean distance; μ R(i+1) is the first-order color moment of the red channel of the i+1th frame of the pending video key frame; μ Ri is the first-order color moment of the red channel of the i-th frame of the pending video key frame; μ G(i+1) is the first-order color moment of the green channel of the i+1th frame of the undetermined video key frame; μGi is the first-order color moment of the green channel of the i-th frame of the pending video key frame; μ B(i+1) is the first-order color moment of the blue channel of the i+1th frame of the undetermined video key frame; μ Bi is the first-order color moment of the blue channel of the i-th frame of the pending video key frame;

[0081] The second Euclidean distance is:

[0082]

[0083] Among them, |d i2 | is the second Euclidean distance; The second-order color moment of the red channel of the i+1th frame of the undetermined video key frame; is the second-order color moment of the red channel of the i-th frame of the pending video key frame; is the second-order color moment of the green channel of the i+1th frame of the undetermined video key frame; is the second-order color moment of the green channel of the i-th frame of the pending video key frame; is the second-order color moment of the blue channel of the i+1th frame of the undetermined video key frame; is the second-order color moment of the blue channel of the i-th frame of the pending video key frame.

[0084] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned classroom experimental teaching video summary extraction method.

[0085] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned classroom experimental teaching video summary extraction method.

[0086] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0087] The present application provides a method, device and medium for extracting summary of classroom experimental teaching videos, which uses a kernel time segmentation algorithm and a depth information extraction model to determine the depth information of the current video frame sequence; inputs the depth information into a trained MLP model to obtain a predicted importance score of each video frame in the current video frame sequence; based on the predicted importance score of each video frame in the current video frame sequence, a K-Means clustering algorithm is used to select multiple video frames from the current video frame as pending video key frames; a fixed threshold method is used to select multiple key frames from the pending video key frames to obtain a video summary of the current video segment. The present application improves the accuracy of extracting summary of classroom experimental teaching videos based on the training of a depth information extraction model and an MLP model. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0089] Figure 1 This is a flow chart of a method for extracting summary of classroom experimental teaching videos in one embodiment of the present application;

[0090] Figure 2 This is a framework diagram of a method for extracting summary of classroom experimental teaching videos in one embodiment of the present application;

[0091] Figure 3 Schematic diagram of the Bi-LSTM network structure in one embodiment of the present application. DETAILED DESCRIPTION

[0092] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0093] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0094] In an exemplary embodiment, Figure 1 As shown, a method for extracting a summary of a classroom experimental teaching video is provided, comprising:

[0095] Step 101: Obtain the classroom experiment teaching video to be extracted.

[0096] Step 102: using a kernel time segmentation algorithm, the time series of the classroom experimental teaching video to be extracted is divided into time series.

[0097] Step 103: Divide the classroom experimental teaching video to be extracted based on the time period sequence to obtain the video segment corresponding to each time period in the time period sequence.

[0098] Step 104: Determine any video segment as the current video segment.

[0099] Step 105: The current video segment is divided into frames to obtain a current video frame sequence.

[0100] Step 106: Input the current video frame sequence into the trained depth information extraction model to obtain the depth information of the current video frame sequence. The depth information extraction model includes a ResNet-50 model and a Bi-LSTM network connected in sequence. The Bi-LSTM network combines the input attention mechanism and the local attention mechanism.

[0101] Step 107: Input the depth information into the trained MLP model to obtain the predicted importance score of each video frame in the current video frame sequence.

[0102] Step 108: According to the predicted importance score of each video frame in the current video frame sequence, a K-Means clustering algorithm is used to select multiple video frames from the current video frames as pending video key frames.

[0103] Step 109: respectively determine the full-channel first-order color moment and the full-channel second-order color moment of each key frame of the pending video.

[0104] Step 1010: According to the full-channel first-order color moment and the full-channel second-order color moment of each frame of the pending video key frame, a fixed threshold method is used to select multiple key frames from the pending video key frames to obtain a video summary of the current video segment.

[0105] Step 102 includes:

[0106] Step 102 - 1: Map the time series of the classroom experimental teaching videos to be extracted to a high-dimensional feature space, calculate the similarity of all subsequences in the time series through a kernel function, and obtain multiple similarity matrices.

[0107] Step 102 - 2 : clustering the similarity matrix using a density-based clustering algorithm to obtain a plurality of first cluster centers.

[0108] Step 102 - 3: assign each time series point in the time series of the extracted classroom experimental teaching video to the cluster where the nearest first cluster center is located, and obtain a cluster assignment vector corresponding to each first cluster center.

[0109] Step 102 - 4 : Divide the time series of the classroom experimental teaching videos to be extracted based on multiple cluster allocation vectors to obtain a time period sequence.

[0110] Step 106 includes:

[0111] Step 106-1: Input the current video frame sequence into the trained ResNet-50 model, and the average pooling layer in the trained ResNet-50 model outputs the deep spatial features of each video frame in the current video frame sequence.

[0112] Step 106-2: Using the trained Bi-LSTM network, determine the first attention weight of each video frame in the current video frame sequence.

[0113] Step 106-3: Perform weighted summation on the deep spatial features according to the first attention weight of each video frame in the current video frame sequence to obtain a motion feature vector.

[0114] Step 106-4: Input the current video frame sequence into the trained Bi-LSTM network to obtain an output feature sequence.

[0115] Step 106 - 5 : According to the output feature sequence, a context relevance score between each output feature in the output feature sequence and the output feature sequence is calculated.

[0116] Step 106-6: Determine a second attention weight corresponding to each output feature based on multiple context relevance scores.

[0117] Step 106-7: Use the local attention mechanism to update the second attention weight to obtain the third attention weight corresponding to each output feature.

[0118] Step 106-8: Based on the third attention weight corresponding to each output feature, perform weighted summation on the output feature sequence to obtain a context vector.

[0119] Step 106-9: sum the context vector and the motion feature vector to obtain depth information.

[0120] Among them, the first attention weight is:

[0121]

[0122] Among them, α i,t is the first attention weight of the i-th video frame at time t. i,t is the correlation score of the spatial features of the i-th video frame at time t, is the deep spatial feature x of the i-th video frame i The transpose of W e is the weight matrix parameter, h t-1 is the hidden state at time t-1 output by the trained Bi-LSTM network. j,t is the correlation score of the spatial features of the jth video frame at time t. N is the number of video frames in the current video frame sequence.

[0123] The contextual relevance scores are:

[0124]

[0125] Among them, l i,tis the context relevance score between the i-th output feature and the output feature sequence at time t. W a and U a are weight parameters to be learned, h i is the i-th output feature. t Output feature sequence for the trained Bi-LSTM network at time t.

[0126] The second attention weight is:

[0127]

[0128] Among them, β i,t is the second attention weight corresponding to the i-th output feature at time t. j,t is the context relevance score between the j-th output feature and the output feature sequence at time t.

[0129] The third attention weight is:

[0130]

[0131] β′ i,t is the third attention weight corresponding to the i-th output feature at time t. σ is the standard deviation of the attention weight. t The local attention mechanism at time t reduces the intermediate amount of computational complexity, and W p are the learned parameters used by the model to predict position.

[0132] Step 108 includes:

[0133] Step 108 - 1 : Arrange all video frames in the current video frame sequence in descending order according to the predicted importance scores.

[0134] Step 108 - 2 : Acquire a preset number of video frames as a first to-be-determined video key frame set.

[0135] Step 108 - 3 : clustering all video frames in the current video frame sequence using the K-Means clustering algorithm to obtain a plurality of second cluster centers.

[0136] Step 108 - 4 : Determine the video frames corresponding to all the second cluster centers as the second to-be-determined video key frame set.

[0137] Step 108 - 5 : Determine the union of the first pending video key frame set and the second pending video key frame set as the pending video key frame.

[0138] The first-order color moment of all channels includes: the first-order color moment of the red channel, the first-order color moment of the green channel, and the first-order color moment of the blue channel. The second-order color moment of all channels includes: the second-order color moment of the red channel, the second-order color moment of the green channel, and the second-order color moment of the blue channel.

[0139] The first-order color moment of the red channel is:

[0140]

[0141] Among them, μ R N is the first-order color moment of the red channel of the key frame of the video to be determined. 1 is the number of pixels in the key frame of the pending video. i is the red pixel value of the i-th pixel in the pending video key frame.

[0142] The first-order color moment of the green channel is:

[0143]

[0144] Among them, μ G is the first-order color moment of the green channel of the key frame of the video to be determined. i is the green pixel value of the i-th pixel in the pending video key frame.

[0145] The second-order color moment of the blue channel is:

[0146]

[0147] Among them, μ B is the first-order color moment of the blue channel of the key frame of the video to be determined. i is the blue pixel value of the i-th pixel in the pending video key frame.

[0148] The second-order color moment of the red channel is:

[0149]

[0150] in, It is the second-order color moment of the red channel of the key frame of the video to be determined.

[0151] The second-order color moment of the green channel is:

[0152]

[0153] in, It is the second-order color moment of the green channel of the key frame of the video to be determined.

[0154] The second-order color moment of the blue channel is:

[0155]

[0156] in, It is the second-order color moment of the blue channel of the key frame of the video to be determined.

[0157] Step 1010 includes:

[0158] Step 1010 - 1 : Determine that the empty set is a key frame set.

[0159] Step 1010 - 2: Set the sequence number of the key frame of the video to be determined i=1.

[0160] Step 1010 - 3 : Determine the first Euclidean distance according to the full-channel first-order color moment of the (i+1)th frame of the pending video key frame and the full-channel first-order color moment of the (i)th frame of the pending video key frame.

[0161] Step 1010 - 4 : Determine the second Euclidean distance according to the full-channel second-order color moment of the (i+1)th frame of the pending video key frame and the full-channel second-order color moment of the (i)th frame of the pending video key frame.

[0162] Step 1010-5: When the key frame condition is met, the i-th frame to be determined video key frame and the i+1-th frame to be determined video key frame are determined as current key frames. The key frame condition is that the first Euclidean distance is greater than the first threshold, and the second Euclidean distance is greater than the second threshold.

[0163] Step 1010-6: Add the current key frame as the last element to the key frame set.

[0164] Step 1010-7: When the key frame condition is not met, determine that the i-th frame of the pending video key frame is a non-key frame.

[0165] Step 1010 - 8 : Increase the value of the pending video key frame sequence number i by 1 and return to step 1010 - 4 until all pending video key frames are traversed and the key frame set is determined to be the video summary of the current video segment.

[0166] The first Euclidean distance is:

[0167]

[0168] Among them, |d i1 | is the first Euclidean distance. μ R(i+1) μ is the first-order color moment of the red channel of the i+1th frame of the undetermined video key frame. Ri μ is the first-order color moment of the red channel of the i-th frame of the pending video key frame. G(i+1) μ is the first-order color moment of the green channel of the i+1th frame of the undetermined video key frame. Gi μ is the first-order color moment of the green channel of the i-th frame of the undetermined video key frame. B(i+1)μ is the first-order color moment of the blue channel of the i+1th frame of the undetermined video key frame. Bi is the first-order color moment of the blue channel of the i-th frame of the pending video key frame.

[0169] The second Euclidean distance is:

[0170]

[0171] Among them, |d i2 | is the second Euclidean distance. It is the second-order color moment of the red channel of the i+1th frame of the pending video key frame. It is the second-order color moment of the red channel of the i-th frame of the pending video key frame. It is the second-order color moment of the green channel of the i+1th frame of the pending video key frame. It is the second-order color moment of the green channel of the i-th frame of the pending video key frame. It is the second-order color moment of the blue channel of the i+1th frame of the pending video key frame. is the second-order color moment of the blue channel of the i-th frame of the pending video key frame.

[0172] This embodiment uses classroom experimental teaching videos as the research object and intelligent video summary optimization model as the research topic. The overall framework diagram is as follows: Figure 2 .

[0173] There are three main sections: First, extract the deep spatial features of video frames based on the Bi-LSTM network. Second, use MLP and K-means clustering algorithms to extract video key frames. Third, use the fixed threshold method for threshold detection to further streamline the key frames to obtain the video summary. The specific steps are as follows:

[0174] Step 1: Split the time segment using the Kernel Time Segmentation (KTS) algorithm.

[0175] Kernel Temporal Segmentation (KTS) is a time series segmentation technique based on the kernel method, which can segment the time series into different segments, each with different characteristics. First, the classroom experimental teaching video (AVI) time series is transformed by kernel function and mapped to a high-dimensional space. Then, a density-based clustering algorithm is used to cluster the time series in the high-dimensional space to obtain a set of time periods. Finally, the final time period segmentation is determined by analyzing the clustering results. The algorithm generates a list of time periods arranged in chronological order. Each time period represents a relatively independent content segment, that is, a collection of video frames. The start time and end time of the interval accurately define the position of the segment in the original video. The details are as follows:

[0176] 1. Map the original time series X to a high-dimensional feature space F, and then use the kernel function (x i ,x j ) Calculate the similarity of all subsequences in the time series and obtain the similarity matrix K: K i,j = k(x i ,x j ).

[0177] 2. Use the clustering algorithm to cluster the similarity matrix K and obtain k cluster centers C =

[0178] {c 1 ,c 2 ,...,c k} to identify different segments. Then all time series points x i Assign it to the cluster center c that is closest to it j In the cluster where it is located, the cluster assignment vector z is obtained i :z i =argmin j=1,2,...,k ||x i -c j ||^2.

[0179] Using clustering to assign vector z i , the time series X is divided into k different segments, and the characteristics of each segment are determined by the position of the corresponding cluster center and the value of the time series point assigned to the cluster. The kernel function uses the Gaussian kernel function.

[0180] Step 2: Extract deep spatial features of video frames based on the Bi-LSTM network.

[0181] The video frames obtained above are input into the ResNet-50 model in sequence, and the output of the AveragePool layer is used as the deep spatial feature information. The feature sequence of all video frames is X = {x 1 ,x 2 ,…,x t ,…,x N}, where x t is the deep spatial feature information extracted from the Nth frame of the original video, where N is the number of frames of the input classroom video. Then, an input attention mechanism is added to calculate the contribution of each region in the input video frame, and then adjust the weight of each region according to the contribution, so that the network pays more attention to the target region.

[0182] The deep spatial features x of the i-th frame i , and the hidden state h of the Bi-LSTM network at the previous moment t-1 Combined, the correlation score of the spatial features of the tth frame at the i-th time step is calculated:

[0183]

[0184] where e i,t is the correlation function calculated using bilinear attention, W e are the weight matrix parameters that need to be learned.

[0185]

[0186] α i,t is the attention weight, which is the i,t Normalized using the softmax function to ensure that the sum of all attention weights is 1. Then use the following formula to perform weighted summation on the deep spatial feature sequence input to obtain the motion feature vector of the t-th time step:

[0187]

[0188] Bi-LSTM is an improvement on unidirectional LSTM. The structure of Bi-LSTM is as follows: Figure 3 As shown in Figure 1, it consists of a forward LSTM layer and a reverse LSTM layer. The forward LSTM layer processes the input sequence from front to back in the normal order of the sequence, while the reverse LSTM layer processes the input sequence from back to front in the reverse order. Therefore, the contextual information before and after the input sequence can be captured at the same time. Specifically, given the input visual feature sequence in is the spatial feature extracted from the original spatial feature of the video frame t after being processed by the input attention mechanism. The forward LSTM starts from the first element x of the input sequence 1 Start processing to x N The backward LSTM starts from the last element x N Process to x 1 .

[0189] h t =f LSTM (h t-1 ,x t ).

[0190] h' t =f LSTM (h' t+1 ,x t ).

[0191] where h t With h' t are the states of the forward LSTM and backward LSTM hidden layers at time t, h t-1 With h' t+1 They are the states of the hidden layers of the forward LSTM at time t-1 and the backward LSTM at time t+1. LSTM() is the calculation process of the LSTM network, which is as follows:

[0192] i t =σ(W i [h t-1 ,x t ]+b i ).

[0193] f t =σ(W f [h t-1 ,x t ]+b f ).

[0194] o t =σ(W o [h t-1 ,x t ]+b o ).

[0195]

[0196] h t =o t tanh(c t ).

[0197] Among them, i t 、f t , o t Represent the output values ​​of the input gate, forget gate and output gate respectively, represents the candidate value of the memory cell at the current moment, c t Indicates the value of the memory cell at the current moment, h t Represents the output at the current moment.

[0198] The output h of the Bi-LSTM network at time t t It is formed by adding the output of the forward LSTM and the backward LSTM. The formula is as follows:

[0199] h t =W ht (h t ,h' t )+b t .

[0200] Where W ht is the weight of Bi-LSTM at time t, (h t , h' t ) is the concatenation of the outputs of the forward LSTM and the backward LSTM at time t, b t is the bias of the Bi-LSTM network output at time t.

[0201] Then, the output feature h of the Bi-LSTM network is calculated according to the method proposed by Bahdanau et al. i Correlation score with the entire output sequence:

[0202]

[0203] Among them l i,t is the additive attention model relevance function, W a and U a is the weight parameter that needs to be learned. i,t The attention weight represents the importance of the output vector of the i-th Bi-LSTM network to the output sequence at time t.

[0204] Then the local attention mechanism is used to reduce the computational complexity, as shown below:

[0205]

[0206] With W p is the model parameter learned to predict the position. N is the number of video frames. Under the action of the sigmoid function, the attention weight is updated as:

[0207]

[0208] After obtaining the attention weights of all time steps, the outputs of all Bi-LSTM {h 1 ,h 2 ,…,h N}, perform weighted summation to obtain the context vector c t , which is different in each time step.

[0209]

[0210] The context vector c t And the motion feature vector after using the input attention mechanism Add together to get the depth information d of the motion video frame t .

[0211] Step 3: Select the video keyframe.

[0212] Use MLP to extract the depth information d of motion video frames t Get the frame-level importance score of the video frame. The output of MLP is a scalar: y t =f l (d t). The MLP consists of an input layer, an output layer, and a hidden layer. The hidden layer has 512 neurons, and the output layer has only one neuron. The loss function uses the mean square error loss function:

[0213] Where N is the number of video frames, represents the true frame-level importance score of the t-th video frame, y t Represents the predicted value of the importance score of the tth video frame. The larger the MSE, the worse the model fit. The smaller the MSE, the better the model fit.

[0214] After obtaining the predicted importance scores of all video frames, all video frames are sorted in descending order according to the frame-level importance scores. However, only using dynamic programming to select key video frames may result in a high recall rate but low precision, resulting in a lack of diversity in the generated summary and a large number of redundant frames.

[0215] In order to eliminate redundant frames and increase the diversity of selected frames, the model is combined with the K-Means clustering algorithm. Given a set Z of all video frames, the number of video frames is N, k clustering centers are set, and the key frames of each cluster are obtained through continuous iteration. The corresponding mathematical expression describing the algorithm is as follows:

[0216] For the input sample set D, an initial partition C is performed on it:

[0217] D = {x 1, x 2 …,x m}.

[0218] C={C 1 ,C 2 …,C k}.

[0219] For the current partition C, the following is iterated:

[0220]

[0221] in:

[0222] In the formula, u i is the cluster center of the i-th cluster Ci, and ‖…‖^2 is the Euclidean distance.

[0223] After the above initial operations, the K-means clustering algorithm can select the most appropriate central elements of K clusters, that is, obtain the k "most representative" key frames in N video frames.

[0224] Step 4: Compare and screen frames to determine the final key frame.

[0225] In order to finally determine the generation of a reasonable and concise video summary, it is necessary to make a final comparison and screening of the acquired key frame sequence. Color moment is a simple and effective color feature representation method based on the statistical moment of image color distribution. Color moment can describe the overall characteristics of image color to a certain extent, including the average distribution of color, the degree of color dispersion, and the skewness of color distribution. In the comparison of adjacent frames, the change of color moment can reflect the change of color content in video. The color moment comparison of two adjacent frames is performed. Because the experimental teaching video frame does not change much but small changes are important, the change gap threshold should be small, so the first-order color moment is set to about 5-10, and the second-order color moment is set to 10-20.

[0226] For an image in RGB color space, let the total number of pixels in the image be N and the red channel pixel value be r i , the green channel pixel value is g i , the blue channel pixel value is b i .

[0227] Then the calculation formulas for the first-order color moment (average color) in the red channel, green channel, and blue channel are:

[0228]

[0229] For two adjacent frames of images, let the first-order color moment of the first frame in the RGB channel be μ R1 ,μ G1 ,μ B1 ), the second frame is (μ R2 ,μ G2 ,μ B2 ), the Euclidean distance is used to measure the difference, and the calculation formula is:

[0230]

[0231] The calculation formula of the second-order color moment (color variance) in the red channel is:

[0232] Among them, μ R is the first-order color moment (average color) of the red channel. Similarly, the color variance of the green channel and the blue channel can be calculated. When comparing the second-order color moments of adjacent frames, for two adjacent frames, let the second-order color moment of the first frame in the RGB channel be The second frame is The difference is measured by Euclidean distance, which is calculated as:

[0233]

[0234] When |d 1|with|d 2 If | is greater than the threshold, it is a key frame and added to the key frame sequence, otherwise it is not. Finally, the key frame sequence is output to obtain the video summary.

[0235] In an exemplary embodiment, a computer device is provided, which may be a server or a terminal. The computer device includes a processor, a memory, an input / output interface (I / O for short) and a communication interface. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for extracting a summary of a classroom experimental teaching video is implemented.

[0236] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0237] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0238] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0239] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0240] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.

[0241] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0242] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application; at the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A method for extracting summary of classroom experimental teaching videos, characterized in that: include: Get the classroom experimental teaching video to be extracted; The kernel time segmentation algorithm is used to divide the time series of the classroom experimental teaching video to be extracted into time series; Divide the classroom experimental teaching video to be extracted based on the time period sequence to obtain a video segment corresponding to each time period in the time period sequence; Determine any video segment as the current video segment; Perform frame processing on the current video segment to obtain a current video frame sequence; Inputting the current video frame sequence into a trained depth information extraction model to obtain depth information of the current video frame sequence; the depth information extraction model includes a ResNet-50 model and a Bi-LSTM network connected in sequence; The Bi-LSTM network combines input attention mechanism and local attention mechanism; Inputting the depth information into the trained MLP model to obtain a predicted importance score for each video frame in the current video frame sequence; According to the predicted importance score of each video frame in the current video frame sequence, a K-Means clustering algorithm is used to select multiple video frames from the current video frames as the pending video key frames; Determine the full-channel first-order color moment and full-channel second-order color moment of each frame of the pending video key frame respectively; According to the full-channel first-order color moment and full-channel second-order color moment of each pending video key frame, a fixed threshold method is used to select multiple key frames from the pending video key frames to obtain the video summary of the current video segment.

2. The classroom experimental teaching video abstract extraction method according to claim 1 is characterized in that: Using the kernel time segmentation algorithm, the time series of the classroom experimental teaching video to be extracted is divided into time series, including: The time series of classroom experimental teaching videos to be extracted are mapped to a high-dimensional feature space, and the similarities of all subsequences in the time series are calculated through the kernel function to obtain multiple similarity matrices; Clustering the similarity matrix using a density-based clustering algorithm to obtain multiple first cluster centers; Assign each time series point in the time series of the extracted classroom experimental teaching video to the cluster where the nearest first cluster center is located, and obtain the cluster assignment vector corresponding to each first cluster center; The time series of the classroom experimental teaching videos to be extracted are divided based on a plurality of the cluster allocation vectors to obtain a time period series.

3. The classroom experimental teaching video abstract extraction method according to claim 1 is characterized in that: Inputting the current video frame sequence into the trained depth information extraction model to obtain the depth information of the current video frame sequence includes: Inputting the current video frame sequence into the trained ResNet-50 model, wherein the average pooling layer in the trained ResNet-50 model outputs the deep spatial features of each video frame in the current video frame sequence; Using the trained Bi-LSTM network, determine the first attention weight of each video frame in the current video frame sequence; The deep spatial features are weighted summed according to the first attention weight of each video frame in the current video frame sequence to obtain a motion feature vector; Inputting the current video frame sequence into the trained Bi-LSTM network to obtain an output feature sequence; According to the output feature sequence, a context relevance score between each output feature in the output feature sequence and the output feature sequence is calculated; Determine a second attention weight corresponding to each output feature based on multiple context relevance scores; Use the local attention mechanism to update the second attention weight to obtain the third attention weight corresponding to each output feature; Based on the third attention weight corresponding to each output feature, the output feature sequence is weighted summed to obtain the context vector; The context vector is summed with the motion feature vector to obtain the depth information.

4. The classroom experimental teaching video abstract extraction method according to claim 3 is characterized in that: The first attention weight is: Among them, α i,t is the first attention weight of the i-th video frame at time t; e i,t is the correlation score of the spatial features of the i-th video frame at time t, is the deep spatial feature x of the i-th video frame i The transpose of W e is the weight matrix parameter, h t-1 is the hidden state at time t-1 output by the trained Bi-LSTM network; e j,t is the correlation score of the spatial features of the j-th video frame at time t; N is the number of video frames in the current video frame sequence; The contextual relevance score is: Among them, l i,t is the context relevance score between the i-th output feature and the output feature sequence at time t; W a and U a are weight parameters to be learned, h i is the i-th output feature; h t Output feature sequence of the trained Bi-LSTM network at time t; The second attention weight is: Among them, β i,t is the second attention weight corresponding to the i-th output feature at time t; l j,t is the context relevance score between the j-th output feature and the output feature sequence at time t; The third attention weight is: β′ i,t is the third attention weight corresponding to the i-th output feature at time t; σ is the standard deviation of the attention weight; p t The local attention mechanism at time t reduces the intermediate amount of computational complexity, and W p are the learned parameters of the model used to predict the position.

5. The classroom experimental teaching video abstract extraction method according to claim 1 is characterized in that: According to the predicted importance score of each video frame in the current video frame sequence, a K-Means clustering algorithm is used to select multiple video frames from the current video frame as the pending video key frames, including: Arrange all video frames in the current video frame sequence in descending order according to the predicted importance scores; Acquiring a preset number of video frames as a first pending video key frame set; Using the K-Means clustering algorithm to cluster all video frames in the current video frame sequence, a plurality of second clustering centers are obtained; Determine the video frames corresponding to all the second cluster centers as the second to-be-determined video key frame set; A union of the first pending video key frame set and the second pending video key frame set is determined as the pending video key frame.

6. The classroom experimental teaching video abstract extraction method according to claim 1 is characterized in that: The full-channel first-order color moment includes: the red channel first-order color moment, the green channel first-order color moment and the blue channel first-order color moment; the full-channel second-order color moment includes: the red channel second-order color moment, the green channel second-order color moment and the blue channel second-order color moment; The first-order color moment of the red channel is: Among them, μ R is the first-order color moment of the red channel of the key frame of the video to be determined; N1 is the number of pixels in the key frame of the video to be determined; r i is the red pixel value of the i-th pixel in the key frame of the video to be determined; The first-order color moment of the green channel is: Among them, μ G is the first-order color moment of the green channel of the key frame of the video to be determined; g i is the green pixel value of the i-th pixel in the key frame of the video to be determined; The second-order color moment of the blue channel is: Among them, μ B is the first-order color moment of the blue channel of the key frame of the video to be determined; b i is the blue pixel value of the i-th pixel in the key frame of the video to be determined; The second-order color moment of the red channel is: in, The second-order color moment of the red channel of the key frame of the video to be determined; The second-order color moment of the green channel is: in, is the second-order color moment of the green channel of the key frame of the video to be determined; The second-order color moment of the blue channel is: in, It is the second-order color moment of the blue channel of the key frame of the video to be determined.

7. The classroom experimental teaching video abstract extraction method according to claim 6 is characterized in that: According to the full-channel first-order color moment and full-channel second-order color moment of each frame of the pending video key frame, multiple key frames are selected from the pending video key frames using a fixed threshold method to obtain a video summary of the current video segment, including: Determine the empty set as the key frame set; Let the key frame sequence number of the pending video i=1; Determine the first Euclidean distance according to the full-channel first-order color moment of the (i+1)th frame of the pending video key frame and the full-channel first-order color moment of the (i)th frame of the pending video key frame; Determine the second Euclidean distance according to the full-channel second-order color moment of the (i+1)th frame of the pending video key frame and the full-channel second-order color moment of the (i)th frame of the pending video key frame; When a key frame condition is met, determining the i-th frame of the pending video key frame and the i+1-th frame of the pending video key frame as the current key frame; the key frame condition is that the first Euclidean distance is greater than the first threshold, and the second Euclidean distance is greater than the second threshold; Add the current key frame as the last element to the key frame set; When the key frame condition is not met, the i-th frame of the pending video key frame is determined to be a non-key frame; Increase the value of the sequence number i of the pending video key frame by 1 and return to step "determine the first Euclidean distance according to the full-channel first-order color moment of the i+1th frame of the pending video key frame and the full-channel first-order color moment of the i-th frame of the pending video key frame" until all the pending video key frames are traversed and the key frame set is determined to be the video summary of the current video segment.

8. The classroom experimental teaching video abstract extraction method according to claim 7 is characterized in that: The first Euclidean distance is: Among them, |d i1 | is the first Euclidean distance; μ R(i+1) is the first-order color moment of the red channel of the i+1th frame of the pending video key frame; μ Ri is the first-order color moment of the red channel of the i-th frame of the pending video key frame; μ G(i+1) is the first-order color moment of the green channel of the i+1th frame of the undetermined video key frame; μ Gi is the first-order color moment of the green channel of the i-th frame of the pending video key frame; μ B(i+1) is the first-order color moment of the blue channel of the i+1th frame of the undetermined video key frame; μ Bi is the first-order color moment of the blue channel of the i-th frame of the pending video key frame; The second Euclidean distance is: Among them, |d i2 | is the second Euclidean distance; The second-order color moment of the red channel of the i+1th frame of the undetermined video key frame; is the second-order color moment of the red channel of the i-th frame of the pending video key frame; is the second-order color moment of the green channel of the i+1th frame of the undetermined video key frame; is the second-order color moment of the green channel of the i-th frame of the pending video key frame; is the second-order color moment of the blue channel of the i+1th frame of the undetermined video key frame; is the second-order color moment of the blue channel of the i-th frame of the pending video key frame.

9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the classroom experimental teaching video abstract extraction method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the classroom experimental teaching video abstract extraction method described in any one of claims 1 to 8 is implemented.