Method, apparatus, device and medium for classifying and extracting knowledge points of online teaching videos
Through the combination of video keyframe recognition model, character recognition model and knowledge point classification model, the problems of large amount of video processing and poor text recognition effect in the prior art are solved, and the fast and accurate classification and extraction of videos are achieved, and the efficiency of teaching resource utilization and user experience are improved.
Patent Information
- Application Number
- CN202411557755.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-04
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-11-04
AI Technical Summary
The prior art has high computational volume and memory requirements in video processing, resulting in limited real-time and efficientness, and poor results when processing complex backgrounds and small-sized texts, making it impossible to accurately identify detailed information.
The video keyframe recognition model, character recognition model and knowledge point classification model are used to automatically perform feature extraction and knowledge point classification. By extracting video keyframes, generating text sequences and generating knowledge point labels, efficient classification and extraction of videos are achieved.
It improves the speed and accuracy of video content labeling and improves the efficiency of teaching resources. Students and teachers can quickly find the learning content they need, improving the flexibility and user experience of online learning.
Smart Images

Figure CN119513362B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video processing, and in particular, to a method for classifying and extracting knowledge points from online teaching videos, a corresponding device, an electronic device, and a computer-readable storage medium. Background Art
[0002] In modern education, online teaching videos have become an important teaching tool. The teaching content in videos is rich and intuitive, which helps to improve learners' interest and comprehension ability. However, existing video teaching resources usually lack systematic organization and efficient classification methods, and teaching videos of different subjects, at different stages, and with different knowledge points are often scattered. Especially for students with specific learning needs, it is difficult to quickly locate the required knowledge points, which seriously affects the utilization efficiency of educational video resources.
[0003] In traditional technologies, for the classification and annotation of videos, most rely on manual operations, and their efficiency and accuracy need to be improved. With the continuous increase in video content, traditional methods are becoming increasingly difficult to meet the requirements.
[0004] When existing technologies process input videos, they perform frame-by-frame recognition on each frame of the image, resulting in a very large amount of computation. In addition, when using convolutional neural networks to extract video features, due to the large number of layers in these models and the large number of convolutional kernels in each layer, the computation amount and memory requirements are very high, thus imposing significant limitations on real-time performance and efficiency. At the same time, additional processing of educational courseware is required; in the OCR text detection process, Faster R-CNN and FCN perform poorly in dealing with complex backgrounds and small-sized texts and cannot accurately identify detailed information. TextBoxes and CTPN have insufficient support for curved texts and slow detection speeds, affecting real-time performance.
[0005] In summary, in adapting to the processing of videos in existing technologies, there are problems such as extremely high computation amount and memory requirements, thus imposing significant limitations on real-time performance and efficiency, and poor performance in dealing with complex backgrounds and small-sized texts and inability to accurately identify detailed information. In consideration of solving these problems, the applicant has made corresponding explorations. Summary of the Invention
[0006] The purpose of the present application is to solve the above problems and provide a method for classifying and extracting knowledge points from online teaching videos, a corresponding device, an electronic device, and a computer-readable storage medium.
[0007] To meet the various purposes of the present application, the following technical solutions are adopted:
[0008] A method for classifying and extracting knowledge points from online teaching videos proposed for one of the purposes of the present application includes:
[0009] Respond to the instruction for classifying and extracting knowledge points from online teaching videos, and obtain a set of online teaching videos, where the set of online teaching videos is constructed by multiple different types of online teaching videos;
[0010] Use a video key frame recognition model that has been trained to a convergent state to extract each video key frame from the video stream corresponding to the online teaching video, so as to construct a video key frame sequence;
[0011] Use a preset character recognition model to extract the text data of each video key frame in the video key frame sequence, so as to generate a text sequence corresponding to each video key frame;
[0012] Input the text sequence corresponding to the video key frame into a knowledge point classification model that has been trained to a convergent state, and generate a knowledge point label corresponding to each video key frame, where the knowledge point label represents the subject and its corresponding knowledge point category;
[0013] Based on the time stamp corresponding to the video key frame and the knowledge point label, cut the online teaching video into multiple video segments, and merge the video segments corresponding to the same knowledge point label to determine the online teaching videos corresponding to each knowledge point label, so as to complete the classification and extraction of the knowledge points of the online teaching video.
[0014] Optionally, before the step of using a video key frame recognition model that has been trained to a convergent state to extract each video key frame from the video stream corresponding to the online teaching video to construct a video key frame sequence, it includes:
[0015] Respond to the preprocessing instruction, and obtain the original frame rate of the video stream corresponding to the online teaching video, where the original frame rate represents the number of image frames per second;
[0016] Based on the original frame rate of the video stream corresponding to the online teaching video, extract the video image frames in the video stream corresponding to the online teaching video at a preset time interval as the target video image frames of the online teaching video.
[0017] Optionally, the step of using a video key frame recognition model that has been trained to a convergent state to extract each video key frame from the video stream corresponding to the online teaching video to construct a video key frame sequence includes:
[0018] Determine the target video image frames of the online teaching video, and input the target video image frames into a video key frame recognition model that has been trained to a convergent state;
[0019] Perform image block embedding on the input target video image frames, obtain multiple image block vectors, and add a classification vector at the same time to form multiple embedding vectors;
[0020] Add position encoding vectors to the multiple embedding vectors to form input vectors, where the position encoding vectors are used to maintain the spatial position information between image patches;
[0021] Stack multiple encoding modules for feature extraction on the input vectors to determine the feature vectors corresponding to each target video image frame, where the encoding modules include multiple multi-head self-attention encoding layers;
[0022] Calculate and determine the Euclidean distance between the feature vectors of adjacent video image frames in the target video image frames, detect whether the Euclidean distance exceeds a preset change rate threshold, and if it exceeds, use the target video image frames that exceed the preset change rate threshold as video key frames to construct a video key frame sequence.
[0023] Optionally, the step of extracting the text data of each video key frame in the key video image sequence by using a preset character recognition model to generate a text sequence corresponding to each video key frame includes:
[0024] Obtain a video key frame containing characters to be recognized;
[0025] Input the video key frame containing characters to be recognized into a preset character recognition model, where the character recognition model is constructed by a DB text detection algorithm module, a text direction classifier, and a CRNN text recognition algorithm module, and a text direction classifier is provided between the output of the DB text detection algorithm module and the input of the CRNN text recognition algorithm module;
[0026] Output the key frame number and text data corresponding to the video key frame to generate a text sequence corresponding to each video key frame.
[0027] Optionally, the steps of training a knowledge point classification model include:
[0028] Obtain the text sequence corresponding to any video key frame of the online teaching video on the online education platform as a training sample;
[0029] Construct encoding vectors from the training samples, call the text feature extractor in the preset knowledge point classification model, perform text feature extraction, and obtain corresponding text features;
[0030] Call a multi-layer fully connected neural network to classify each text feature to obtain a classification result, use the subject and its corresponding knowledge point category pointed to by the training sample as the supervision label of the classification result, and backpropagate to correct the weight parameters of the text feature extractor until its loss function converges to complete the training.
[0031] Optionally, the step of splitting the online teaching video into multiple video segments based on the timestamps corresponding to the video key frames and the knowledge point tags, and merging the video segments corresponding to the same knowledge point tag to determine the online teaching videos corresponding to the respective knowledge point tags includes:
[0032] Obtain the video key frames in the online teaching video and record the timestamps of each video key frame in the online teaching video, where the timestamps are used to identify the exact positions of each key frame;
[0033] Based on the timestamps corresponding to the video key frames and the knowledge point tags, read the positions of each video key frame in the online teaching video to determine the video segmentation points;
[0034] Split the online teaching video based on the video segmentation points to determine the video segments corresponding to the respective knowledge point tags, where the video segments correspond one-to-one with the knowledge point tags;
[0035] Merge the video segments corresponding to the same knowledge point tag to determine the online teaching videos corresponding to the respective knowledge point tags, so as to complete the classification and extraction of the knowledge points of the online teaching video.
[0036] Optionally, the basic network architecture of the video key frame recognition model is the ViT model; the basic network architecture of the knowledge point classification model is the BERT model.
[0037] An online teaching video knowledge point classification and extraction device provided to meet another object of the present application includes:
[0038] A teaching video acquisition module configured to respond to an instruction for classifying and extracting knowledge points of an online teaching video and acquire an online teaching video set, where the online teaching video set is constructed by a plurality of different types of online teaching videos;
[0039] A key frame sequence construction module configured to extract each video key frame from the video stream corresponding to the online teaching video by using a video key frame recognition model trained to a convergent state to construct a video key frame sequence;
[0040] A text sequence generation module configured to extract the text data of each video key frame in the video key frame sequence by using a preset character recognition model to generate a text sequence corresponding to each video key frame;
[0041] A knowledge point tag generation module configured to input the text sequence corresponding to the video key frame into a knowledge point classification model trained to a convergent state to generate a knowledge point tag corresponding to each video key frame, where the knowledge point tag represents the subject and its corresponding knowledge point category;
[0042] The classification and extraction module is configured to cut the online teaching video into multiple video segments based on the timestamps corresponding to the video key frames and knowledge point tags, and merge the video segments corresponding to the same knowledge point tag to determine the online teaching videos corresponding to each knowledge point tag, so as to complete the classification and extraction of the knowledge points of the online teaching video.
[0043] An electronic device provided to meet another object of the present application includes a central processing unit and a memory. The central processing unit is used to call and run a computer program stored in the memory to execute the steps of the method for classifying and extracting knowledge points of the online teaching video of the present application.
[0044] A computer-readable storage medium provided to meet another object of the present application stores a computer program implemented based on the method for classifying and extracting knowledge points of the online teaching video in the form of computer-readable instructions. When the computer program is called and run by a computer, it executes the steps included in the corresponding method.
[0045] Compared with the prior art, in the prior art for video processing, there are problems such as very high computational complexity and memory requirements, resulting in great limitations in real-time performance and efficiency, and poor effects in processing complex backgrounds and small-size texts, and inability to accurately identify detailed information. The present application includes but is not limited to the following beneficial effects:
[0046] First, for the method for classifying and extracting knowledge points of the online teaching video proposed in the present application, for traditional teaching videos, manual content annotation and classification are required, which is not only time-consuming and laborious but also prone to human errors. The present application automatically performs feature extraction and knowledge point classification through a video key frame recognition model, a character recognition model, and a knowledge point classification model, and can quickly and accurately annotate video content, improving the utilization efficiency of teaching resources. Students and teachers can quickly find the online teaching videos corresponding to the required learning and teaching content through the classified knowledge points, saving time, improving the efficiency of learning and teaching, effectively improving the flexibility of online learning, and enhancing user usage efficiency.
[0047] Second, the video key frame recognition model proposed in the present application is trained using a deep learning model with an attention mechanism, greatly reducing the number of videos converted into images, alleviating the computational pressure on subsequent modules, and improving the overall operation efficiency. At the same time, the generated key frame images have low similarity and strong representativeness, and can be used as an overview of the key content of the video. Through this method, ordinary computers can make full use of their computing power to run more advanced deep learning algorithms, significantly improving the performance of the video frame feature extraction algorithm.
[0048] Thirdly, the character recognition model proposed in this application is constructed by a DB text detection algorithm module, a text direction classifier, and a CRNN text recognition algorithm module. A text direction classifier is provided between the output of the DB text detection algorithm module and the input of the CRNN text recognition algorithm module to handle text recognition in different directions and improve the accuracy of text recognition. At the same time, the CRNN text recognition algorithm is suitable for processing complex texts and is applicable to complex and irregular texts, such as handwritten texts or texts with different fonts in lectures or teaching videos. It has strong robustness and is applicable to processing the video key frames in the course videos recorded during the teaching lectures, academic lectures, academic conferences, and scientific research training of this application, and the knowledge points of the specific content to which the video key frames classified and extracted according to the recognition results belong;
[0049] Fourthly, this application realizes the accurate segmentation of different knowledge points from long videos. This method ensures that the finally generated video file has good playback continuity through an optimized video merging technology, meeting the requirements of actual teaching scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The above and / or additional aspects and advantages of this application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0051] Figure 1 is a schematic flowchart of the method for classifying and extracting knowledge points from network teaching videos in an embodiment of this application;
[0052] Figure 2 is a schematic flowchart of preprocessing network teaching videos in an embodiment of this application;
[0053] Figure 3 is a schematic flowchart of constructing a video key frame sequence in an embodiment of this application;
[0054] Figure 4 is a schematic flowchart of generating a text sequence corresponding to each video key frame in an embodiment of this application;
[0055] Figure 5 is a schematic flowchart of training a knowledge point classification model in an embodiment of this application;
[0056] Figure 6 is a schematic flowchart of determining the network teaching videos corresponding to each knowledge point label in an embodiment of this application;
[0057] Figure 7 is a schematic block diagram of the device for classifying and extracting knowledge points from network teaching videos in an embodiment of this application;
[0058] Figure 8 is a schematic structural diagram of a computer device in an embodiment of this application. Detailed implementation manners
[0059] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and should not be construed as a limitation to the present application.
[0060] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0061] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.
[0062] Those skilled in the art can understand that the "client", "terminal", and "terminal device" used herein include both devices with a wireless signal receiver that only has the ability to receive and no ability to transmit, and devices with receiving and transmitting hardware that can perform two-way communication on a two-way communication link. Such devices may include: cellular or other communication devices such as personal computers and tablet computers, which have a single-line display or a multi-line display or a cellular or other communication device without a multi-line display; PCS (Personal Communications Service), which can combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant), which may include a radio frequency receiver, a pager, Internet / intranet access, a web browser, a notepad, a calendar, and / or a GPS (Global Positioning System) receiver; conventional laptop and / or palm computers or other devices, which are conventional laptop and / or palm computers or other devices with and / or including a radio frequency receiver. The "client", "terminal", and "terminal device" used herein can be portable, transportable, installed in a vehicle (air, sea, and / or land), or suitable for and / or configured to run locally, and / or run in a distributed manner at any other location on the earth and / or in space. The "client", "terminal", and "terminal device" used herein can also be a communication terminal, an Internet access terminal, a music / video playback terminal, such as a PDA, a MID (Mobile Internet Device), and / or a mobile phone with music / video playback function, or can also be a smart TV, a set-top box, and other devices.
[0063] The hardware referred to by names such as "server", "client", and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer, and is a hardware device with the necessary components disclosed by the von Neumann principle, including a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device. The computer program is stored in its memory, and the central processing unit loads the program stored in the external memory into the memory for execution, executes the instructions in the program, and interacts with the input / output devices to complete specific functions.
[0064] It should be noted that the concept of "server" in this application can similarly be extended to apply to server clusters. According to the network deployment principles understood by those skilled in the art, the various servers should be logically divided. Physically, these servers can either be independent of each other but can be invoked through interfaces, or integrated into a single physical computer or a set of computer clusters. Those skilled in the art should understand this flexibility and should not be restricted by this when implementing the network deployment method of this application.
[0065] One or several technical features of this application, unless explicitly specified, can either be deployed on the server and accessed by the client remotely invoking the online service interface provided by the server, or directly deployed and run on the client for access.
[0066] The neural network models cited or possibly cited in this application, unless explicitly specified, can either be deployed on a remote server and remotely invoked on the client, or deployed on a client capable of handling the device and directly invoked. In some embodiments, when it runs on the client, its corresponding intelligence can be obtained through transfer learning to reduce the requirements for the client's hardware operating resources and avoid over-occupying the client's hardware operating resources.
[0067] All kinds of data involved in this application, unless explicitly specified, can either be remotely stored on the server or stored on the local terminal device, as long as it is suitable for being invoked by the technical solution of this application.
[0068] Those skilled in the art should be aware that although the various methods of this application are described based on the same concept and thus show commonality with each other, unless otherwise specified, these methods can be executed independently. Similarly, for the various embodiments disclosed in this application, they are all proposed based on the same inventive concept. Therefore, for the same expressed concepts, as well as concepts that are only appropriately transformed for convenience although the concept expressions are different, they should be equivalently understood.
[0069] For the various embodiments to be disclosed in this application, unless explicitly stated to be mutually exclusive, the relevant technical features involved in each embodiment can be cross-combined to flexibly construct new embodiments, as long as such combination does not deviate from the creative spirit of this application and can meet the requirements in the prior art or solve certain deficiencies in the prior art. Those skilled in the art should be aware of this flexibility.
[0070] Please refer to Figure 1 , in one embodiment of the method for classifying and extracting knowledge points from network teaching videos of this application, it includes:
[0071] Step S10: In response to an instruction for classifying and extracting knowledge points from online teaching videos, obtain an online teaching video set, where the online teaching video set is constructed by multiple different types of online teaching videos;
[0072] The online teaching video management terminal can respond to an instruction for classifying and extracting knowledge points from online teaching videos to obtain an online teaching video set, where the online teaching video set is constructed by multiple different types of online teaching videos; specifically, the online teaching videos can be course videos recorded during teaching lectures, academic lectures, academic conferences, and scientific research training, etc. When importing a single online teaching video, multiple video formats are supported, including but not limited to online teaching videos in video formats such as MP4, AVI, and WMV. At the same time, it can perform database backend upload, support the upload of a large amount of video data, and ensure the safe and stable storage of video data on the server of the online teaching video management terminal.
[0073] In some embodiments, refer to Figure 2 , before the step of extracting each video key frame from the video stream corresponding to the online teaching video by using a video key frame recognition model that has been trained to a convergent state to construct a video key frame sequence, it includes:
[0074] Step S101: In response to a preprocessing instruction, obtain the original frame rate of the video stream corresponding to the online teaching video, where the original frame rate represents the number of image frames per second;
[0075] Step S102: Based on the original frame rate of the video stream corresponding to the online teaching video, extract video image frames from the video stream corresponding to the online teaching video at a preset time interval as the target video image frames of the online teaching video.
[0076] Specifically, before extracting each video key frame from the video stream corresponding to the online teaching video by using a video key frame recognition model that has been trained to a convergent state, the online teaching video management terminal can respond to an instruction for preprocessing the online teaching video to obtain the original frame rate (fps) of the video stream corresponding to the online teaching video, where the original frame rate (fps) represents the number of image frames per second; then, according to the original frame rate of the online teaching video, use the fixed interval method to downsample the video image frames to reduce the number of video image frames, and extract video image frames from the video stream corresponding to the online teaching video at a preset time interval as the target video image frames of the online teaching video. For example, if the target frame rate is 10 frames per second, then 10 frames are extracted per second.
[0077] The above steps help to reduce the amount of video data, improve the efficiency of subsequent video key frame processing, and save computing resources.
[0078] Step S20: Use the video key-frame recognition model that has been trained to a convergent state to extract each video key frame from the video stream corresponding to the network teaching video, so as to construct a video key-frame sequence;
[0079] After determining the target video image frame of the network teaching video, use the video key-frame recognition model that has been trained to a convergent state to extract each video key frame from the video stream corresponding to the network teaching video, so as to construct a video key-frame sequence; wherein, the basic network architecture of the video key-frame recognition model is the ViT model.
[0080] Specifically, please refer to Figure 3 , the steps of using the video key-frame recognition model that has been trained to a convergent state to extract each video key frame from the video stream corresponding to the network teaching video, so as to construct a video key-frame sequence, include:
[0081] Step S201: Determine the target video image frame of the network teaching video, and input the target video image frame into the video key-frame recognition model that has been trained to a convergent state;
[0082] After determining the target video image frame of the network teaching video, input the target video image frame into the video key-frame recognition model that has been trained to a convergent state, so as to extract the video key frame in the target video image frame of the network teaching video.
[0083] Step S202: Perform image patch embedding on the input target video image frame to obtain a plurality of image patch vectors, and at the same time add a classification vector to form a plurality of embedding vectors;
[0084] The basic network architecture of the video key-frame recognition model is the ViT model, and the ViT model is a model that is further improved based on the Transformer model applied to NLP problems and acts on the image field. It first converts each image in the image field into a word structure in natural language processing, which is called image patch embedding.
[0085] Specifically, the input target video image frame is normalized to obtain a picture of a standard size, which can be regarded as a complete sentence; then it is segmented into small blocks of a fixed size, called Patches. Flattening the pixel values of each small block can become a word in the sentence. Subsequently, each Patch is compressed into a vector of a certain dimension through a fully connected network, thereby obtaining multiple image block vectors. This process is the Patch Embedding. In the embodiments of the present application, the standard size is 224*224, the fixed size of the Patch is 16*16, and the dimension of the vector is 768. Those skilled in the art can determine the above parameters according to actual scenarios as needed, and no limitation is made here.
[0086] After obtaining multiple image block vectors, a classification vector of the same dimension is concatenated. It is not difficult to understand that this vector is used for learning class information during model training, and the classification vector is a learnable embedding vector.
[0087] Step S203: Add a position encoding vector to the multiple embedding vectors to form an input vector, where the position encoding vector is used to maintain the spatial position information between the image blocks;
[0088] Multiple embedding vectors can be obtained from the above step S202, but the embedding vectors lack position information, that is, except for the classification vector, other vectors have their corresponding position information in the picture. Therefore, in order to maintain the spatial position information between each Patch in the input image, it is necessary to add a position encoding vector to these multiple embedding vectors.
[0089] Specifically, a one-dimensional learnable position embedding vector is directly used, and the position embedding vector is directly added to the embedding vector to form an input vector. Therefore, the input target video image frame has completed the vector embedding work here and can be input into the Transformer module for training and feature learning.
[0090] Step S204: Stack multiple encoding modules for feature extraction for the input vector to determine the feature vector corresponding to each target video image frame, where the encoding module includes multiple multi-head self-attention encoding layers;
[0091] Stack multiple encoding modules for the above input vector to extract features and determine the feature vector corresponding to each target video image frame. Among them, the encoding module mainly includes multiple multi-head self-attention encoding layers and multi-layer perceptrons. Specifically, the encoding module includes two parts. The first part is Layer Norm -> Multi-Head Attention -> Dropout -> Short Cut. The second part is Layer Norm -> Multi-Layer Perception -> Dropout -> ShortCut. The multi-layer perceptron includes Linear -> GELU -> Dropout -> Linear -> Dropout. The layer normalization is to normalize the specified dimension of a single data. The feature extraction enables the classification vector to fuse the semantic features of all image patches to determine the feature vector corresponding to each target video image frame.
[0092] Step S205: Calculate and determine the Euclidean distance between the feature vectors of adjacent video image frames in the target video image frame, and detect whether the Euclidean distance exceeds a preset change rate threshold. If it exceeds, regard the target video image frame that exceeds the preset change rate threshold as a video key frame to construct a video key frame sequence.
[0093] Specifically, after determining the feature vector corresponding to each target video image frame, construct a two-dimensional array: store the feature vector extracted from each frame in the two-dimensional array features_array, where each row represents the feature vector of a frame, and the feature vector is obtained from the last hidden state of the ViT model; store the feature vectors in the two-dimensional array features_array in a list named features for subsequent processing.
[0094] Calculate the Euclidean distance between the feature vectors of adjacent video image frames in the target video image frame to measure the change rate of features. The calculation formula for the Euclidean distance (the difference between the feature vectors of adjacent video image frames) between the feature vectors of adjacent video image frames is as follows:
[0095]
[0096] Among them, a and b are two feature vectors, and d is the dimension of the feature vector.
[0097] Further, according to the Euclidean distance between the feature vectors of adjacent video image frames, that is, the difference between the feature vectors of adjacent video image frames, an array of shape (n-1, d) is obtained, where n is the number of frames and d is the dimension of the feature vector.
[0098] Further, video key frames are selected from the target video image frames according to a preset change rate threshold (threshold), and the preset change rate threshold (threshold) takes values between 0.1 and 10. In this embodiment, the preset change rate threshold (threshold) is taken as 2, and the target video image frames with a change rate r greater than the preset change rate threshold are considered video key frames.
[0099] Furthermore, the image sequence numbers frame_id of the detected video key frames are output to identify these frames as video key frames in the network teaching video.
[0100] Step S30: Extract the text data of each video key frame in the video key frame sequence by using a preset character recognition model to generate a text sequence corresponding to each video key frame;
[0101] After extracting each video key frame from the video stream corresponding to the network teaching video by using a video key frame recognition model trained to a converged state to construct a video key frame sequence, extract the text data of each video key frame in the video key frame sequence by using a preset character recognition model to generate a text sequence corresponding to each video key frame; wherein, the character recognition model is constructed by a DB text detection algorithm module, a text direction classifier, and a CRNN text recognition algorithm module, and a text direction classifier is provided between the output of the DB text detection algorithm module and the input of the CRNN text recognition algorithm module.
[0102] Further, please refer to Figure 4 , the steps of extracting the text data of each video key frame in the key video image sequence by using a preset character recognition model to generate a text sequence corresponding to each video key frame include:
[0103] Step S301: Obtain a video key frame containing characters to be recognized;
[0104] Step S302: Input the video key frame containing characters to be recognized into a preset character recognition model, where the character recognition model is constructed by a DB text detection algorithm module, a text direction classifier, and a CRNN text recognition algorithm module, and a text direction classifier is provided between the output of the DB text detection algorithm module and the input of the CRNN text recognition algorithm module;
[0105] Step S303: Output the key frame numbers and text data corresponding to the video key frames to generate a text sequence corresponding to each video key frame.
[0106] After constructing the video key frame sequence in the network teaching video, obtain the video key frames containing characters to be recognized from the video key frame sequence; input the video key frames containing characters to be recognized into a preset character recognition model, where the character recognition model is constructed by a DB text detection algorithm module, a text direction classifier, and a CRNN text recognition algorithm module, and a text direction classifier is provided between the output of the DB text detection algorithm module and the input of the CRNN text recognition algorithm module; output the key frame numbers and text data corresponding to the video key frames to generate a text sequence corresponding to each video key frame.
[0107] Specifically, the DB (Differentiable Binarization) text detection algorithm is used to detect text regions from the video key frames of this application. Its core idea is to transform the text detection problem into a binarization task and optimize it through deep learning methods. Its processing flow includes:
[0108] First, a convolutional neural network (CNN) is used to extract features from the input image. Commonly used networks include ResNet, VGG, etc., all of which are mature neural networks and can be selected for use; among them, the feature map extracted by the network contains local information and high-level feature representations of the image.
[0109] Furthermore, the feature map is input into a segmentation network, which is responsible for generating a binarized text region prediction map. In the image output by the segmentation network, the text region is marked as the foreground (usually white), and the background is marked as the background (usually black). Through a differentiable binarization loss function, the network can generate clear text region boundaries.
[0110] Through further contour detection or morphological operations, extract the connected regions in the binarized image, locate the bounding boxes of the text regions to determine the text image regions.
[0111] As can be seen from the above steps, the DB text detection algorithm transforms the text detection task into a binarization problem, uses a deep learning network for feature extraction and segmentation, effectively recognizes text image regions, and its advantage lies in being able to handle various complex scenes and text styles. At the same time, through a differentiable loss function and post-processing techniques, the detection accuracy and stability are improved.
[0112] In some embodiments, a text direction classifier is provided between the output of the DB text detection algorithm module and the input of the CRNN text recognition algorithm module. The text direction classifier is used to predict the direction of the text region, perform corresponding rotation correction, and use the text image region output by the above-mentioned DB text detection algorithm module as the input of the text direction classifier. A convolutional neural network (CNN) is used to extract features, denoted as f = CNN(I). The convolutional neural network (CNN) extracts effective features from the text image region for subsequent classification tasks;
[0113] The Softmax layer is used to classify the extracted features, and the output D = Softmax(W·f + b) is calculated, where W and b are the weights and biases obtained through training. The classification result D includes four directions: forward (0 degrees), counterclockwise rotation by 90 degrees, counterclockwise rotation by 180 degrees, and counterclockwise rotation by 270 degrees.
[0114] The image is rotationally corrected according to the classification result D to adjust the text region to be forward.
[0115] In some embodiments, the CRNN text recognition algorithm module recognizes the text region after rotation correction. The CRNN (Convolutional Recurrent Neural Network) text recognition algorithm combines the advantages of the convolutional neural network (CNN) and the recurrent neural network (RNN) and is used to handle text recognition tasks in sequential data. Its workflow includes:
[0116] In the convolutional layer, a convolutional neural network is used to extract features from the image. The convolutional layer is used to extract local features and spatial information from the input video key frames;
[0117] In the pooling layer, the dimension of the feature map is reduced to reduce the computational amount while retaining important feature information. In the recurrent layer, the features extracted by the convolutional layer are input into the recurrent neural network (such as LSTM or GRU). The recurrent layer is used to handle the temporal dependence of sequential data and is used to understand the context relationship between characters. The connectionist temporal classification (CTC) loss function is used for training, allowing the model to learn without explicit alignment in the input sequence, thereby performing character prediction and sequence decoding.
[0118] As can be seen from the above steps, the CRNN (Convolutional Recurrent Neural Network) text recognition algorithm extracts features through the convolutional layer, then processes the sequence characteristics through the recurrent layer, and finally performs text recognition through CTC decoding. CRNN combines the advantages of the convolutional neural network (CNN) and the recurrent neural network (RNN). The convolutional neural network (CNN) is good at extracting local features from images, while the recurrent neural network (RNN) is good at processing sequence data. The CRNN text recognition algorithm combines the two and can effectively process and recognize character sequences. Different from traditional OCR methods, the CRNN text recognition algorithm can process variable-length text sequences without the need to pre-fix the length of the character sequence. The CRNN text recognition algorithm can directly perform end-to-end training from images to text, simplifying the model training process.
[0119] Furthermore, the CRNN text recognition algorithm is suitable for processing complex texts, applicable to complex and irregular texts, such as handwritten texts or texts with different fonts in lectures or teaching videos, and has strong robustness, suitable for processing video key frames in the course videos recorded during the teaching lectures, academic lectures, academic conferences, and scientific research training of this application.
[0120] Step S40: Input the text sequence corresponding to the video key frame into the knowledge point classification model that has been trained to a convergent state to generate the corresponding knowledge point label for each video key frame, where the knowledge point label represents the subject and its corresponding knowledge point category.
[0121] After extracting the text data of each video key frame in the video key frame sequence by using a preset character recognition model to generate the text sequence corresponding to each video key frame, input the text sequence corresponding to the video key frame into the knowledge point classification model that has been trained to a convergent state to generate the corresponding knowledge point label for each video key frame, where the knowledge point label represents the subject and its corresponding knowledge point category. Specifically, the basic network architecture of the knowledge point classification model is the BERT model. The subjects include Chinese, mathematics, physics, biology, art, and music subjects, etc. The knowledge point categories corresponding to the mathematics subject include algebra, geometry, calculus, and statistics, etc. The knowledge point categories corresponding to physics include mechanics, thermodynamics, electromagnetism, and optics, etc. Among them, one subject corresponds to multiple knowledge point categories, and the knowledge point categories between different subjects are different.
[0122] In a specific embodiment, please refer to Figure 5 , the steps of training the knowledge point classification model include:
[0123] Step S401: Obtain the text sequence corresponding to any video key frame of the online teaching video in the online education platform as a training sample;
[0124] As mentioned above, the basic network architecture of the knowledge point classification model of the present application is the BERT model. The BERT model can be used to implement the text feature extractor in this embodiment. For the BERT model, it needs to be pre-trained until it converges to efficiently serve the technical solutions of various embodiments of the present application.
[0125] Preferably, the text sequence corresponding to any video key frame of the online teaching video in the online education platform is regarded as a training sample. Therefore, in this embodiment, the text sequence corresponding to any of the video key frames of the online teaching video in the online education platform can be obtained for training. Among them, the text sequences corresponding to the handwritten text or the text with different fonts in the lecture or teaching video are all used as training samples, and the subject and its corresponding knowledge point category pointed to by the training sample can be used as the supervision label after the BERT model is connected to the classifier.
[0126] Step S402: Construct an encoded vector from the training sample, and call the text feature extractor in the preset knowledge point classification model to perform text feature extraction to obtain corresponding text features;
[0127] Obtain the text sequence corresponding to any video key frame of the online teaching video in the online education platform, and construct the text sequence corresponding to the any video key frame into an encoded vector to meet the input requirements of the BERT model. The BERT model converts each word in the text sequence corresponding to any video key frame into a one-dimensional vector by querying the word vector table as the model input. In addition, the model input includes two other parts: the text vector and the position vector. The value of the text vector is automatically learned during the model training process and is used to depict the global semantic information of the text and fuse it with the semantic information of single words / terms. The position vector: Since the semantic information carried by words / terms at different positions in the text is different. For example, "I miss her" and "She misses me". Therefore, the BERT model attaches a different vector to words / terms at different positions for distinction. The BERT model takes the sum of the word vector, the text vector, and the position vector as the encoded vector required for model input. The training sample can be pre-constructed into the corresponding encoded vector to provide input for the BERT model. The BERT model performs text feature extraction on the input encoded vector according to its own implementation logic and finally obtains the corresponding text features.
[0128] Step S403: Call a multi-layer fully connected neural network to classify each text feature to obtain a classification result. Use the subject and its corresponding knowledge point category pointed to by the training sample as the supervision label of the classification result, and backpropagate to correct the weight parameters of the text feature extractor until its loss function converges to complete the training.
[0129] To implement the task training of the BERT model, the text features output by it are input into a multi-layer fully connected neural network for classification to obtain the scores corresponding to the classifications of each subject and its corresponding knowledge point category; since the subject and its corresponding knowledge point category are used as supervision labels in the multi-layer fully connected neural network, therefore, the multi-layer fully connected neural network will backpropagate to the BERT model according to the supervision of the supervision label to correct its weight parameters until the entire loss function converges, and finally complete the training. It can be understood that the loss function can be implemented based on the cross-entropy loss function.
[0130] After the training convergence state of the knowledge point classification model, input the text sequence corresponding to the video key frame into the knowledge point classification model that has been trained to the convergence state, and then the knowledge point label corresponding to each video key frame can be generated, and the subject and its corresponding knowledge point category corresponding to each video key frame can be determined.
[0131] This embodiment realizes the training process of the knowledge point classification model required by the technical solution of the present application. By training a knowledge point classification model to generate the knowledge point label corresponding to each video key frame and determine the subject and its corresponding knowledge point category corresponding to each video key frame, it can be seen that since the text sequence corresponding to any video key frame of the online teaching video in the online education platform is used for training during the training process, the BERT model is easier to converge. Thus, the obtained text feature extractor has stronger representation learning ability and can more accurately match the subject and its corresponding knowledge point category suitable for the text sequence corresponding to the video key frame, which helps to improve the content matching degree of the online teaching video and enhance the user interaction experience.
[0132] Step S50: Based on the time stamp and knowledge point label corresponding to the video key frame, cut the online teaching video into multiple video segments, and merge the video segments corresponding to the same knowledge point label to determine the online teaching videos corresponding to each knowledge point label, so as to complete the classification extraction of the knowledge points of the online teaching video.
[0133] After inputting the text sequence corresponding to the video key frames into the knowledge point classification model that has been trained to a convergent state, and generating the corresponding knowledge point labels for each video key frame, based on the time stamps and knowledge point labels corresponding to the video key frames, the online teaching video is segmented into multiple video segments, and the video segments corresponding to the same knowledge point label are merged to determine the online teaching videos corresponding to each knowledge point label, so as to complete the classification and extraction of the knowledge points of the online teaching video.
[0134] In a specific embodiment, please refer to Figure 6 , based on the time stamps and knowledge point labels corresponding to the video key frames, the steps of segmenting the online teaching video into multiple video segments and merging the video segments corresponding to the same knowledge point label to determine the online teaching videos corresponding to each knowledge point label include:
[0135] Step S501: Obtain the video key frames in the online teaching video, and record the time stamps of each video key frame in the online teaching video, where the time stamps are used to identify the exact positions of each key frame;
[0136] Step S502: Based on the time stamps and knowledge point labels corresponding to the video key frames, read the positions of each video key frame in the online teaching video to determine the video segmentation points;
[0137] Step S503: Segment the online teaching video based on the video segmentation points to determine the video segments corresponding to each knowledge point label, where the video segments correspond one-to-one with the knowledge point labels;
[0138] Step S504: Merge the video segments corresponding to the same knowledge point label to determine the online teaching videos corresponding to each knowledge point label, so as to complete the classification and extraction of the knowledge points of the online teaching video.
[0139] Specifically, the online teaching video management terminal can obtain the video key frames in the online teaching video and record the time stamps of each key frame in the video for subsequent processing and segmentation. The time stamps are used to identify the exact positions of each key frame for use in the video segmentation process.
[0140] Further, for the timestamp information recorded in the above steps, read the positions of each key frame in the online teaching video to determine the segmentation points; a video editing tool (such as FFmpeg) can be used to segment the original video into multiple small video segments according to the timestamp information. Each small video segment corresponds to a specific knowledge point. For example, from time point A to time point B is a knowledge point segment, and from time point B to time point C is another knowledge point segment; name and store the segmented small video segments. Each video segment is named according to its corresponding knowledge point and stored in a preset folder, such as "Subject 1_Knowledge Point 1_Part 1.mp4", "Subject 1_Knowledge Point 1_Part 2.mp4", etc.
[0141] Sort all relevant video segments according to knowledge points, call the preset video editing tool, and merge all video segments belonging to the same knowledge point into a complete video file; store the merged complete video file in a specified directory for subsequent release and use. For example, merge "Subject 1_Knowledge Point 1_Part 1.mp4" and "Subject 1_Knowledge Point 1_Part 2.mp4" into "Subject 1_Knowledge Point 1_Complete Video.mp4" and store it in the "Subject 1_Knowledge Point 1" folder.
[0142] As can be seen from the above steps, by using the knowledge point classification model to determine the subject and the corresponding knowledge point category of each video key frame in the online teaching video, merge the video segments corresponding to the same subject and the same knowledge point category to determine the online teaching video corresponding to each knowledge point label. The educational video is efficiently processed and sorted, and the complete video files of each knowledge point corresponding to each subject are output, so as to facilitate students' targeted learning and improve the utilization rate of educational resources.
[0143] As can be seen from the above embodiments, compared with the prior art, for the processing of videos in the prior art, there are problems such as very high computational complexity and memory requirements, resulting in great limitations in real-time performance and efficiency, and poor effects in processing complex backgrounds and small-size texts, and unable to accurately identify detailed information. The present application includes, but is not limited to, the following beneficial effects:
[0144] First, the method for classifying and extracting knowledge points from online teaching videos proposed in this application addresses the problems of traditional teaching videos that require manual content annotation and classification, which is not only time-consuming and laborious but also prone to human errors. This application automatically extracts features and classifies knowledge points through a video key-frame recognition model, a character recognition model, and a knowledge point classification model, enabling fast and accurate annotation of video content and improving the utilization efficiency of teaching resources. Students and teachers can quickly find the online teaching videos corresponding to the learning and teaching content they need through the classified knowledge points, saving time, improving the efficiency of learning and teaching, effectively enhancing the flexibility of online learning, and improving user efficiency.
[0145] Second, the video key-frame recognition model proposed in this application is trained using a deep learning model with an attention mechanism, greatly reducing the number of videos converted into images, alleviating the computational pressure on subsequent modules, and improving the overall operation efficiency. At the same time, the generated key-frame images have low similarity and strong representativeness, which can be used as an overview of the key content of the video. Through this method, ordinary computers can make full use of their computing power to run more advanced deep learning algorithms, significantly improving the performance of the video frame feature extraction algorithm.
[0146] Third, the character recognition model proposed in this application is constructed by a DB text detection algorithm module, a text direction classifier, and a CRNN text recognition algorithm module. A text direction classifier is provided between the output of the DB text detection algorithm module and the input of the CRNN text recognition algorithm module to handle text recognition in different directions and improve the accuracy of text recognition. At the same time, the CRNN text recognition algorithm is suitable for processing complex texts, such as handwritten texts or texts with different fonts in lectures or teaching videos, and has strong robustness. It is applicable to processing the video key frames in the course videos recorded during teaching lectures, academic lectures, academic conferences, and scientific research training in this application, and classifying and extracting the knowledge points of the specific content to which the video key frames belong according to the recognition results.
[0147] Fourth, this application realizes accurate segmentation of different knowledge points from long videos. Through an optimized video merging technology, it ensures that the finally generated video file has good playback continuity and meets the requirements of actual teaching scenarios.
[0148] Please refer to Figure 7, A device for classifying and extracting knowledge points from online teaching videos provided to meet one of the purposes of this application, includes a teaching video acquisition module 1100, a key frame sequence construction module 1200, a text sequence generation module 1300, a knowledge point label generation module 1400, and a classification and extraction module 1500. Among them, the teaching video acquisition module 1100 is configured to respond to an instruction for classifying and extracting knowledge points from online teaching videos, and acquire an online teaching video set, where the online teaching video set is constructed by multiple different types of online teaching videos; the key frame sequence construction module 1200 is configured to extract each video key frame from the video stream corresponding to the online teaching video by using a video key frame recognition model that has been trained to a convergent state, so as to construct a video key frame sequence; the text sequence generation module 1300 is configured to extract the text data of each video key frame in the video key frame sequence by using a preset character recognition model, so as to generate a text sequence corresponding to each video key frame; the knowledge point label generation module 1400 is configured to input the text sequence corresponding to the video key frame into a knowledge point classification model that has been trained to a convergent state, and generate a knowledge point label corresponding to each video key frame, where the knowledge point label represents the subject and its corresponding knowledge point category; the classification and extraction module 1500 is configured to cut the online teaching video into multiple video segments based on the time stamp and knowledge point label corresponding to the video key frame, and merge the video segments corresponding to the same knowledge point label to determine the online teaching videos corresponding to each knowledge point label, so as to complete the classification and extraction of knowledge points from online teaching videos.
[0149] Based on any embodiment of this application, please refer to Figure 8 , Another embodiment of this application further provides an electronic device, which can be implemented by a computer device, as shown in Figure 8 . The internal structure schematic diagram of the computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected through a system bus. Among them, the computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The control information sequence can be stored in the database. When the computer-readable instructions are executed by the processor, the processor can implement a method for classifying and extracting knowledge points from online teaching videos. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device can store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the method for classifying and extracting knowledge points from online teaching videos of this application. The network interface of the computer device is used to connect and communicate with the terminal. Those skilled in the art can understand, Figure 8The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0150] In this embodiment, the processor is used to execute Figure 7 the specific functions of each module and its sub-modules in. The memory stores the program codes and various types of data required to execute the above-mentioned modules or sub-modules. The network interface is used for data transmission between the user terminal and the server. The memory in this embodiment stores the program codes and data required to execute all modules / sub-modules in the network teaching video knowledge point classification extraction device of this application, and the server can call the program codes and data of the server to execute the functions of all sub-modules.
[0151] This application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the network teaching video knowledge point classification extraction method described in any embodiment of this application.
[0152] This application also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by one or more processors, the steps of the network teaching video knowledge point classification extraction method described in any embodiment of this application are implemented.
[0153] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments of this application can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0154] The above are only some embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.
[0155] In summary, the present application automatically performs feature extraction and knowledge point classification through a video key frame recognition model, a character recognition model, and a knowledge point classification model, and can quickly and accurately annotate video content, improving the utilization efficiency of teaching resources; students and teachers can quickly find the online teaching videos corresponding to the required learning and teaching content through the classified knowledge points, saving time, improving the efficiency of learning and teaching, effectively improving the flexibility of online learning, and enhancing the user usage efficiency.
Claims
1. A method for classifying and extracting knowledge points from online teaching videos, characterized in that: include: In response to an instruction to classify and extract knowledge points from online teaching videos, a collection of online teaching videos is obtained, wherein the collection of online teaching videos is constructed from a plurality of online teaching videos of different types; The video key frame recognition model that has been trained to a convergence state is used to extract each video key frame from the video stream corresponding to the online teaching video to construct a video key frame sequence, which includes: Determine a target video image frame of the online teaching video, and input the target video image frame into a video key frame recognition model that has been trained to a convergence state; Perform image block embedding on the input target video image frame to obtain multiple image block vectors, and add a classification vector to form multiple embedding vectors; Adding a position encoding vector to the multiple embedding vectors to form an input vector, wherein the position encoding vector is used to maintain spatial position information between image blocks; Stacking a plurality of encoding modules for the input vector to perform feature extraction to determine a feature vector corresponding to each target video image frame, wherein the encoding module includes a plurality of multi-head self-attention encoding layers; Calculate and determine the Euclidean distance between feature vectors of adjacent video image frames in the target video image frame, detect whether the Euclidean distance exceeds a preset change rate threshold, and if so, use the target video image frame that exceeds the preset change rate threshold as a video key frame to construct a video key frame sequence; Extracting text data of each video key frame in the video key frame sequence using a preset character recognition model to generate a text sequence corresponding to each video key frame; Inputting the text sequence corresponding to the video key frame into the knowledge point classification model that has been trained to a convergence state, generating a corresponding knowledge point label for each video key frame, wherein the knowledge point label represents the subject and its corresponding knowledge point category; Based on the timestamps and knowledge point labels corresponding to the video key frames, the online teaching video is cut into multiple video clips, and the video clips corresponding to the same knowledge point label are merged to determine the online teaching video corresponding to each knowledge point label, so as to complete the classification extraction of knowledge points in the online teaching video.
2. The method for classifying and extracting knowledge points from online teaching videos according to claim 1, characterized in that: Before the step of extracting each video key frame from the video stream corresponding to the online teaching video using the video key frame recognition model that has been trained to a convergence state to construct a video key frame sequence, the method includes: In response to the preprocessing instruction, an original frame rate of a video stream corresponding to the online teaching video is obtained, wherein the original frame rate represents the number of image frames per second; Based on the original frame rate of the video stream corresponding to the online teaching video, video image frames in the video stream corresponding to the online teaching video are extracted at preset time intervals as target video image frames of the online teaching video.
3. The method for classifying and extracting knowledge points from online teaching videos according to claim 1, characterized in that: The step of extracting text data of each video key frame in the video key frame sequence by using a preset character recognition model to generate a text sequence corresponding to each video key frame comprises: Get the video key frame containing the characters to be recognized; Input the video key frame containing the characters to be recognized into a preset character recognition model, wherein the character recognition model is constructed by a DB text detection algorithm module, a text direction classifier and a CRNN text recognition algorithm module, and a text direction classifier is provided between the output of the DB text detection algorithm module and the input of the CRNN text recognition algorithm module; The key frame sequence number and text data corresponding to the video key frame are output to generate a text sequence corresponding to each video key frame.
4. The method for classifying and extracting knowledge points from online teaching videos according to claim 1, characterized in that: The steps for training the knowledge point classification model include: Obtaining a text sequence corresponding to any video key frame of an online teaching video in an online education platform as a training sample; Constructing a coding vector from the training sample, calling a text feature extractor in a preset knowledge point classification model to extract text features and obtain corresponding text features; A multi-layer fully connected neural network is called to classify each text feature to obtain the classification result. The subject pointed to by the training sample and its corresponding knowledge point category are used as the supervision label of the classification result. The weight parameters of the text feature extractor are corrected by back propagation until its loss function converges to complete the training.
5. The method for classifying and extracting knowledge points from online teaching videos according to claim 1, characterized in that: The steps of cutting the online teaching video into multiple video clips based on the timestamps and knowledge point labels corresponding to the video key frames, and merging the video clips corresponding to the same knowledge point label to determine the online teaching videos corresponding to the respective knowledge point labels include: Obtaining video key frames in the online teaching video, and recording the timestamp of each video key frame in the online teaching video, wherein the timestamp is used to identify the precise position of each key frame; Based on the timestamps and knowledge point labels corresponding to the video key frames, the positions of the video key frames in the online teaching video are read to determine the video segmentation points; Segmenting the online teaching video based on the video segmentation points to determine video segments corresponding to each knowledge point label, wherein the video segments correspond to the knowledge point labels one by one; The video clips corresponding to the same knowledge point label are merged to determine the online teaching videos corresponding to each knowledge point label, so as to complete the classification extraction of knowledge points in the online teaching videos.
6. The method for classifying and extracting knowledge points from online teaching videos according to any one of claims 1 to 5, characterized in that: The basic network architecture of the video key frame recognition model is the ViT model; the basic network architecture of the knowledge point classification model is the BERT model.
7. A network teaching video knowledge point classification and extraction device, characterized in that: include: A teaching video acquisition module, configured to respond to an instruction to classify and extract knowledge points from online teaching videos, and acquire an online teaching video set, wherein the online teaching video set is constructed from a plurality of online teaching videos of different types; The key frame sequence construction module is configured to extract each video key frame from the video stream corresponding to the online teaching video using a video key frame recognition model that has been trained to a convergence state to construct a video key frame sequence, which includes: Determine a target video image frame of the online teaching video, and input the target video image frame into a video key frame recognition model that has been trained to a convergence state; Perform image block embedding on the input target video image frame to obtain multiple image block vectors, and add a classification vector to form multiple embedding vectors; Adding a position encoding vector to the multiple embedding vectors to form an input vector, wherein the position encoding vector is used to maintain spatial position information between image blocks; Stacking a plurality of encoding modules for the input vector to perform feature extraction to determine a feature vector corresponding to each target video image frame, wherein the encoding module includes a plurality of multi-head self-attention encoding layers; Calculating and determining the Euclidean distance between feature vectors of adjacent video image frames in the target video image frame, detecting whether the Euclidean distance exceeds a preset change rate threshold, and if so, taking the target video image frame exceeding the preset change rate threshold as a video key frame to construct a video key frame sequence; A text sequence generation module is configured to extract text data of each video key frame in the video key frame sequence using a preset character recognition model to generate a text sequence corresponding to each video key frame; A knowledge point label generation module is configured to input the text sequence corresponding to the video key frame into the knowledge point classification model that has been trained to a convergence state, and generate a corresponding knowledge point label for each video key frame, wherein the knowledge point label represents a subject and its corresponding knowledge point category; The classification extraction module is configured to cut the online teaching video into multiple video clips based on the timestamps and knowledge point labels corresponding to the video key frames, merge the video clips corresponding to the same knowledge point labels to determine the online teaching videos corresponding to the various knowledge point labels, so as to complete the classification extraction of knowledge points in the online teaching video.
8. An electronic device, comprising a central processing unit and a memory, characterized in that: The central processing unit is used to call and run the computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: It stores a computer program implemented according to the method described in any one of claims 1 to 6 in the form of computer-readable instructions, and when the computer program is called and executed by a computer, the steps included in the corresponding method are executed.
Citation Information
Patent Citations
Model training method, video category detection method and device, electronic device and computer readable medium
CN110119757A
Intelligent sports video classification method and system based on multi-attribute learning
CN117271831A