Video marking method, device and equipment based on courseware page classification and medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU SHIYUAN ELECTRONICS CO LTD
- Filing Date
- 2023-10-10
- Publication Date
- 2026-08-07
AI Technical Summary
[0005]本发明提供了一种基于课件页面分类的视频标记方法、装置、设备和介质,以解决现有的自动标识视频大纲的方案,不能对视频大纲进行精准提取和标识的技术问题
[0005] This invention provides a video tagging method, apparatus, device, and medium based on courseware page classification to solve the technical problem that existing automatic video outline tagging schemes cannot accurately extract and tag video outlines.
Smart Images

Figure CN119815154B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of educational multimedia technology, and in particular to video tagging methods, apparatus, devices, and media based on courseware page classification. Background Technology
[0002] With the continuous development and popularization of information technology, educational forms relying on information technology, such as computer-assisted instruction, distance learning, self-directed online learning, and electronic classrooms, have made significant progress in the field of education. Classroom video recording has become commonplace in modern teaching. Recorded classroom videos can be used for students' remote learning, after-class study, and for teachers' post-class review and reflection, playing a significant role in achieving educational equity and improving teaching quality.
[0003] Classroom videos are typically around forty minutes long. Students who didn't participate in the recording usually need to watch the entire video to understand the content. Students who did participate also need to repeatedly navigate back and forth using the progress bar to find specific points they didn't understand or weren't interested in after class. Teachers who participated in the recording also need to repeatedly jump around the video's progress bar to locate specific teaching segments when they want to reflect on their teaching. To achieve quick retrieval, classroom videos are usually outlined manually or by using algorithms to identify the video content. Manual addition is inefficient; therefore, algorithms are often used to identify the video content and add an outline for efficient and accurate classroom video outlining.
[0004] When the inventors were researching existing solutions for automatically segmenting video segments and automatically identifying the corresponding video outlines in non-educational scenarios, they found that existing automatic video outline identification solutions usually require a certain amount of video annotation data to train the model. The model then identifies the video subtitle content and screen to complete the segmentation. Since the boundaries of video segments are usually quite blurry, existing automatic video outline identification solutions cannot accurately extract and identify the video outlines. Summary of the Invention
[0005] This invention provides a video tagging method, apparatus, device, and medium based on courseware page classification to solve the technical problem that existing automatic video outline tagging schemes cannot accurately extract and tag video outlines.
[0006] In a first aspect, embodiments of this application provide a video tagging method based on courseware page classification, the video tagging method based on courseware page classification includes:
[0007] Get the teaching materials played during the classroom video recording process and the page switching events of the teaching materials;
[0008] Based on the pre-trained teaching segment recognition model, the teaching segment corresponding to each page of the teaching courseware is identified, and the teaching segment information corresponding to each page of the courseware is obtained.
[0009] Based on page switching events, identify the target courseware pages that have been displayed and the corresponding display time periods for each target courseware page;
[0010] Based on the teaching segment information corresponding to the target courseware page, mark the video outline information for the corresponding display time period in the classroom video.
[0011] Secondly, embodiments of this application also provide a video tagging device based on courseware page classification, the video tagging device based on courseware page classification includes:
[0012] The classroom data acquisition unit is used to acquire teaching materials played during classroom video recording and page switching events of the teaching materials;
[0013] The courseware page classification unit is used to identify the teaching segment corresponding to each courseware page based on the pre-trained teaching segment recognition model, and obtain the teaching segment information corresponding to each courseware page.
[0014] The target page confirmation unit is used to confirm the displayed target courseware pages and the corresponding display time periods for each target courseware page based on page switching events.
[0015] The outline information marking unit is used to mark the video outline information of the corresponding display time period in the classroom video according to the teaching segment information corresponding to the target courseware page.
[0016] Thirdly, embodiments of this application also provide an electronic device, which includes:
[0017] One or more processors;
[0018] Memory, used to store one or more computer programs;
[0019] When one or more computer programs are executed by one or more processors, electronic devices enable video tagging methods based on courseware page classification, as described in the first aspect.
[0020] Fourthly, embodiments of this application also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the video tagging method based on courseware page classification as described in the first aspect.
[0021] The aforementioned video tagging method, apparatus, device, and medium based on courseware page classification includes the following steps: The method acquires the teaching courseware played during classroom video recording and page switching events of the teaching courseware; based on a pre-trained teaching segment recognition model, it identifies the teaching segment corresponding to each courseware page to obtain teaching segment information for each courseware page; it confirms the displayed target courseware pages and the corresponding display time periods based on page switching events; and it tags the video outline information of the corresponding display time period in the classroom video based on the teaching segment information corresponding to the target courseware pages. Based on common teaching segments in teaching activities, each page of the teaching materials corresponding to a teaching activity is mapped to a specific teaching segment. This mapping is achieved by predicting the pages in the teaching materials using a pre-trained teaching segment recognition model. Furthermore, data sources and events causing changes in the main content displayed in the classroom video are recorded, confirming the timing of these changes and the subsequent teaching segment of the displayed teaching material page. These teaching segments serve as the outline information for the display time of the corresponding teaching material page. Without requiring the identification of the classroom video content itself, this method achieves precise structural division and outline information identification of classroom videos based on the recognition of routine segments in teaching activities. Attached Figure Description
[0022] Figure 1 A flowchart illustrating a video tagging method based on courseware page classification, provided as an embodiment of this application.
[0023] Figure 2 This is a schematic diagram of the screen structure in a classroom video.
[0024] Figure 3 This is a schematic diagram illustrating the generation process of the data processing-related model in the video tagging method based on courseware page classification provided in this application embodiment.
[0025] Figure 4 This is a schematic diagram illustrating the relationship between page switching events and the timeline in the video tagging method based on courseware page classification provided in this application embodiment.
[0026] Figure 5 This is a schematic diagram showing the video outline generated by the video tagging method based on courseware page classification provided in this application embodiment.
[0027] Figure 6 This is a schematic diagram of the structure of a video tagging device based on courseware page classification provided in an embodiment of this application.
[0028] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0029] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and not for limiting the invention. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention and not all of the structures.
[0030] It should be noted that, due to space limitations, this application specification does not exhaustively list all possible implementation methods. Those skilled in the art should be able to conceive after reading this application specification that, as long as the technical features do not contradict each other, any combination of technical features can constitute an optional implementation method.
[0031] The embodiments of the present invention will be described in detail below.
[0032] With the development of information technology, the forms of computer-aided instruction are constantly being enriched, and the effectiveness is continuously improving. For example, interactive whiteboards can be configured in teaching scenarios to enable the electronic display and recording of teaching information. The interactive whiteboard mentioned in this solution can be an integrated device that uses touch technology to control the content displayed on the display board and realize human-computer interaction. It integrates one or more functions such as a projector, electronic whiteboard, screen, audio system, television, and video conferencing terminal.
[0033] Generally, an interactive flat panel includes at least one screen. For example, an interactive flat panel is equipped with a touch-enabled screen, which can be a capacitive, resistive, or electromagnetic screen. In this embodiment, the user can further perform touch operations by touching the screen with a finger or stylus. Accordingly, the interactive flat panel detects the touch location and responds accordingly to achieve the touch function. Typically, the interactive flat panel is equipped with at least one operating system, which includes, but is not limited to, Android, Linux, and Windows systems.
[0034] In teaching scenarios, interactive whiteboards serve as interactive devices primarily used to support the electronic display and recording of teaching information. The software within these whiteboards mainly fulfills the display and recording functionalities required in teaching settings. For example, an interactive whiteboard can install at least one application with document presentation capabilities. This application can be a pre-installed application on the operating system; alternatively, it can also install applications downloaded from third-party devices or servers. Optionally, in addition to displaying pre-edited content, the document presentation program can offer other editing functions during the presentation process, such as inserting tables, images, and graphs, and drawing tables and graphs. This allows for real-time display of user information and input, thereby improving the information display and interactive effects during teaching-related document presentations. For instance, an interactive whiteboard can install at least one application with whiteboard functionality, enabling teachers to write, record, modify, and save information in real-time during teaching scenarios.
[0035] When used alone, interactive whiteboards are mainly suitable for teachers to electronically display teaching information, such as teaching courseware, simulation animations, and simulation software, to students in the same teaching scenario during offline teaching; as well as to electronically record the blackboard writing, notes, annotations, and other content generated by teachers during teaching activities.
[0036] In addition to the existing interactive whiteboard in the teaching setting, a camera can be set up to record classroom videos during the teaching process. The camera lens is pointed towards the side where the interactive whiteboard is located. During specific teaching activities, the interactive whiteboard displays information based on the teacher's actions. In terms of the relative positions of relevant elements during the teaching process, the camera's main shooting targets are the teacher and the interactive whiteboard, with the wall where the interactive whiteboard is located serving as the teacher's lecture background. When using the classroom videos recorded by the camera later, because the videos are only filmed from one shooting area, they lack structured processing corresponding to the actual teaching activities during recording. Directly using the classroom videos requires adjusting the playback progress to locate the possible target teaching moments. For example, students who did not participate in the classroom recording usually need to watch the entire video to roughly understand the content of the lesson; students who participated in the classroom recording also need to repeatedly jump back and forth using the progress bar to locate certain knowledge points that they did not understand or were interested in after class; and teachers who participated in the classroom recording need to repeatedly jump back and forth using the video progress bar to find specific teaching segments when they want to reflect on their teaching after class. To improve the efficiency of subsequent browsing of classroom videos, we can consider using the segmented content structure of the actual content recorded in the classroom video as the outline information of each segment, thus obtaining the video outline. Based on the video outline, we can quickly identify the general content of each part of the classroom video and quickly jump to view a specific part. The video outline generated by the automatic content recognition scheme in general video processing scenarios usually cannot accurately identify the segmented content structure of the classroom video with reference to the teaching content.
[0037] To address the above technical problems, this application proposes a video tagging method based on courseware page classification. According to common specific teaching segments in teaching activities, each courseware page corresponding to a teaching activity is mapped to a specific teaching segment. This mapping is obtained by predicting courseware pages in the courseware using a pre-trained teaching segment recognition model. Furthermore, data sources and change events that cause changes in the main content displayed in the classroom video are recorded, confirming the time of occurrence of the change event and the teaching segment of the courseware page displayed afterward. The teaching segment is used as the outline information for the display period of the corresponding courseware page. Without requiring identification of the classroom video content, this method achieves accurate structural division and outline information identification of classroom videos based on the identification of common segments in teaching activities. The various embodiments of this invention are described in detail below.
[0038] Figure 1This application provides a flowchart of a video tagging method based on courseware page classification. This method is applied to electronic devices that process relevant images from a teaching environment, and is implemented by electronic devices such as recording hosts or cloud servers that store and manage teaching data. Figure 1 As shown, the video tagging method based on courseware page classification includes steps S110-S140:
[0039] Step S110: Obtain the teaching materials played during the classroom video recording process and the page switching events of the teaching materials.
[0040] Taking a recording host as an example, in a teaching scenario where an interactive whiteboard already exists, a camera and a recording host can be additionally set up, both connected to the recording host. The camera lens faces the side where the interactive whiteboard is located. Considering the relative positions of relevant elements during the teaching process, the camera's shooting targets during classroom video recording in one embodiment include the teacher and the interactive whiteboard. The interactive whiteboard is used to display teaching materials, i.e., the teaching materials are displayed within the recording area of the classroom video. The images captured by the camera are transmitted to the recording host in real time via the corresponding connection. While the interactive whiteboard displays information on the screen according to the teacher's actions, it also transmits page switching events in real time to the recording host via the corresponding connection. Specifically, page switching events refer to events such as entering slideshow mode using the teaching material editing software, switching from one slide to another in slideshow mode, or ending slideshow mode. In slideshow mode, the teaching material pages are displayed in full screen.
[0041] During teacher-organized teaching activities, various operations may occur on the interactive whiteboard, such as annotation and blackboard writing. The teaching materials are typically structured in pages, with one page displayed at a time. Content switching can be automatic or triggered by the teacher using a laser pointer, touchscreen, mouse, or keyboard shortcuts. In this embodiment, all events causing page switching in the teaching materials are defined as page switching events. Based on the above image transmission and page switching events, the recording host receives the original data (or at very small intervals, equivalent to synchronization) from the camera that causes changes in the interactive whiteboard display. This can be considered as achieving simultaneous recording of page switching events and classroom video on the same timeline.
[0042] like Figure 2As shown, when the teaching scenario is classroom 10, the classroom video is captured by a camera pointing towards the teacher's lecturing background. Structurally, the classroom video includes a foreground primarily composed of the teacher 30 and the interactive whiteboard 20, and a background primarily composed of the wall. The content of the classroom video is the video frame. Although changes in the content displayed on the interactive whiteboard 20 and the movement of the teacher 30 may cause variations in the details of the video frame, the camera's state and the layout of classroom 10 remain largely unchanged, resulting in a relatively constant structure of the video frame. If the classroom video is acquired by a recording host, the recording host can transmit data to the data processing hardware used for presenting the teaching materials via various existing connection methods (wired or wireless). For example, if the hardware device for presenting the teaching materials is an interactive whiteboard, the recording host acquires the teaching materials from the interactive whiteboard; if the hardware device is a laptop connected to the interactive whiteboard, the recording host acquires the teaching materials from the laptop. Alternatively, the hardware device for presenting the teaching materials may also be a projector, in which case the recording host acquires the teaching materials from the computer currently connected to the projector.
[0043] It should be understood that using a recording host is only one possible implementation method. Alternatively, a cloud server can be used. If the classroom video is processed in real-time, page switching events are also generated and received in real-time, effectively recording page switching events and the classroom video on the same timeline. If the video outline is generated after the classroom video recording is complete, considering that the start time of the classroom video recording and the start time of the teacher's operation of the teaching materials may not be synchronized, to ensure that page switching events are aligned with changes in the interactive whiteboard content displayed in the classroom video, in addition to generating a classroom video with the recording length as the timeline, the system time at which the classroom video recording began also needs to be specifically recorded during the recording of the classroom video and the recording of page switching events. Correspondingly, the system time also needs to be recorded for each page switching event. This ensures that when the teaching segment information in the target teaching material page corresponding to the page switching event is used as the video outline information for the classroom video, the final video outline information and the changes in the actual screen in the classroom video can accurately correspond due to the system time. In the specific implementation of acquiring and recording basic data and information, each page switching event needs to record at least the occurrence time of the page switching event and the page information of the teaching material page displayed after the switch. For the entire processing of this application embodiment, the outline information of the classroom video is generated with reference to the courseware page displayed after each page switching event. The page switching event and the courseware page displayed after the switch are directly obtained from the data source. Therefore, the position used to display the teaching courseware does not necessarily need to be within the recording area of the classroom video. Figure 2The image shown is merely an example of a captured image. If the camera is tracking a target, it may no longer capture the display of the teaching materials, but the page switching event for displaying the teaching materials can always be received, and this solution can still be implemented accordingly.
[0044] In addition, classroom videos and teaching materials can be acquired separately. For example, before or after acquiring classroom videos from a camera, users can upload teaching materials through other electronic devices, and the recording host can acquire the teaching materials accordingly. That is, in this application embodiment, there is no requirement that the acquisition of classroom videos and teaching materials have a necessary temporal sequence or synchronous relationship. It is only necessary that there are classroom videos and corresponding teaching materials before the specific data processing process of the video outline generation method in this application embodiment is implemented.
[0045] If the acquisition of classroom video is completed by the recording host, the recording host can transmit data to the camera via HDMI (High Definition Multimedia Interface) cable or other multimedia protocol connection cable. That is, classroom video can be sent from the camera to the recording host through any of the various multimedia protocol connection cables.
[0046] Step S120: Based on the pre-trained teaching segment recognition model, identify the teaching segment corresponding to each page of the teaching courseware to obtain the teaching segment information corresponding to each page.
[0047] In classroom teaching activities, from a macro perspective, it involves teachers and students exchanging information in the same space. However, the teaching activities themselves have clear communication goals; more specifically, they are exchange activities with clear teaching objectives. To achieve these objectives, teachers usually design the teaching process, striving to achieve the established teaching goals through rigorous and meticulous teaching steps. Teaching activities in different subjects and at different grade levels have their own subject-specific characteristics. Taking elementary school Chinese as an example: common teaching steps in elementary school Chinese teaching include introducing the new lesson, learning objectives, vocabulary learning, character introductions, background information, discussion, exercises, lesson summary, consolidating new knowledge, knowledge expansion, and assigning homework.
[0048] To ensure the continuity of teaching activities, teachers generally do not verbally indicate which teaching segment will proceed next. Instead, they use detailed design (such as short stories and transition words) to connect different teaching segments. Although this content is recorded in the classroom video, most existing automatic content recognition solutions can only recognize surface information such as images and audio, and cannot extract and structure the information recorded in the classroom video. The randomness of the specific information recorded in the classroom video during its generation further complicates information extraction and structuring.
[0049] In this embodiment, to ensure accurate structural division of the classroom video, considering that teaching activities are often based on teaching courseware and the recording target of classroom videos also includes the display of teaching courseware, the focus is on structural division of the teaching courseware to achieve structural division of the classroom video. Firstly, the teaching courseware is designed based on the teaching approach for conducting teaching activities; that is, the courseware pages in the teaching courseware must correspond to the designed teaching segments. Secondly, the conduct of teaching activities based on the teaching courseware will inevitably revolve around the display of the teaching courseware. Teachers may adjust teaching details at any time during the teaching process, but the structure of the teaching courseware remains unchanged amidst other changes. Based on the above analysis of specific application scenarios, this embodiment uses the correspondence between courseware pages and teaching segments in the teaching courseware as a reference for generating the outline information of the classroom video.
[0050] In editing teaching courseware, the specific content on the courseware pages is designed specifically for each teaching segment. To reduce the output of invalid information, the courseware pages themselves do not specifically record which teaching segment they correspond to. In this embodiment, a pre-trained teaching segment recognition model is used to identify the corresponding teaching segment on each courseware page to obtain the teaching segment information corresponding to each courseware page.
[0051] The teaching segment recognition model can be trained as follows: First label information is added to each page of multiple first sample coursewares according to predefined teaching segments. Each first sample courseware and its corresponding first label information constitute a first sample dataset. A pre-defined language model is trained based on this dataset to obtain a teaching segment recognition model that takes the courseware as input and outputs a model based on the probability distribution of different label information corresponding to each page of the courseware. The teaching segment information is confirmed based on the basic output.
[0052] For example, for elementary school Chinese teaching materials, using the predefined teaching steps mentioned earlier, a teaching step tag is added to each page of the teaching materials; this is the first tag information. For multiple teaching materials, after manually adding the first tag information to each page, the first sample teaching material dataset D = (d1, d2, ..., d...) can be obtained. n ), where d i =(a i ,b i ), where a i For the content of a certain teaching courseware page, b i The content of the teaching courseware page a i All corresponding first tag information.
[0053] Using the first sample courseware dataset D, a model for identifying teaching segments corresponding to courseware pages is trained based on a pre-built language model; this is known as the teaching segment identification model. The input to the teaching segment identification model is the entire courseware, specifically all its pages. The basic output is the probability distribution of the teaching segments corresponding to each courseware page, i.e., the probability distribution of different label information corresponding to each courseware page. Based on this basic output, the label information with the highest probability can be used as the teaching segment information corresponding to the courseware page.
[0054] Taking BERT as a pre-built language model as an example, assuming the number of teaching segments (i.e., the number of first-label information) is N, all courseware pages can be input into the BERT model. The embedding corresponding to the CLS layer in the last layer of BERT is taken as the courseware page encoding, and then a fully connected layer (with softmax as the activation function) is added. The number of neurons in the output layer is N, corresponding to the probability distribution of N first-label information. The first-label information with the highest probability among the N probability distributions is taken as the final output, that is, the teaching segment information corresponding to the courseware page.
[0055] Let "bert(x)" represent encoding text x using BERT, "[CLS]" represent taking the embedding corresponding to the CLS of the last layer of BERT, f(x) represent the CLS encoding result of text x after BERT, h(x) represent the teaching segment prediction model for the courseware page, out(x) represent the teaching segment prediction result of text x in the courseware page, and W and b represent the weight parameters and bias parameters of the fully connected layer, respectively. When the input courseware content is x... i At that time, the above process can be expressed as:
[0056] f(x i ) = bert(x i [CLS]
[0057] h(x i = softmax(Wf(x) i )+b)
[0058] out(x i ) = argmax(h(x i ))
[0059] After training a certain number of rounds using the first sample courseware dataset D, a teaching segment recognition model can be obtained to predict the teaching segment corresponding to a courseware page.
[0060] Of course, when generating the first courseware sample dataset, a single courseware page and the first label information corresponding to that courseware page can also be used as a data element. The overall training method is roughly the same, so it will not be repeated here.
[0061] Step S130: Based on the page switching event, confirm the target courseware pages that have been displayed and the display time period corresponding to each target courseware page.
[0062] The images in a classroom video are video images, and an important part of the video images is the display of the teaching courseware. Correspondingly, the changes in the content of the video images are strongly related to the operations performed on the teaching courseware during the specific teaching process. That is, operations such as annotation, blackboard writing, and page switching on the teaching courseware will cause changes to the content of the recorded classroom video.
[0063] For teaching courseware, each time a courseware page is displayed, each courseware page usually has a core content. Annotation and blackboard writing operations performed during the display of a courseware page are generally centered on the core content of that courseware page. Therefore, although annotation and blackboard writing operations may cause changes to the content on the screen in the classroom video, the changes are still considered to occur within the display phase of a courseware page. Based on this, when generating the video outline for the classroom video, the display period of each courseware page is used as a reference for dividing the segmented content structure of the video outline, and the information summary of the courseware page is used as the outline information for the content structure of each segment.
[0064] The display changes of courseware pages are determined by page switching operations, such as page turning (including forward and backward page turning) and page jumps. It should be understood that starting and ending the display of a courseware page are also types of page switching operations. In this embodiment, based on the specific information recorded for each page switching event, it is possible to determine which courseware pages enter the display state due to the teacher's page switching operations during the teaching process, and the corresponding display period for each courseware page. In this embodiment, considering that not all courseware pages of a teaching courseware may necessarily enter the display state, or not all courseware pages may necessarily be effectively displayed, only the courseware pages that are effectively displayed are defined as target courseware pages. The page switching event statistics yield the target courseware pages and the corresponding display period for each target courseware page. Effective display refers to display where the display duration reaches a preset threshold value.
[0065] For example, when a teaching courseware is used during a lesson, the system monitors page switching events during the teacher's lecture and records the specific time of each page switching event and the page information displayed after the event. Alternatively, recording page switching events can be a comprehensive information processing process. For instance, based on page changes and specific times, the system records the display period of the most recently displayed target courseware page, i.e., recording the start and end times of the display. If the interval between two adjacent switching operations is very short (e.g., 1 or 2 seconds), the previous switching operation can be considered a transition to displaying the target courseware page (e.g., quickly flipping from page 1 to page 3, with page 2 as a transition). In this case, this series of page switching operations can be corrected into a single page switching event, and the courseware page whose final display duration reaches a preset threshold is taken as the target courseware page. This processing results in a target courseware page display period that actually shows multiple courseware pages, but only one page is displayed normally; the others are transitional displays with no actual display significance, and there is no need to generate corresponding outline information.
[0066] The specific method for confirming the display time period can be designed according to the time information recording method in the page switching event. For example, the moment when the classroom video starts recording can be taken as the inspiration moment for the first segment of content structure, and the courseware page before the first page switching event can be taken as the first target courseware page. Then, the first target courseware page may be the first page of courseware used in this teaching activity (the teacher has already operated the teaching courseware and started displaying the courseware page before the start of the classroom video recording, and the courseware page is displayed in full screen from the beginning of the classroom video), or the first target courseware page may be empty (the slideshow mode is entered after the start of the classroom video recording, and there is no courseware page displayed in full screen at the beginning of the classroom video). Alternatively, the display time period can be divided directly based on the existence of a clear target courseware page. That is, only when a clear target courseware page is displayed is it necessary to confirm the corresponding display time period, and correspondingly, there is a need to confirm the outline information in the corresponding segment of content structure in the classroom video. When the page switching event and the classroom video are recorded on the same timeline, the display time period can be accurately divided at any time the slideshow mode is entered.
[0067] like Figure 4 As shown, in a certain teaching activity, the recording of the classroom video starts at time t0. At time t1, the teacher performs a slideshow playback operation, and a page switching event occurs, entering slideshow playback mode. The first page of the 15-page teaching courseware is displayed. Subsequently, at t2, t3, and t4, the video switches to the second, third, and fourth pages respectively. During this process, the classroom video starting from time t0 is generated, along with the page switching events corresponding to times t1, t2, t3, and t4.
[0068] Another exemplary implementation assumes that the teacher enters slideshow mode before starting the recording of the class video. Twenty seconds after the recording begins, the teacher switches from page 1 to page 2, speaks for two minutes, then switches to page 5, and ends the recording after 10 minutes. Each page switch event is recorded in the format (page start time (in seconds), page end time (in seconds), courseware page number). Therefore, the above page switch events can be recorded as: [(0,20,1),(20,140,2),(140,740,5)]. Alternatively, each page switch event can be recorded as (page switch time (in seconds), page number before switch, page number after switch). Therefore, the above page switch events can be recorded as: [(20,1,2),(140,2,5),(740,5,0)]. For the former recording method, the first two numbers in the page switch event recording format can directly identify the display period, and the third number can identify the target courseware page corresponding to that display period. For the latter recording method, secondary processing is required to confirm the display time period and the corresponding target courseware page. The specific processing result is the same as that of the former recording method, depending on the time of the page switching event and the target courseware page before and after the switch.
[0069] It should be understood that there is no strict order of execution for identifying the teaching segments on the courseware pages and confirming the display time periods. Steps S120 and S130 only indicate the identification of data processing content and do not indicate any restriction on the execution order. In the specific implementation process, processing can be performed in real time according to the data reception status, or in parallel after reception, or any one can be processed first after reception. The final processing result should be able to confirm the display time period when step S140 is executed, as well as the teaching segment information of the courseware page corresponding to each display time period.
[0070] Step S140: Mark the video outline information of the corresponding display time period in the classroom video according to the teaching segment information corresponding to the target courseware page.
[0071] In one alternative implementation, the first label information with the highest probability among N probability distributions can be directly used as the final output, and then used as the corresponding teaching segment information to mark the corresponding display time period on the courseware page, thus completing the marking of video partner information.
[0072] In another specific implementation, there is a correlation between changes in the teaching segment information on the courseware page and changes in the courseware page sequence. For example, the first few pages of the teaching courseware often belong to the introduction of new lessons, while the last page often belongs to the assignment of homework. To further improve the accuracy of the teaching segment recognition results on the courseware page, this embodiment first corrects the teaching courseware and the corresponding teaching segment information based on a pre-trained segment sequence error correction model when marking the video outline information; then, the corrected teaching segment information corresponding to the target courseware page is marked as the video outline information of the corresponding display period in the classroom video. In other words, a segment sequence error correction model is proposed, which corrects the predicted teaching segment information before marking the video outline information.
[0073] The step sequence error correction model is trained as follows: Second label information is acquired for each page of multiple second sample courseware based on predefined teaching steps, along with third label information for each second sample courseware. Each second sample courseware, its corresponding second and third label information, is used as a second sample data set. The third label information is the label with the highest probability among the label information obtained by the teaching step recognition model for each courseware page. A pre-set neural network model is trained based on the second sample courseware dataset to obtain a step sequence error correction model that takes the teaching courseware and its third label information as input, and outputs the correction result of the third label information corresponding to each page of the teaching courseware.
[0074] The sample data used to train the sequence error correction model can be constructed based on the first sample courseware dataset that has already been labeled before training the teaching process recognition model. Assume a manually labeled n-page courseware X, X = (x1, x2, x3, ..., x...). n ), x i This is the i-th page of the courseware, and its corresponding manually annotated result is Y, where Y = (y1, y2, y3, ..., y4). n ), y i This is the first label information of the i-th page of the courseware. Using a trained teaching segment recognition model, the predicted label P is obtained for the courseware X, where P = (p1, p2, p3, ..., p...). n ), p i This is the third label information of the i-th page of courseware content, which is the label information with the highest probability among the label information obtained by the courseware page input teaching link recognition model. At this time, X, P, and Y constitute a training data for the courseware page teaching link error correction model, which is also the second sample courseware dataset.
[0075] When training a sequence correction model based on the second sample courseware dataset, the length of the entire courseware content is much longer than the length of a single page, and also far exceeds the maximum supported length of the pre-defined language model (usually 512 or 1024), posing a significant challenge to the learning of the sequence correction model. To reduce the difficulty of understanding extremely long texts, decrease the model's dependence on GPU memory, and more easily capture the relationship between courseware pages and teaching segments, as well as the relationship between the segment sequence and the sequence of courseware content, the encoding part f(x) of the teaching segment recognition model can be reused to encode the text of the input courseware X, obtaining EX, where EX = (ex1, ex2, ..., ex...). n ), where, ex i =f(x) i ), indicating the courseware page x i The text encoding. The encoding of the predicted teaching segment label P is denoted as EP, that is, each label p i The vector is mapped to a k-dimensional vector (k defaults to 32) through a fully connected layer, denoted as ep. i The encoding of each courseware page (exi) and its corresponding label (epi) is concatenated to form the final encoding (EH) of the courseware page. This final encoding is then passed through a bidirectional LSTM (Long Short-Term Memory) network (E-LSTM), followed by a fully connected layer (assuming weights M, biases Z, and softmax as the activation function). The output layer has N neurons, corresponding to the probability distribution of N teaching segments. The teaching segment with the highest probability among the N teaching segment information probability distributions is taken as the final corrected teaching segment information output of the corresponding courseware page. Assume the final corrected output teaching segment information sequence of courseware X is C, where C = (c1, c2, c3, ..., c...). n ), where c i Indicates the content x of the courseware i And predicted teaching process information p i The teaching process information after model correction; BiLSTM(x) represents inputting vector x into a bidirectional LSTM network, then the above process can be represented as:
[0076] EX = (ex1, ex2, ..., ex...) n )
[0077] EP = (ep1, ep2, ..., ep) n )
[0078] EH=(concat(ex1,ep1),concat(ex2,ep2),...,concat(ex n ,ep n ))
[0079] LSTM = BiLSTM(EH)
[0080] C=argmax(softmax(M(E-LSTM)+Z))
[0081] After training the network for a certain number of rounds using the second sample courseware dataset, a step sequence error correction model can be obtained.
[0082] The above process of constructing a training sample set and conducting training can be referenced in section 3. For a teaching courseware consisting of 15 pages, the teaching segment recognition model may not be able to accurately predict the teaching segment information after recognizing the teaching segment page by page. For example, it may predict the initial courseware page as a class summary. In this case, a training sample set can be further constructed to correct the segment sequence and obtain accurate teaching segment information.
[0083] When marking video outline information, adjacent display periods with the same teaching segment information can be merged. For example, if multiple consecutive display periods are "vocabulary learning", then multiple display periods can be merged into one display period and marked as "vocabulary learning".
[0084] After obtaining accurate teaching segment information in steps S110-S130, this is equivalent to obtaining the video outline information for the display time slot corresponding to each target courseware page used in a recorded classroom video of a lesson. Assume the target courseware pages used by the teacher are X = (x1, x2, x5), x... i This is the i-th page of the courseware. Using the prediction and correction methods for teaching segment information described above, the sequence of teaching segment information used on each page is obtained, which is also the video outline information sequence C, C = (c1, c2, c5), where c... i This refers to the video outline information for the display time segment corresponding to the i-th page of the courseware in the classroom video.
[0085] like Figure 5 As shown, the video outline information 303 and the classroom video 302 are displayed in the same video playback window 301. The video outline information 302 can display the teaching segment information corresponding to each target courseware page, serving as the outline information for the corresponding content segmentation structure, and presenting the start and end time information of the display period (including the start time and end time). When a trigger operation for a certain outline information is received in the video outline information 303, the classroom video 302 jumps to the start time of the corresponding display period.
[0086] The aforementioned video tagging method based on courseware page classification acquires the teaching courseware played during classroom video recording and page switching events of the teaching courseware; based on a pre-trained teaching segment recognition model, it identifies the teaching segment corresponding to each courseware page, obtaining the teaching segment information corresponding to each courseware page; it confirms the displayed target courseware pages and the corresponding display time periods based on page switching events; and it tags the video outline information of the corresponding display time periods in the classroom video based on the teaching segment information corresponding to the target courseware pages. According to common specific teaching segments in the teaching activities, each courseware page of the corresponding teaching activity is mapped to a teaching segment. Specifically, the mapping is predicted using a pre-trained teaching segment recognition model. Furthermore, it records the data sources and change events that cause changes in the main displayed content in the classroom video, confirming the time of occurrence of the change event and the teaching segment of the courseware page displayed afterward. The teaching segment is used as the outline information of the corresponding courseware page's display time period. Without needing to identify the classroom video content, it achieves accurate structural division and outline information identification of classroom videos based on the identification of common segments in teaching activities.
[0087] Figure 6 This is a schematic diagram of a video tagging device based on courseware page classification, provided as an embodiment of this application. Figure 6 As shown, the video tagging device based on courseware page classification includes a classroom data acquisition unit 210, a courseware page classification unit 220, a target page confirmation unit 230, and an outline information tagging unit 240.
[0088] The system includes: a classroom data acquisition unit 210, used to acquire teaching courseware played during classroom video recording and page switching events of the teaching courseware; a courseware page classification unit 220, used to identify the teaching segment corresponding to each courseware page based on a pre-trained teaching segment recognition model, and obtain the teaching segment information corresponding to each courseware page; a target page confirmation unit 230, used to confirm the displayed target courseware pages and the display time period corresponding to each target courseware page based on the page switching events; and an outline information marking unit 240, used to mark the video outline information of the corresponding display time period in the classroom video based on the teaching segment information corresponding to the target courseware page.
[0089] Based on the above embodiments, the teaching process identification model is trained in the following manner:
[0090] Obtain the first tag information added to each page of multiple first sample coursewares according to the predefined teaching steps. Each first sample courseware and its corresponding first tag information are used as a first sample data to obtain the first sample courseware dataset.
[0091] The pre-set language model is trained based on the first sample courseware dataset to obtain a teaching process recognition model that takes the teaching courseware as input and outputs the probability distribution of different label information corresponding to each courseware page in the teaching courseware.
[0092] The information for the teaching process is confirmed based on the basic output.
[0093] Based on the above embodiments, the outline information marking unit 240 includes:
[0094] The information correction module is used to correct the teaching materials and corresponding teaching process information based on the pre-trained process sequence error correction model.
[0095] The information tagging module is used to tag the corrected teaching segment information corresponding to the target courseware page as the video outline information of the corresponding display time segment in the classroom video.
[0096] Based on the above embodiments, the link sequence error correction model is trained in the following manner:
[0097] Obtain the second label information added to each page of multiple second sample coursewares according to the predefined teaching links, as well as the third label information of each second sample courseware. Each second sample courseware and its corresponding second label information and third label information are used as a second sample data to obtain the second sample courseware dataset. The third label information is the label information with the highest probability among the label information obtained by the teaching link recognition model for each courseware page.
[0098] The pre-set neural network model is trained based on the second sample courseware dataset to obtain a sequence error correction model that takes the teaching courseware and its third label information as input, and outputs the correction result of the third label information corresponding to each courseware page in the teaching courseware.
[0099] Based on the above embodiments, the step sequence error correction model and the teaching step identification model reuse the encoding part of the teaching courseware.
[0100] Based on the above embodiments, the outline information marking unit 240 includes:
[0101] The display time period merging module is used to merge adjacent display time periods with the same teaching segment information.
[0102] Based on the above embodiments, page switching events and classroom videos are recorded on the same timeline.
[0103] The video tagging device based on courseware page classification provided in this application embodiment is included in an electronic device and can be used to execute the corresponding video tagging method based on courseware page classification provided in the above embodiment, and has corresponding functions and beneficial effects.
[0104] It is worth noting that in the above embodiments of the video tagging device based on courseware page classification, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.
[0105] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 7 As shown, the electronic device includes a processor 310 and a memory 320, and may also include an input device 330, an output device 340, and a communication device 350; the number of processors 310 in the electronic device may be one or more. Figure 7 Taking a processor 310 as an example; the processor 310, memory 320, input device 330, output device 340, and communication device 350 in the electronic device can be connected via a bus or other means. Figure 7 Taking the example of a connection between China and Israel via a bus.
[0106] The memory 320, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the video tagging method based on courseware page classification in this embodiment. The processor 310 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 320, thereby implementing the aforementioned video tagging method based on courseware page classification.
[0107] The memory 320 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 320 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 320 may further include memory remotely located relative to the processor 310, which can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0108] Input device 330 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the electronic device. Output device 340 may include display devices such as a display screen.
[0109] The aforementioned electronic device includes a video tagging device based on courseware page classification, which can be used to execute any video tagging method based on courseware page classification, and has corresponding functions and beneficial effects.
[0110] This application also provides a computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program performs relevant operations in the video tagging method based on courseware page classification provided in any embodiment of this application, and has corresponding functions and beneficial effects.
[0111] Those skilled in the art will understand that embodiments of this application may be provided as methods, systems, or computer program products.
[0112] Therefore, this application may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, produce implementations of the flowchart... Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0113] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory. Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0114] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0115] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0116] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. A video tagging method based on courseware page classification, characterized in that, include: Acquire the teaching materials played during the classroom video recording process and the page switching events of the teaching materials; Based on the pre-trained teaching segment recognition model, the teaching segment corresponding to each page of the teaching courseware is identified to obtain the teaching segment information corresponding to each page. Based on the page switching event, confirm the target courseware pages that have been displayed and the display time period corresponding to each target courseware page; The video outline information for the corresponding display time period in the classroom video is marked according to the teaching segment information corresponding to the target courseware page; The video outline information is used to display in the same video playback window as the classroom video. When a trigger operation is received on a certain outline information in the video outline information, the classroom video jumps to the start time of the corresponding display period for playback.
2. The video tagging method based on courseware page classification according to claim 1, characterized in that, The teaching segment identification model was trained in the following way: Obtain the first tag information added to each page of multiple first sample coursewares according to the predefined teaching steps. Each first sample courseware and its corresponding first tag information are used as a first sample data to obtain the first sample courseware dataset. The pre-set language model is trained based on the first sample courseware dataset to obtain a teaching process recognition model that takes the teaching courseware as input and outputs the probability distribution of different label information corresponding to each courseware page in the teaching courseware. The teaching process information is confirmed based on the basic output.
3. The video tagging method based on courseware page classification according to claim 1 or 2, characterized in that, The step of marking the video outline information of the corresponding display time period in the classroom video according to the teaching segment information corresponding to the target courseware page includes: The teaching materials and corresponding teaching process information are corrected based on a pre-trained process sequence error correction model. The corrected teaching segment information corresponding to the target courseware page is marked as the video outline information of the corresponding display time period in the classroom video.
4. The video tagging method based on courseware page classification according to claim 3, characterized in that, The sequence error correction model is trained in the following manner: The second label information added to each page of multiple second sample coursewares according to the predefined teaching links, and the third label information of each second sample courseware are obtained. Each second sample courseware and its corresponding second label information and third label information are used as a second sample data to obtain the second sample courseware dataset. The third label information is the label information with the highest probability among the label information obtained by the teaching link recognition model for each courseware page. The pre-set neural network model is trained based on the second sample courseware dataset to obtain a sequence error correction model that takes the teaching courseware and its third label information as input and the correction result of the third label information corresponding to each courseware page as output.
5. The video tagging method based on courseware page classification according to claim 4, characterized in that, The sequence error correction model and the teaching segment identification model reuse the encoding part of the teaching courseware.
6. The video tagging method based on courseware page classification according to claim 1 or 2, characterized in that, The step of marking the video outline information of the corresponding display time period in the classroom video according to the teaching segment information corresponding to the target courseware page includes: Adjacent time slots with identical teaching information will be merged.
7. The video tagging method based on courseware page classification according to claim 1 or 2, characterized in that, The page switching events and the classroom videos are recorded on the same timeline.
8. A video tagging device based on courseware page classification, characterized in that, include: The classroom data acquisition unit is used to acquire the teaching courseware played during the classroom video recording process and the page switching events of the teaching courseware; The courseware page classification unit is used to identify the teaching segment corresponding to each courseware page based on the pre-trained teaching segment recognition model, and obtain the teaching segment information corresponding to each courseware page. The target page confirmation unit is used to confirm the displayed target courseware pages and the display time period corresponding to each target courseware page based on the page switching event. The outline information marking unit is used to mark the video outline information of the corresponding display period in the classroom video according to the teaching segment information corresponding to the target courseware page; The video outline information is used to display in the same video playback window as the classroom video. When a trigger operation is received on a certain outline information in the video outline information, the classroom video jumps to the start time of the corresponding display period for playback.
9. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more computer programs; When the one or more computer programs are executed by the one or more processors, the electronic device implements the video tagging method based on courseware page classification as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the video tagging method based on courseware page classification as described in any one of claims 1-7.
Citation Information
Patent Citations
Video direct broadcasting and video-on-demand system for teaching
CN107277598A
Test question video generation method based on courseware attached video, storage medium and equipment
CN115883867A