A method and device for analyzing video teaching content

By drawing frames to identify teaching videos and building a teaching knowledge tree, the problem of difficulty in capturing the semantic information of teaching content in the existing technology is solved, and the efficient organization and retrievalability of video content is achieved.

CN118447435BActive Publication Date: 2025-05-06DMAI (GUANGZHOU) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410671446.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-28
Publication Date
2025-05-06
Estimated Expiration
2044-05-28

AI Technical Summary

Technical Problem

Existing video analysis methods are difficult to effectively capture semantic information in teaching content videos, such as chapter structure and knowledge points, which leads to difficulties in organizing, classifying and retrieving video content.

Method used

By extracting frames to the teaching video to be analyzed, the teaching directory page, the teaching title page and the teaching content page are identified, and the teaching knowledge tree is constructed based on this information as the analysis result of the video.

Benefits of technology

It improves the retrievalability of video content, helps learners quickly locate and understand key knowledge points in the video, and enhances the organization and usability of teaching content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118447435B_ABST
    Figure CN118447435B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for analyzing video teaching content, comprising: firstly obtaining a teaching video to be analyzed, and obtaining candidate video frames with timestamps by frame extraction processing. Then, the teaching catalog page, teaching title page and teaching content page are identified from these frames. Finally, a teaching knowledge tree is constructed based on this information as the analysis result of the video. Such a design not only improves the retrievability of the video content, but also helps learners to quickly locate and understand the key knowledge points in the video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video processing technology, and in particular to a method and device for analyzing video teaching content. Background Art

[0002] With the development of Internet technology, online teaching videos have become an important learning resource. However, with the surge in the number of videos, how to effectively organize, classify and retrieve video content has become a challenge. Existing video analysis methods often focus on extracting underlying features of videos, such as color, texture, etc., but for teaching content videos, these methods are difficult to capture their semantic information, such as chapter structure, knowledge points, etc. Therefore, it is particularly important to develop a method that can perform in-depth analysis of video teaching content. Summary of the invention

[0003] The purpose of the present invention is to provide a method and device for analyzing video teaching content.

[0004] In a first aspect, an embodiment of the present invention provides a method for analyzing video teaching content, comprising:

[0005] Get the teaching video to be analyzed;

[0006] Performing frame extraction processing on the teaching video to be analyzed to obtain a plurality of candidate teaching video frames, wherein the candidate teaching video frames are configured with timestamps;

[0007] Determining a teaching catalog page, a teaching title page, and a teaching content page from the plurality of candidate teaching video frames respectively;

[0008] According to the teaching catalog page, teaching title page and teaching content page, a teaching knowledge tree is constructed as the analysis result of the teaching video to be analyzed.

[0009] In the embodiment of the present invention, the frame extraction process is performed on the teaching video to be analyzed to obtain a plurality of candidate teaching video frames, including:

[0010] Extract frames from the teaching video to be analyzed at preset time intervals to obtain multiple initial teaching video frames;

[0011] The multiple initial teaching video frames are deduplicated to obtain the multiple candidate teaching video frames.

[0012] In the embodiment of the present invention, the deduplication processing of the multiple initial teaching video frames to obtain the multiple candidate teaching video frames includes:

[0013] Calculating structural similarity indexes of two adjacent initial teaching video frames in the plurality of initial teaching video frames;

[0014] If there is a structure similarity index greater than a preset similarity index threshold, the initial teaching video frame located in the latter order is deleted, and the structure similarity indexes of the initial teaching video frames adjacent to each other in the remaining initial teaching video frames are recalculated until the deduplication process is completed;

[0015] If there is no structure similarity index greater than the preset similarity index threshold, the deduplication process is terminated;

[0016] The initial teaching video frames after multiple duplications are removed are used as the multiple candidate teaching video frames.

[0017] In an embodiment of the present invention, determining the teaching catalog page from the plurality of candidate teaching video frames includes:

[0018] Extracting text content of a target candidate teaching video frame, wherein the target candidate teaching video frame is any candidate teaching video frame among the multiple candidate teaching video frames;

[0019] Analyze the text content to obtain keywords included in the text content;

[0020] In the case where the keyword includes a preset catalog keyword, the content corresponding to the target candidate teaching video is used as the teaching catalog page.

[0021] In an embodiment of the present invention, determining the teaching title page from the plurality of candidate teaching video frames includes:

[0022] Acquire text font information of a target candidate teaching video frame, wherein the target candidate teaching video frame is any candidate teaching video frame among the multiple candidate teaching video frames, and the text font information includes font size, font position and keyword information;

[0023] If the title content is determined from the text font information according to the font size, font position and keyword information, the content corresponding to the target candidate teaching video is used as the teaching title page.

[0024] In an embodiment of the present invention, determining the teaching content page from the plurality of candidate teaching video frames includes:

[0025] The contents other than the teaching catalog page and the teaching title page are regarded as the teaching content page.

[0026] In an embodiment of the present invention, the step of constructing a teaching knowledge tree as the analysis result of the teaching video to be analyzed according to the teaching catalog page, the teaching title page and the teaching content page includes:

[0027] Obtaining the first keyword of all the teaching catalog pages, and using the first keyword as the index basis of the teaching knowledge tree;

[0028] Acquire a second keyword of each teaching title page, and determine a key knowledge point corresponding to each teaching title page according to the second keyword;

[0029] Determine the association relationship between each of the teaching content pages and each of the teaching catalog pages and each of the teaching title pages, and determine the positional relationship between the teaching catalog pages, teaching title pages and teaching content pages in constructing the teaching knowledge tree;

[0030] The teaching knowledge tree is constructed according to the index basis, the key knowledge points, the position relationship and the timestamp, and there is a binding relationship between the key knowledge points and the timestamp.

[0031] In an embodiment of the present invention, the method further includes:

[0032] In response to the video frame modification operation, performing at least one of adding, deleting, and adjusting the order of the plurality of candidate teaching video frames;

[0033] In response to the text adjustment operation, modify the text content included in the teaching catalog page, the teaching title page and the teaching content page;

[0034] In response to the knowledge tree saving operation, the teaching knowledge tree is saved and bound to the video identifier of the teaching video to be analyzed.

[0035] In the embodiment of the present invention, the step of obtaining the teaching video to be analyzed includes:

[0036] Set up synchronization mechanism with the preset teaching video database;

[0037] When it is detected that a new teaching video to be processed is stored in the preset teaching video database, it is determined whether the teaching video to be processed is a recorded demonstration video and whether the video duration exceeds a preset minimum duration;

[0038] If yes, the teaching video to be processed is used as the teaching video to be analyzed;

[0039] If not, the to-be-processed teaching video is deleted from the preset teaching video database.

[0040] In a second aspect, an embodiment of the present invention provides an analysis device for video teaching content, comprising:

[0041] An acquisition module is used to acquire the teaching video to be analyzed;

[0042] The analysis module is used to extract frames from the teaching video to be analyzed to obtain multiple candidate teaching video frames, each of which is configured with a timestamp; determine a teaching catalog page, a teaching title page, and a teaching content page from the multiple candidate teaching video frames; and construct a teaching knowledge tree as the analysis result of the teaching video to be analyzed based on the teaching catalog page, the teaching title page, and the teaching content page.

[0043] Compared with the prior art, the beneficial effects provided by the present invention include: adopting a method and device for analyzing video teaching content disclosed by the present invention, by obtaining the teaching video to be analyzed, and obtaining candidate video frames with timestamps through frame extraction processing. Then, the teaching catalog page, teaching title page and teaching content page are identified from these frames. Finally, a teaching knowledge tree is constructed based on this information as the analysis result of the video. Such a design not only improves the searchability of the video content, but also helps learners quickly locate and understand the key knowledge points in the video. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative work.

[0045] Figure 1 A schematic diagram of the structure provided by an embodiment of the present invention;

[0046] Figure 2 A schematic diagram of the structure provided by an embodiment of the present invention;

[0047] Figure 3 A schematic diagram of the structure provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0049] The specific implementation modes of the present invention are described in detail below in conjunction with the accompanying drawings.

[0050] In order to solve the technical problems in the aforementioned background technology, Figure 1This is a flow chart of a method for analyzing video teaching content provided in an embodiment of the present disclosure. The method for analyzing video teaching content is introduced in detail below.

[0051] Step S201, obtaining a teaching video to be analyzed;

[0052] Step S202, extracting frames from the teaching video to be analyzed to obtain a plurality of candidate teaching video frames, wherein the candidate teaching video frames are configured with timestamps;

[0053] Step S203, determining a teaching catalog page, a teaching title page, and a teaching content page from the plurality of candidate teaching video frames respectively;

[0054] Step S204, constructing a teaching knowledge tree as the analysis result of the teaching video to be analyzed according to the teaching catalog page, the teaching title page and the teaching content page.

[0055] In an embodiment of the present invention, illustratively, the server retrieves and obtains a teaching video to be analyzed from the video library. This video is an online course recording on "Basics of Probability Theory in Mathematics", which is about 45 minutes long and in MP4 format. The server uses professional video processing software to extract frames for this teaching video. The frequency of extracting frames is set to one frame per second to ensure that all key information in the video can be captured. After the frame extraction process, the server obtains 2700 candidate teaching video frames with timestamps. The server runs an image recognition algorithm to analyze these 2700 candidate teaching video frames one by one. Through analysis, the server determines the following key frames: In the frame with a timestamp of 00:03, a clear teaching catalog page is identified, which lists the main content and chapter arrangement of this lesson. In the frame with a timestamp of 00:10, a teaching title page is identified, which clearly reads the title "Basics of Probability Theory in Mathematics". In frames with timestamps of 05:20, 10:40, 15:50 and other time points, teaching content pages containing specific mathematical formulas and examples were identified. The server began to build a teaching knowledge tree based on the identified teaching catalog page, teaching title page and teaching content page. First, take "Basics of Probability Theory in Mathematics" as the root node, and then add each chapter as a child node based on the content of the teaching catalog page. Then, based on the identified teaching content page, add specific mathematical formulas, concepts, examples, etc. to the child nodes of the corresponding chapters. Finally, the server built a complete teaching knowledge tree that clearly shows the knowledge structure and content hierarchy of this lesson. This teaching knowledge tree can not only help learners better understand the course content, but also serve as a reference for teachers to prepare lessons and evaluate teaching.

[0056] In the embodiment of the present invention, the aforementioned step S202 may be implemented through the following example execution.

[0057] Extract frames from the teaching video to be analyzed at preset time intervals to obtain multiple initial teaching video frames;

[0058] The multiple initial teaching video frames are deduplicated to obtain the multiple candidate teaching video frames.

[0059] In an embodiment of the present invention, illustratively, in the previous scenario, the server has obtained an online course recording on "Basics of Probability Theory in Mathematics". Now, the server will extract frames for this 45-minute teaching video at a preset time interval, for example, every 5 seconds. This means that during the entire video playback process, the server will extract a frame every 5 seconds as the initial teaching video frame. Therefore, in this 45-minute (i.e., 2700 seconds) video, the server will extract about 540 initial teaching video frames. Among the 540 extracted initial teaching video frames, there may be some duplicate frames, especially when there are still pictures in the video or when the teacher switches slides during the explanation. In order to improve the accuracy and efficiency of subsequent analysis, the server will perform deduplication processing on these initial teaching video frames. Specifically, the server will run an image similarity algorithm to compare the similarity between adjacent initial teaching video frames. If two adjacent frames are found to be highly similar (for example, the similarity exceeds 95%), they are considered to be duplicate frames and one of them is deleted. Through this process, the server can remove redundant and repeated initial teaching video frames, thereby obtaining a set of streamlined and non-repetitive candidate teaching video frames. For example, out of the original 540 initial teaching video frames, after deduplication processing, there may be about 400 candidate teaching video frames remaining. These candidate teaching video frames will serve as the basic data for subsequent steps (such as determining the teaching catalog page, teaching title page, and teaching content page, etc.). Through such frame extraction and deduplication processing, the server can more effectively analyze the teaching video, extract key teaching information, and provide a solid foundation for building an accurate teaching knowledge tree.

[0060] In the embodiment of the present invention, the aforementioned step of performing deduplication processing on the multiple initial teaching video frames to obtain the multiple candidate teaching video frames can be implemented through the following example.

[0061] Calculating structural similarity indexes of two adjacent initial teaching video frames in the plurality of initial teaching video frames;

[0062] If there is a structure similarity index greater than a preset similarity index threshold, the initial teaching video frame located in the latter order is deleted, and the structure similarity indexes of the initial teaching video frames adjacent to each other in the remaining initial teaching video frames are recalculated until the deduplication process is completed;

[0063] If there is no structure similarity index greater than the preset similarity index threshold, the deduplication process is terminated;

[0064] The initial teaching video frames after multiple duplications are removed are used as the multiple candidate teaching video frames.

[0065] In an embodiment of the present invention, illustratively, after the previous frame extraction process, the server obtains 540 initial teaching video frames. Now, the server starts to calculate the structural similarity index (SSIM) between two adjacent frames in these initial teaching video frames. SSIM is an indicator to measure the similarity between two images, taking into account three aspects: brightness, contrast and structure. For example, the server first calculates the SSIM between the 1st frame and the 2nd frame, then the SSIM between the 2nd frame and the 3rd frame, and so on, until the SSIM of the last pair of adjacent frames (the 539th frame and the 540th frame) is calculated. The server sets a similarity index threshold, such as 0.98, to determine whether two adjacent frames are too similar. After calculating the SSIM of all adjacent frames, the server checks whether these SSIM values ​​exceed the set threshold. Assuming that the SSIM value between the 20th frame and the 21st frame is 0.99 during the calculation process, which exceeds the set threshold of 0.98, the server will judge that the two frames are too similar and decide to delete one of them. In this example, the server chooses to delete the 21st frame in the subsequent sequence. After deleting the 21st frame, the original 22nd frame now becomes the new 21st frame. The server needs to recalculate the SSIM value between the new 20th frame and the 21st frame to ensure the accuracy of the deduplication process. This process will continue until the SSIM value of no adjacent frames exceeds the set threshold. If after a certain round of calculation, the server finds that the SSIM values ​​of all adjacent frames do not exceed the set threshold, then it will determine that the current initial teaching video frame set is sufficient for deduplication, and the deduplication process will end at this time. After the above deduplication process, the original 540 initial teaching video frames may be reduced to 400 or less initial teaching video frames after deduplication. These initial teaching video frames after deduplication are now called "candidate teaching video frames", which will be used for subsequent teaching content analysis, such as determining the teaching catalog page, teaching title page, and teaching content page. Through such a deduplication process, the server can ensure that each frame in the candidate teaching video frame set contains unique and valuable teaching information, providing a high-quality data foundation for subsequent teaching content analysis.

[0066] In the embodiment of the present invention, the aforementioned step of determining the teaching catalog page from the multiple candidate teaching video frames can be implemented through the following example.

[0067] Extracting text content of a target candidate teaching video frame, wherein the target candidate teaching video frame is any candidate teaching video frame among the multiple candidate teaching video frames;

[0068] Analyze the text content to obtain keywords included in the text content;

[0069] In the case where the keyword includes a preset catalog keyword, the content corresponding to the target candidate teaching video is used as the teaching catalog page.

[0070] In the embodiment of the present invention, illustratively, the server has extracted frames from the teaching video to be analyzed and removed duplicates to obtain multiple candidate teaching video frames. Now, the server begins to determine the teaching catalog page from these candidate teaching video frames. First, the server selects a target candidate teaching video frame, which is the slide shown by the teacher at the beginning of the lecture, and is likely to be a teaching catalog page. The server uses OCR (optical character recognition) technology to extract the text content in this frame. For example, OCR technology recognizes the text in the frame: "Chapter 1 Introduction", "Chapter 2 Probability Theory Foundation", "Chapter 3 Random Variables and Their Distribution", etc. The server analyzes the extracted text content to obtain keywords. In this process, the server may use natural language processing technology, such as word frequency analysis, TF-IDF and other methods to extract keywords. In our example, the keywords identified by the server may include "Introduction", "Probability Theory Foundation", "Random Variables", etc., which are closely related to the teaching content. The server has a preset directory keyword list, which may include common directory structure words such as "chapter", "section", and "part". Now, the server compares the extracted keywords with this preset list. In our example, "chapter" is a preset catalog keyword, and the keywords extracted by the server from the text content do include "chapter", which indicates that this candidate teaching video frame is likely to be a teaching catalog page. Since the extracted keywords match the preset catalog keywords, the server determines that this candidate teaching video frame is a teaching catalog page. Therefore, it saves the content of this frame as a teaching catalog page for subsequent use. Through this process, the server can accurately determine the teaching catalog page from multiple candidate teaching video frames, providing important information for subsequent teaching content analysis and knowledge tree construction.

[0071] In the embodiment of the present invention, the aforementioned step of determining the teaching title page from the plurality of candidate teaching video frames can be implemented through the following example.

[0072] Acquire text font information of a target candidate teaching video frame, wherein the target candidate teaching video frame is any candidate teaching video frame among the multiple candidate teaching video frames, and the text font information includes font size, font position and keyword information;

[0073] If the title content is determined from the text font information according to the font size, font position and keyword information, the content corresponding to the target candidate teaching video is used as the teaching title page.

[0074] In the embodiment of the present invention, the server continues to process the candidate teaching video frames obtained after extracting frames from the teaching video and removing duplicates. In order to determine the teaching title page, the server selects a target candidate teaching video frame for analysis. This frame may be a slide containing the course title displayed before the teacher starts the formal explanation. The server uses OCR technology and image processing technology to extract the text font information in this frame. This information includes font size, font position and keywords. For example, OCR technology may recognize that there is a line of large font text in the frame: "Basics of Probability Theory in Mathematics", and at the same time recognize that the font size of this line of text is significantly larger than other text and is located in the center of the slide. The server analyzes based on the extracted text font information. First, it will focus on those texts whose font size is significantly larger than the surrounding text, because these texts are likely to be titles. Second, the server will consider the position of the text, and the title will usually be placed in a prominent position on the slide, such as the center or top. Finally, the server will also combine keyword information to make a judgment, for example, certain specific words or phrases may be more likely to be considered as part of the title. In our example, the line of text "Basics of Probability Theory in Mathematics" is likely to be recognized as title content because of its large font and location in the center. Based on the above analysis, the server determines that the "Basics of Probability Theory in Mathematics" contained in this candidate teaching video frame is a teaching title. Therefore, it saves the content of this frame as a teaching title page for subsequent use. Through this process, the server can accurately determine the teaching title page from multiple candidate teaching video frames, which provides key information for constructing a teaching knowledge tree and subsequent teaching content analysis.

[0075] In the embodiment of the present invention, the aforementioned step of determining the teaching content page from the multiple candidate teaching video frames can be implemented through the following example.

[0076] The contents other than the teaching catalog page and the teaching title page are regarded as the teaching content page.

[0077] In an embodiment of the present invention, exemplarily, in the previous steps, the server has successfully identified and extracted the teaching catalog page and the teaching title page from multiple candidate teaching video frames. Now, the server needs to determine the teaching content page. The teaching content page usually contains specific knowledge points, examples, explanations, etc. of the course, which is all the content in the teaching video except the catalog page and the title page. The server first excludes the candidate teaching video frames that have been determined as teaching catalog pages and teaching title pages. These frames have been identified as specific teaching structural elements and therefore do not belong to the teaching content page. Next, the server classifies the remaining candidate teaching video frames as teaching content pages. These frames contain the specific content explained by the teacher, such as formula derivation, case analysis, experimental operation, etc. For example, in an online course on "Basics of Probability Theory in Mathematics", the server has identified the teaching catalog page and the teaching title page. Among the remaining candidate teaching video frames, one frame shows that the teacher is explaining a specific probability calculation formula, and another frame is the teacher showing the steps to solve a probability problem. These frames are determined by the server as teaching content pages because they contain specific teaching knowledge points and explanation content. Through this step, the server can accurately determine the teaching content page from multiple candidate teaching video frames, which provides basic data for subsequent in-depth analysis of the teaching content and knowledge extraction.

[0078] In the embodiment of the present invention, the aforementioned step S204 may be implemented through the following example execution.

[0079] Obtaining the first keyword of all the teaching catalog pages, and using the first keyword as the index basis of the teaching knowledge tree;

[0080] Acquire a second keyword of each teaching title page, and determine a key knowledge point corresponding to each teaching title page according to the second keyword;

[0081] Determine the association relationship between each of the teaching content pages and each of the teaching catalog pages and each of the teaching title pages, and determine the positional relationship between the teaching catalog pages, teaching title pages and teaching content pages in constructing the teaching knowledge tree;

[0082] The teaching knowledge tree is constructed according to the index basis, the key knowledge points, the position relationship and the timestamp, and there is a binding relationship between the key knowledge points and the timestamp.

[0083] In an embodiment of the present invention, illustratively, the server has determined the teaching catalog pages, and now it extracts the first keyword from these pages. For example, in a course about "Programming Basics", the teaching catalog page may contain multiple chapter titles, such as "Chapter 1: Overview of Programming Languages", "Chapter 2: Variables and Data Types", etc. The server uses these chapter titles as the first keyword, and they will become the index basis when constructing the teaching knowledge tree. Next, the server extracts the second keyword from each teaching title page. These keywords represent the core knowledge points explained by each title page. For example, under "Chapter 2: Variables and Data Types", a teaching title page may be "2.1 Definition and Declaration of Variables", and the server uses "Definition and Declaration of Variables" as the second keyword, which indicates the key knowledge points explained by the page. The server further analyzes the teaching content pages to determine their association with the teaching catalog pages and teaching title pages. This is usually achieved by comparing text content, keywords or timestamps. For example, if a teaching content page explains "Definition and Declaration of Variables" in detail, it will be associated with the corresponding teaching title page "2.1 Definition and Declaration of Variables", and then also associated with the teaching catalog page "Chapter 2: Variables and Data Types". Finally, the server constructs a teaching knowledge tree based on the collected information. The root node of the tree is the course name, such as "Programming Basics". The second-level node of the tree is the first keyword of the teaching catalog page, that is, the title of each chapter. The third-level node is the second keyword of the teaching title page, representing each subsection or specific knowledge point. The leaf nodes are the teaching content pages related to these knowledge points, and they are placed in the correct position in the tree through association relationships. In addition, each node (especially the key knowledge point nodes) is bound to a timestamp. This means that when a user clicks on a certain knowledge point, the server can directly locate the corresponding time period in the teaching video according to the timestamp, which is convenient for users to quickly find and learn specific content. Through this process, the server constructs a structured teaching knowledge tree, which clearly shows the hierarchy of the course content and the relationship between knowledge points, greatly improving the user's learning efficiency and experience.

[0084] In the embodiments of the present invention, the following implementation modes are also provided.

[0085] In response to the video frame modification operation, performing at least one of adding, deleting, and adjusting the order of the plurality of candidate teaching video frames;

[0086] In response to the text adjustment operation, modify the text content included in the teaching catalog page, the teaching title page and the teaching content page;

[0087] In response to the knowledge tree saving operation, the teaching knowledge tree is saved and bound to the video identifier of the teaching video to be analyzed.

[0088] In an embodiment of the present invention, exemplarily, the server receives a video frame modification operation request, which comes from an operation manager of a teaching platform, and he hopes to improve the teaching knowledge tree automatically generated before by adjusting the candidate teaching video frame. Specifically, the operation manager of the teaching platform finds that a key teaching content page is omitted, so he hopes to add it to the teaching knowledge tree through a new operation. According to the request of the operation manager of the teaching platform, the server adds the omitted teaching content page in multiple candidate teaching video frames, and reconstructs the teaching knowledge tree to ensure that the new teaching content page is correctly added to the tree. Subsequently, the server receives another text adjustment operation request. This time, the operation manager of the teaching platform finds that the text content of a certain teaching title page is wrong and needs to be modified. He corrects the wrong title content to the correct title through the text adjustment operation. The server updates the text content of the teaching title page according to the modification of the operation manager of the teaching platform, and reconstructs the teaching knowledge tree to ensure that all the text contents are accurate. After completing the above modification, the operation manager of the teaching platform is satisfied with the modified teaching knowledge tree and hopes to save it for subsequent use. Therefore, he initiated a knowledge tree save operation. In response to this save operation, the server saved the modified teaching knowledge tree and bound it to the video identifier of the corresponding teaching video to be analyzed. In this way, whenever this teaching video is accessed, the server will provide the latest version of the teaching knowledge tree bound to it. Through the above steps, the server can not only flexibly modify and improve the teaching knowledge tree according to the user's feedback, but also ensure that the teaching knowledge tree provided each time is the latest version that matches the corresponding teaching video content.

[0089] In the embodiment of the present invention, the aforementioned step S201 can be implemented through the following example execution.

[0090] Set up synchronization mechanism with the preset teaching video database;

[0091] When it is detected that a new teaching video to be processed is stored in the preset teaching video database, it is determined whether the teaching video to be processed is a recorded demonstration video and whether the video duration exceeds a preset minimum duration;

[0092] If yes, the teaching video to be processed is used as the teaching video to be analyzed;

[0093] If not, the to-be-processed teaching video is deleted from the preset teaching video database.

[0094] In an embodiment of the present invention, illustratively, the server first establishes a synchronization mechanism with a preset teaching video database. This means that whenever a new video file is stored in the database, the server will be notified immediately. For example, an educational institution has a database specifically for storing teaching videos. Whenever a teacher uploads a new teaching video, the database will notify the server through a synchronization mechanism that a new video file is stored. When the server receives the notification from the database, it will immediately detect the newly stored teaching video to be processed. This step is to confirm whether the video meets the criteria for further analysis. For example, the server may detect a newly uploaded video file named "Python Programming Basics". The server will then determine whether the new video is a recorded demonstration video and whether the video duration exceeds the preset minimum duration. The recorded demonstration video usually contains the teacher's explanation and specific operation demonstration, and is an important resource for building a teaching knowledge tree. At the same time, the video duration is also an important judgment criterion, because a video that is too short may not contain enough teaching content. For example, the server may determine that the video "Python Programming Basics" is a recorded demonstration video, and the duration is 1 hour and 20 minutes, which exceeds the preset minimum duration standard (such as 30 minutes). If the teaching video to be processed meets the above conditions, that is, it is a recorded demonstration video and exceeds the preset minimum duration, the server will mark it as a teaching video to be analyzed. In our example, the "Python Programming Basics" video will be selected by the server as the teaching video to be analyzed. If the teaching video to be processed does not meet the above conditions, the server will delete it from the preset teaching video database. For example, if the uploaded video is too short or is not a recorded demonstration video, the server will perform a deletion operation to save storage space and ensure the accuracy of subsequent analysis. Through the above steps, the server can effectively screen out recorded demonstration videos that meet the analysis requirements and provide a high-quality data foundation for the subsequent construction of the teaching knowledge tree.

[0095] In order to more clearly describe the solution of the embodiment of the present invention, a more complete implementation of the embodiment of the present invention is provided below.

[0096] 1. Get the video to be analyzed:

[0097] Users can obtain the video to be analyzed in two ways.

[0098] One is to directly upload local video files;

[0099] The second is to submit the video online through network transmission.

[0100] To ensure the accuracy and effectiveness of the analysis, the submitted videos must meet certain conditions, such as screen recordings containing PPTs, excluding lectures and promotional videos, to ensure that the teaching information can be accurately captured. The video content should be teaching-related, and the recommended length is between 5 and 45 minutes, but longer videos can also be accommodated.

[0101] 2. Screenshot of PPT page with frame extraction and recognition:

[0102] The system uses the ffmpeg tool to extract frames from the video at the same preset time interval to generate a series of static video frames with timestamps.

[0103] Frame extraction is to extract frames from the video by sampling at equal time intervals. Assuming the total length of the video is L seconds and the extraction interval is t seconds, the total number of frames extracted N can be calculated by the following formula:

[0104]

[0105] in, Denotes rounding down. Deduplication is achieved by calculating the visual difference between consecutive frames. The visual difference can be measured by calculating the structural similarity index (SSIM) between consecutive frames. SSIM is a measure of the visual similarity between two images, and its value ranges from -1 to 1, with 1 indicating that they are exactly the same. A threshold τ (for example, 0.85) is set. When SSIM is greater than τ, the two frames are considered similar and the subsequent frames are removed.

[0106] After the frame extraction is completed, there should be more similar content. Through image similarity algorithms, such as structural similarity index (SSIM) or average hashing (aHash), continuous frames are compared and repeated frames with high similarity are removed to reduce the subsequent calculation burden.

[0107] 3. Automatically identify the catalog page:

[0108] Using OCR technology, the system automatically identifies the table of contents from the keyframes. Usually, the table of contents contains the outline or chapter titles of the entire presentation. Text analysis can be used to extract the text content in the keyframes through OCR (Optical Character Recognition) technology and analyze the text to find keywords that may indicate a table of contents, such as "table of contents", "outline", "chapter", etc.

[0109] If the frame extraction process captures multiple similar catalog pages, the system retains the clearest (highest confidence) catalog page based on the confidence score. During the OCR process, the image is converted to a grayscale color space, and necessary noise reduction processing is performed. The text detection algorithm is applied to locate the text area, and finally the detected text blocks are recognized using OCR engines such as Tesseract. For the recognized text, the system further performs keyword extraction and semantic analysis to provide support for subsequent knowledge extraction and video indexing. If there is a recognition error or missed recognition, the user can manually intervene to mark and correct the content of the main catalog page.

[0110] 4. Automatically identify the title page:

[0111] The system further identifies image images containing first-level, second-level, and third-level titles, and automatically locates and frames the title area.

[0112] The title page is generally identified through text analysis or style matching. Usually, the title page contains the theme or title of the entire presentation, and these words often have a large font size and a prominent position. By extracting the text content in the key frame through OCR technology, the font size, position and keywords of the text can be analyzed to determine whether it is a title page.

[0113] In addition to applying the above-mentioned OCR technology, the recognition and transcription of title text will also utilize syntactic analysis and dependency analysis in natural language processing (NLP) to understand the text structure and semantics.

[0114] On this basis, the system extracts key knowledge points based on the title content, which may involve summarizing and simplifying the paragraph content. If there is a professional knowledge point vocabulary, the system can directly refer to the vocabulary to extract keywords; if there is no ready-made vocabulary, it is necessary to use the large language model LLM to summarize and condense the knowledge points.

[0115] 5. Identify non-directory pages, i.e. content pages:

[0116] Excluding the table of contents and title page, other pages are considered content pages.

[0117] For non-catalog page content, the system will calculate its relevance to the catalog page to determine its position in the video knowledge structure.

[0118] In addition, the system will deduplicate content pages, remove duplicate or very similar frames, and retain only the most representative and relevant frame content.

[0119] 6. Manual assisted information verification:

[0120] Although automation technology has greatly improved processing efficiency, manual verification is still indispensable. Users need to review the results of the system's automatic recognition. Once misidentification or missing information is found, manual corrections and adjustments should be made in a timely manner to ensure that every title page and its content are accurate.

[0121] 7. Generate video chapter knowledge tree content:

[0122] After the above steps, the system finally merges the knowledge elements of all key frames into a knowledge tree with a single root node, which intuitively displays the knowledge structure of the entire video and facilitates users to quickly understand and navigate.

[0123] 8. Generate knowledge points with annotated timeline information at the same time:

[0124] The system not only provides a detailed description of the knowledge point, but also combines each knowledge point with the corresponding timestamp to form a timeline annotation. This allows users to easily jump to a specific segment in the video, improving the pertinence and efficiency of learning.

[0125] In order to implement the above overall solution, the following technical solution provides an architecture of an analysis system for video teaching content.

[0126] Video acquisition unit:

[0127] Upload local video: Users can upload video files with PPT screen recording through the interface, supporting multiple common video formats such as avi, mp4, flv, etc. This provides users with flexibility and allows them to easily upload video files of different formats to the system.

[0128] Data synchronization video acquisition: The system has set up a timed synchronization mechanism, which can automatically transfer qualified videos to the video acquisition unit. This mechanism ensures that the video resources in the system are always synchronized with the original data without manual intervention.

[0129] Access screening: Currently, there is no automated access screening mechanism for uploaded videos, so manual access screening is required. In order to speed up the access process, we can consider adding an automatic judgment mechanism for PPT screen recording videos to reduce manual operations.

[0130] Video length judgment: The system will automatically judge the video length. If the video is too short (less than 5 minutes), it will prompt that the video upload failed and will not be stored. This is to ensure that the videos stored in the system have certain content and quality.

[0131] Video storage unit:

[0132] The uploaded video resources will be automatically stored in the video storage unit, and users can intuitively view the results of analyzed and unanalyzed videos. At the same time, users can also add, delete, modify and query videos to meet different management needs.

[0133] Annotation status: The system will display the annotation status of the video, including AI analysis in progress and analysis completed. This helps users understand the progress of video processing and take appropriate actions in a timely manner.

[0134] Publishing status: The system will display the publishing status of the video, including published or unpublished. This can help users understand whether the video is ready for others to watch or use.

[0135] Update time: The system will record the latest editing time so that users can understand the latest updates of the video.

[0136] Editing function: Users can edit individual videos by clicking on the content annotation unit. This provides users with more control and allows them to modify and adjust the video as needed.

[0137] Preview results: The system provides a preview function, and users can preview the effect of the video after editing. This helps users confirm whether the editing results meet expectations and make necessary adjustments.

[0138] Content annotation unit:

[0139] (1) Catalog page recognition: Using OCR technology, the system automatically recognizes the catalog page from the video frame. Then, OCR recognition and keyword combing are performed on the text content in the catalog page to provide support for subsequent knowledge extraction and video indexing. If there are recognition errors or missed recognitions, users can manually intervene to mark and correct the content of the main catalog page.

[0140] (2) Title page identification:

[0141] The system further identifies images containing first-level, second-level, and third-level titles, automatically locates and frames the title area, and then recognizes and transcribes the title text to make it editable text.

[0142] On this basis, the system extracts key knowledge points based on the title content, which may involve summarizing and simplifying paragraph content.

[0143] If there is a vocabulary of professional knowledge points as input, the system can directly extract keywords by referring to the vocabulary; if there is no ready-made vocabulary, it is necessary to use large language models LLM such as GPT, Wenxin model, etc. to summarize and condense knowledge points after semantic understanding.

[0144] (3) Content page identification: For non-catalog page content, the system calculates its relevance to the catalog page to determine its position in the video knowledge structure. At the same time, the content page is deduplicated, and duplicate or very similar frames are removed, retaining only the most representative and relevant frame content.

[0145] All the pages identified above have a start time point and an end time point. When removing duplicate pages, the time periods in which multiple pages appear need to be merged.

[0146] Content Modification Unit:

[0147] Content modification is responsible for modifying and optimizing the identified and marked content.

[0148] (1) Add key frame images:

[0149] Users can add keyframe images by taking screenshots during video playback. This function allows users to supplement keyframes that may be missed by the system during automatic recognition, thereby ensuring the integrity of the video content.

[0150] Through the screenshot tool or interface, select specific frames in the video as new key frames and add them to the system.

[0151] (2) Delete key frame images:

[0152] The system removes duplicate keyframe images, but it is inevitable that there will be highly similar content. If the system generates duplicate or highly similar keyframe images during the automatic recognition process, users can use the delete function to remove these redundant content.

[0153] (3) Adjust the order of key frame images:

[0154] In some cases (such as the order in which the lecturer explains), the system may incorrectly sort the keyframe images. Users can use the function of adjusting the order to manually correct the order of the keyframes to ensure the coherence and logic of the video content.

[0155] Through the drag-and-drop interface or sequencing tools, users can easily change the order of keyframes to match the actual video content.

[0156] (4) Correct title content recognition errors:

[0157] During the title page recognition process, text recognition errors or inaccurate transcription may occur. The content correction unit provides a function to correct the title content, and users can manually edit and modify incorrect title text.

[0158] By providing text editing tools, users can correct title content directly in the system to ensure the accuracy and consistency of the title.

[0159] Result preview unit:

[0160] Preview structured video chapter structure: you can see the complete video chapter structure effect of a single video;

[0161] Preview video summary: Generate video summary content based on the chapter structure content;

[0162] Preview the knowledge point positioning effect: output a knowledge point content list with timestamps based on the previous key frames and chapter results with time information.

[0163] The result preview unit is described in detail below.

[0164] Preview structured video chapter structure: This feature allows users to intuitively see the complete video chapter structure effect of a single video. Chapter structure generally refers to the organization and layout of video content, including the division of various parts or chapters. By previewing the chapter structure, users can quickly understand the overall framework and logical flow of the video, which helps with subsequent in-depth analysis and processing of the video content. This structured preview is particularly suitable for long videos or videos with complex content, as it can help users grasp the main idea and key points of the video more quickly.

[0165] Preview video summary: Based on the chapter structure of the video, the system can automatically generate a video summary. This summary is a simplification and refinement of the main content of the video, aiming to help users quickly understand the core information of the video. By previewing the video summary, users can quickly determine whether the video meets their interests or needs without having to watch the entire video. In the era of information explosion, this summary preview function is of great significance for screening and filtering a large amount of video content.

[0166] Preview knowledge point positioning effect: This function uses the previous key frame and chapter structure analysis results to output a knowledge point content list with timestamps. By previewing this list, users can quickly locate key information or important knowledge points in the video, greatly improving browsing and learning efficiency. For education, training or research videos, this knowledge point positioning function is particularly important because it can help users get to the core content directly, saving time and energy.

[0167] Therefore, these preview functions in the "result preview unit" play an important role in video processing and analysis. They not only improve the efficiency of users browsing video information, but also provide users with a more convenient and efficient way to grasp and understand video content. These functions are undoubtedly of great significance in quickly browsing video information and grasping video content. In addition, combined with the article abstracts generated by large model technology, users' understanding and application of video content can be further enriched. For example, users can quickly judge the theme and value of the video based on the abstract, or quickly find the content part that they are interested in or need to be studied in depth based on the knowledge point positioning. These functions are of great practical value for users who need to quickly obtain information from a large number of videos.

[0168] The positioning of video timestamps provides a powerful technical support for the knowledge point cutting of video content. Below, we will elaborate on the role and significance of video timestamp positioning in knowledge point cutting.

[0169] 1. Accurate positioning and fast retrieval

[0170] Through timestamp positioning, we can accurately mark the start and end time of each key knowledge point in the video. This is extremely useful for both learners and educators. Learners can jump directly to the knowledge points they are interested in or need to review, while educators can quickly find and display the key points in the teaching.

[0171] 2. Personalized learning experience

[0172] Everyone's learning needs and pace are different. With timestamp positioning, learners can selectively watch and learn specific parts of the video according to their needs, thus achieving a more personalized learning experience.

[0173] 3. Improve learning efficiency

[0174] Without timestamp positioning, learners may need to spend a lot of time to find and locate key information in the video. With accurate timestamps, learners can quickly find and focus on the core content of the video, greatly improving learning efficiency.

[0175] 4. Rich teaching resources cutting and reorganization

[0176] Educators can use timestamp positioning technology to cut video content into multiple independent knowledge point segments. These segments can not only be used as separate teaching resources, but can also be recombined as needed to create more flexible and diverse teaching content.

[0177] 5. Promote the reuse and innovation of video content

[0178] Precise timestamp positioning makes it easier to reuse video content. Educators or content creators can easily extract specific parts of the video for secondary creation or integration to generate new teaching content or products.

[0179] 6. Improve interactivity and engagement

[0180] Through timestamp positioning, educators can embed interactive elements in online teaching platforms, such as questions, discussions, or quizzes, which can be associated with specific time points in the video. This interactivity can not only increase learner engagement, but also help educators better understand learners' mastery.

[0181] 7. Expand the application scenarios of video teaching

[0182] With the rise of online education, video timestamp positioning technology has made video teaching no longer limited to the traditional classroom teaching environment. Whether it is corporate training, distance education or self-study, accurate timestamp positioning can provide learners with a more flexible and efficient learning experience.

[0183] Therefore, video timestamp positioning plays an indispensable role in knowledge point segmentation. It not only improves the efficiency and flexibility of learning, but also provides educators and learners with richer and more diverse teaching resources and learning methods. With the continuous development of technology, we have reason to believe that video timestamp positioning will play a more important role in the future education field.

[0184] Save content result unit:

[0185] After the chapter results have been confirmed to be correct, you can save the results. After saving, the production process of a video is completed. After a single production is completed, you can enter the next video production process.

[0186] In addition, the embodiment of the present invention can also analyze the audio signal to solve the problem of inconsistent sound and picture. Through the deep integration analysis of audio and video content, the system can more accurately identify and mark the knowledge points and important information in the video, providing users with a more complete learning aid tool.

[0187] Such a design has the following beneficial effects, improving learning efficiency: By quickly locating and summarizing the key knowledge points of teaching videos, this patent can effectively improve learners' learning efficiency. It is also possible to improve learning transfer capabilities and personalized recommendations by constructing knowledge graphs of multiple videos, ultimately improving learning efficiency. Intelligent retrieval and browsing: Supporting intelligent retrieval of large-scale video data, users can quickly find or locate target video clips, saving browsing time. Course production tool optimization: Provides efficient editing tools for video content producers, reversely optimizes the production process based on end-user needs, and improves production efficiency. Automatic decomposition of knowledge points: Automatically decompose video content into various knowledge point fragments to facilitate subsequent personalized learning and content reorganization. Quick positioning and navigation: The generated knowledge tree and annotated timeline enable users to quickly navigate in the video and find the precise clips they need. Effectiveness of video resources: Quick summary and induction of training platform course videos to maximize resource utility.

[0188] Please refer to Figure 2 , Figure 2 An analysis device 110 for video teaching content provided in an embodiment of the present invention includes:

[0189] An acquisition module 1101 is used to acquire a teaching video to be analyzed;

[0190] The analysis module 1102 is used to extract frames from the teaching video to be analyzed to obtain multiple candidate teaching video frames, each of which is configured with a timestamp; determine the teaching catalog page, teaching title page and teaching content page from the multiple candidate teaching video frames respectively; and construct a teaching knowledge tree as the analysis result of the teaching video to be analyzed based on the teaching catalog page, teaching title page and teaching content page.

[0191] It should be noted that the implementation principle of the aforementioned analysis device 110 for video teaching content can refer to the implementation principle of the aforementioned analysis method for video teaching content, which will not be repeated here. It should be understood that the division of the various modules of the above device is only a division of logical functions. In actual implementation, they can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software calling through processing elements; they can also be all implemented in the form of hardware; some modules can also be implemented in the form of software calling through processing elements, and some modules can be implemented in the form of hardware. For example, the analysis device 110 for video teaching content can be a separately established processing element, or it can be integrated in a chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a processing element of the above device. The function of the above analysis device 110 for video teaching content. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each module above can be completed by an integrated logic circuit of hardware in a processor element or an instruction in the form of software.

[0192] For example, the above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application specific integrated circuits (ASIC), or one or more microprocessors (digital signal processors, DSP), or one or more field programmable gate arrays (FPGA), etc. For another example, when a module above is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0193] The embodiment of the present invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned analysis device 110 for video teaching content. Figure 3 As shown, Figure 3The computer device 100 provided in the embodiment of the present invention is a structural block diagram. The computer device 100 includes an analysis device 110 for video teaching content, a memory 111, a processor 112 and a communication unit 113.

[0194] To achieve data transmission or interaction, the memory 111, the processor 112 and the communication unit 113 are electrically connected to each other directly or indirectly. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The analysis device 110 for video teaching content includes at least one software function module that can be stored in the memory 111 in the form of software or firmware or solidified in the operating system (OS) of the computer device 100. The processor 112 is used to execute the analysis device 110 for video teaching content stored in the memory 111, such as the software function module and computer program included in the analysis device 110 for video teaching content.

[0195] An embodiment of the present invention provides a readable storage medium, which includes a computer program. When the computer program is running, the computer device where the readable storage medium is located is controlled to execute the aforementioned analysis method for video teaching content.

[0196] For illustrative purposes, the foregoing description is made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the present disclosure to the precise form disclosed. Numerous modifications and variations are possible in accordance with the above teachings. These embodiments are selected and described in order to best illustrate the principles of the present disclosure and its practical application, so that those skilled in the art can best utilize the present disclosure and utilize various embodiments with different modifications to suit the intended specific application.

Claims

1. A method for analyzing video teaching content, characterized in that: include: Get the teaching video to be analyzed; Performing frame extraction processing on the teaching video to be analyzed to obtain a plurality of candidate teaching video frames, wherein the candidate teaching video frames are configured with timestamps; Determining a teaching catalog page, a teaching title page, and a teaching content page from the plurality of candidate teaching video frames respectively; According to the teaching catalog page, the teaching title page and the teaching content page, a teaching knowledge tree is constructed as the analysis result of the teaching video to be analyzed; The step of constructing a teaching knowledge tree as the analysis result of the teaching video to be analyzed according to the teaching catalog page, the teaching title page and the teaching content page includes: Obtaining the first keyword of all the teaching catalog pages, and using the first keyword as the index basis of the teaching knowledge tree; Acquire a second keyword of each teaching title page, and determine a key knowledge point corresponding to each teaching title page according to the second keyword; Determine the association relationship between each of the teaching content pages and each of the teaching catalog pages and each of the teaching title pages, and determine the positional relationship between the teaching catalog pages, teaching title pages and teaching content pages in constructing the teaching knowledge tree; The teaching knowledge tree is constructed according to the index basis, the key knowledge points, the position relationship and the timestamp, and there is a binding relationship between the key knowledge points and the timestamp.

2. The method according to claim 1, characterized in that The extracting frame of the teaching video to be analyzed to obtain a plurality of candidate teaching video frames includes: Extract frames from the teaching video to be analyzed at preset time intervals to obtain multiple initial teaching video frames; The multiple initial teaching video frames are deduplicated to obtain the multiple candidate teaching video frames.

3. The method according to claim 2, characterized in that The performing deduplication processing on the multiple initial teaching video frames to obtain the multiple candidate teaching video frames includes: Calculating structural similarity indexes of two adjacent initial teaching video frames in the plurality of initial teaching video frames; If there is a structure similarity index greater than a preset similarity index threshold, the initial teaching video frame located in the latter order is deleted, and the structure similarity indexes of the initial teaching video frames adjacent to each other in the remaining initial teaching video frames are recalculated until the deduplication process is completed; If there is no structure similarity index greater than the preset similarity index threshold, the deduplication process is terminated; The initial teaching video frames after multiple duplications are removed are used as the multiple candidate teaching video frames.

4. The method according to claim 1, characterized in that: Determining the teaching catalog page from the plurality of candidate teaching video frames includes: Extracting text content of a target candidate teaching video frame, wherein the target candidate teaching video frame is any candidate teaching video frame among the multiple candidate teaching video frames; Analyze the text content to obtain keywords included in the text content; In the case where the keyword includes a preset catalog keyword, the content corresponding to the target candidate teaching video is used as the teaching catalog page.

5. The method according to claim 1, characterized in that Determining the teaching title page from the plurality of candidate teaching video frames comprises: Acquire text font information of a target candidate teaching video frame, wherein the target candidate teaching video frame is any candidate teaching video frame among the multiple candidate teaching video frames, and the text font information includes font size, font position and keyword information; If the title content is determined from the text font information according to the font size, font position and keyword information, the content corresponding to the target candidate teaching video is used as the teaching title page.

6. The method according to claim 1, characterized in that Determining the teaching content page from the plurality of candidate teaching video frames includes: The contents other than the teaching catalog page and the teaching title page are used as the teaching content page.

7. The method according to claim 1, characterized in that The method further comprises: In response to the video frame modification operation, performing at least one of adding, deleting, and adjusting the order of the plurality of candidate teaching video frames; In response to the text adjustment operation, modify the text content included in the teaching catalog page, the teaching title page and the teaching content page; In response to the knowledge tree saving operation, the teaching knowledge tree is saved and bound to the video identifier of the teaching video to be analyzed.

8. The method according to claim 1, characterized in that The step of obtaining the teaching video to be analyzed includes: Set up synchronization mechanism with the preset teaching video database; When it is detected that a new teaching video to be processed is stored in the preset teaching video database, it is determined whether the teaching video to be processed is a recorded demonstration video and whether the video duration exceeds a preset minimum duration; If yes, the teaching video to be processed is used as the teaching video to be analyzed; If not, the to-be-processed teaching video is deleted from the preset teaching video database.

9. An analysis device for video teaching content, characterized in that: include: An acquisition module is used to acquire the teaching video to be analyzed; An analysis module, used for performing frame extraction processing on the teaching video to be analyzed to obtain a plurality of candidate teaching video frames, wherein the candidate teaching video frames are configured with timestamps; Determining a teaching catalog page, a teaching title page, and a teaching content page from the plurality of candidate teaching video frames respectively; According to the teaching catalog page, the teaching title page and the teaching content page, a teaching knowledge tree is constructed as the analysis result of the teaching video to be analyzed; The analysis module is specifically used for: Obtaining the first keyword of all the teaching catalog pages, and using the first keyword as the index basis of the teaching knowledge tree; obtaining the second keyword of each teaching title page, and determining the key knowledge point corresponding to each teaching title page according to the second keyword; Determine the association relationship between each of the teaching content pages and each of the teaching catalog pages and each of the teaching title pages, and determine the positional relationship between the teaching catalog pages, teaching title pages and teaching content pages in constructing the teaching knowledge tree; The teaching knowledge tree is constructed according to the index basis, the key knowledge points, the position relationship and the timestamp, and there is a binding relationship between the key knowledge points and the timestamp.

Citation Information

Patent Citations

  • Teaching video knowledge graph construction method and device, equipment and storage medium

    CN117973526A