Large model-based video processing method and apparatus, device, and storage medium
By automatically generating video handouts using large-scale model technology, the quality and efficiency issues of manually generated handouts have been resolved, resulting in high-quality, personalized handout documents that improve student learning experience and teaching efficiency.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING BAIDU NETCOM SCI & TECH CO LTD
- Filing Date
- 2025-09-16
- Publication Date
- 2026-07-30
AI Technical Summary
Current video lecture generation relies on manual methods, resulting in inconsistent quality, low efficiency, and difficulty in personalization, especially when dealing with complex knowledge points and logical relationships, where accuracy is insufficient.
Using large-scale modeling technology, lecture notes are generated by recognizing screenshots and subtitles in videos, and exercises are inserted into the lecture notes to achieve automated generation of lecture notes documents. Combined with graphic presentation, this improves learning interest and efficiency.
The generated lecture notes are of high quality, standardized, and engaging, reducing the burden on teachers and improving students' learning experience and teaching efficiency.
Smart Images

Figure CN2025121703_30072026_PF_FP_ABST
Abstract
Description
Video processing methods, devices, equipment, and storage media based on large models
[0001] This disclosure claims priority to Chinese Patent Application No. 202510125424.9, filed on January 26, 2025, entitled “Video Processing Method, Apparatus, Device and Storage Medium Based on Large Model”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure relates to deep learning in artificial intelligence, and more particularly to a video processing method, apparatus, device, and storage medium based on a large model. Background Technology
[0003] Currently, the lecture notes for videos watched by users during their learning process mainly rely on manual output. That is, lecture note editors need to edit the notes based on the video content after watching it. Obviously, this method of editing lecture notes is easily limited by individual abilities, resulting in inconsistent quality and efficiency. Summary of the Invention
[0004] This disclosure provides a video processing method, apparatus, device, and storage medium based on a large model.
[0005] According to a first aspect of this disclosure, a video processing method based on a large model is provided, comprising:
[0006] The video to be processed is identified to obtain exercise information and knowledge information; wherein, the exercise information includes the text of the exercises in the courseware screenshot of the video to be processed; the knowledge information includes the knowledge text in the courseware screenshot of the video to be processed and the subtitle text corresponding to the audio of the video to be processed; the courseware screenshot is a screenshot of the video to be processed;
[0007] The knowledge information is input into a preset lecture handout model to generate the lecture handout content text, and the lecture handout content text is divided to obtain at least one content block.
[0008] Insert the exercises from the exercise information into at least one of the content blocks to obtain the presentation file of the lecture notes for the video to be processed.
[0009] According to a second aspect of this disclosure, a video processing apparatus based on a large model is provided, comprising:
[0010] An extraction unit is used to identify the video to be processed and obtain exercise information and knowledge information; wherein, the exercise information includes the text of the exercises in the courseware screenshot of the video to be processed; the knowledge information includes the knowledge text in the courseware screenshot of the video to be processed and the subtitle text corresponding to the audio of the video to be processed; the courseware screenshot is a screenshot of the video to be processed;
[0011] The first generation unit is used to input the knowledge information into a preset lecture handout model, generate the lecture handout content text, and divide the lecture handout content text to obtain at least one content block.
[0012] The second generation unit is used to insert the exercises from the exercise information into at least one of the content blocks to obtain the handout file to be displayed for the video to be processed.
[0013] According to a third aspect of this disclosure, an electronic device is provided, comprising:
[0014] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method of the first aspect.
[0015] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method described in the first aspect.
[0016] According to a fifth aspect of this disclosure, a computer program product is provided, the computer program product comprising: a computer program stored in a readable storage medium, wherein at least one processor of an electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program to cause the electronic device to perform the method described in the first aspect.
[0017] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0018] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0019] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure;
[0020] Figure 2 is a schematic diagram according to a second embodiment of the present disclosure;
[0021] Figure 3 is a schematic diagram according to a third embodiment of the present disclosure;
[0022] Figure 4 is a schematic diagram according to the fourth embodiment of the present disclosure;
[0023] Figure 5 is a schematic diagram according to the fifth embodiment of the present disclosure;
[0024] Figure 6 is a schematic diagram according to a sixth embodiment of the present disclosure;
[0025] Figure 7 shows a schematic block diagram of an electronic device for implementing embodiments of the present disclosure. Detailed Implementation
[0026] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0027] With the development of the internet, more and more students prefer to learn new knowledge or consolidate existing skills through video courses. Currently, the generation of lecture notes for video courses mainly relies on manual methods. Teachers or teaching assistants often need to manually watch the video content, extract key knowledge points, and take relevant screenshots from the video. Then, they combine this with subtitles and video content to create detailed lecture notes. However, this method of creating lecture notes is easily limited by individual understanding and expression abilities, making it difficult to guarantee standardization and high quality, resulting in inconsistent quality. Furthermore, when the video content is long and contains many knowledge points, manually generating lecture notes requires a significant amount of time and effort, leading to inefficiency.
[0028] To address the shortcomings of manually generated handouts, several rule-based or template-based automated handout generation methods have been proposed. However, these methods often only handle simple text and image information, struggling to accurately understand and analyze complex knowledge points and logical relationships within videos, and failing to accurately extract and interpret key information. Consequently, the generated handouts suffer from deficiencies in accuracy and completeness. Furthermore, this technology struggles to generate personalized handouts based on students' learning progress and comprehension abilities, thus failing to meet the needs of diverse learners.
[0029] To overcome the aforementioned limitations, this disclosure innovatively introduces advanced large-scale model technology. Based on the powerful data processing capabilities and deep learning algorithms of large-scale model technology, it can understand and analyze massive amounts of video content and related text information, automatically parse knowledge points in video courses, identify key information, and deeply understand their inherent logical relationships and hierarchical structure, thereby achieving the automatic generation of video lecture notes. Specifically, the large-scale model-based video processing method of this disclosure utilizes existing large language model capabilities, through prompt words and some pre-processing and post-processing rules, to achieve the generation of learning lecture notes.
[0030] Furthermore, this publication not only generates lecture notes with detailed explanations of each knowledge point, but also cleverly incorporates corresponding video screenshots and combines knowledge points with exercises. The screenshots visually demonstrate how knowledge points are presented in the videos, helping students better understand and memorize the course content. This combination of text and images also makes the lecture notes more engaging and interesting, increasing student interest and participation. The arrangement of exercises after each knowledge point helps users consolidate and identify gaps in their knowledge after completing the learning.
[0031] From a practical application perspective, the lecture notes provided by this publication not only effectively assist students in reviewing and consolidating their learning after class, but also greatly reduce the workload of teachers in creating lecture notes. Teachers can then devote more time and energy to optimizing teaching content and innovating teaching methods, thereby improving overall teaching efficiency and quality.
[0032] The implementation of this disclosure can be accomplished through three functional modules. These three functional modules can be a preprocessing module, a content generation module, and a post-processing module.
[0033] The preprocessing module processes the images and audio of the video to be processed, obtaining courseware screenshots and subtitle text. The courseware screenshots can be deduplicated. Each courseware screenshot corresponds to a specific time segment displayed in the video. The preprocessing module can also map the courseware screenshots and subtitle text, and using a large model, fuse the knowledge text identified in the courseware screenshots and the subtitle text to obtain knowledge information.
[0034] The content generation module utilizes a large model to categorize the text information in courseware screenshots into knowledge and questions, obtaining both exercise and knowledge information. This knowledge information is then fused with the subtitle text to arrive at the final knowledge information. Furthermore, the large model can extract exercises from the exercise information. Additionally, based on the knowledge information, the model can generate the lecture content text and divide it into multiple content blocks. Then, based on chronological order, the exercise information can be fused with these content blocks to obtain the final lecture notes.
[0035] The post-processing module can use rules to correct the lecture notes generated in the previous step, thereby ensuring that the lecture notes document meets the requirements of the Markdown format, and call pandoc and xelatex to render and generate Word and PDF files, and finally output lecture notes that can be sent to users.
[0036] This disclosed method for generating lecture notes can efficiently and automatically generate high-quality lecture note documents for videos, completely revolutionizing the traditional, cumbersome process that relies on manual operation by teachers. Furthermore, based on advanced large-scale modeling technology, this method can accurately analyze the knowledge points in video courses, ensuring not only standardized and high-quality documents but also a deep understanding of the knowledge points and a clear presentation of their logical relationships. In addition, the lecture notes generated by this method incorporate video screenshots and questions, making the lecture note documents more intuitive and engaging, greatly enhancing the student learning experience and effectiveness. This disclosed lecture note document not only effectively assists students in their after-class review but also significantly reduces the workload of teachers, providing users with superior educational products and services.
[0037] This disclosure provides a video processing method, apparatus, device, and storage medium based on a large model, which is applied to deep learning in the field of artificial intelligence to improve the quality and efficiency of lecture notes generation.
[0038] It should be noted that the head model in this embodiment is not a head model specific to any particular user and does not reflect the personal information of any particular user. It should also be noted that the two-dimensional face image in this embodiment comes from a publicly available dataset.
[0039] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0040] To enable readers to have a deeper understanding of the implementation principle of this disclosure, the embodiment shown in Figure 1 will be further detailed in conjunction with Figures 2-4 below.
[0041] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure. As shown in Figure 1, the present disclosure provides a video processing method based on a large model, the method comprising:
[0042] 101. Recognize the video to be processed to obtain exercise information and knowledge information. The exercise information includes the text of the exercises in the courseware screenshot of the video to be processed. The knowledge information includes the knowledge text in the courseware screenshot of the video to be processed and the subtitle text corresponding to the audio of the video to be processed. The courseware screenshot is a screenshot of the video to be processed.
[0043] In this embodiment, a video to be processed is first obtained. Optionally, the video to be processed can be used to assist users in learning knowledge. Optionally, the playback content of the video to be processed can include courseware. Optionally, the video can also include an image of a teacher. Optionally, the audio of the video to be processed is the teacher explaining the courseware. Optionally, the courseware can be a document such as PPT or PDF.
[0044] After identifying the video to be processed, exercise information and knowledge information can be extracted from it. The exercise information can be exercise text obtained by recognizing screenshots of the courseware played by the video. The knowledge information can include knowledge text obtained by recognizing screenshots of the courseware played by the video, and subtitle text obtained by processing the audio of the video. Optionally, the knowledge information can be obtained by fusing the knowledge text and the subtitle text.
[0045] Optionally, after determining the video to be processed, screenshots can be taken from the video at a preset frame rate, and adjacent similar screenshots can be merged to obtain the final courseware screenshot. The text in the courseware screenshot can be recognized to obtain the exercise text and knowledge text. The exercise text refers to the exercises appearing in the courseware screenshot, and the knowledge text refers to all text in the courseware screenshot other than the exercise text. Optionally, the classification of the exercise text and the knowledge text can be obtained by classifying the text in the courseware screenshot using a preset classification model after the text has been recognized.
[0046] Optionally, a video database may also be included. This video database can contain multiple videos. The video to be processed can be retrieved by inputting its video ID and other identifying information.
[0047] 102. Input the knowledge information into the preset lecture handout model to generate the lecture handout content text, and divide the lecture handout content text to obtain at least one content block.
[0048] In this embodiment, after obtaining the knowledge information from the video to be processed according to step 101, the knowledge information can be input into a preset lecture handout model for processing to generate the lecture handout content text. Optionally, the processing may include extracting titles from the knowledge information and hierarchically dividing the extracted titles.
[0049] After obtaining the lecture notes, the notes can be divided into multiple content blocks. Each content block can include a portion of continuous text. Optionally, this content block division process can be implemented using a large model. Optionally, each content block can correspond to a single knowledge point.
[0050] Optionally, to preserve the validity of the courseware screenshots, the large model can process the knowledge information corresponding to each screenshot to obtain the lecture content text for each screenshot. Based on the playback order of all courseware screenshots in the video to be processed, the lecture content text of these screenshots can be combined to obtain the lecture content text for the entire video to be processed. Optionally, each content block may include the content text corresponding to at least one courseware screenshot.
[0051] 103. Insert the exercises from the exercise information into at least one content block to obtain the presentation file of the lecture notes for the video to be processed.
[0052] In this embodiment, exercises can be extracted from the exercise information obtained in step 101. These exercises can then be inserted into the content block obtained in step 102. This content block containing the inserted exercises can form a presentation file of the lecture notes for the video to be processed. Optionally, this presentation file is a lecture note document that can be output to the user for reading.
[0053] Optionally, exercise information can be inserted into the matching content block to help users consolidate their knowledge of the content block through the exercises within the exercise information. Optionally, the exercise can be inserted at the end of the content block to achieve a knowledge point-exercise handout structure, better assisting users in learning knowledge and consolidating it through exercises.
[0054] Optionally, the matching between the exercise and the content block can be achieved based on their relevance. Alternatively, the matching can be achieved based on the time corresponding to the content block in the video to be processed, and the time corresponding to the exercise in the video to be processed. Optionally, the time corresponding to the content block in the video to be processed can be determined based on the courseware screenshots included in the content block. Optionally, the time corresponding to the exercise in the video to be processed can be determined based on the courseware screenshots containing the exercise.
[0055] In this embodiment, the exercise information and knowledge information are identified from the video to be processed; the knowledge information is input into a preset lecture handout model for processing to generate lecture handout content text; the lecture handout content text is divided to obtain at least one content block; the exercises in the exercise information are inserted into at least one content block to obtain the final lecture handout with display file. This method realizes the automatic generation of lecture handouts from the video to be processed and improves the effectiveness of lecture handout instructions. Furthermore, by additionally inserting the exercises in the exercise information into the lecture handout content, this disclosure achieves more accurate placement of the exercises in the exercise information, which helps users consolidate their knowledge points after learning the knowledge points in the content block through the exercises in the exercise information, thereby improving learning effectiveness and enhancing the effectiveness and practicality of the lecture handouts.
[0056] Based on the above embodiments, the specific process of extracting exercise information and knowledge information from the video to be processed may include:
[0057] 1011. Using a text recognition model, identify the courseware screenshots from the video to be processed to obtain exercise information and knowledge text.
[0058] In this embodiment, after acquiring the video to be processed, screenshots of courseware can be obtained from the video. Optionally, these screenshots can be obtained by taking screenshots of the video to be processed at a preset frequency and then removing duplicates. Optionally, each screenshot can correspond to a time information, including the appearance time and end time of the screenshot in the video to be processed.
[0059] Optionally, deduplication of screenshots can be achieved by calculating the similarity of the text within the screenshots. Specifically, after recognizing the text information in two adjacent screenshots, the similarity of the text information in the two screenshots can be calculated. If the similarity is less than a similarity threshold, the two screenshots are determined to be dissimilar and belong to two different courseware screenshots. If the similarity is greater than or equal to the similarity threshold, the two screenshots are determined to be similar and belong to the same courseware screenshot. Optionally, after identifying multiple consecutive and similar screenshots, one screenshot can be selected as the courseware screenshot. Optionally, the first or last screenshot among the multiple consecutive and similar screenshots can be used as the courseware screenshot.
[0060] After obtaining a screenshot of the courseware, the text information in the screenshot can be obtained through text recognition. Optionally, this text recognition can be achieved using a preset text recognition model. Optionally, after obtaining the text information, it can be classified to obtain the exercise text and knowledge text. The exercise text is the exercise information.
[0061] Optionally, the screenshot of the courseware may have third-party time information. Optionally, the third-party time information may be used to indicate the time when the screenshot of the courseware appears in the video to be processed. Alternatively, the third-party time information may be used to indicate the time when the screenshot of the courseware appears and the time when it ends in the video to be processed.
[0062] 1012. Use an audio recognition model to recognize the audio of the video to be processed and obtain the corresponding subtitle text.
[0063] In this embodiment, after acquiring the video to be processed, the audio of the video can also be obtained. Then, an audio recognition model can be used to recognize the audio and obtain subtitle text. Optionally, the smallest unit in the subtitle text can be a sentence.
[0064] Optionally, the subtitle text may have fourth time information. Optionally, each sentence in the subtitle text may correspond to a fourth time information. Optionally, the fourth time information is the timestamp of that sentence. Optionally, the fourth time information may be used to indicate the time when that sentence in the subtitle text appears in the audio. Alternatively, the fourth time information may be used to determine the time when that sentence in the subtitle text appears and ends in the audio.
[0065] 1013. Merge the knowledge text and subtitle text to obtain knowledge information.
[0066] In this embodiment, after obtaining the knowledge text of the courseware screenshot and the subtitle text of the audio according to step 1011, the courseware screenshot and the subtitle text can be aligned based on the third time information of the courseware screenshot and the fourth time information of the subtitle text. Optionally, after alignment, the subtitle text corresponding to the courseware screenshot can be determined. Optionally, the subtitle text corresponding to the courseware screenshot can be the content spoken by the teacher during the playback of the courseware screenshot in the video to be processed. Optionally, the subtitle text corresponding to the courseware screenshot can be the teacher's explanation content relative to the courseware screenshot. Therefore, based on the knowledge text in the courseware screenshot and the subtitle text corresponding to the courseware screenshot, the knowledge information corresponding to the courseware screenshot can be fused to obtain the knowledge information of the courseware screenshot. Optionally, the knowledge information of all the courseware screenshots can form the knowledge information of the video to be processed. Optionally, each piece of knowledge information can correspond to one courseware screenshot. Optionally, the time information corresponding to the knowledge information can be the third time information of the courseware screenshot.
[0067] Optionally, the fusion can be achieved using a pre-defined fusion model. Optionally, the fusion of the knowledge text and the subtitle text can be achieved by regenerating some text based on the knowledge text and the subtitle text, which is the knowledge information.
[0068] In this embodiment, by identifying the exercise information and knowledge information in the video to be processed, the automatic extraction of the exercise information and knowledge information is realized, which improves the accuracy of the extraction of the exercise information and knowledge information and provides a foundation for subsequent processing.
[0069] In one example, the specific process of fusing knowledge text and caption text to obtain knowledge information may include:
[0070] 10121. Based on the third time information of the knowledge text, determine the subtitle text corresponding to the fourth time information that matches the third time information.
[0071] In this embodiment, the third time information corresponding to the courseware screenshot where the knowledge text is located can be determined first. Optionally, the third time information may include a start time and an end time. The start time is the time when the courseware screenshot begins to appear in the video to be processed, and the end time is the time when the video switches to the next courseware screenshot after the courseware screenshot appears at that location.
[0072] Since the fourth time information in the subtitle text is the timestamp corresponding to the sentence within the subtitle text, and considering that a teacher typically needs to elaborate on the content of a screenshot of courseware using many sentences, the subtitle text corresponding to the fourth time information that matches the third time information can be determined by checking whether this fourth time information falls within the range of the third time information. That is, if the fourth time information indicates that the start time of the sentence in the subtitle text is within the time period indicated by the third time information, then the fourth time information matches the third time information. In other words, that sentence in the subtitle text matches the knowledge text.
[0073] Optionally, the final subtitle text corresponding to the knowledge text in a screenshot of a courseware can be multiple consecutive sentences in that subtitle text.
[0074] 10122. Merge the knowledge text and its matching subtitle text to obtain knowledge information.
[0075] In this embodiment, after determining the knowledge text and its corresponding subtitle text, the knowledge text and the subtitle text are fused. Optionally, the fusion of the knowledge text and the subtitle text can be achieved by inputting them into a preset fusion model. Optionally, the fusion yields knowledge information. This knowledge information may include multiple consecutive and coherent sentences. Each screenshot of the courseware may correspond to a segment of knowledge information. This knowledge information may correspond to third-time information.
[0076] In this embodiment, by fusing knowledge text and subtitle text, the text on the courseware and the audio content of the teacher's explanation are integrated, which improves the effectiveness and completeness of the integrated knowledge information and provides a foundation for subsequent processing.
[0077] Figure 2 is a schematic diagram according to a second embodiment of the present disclosure. Based on the embodiment shown in Figure 1, as shown in Figure 2, the present disclosure provides a video processing method based on a large model. The specific process of inserting exercises from exercise information into content blocks to obtain the lecture notes to be displayed may include:
[0078] S201. Identify each exercise in the exercise information to obtain at least one exercise; and match the content block corresponding to the obtained exercise from at least one content block.
[0079] In this embodiment, the exercise information may include the exercise portions from all the courseware screenshots of the video to be processed. Optionally, multiple exercises in the exercise information are not extracted individually. Therefore, after obtaining the exercise information, it can be further processed to extract the exercises from it.
[0080] Furthermore, each exercise can be matched with a content block to determine the content block that matches the exercise. Optionally, the content block that matches the exercise can be determined based on the content similarity between the exercise and the content block. Alternatively, the content block corresponding to the exercise can be determined based on a screenshot of the courseware where the exercise is located and a screenshot of the courseware corresponding to the content block.
[0081] S202. Insert the obtained exercises into the content block corresponding to the obtained exercises to obtain the file to be displayed.
[0082] In this embodiment, after determining the correspondence between exercises and content blocks, the exercises can be inserted into the content block. Optionally, the exercises can be inserted at the end of the content block. Optionally, one content block can correspond to one knowledge point. Optionally, the way the exercises are inserted can form a knowledge point-exercise sequence, so that users can consolidate their knowledge points through exercises after completing their learning, thereby improving their learning efficiency.
[0083] In one example, the exercise has first time information. This first time information indicates the time when the courseware screenshot containing the exercise appears on the video to be processed. Optionally, the first time information may include a start time and an end time. The start time is the time when the courseware screenshot appears. The end time is the time when the courseware screenshot switches to the next courseware screenshot after its appearance. The content block has second time information. Optionally, each content block may correspond to at least one courseware screenshot. The second time information of the content block can be determined based on the third time information of the courseware screenshot corresponding to the content block.
[0084] Optionally, the exercise and the content block can be matched based on the first time information of the exercise and the second time information of the content block. Optionally, if the first time information of the exercise and the second time information of the content block have an intersection, then the exercise and the content block are determined to be matched.
[0085] In this embodiment, the insertion of exercises is achieved by matching the exercises and the content block, ensuring the layout of knowledge points and exercises, and improving the user experience when using the lecture notes.
[0086] In one example, the specific way to extract the exercise from the exercise information may include:
[0087] First, based on the exercise number in the exercise information, the exercises in the exercise information are split into at least one exercise.
[0088] In this embodiment, since exercises are typically assigned a question number for easy identification during explanation or answer provision, the exercise text within the exercise information can be split based on this question number to obtain at least one exercise.
[0089] Secondly, the format of the exercises in the exercise information is recognized to obtain at least one exercise.
[0090] In this embodiment, exercises typically include multiple-choice questions, fill-in-the-blank questions, and problem-solving questions. The question types are usually relatively fixed for different knowledge areas. Therefore, based on the question type, the format of the exercise can be determined. Consequently, based on this format, the exercise can be segmented to obtain at least one exercise. For example, for multiple-choice questions, the stem and options can be extracted to obtain one exercise.
[0091] In this embodiment, the extraction of exercises from the exercise information is achieved through the two methods described above, thereby improving the efficiency of exercise acquisition.
[0092] Based on the above embodiments, the process of determining the matching relationship between exercises and content blocks may further include the following steps:
[0093] 2011. Obtain the timestamp of the subtitle text corresponding to the content block, and determine the time information corresponding to the content block based on the timestamp.
[0094] In this embodiment, the knowledge information for generating the content block can be determined based on the content block. Then, the subtitle text contained within this knowledge information can be obtained. This subtitle text may include a timestamp. Based on the timestamps of all subtitle texts corresponding to a content block, the time information corresponding to that content block can be generated. Since there is no direct correspondence between the segmentation of a content block and a screenshot of the courseware, multiple content blocks may appear in a single screenshot. Therefore, using timestamps to determine the time information of the content block is more accurate.
[0095] 2012. Based on the screenshot of the courseware containing the exercise information and the corresponding time information in the screenshot, determine the time information corresponding to the exercise information. The courseware screenshot is a screenshot of the video to be processed. The time information of the courseware screenshot is determined based on the time when the screenshot appears in the video to be processed.
[0096] In this embodiment, since the exercise information is directly extracted from screenshots of the courseware, the corresponding courseware screenshot can be determined based on the exercise information. Each courseware screenshot can correspond to a time information. This time information indicates the time when the courseware screenshot appears in the video to be processed. Therefore, the time information of the courseware screenshot can be used as the time information corresponding to the exercise information.
[0097] 2013. Based on the time information of the content blocks and the time information of the exercises, determine the matching relationship between the exercise information and the content blocks.
[0098] In this embodiment, the matching relationship between content blocks and exercise information can be determined by matching the time information of the content blocks and the time information of the exercise information. Optionally, by comparing and analyzing this time information, the time overlap or proximity relationship between the exercise information and the content blocks can be identified, thereby determining the content blocks and exercise information that have time overlap or proximity relationship as matched content blocks and exercise information.
[0099] Optionally, a match between a content block and exercise information can be determined when their time information overlaps. Specifically, based on the time information of each content block, it is determined whether there is any exercise information that overlaps with the time information of that content block. If such exercise information exists, it is determined that the content block and the exercise information overlap. Therefore, a matching relationship can be established between the overlapping content blocks and the exercise information. Optionally, a content block can be matched with one or more exercise information. Optionally, if the time information of the content block overlaps with the time information of the exercise information, it indicates that the exercise information was explained simultaneously during the explanation of the content block in the video to be processed.
[0100] Optionally, if there is a problem information that does not intersect with any content block, then the content block whose time information is closest to that of the problem information can be determined as the content block that matches the problem information.
[0101] In this embodiment, the time information corresponding to a content block can be generated based on the timestamps of all subtitle texts corresponding to that content block; the time information of the courseware screenshot can be used as the time information corresponding to the exercise information; the matching relationship between the content block and the exercise information can be determined by matching the time information of the content block and the time information of the exercise information, thereby realizing the automatic association between the content block and the exercise information. This enables learners to quickly locate relevant questions and content blocks while watching videos, enhancing the efficiency and effectiveness of learning, and improving the relevance and navigation of the content blocks and exercise information in the video to be processed.
[0102] Figure 3 is a schematic diagram according to a third embodiment of the present disclosure. Based on the embodiments shown in Figures 1 and 2, as shown in Figure 3, the present disclosure provides a video processing method based on a large model. The specific process of processing knowledge information to generate lecture content and dividing the lecture content into multiple content blocks may include:
[0103] 301. Input the knowledge information into the preset lecture model according to the playback order of the videos to be processed, and extract multiple knowledge topics. Among them, the knowledge topic represents the title obtained from the extraction of some knowledge information.
[0104] In this embodiment, each courseware screenshot can correspond to the fused knowledge information. Therefore, the knowledge information of each courseware screenshot can be input into a preset lecture note model according to the playback order of the video to be processed, resulting in multiple knowledge topics output by the preset lecture note model. Specifically, the preset lecture note model can be a large model. Optionally, the preset lecture note model can segment the knowledge information to separate the knowledge content corresponding to different knowledge topics. Simultaneously, for the segmented knowledge content, the preset lecture note model can extract the knowledge topic corresponding to each piece of knowledge content. Optionally, the knowledge topic can be the title corresponding to that piece of knowledge content.
[0105] 302. Input the knowledge topics into the preset lecture model according to the playback order of the videos to be processed, and generate a knowledge outline. The knowledge outline represents the hierarchical relationship of the knowledge topics.
[0106] In this embodiment, after obtaining the knowledge topics through step 301, these knowledge topics can be input again into the preset lecture handout model according to the playback order of the video to be processed. The preset lecture handout model can generate a knowledge outline based on the knowledge topics. The knowledge outline can include the hierarchical relationship of each knowledge topic. The hierarchical relationship can include the level corresponding to each knowledge topic, as well as the correspondence between multiple knowledge topics.
[0107] 303. Input the knowledge information corresponding to each knowledge topic into the preset lecture handout model to generate the explanation content for each knowledge topic.
[0108] In this embodiment, after determining the knowledge topics, the knowledge information corresponding to each knowledge topic can be input into a preset lecture note model. This preset lecture note model will automatically generate explanation content for each knowledge topic based on the input knowledge information. This explanation content aims to clearly and systematically present the core concepts and details of the knowledge topic, ensuring that learners can understand each topic more quickly.
[0109] 304. Based on the knowledge topic, combine the lecture content with the knowledge outline to obtain the lecture notes.
[0110] In this embodiment, after obtaining the knowledge outline and the corresponding explanation content for each knowledge topic according to the above steps, the explanation content for each knowledge topic can be inserted into the corresponding position in the knowledge outline. This process of inserting the explanation content is the process of refining the lecture notes based on the knowledge outline, ultimately resulting in complete lecture notes. This ensures that the final lecture notes provide a comprehensive and well-organized learning text, helping learners systematically understand and master each knowledge topic.
[0111] 305. Based on the knowledge outline of the lecture notes, input the highest-level advanced knowledge topics and their lecture notes into the preset segmentation model to obtain the fused content of each advanced knowledge topic.
[0112] In this embodiment, multiple knowledge topics at the highest level of the knowledge outline can be obtained, and these topics can be named advanced knowledge topics. Then, the lecture notes corresponding to each advanced knowledge topic can be obtained from the lecture notes obtained in step 304. Optionally, the lecture notes corresponding to each advanced knowledge topic may include the explanation content of the advanced knowledge topic, its subordinate knowledge topics, and their explanation content. Furthermore, each advanced knowledge topic and its corresponding lecture notes can be input into a preset segmentation model. This preset segmentation model can merge the lecture notes of each advanced knowledge topic to obtain the merged content of each advanced knowledge topic.
[0113] 306. Input the integrated content of advanced knowledge topics into the preset segmentation model in the order of the knowledge outline to obtain multiple content blocks.
[0114] In this embodiment, the advanced knowledge topics and their integrated content can be sorted according to the knowledge outline. Then, the integrated content of all advanced knowledge topics can be input into a preset segmentation model. This preset segmentation model can merge the input integrated content to obtain multiple content blocks. Each content block can include at least one advanced knowledge topic. The preset segmentation model can segment based on the consistency of the input information, grouping at least one advanced knowledge topic with high content relevance into a single content block. Specifically, the preset segmentation model can calculate the relevance of the lecture content of adjacent advanced knowledge topics. If the calculated relevance is greater than a relevance threshold, it can be determined that the two advanced knowledge topics belong to the same content block. Otherwise, it can be determined that the two advanced knowledge topics belong to two separate content blocks.
[0115] In this embodiment, by using a preset large-scale lecture handout model, the automatic generation of lecture handout content from knowledge information is achieved, improving the efficiency of lecture handout generation. Furthermore, during the lecture handout content generation process, optimization of knowledge topics and lecture handouts further improves the effectiveness and quality of the lecture handout content. Moreover, by using a preset segmentation model, the lecture handout content is divided into multiple content blocks, providing a good foundation for adding subsequent exercise information.
[0116] In one example, after obtaining the multiple knowledge topics, the knowledge topics can be optimized using a preset knowledge topic optimization strategy. Optionally, when the optimization is implemented through a topic optimization model, the optimized knowledge topic can be obtained directly by inputting the knowledge topic into the topic optimization model. The optimized knowledge topic can be a knowledge topic obtained by splitting and merging the original knowledge topic. Alternatively, when the optimization is implemented through prediction rules, therefore, based on the above step 301, it may also include:
[0117] 3011. If the knowledge topic does not match the preset video topic of the video to be processed, delete the knowledge topic.
[0118] In this embodiment, a preset topic optimization model can be used to judge the input knowledge topic and determine whether it matches the preset video topic of the video to be processed. That is, whether the knowledge topic is included in the scope of the video topic. If the knowledge topic matches the video topic, no processing is required. If the knowledge topic does not match the video topic, the topic can be deleted.
[0119] 3012. If there are two knowledge topics and their corresponding knowledge information with a content matching degree greater than the first threshold, then merge the knowledge topics and the knowledge information corresponding to the knowledge topics.
[0120] In this embodiment, knowledge information from two adjacent knowledge topics can be matched. Optionally, the matching degree of the knowledge information from the two adjacent knowledge topics can be calculated. Optionally, the matching degree can be calculated using a preset text matching algorithm. If the matching degree is greater than a first threshold, it indicates that the two adjacent knowledge topics describe the same content, and the knowledge topics and their corresponding knowledge information can be merged.
[0121] Optionally, the knowledge information of any two non-adjacent knowledge topics can also be matched. If the matching degree of the knowledge information of the two knowledge topics is greater than a first threshold, the knowledge information of the two knowledge topics can be deduplicated to obtain deduplicated knowledge information. Based on the deduplicated knowledge information, the knowledge topic can be regenerated.
[0122] In this embodiment, the accuracy and effectiveness of the knowledge topic are improved by further verifying the knowledge topic through these two methods.
[0123] In one example, after obtaining the knowledge outline, it can be optimized using a preset knowledge outline optimization strategy. Optionally, when the optimization is implemented through a large-scale outline optimization model, the optimized knowledge outline can be obtained directly by inputting the knowledge outline into the model. The optimized knowledge outline can be a knowledge outline obtained by adjusting the hierarchy of knowledge topics in the original knowledge outline. Alternatively, when the optimization is implemented through prediction rules, in addition to step 302 above, it can also include:
[0124] 3021. Based on the hierarchical relationship of knowledge topics in the knowledge outline, calculate the hierarchical parameters between the explanation content of the upper-level knowledge topic and the explanation content of the lower-level knowledge topic.
[0125] In this embodiment, the hierarchical relationship of each knowledge topic can be determined from the knowledge outline. Then, the upper-level knowledge topic and its corresponding lower-level knowledge topic can be obtained from the knowledge outline. The upper-level and lower-level knowledge topics are knowledge topics with a hierarchical relationship. The upper-level and lower-level knowledge topics can be non-adjacent. For example, when the level of the upper-level knowledge topic is 1, the levels of the lower-level knowledge topics can be 1.1, 1.2, 1.3, etc. The level 1.1 can also include knowledge topics at levels such as 1.1.1, 1.1.2, etc., existing between 1.1 and 1.2. After determining the upper-level and lower-level knowledge topics, the hierarchical parameter between the two explanatory contents can be calculated based on the explanatory content of the upper-level and lower-level knowledge topics. Optionally, the explanatory content can be a summary and descriptive content generated based on the knowledge content of the knowledge topic, making it easier for users to understand. Optionally, the hierarchical parameter can include relevance indicators, priority indicators, etc. Among them, the relevance index can be used to indicate the relevance between the knowledge content of the upper-level knowledge topic and the lower-level knowledge topic. The priority index can be used to indicate the probability that the upper-level knowledge topic is above the lower-level knowledge topic.
[0126] 3022. If the hierarchical parameter is less than the second threshold, the hierarchical relationship between the upper-level knowledge topic and the lower-level knowledge topic is redefined.
[0127] In this embodiment, after calculating the hierarchical parameter according to step 3021, the hierarchical parameter can be compared with the second threshold. If the hierarchical parameter is greater than or equal to the second threshold, it indicates that the upper-level knowledge topic and the lower-level knowledge topic meet expectations. If the hierarchical parameter is less than the second threshold, it indicates that there is a problem with the setting of the upper-level knowledge topic and the lower-level knowledge topic. Therefore, when the hierarchical parameter is less than the second threshold, the hierarchical relationship between the upper-level knowledge topic and the lower-level knowledge topic can be redefined.
[0128] Optionally, when the hierarchical parameter includes multiple values, it can correspond to multiple second thresholds.
[0129] Optionally, the specific process of redefining the hierarchical relationship between higher-level and lower-level knowledge topics may include:
[0130] First, determine if the first correlation between the upper-level knowledge topic and the lower-level knowledge topic is less than a second threshold. If the first correlation is less than the second threshold, it means that the upper-level knowledge topic does not contain the lower-level knowledge topic. In this case, further calculate the second correlation between the lower-level knowledge topic and the upper-level knowledge topic of the upper-level knowledge topic. If the second correlation is greater than or equal to the second threshold, modify the level of the lower-level knowledge topic so that it has the same level as the upper-level knowledge topic. Otherwise, if the second correlation is less than the second threshold, calculate the correlation between the lower-level knowledge topic and the next higher-level knowledge topic. Continue this process. If, after calculating the correlation between the lower-level knowledge topic and the top-level knowledge topic, the resulting third correlation is less than the second threshold, modify the level of the lower-level knowledge topic to the same level as the top-level knowledge topic.
[0131] 3023. Update the knowledge outline based on the newly determined hierarchical relationships.
[0132] In this embodiment, after determining that the level of the knowledge topic has changed, the knowledge outline can be updated according to the changed knowledge topic to obtain the updated knowledge outline.
[0133] In this embodiment, the effectiveness of the lecture notes is improved by optimizing the knowledge outline.
[0134] In one example, the specific process of segmenting multiple high-level knowledge topics into multiple content blocks may include:
[0135] 3061. Based on the knowledge outline, calculate the content relevance of the fused content of two adjacent advanced knowledge topics.
[0136] In this embodiment, the order of high-level knowledge topics is first determined based on the knowledge outline, and then two adjacent high-level knowledge topics are obtained. Next, appropriate analysis tools or algorithms can be used to calculate the content relevance of the fused content of these two adjacent high-level knowledge topics. For example, the final content relevance can be obtained by analyzing the topic keywords, concept overlap, and information complementarity of the two fused contents.
[0137] 3062. If the correlation between two adjacent advanced knowledge topics is greater than or equal to the third threshold, then the two knowledge topics are determined to belong to the same content block.
[0138] In this embodiment, after calculating the content relevance according to step 3061, the content relevance can be compared with a third threshold. If the content relevance is greater than or equal to the third threshold, it indicates that the two adjacent high-level knowledge topics have a high relevance and belong to the same content block.
[0139] 3063. If the correlation between two adjacent advanced knowledge topics is less than the third threshold, then the two knowledge topics are determined to belong to two content blocks.
[0140] In this embodiment, otherwise, if step 3062 determines that the content relevance is less than the third threshold, then it can be determined that the relevance between two adjacent high-level knowledge topics is low, and thus it can be determined that the two adjacent high-level knowledge topics belong to two content blocks. That is, among two adjacent high-level knowledge topics, the high-level knowledge topic that appears first according to the knowledge outline belongs to the last high-level knowledge topic of the previous content block. And the high-level knowledge topic that appears later according to the knowledge outline belongs to the first high-level knowledge topic of the next content block.
[0141] Figure 4 is a schematic diagram according to the fourth embodiment of this disclosure. Based on the embodiments shown in Figures 1 to 3, as shown in Figure 4, this disclosure provides a specific implementation of a video processing method based on a large model. This specific implementation process can be divided into three stages, which can correspond to three functional modules. These three functional modules can respectively include a preprocessing stage, a content generation stage, and a post-processing stage. The three modules corresponding to these three stages can be a preprocessing module, a content generation module, and a post-processing module, respectively.
[0142] 401. Execute the preprocessing module during the preprocessing stage. This preprocessing module mainly processes the input data required for lecture note generation. The data includes the subtitle text of the video to be processed, screenshots of the courseware, and the text information of the courseware screenshots. The input model consists of the subtitle text and the text information of the courseware screenshots.
[0143] 4011. The input is a video ID. The system queries the video to be processed using this ID. Then, based on this video, it obtains the subtitle text and courseware screenshots. The subtitle text may also include a timestamp. The courseware screenshots may also include a timestamp.
[0144] 4012. Perform deduplication and merging on adjacent courseware screenshots that meet a certain threshold of content similarity. That is, retain only one of multiple similar courseware screenshots. Based on these multiple similar screenshots, the time information of the courseware's spike can be generated according to its minimum and maximum timestamps. Since the courseware screenshots are extracted from the video to be processed, there may be duplicate screenshots. A sliding window algorithm can be used for intelligent filtering to identify and merge adjacent courseware with highly similar text information, effectively removing redundancy and ensuring the conciseness and accuracy of the information.
[0145] 4013. Based on the time information in the courseware screenshots, select all subtitle text within that time range and align the courseware screenshots with the subtitle text. Each page of courseware screenshots covers a specific time segment in the video. The subtitle text within this time segment is the corresponding subtitle text in the courseware.
[0146] 4014. Utilize a large-scale model to fuse the text information from courseware screenshots with their corresponding subtitle text. After obtaining the correspondence between courseware screenshots and subtitle text, the text information from the screenshots can be extracted and input into the large-scale model along with the subtitle text for information fusion, resulting in complete classroom knowledge information. This is because the content in courseware screenshots is generally prepared in advance by the teacher and contains relatively complete written information. However, when the teacher explains the content covered in the screenshots, they may include additional interpretations and key points. Therefore, only by combining these two elements can a relatively complete set of classroom knowledge information be obtained.
[0147] 402. Execute the content generation module during the content generation phase.
[0148] 4021. Use a large model to classify the text information in the courseware screenshots to obtain exercise information and content information. This content information, after being merged with the subtitle text, yields knowledge information. Based on the content structure of the video to be processed and the students' key needs, the content can be simply divided into two categories: exercise information and knowledge information. To ensure the quality and completeness of these two parts, they need to be processed differently.
[0149] 4022. Use a large model to process knowledge information and generate lecture notes. Input text classified as knowledge information into the large model to generate hierarchical knowledge content, resulting in lecture notes.
[0150] 4023. Use a large model to extract questions from the exercise information. Input the text information of the courseware screenshots categorized as questions into the large model to extract the question content.
[0151] 4024. Sort the knowledge information and exercise information according to the timestamp to ensure that the order is consistent with the teacher's explanation order, restore the original explanation logic of the video to be processed, and generate the final handout.
[0152] 403. Execute the post-processing module during the post-processing stage. This post-processing module mainly utilizes the lecture content obtained from the content generation module to perform legalization processing and render and generate lecture file.
[0153] 4031. Correct the format of the handouts using preset rules. The format is Markdown. Specific rules include: adding line breaks before and after Markdown image tags; verifying the existence of image names generated by the model; retaining only the last one if multiple identical image tags are generated consecutively; adding line breaks before and after second-level headings; using $ for enclosing LaTeX formulas, etc. Since this disclosure primarily uses large models for content generation in Markdown format, some content generated by these large models may not meet the requirements for Markdown rendering. Therefore, this step legitimizes the content according to image rendering rules, formula rendering rules, etc., to ensure correct file rendering in the next step.
[0154] 4032. Using Pandoc to render Markdown files into Word and PDF formats. This disclosure mainly uses Pandoc to generate PDF and Word files. The generation of PDF files involves two steps: first, converting the Markdown to a TeX file, and then using XelaTeX to render the TeX file into a PDF file. The advantage of this approach is that different TeX templates can be customized according to needs, ultimately generating lecture notes files in various styles.
[0155] In summary, the preprocessing module enables the generation of high-quality courseware screenshots and text content, which are then integrated with subtitles to achieve a comprehensive description of the classroom content. This results in a more complete video content for processing, laying the foundation for subsequent processing. Furthermore, the content generation and post-processing modules ensure accurate and efficient generation of lecture notes, guaranteeing their comprehensiveness and vividness, thereby improving their effectiveness and auxiliary impact.
[0156] Figure 5 is a schematic diagram according to a fifth embodiment of the present disclosure. As shown in Figure 5, the present disclosure provides a video processing apparatus 500 based on a large model, the apparatus comprising:
[0157] Extraction unit 501 is used to identify the video to be processed and obtain exercise information and knowledge information; wherein, the exercise information includes the text of the exercises in the courseware screenshot of the video to be processed; the knowledge information includes the knowledge text in the courseware screenshot of the video to be processed and the subtitle text corresponding to the audio of the video to be processed; the courseware screenshot is a screenshot of the video to be processed.
[0158] The first generation unit 502 is used to input the knowledge information into a preset lecture handout model, generate the lecture handout content text, and divide the lecture handout content text to obtain at least one content block.
[0159] The second generation unit 503 is used to insert the exercises in the exercise information into at least one of the content blocks to obtain the handout file to be displayed for the video to be processed.
[0160] The apparatus in the embodiment can execute the technical solutions in the above method, and its specific implementation process and technical principles are the same, so they will not be repeated here.
[0161] Figure 6 is a schematic diagram according to a sixth embodiment of the present disclosure. As shown in Figure 6, the present disclosure provides a video processing apparatus 600 based on a large model, the apparatus comprising:
[0162] Extraction unit 601 is used to identify the video to be processed and obtain exercise information and knowledge information; wherein, the exercise information includes the text of the exercises in the courseware screenshot of the video to be processed; the knowledge information includes the knowledge text in the courseware screenshot of the video to be processed and the subtitle text corresponding to the audio of the video to be processed; the courseware screenshot is a screenshot of the video to be processed.
[0163] The first generation unit 602 is used to input the knowledge information into a preset lecture handout model, generate the lecture handout content text, and divide the lecture handout content text to obtain at least one content block.
[0164] The second generation unit 603 is used to insert the exercises in the exercise information into at least one of the content blocks to obtain the handout file to be displayed for the video to be processed.
[0165] Optionally, the second generating unit 603 includes:
[0166] The identification module 6031 is used to identify each exercise in the exercise information to obtain at least one exercise; and to match the content block corresponding to the obtained exercise from the at least one content block.
[0167] The insertion module 6032 is used to insert the obtained exercises into the content block corresponding to the obtained exercises to obtain the file to be displayed.
[0168] Optionally, the exercise has first time information, which indicates the time when the screenshot of the courseware containing the exercise appears on the video to be processed; the content block has second time information.
[0169] The recognition module 6031 includes:
[0170] The matching submodule 60311 is used to determine the content block that matches the obtained exercise based on the first time information of the obtained exercise and the second time information of the content block.
[0171] Optionally, the identification module 6031 includes:
[0172] The splitting submodule is used to split the exercises in the exercise information according to the exercise number to obtain at least one exercise; or, to perform format recognition processing on the exercises in the exercise information to obtain at least one exercise.
[0173] Optionally, the first generating unit 602 includes:
[0174] The topic generation module 6021 is used to input the knowledge information into a preset lecture handout model according to the playback order of the video to be processed, and extract multiple knowledge topics; wherein, the knowledge topic represents the title extracted from part of the knowledge information.
[0175] The outline generation module 6023 is used to input the knowledge topics into the preset lecture handout model according to the playback order of the videos to be processed, and generate a knowledge outline; wherein, the knowledge outline represents the hierarchical relationship of the knowledge topics.
[0176] The explanation generation module 6025 is used to input the knowledge information corresponding to each knowledge topic into the preset lecture handout model and generate the explanation content for each knowledge topic.
[0177] The combination module 6026 is used to combine the explanation content with the knowledge outline according to the knowledge topic to obtain the content text of the lecture notes.
[0178] Optionally, the first generation unit 602 further includes a topic optimization module 6022, used for...
[0179] The first optimization submodule 60221 is used to delete the knowledge topic if the knowledge topic does not match the preset video topic of the video to be processed.
[0180] The second optimization submodule 60222 is used to merge the knowledge topics and the knowledge information corresponding to the knowledge topics if there are two adjacent knowledge topics and their corresponding knowledge information with a content matching degree greater than a first threshold.
[0181] Optionally, the first generation unit 602 further includes an outline optimization module 6024. The outline optimization module 6024 includes:
[0182] The calculation submodule 60241 is used to calculate the hierarchical parameters between the upper-level knowledge topic and the lower-level knowledge topic corresponding to the upper-level knowledge topic in the hierarchical relationship of the knowledge topics in the knowledge outline.
[0183] The judgment submodule 60242 is used to redetermine the hierarchical relationship between the upper-level knowledge topic and the lower-level knowledge topic if the hierarchical parameter is less than the second threshold.
[0184] The update submodule 60243 is used to update the knowledge outline based on the redefined hierarchical relationship.
[0185] Optionally, the first generating unit 602 includes:
[0186] Extraction module 6027 is used to divide the content text of the lecture notes into multiple high-level knowledge topic blocks according to the highest-level high-level knowledge topics in the knowledge outline.
[0187] The segmentation module 6028 is used to combine the high-level knowledge topic blocks to obtain multiple content blocks; wherein each content block includes at least one of the high-level knowledge topic blocks.
[0188] Optionally, the segmentation module 6028 includes:
[0189] The calculation submodule 60281 is used to calculate the content relevance between two adjacent high-level knowledge topic blocks.
[0190] The first judgment submodule 60282 is used to combine two knowledge topic blocks into one content block if the content relevance of two adjacent high-level knowledge topic blocks is greater than or equal to a third threshold.
[0191] Optionally, the extraction unit 601 includes:
[0192] The first recognition module 6011 is used to use a text recognition model to recognize courseware screenshots from the video to be processed, and obtain exercise information and knowledge text.
[0193] The second recognition module 6012 is used to recognize the audio of the video to be processed using an audio recognition model, and obtain the subtitle text corresponding to the audio.
[0194] The fusion module 6013 is used to fuse the knowledge text and the subtitle text to obtain the knowledge information.
[0195] Optionally, the knowledge text has third time information, which indicates the time when the courseware screenshot containing the knowledge text appears on the video to be processed; the subtitle text has fourth time information, which indicates the time point when the subtitle text appears in the audio.
[0196] Fusion module 6013 includes:
[0197] The matching submodule 60121 is used to determine the subtitle text corresponding to the fourth time information that matches the third time information based on the third time information of the knowledge text.
[0198] The fusion submodule 60122 is used to fuse the knowledge text and its matching subtitle text to obtain the knowledge information.
[0199] The apparatus described above can execute the technical solution in the above method. Its specific implementation process and technical principles are the same, and will not be repeated here.
[0200] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0201] According to embodiments of this disclosure, this disclosure also provides a computer program product comprising: a computer program stored in a readable storage medium, at least one processor of an electronic device being able to read the computer program from the readable storage medium, and the at least one processor executing the computer program causing the electronic device to perform the scheme provided in any of the above embodiments.
[0202] Figure 7 illustrates a schematic block diagram of an electronic device 700 for implementing embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0203] As shown in Figure 7, device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 702 or a computer program loaded from storage unit 708 into random access memory (RAM) 703. RAM 703 can also store various programs and data required for the operation of device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via bus 704. Input / output (I / O) interface 705 is also connected to bus 704.
[0204] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, etc.; output unit 707, such as various types of monitors, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0205] The computing unit 701 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as large-model-based video processing methods. For example, in some embodiments, the large-model-based video processing method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the large-model-based video processing method described above can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to perform large-model-based video processing methods by any other suitable means (e.g., by means of firmware).
[0206] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0207] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0208] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0209] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0210] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0211] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.
[0212] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0213] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A video processing method based on a large model, comprising: The video to be processed is identified to obtain exercise information and knowledge information; wherein, the exercise information includes the text of the exercises in the courseware screenshot of the video to be processed; the knowledge information includes the knowledge text in the courseware screenshot of the video to be processed and the subtitle text corresponding to the audio of the video to be processed; the courseware screenshot is a screenshot of the video to be processed; The knowledge information is input into a preset lecture handout model to generate the lecture handout content text, and the lecture handout content text is divided to obtain at least one content block. Insert the exercises from the exercise information into at least one of the content blocks to obtain the presentation file of the lecture notes for the video to be processed.
2. The method of claim 1, wherein, The step of inserting the exercises from the exercise information into at least one of the content blocks to obtain the presentation file of the lecture notes for the video to be processed includes: Identify each exercise in the exercise information to obtain at least one exercise; and match the content block corresponding to the obtained exercise from the at least one content block; The obtained exercises are inserted into the content blocks corresponding to the obtained exercises to obtain the file to be displayed.
3. The method of claim 2, wherein, The exercise has first time information, which indicates the time when the screenshot of the courseware containing the exercise appears on the video to be processed; The content block has second time information; The step of matching the content block corresponding to the obtained exercise from the at least one content block includes: Based on the first time information of the obtained exercises and the second time information of the content blocks, the content blocks that match the obtained exercises are determined.
4. The method of claim 2 or 3, wherein, The process of identifying each exercise in the exercise information to obtain at least one exercise includes: Based on the exercise number in the exercise information, the exercises in the exercise information are split to obtain at least one exercise; Alternatively, the exercises in the exercise information can be processed to identify the format, thereby obtaining at least one exercise.
5. The method of any one of claims 1-4, wherein, The step of inputting the knowledge information into a preset lecture handout model to generate the lecture handout content text includes: The knowledge information is input into a preset lecture model according to the playback order of the video to be processed, and multiple knowledge topics are extracted; wherein, the knowledge topic represents the title obtained by extracting part of the knowledge information; The knowledge topics are input into the preset lecture model according to the playback order of the videos to be processed to generate a knowledge outline; wherein, the knowledge outline represents the hierarchical relationship of the knowledge topics; The knowledge information corresponding to each knowledge topic is input into the preset lecture handout model to generate the explanation content for each knowledge topic. Based on the knowledge topic, the explanation content is combined with the knowledge outline to obtain the text of the lecture notes.
6. The method of claim 5, wherein, After inputting the knowledge information into a preset lecture model according to the playback order of the videos to be processed, and extracting multiple knowledge topics, the process further includes: If the knowledge topic does not match the preset video topic of the video to be processed, then the knowledge topic is deleted; If there are two adjacent knowledge topics and their corresponding knowledge information with a content matching degree greater than a first threshold, then the knowledge topics and the knowledge information corresponding to the knowledge topics are merged.
7. The method of claim 6, wherein, After inputting the knowledge topics into the preset lecture outline model according to the playback order of the videos to be processed, the process further includes: Based on the hierarchical relationship of the knowledge topics in the knowledge outline, calculate the hierarchical parameters between the upper-level knowledge topics and the lower-level knowledge topics corresponding to the upper-level knowledge topics in the hierarchical relationship; If the hierarchical parameter is less than the second threshold, the hierarchical relationship between the upper-level knowledge topic and the lower-level knowledge topic is redefined. The knowledge outline is updated based on the redefined hierarchical relationships.
8. The method according to any one of claims 5-7, wherein, The process of dividing the text content of the lecture notes to obtain at least one content block includes: Based on the highest-level advanced knowledge topics in the knowledge outline, the content text of the lecture notes is divided into multiple advanced knowledge topic blocks; The advanced knowledge topic blocks are combined to obtain multiple content blocks; wherein each content block includes at least one advanced knowledge topic block.
9. The method according to claim 8, wherein, The combination of the high-level knowledge topic blocks to obtain multiple content blocks includes: Calculate the content relevance between two adjacent high-level knowledge topic blocks; If the content relevance of two adjacent advanced knowledge topic blocks is greater than or equal to a third threshold, then the two knowledge topic blocks are combined into one content block.
10. The method of any one of claims 1-9, wherein, The process of identifying the video to be processed yields exercise information and knowledge information, including: Using a text recognition model, the courseware screenshots from the video to be processed are identified to obtain exercise information and knowledge text; Using an audio recognition model, the audio of the video to be processed is recognized to obtain the subtitle text corresponding to the audio. The knowledge text and the subtitle text are fused together to obtain the knowledge information.
11. The method according to claim 10, wherein, The knowledge text has third time information, which indicates the time when the screenshot of the courseware containing the knowledge text appears on the video to be processed; The subtitle text has fourth time information, which indicates the time point in the speech corresponding to the subtitle text; The knowledge text and the subtitle text are fused to obtain the knowledge information, including: Based on the third time information of the knowledge text, determine the subtitle text corresponding to the fourth time information that matches the third time information; The knowledge text and its matching subtitle text are merged to obtain the knowledge information.
12. A video processing apparatus based on a large model, comprising: An extraction unit is used to identify the video to be processed and obtain exercise information and knowledge information; wherein, the exercise information includes the text of the exercises in the courseware screenshot of the video to be processed; the knowledge information includes the knowledge text in the courseware screenshot of the video to be processed and the subtitle text corresponding to the audio of the video to be processed; the courseware screenshot is a screenshot of the video to be processed; The first generation unit is used to input the knowledge information into a preset lecture handout model, generate the lecture handout content text, and divide the lecture handout content text to obtain at least one content block. The second generation unit is used to insert the exercises from the exercise information into at least one of the content blocks to obtain the handout file to be displayed for the video to be processed.
13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-11.
14. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-11.
15. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1-11.