Foreign language teaching video generation method and apparatus

By generating foreign language teaching videos using large language models and TTS technology, the problem of traditional teaching videos failing to meet personalized needs is solved, enabling the generation of personalized foreign language teaching videos and improving learning efficiency and interest.

WO2026041000A1PCT designated stage Publication Date: 2026-02-26JIANG QIUSHI

Patent Information

Application Number
PCT/CN2025/115594
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-20
Filing Date
2025-08-19
Publication Date
2026-02-26

AI Technical Summary

Technical Problem

Traditional foreign language teaching videos cannot meet the needs for personalization and interactivity, and existing automatic subtitle tools ignore the deeper meaning of language learning, resulting in insufficient learning efficiency and interest.

Method used

By acquiring video subtitle text, using a large language model to generate foreign language teaching explanation text, and combining it with TTS technology to generate explanation audio, the temporal information relationship is determined, and personalized foreign language teaching videos are generated.

Benefits of technology

It automatically generates personalized foreign language teaching videos to meet learners' specific needs, improve learning efficiency and interest, and provide in-depth language comprehension.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025115594_26022026_PF_FP_ABST
    Figure CN2025115594_26022026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a foreign language teaching video generation method and apparatus. The foreign language teaching video generation method comprises: acquiring video subtitle text corresponding to a video to be processed; on the basis of the video subtitle text and a preset explanation generation rule, using a large language model to generate foreign language teaching explanation text of the video subtitle text; generating a corresponding explanation audio on the basis of the foreign language teaching explanation text; determining a time information relationship between the foreign language teaching explanation text and the explanation audio; on the basis of the time information relationship between the foreign language teaching explanation text and the explanation audio, generating a display style corresponding to the explanation audio; and on the basis of the explanation audio, the display style corresponding to the explanation audio, and the time information relationship between the explanation text and the explanation audio, generating an explanation video corresponding to the video to be processed. In the present invention, a target teaching video corresponding to a video to be processed can be automatically generated, and the generated teaching video is highly targeted, so that the personalized requirements of users can be met.
Need to check novelty before this filing date? Find Prior Art

Description

Foreign language teaching video generation method and device

[0001] The present application claims priority to the Chinese patent application No. 202411144481.3, filed on August 20, 2024, entitled "Foreign language teaching video generation method, generation device and computer program product", the entire content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] The present application relates to the technical field of video processing, in particular to a foreign language teaching video generation method and device based on a large language model. BACKGROUND

[0003] The current foreign language learning market faces a series of challenges, especially for those learners who seek personalized, interactive and interesting learning experiences. Traditional teaching videos are often based on fixed teaching materials. These materials, although structured and systematic, do not meet the interests and needs of all learners. Due to the universality of the teaching materials, they often fail to take into account the specific levels of different learners, resulting in some learners feeling that the content is too simple and thus boring, while others may feel that the content is too difficult and thus frustrating. This one-size-fits-all approach limits the flexibility and efficiency of foreign language learning, making it difficult for many learners to find teaching videos that are both suitable for their current level and interesting, ultimately affecting the sustainability and effectiveness of learning.

[0004] In addition, although automatic captioning and translation tools can compensate for language understanding barriers to some extent, their functions are still relatively limited. These tools mainly focus on the direct translation of vocabulary, ignoring more complex and subtle aspects of language learning, such as the analysis of grammatical structure, the polysemy of vocabulary and the understanding of cultural background. Therefore, learners often only get the surface literal meaning, but cannot deeply understand the underlying meaning and usage scenarios of the language, which is far from enough to develop real language use ability.

[0005] Therefore, it is necessary to automatically generate personalized foreign language teaching videos according to the actual needs and learning progress of learners. SUMMARY

[0006] The present application provides a foreign language teaching video generation method, generation device and computer program product to generate personalized teaching videos to meet the actual needs of learners.

[0007] The present application provides a foreign language teaching video generation method, which includes the following steps:

[0008] Obtaining the video caption text corresponding to the to-be-processed video;

[0009] Based on the video subtitle text and the preset explanation generation rule, a foreign language teaching explanation text of the video subtitle text is generated by using a large language model;

[0010] Based on the foreign language teaching explanation text, corresponding explanation audio is generated;

[0011] The time information relationship between the foreign language teaching explanation text and the explanation audio is determined;

[0012] According to the time information relationship between the foreign language teaching explanation text and the explanation audio, a display style corresponding to the explanation audio is generated;

[0013] According to the explanation audio, the display style corresponding to the explanation audio, and the time information relationship between the explanation text and the explanation audio, an explanation video corresponding to the to-be-processed video is generated.

[0014] According to the foreign language teaching video generation method provided by the application, the video subtitle text corresponding to the to-be-processed video is obtained directly from the to-be-processed video; or

[0015] A to-be-processed video provided by a user is obtained;

[0016] The to-be-processed video is segmented and preprocessed;

[0017] The subtitle text corresponding to the preprocessed video is obtained and stored.

[0018] According to the foreign language teaching video generation method provided by the application, the preset explanation generation rule includes system role identity setting, content generation requirement, and format requirement of the generated result; the content generation requirement includes translation language style, sentence explanation rule, and explanation style.

[0019] According to the foreign language teaching video generation method provided by the application, the foreign language teaching explanation text is processed into a language style acceptable to a TTS model, and a voice synthesis technology is called to generate the explanation audio.

[0020] The foreign language teaching explanation text is processed into a language style acceptable to a TTS model, and a voice synthesis technology is called to generate the explanation audio.

[0021] Further, the foreign language teaching explanation text is processed into a language style acceptable to a TTS model, which includes language marking of the foreign language teaching explanation text, and selection of a suitable TTS model according to different languages;

[0022] The language marking process of the foreign language teaching explanation text is as follows:

[0023] The large language model is prompted to distinguish the explanation text according to languages, and a separator is inserted at the position of language change;

[0024] The text cut by the delimiter is obtained, the delimiter is replaced by an SSML format tag, and an SSML document is generated;

[0025] The generated SSML document is transmitted to a TTS model to generate the explanation audio.

[0026] According to the foreign language teaching video generation method provided by the application, the determination of the time information relationship between the foreign language teaching explanation text and the explanation audio comprises:

[0027] The explanation text is structured and divided into semantic segments according to a predetermined rule;

[0028] The audio text content and the time information of the explanation audio are obtained;

[0029] According to the correspondence between the semantic segments and the audio text content and the time information, each semantic segment is assigned a corresponding time interval;

[0030] The explanation text is time stamped according to the assigned time interval.

[0031] Further, the specific process of generating the display style corresponding to the explanation audio according to the time information relationship between the foreign language teaching explanation text and the explanation audio comprises:

[0032] A format file is generated based on the time stamped explanation text, the correspondence between the semantic segments and the audio text content;

[0033] The display style of the foreign language teaching explanation text is generated according to the format file, and the display style is used to control the display mode of the foreign language teaching explanation text in the explanation video, and the display mode comprises the time sequence presentation, position adjustment, visual style or interactive effect setting of the semantic segments.

[0034] According to the foreign language teaching video generation method provided by the application, the method further comprises:

[0035] The explanation video and the to-be-processed video are spliced to generate an integrated video;

[0036] At least two integrated videos are spliced together to generate a set of foreign language teaching videos.

[0037] The application further provides a foreign language teaching video generation device, which comprises a video processing module, a content generation module, a TTS voice generation module, a time information determination module, a subtitle and visual content generation module and a video generation module.

[0038] The video processing module is configured to process the obtained video to be processed to obtain video subtitle text; the content generation module is configured to generate foreign language teaching explanation text of the video subtitle text based on the video subtitle text and a preset explanation generation rule, and using a large language model; the TTS voice generation module is configured to generate corresponding explanation audio based on the foreign language teaching explanation text; the time information determination module is configured to determine a time information relationship between the foreign language teaching explanation text and the explanation audio; the subtitle and visual content generation module is configured to generate a display style corresponding to the explanation audio based on the time information relationship between the foreign language teaching explanation text and the explanation audio; and the video generation module is configured to generate a teaching video corresponding to the video to be processed based on the explanation audio, the display style corresponding to the explanation audio, and the time information relationship between the explanation text and the explanation audio.

[0039] According to the above specific embodiments of the present application, at least the following beneficial effects are achieved: using the foreign language teaching video generation method provided by the present application, the user can generate foreign language teaching material video using any video of interest, without manual operation, and the target teaching video segment corresponding to the video to be processed can be automatically generated, and the generated teaching video is highly targeted and can meet the personalized needs of the user. BRIEF DESCRIPTION OF DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the present application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings described below are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0041] Fig. 1 is a flowchart of the foreign language teaching video generation method provided by the present application;

[0042] Fig. 2 is a style diagram of generating explanation video clips based on original video clip text content in the foreign language teaching video generation method provided by the present application;

[0043] Fig. 3 is a flowchart of splicing original video clips and explanation video clips in the foreign language teaching video generation method provided by the present application;

[0044] Fig. 4 is a specific teaching explanation video generation flowchart provided by the foreign language teaching video generation method provided by the present application;

[0045] Fig. 5 is a structural diagram of the foreign language teaching video generation device provided by the present application.

[0046] Reference signs: 11, video processing module; 12, content generation module; 13, TTS voice generation module; 14, time information determination module; 15, subtitle and visual content generation module; 16, video generation module. DETAILED DESCRIPTION

[0047] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in combination with the drawings in the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0048] The foreign language teaching video generation method and device provided by the present application will be described below in combination with FIGS. 1-5.

[0049] As shown in FIG. 1, the foreign language teaching video generation method provided by the present application includes the following steps:

[0050] S1, obtaining a video subtitle text corresponding to a to-be-processed video.

[0051] S2, generating a foreign language teaching explanation text of the video subtitle text by using a large language model based on the video subtitle text and a preset explanation generation rule.

[0052] S3, generating a corresponding explanation audio based on the foreign language teaching explanation text.

[0053] S4, determining a time information relationship between the foreign language teaching explanation text and the explanation audio.

[0054] S5, generating a display style corresponding to the explanation audio according to the time information relationship between the foreign language teaching explanation text and the explanation audio.

[0055] S6, generating an explanation video corresponding to the to-be-processed video according to the explanation audio, the display style corresponding to the explanation audio, and the time information relationship between the explanation text and the explanation audio.

[0056] In the above step S1, after obtaining the to-be-processed video provided by the user, the video can be preprocessed by segmentation, and then the subtitle text corresponding to the segmented to-be-processed video is obtained and stored. The subtitle text contains text content and time information.

[0057] The to-be-processed video refers to any video for which an explanation video segment needs to be generated. For example, the to-be-processed video can be a movie, a TV series, an animation video, etc. The present application does not limit the type of to-be-processed video.

[0058] The application does not limit the pre-processing manner of the video to be processed.

[0059] In a possible implementation of the application, if the video to be processed is a long video file, the long video can be cut into two or more short-time video segments based on a preset time length. For example, each short-time video segment can be 3 minutes; or the long video is accurately cut based on the time information of the subtitle text corresponding to the video to be processed, to ensure the integrity of the content of each video segment.

[0060] The application does not limit the manner of obtaining the subtitle text corresponding to the video to be processed.

[0061] In a possible implementation of the application, the audio of the video to be processed can be separated by audio extraction, the audio file can be subjected to speech recognition by ASR speech-to-text technology, the corresponding audio text can be obtained, and the audio text can be used as the video subtitle text of the video to be processed. The corresponding video subtitle text can also be directly obtained according to the subtitle file of the video to be processed. The application does not limit the audio extraction tool and speech recognition technology used.

[0062] In a possible implementation of the application, the original video can be cut into segments based on SRT (SubRip Text, subtitle file) using the text and corresponding timestamp information. The cut segments can be used as material files for subsequent splicing with the explanation video. As shown in Table 1, after processing, the corresponding relationship between the text sentence, timestamp, and corresponding original video segment can be obtained, and the corresponding relationship information can be stored as subsequent data input.

[0063] Table 1: Corresponding relationship between original video subtitle text and timestamp

[0064] In the above step S2, the video subtitle text is processed by calling a large language model, and detailed explanation content for each sentence is automatically generated according to a preset explanation generation rule.

[0065] The preset explanation generation rule is not limited. The preset explanation generation rule can include analysis of the structure, syntax, and word meaning of a sentence, or can be combined in a configuration form according to the learning preferences of a user.

[0066] Specifically, the preset explanation generation rule can include but is not limited to system role identity setting, content generation requirement, and format requirement of the generated result. The content generation requirement includes translation style, sentence explanation rule, and explanation style.

[0067] In a possible implementation of the present application, the mother tongue and the target learning language can be selected on the user front page, or the target learning language can be automatically obtained by video sampling, and then the selected content is transmitted as a parameter to the generation rule. The system supports multiple language combinations, for example, the user's mother tongue can be Chinese and the target learning language can be Egyptian dialect, or the user's mother tongue can be Korean and the target learning language can be English. The support range is jointly affected by the support range of ASR (Automatic Speech Recognition), large language model, and TTS (Text To Speech).

[0068] In a possible implementation of the present application, based on the preset explanation generation rule, a prompt word can be set, wherein the prompt word can include but is not limited to a role setting sentence, a translated language feeling style, a sentence explanation rule, an explanation style, and a setting sentence of a format rule.

[0069] Role setting: The system role identity can be set as a language teaching expert, or the rule can be superimposed by obtaining the user's selected mother tongue and target language.

[0070] For example, "You are a [Japanese] language teaching expert for [Chinese] mother tongue students. Please consider the language feeling habits of the students' mother tongue when generating the explanation", wherein [Chinese] and [Japanese] are the received transmission parameters. In this way, the large language model can generate an explanation method that is more suitable for the language user to easily understand according to the language characteristics of the student's mother tongue when outputting.

[0071] Translated language feeling style: The meaning of the sentence is translated, so that the user has a cognitive concept of the meaning of the sentence to be learned. The setting can require the language model to ensure that the translation is faithful to the meaning and language feeling of the original text. In a possible implementation of the present application, the context of the sentence can also be transmitted to the language model as background information, so that the translation is more consistent with the original context.

[0072] Sentence explanation rule: The teaching explanation method based on the original text is set in the prompt word, which guides the user to learn the grammar knowledge points involved in the video content in a certain way and angle. The specific explanation rule design is not limited in the present application.

[0073] In a possible implementation method, the explanation rule can include a word order explanation rule, a grammar function explanation rule, a context explanation rule, an example and a metaphor rule, etc.

[0074] Among them, the word order explanation rule refers to explaining the arrangement of words in a sentence and explaining the influence of the position of different words in the sentence on the meaning of the sentence. For example, the word order of subject-predicate-object may have different importance in different languages, and the logic needs to be explained.

[0075] The grammatical function explanation rule refers to explaining the specific grammatical function of words, phrases or sentences in the sentence, which helps users understand the structure and logic of the sentence.

[0076] The context explanation rule refers to considering the specific context in which the sentence is placed, including the context and the intention of the language user. When explaining, relevant background information can be provided to help understand the meaning and usage of the sentence.

[0077] The example and metaphor rule refers to explaining language phenomena through examples or metaphors, which can make abstract concepts more specific and easier to understand. This method helps listeners or readers to associate abstract concepts with specific situations.

[0078] As a possible implementation, different target learning languages can also be matched by rule routing. Since the structures of different languages are different, the language types can be pre-set in advance to achieve more refined explanation effects. Then, according to the target learning language selected by the user, the sentence explanation rules in the database are called to splice into the overall prompt words to be generated.

[0079] Explanation style: directly affects the degree of user understanding of information. Clear and concise explanation style can be more easily understood by the recipient.

[0080] Format rule of generated content: Since audio generation will be performed later, the large language model can be required to generate the format at this stage to facilitate subsequent TTS text preprocessing. The main purpose of the format requirement is to optimize the auditory effect after generating audio. Influenced by the structure of the training data of the large language model used, if the format is not required, the model will generate the format required during training by default, which may not be consistent with the required format style. For example, it may generate 1., 2., 3. sequence format before each piece of content, or use unexpected symbols to wrap some text parts, such as wrapping foreign language parts with 「」. These additional contents may cause interruptions, sudden reading of symbol names, and other unexpected effects after TTS generates audio. Therefore, the format requirement of the generated content needs to be set in advance at this stage, or a large language model that has been trained to generate content according to the expected format rule is used. And before generating audio, data format cleaning processing should be performed to remove characters or format styles that do not meet the requirements to prevent auditory interruptions.

[0081] In a possible implementation of the present application, when a general large language model without fine-tuning is used for generation, a few-shot or hard requirement method can be used to constrain the format. The few-shot method can be used in the prompt to provide the desired generation format and format sample to the large language model, or the hard requirement method, such as "do not generate text with xx symbol", is used to prevent the large language model from continuously generating incorrect format content according to the training style. In actual operation, both methods will occur.

[0082] In a possible implementation of the present application, the large language model can be trained according to the expected format rule. In the long run, this method is more stable and controllable, but it still needs to prepare enough formatted data at the beginning of the project to fine-tune the model.

[0083] In a possible implementation of the present application, the prompt can also be packaged in components, and the combination range of each prompt can be customized by the user.

[0084] In the above step S3, the generated explanation text is processed into a language style acceptable to the TTS model, and then a voice synthesis technology (TTS) is called to generate the explanation audio. In order to enhance the accuracy effect, the voice synthesis technology for multi-language effect enhancement can be used for the text with frequent cross of multiple languages. The present application does not limit the voice synthesis technology used.

[0085] In a possible implementation of the present application, the text is first language-annotated, and then a suitable TTS model is selected according to different languages to improve the pronunciation accuracy of the final audio.

[0086] Due to the need for fine audio generation, the explanation text needs to be processed into a more fine marked language acceptable to the TTS before calling the TTS model. In the text-to-speech (TTS) technology, SSML (Speech Synthesis Markup Language) is a common markup language used to define various properties of speech, such as voice model, language, speech speed, etc. The selected markup language is related to the selected TTS model, and the processing steps of converting the explanation text into a markup language are described as follows:

[0087] (1) The large language model is required to distinguish the explanation text by language through the prompt, and a separator is inserted at the position where the language changes, for example, the explanation text is a mixture of Chinese and Japanese characters, and the non-Chinese or Japanese part is wrapped with a separator in the prompt.

[0088] (2) After obtaining the text cut by the separator, the separator is replaced with an SSML format tag to generate an SSML document.

[0089] (3) passing the generated SSML document to a TTS model to generate the narration audio.

[0090] Whether to adopt this language-based format processing and selection of TTS model capabilities. In a possible embodiment of the present application, a multi-language mixed TTS model can also be adopted. If a multi-language mixed TTS model is selected, it is more simplified, and only the required parameters need to be declared in the table header, without specific tag or parameter declaration processing according to the language in the text content.

[0091] In the above step S4, after obtaining the audio file of the narration audio, the time mapping relationship between the narration text and the narration audio is constructed, so that the narration text content can be played synchronously, precisely controlled and dynamically interacted during the display process. The establishment process of the corresponding relationship mainly includes: obtaining the time information of the audio content, structuring the narration text, and realizing the corresponding relationship between the narration unit and the time.

[0092] The step will be described in detail as follows:

[0093] 1. Obtaining the time information of the audio content

[0094] The main purpose of obtaining the time information of the audio content is to obtain the time information of the narration text in the narration audio, and the method of obtaining the time information is not limited. In a possible embodiment of the present application, the time corresponding relationship between the narration text and the audio can be established by the following two types of technical paths, i.e. the first type of path is an indirect method based on speech recognition and text matching, and the second type of path is a direct alignment method based on acoustic model or boundary detection. They will be described in detail as follows.

[0095] (1) Indirect method based on speech recognition and text matching

[0096] The speech recognition technology (ASR) is used to convert the narration audio into text, and a text file with audio text content and time stamp is generated. In the text file, the audio text content has a time stamp in units of sentences, and the time stamp expresses the start time information and the end time information of the sentence. The audio text content is not used for display on the screen, but is used for similarity matching with the narration text to match the time stamp for the narration text. The matching of the audio text content and the narration text can be achieved by the following two ways.

[0097] 1) Text similarity calculation

[0098] Based on cosine similarity, edit distance, vector matching, etc., the explanation text and the recognized audio text content are matched piece by piece, and those with a matching score higher than a threshold are considered to correspond, and then a timestamp is obtained for the explanation text.

[0099] In a possible embodiment of the present application, the Levenshtein distance algorithm can be used when processing text matching. The algorithm calculates the Levenshtein distance between two strings, and then converts the distance into a similarity score by the following formula:

[0100] Where max_length represents the length of the longer string of the two strings, and distance represents the Levenshtein distance between the two strings.

[0101] The basic idea of this algorithm is that the smaller the edit distance between two strings (i.e., the higher the similarity), the higher their similarity score. The edit distance refers to the minimum number of editing operations required to convert one string to another, including insertion, deletion, and replacement of characters. It is suitable for comparing the similarity between two texts, such as spelling correction, search suggestions, etc., and can consider character-level editing operations, suitable for short text and character-level matching.

[0102] It should be noted that the above Levenshtein distance algorithm is only one possible text similarity calculation method in the present application, which is used to illustrate the realizability of the present application, and is not a limiting technical solution. In other embodiments of the present application, various text similarity calculation methods can be used, such as cosine similarity, edit distance, Jaccard similarity coefficient, deep semantic matching algorithm (such as Siamese network, Transformer encoder output), etc. In actual scenarios, appropriate similarity calculation methods can be selected according to data characteristics, and the present application is not limited in the embodiments.

[0103] 2) Semantic vector matching

[0104] The explanation text and the audio text content can be converted into semantic vectors through a pre-trained language model or a semantic embedding model (such as BERT, Sentence-BERT, etc.), and then aligned at the sentence or paragraph level based on semantic similarity (such as cosine similarity, Euclidean distance, or other vector space-based similarity measurement methods). In the embodiments of the present application, the specific model type and alignment strategy are not limited, and can be flexibly selected according to the accuracy requirements and resource constraints.

[0105] (2) Direct alignment based on acoustic model or boundary detection

[0106] This way obtains the time boundary information of each unit (such as a word or a phoneme, etc.) in the text by directly aligning the explanation text with the audio signal, without the need for intermediate transcription subtitles.

[0107] Common techniques include but are not limited to: 1) Forced Alignment: aligning a piece of audio with a known text, outputting the start and end times of each word or phoneme. This method calculates the start and end times of each word or phoneme in the text in the audio based on an acoustic model (such as HMM, DNN-HMM, or Transformer-ASR), obtaining a high-precision time correspondence for driving the generation of display styles. 2) Word boundary recognition: the timestamp information output by the automatic speech recognition system can be combined with the phoneme boundary information in the acoustic model to infer the start and end positions of the words, and then obtain accurate timestamps; the features such as energy change, pause interval, and speech flow interruption of the audio frame can also be analyzed to identify the boundaries of each word or semantic unit in the explanation audio, thereby assisting in establishing the time correspondence of the explanation unit. 3) Frame-level signal analysis: based on audio energy, pause, speech activity, and other frame-level features to detect the potential boundary positions of the explanation unit.

[0108] This way is suitable for scenarios that require high time accuracy, such as word-by-word playback, highlight control, etc.

[0109] The above two types of methods can be flexibly selected according to the actual application scenario, system resources, and synchronization accuracy requirements, and can also be combined to improve the alignment accuracy. In actual scenarios, the appropriate technical method can be selected according to the data characteristics or the structure of the explanation unit to be divided, and the invention embodiments are not limited.

[0110] 2. Structured division of explanation text

[0111] After obtaining the basic time information, the explanation text is processed in a structured manner. The system processes the explanation text in a structured manner and divides it into the required granularity of explanation units (such as paragraphs, sentences, phrases, etc. semantic fragments) as needed. The division methods include but are not limited to:

[0112] 1) Sentence division based on punctuation marks or line break rules;

[0113] The text can be preliminarily divided into sentence units according to punctuation marks such as periods, commas, semicolons, line breaks, and format symbols in the explanation text.

[0114] 2) Division of the semantic structure of the explanation content based on natural language processing techniques (such as identifying subject explanations, clause explanations, etc. segments);

[0115] The semantic segments in the explanation content can be recognized by natural language processing technologies such as syntax analysis, dependency syntax tree, and semantic role labeling, for example, subject-predicate-object structure, modifier segment, and explanatory clause. The processing can be based on a rule engine, a statistical model, or a pre-trained language model (such as BERT or RoBERTa) for context understanding to achieve more semantic granularity unit division. In the embodiments of the present application, the specific NLP model or implementation method is not limited.

[0116] 3) Business logic division based on interaction requirements or teaching strategy settings.

[0117] The structural division rules can be set in various ways such as rule templates, semantic understanding, and user behavior data. For example, combined with semantic segment recognition technology, the teaching steps (such as guidance, explanation, practice, and summary) in the content are recognized; or the structural boundary is optimized according to the user interaction behavior (such as pause, playback, and high-frequency click). In the embodiments of the present application, the specific implementation method is not limited, and can be flexibly configured according to the specific application scene and system capability.

[0118] 3. Realize the correspondence between the explanation unit and the time

[0119] After completing the structural division, the system combines and processes the time information according to the granularity type of the explanation unit, which is used to drive the content display: 1) word-level explanation unit: directly use the time stamp obtained by word boundary recognition or forced alignment; 2) phrase / syntactic component explanation unit: the system can calculate its interval through the earliest start time and the latest end time of the words it contains; 3) sentence / paragraph-level explanation unit: integrate the time interval of the contained phrase or sentence unit, and if necessary, context smoothing processing can be introduced; in addition, the system can introduce frame-level audio feature analysis (such as pause, speech energy change, etc.) for boundary fine-tuning or error compensation to improve the stability and accuracy of time mapping.

[0120] Based on the time correspondence between the explanation text and the audio established in the foregoing, each explanation unit is allocated with a corresponding time interval, and the alignment method used can be flexibly selected according to the text granularity and audio features. In some scenarios, content alignment methods (such as text similarity calculation, semantic segment alignment, forced alignment, word boundary recognition, etc.) are used for fine-grained time matching of each explanation unit to improve synchronization accuracy.

[0121] After the correspondence between the explanation text and the time information is obtained in step S5, the display style of the explanation text can be generated. The text display aims to enhance understanding and memory, and keep the audio and the explanation content synchronized. After the display content is generated, the correspondence between the text, the display content, and the time information can be obtained, and the correspondence can be stored. Thus, in subsequent generation of the video, the display of the explanation text can be performed according to the time, that is, the explanation text content is displayed in the audio playing time corresponding to the text. The display style can be a subtitle style or a picture style, which is not limited in the present application.

[0122] The format file is generated based on the timestamped explanation text, the correspondence between the semantic segment and the audio text content. The format file includes the divided teaching explanation text and the corresponding start time and end time. The display style of the foreign language teaching explanation text is generated according to the format file. The display style is used to control the display mode of the foreign language teaching explanation text in the explanation video. The display mode includes the time sequence presentation, position adjustment, visual style, or interactive effect setting of the semantic segment.

[0123] Specifically, according to the time information relationship between the foreign language teaching explanation text and the explanation audio, the specific process of generating the display style corresponding to the explanation audio is as follows: the display style of the foreign language teaching explanation text is generated according to the format file, and the display style for controlling the display time sequence of the explanation unit is generated based on the time correspondence. The display style can include but is not limited to the appearance / disappearance time sequence of the text, the highlight block, the subtitle style, the bubble annotation, the position animation, or other ways for enhancing the visualization effect of teaching.

[0124] The display style can be a subtitle style or a picture style, etc. As a preferred embodiment, the original video subtitle text and the corresponding original video screenshot can be used as the background display of the explanation video, and the display style is displayed on the background display.

[0125] As a specific example, as shown in FIG. 2, the display style is the grammar explanation text, which is displayed on the background of the original video subtitle text.

[0126] In a possible embodiment of the present application, the generated subtitles or pictures can be burned into the video, which can facilitate the user to view at any time without network environment.

[0127] In a possible embodiment of the present application, a front-end mask layer can be added on the page to allow the user to set the control to turn on or off the subtitles (or other visual styles). Thus, the user can view the foreign language explanation content as needed when learning, and enhance the association memory.

[0128] In step S6, the audio and display style are combined based on the time information relationship between the explanation text and the audio to generate an explanation video segment, so that each sentence of the explanation text is accompanied by detailed linguistic audio and screen display content, which helps to improve the efficiency of language learning.

[0129] In a possible implementation of the present application, the explanation video segment can be spliced with the original video clip, so that the original video clip is played first, and then the linguistic explanation for the sentence in the video is played, thereby enhancing user memory.

[0130] For ease of understanding, as shown in FIG. 3, the original video clip and the explanation video generated based on the text in the original video clip can be spliced to obtain a spliced video. The front part of the spliced video is the original video clip, and the rear part is the explanation video, which can enhance user memory.

[0131] In a possible implementation of the present application, after the spliced video is generated, a plurality of spliced videos can be spliced to output a set of foreign language teaching videos.

[0132] As can be seen from the entire process of the present application, the user only needs to provide the video to be processed, and the rest of the process will be automatically processed by the present application. Finally, a video accompanied by an original sound clip and detailed explanation is output, and the user does not need to intervene in the entire process.

[0133] For ease of understanding of the entire process, as shown in FIG. 4, a specific teaching explanation video generation process is as follows: extracting audio from a video to be processed to obtain an audio file; performing speech recognition on the audio file to obtain video subtitle text; inputting the video subtitle text into a large language text generation model to obtain teaching explanation text; performing speech synthesis on the teaching explanation text to obtain teaching explanation audio; and generating a teaching explanation video according to the teaching explanation audio.

[0134] Using the foreign language teaching video generation method provided by the present application, the user can generate foreign language teaching material videos using any video of interest without manual operation, and the target teaching video segment corresponding to the video to be processed can be automatically generated. The user's native language and the target learning language can be arbitrarily selected, so the scalability is very high. The user can select a teaching structure combination to expand more personalized learning content. When generating explanation audio, the high-frequency cross-mixed language is optimized for effect, which can ensure that the pronunciation of the explanation content and the original sound is accurate and clear. The explanation content is automatically displayed in the video along the time axis, which facilitates the user to intuitively understand the text content when listening to the explanation. The present application can be applied to a mobile terminal and is adapted to the size of the mobile terminal. The overall end-to-end accuracy of the present application is controllable.

[0135] In the example embodiment, based on the foreign language teaching video generation method provided by the present application, as shown in FIG. 5, the present application further provides a foreign language teaching video generation device, which comprises a video processing module 11, a content generation module 12, a TTS speech generation module 13, a time information determination module 14, a subtitle and visual content generation module 15, and a video generation module 16.

[0136] The video processing module 11 is configured to process the obtained to-be-processed video to obtain a video subtitle text. The content generation module 12 is configured to generate a foreign language teaching explanation text of the video subtitle text by using a large language model according to the video subtitle text and a preset explanation generation rule. The TTS speech generation module 13 is configured to generate corresponding explanation audio according to the foreign language teaching explanation text. The time information determination module 14 is configured to determine the time information relationship between the foreign language teaching explanation text and the explanation audio. The subtitle and visual content generation module 15 is configured to generate a display style corresponding to the explanation audio according to the time information relationship between the foreign language teaching explanation text and the explanation audio. The video generation module 16 is configured to generate an explanation video corresponding to the to-be-processed video according to the explanation audio, the display style corresponding to the explanation audio, and the time information relationship between the explanation text and the explanation audio.

[0137] Specifically, the video processing module 11 cuts the video to cut long videos into video segments with shorter time; decodes each video segment to extract video stream and audio stream; and converts the audio file into a timestamped subtitle text by using the ASR speech-to-text technology. The video processing module can improve the conversion accuracy by performing processing such as splitting the speech, filtering the noise, and identifying the dialect.

[0138] The content generation module 12 processes the extracted subtitle text by using a large language model, and automatically generates detailed explanation content for each sentence according to a preset explanation generation rule. The preset explanation generation rule can include analyzing the structure, syntax, and word meaning of the sentence, or can be combined in a configuration form according to the user's learning preference.

[0139] The TTS speech generation module 13 generates explanation audio by using a text-to-speech technology (TTS); and performs language identification on the generated explanation content to provide subsequent TTS speech synthesis accuracy. When multiple language pronunciations are involved, a suitable TTS model can be selected according to different languages to improve accuracy, or a good multi-language mixed TTS model can be used.

[0140] The time information determination module 14 obtains the corresponding explanation audio subtitle file based on the audio file of the explanation audio. Based on the correspondence between the subtitle text of the explanation audio and the explanation text, the time information of the explanation text in the explanation audio is determined, so that the explanation text and the explanation audio can be time-synchronized.

[0141] The subtitle and visual content generation module 15 generates synchronized subtitles according to the explanation content, or displays in other visual styles, such as images, front masks, etc. The design of the visual content aims to enhance understanding and memory while maintaining synchronization with the explanation content.

[0142] The video generation module 16 splices the processed audio, subtitles and visual content with the original video slices to output the final teaching video, in which each sentence is accompanied by detailed linguistic explanation and visual auxiliary materials.

[0143] It should be noted that the foreign language teaching video generation device and the foreign language teaching video generation method provided in the above embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0144] On the other hand, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and when the computer program is executed by a processor, the processor can execute the foreign language teaching video generation method provided by each method, the method comprises:

[0145] Obtaining the video subtitle text corresponding to the to-be-processed video;

[0146] Based on the video subtitle text and the preset explanation generation rule, a large language model is used to generate a foreign language teaching explanation text of the video subtitle text;

[0147] Generating corresponding explanation audio based on the foreign language teaching explanation text;

[0148] Determining the time information relationship between the foreign language teaching explanation text and the explanation audio;

[0149] Generating a display style corresponding to the explanation audio according to the time information relationship between the foreign language teaching explanation text and the explanation audio;

[0150] Generating an explanation video corresponding to the to-be-processed video according to the explanation audio, the display style corresponding to the explanation audio, and the time information relationship between the explanation text and the explanation audio.

[0151] The device embodiments described above are only schematic, wherein the units illustrated as separate components can or can not be physically separated, and the components illustrated as units can or can not be physical units, i.e., they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the present embodiment scheme according to actual needs. Those skilled in the art can understand and implement it without creative labor.

[0152] Those skilled in the art can clearly understand the implementation of the various embodiments by means of software and necessary general hardware platforms through the description of the above embodiments, and of course, the embodiments can also be implemented by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, and the computer software product can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the method described in each embodiment or some parts of the embodiment.

[0153] Finally, it should be noted that: the above examples are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing examples, or make equivalent replacement for some technical features thereof; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A foreign language teaching video generation method, characterized by, The method comprises the following steps: obtaining video subtitle text corresponding to a to-be-processed video; generating foreign language teaching explanation text of the video subtitle text based on the video subtitle text and a preset explanation generation rule by using a large language model; generating corresponding explanation audio based on the foreign language teaching explanation text; determining the time information relationship between the foreign language teaching explanation text and the explanation audio; generating a display style corresponding to the explanation audio according to the time information relationship between the foreign language teaching explanation text and the explanation audio; generating a to-be-processed video corresponding to the explanation video according to the explanation audio, the display style corresponding to the explanation audio, and the time information relationship between the explanation text and the explanation audio.

2. The foreign language teaching video generation method of claim 1, wherein The method comprises the following steps: obtaining video subtitle text corresponding to a to-be-processed video; or obtaining a to-be-processed video provided by a user; segmenting and preprocessing the to-be-processed video; 3. The foreign language teaching video generation method of claim 1, wherein obtaining and storing the subtitle text corresponding to the preprocessed video.

4. The foreign language teaching video generation method of claim 1, wherein The preset explanation generation rule comprises system role identity setting, content generation requirement, and format requirement of the generated result; the content generation requirement comprises translation sense style, sentence explanation rule, and explanation style. The method comprises the following steps:

5. The foreign language teaching video generation method of claim 4, wherein, processing the foreign language teaching explanation text into a language style acceptable to a TTS model, and then calling a voice synthesis technology to generate the explanation audio. The method comprises the following steps: performing language type annotation on the foreign language teaching explanation text, and selecting a suitable TTS model according to different language types; The process of performing language type annotation on the foreign language teaching explanation text comprises the following steps: requiring the large language model to distinguish the explanation text by language type through a prompt word, and inserting a separator at the position where the language type changes; 6. The foreign language teaching video generation method of claim 1, wherein obtaining the text cut by the separator, replacing the separator with an SSML format tag to generate an SSML document; passing the generated SSML document to the TTS model to generate the explanation audio. The method comprises the following steps: performing structured processing on the explanation text, and dividing it into semantic segments according to a predetermined rule; obtaining the audio text content and time information of the explanation audio; 7. The foreign language teaching video generation method of claim 6, wherein, allocating a corresponding time interval to each semantic segment based on the correspondence between the semantic segments and the audio text content and the time information; time stamping the explanation text according to the allocated time interval. The method comprises the following steps:

8. The foreign language teaching video generation method of claim 1, wherein, generating a format file based on the time-stamped explanation text, the correspondence between the semantic segments and the audio text content; generating a display style of the foreign language teaching explanation text according to the format file, wherein the display style is used to control the display mode of the foreign language teaching explanation text in the explanation video, and the display mode comprises time sequence presentation, position adjustment, visual style, or interactive effect setting of the semantic segments. The method further comprises the following steps: splicing the explanation video and the to-be-processed video to generate an integrated video. Splice at least two integrated videos together to generate a set of foreign language teaching videos.

9. A foreign language education video generation device characterized by comprising: The system comprises a video processing module, a content generation module, a TTS speech generation module, a time information determination module, a subtitle and visual content generation module, and a video generation module. The video processing module is configured to process the obtained video to be processed to obtain video subtitle text; the content generation module is configured to generate foreign language teaching explanation text of the video subtitle text by using a large language model according to the video subtitle text and a preset explanation generation rule; the TTS speech generation module is configured to generate corresponding explanation audio according to the foreign language teaching explanation text; the time information determination module is configured to determine the time information relationship between the foreign language teaching explanation text and the explanation audio; the subtitle and visual content generation module is configured to generate a display style corresponding to the explanation audio according to the time information relationship between the foreign language teaching explanation text and the explanation audio; and the video generation module is configured to generate an explanation video corresponding to the video to be processed according to the explanation audio, the display style corresponding to the explanation audio, and the time information relationship between the explanation text and the explanation audio.

Citation Information

Patent Citations

  • Teaching video automatic subtitle processing method and system

    CN111986656A

  • Financial marketing video generation method and device, computer equipment and storage medium

    CN114401377A

  • Video labeling method and device, electronic equipment and storage medium

    CN118277609A

  • Video explanation content generation method and neural network model training method

    CN118298358A

  • Foreign language teaching video generation method and device and computer program product

    CN119600126A

Cited By

  • Teaching resource cloud platform

    CN122154685A