Course file generation method and device, terminal equipment and storage medium
By mixing and recognizing the audio data collected by the microphone with the audio data from the playing video, accurate course text is generated, solving the problem of inaccurate subtitles in micro-lesson videos and improving the accuracy and real-time performance of the subtitles.
Patent Information
- Application Number
- CN202311416081.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-27
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2043-10-27
AI Technical Summary
The accuracy of subtitles in existing micro-lesson videos is low, which affects students' learning outcomes.
By mixing audio data captured by the microphone in the terminal device with the audio data of the playing video, mixed audio data is generated. The audio data captured by the microphone is then identified and processed to generate course text, removing function words to ensure the accuracy of the subtitles.
It improves the accuracy and real-time generation of subtitles in micro-lesson videos, ensuring that subtitles only include the content being explained, reducing interference from audio data played on terminal devices, and enhancing the learning experience.
Smart Images

Figure CN119906796B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more specifically, to a method for generating course documents, a device for generating course documents, a terminal device, and a computer-readable storage medium in the field of computer technology. Background Technology
[0002] With the continuous development of artificial intelligence, traditional teaching methods are also evolving, giving rise to online teaching methods, such as micro-lecture videos. Micro-lecture videos are popular with students due to their flexibility and ability to be viewed repeatedly. To improve the quality and efficiency of students' learning through micro-lectures, subtitles can be added to the videos to enhance the readability of the teacher's explanations.
[0003] In related technologies, after the micro-lesson video is recorded, corresponding explanatory subtitles can be added to the video; however, the accuracy of these subtitles is currently low. Therefore, improving the accuracy of subtitles in micro-lesson videos has become an urgent problem to be solved. Summary of the Invention
[0004] This application provides a method for generating course documents, a device for generating course documents, a terminal device, and a computer-readable storage medium. This method can improve the accuracy of subtitles in micro-lesson videos.
[0005] Firstly, a method for generating course files is provided. This method includes: acquiring audio data collected by a microphone in a terminal device, and audio and image data of a video played in the terminal device, wherein the audio data collected by the microphone does not include the audio data of the video; mixing the audio data collected by the microphone and the audio data of the video to obtain mixed data, wherein the mixing process integrates the audio data collected by the microphone and the audio data of the video, and the mixed data includes both the audio data collected by the microphone and the audio data of the video; recording the mixed data and image data to obtain a course video; performing recognition processing on the audio data collected by the microphone to obtain course text; and generating a course file based on the course video and the course text.
[0006] In the embodiments of this application, during the generation of course files (i.e., micro-lesson videos), audio data is obtained by mixing audio data collected by the microphone and audio data of the video played on the terminal device. The mixed audio data and the image data of the terminal are then recorded to obtain the corresponding course video. Since the audio data collected by the microphone does not include the audio data of the video, the audio data of the video played on the terminal device will not be recognized when the text generated by recognizing the audio data collected by the microphone. This prevents the audio data of the video played on the terminal device from interfering with the text recognized by the audio data collected by the microphone. As a result, the recognized course text only includes the text corresponding to the audio data collected by the microphone, and does not include the text corresponding to the audio data of the video played on the terminal device, thereby improving the accuracy of the subtitles in the micro-lesson video.
[0007] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: after acquiring audio data collected by a microphone in a terminal device, the method further includes: copying the audio data collected by the microphone to obtain first data and second data, wherein the first data and second data are identical; mixing the audio data collected by the microphone and the audio data of a video playback to obtain mixed data, including: mixing the first data and the audio data of the video playback to obtain mixed data; and recognizing the audio data collected by the microphone to obtain course text, including: recognizing the second data to obtain course text; wherein the mixing and recognition processes are performed simultaneously.
[0008] In the embodiments of this application, by copying and processing the audio data collected by the microphone to obtain two identical audio data (i.e., the first data and the second data), it is possible to simultaneously perform recognition processing on the second data while mixing the first data and the audio data of the video being played, thereby synchronously generating course videos and course text, and improving the real-time generation of subtitles in the course files.
[0009] Combining the first aspect and the above implementation methods, in some implementation methods of the first aspect, the second data is subjected to identification processing. The method includes: obtaining function words in the second data; removing function words from the second data; and performing identification processing on the second data after removing function words.
[0010] In the embodiments of this application, since function words are removed from the second data during the recognition process, the second data being recognized can be free of function words and only include explanatory text related to the course, thereby further improving the accuracy of the subtitles in the course file.
[0011] In conjunction with the first aspect and the above implementation methods, in some implementation methods of the first aspect, the method further includes: removing function words from the course text to obtain the target text; generating a course file based on the course video and the course text, including: generating a course file based on the target text and the course video.
[0012] In the embodiments of this application, by removing function words from the course text, the target text containing the removed function words contains only explanatory text related to the course, thereby further improving the accuracy of the subtitles in the course file.
[0013] Combining the first aspect and the above implementation methods, in some implementation methods of the first aspect, the course file includes multiple course sub-files, the course text includes multiple course sub-texts, and the course video includes multiple course sub-videos; the course file is generated based on the course video and the course text, and the method includes: segmenting the course text to obtain multiple course sub-texts; editing the course video based on the multiple course sub-texts to obtain multiple course sub-videos; and generating multiple course sub-files based on the multiple course sub-texts and the multiple course sub-videos.
[0014] In the embodiments of this application, multiple course sub-files can be generated based on multiple course sub-texts and multiple course sub-videos. Since course sub-files can be generated based on course sub-texts and corresponding course sub-videos, the generated course files can meet the user's listening needs or be edited for key course content, thus improving the flexibility of course sub-file generation.
[0015] In conjunction with the first aspect and the above implementation methods, in some implementation methods of the first aspect, the method further includes: detecting whether the sorting of multiple word segments of the course text meets the preset sorting conditions; if it is detected that the sorting of multiple word segments does not meet the preset sorting conditions, adjusting the sorting of multiple word segments to meet the preset sorting conditions; or, if it is detected that the sorting of multiple word segments does not meet the preset sorting conditions, sending a reminder message to the user, so that the user adjusts the sorting of multiple word segments to meet the preset sorting conditions according to the reminder message; wherein, the reminder message is used to prompt the user that the sorting of multiple word segments does not meet the preset sorting conditions.
[0016] In the embodiments of this application, if it is detected that the sorting of multiple word segments does not meet the preset sorting conditions, the sorting of multiple word segments can be adjusted to meet the preset sorting conditions. That is, when the sorting of multiple word segments is disordered, the disordered sorting is adjusted to the correct sorting, thereby making the sentences in the course text more fluent and further improving the accuracy of the subtitles in the micro-lesson video.
[0017] In conjunction with the first aspect and the above implementation methods, in some implementation methods of the first aspect, the method of acquiring audio data collected by the microphone in the terminal device includes: acquiring audio data collected by the microphone according to a preset duration; or, acquiring audio data collected by the microphone according to a preset page number interval, wherein the preset page number interval is determined by the size of the image data.
[0018] In the embodiments of this application, the audio data collected by the microphone is obtained according to a preset duration or a preset page number interval. This avoids the situation where the volume data collected by the microphone is obtained all at once after the course video recording is completed, which can easily result in an excessive amount of audio data and reduce the speed of audio data recognition. Therefore, obtaining the audio data collected by the microphone according to a preset duration or a preset page number interval can reduce the amount of processing required for each audio data recognition, thereby improving the efficiency of audio data recognition.
[0019] In combination with the first aspect and the above implementation methods, in some implementation methods of the first aspect, the method further includes: resampling the second data.
[0020] In the embodiments of this application, the second data can be resampled, which reduces the sampling rate of the second data after resampling, making it more suitable for fast speech recognition scenarios. Since the size of the second data is reduced by resampling, the recognition of the second data is faster, thereby improving the efficiency of subtitle generation in micro-lesson videos.
[0021] Combining the first aspect and the above implementation methods, in some implementation methods of the first aspect, a course file is generated based on the course video and course text. The method includes: obtaining a first timestamp of the audio data collected by the microphone and obtaining a second timestamp of the mixed data; determining the difference between the first timestamp and the second timestamp; controlling the audio data collected by the microphone and the mixed data to be aligned based on the difference; and generating a course file based on the image data, the course text, and the aligned mixed data.
[0022] In the embodiments of this application, after aligning the audio data and mixing data collected by the microphone, it can be shown that the timelines of the audio data and mixing data collected by the microphone are in one-to-one correspondence. Since the mixing data and the image data in the video played by the terminal device are recorded simultaneously, it can be further shown that the timelines of the image data, course text and the aligned mixing data in the video played by the terminal device are also in one-to-one correspondence. This ensures that the time of the image data, audio data and subtitles in the course file is synchronized, avoiding the possibility of the image data, audio data and subtitles in the course file being out of sync.
[0023] Secondly, a course document generation apparatus is provided, comprising: an acquisition module for acquiring audio data collected by a microphone in a terminal device and audio and image data of a video played in the terminal device, wherein the audio data collected by the microphone does not include the audio data of the video; a mixing module for mixing the audio data collected by the microphone and the audio data of the video to obtain mixed data, wherein the mixing process integrates the audio data collected by the microphone and the audio data of the video, and the mixed data includes the audio data collected by the microphone and the audio data of the video; a recording module for recording the mixed data and the image data to obtain a course video; a recognition module for recognizing the audio data collected by the microphone to obtain course text; and a generation module for generating a course document based on the course video and the course text.
[0024] In conjunction with the second aspect, in some implementations of the second aspect, the device further includes: a copying module, which, after acquiring audio data collected by the microphone in the terminal device, performs copying processing on the audio data collected by the microphone to obtain first data and second data, wherein the first data and second data are identical; performs mixing processing on the audio data collected by the microphone and the audio data of the played video to obtain mixed data, wherein the mixing module is specifically used to: perform mixing processing on the first data and the audio data of the played video to obtain mixed data; and performs recognition processing on the audio data collected by the microphone to obtain course text, wherein the recognition module is specifically used to perform recognition processing on the second data to obtain course text; wherein the mixing processing and recognition processing are performed simultaneously.
[0025] Combining the second aspect and the above implementation methods, in some implementation methods of the second aspect, the second data is subjected to identification processing. The identification module is specifically used to: obtain function words in the second data; remove function words from the second data; and perform identification processing on the second data after removing function words.
[0026] In conjunction with the second aspect and the above implementation methods, in some implementation methods of the second aspect, the device further includes: a removal module for removing function words from the course text to obtain the target text; and generating a course file based on the course video and the course text. The generation module is specifically used to generate the course file based on the target text and the course video.
[0027] Combining the second aspect and the above implementation methods, in some implementation methods of the second aspect, the course file includes multiple course sub-files, the course text includes multiple course sub-texts, and the course video includes multiple course sub-videos; based on the course video and the course text, a course file is generated. The generation module is specifically used for: segmenting the course text to obtain multiple course sub-texts; editing the course video based on the multiple course sub-texts to obtain multiple course sub-videos; and generating multiple course sub-files based on the multiple course sub-texts and multiple course sub-videos.
[0028] In conjunction with the second aspect and the above implementation methods, in some implementation methods of the second aspect, the device further includes: a detection module, used to detect whether the sorting of multiple word segments of the course text conforms to preset sorting conditions; if the sorting of multiple word segments is detected to be inconsistent with the preset sorting conditions, the sorting of multiple word segments is adjusted to conform to the preset sorting conditions; or, if the sorting of multiple word segments is detected to be inconsistent with the preset sorting conditions, a reminder message is sent to the user, so that the user adjusts the sorting of multiple word segments to conform to the preset sorting conditions according to the reminder message; wherein, the reminder message is used to prompt the user that the sorting of multiple word segments does not conform to the preset sorting conditions.
[0029] In conjunction with the second aspect and the above implementation methods, in some implementation methods of the second aspect, the audio data collected by the microphone in the terminal device is obtained. Specifically, the acquisition module is used to: acquire the audio data collected by the microphone according to a preset duration; or, acquire the audio data collected by the microphone according to a preset page number interval, wherein the preset page number interval is determined by the size of the image data.
[0030] In conjunction with the second aspect and the above-described implementations, in some implementations of the second aspect, the device further includes: a resampling module for resampling the second data.
[0031] Combining the second aspect and the above implementation methods, in some implementation methods of the second aspect, a course file is generated based on the course video and course text. The generation module is specifically used to: obtain the first timestamp of the audio data collected by the microphone and the second timestamp of the mixing data; determine the difference between the first timestamp and the second timestamp; control the alignment processing of the audio data collected by the microphone and the mixing data based on the difference; and generate the course file based on the image data, the course text, and the aligned mixing data.
[0032] Thirdly, a terminal device is provided, including a memory and a processor. The memory is used to store executable program code, and the processor is used to call and run the executable program code from the memory, causing the terminal device to perform the methods of the first aspect or any possible implementation thereof.
[0033] Fourthly, a computer program product is provided, comprising: computer program code, which, when run on a computer, causes the computer to perform the methods described in the first aspect or any possible implementation thereof.
[0034] Fifthly, a computer-readable storage medium is provided that stores computer program code, which, when executed on a computer, causes the computer to perform the methods described in the first aspect or any possible implementation thereof. Attached Figure Description
[0035] Figure 1 This is a schematic diagram of a scenario for generating micro-lesson videos provided in an embodiment of this application;
[0036] Figure 2 This is a schematic diagram illustrating the principle of course file generation provided in the embodiments of this application;
[0037] Figure 3 This is a flowchart illustrating a method for generating course files according to an embodiment of this application;
[0038] Figure 4 This is a schematic diagram illustrating the display method of subtitles on the screen of a terminal device according to an embodiment of this application;
[0039] Figure 5 This is a flowchart illustrating another method for generating course files provided in an embodiment of this application;
[0040] Figure 6 This is a schematic diagram of the structure of the course document generation device provided in the embodiments of this application;
[0041] Figure 7 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application. Detailed Implementation
[0042] The technical solutions in this application will be clearly and thoroughly described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. "And / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.
[0043] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0044] Figure 1 This is a schematic diagram of a scenario for generating micro-lesson videos provided in an embodiment of this application.
[0045] For example, such as Figure 1 As shown, Figure 1 (a) includes a first interface 110, which can represent a trigger interface for starting the recording of micro-lesson videos. This interface can include a "Start Recording" button to trigger the recording of micro-lesson videos. Figure 1 (b) includes a second interface 120, which can represent the interface during the recording of micro-lesson videos. This interface can include information such as the recording duration. Figure 1 (c) includes a third interface 130, which can represent a trigger interface for ending the recording of the micro-lesson video. This interface can include a "End Recording" button that triggers the end of the recording of the micro-lesson video. Figure 1 (d) includes a fourth interface 140, which can represent the interface for playing the micro-lesson video after it has been recorded. This interface includes the micro-lesson video and corresponding subtitles. Furthermore, the first interface 110, the second interface 120, the third interface 130, and the fourth interface 140 are all displayed on a terminal device. The terminal device can be used to collect the lecturer's (e.g., a teacher's) audio of the course explanation, as well as to record its own audio data (which may include the course content or audio data played by other software) and image data, and generate a micro-lesson video based on this audio and image data.
[0046] It should be understood that terminal devices include, but are not limited to, smartphones, tablets, laptops, and other devices such as video playback devices. Furthermore, the operating system of the terminal includes, but is not limited to, systems deeply developed based on the Android system, Apple's iOS system, systems deeply developed based on the iOS system, or other systems. In other words, any device capable of acquiring audio and image data, and processing audio and image data to generate micro-lesson videos, can be considered a terminal device; this application does not limit this.
[0047] In related technologies, when creating a micro-lecture video, the lecturer can click the start recording icon on the terminal device (e.g., first interface 110) to start the terminal device from the current moment to collect the lecturer's audio of explaining the course, as well as the audio and image data played by the lecturer (e.g., second interface 120); after recording is completed, the lecturer can click the end recording button on the end recording icon on the terminal device (e.g., third interface 130) to end the recording of the micro-lecture video.
[0048] To facilitate better learning for learners (e.g., students) through micro-lecture videos, audio data related to the content being explained can be extracted from the videos. This extracted audio data can then be recognized to obtain corresponding text. This text is then added to the micro-lecture video to generate corresponding subtitles, as shown in the fourth interface 140. While the micro-lecture video plays on the left side of the terminal device, the lecturer's subtitles can be displayed on the right side, allowing learners to understand the course content more clearly and intuitively, thus improving learning effectiveness. However, because the micro-lecture video may include audio data played by the terminal device itself, the subtitles may contain subtitles corresponding to the terminal device's own audio data, not just the lecturer's explanation of the course content. This could easily interfere with learners' viewing of the course explanation subtitles.
[0049] To address the aforementioned issues, this application provides a method for generating course documents, a device for generating course documents, a terminal device, and a computer-readable storage medium, which are described in detail below.
[0050] Figure 2 This is a schematic diagram illustrating the principle of course file generation provided in the embodiments of this application.
[0051] For example, such as Figure 2 As shown, Figure 2 Specifically, it includes the following:
[0052] The audio data collected by the microphone in terminal device S201, the audio data played by video in terminal device S202, and the image data played by video in terminal device S203.
[0053] S204 encodes the audio data collected by the microphone using audio encoder A, and then inputs the encoded audio data collected by the microphone into the network speech recognition service module 205 or the local speech recognition module 206 for recognition, thereby obtaining the text corresponding to the audio data collected by the microphone.
[0054] Furthermore, in step S207, the audio data captured by the microphone and the audio data played from the video are mixed. The mixed audio data (which can be referred to as "mixed data") is input into audio encoder B. In step S208, the mixed data is encoded by audio encoder B, and in step S209, the encoded mixed data is output. Simultaneously, in step S210, the input image data is encoded by video encoder, and in step S211, the encoded image data is output. Then, in step S212, the encoded mixed data and image data are used to generate a micro-lesson video. Further, in step S213, a micro-lesson file is generated from the micro-lesson video and text.
[0055] Figure 3 This is a flowchart illustrating a method for generating course files according to an embodiment of this application. This method can be... Figure 1 Executed by the terminal device in the process.
[0056] For example, such as Figure 3 As shown, the method 300 includes the following implementation process:
[0057] S310: Acquire audio data collected by the microphone in the terminal device and audio and image data of the video being played in the terminal device.
[0058] The audio data collected by the microphone does not include the audio data of the video playback; in other words, the audio data collected by the microphone only includes the narration.
[0059] It should be understood that the audio data used to play videos on a terminal device refers to the audio data directly obtained from the terminal device's internal system, not the audio data collected through a microphone; that is, the audio data used to play videos does not need to be collected through a microphone, but can be directly retrieved from within the terminal device's internal system.
[0060] Optionally, if the audio data collected by the microphone includes both the lecturing voice and the audio played by the terminal device, the lecturing voice can be extracted separately based on the teacher's voice characteristics.
[0061] Optionally, the audio data captured by the microphone may include the audio of one or more teachers lecturing on the course at the current moment, and this audio may also include the audio of students interacting with teachers during the lecture; the playback time of the video on the terminal device is synchronized with the playback audio. Therefore, the audio and image data of the video played on the terminal device are also synchronized with the playback audio. Furthermore, the audio data of the video playback may include its own playing audio (e.g., background music or message notifications), and this audio data can be directly obtained from within the terminal device.
[0062] Optionally, audio data collected by the microphone can be acquired according to a preset duration; or, audio data collected by the microphone can be acquired according to a preset page number interval.
[0063] The preset page number interval is determined by the size of the image data. Since the size of the image data can correspond to the number of pages—for example, an image data size of 60kb corresponds to 5 pages when the image data size is 300kb—the preset duration can represent a fixed duration, such as 5 seconds or 10 seconds. This embodiment of the application does not limit this.
[0064] For example, you can acquire audio data collected by a microphone over a fixed duration of 5 seconds (which can be denoted as "1.aac, 2.aac...n.aac"), where the duration of the previous acquisition and the next acquisition are consecutive; or, you can acquire audio data collected by a microphone for every 5 pages (which can be denoted as "P1.aac, P2.aac...Pn.aac").
[0065] In the embodiments of this application, the audio data collected by the microphone is obtained according to a preset duration or a preset page number interval. This avoids the situation where the volume data collected by the microphone is obtained all at once after the course video recording is completed, which can easily result in an excessive amount of audio data and reduce the speed of audio data recognition. Therefore, obtaining the audio data collected by the microphone according to a preset duration or a preset page number interval can reduce the amount of processing required for each audio data recognition, thereby improving the efficiency of audio data recognition.
[0066] The S320 mixes the audio data captured by the microphone and the audio data from the playing video to obtain mixed audio data.
[0067] Among them, audio mixing can be used to integrate audio data captured by the microphone and audio data played from the video, and the mixed data can include audio data captured by the microphone and audio data played from the video.
[0068] For example, audio data captured by the microphone and audio data from the video playback can be integrated into a stereo track or a mono track, so that the mixed data obtained after mixing processing can include audio data captured by the microphone and audio data from the video playback.
[0069] The S330 records the mixed audio data and image data to obtain course videos.
[0070] For example, the audio mixing data and the image data in the video played on the terminal can be recorded to obtain the video at the current moment (which can be referred to as "course video"). As the recording time increases, a course video of a certain duration can be obtained.
[0071] S340 processes and identifies the audio data collected by the microphone to obtain the course text.
[0072] For example, the audio data corresponding to the lecturing sound collected by the microphone can be identified and processed to obtain one or more texts corresponding to the audio data, and these texts can be combined into text (which can be referred to as "course text").
[0073] It should be understood that in order to ensure the synchronous generation of course videos and course texts, S320 and S340 can be performed simultaneously; in addition, in the embodiments of this application, it is not excluded that the execution steps of S320 and S340 are to execute S320 first and then S340, or to execute S340 first and then S320.
[0074] Optionally, after acquiring the audio data collected by the microphone, the audio data collected by the microphone can be copied to obtain first data and second data, wherein the first data and second data are identical. The audio data collected by the microphone and the audio data from the played video are mixed to obtain mixed data, including: mixing the first data and the audio data from the played video; and the audio data collected by the microphone is recognized to obtain course text, including: recognizing the second data to obtain course text; wherein the mixing and recognition processes are performed simultaneously.
[0075] By copying the audio data captured by the microphone, two identical audio data sets can be obtained, such as identical first and second data sets. These data sets (e.g., the first data set) can then be mixed with the audio data from the playing video to obtain mixed audio data. Simultaneously, the other data set (e.g., the second data set) can be input to... Figure 2 The network speech recognition service module or local speech recognition module shown can be used to recognize the course text corresponding to the second data.
[0076] In this embodiment of the application, by copying and processing the audio data collected by the microphone to obtain two identical audio data (i.e., the first data and the second data), it is possible to simultaneously perform recognition processing on the second data while mixing the first data and the audio data of the video being played, thereby synchronously generating course videos and course text, and improving the real-time generation of subtitles in the course files.
[0077] Since large audio data can slow down the recognition process, the second data can be resampled. This reduces the sampling rate of the resampled second data, making it more suitable for fast speech recognition scenarios. Because the size of the second data is reduced by resampling, the recognition of the second data is faster, thereby improving the efficiency of subtitle generation in micro-lesson videos.
[0078] Optionally, noise reduction processing can be performed on the second data to prevent interference from the microphone capturing audio data that is not the narration sound, thus avoiding interference with the recognition processing.
[0079] Optionally, the second data can be optimized, for example, by converting the dialect into standard Mandarin, or by optimizing the loudness of the second data to make the loudness the same, so as to avoid uneven sound volume in the final output.
[0080] In one possible implementation, when performing recognition processing on the second data, function words in the second data can also be obtained; function words in the second data can be removed; and recognition processing can be performed on the second data after removing function words.
[0081] For example, during the recognition processing of the second data, one or more function words, such as interjections like "ah," "la," and "ne," can be obtained. These function words are then removed from the second data. The second data after removing function words is then processed, meaning the second data after removing function words only contains audio data related to the course. Specifically, without removing function words from the second data, the recognized course text might be "This lesson we will learn addition ah," while with function word removal, the recognized course text is simply "This lesson we will learn addition," meaning the recognized course text only contains explanatory text related to the course.
[0082] In the embodiments of this application, since function words are removed from the second data during the recognition process, the second data being recognized can be free of function words and only include explanatory text related to the course, thereby further improving the accuracy of the subtitles in the course file.
[0083] In another possible implementation, the system detects whether the order of multiple word segments in the course text meets the preset ordering conditions. If the order of multiple word segments does not meet the preset ordering conditions, the order of multiple word segments is adjusted to meet the preset ordering conditions. Alternatively, if the order of multiple word segments does not meet the preset ordering conditions, a reminder message is sent to the user, so that the user can adjust the order of multiple word segments to meet the preset ordering conditions according to the reminder message.
[0084] The preset sorting conditions can represent the order of the subject, predicate, and object, and can be obtained by training a convolutional network model. This application does not limit this.
[0085] The notification message is used to inform the user that the sorting of multiple word segments does not meet the preset sorting criteria. This notification message can be displayed on the terminal screen showing the parts of the word segments that do not meet the preset sorting criteria, or it can be provided through voice prompts, text prompts, etc., and this embodiment of the application does not limit this.
[0086] For example, a trained convolutional network model can be used to detect whether the order of multiple word segments (e.g., ) in the course text meets the preset ordering conditions. If it is detected that the order of "learning addition in this lesson we" does not meet the preset ordering conditions, the order of "learning addition in this lesson we" can be adjusted to "this lesson we learn addition". That is, when the order of multiple word segments is disordered, the disordered order is adjusted to the correct order, so that the order adjustment of multiple word segments can meet the preset ordering conditions.
[0087] Alternatively, if the sorting of "Learning addition in this lesson" is detected to be inconsistent with the preset sorting conditions, a text message can be sent to the user, allowing the user to adjust the sorting of multiple word segments based on the text message. That is, when the sorting of multiple word segments is disordered, the disordered sorting can be adjusted to the correct sorting, so that the sorting adjustment of multiple word segments can meet the preset sorting conditions.
[0088] In the embodiments of this application, if it is detected that the sorting of multiple word segments does not meet the preset sorting conditions, the sorting of multiple word segments can be adjusted to meet the preset sorting conditions. That is, when the sorting of multiple word segments is disordered, the disordered sorting is adjusted to the correct sorting, thereby making the sentences in the course text more fluent and further improving the accuracy of the subtitles in the micro-lesson video.
[0089] S350 generates course files based on course videos and course texts.
[0090] For example, the obtained course videos and course texts can be combined into a single complete file (which can be referred to as a "course file"), thereby generating a course file (which can also be referred to as a "micro-lecture video").
[0091] Furthermore, the generated course files can be saved on the terminal device and / or uploaded to the cloud server.
[0092] Optionally, after obtaining the course text, function words in the course text can be removed to obtain the target text; a course file can be generated based on the course video and the course text, including: generating a course file based on the target text and the course video.
[0093] For example, after obtaining the course text, the course file is inspected. If one or more function words are detected, such as interjections like "ah," "la," and "ne," these function words are removed from the course text. The course text with function words removed is then used as the target text. Specifically, without function word removal, the course text might read "This lesson we will learn addition ah," while with function word removal, the course text reads "This lesson we will learn addition," meaning the target text only contains explanatory text related to the course.
[0094] In the embodiments of this application, by removing function words from the course text, the target text containing the removed function words contains only explanatory text related to the course, thereby further improving the accuracy of the subtitles in the course file.
[0095] Optionally, a first timestamp of the audio data captured by the microphone and a second timestamp of the mixed data can be obtained; the difference between the first timestamp and the second timestamp can be determined; the audio data captured by the microphone and the mixed data can be aligned according to the difference; and a course file can be generated according to the image data, course text and the aligned mixed data.
[0096] For example, if a certain time point (e.g., a first timestamp) in the audio data collected by the microphone is obtained as 02:15, and a certain time point (e.g., a second timestamp) in the mixed audio data is obtained as 02:14, the difference between the first timestamp and the second timestamp can be determined to be 00:01. If the difference is less than or equal to a preset difference (e.g., 00:05), the audio data collected by the microphone and the mixed audio data can be aligned. After the time points of the audio data collected by the microphone and the mixed audio data are aligned, it can be said that the time points of the course text and the time points of the aligned mixed audio data are also aligned. Since the mixed audio data and the image data in the video played by the terminal device are recorded simultaneously, it can be said that the image data in the video played by the terminal device, the course text, and the aligned mixed audio data are also aligned. Thus, the image data in the video played by the terminal device, the course text, and the processed mixed audio data are synthesized to generate a course file.
[0097] In the embodiments of this application, after aligning the audio data and mixing data collected by the microphone, it can be shown that the timelines of the audio data and mixing data collected by the microphone are in one-to-one correspondence. Since the mixing data and the image data in the video played by the terminal device are recorded simultaneously, it can be further shown that the timelines of the image data, course text and the aligned mixing data in the video played by the terminal device are also in one-to-one correspondence. This ensures that the time of the image data, audio data and subtitles in the course file is synchronized, avoiding the possibility of the image data, audio data and subtitles in the course file being out of sync.
[0098] Optionally, when playing the course file, the course text can be played in correspondence with the course video, that is, the timeline of the course text and the timeline of the course video are in one-to-one correspondence. When playing the course file, the course file can be used as annotations to explain the content corresponding to the course video; it can also be used as subtitles or bullet comments to explain the content corresponding to the course video. The comparison of the embodiments in this application is not limited.
[0099] For example, a course file may include multiple course sub-files, a course text may include multiple course sub-texts, and a course video may include multiple course sub-videos; multiple course sub-texts are obtained by segmenting the course text; multiple course sub-videos are obtained by editing the course video based on the multiple course sub-texts; and multiple course sub-files are generated based on the multiple course sub-texts and multiple course sub-videos.
[0100] For example, the course text can be segmented according to a preset duration or preset page number interval to obtain course sub-text 1 (which can be denoted as "1.txt"), course sub-text 2 (which can be denoted as "2.txt"), course sub-text 3 (which can be denoted as "3.txt"), ..., course sub-text n (which can be denoted as "n.txt"). Then, based on the obtained course sub-text 1, course sub-text 2, course sub-text 3, ..., course sub-text n, the course video can be edited to obtain multiple course sub-videos corresponding to each course sub-text, namely course sub-video 1, course sub-video 2, course sub-video 3, ..., course sub-video n. It should be understood that the time points corresponding to course sub-text n and course sub-video n are corresponding.
[0101] Furthermore, course sub-text 1, course sub-text 2, course sub-text 3...course sub-text n and their corresponding course sub-videos 1, course sub-video 2, course sub-video 3...course sub-video n can be synthesized to obtain course sub-file 1, course sub-file 2, course sub-file 3...course sub-file n. For example, course sub-text 1 and course sub-video 1 can be synthesized to obtain course sub-file 1, and so on, which will not be elaborated here.
[0102] Optionally, after generating the course file, the course file can be segmented according to a preset duration or a preset page number interval to obtain multiple course sub-files.
[0103] In the embodiments of this application, multiple course sub-files can be generated based on multiple course sub-texts and multiple course sub-videos. Since course sub-files can be generated based on course sub-texts and corresponding course sub-videos, the generated course files can meet the user's listening needs or be edited for key course content, thus improving the flexibility of course sub-file generation.
[0104] For example, when playing a course file, the subtitles in the course file can be as follows: Figure 1 The fourth interface 140, shown in (d) above, is displayed on the right side of the terminal device screen, or as shown below. Figure 4 The fifth interface 410 shown in (a) is displayed on the left side of the terminal device screen; it can also be displayed as shown in (a). Figure 4 The sixth interface 420 shown in (b) is displayed at the bottom of the terminal device screen, or it may also be displayed at the top of the terminal device screen; or it may also be displayed as shown in (b). Figure 4 The seventh interface 430 shown in (c) and Figure 4 The eighth interface 440 shown in (d) displays subtitles in chronological order, and the display time of the subtitles corresponds to the same time as the audio data output by the microphone.
[0105] It should be understood that the display method of subtitles on the terminal device screen can also be other methods, and the screen ratio of subtitles when displayed on the terminal device screen and the screen ratio of images when the course file is played can be adjusted according to user needs. This application embodiment does not limit this.
[0106] exist Figure 3In the method 300 shown, during the generation of course files (i.e., micro-lesson videos), audio data is obtained by mixing audio data collected by the microphone and audio data of the video played on the terminal device. The mixed audio data and the image data of the terminal are then recorded to obtain the corresponding course video. Since the audio data collected by the microphone does not include the audio data of the video, the audio data of the video played on the terminal device will not be recognized when the text generated by recognizing the audio data collected by the microphone. This prevents the audio data of the video played on the terminal device from interfering with the text recognized by the audio data collected by the microphone. As a result, the recognized course text only includes the text corresponding to the audio data collected by the microphone, and does not include the text corresponding to the audio data of the video played on the terminal device, thereby improving the accuracy of the subtitles in the micro-lesson video.
[0107] It should be noted that the solution provided in this application embodiment is only for the purpose of generating course files, but it does not affect the implementation of this solution in similar scenarios, such as live streaming or song recording.
[0108] Figure 5 This is a flowchart illustrating another method for generating course files provided in this application embodiment. This method can be... Figure 1 Executed by the terminal device in the process.
[0109] For example, such as Figure 5 As shown, the method 500 includes the following implementation process:
[0110] S501, The user has triggered the start recording button for the course file.
[0111] For example, the terminal device can detect when the lecturer clicks on a button such as... Figure 1 The first interface 110 shown in (a) has a start recording button, and when the start recording button is detected to be triggered, the course file generation process begins.
[0112] S502 acquires audio data captured by the microphone, S503 acquires audio data from video playback on the terminal device, and S504 acquires image data from video playback on the terminal device. It should be understood that S502, S503, and S504 are performed synchronously.
[0113] For example, the microphone in the terminal device can capture the narration audio at the current moment, such as Pulse-Code Modulation (PCM) data; the terminal device's own audio data during video playback, such as background music; and the terminal device's own image data during video playback, such as color-coded data. After acquiring the audio data captured by the microphone, S504 can be executed.
[0114] S505 copies and processes the audio data captured by the microphone to obtain the first data and the second data.
[0115] For example, the acquired audio data from the microphone is copied to obtain identical first and second data.
[0116] S506, input the second data into audio encoder A.
[0117] For example, the second data is input into audio encoder A, and the second data encoded by audio encoder A is input as follows: Figure 2 The network speech recognition service module or local speech recognition module shown can execute S506.
[0118] S507 performs recognition processing on the encoded second data to obtain the course text.
[0119] For example, text recognition processing is performed on the second data encoded by audio encoder A to obtain the text corresponding to the second data, thereby obtaining the course text corresponding to the second data. For example, the second data is output in segments according to a preset duration or preset page number interval, corresponding to multiple recognized course sub-texts, course sub-text 1, course sub-text 2, course sub-text 3... course sub-text n.
[0120] After copying the audio data collected by the microphone to obtain the first data, S508 can be executed.
[0121] S508 performs audio mixing processing on the first data and the audio data of the video playback to obtain mixed data.
[0122] For example, the first data and the audio data of the video playback can be mixed, that is, the first data and the audio data of the video playback are integrated to obtain mixed data.
[0123] S509 inputs the mixing data into the audio encoder B.
[0124] For example, the mixed data obtained after mixing can be input into audio encoder B.
[0125] It's important to note that audio encoder A and audio encoder B can use different configurations. For example, audio encoder A can use a 16000 sampling rate mono encoding, resulting in a smaller file size for the encoded second data, making it more suitable for network transmission and speech recognition. Audio encoder B, on the other hand, can use a 44100 sampling rate stereo encoding to ensure the clarity of the audio signal, thereby improving the playback quality of the mixed data. In other words, audio encoder B has a higher sampling rate than audio encoder A.
[0126] S510 inputs image data into the video encoder.
[0127] For example, the image data of the video playback obtained in S503 can be input into the video encoder.
[0128] S511 records the encoded mixed audio data and image data to obtain course videos.
[0129] For example, the mixed audio data encoded by the audio encoder B in S509 and the image data encoded by the video encoder in S510 are recorded to obtain the course video.
[0130] It should be understood that S506 and S507, along with S508, S509 and S510, and S511, are performed synchronously to ensure the real-time generation of course videos and course texts.
[0131] S512 combines course text and course video to generate course files.
[0132] For example, the course text obtained in S507 and the course video obtained in S511 can be combined to obtain a course file, which can be saved and displayed.
[0133] S513, the course file is segmented to obtain multiple course sub-files.
[0134] For example, after generating the course file, it can be edited according to the learning needs or the key content of the course to generate multiple course sub-files.
[0135] S514 saves and displays multiple course sub-files.
[0136] For example, after generating multiple course sub-files, the multiple course sub-files can be saved and / or displayed.
[0137] It should be understood that Figure 5 All steps involved in Figure 3 The corresponding steps have already been introduced and will not be repeated here.
[0138] Figure 6 This is a schematic diagram of the structure of the course document generation device provided in the embodiments of this application.
[0139] For example, such as Figure 6 As shown, the device 600 includes:
[0140] Acquisition module 610: used to acquire audio data collected by the microphone in the terminal device and audio and image data of the video played in the terminal device, wherein the audio data collected by the microphone does not include the audio data of the video played.
[0141] Mixing module 620: Used to mix the audio data collected by the microphone and the audio data of the video playback to obtain mixed data. The mixing process is used to integrate the audio data collected by the microphone and the audio data of the video playback. The mixed data includes the audio data collected by the microphone and the audio data of the video playback.
[0142] Recording module 630: Used to record mixed audio data and image data to obtain course videos;
[0143] Recognition module 640: Used to process and recognize the audio data collected by the microphone to obtain the course text;
[0144] Generation module 650: Used to generate course files based on course videos and course text.
[0145] Optionally, the device 600 further includes: a copy module 660, which, after acquiring audio data collected by the microphone in the terminal device, copies the audio data collected by the microphone to obtain first data and second data, wherein the first data and second data are identical; a mixing module 620, which specifically performs mixing processing on the audio data collected by the microphone and the audio data of the played video to obtain mixed data; and a recognition module 640, which specifically performs recognition processing on the audio data collected by the microphone to obtain course text; wherein the mixing processing and recognition processing are performed simultaneously.
[0146] In one possible implementation, the second data is processed for identification. Specifically, the identification module 640 is used to: obtain function words in the second data; remove function words from the second data; and process the second data after removing function words for identification.
[0147] Optionally, the device 600 further includes: a removal module 670 for removing function words from the course text to obtain the target text; and a generation module 650 for generating a course file based on the course video and the course text.
[0148] In one possible implementation, the course file includes multiple course sub-files, the course text includes multiple course sub-texts, and the course video includes multiple course sub-videos. Based on the course video and the course text, a course file is generated. Specifically, the generation module 650 is used to: segment the course text to obtain multiple course sub-texts; edit the course video based on the multiple course sub-texts to obtain multiple course sub-videos; and generate multiple course sub-files based on the multiple course sub-texts and multiple course sub-videos.
[0149] Optionally, the device 600 further includes: a detection module 680, used to detect whether the sorting of multiple word segments in the course text meets preset sorting conditions; if the sorting of multiple word segments is detected to not meet the preset sorting conditions, the sorting of multiple word segments is adjusted to meet the preset sorting conditions; or, if the sorting of multiple word segments is detected to not meet the preset sorting conditions, a reminder message is sent to the user, so that the user adjusts the sorting of multiple word segments to meet the preset sorting conditions according to the reminder message; wherein, the reminder message is used to prompt the user that the sorting of multiple word segments does not meet the preset sorting conditions.
[0150] In one possible implementation, the acquisition module 610 is specifically used to: acquire audio data collected by the microphone in the terminal device according to a preset duration; or, acquire audio data collected by the microphone according to a preset page number interval, wherein the preset page number interval is determined by the size of the image data.
[0151] Optionally, the device 600 further includes a resampling module 690 for resampling the second data.
[0152] In one possible implementation, a course file is generated based on the course video and course text. The generation module 650 is specifically used to: obtain a first timestamp of the audio data collected by the microphone and a second timestamp of the mixed data; determine the difference between the first timestamp and the second timestamp; control the alignment of the audio data collected by the microphone and the mixed data based on the difference; and generate the course file based on the image data, the course text, and the aligned mixed data.
[0153] It should be noted that the course document generation device provided in the above embodiments is only illustrated by the division of the above functional modules when executing the course document generation method. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0154] Furthermore, the course document generation apparatus and course document generation method embodiments provided in the above embodiments belong to the same concept. Therefore, for details not disclosed in the apparatus embodiments of this specification, please refer to the course document generation method embodiments described above in this specification, which will not be repeated here.
[0155] Figure 7 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application.
[0156] For example, such as Figure 7 As shown, the terminal device 700 includes a memory 710 and a processor 720. The memory 710 stores executable program code 7101, and the processor 720 is used to call and execute the executable program code 7101 to perform a method for generating course files.
[0157] This application can divide the terminal device into functional modules based on the above method example. For example, each module can correspond to a separate functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0158] When each functional module is divided according to its corresponding function, the terminal device may include: an acquisition module, a mixing module, a recording module, a recognition module, and a generation module, etc. It should be noted that all relevant content of each step involved in the above method embodiments can be referenced from the functional descriptions of the corresponding functional modules, and will not be repeated here.
[0159] The terminal device provided in this application is used to execute the above-mentioned method for generating course files, and thus can achieve the same effect as the above-mentioned implementation method.
[0160] When using integrated units, the terminal device may include a processing module and a storage module. The processing module is used to control and manage the actions of the terminal device. The storage module is used to support the execution of program code and data by the terminal device.
[0161] The processing module may be a processor or a controller, which can implement or execute various exemplary logic blocks, modules, and circuits as disclosed in this application. The processor may also be a combination of functions that implement computing capabilities, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, etc., and the storage module may be a memory.
[0162] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described in the foregoing embodiments. The computer-readable storage medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs (Digital Video Discs), CD-ROMs (Compact Disc Read-Only Memory), microdrives, magneto-optical disks, ROMs (Read-Only Memory), RAMs (Random Access Memory), EPROMs (Erasable Programmable Read-Only Memory), EEPROMs (Electrically Erasable Programmable Read Only Memory), DRAMs (Dynamic Random Access Memory), VRAMs (Video Random Access Memory), flash memory devices, magnetic cards or optical cards, nanosystems (including molecular memory ICs), or any type of medium or device suitable for storing instructions and / or data.
[0163] This application also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned steps to implement a method for generating course files as described in the above embodiments.
[0164] In addition, the terminal device provided in the embodiments of this application may specifically be a chip, component or module. The terminal device may include a connected processor and a memory. The memory is used to store instructions. When the terminal device is running, the processor may call and execute the instructions to make the chip execute a course file generation method in the above embodiments.
[0165] The terminal device, computer-readable storage medium, computer program product or chip provided in this application are all used to execute the corresponding methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0166] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0167] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0168] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method of generating a course file, characterized by, The method comprises: acquiring audio data collected by a microphone in a terminal device and audio data and image data of a video played in the terminal device, wherein the audio data collected by the microphone does not include the audio data of the played video; performing audio mixing processing on the audio data collected by the microphone and the audio data of the played video to obtain mixed audio data, wherein the audio mixing processing is used for integrating the audio data collected by the microphone and the audio data of the played video, and the mixed audio data includes the audio data collected by the microphone and the audio data of the played video; recording the mixed audio data and the image data to obtain a course video; performing recognition processing on the audio data collected by the microphone to obtain a course text; generating a course file according to the course video and the course text.
2. The method of claim 1, wherein, After the acquiring of the audio data collected by the microphone in the terminal device, the method further comprises: performing copy processing on the audio data collected by the microphone to obtain first data and second data, wherein the first data and the second data are the same; the performing of the audio mixing processing on the audio data collected by the microphone and the audio data of the played video to obtain mixed audio data comprises: performing the audio mixing processing on the first data and the audio data of the played video to obtain the mixed audio data; and the performing of the recognition processing on the audio data collected by the microphone to obtain a course text comprises: performing the recognition processing on the second data to obtain the course text; wherein the audio mixing processing and the recognition processing are performed synchronously.
3. The method of claim 2, wherein, The performing of the recognition processing on the second data comprises: acquiring function words in the second data; removing the function words in the second data; performing the recognition processing on the second data from which the function words are removed.
4. The method of claim 1, wherein, The method further comprises: removing function words in the course text to obtain a target text; the generating of the course file according to the course video and the course text comprises: generating the course file according to the target text and the course video.
5. The method according to any one of claims 1 to 3, characterized in that, The course file includes a plurality of course sub-files, the course text includes a plurality of course sub-texts, and the course video includes a plurality of course sub-videos; the generating of the course file according to the course video and the course text comprises: performing segmentation processing on the course text to obtain a plurality of the course sub-texts; performing editing processing on the course video according to a plurality of the course sub-texts to obtain a plurality of the course sub-videos; generating a plurality of the course sub-files according to a plurality of the course sub-texts and a plurality of the course sub-videos.
6. The method according to any one of claims 1 to 4, characterized in that, The method further comprises: detecting whether the order of a plurality of segmented words of the course text meets a preset order condition; if it is detected that the order of the plurality of segmented words does not meet the preset order condition, adjusting the order of the plurality of segmented words to meet the preset order condition; or if it is detected that the order of the plurality of segmented words does not meet the preset order condition, sending a reminder information to a user to enable the user to adjust the order of the plurality of segmented words to meet the preset order condition according to the reminder information. The reminding information is used to prompt the user that the sorting of the multiple word segmentation results does not conform to the preset sorting condition.
7. The method of claim 1, wherein, The audio data collected by the microphone in the terminal device is acquired, including: The audio data collected by the microphone is acquired according to a preset time length; or, The audio data collected by the microphone is acquired according to a preset page interval, wherein the preset page interval is determined by the size of the image data.
8. The method of claim 2, wherein, Further comprising: The second data is resampled.
9. The method according to any one of claims 1 to 4 or 7 or 8, characterized in that, The course file is generated according to the course video and the course text, including: The first timestamp of the audio data collected by the microphone is acquired, and the second timestamp of the mixed audio data is acquired; The difference between the first timestamp and the second timestamp is determined; The audio data collected by the microphone and the mixed audio data are aligned according to the difference; The course file is generated according to the image data, the course text and the mixed audio data after alignment.
10. An apparatus for generating a course file, characterized by comprising: The device comprises: An acquisition module is configured to acquire audio data collected by a microphone in a terminal device and audio data and image data of a played video in the terminal device, wherein the audio data collected by the microphone does not include the audio data of the played video; A mixing module is configured to mix the audio data collected by the microphone and the audio data of the played video to obtain mixed audio data, wherein the mixing is configured to integrate the audio data collected by the microphone and the audio data of the played video, and the mixed audio data includes the audio data collected by the microphone and the audio data of the played video; A recording module is configured to record the mixed audio data and the image data to obtain a course video; An identification module is configured to identify the audio data collected by the microphone to obtain a course text; A generation module is configured to generate a course file according to the course video and the course text.
11. A terminal device, comprising: The terminal device comprises: A memory is configured to store executable program codes; A processor is configured to call and run the executable program codes from the memory, so that the terminal device executes the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, when the computer program is executed, the method according to any one of claims 1 to 9 is realized.
Citation Information
Patent Citations
Video file recording method, audio file recording method and mobile terminal
CN107316642A
Method for producing lecture text data in mobile communications terminal and mobile communications terminal using same
KR1020140062247A