English video role learning processing method and device based on intelligent terminal, and terminal
Through AI audio recognition technology, video roles are divided and mute control is performed on smart terminals. Combined with real-time evaluation and personalized guidance, the problems of insufficient learning interest and poor interactivity in the existing system are solved, and an efficient English learning experience is achieved.
Patent Information
- Application Number
- CN202510752023.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-16
AI Technical Summary
Existing smart terminal English learning systems lack role learning processing functions, resulting in insufficient user interest in learning, poor interactivity and fun, inability to train for dialogues with different roles, and lack of real-time pronunciation evaluation and personalized guidance.
Video roles are divided through AI audio recognition technology. After the user selects a specified role, the audio of that role is muted and subtitles are displayed. The user's dubbing is compared with the original audio in real time, and ratings and improvement suggestions are provided. The system also automatically searches for pronunciation guidance videos based on non-standard pronunciation.
It improves users' interest and efficiency in English learning, realizes personalized pronunciation training and immersive learning in multi-role dialogue scenarios, and enhances the interactivity of learning and real-time feedback mechanism.
Smart Images

Figure CN120658908A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent terminals and artificial intelligence technology, and in particular to an English video character learning and processing method and device based on an intelligent terminal, an intelligent terminal, and a storage medium. Background Art
[0002] With the development of science and technology and the continuous improvement of people's living standards, the use of various smart terminals such as smart TVs is becoming more and more popular. Users often use smart terminals to play English videos for English learning.
[0003] Existing smart terminal English learning methods, such as those related to children's spoken English, primarily take two forms: one involves oral reading following specific sentences, and the other is based on large AI models and conducted through dialogue. The former is relatively traditional, while the latter requires a higher level of English proficiency for children. Both approaches lack appeal and are designed solely for learning's sake. They lack the ability to handle character learning in English videos, making it difficult for users to practice specific characters, hindering their interest in learning English. Traditional English learning methods often employ a one-way, indoctrination-based approach, where learners passively absorb knowledge. This lacks interactivity and interest, making it difficult to stimulate learning enthusiasm. Particularly in the field of children's English education, existing learning methods often fail to address children's natural tendency to imitate and role-play, resulting in poor learning outcomes. Furthermore, existing English learning systems often fail to provide targeted training for dialogues involving different characters, making it difficult for learners to receive personalized pronunciation guidance and improvement suggestions, resulting in low learning efficiency. In video learning scenarios, existing technologies lack effective mechanisms for multi-character dialogues, cannot implement silent dubbing for specific characters, and lack the ability to compare and evaluate learners' pronunciation accuracy in real time. These problems seriously restrict the fun and effectiveness of English learning through videos. In view of the above problems, the existing technology needs to be improved.
[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0005] To address the above technical issues, the present invention provides a method and apparatus for processing English video character learning based on a smart terminal, as well as a smart terminal and storage medium. This invention addresses the technical problem that existing smart terminals lack English video character learning capabilities, making it difficult for users to choose the right characters to practice learning and thus hindering their interest in English learning. The present invention has the advantage of increasing users' interest and efficiency in English learning.
[0006] The present application provides an English video role learning and processing method based on a smart terminal, and the technical solution is as follows: an English video role learning and processing method based on a smart terminal, comprising: obtaining a target English video, and performing audio analysis on the target English video to extract audio information; dividing the extracted audio information into roles through AI audio recognition, and at the same time translating and extracting subtitles corresponding to the audio of the divided roles, and marking the subtitles of the corresponding segments of the target English video with role labels of different colors; when receiving an instruction to play the target English video, confirming that the user selects a designated role for dubbing; controlling the playback of the target English video according to the confirmed designated role for dubbing, and when the dubbing screen of the designated role is played, controlling the audio of the designated role to be muted and only displaying the screen and corresponding subtitles, and controlling the real-time acquisition of the dubbing audio of the current designated role; at the same time, controlling the normal playback of the audio of other roles; comparing the obtained dubbing audio of the current designated role with the original dubbing audio of the corresponding role in the target English video, and outputting the comparison score results and improvement suggestions.
[0007] Furthermore, the present application also proposes that, when an instruction to play a target English video is received, a pop-up window is used to prompt the user whether to choose to play a role in the video and perform spoken English dubbing; and the selectable dubbing roles are displayed; the user's operation instruction is received, and the user's selection of the designated dubbing role is confirmed; according to the designated dubbing role selected by the user, the audio of the corresponding designated role in the target English video is muted, and the audio of other roles is normal.
[0008] Furthermore, the present application also proposes to compare the obtained dubbing audio of the currently designated character with the original dubbing audio of the corresponding character in the target English video, detect whether the dubbing audio of the currently designated character is correct and output the comparison score results and improvement suggestions for further practice; and according to the preset memory rules, regularly remind users to practice repeatedly.
[0009] Furthermore, the present application also proposes that when it is detected that the pronunciation of the dubbing audio of the currently specified character is not standard compared with the original dubbing audio of the corresponding character in the target English video, the pronunciation guidance practice video corresponding to the audio content with non-standard pronunciation is automatically searched through AI, and output for playback for the user to perform improvement practice.
[0010] Furthermore, the present application also proposes to classify the target English videos for role division according to the user's age; when receiving an instruction to play the target English video, display the currently dubbed characters; and receive the user's operation instruction to select multiple designated dubbing characters.
[0011] Furthermore, the present application also proposes to obtain extracted audio information, and use AI audio recognition to identify the audio information of different people, divide the roles according to the identified audio information of different people, and divide the audio into different roles; translate the audio of the divided different roles into corresponding subtitles; extract the subtitles corresponding to the audio of the divided roles, and mark the subtitles of the corresponding roles on the corresponding clip screen of the target English video, and mark the role labels of different colors on the subtitles of the corresponding clip screen.
[0012] Furthermore, the present application also proposes to control the start of playback of the target English video according to the designated dubbing role selected by the user; when the dubbing screen of the designated role is played, the dubbing screen of the designated role is controlled to be displayed, and the audio of the designated role is controlled to be muted and the corresponding subtitles marked with different color role labels are displayed; and the audio of other characters is controlled to play normally.
[0013] Furthermore, the present application also proposes an English video role learning and processing device based on an intelligent terminal, comprising: an acquisition module for acquiring a target English video, performing audio analysis on the target English video, and extracting audio information; a role division module for dividing the extracted audio information into roles through AI audio recognition, and at the same time translating and extracting subtitles corresponding to the audio of the divided roles, and marking the subtitles of the corresponding segments of the target English video with role labels of different colors; a dubbing role selection module for confirming the user's selection of a designated role for dubbing upon receiving an instruction to play the target English video; a dubbing control module for controlling the playback of the target English video according to the confirmed designated role for dubbing, and when the dubbing screen of the designated role is played, controlling the audio of the designated role to be muted and only displaying the screen and corresponding subtitles, and controlling the real-time acquisition of the dubbing audio of the current designated role; and at the same time controlling the normal playback of the audio of other roles; a result output and suggestion module for comparing the acquired dubbing audio of the current designated role with the original dubbing audio of the corresponding role in the target English video, and outputting the comparison score results and improvement suggestions.
[0014] Furthermore, the present application also proposes an intelligent terminal comprising a memory and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by one or more processors, wherein the one or more programs include those for executing the above method.
[0015] Furthermore, the present application also proposes a computer-readable storage medium, which enables the electronic device to perform the above method when instructions in the storage medium are executed by a processor of the electronic device.
[0016] From the above, it can be seen that the present application provides an English video role learning and processing method, device, smart terminal and storage medium based on a smart terminal, which realizes interactive learning between users and video characters through role division, silent dubbing control and real-time comparative evaluation mechanism, stimulates learning interest and improves pronunciation training effect, and has the advantage of improving users' interest and efficiency in English learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 1 is a flow chart of an English video role learning processing method based on a smart terminal provided in Example 1 of the present invention.
[0019] Figure 2 This is a schematic diagram of an example interface for starting a conversation in the English video role learning processing method based on a smart terminal provided by an embodiment of the present invention.
[0020] Figure 3 This is a schematic diagram of an example interface of a specific English speaking zone of the English video role learning processing method based on a smart terminal provided by an embodiment of the present invention.
[0021] Figure 4 4 is a flow chart of an English video role learning processing method based on a smart terminal provided in Example 2 of the present invention.
[0022] Figure 5 This is a diagram of triggering during movie watching, illustrating an example of an English video role learning processing method based on a smart terminal provided by an embodiment of the present invention.
[0023] Figure 6 A functional block diagram of an embodiment of an English video character learning and processing device based on a smart terminal provided by the present invention.
[0024] Figure 7 This is a block diagram of the internal structure of the smart terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0026] It should be noted that if the embodiments of the present invention involve directional indications (such as up, down, left, right, front, back, etc.), the directional indications are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.
[0027] Existing English learning software typically uses sentence-reading or AI dialogue modes, but these methods are not attractive enough for children and fail to stimulate long-term learning interest. Traditional reading-reading methods lack interactive scenarios, making them boring for children. AI dialogue modes also require a high level of language proficiency, making them difficult for younger users to adapt to. For example, in cartoon learning scenarios, while children may be interested in watching the content, they are unable to participate in character interaction, resulting in a disconnect between language input and output.
[0028] To solve the above problems, the inventors observed that children showed a strong interest in role-playing in pretend games, and thought about how to integrate this nature into English learning. The traditional video playback mode only outputs information in one direction, and users passively receive it, which cannot form effective language training. If the lines of specific characters can be replaced with user dubbing, it can not only retain the fun of the video but also create language practice opportunities. The technical difficulty lies in accurately identifying the voices of video characters, dynamically controlling the audio channels, and establishing a real-time pronunciation evaluation system. By introducing AI voice processing technology, the separation of character voices and the simultaneous comparison of user dubbing can be achieved, and an immersive learning environment can be constructed.
[0029] Therefore, the present application proposes an English video character learning and processing method based on a smart terminal, which is described in detail in the following embodiments.
[0030] Example 1 like Figure 1 As shown, an English video character learning and processing method based on an intelligent terminal in embodiment 1 of the present invention includes the following steps: Step S100: obtaining a target English video, and performing audio analysis on the target English video to extract audio information; Step S200: The extracted audio information is divided into roles through AI audio recognition, and subtitles corresponding to the audio of the divided roles are translated and extracted, and the subtitles of the corresponding segments of the target English video are marked with different color role labels; Step S300: upon receiving an instruction to play the target English video, confirming the user's selection of a designated role for dubbing; Step S400: Control the playback of the target English video based on the confirmed dubbing of the designated character. When the dubbing scene of the designated character is played, mute the audio of the designated character and only display the scene and corresponding subtitles. Also, control the real-time acquisition of the dubbing audio of the current designated character. At the same time, control the normal playback of the audio of other characters. Step S500: Compare the obtained dubbing audio of the currently designated character with the original dubbing audio of the corresponding character in the target English video, and output a comparison score result and improvement suggestions.
[0031] This application proposes acquiring a target English video and analyzing its audio information, using AI recognition to segment the character's voice, and simultaneously generating color-coded subtitles. After the user selects a character to dub, playback automatically mutes the character's original voice, preserving the image and subtitles for the user to dub in real time. The system then collects the user's dubbing data, compares it with the original audio, and outputs a score and improvement suggestions.
[0032] Among them, audio analysis refers to the extraction of spectral features of the audio track of the video file, which can be implemented specifically by the Mel-frequency cepstral coefficient algorithm to distinguish the voice characteristics of different characters. Role segmentation refers to the division of continuous audio into segments of different speakers based on voiceprint features. It can be implemented specifically by using a speaker clustering algorithm based on a Gaussian mixture model to ensure that each character's lines are independently identified. Subtitle color labeling refers to assigning specific colors to the text content corresponding to different characters. It can be implemented by superimposing a color code layer on the video decoding layer, making it easier for users to intuitively distinguish between the subjects of the dialogue. Audio mute control refers to dynamically shielding the output signal of a specified channel for the audio of a specified character. It can be implemented specifically by using digital audio filtering technology to keep the voices of other characters playing normally. Real-time dubbing acquisition refers to obtaining the user's voice stream through the terminal microphone. It can be implemented specifically by using audio cache queue technology to ensure synchronization with the video screen.
[0033] Specifically, the system first analyzes the audio track data of the video file and uses voiceprint recognition technology to segment the dialogue into independent character segments. Each segment generates corresponding subtitle text and adds color labels to distinguish the speakers, such as Figure 2 As shown in the figure below. After the user selects a character to play, the player enters silent mode upon detecting the character's image and simultaneously activates the recording function to capture the user's voiceover. The audio of other characters in the original video continues to play normally to maintain scene coherence. After pre-processing, the user's voiceover data is compared with the original voice of the corresponding character at the phoneme level to evaluate pronunciation accuracy and rhythm matching.
[0034] Compared to existing technologies, which only offer one-way voice input or mechanical reading, this solution builds character interaction scenarios, transforming the learning process into a gamified experience. Traditional methods cannot independently control the voice of a specific character. This solution uses precise voiceprint recognition and audio processing technology to create a personalized practice space while maintaining the integrity of the video. Compared to general pronunciation assessment systems, this solution provides targeted improvement suggestions based on the specific character context, improving the effectiveness of error correction.
[0035] Through the above technical solution, this application realizes scenario-based immersive training for English learning, and users can enhance their language proficiency through role-playing. The system automatically isolates the voice signal of the target practice role, eliminates background conversation interference, and improves the concentration of pronunciation training. The real-time comparative feedback mechanism helps users correct pronunciation deviations in a timely manner, and color-labeled subtitles assist in understanding the conversation context, forming an audio-visual collaborative learning effect. This solution is particularly suitable for children's English enlightenment education, transforming animation viewing into active language practice, significantly improving learning participation and sustainability.
[0036] The present application further proposes that when an instruction to play a target English video is received, a pop-up window is used to prompt the user whether to choose to play a role in the video and perform spoken English dubbing; and the selectable dubbing roles are displayed; the user's operation instruction is received, and the user's selection of the designated dubbing role is confirmed; according to the designated dubbing role selected by the user, the audio of the corresponding designated role in the target English video is muted, and the audio of other characters is normal.
[0037] Among them, pop-up prompts refer to the interactive interface that pops up in the form of an independent window on the video playback interface. Specifically, this can be achieved by using a modal window to cover part of the screen area to ensure that users give priority to character selection operations. User selection of dubbing roles refers to selecting the target role from a candidate list through touch clicks or voice commands. Specifically, this can be achieved by using button controls to arrange the character portraits and names, which is convenient for users to intuitively identify and quickly operate. Setting audio muting refers to silencing the original audio track of the selected character. Specifically, this can be achieved by separating multi-track audio and adjusting the gain value of the specified audio track to zero, while retaining the normal playback status of the remaining audio tracks.
[0038] Specifically, when the user starts the target video, the system immediately triggers a pop-up interface, presenting the option of entering role-playing mode. If the user confirms to enter, a list of all characters in the video that can be dubbed is displayed, such as arranging character images and names in the form of horizontally sliding cards. After the user completes the selection by clicking on the target character card, the player automatically mutes the original dialogue clips of the character, while keeping the background music and other character voices audible. For example, in an animated scene, if the user chooses to play the protagonist, only the lip animation and subtitles are displayed when the protagonist's lines appear. The user needs to record his or her own dubbing in real time, while other supporting character dialogues continue to play normally.
[0039] Compared with existing technologies, traditional English learning software lacks interactive role-playing mechanisms, requiring users to manually pause the video and find lines to follow along, resulting in a fragmented learning process. This solution actively guides users into role-playing mode through pop-up windows and automatically handles track switching during playback, eliminating the need for users to intervene in technical details and making the learning process more coherent. Existing systems that support multi-track control usually require users to adjust parameters in complex settings menus, but this solution automatically links character selection with track control, significantly lowering the operational barrier.
[0040] Through the above technical solution, this application realizes the rapid activation and precise control of role-playing functions, solving the problems of traditional English learning software with a single interactive mode and cumbersome operation steps. After the user directly selects the target character through the visual interface, the system automatically adapts to the silent area and retains the ambient sound effects, making the dubbing practice closer to the real dialogue scene. For example, when children practice the lines of animated characters, they can still hear the response sentences of other characters, thereby maintaining the context of the dialogue and enhancing the immersion of language input.
[0041] This application further proposes to compare the obtained dubbing audio of the currently specified character with the original dubbing audio of the corresponding character in the target English video, detect whether the dubbing audio of the currently specified character is correct, and output the comparison score results and improvement suggestions for further practice, and regularly remind users to practice repeatedly according to preset memory rules.
[0042] Among them, the comparative scoring result refers to the quantitative evaluation of the pronunciation accuracy of the user's dubbing and the original voice through speech recognition technology. Specifically, it can be achieved by calculating the phoneme matching degree using a dynamic time warping algorithm, which is used to objectively reflect the degree of difference between the user's pronunciation and the standard pronunciation. Improvement suggestions refer to targeted practice content generated based on pronunciation difference analysis. Specifically, it can be achieved by extracting the phonetic features corresponding to incorrect pronunciations through natural language processing technology and associating them with a preset teaching resource library, which is used to guide users to strengthen training on weak links. Memory rules refer to a periodic practice reminder mechanism set based on the Ebbinghaus forgetting curve. Specifically, it can be achieved by using a time-triggered message push module, which is used to arrange review plans according to memory rules to consolidate learning effects.
[0043] Specifically, after the user completes the dubbing of the role, the system automatically compares the real-time recording file with the original sound frame by frame for acoustic features, and identifies words or phrases whose pronunciation deviation exceeds the threshold. For detected pronunciation errors, the system matches demonstration video clips containing the same phonemes from the preset resource library, such as pronunciation lip shape demonstration animations or tongue position diagrams. At the same time, the system sets the practice cycle according to the forgetting curve model. For example, it automatically pushes review reminders on the 1st, 3rd, and 7th days after the first practice, and generates special practice tasks on the interface that contain a summary of historical incorrect pronunciations.
[0044] Compared to existing technologies, traditional English learning systems only provide instant pronunciation scoring and lack a continuous improvement mechanism, failing to effectively address the degradation of learning outcomes caused by the user's forgetting curve. Existing shadowing training models often use fixed-interval reminders and fail to tailor content recommendations based on individual pronunciation deficiencies. This solution, by establishing a database of pronunciation errors and linking it with a memory model, enables personalized practice planning based on the user's actual learning progress.
[0045] Through the above technical solution, this application can effectively improve the accuracy of user pronunciation correction and provide precise learning guidance by associating incorrect pronunciations with specialized practice resources. The periodic reminder mechanism can maintain the continuity of user learning, avoiding the phenomenon of pronunciation habits regressing due to practice interruptions in traditional methods, and forming a virtuous cycle system for pronunciation correction.
[0046] This application further proposes that when it is detected that the pronunciation of the dubbing audio of the currently specified character is not standard compared with the original dubbing audio of the corresponding character in the target English video, AI is used to automatically search for the pronunciation guidance practice video corresponding to the audio content with non-standard pronunciation, and output it for playback for the user to perform improvement practice.
[0047] Non-standard audio content refers to content where the user's dubbing differs from the original in terms of phonemes, stress, or intonation. This can be achieved by using a speech recognition engine to compare acoustic features, such as extracting spectral features using Mel-frequency cepstral coefficients and then performing dynamic time warping. Pronunciation guidance videos refer to demonstration teaching clips tailored to specific pronunciation issues. These can be achieved using standardized teaching materials pre-stored in the video resource library, such as animated tutorials that include lip movements and pronunciation breakdown steps.
[0048] Specifically, when the voice comparison module detects that the user's pronunciation deviates from the original audio track in a specific syllable, it triggers the keyword extraction mechanism to analyze the word or phoneme corresponding to the incorrect pronunciation, and then calls the retrieval interface to match the associated pronunciation guidance content in the pre-stored teaching resource library. For example, if the user mispronounces "th", the system will automatically call a demonstration video containing the pronunciation skills of the consonant combination and push it to the user interface through an independent playback window. This process uses asynchronous loading to ensure that the call of video resources does not affect the playback smoothness of the original video.
[0049] Compared to existing technologies, traditional English learning systems only provide pronunciation scoring but lack targeted error correction solutions, forcing users to search for practice materials on their own. This solution intelligently links incorrect pronunciations with teaching resources, providing instant error correction guidance and effectively shortening the user's journey from error discovery to solution.
[0050] Through the above technical solution, this application can provide users with precise practice resources targeting their pronunciation weaknesses, avoiding the inefficiency caused by blind practice in traditional methods and significantly improving the effectiveness of pronunciation correction. At the same time, the automated resource matching mechanism reduces the user's operational burden of searching for learning materials, which helps maintain the continuity of the learning process.
[0051] The present application further proposes that when an instruction to play a target English video is received, the target English video with divided roles is classified according to the user's age, the currently dubbed roles are displayed, and the user's operation instruction is received to select multiple designated dubbing roles.
[0052] Among them, user age classification refers to matching video content with user cognitive levels through age recognition algorithms. Specifically, user registration information or age data input on the device can be used for classification to ensure that the role division is consistent with the user's language learning stage. Displaying the currently available dubbing roles refers to dynamically presenting a list of roles suitable for the user's age group through an interactive interface. Specifically, this can be achieved through a sliding menu or a layered display interface, allowing users to quickly locate the target role. Selecting multiple designated roles for dubbing means allowing multiple role switching during a single learning process. Specifically, this can be achieved through check boxes or multi-touch operations, realizing the construction of complex oral training scenarios.
[0053] Specifically, after obtaining the video role classification data, the system first uses the user age parameter to filter the video library. For example, for users aged 5-8, simple dialogue characters in cartoons can be displayed first; for users aged 9-12, characters with complex sentence patterns in situation dramas can be displayed first; Figure 3 As shown, the left side displays optional user age parameters of 2, 3, 4, 5, and 6 years old. When the age parameter on the left is selected, the right interface will display the video animation corresponding to the selected age.
[0054] In this application, when a playback command is triggered, the interface generates a dynamic character selection panel that only displays appropriate character options based on age classification results. Users can mark multiple target characters at the same time by long-pressing or selecting an area. The system generates a corresponding dubbing task sequence based on the multi-select command. When the video plays to a dialogue clip of any selected character, the system automatically switches to silent mode and activates the recording function, enabling multi-character alternation training.
[0055] Compared to existing technologies, traditional English learning systems typically use a single character follow-up model, failing to tailor training content to learners' age characteristics and lacking multi-character interaction mechanisms. Existing video learning tools often neglect age-appropriate character selection, resulting in younger users being exposed to complex dialogues beyond their cognitive level, while older users may be forced to repeat basic content. This solution, through age-specific content filtering and a multi-character selection mechanism, ensures the appropriateness of learning materials while increasing the diversity of training dimensions.
[0056] Through the above technical solutions, this application realizes age-appropriate content push and composite role training, solving the problem of mismatch between content and user cognitive level in traditional English video learning. Through the multi-role selection mechanism, users can experience different dialogue scenarios in a single learning process, enhancing the multi-dimensional training of language application ability. The age classification mechanism effectively avoids the situation where learning content is too difficult or too easy, so that users of different age groups can obtain matching language training materials.
[0057] The present application further proposes to divide the extracted audio information into roles through AI audio recognition, and at the same time translate and extract subtitles corresponding to the audio of the divided roles, and mark the subtitles of the corresponding clips of the target English video with role labels of different colors. Specifically, it includes: obtaining the extracted audio information, and using AI audio recognition to identify the audio information of different people, dividing the roles according to the identified audio information of different people, and dividing the audio into audio of different roles; translating the audio of different roles into corresponding subtitles; extracting subtitles corresponding to the audio of the divided roles, and marking the subtitles of the corresponding roles on the corresponding clip screen of the target English video, and marking role labels of different colors on the subtitles of the corresponding clip screen.
[0058] Among them, AI audio recognition refers to the technology of using machine learning algorithms to analyze the voiceprint features in audio to distinguish different speakers. Specifically, it can be implemented by using a voiceprint recognition model based on deep neural networks, and the speaker identity is classified by extracting the spectral features of the audio. Role division refers to the classification of segments of different speakers into independent roles based on the voiceprint features in the audio. Specifically, a clustering algorithm can be used to classify audio segments with similar voiceprint features as the same role. Translated subtitles refer to the conversion of English dialogues in audio into text information in the target language. Specifically, it can be implemented by using a machine translation model based on a neural network, and the text after speech recognition is translated into a preset language in real time. Different color role labels refer to the assignment of a unique color identifier to each role. Specifically, this can be achieved by embedding color coding in the font color or background color of the subtitle text, so that the subtitles of the same role are visually consistent.
[0059] Specifically, the audio information extracted from the target English video is preprocessed and input into the AI audio recognition model. The model identifies the differences in voiceprints of different speakers by analyzing the spectral characteristics of the audio signal and divides the continuous audio stream into independent segments belonging to different characters. For each segmented audio segment, speech recognition technology is used to convert it into text, and subtitles in the target language are generated through the translation model. When the subtitles are displayed synchronously with the video screen, a specific color is assigned to the subtitles of each character according to the preset color mapping rules. For example, the subtitles of character A are displayed in red, and the subtitles of character B are displayed in blue. During video playback, when a character starts speaking, the corresponding subtitles will be dynamically superimposed on the bottom of the screen in the color exclusive to that character, while the subtitles of other characters will be displayed in different colors.
[0060] Compared to existing technologies, traditional English learning videos typically only provide single-color subtitles or fail to differentiate between different characters, making it difficult for users to quickly distinguish the conversations between different characters. While some existing systems support subtitle generation, they lack dynamic character-based annotation capabilities and are unable to use visual cues to help users track the language expressions of specific characters. This solution uses AI-driven character segmentation and color labeling technology to achieve accurate subtitle differentiation in multi-character dialogue scenarios, significantly reducing the cognitive burden on users to understand complex conversations.
[0061] Through the above technical solution, this application solves the problem of low learning efficiency caused by confusing character dialogues in traditional English learning videos. The color-coded subtitles help users quickly associate characters with corresponding language content, improving learning focus in multi-character scenarios. At the same time, dynamically generated character label subtitles enhance the readability of video content, allowing users to intuitively distinguish the pronunciation characteristics and language styles of different characters, providing a clear reference for subsequent dubbing practice.
[0062] This application further proposes an English video character learning and processing method based on a smart terminal. After the user selects a designated character for dubbing, the target English video starts to play; when the dubbing screen of the designated character is played, the dubbing screen of the character is displayed and its audio is muted, and the corresponding subtitles marked with different color character labels are displayed at the same time, and the audio of other characters continues to play normally.
[0063] Audio muting silences the original audio signal of a specific character in the video. This can be achieved through audio signal filtering or volume zeroing, providing users with an environment for independent dubbing. Color-coded character labels assign differentiated color identifiers to subtitle text based on the character classification results. This can be achieved through RGB color coding or preset color templates. This allows users to intuitively distinguish the lines corresponding to different characters, allowing users to quickly identify the clips currently requiring dubbing.
[0064] Specifically, after the user selects a character, the video player will detect in real time during operation whether the current playback screen contains the dubbing content of that character. For example, in an animated scene, if the user chooses to dub character A, when character A's lines appear, the system automatically blocks the original audio and displays subtitles with red labels. At this time, the user can record their own dubbing into the microphone. At the same time, the audio of the dialogue between characters B and C is still output normally to ensure the continuity of the video context. The color label of the subtitle is bound to the character, for example, red for character A and blue for character B, which helps users determine the current paragraph that needs dubbing through visual cues.
[0065] Compared to existing technologies, traditional English learning systems typically use global muting or single-character replacement modes, making it impossible to dynamically separate and control the audio of multiple characters. For example, if a user chooses to voice a character, the audio of other characters may be completely muted, resulting in a distorted learning scene. However, this solution accurately identifies specific character images while retaining the audio of other characters. This maintains the real context of the video and strengthens the user's focus on the target segment through color labeling, making the role-playing process more interactive and immersive.
[0066] Through the above technical solution, this application can precisely control the timing of muting a specific character without interrupting the overall video playback, and use color-coded subtitles to prompt users to practice dubbing in a timely manner. The continued voice of other characters can help users maintain the rhythm of the conversation and avoid the loss of context caused by complete silence, thereby improving the accuracy of English oral imitation and the ability to adapt to different scenarios.
[0067] The present invention is further described in detail below through another specific application embodiment.
[0068] Example 2 like Figure 4 As shown, the second embodiment provides an English video character learning processing method based on a smart terminal, including: S10: Use AI to extract audio from English cartoons, divide them by characters, and extract subtitles at the same time, and then enter S11 or S21; S11, the user chooses to play the English cartoon normally and enters S12; S12, play to a specific clip, ask the user to use English chat to perform just the plot, and then enter S23; S21. The user chooses to enter the English speaking area and proceeds to S22; S22, choose the appropriate age and corresponding script, such as Figure 3 As shown, then enter S23; S23, receiving the user's operation instruction to select the favorite role, such as Figure 5 As shown, the user selects the first character to learn dubbing and enters S24; S24, the English animation plays normally. When the character that requires the user to speak is played, it is paused and the process goes to S25; S25, obtaining the corresponding spoken language read by the user following the text subtitles of the smart terminal such as the TV, and proceeding to S26; S26, the animation continues to play and enters S27; S27, determine whether the current playback is finished, and then enter S28; S28: Score the user's speech and conduct intensive practice. Then end.
[0069] In this specific application embodiment, it is possible to realize the platform's children's cartoons with English pronunciation, such as "Peppa Pig", by means of AI, extracting audio suitable for learning, translating the corresponding subtitles, and then labeling the clips on the original film. When the child watches Peppa Pig normally, after the corresponding marked clip is played, the user will be asked whether to play the role in it and speak English. The user selects a role, such as Peppa, and the system will mute Peppa's voice in this clip, leaving only subtitles. At the same time, the voice of another character, such as Peppa's father, will be normal. In this way, after Peppa Pig's father's voice is played, it is the child's turn to speak along with the subtitles. AI will detect whether the spoken language is correct, and thus determine whether it is necessary to continue practicing or watch the subsequent content normally. Among them, the interface is triggered during the movie, for example Figure 5 As shown, users can select the role they need to practice dubbing. The example interface for starting a conversation is as follows Figure 2 As shown, users can follow the reading progress according to the subtitles marked with the character color. Figure 3 As shown, users can select video animations for different age groups according to their age.
[0070] The specific application embodiment of the present invention uses children's favorite cartoons in a "role-playing" manner to restore children's mentality of playing house, and packages English learning as a game to enhance children's learning interest.
[0071] Exemplary devices like Figure 6 As shown, an embodiment of the present invention provides an English video character learning and processing device based on an intelligent terminal, the device comprising: An acquisition module 310 is configured to acquire a target English video, perform audio analysis on the target English video, and extract audio information; The role classification module 320 is used to classify the extracted audio information into roles through AI audio recognition, translate and extract subtitles corresponding to the audio of the divided roles, and mark the subtitles of the corresponding segments of the target English video with different color role labels; The dubbing role selection module 330 is used to confirm the user's selection of a designated role for dubbing upon receiving an instruction to play the target English video; The dubbing control module 340 is used to control the playback of the target English video according to the confirmed dubbing of the designated character. When the dubbing scene of the designated character is played, the audio of the designated character is muted and only the picture and corresponding subtitles are displayed. The dubbing audio of the current designated character is obtained in real time. At the same time, the audio of other characters is played normally. The result output and suggestion module 350 is used to compare the obtained dubbing audio of the currently specified character with the original dubbing audio of the corresponding character in the target English video, and output the comparison score result and improvement suggestions.
[0072] The acquisition module is a unit that extracts the target English video from a storage device or network resource and performs audio analysis. Specifically, it can be implemented using an audio decoder and signal processor to separate the audio track in the video into a processable digital signal. The role classification module is a unit that distinguishes different speakers based on a voiceprint recognition algorithm. For example, it uses a convolutional neural network to extract voice features and perform cluster analysis to implement role classification and generate color-coded subtitle data. The dubbing role selection module is an interactive unit that responds to user input and determines the roles involved in dubbing. It can use a touch screen interface combined with button controls to implement role list display and selection functions. The dubbing control module is a logic unit that dynamically adjusts the playback status of the audio track based on the selected role. For example, it uses an audio routing controller to mute the specified role's channel and activate the microphone to collect the user's dubbing audio. The result output and suggestion module is an algorithmic unit that performs voice comparison and generates feedback. Specifically, it can use a dynamic time warping algorithm to calculate pronunciation similarity and generate improvement suggestions based on a preset rule library.
[0073] Specifically, the device pre-processes the original video through the acquisition module to separate the analyzable audio stream. The role division module uses voiceprint feature extraction technology to identify the voice segments of different characters, and generates multi-language subtitles and adds visual identification. When the user selects a dubbing role, the dubbing control module switches the audio channel status in real time during video playback, silencing the original voice of the specified character and activating the recording function, while the dialogues of other characters continue to play normally. The audio data of the user's dubbing is synchronously transmitted to the result output and suggestion module. After the timeline is aligned with the original audio track, the pronunciation accuracy, intonation and rhythm differences are analyzed by the speech recognition engine, and finally a visual report containing scores and practice suggestions is generated.
[0074] Compared to existing technologies, traditional English learning devices only offer one-way voice-following functionality, failing to dynamically integrate video clips with role-playing. This device, through character segmentation and audio track control technology, allows users to precisely locate specific characters' dialogue segments for imitation practice while watching a video, while retaining the original voices of other characters to maintain contextual integrity. While existing systems' feedback mechanisms are often limited to simple accuracy statistics, this device employs a multi-level voice comparison algorithm that identifies deviations in pronunciation details and associates them with targeted practice resources.
[0075] Through the above technical solutions, this application effectively addresses the poor interactivity and limited feedback of traditional English learning tools, enabling users to improve their speaking skills in simulated real-life conversations. The role-based learning model enhances the focus and fun of practice. The multi-dimensional speech analysis function helps users quickly identify pronunciation weaknesses. The color-coded subtitles and mute switching mechanism work together to reduce learning distractions in multi-role conversation scenarios.
[0076] Based on the above embodiment, the present invention also provides an intelligent terminal, whose principle block diagram can be shown as follows: Figure 7 The intelligent terminal includes a processor, a memory, a network interface, a display screen, and a database connected via a system bus.
[0077] An intelligent terminal of the present application further includes one or more programs, wherein one or more programs are stored in a memory and configured to be executed by one or more processors. The one or more programs include a method for executing an English video role learning method based on the intelligent terminal.
[0078] The memory refers to a storage device used to store the target English video and the audio information, role division data, subtitle label information, and program code generated during its processing. Specifically, it can be implemented using a solid-state drive or flash memory chip. Its function is to provide the terminal with persistent data storage capabilities to achieve continuous operation of the video processing process. The processor refers to an arithmetic unit that executes program instructions to control video playback, audio analysis, role division, and user interaction functions. Specifically, it can be implemented using a multi-core central processing unit or graphics processing unit. Its function is to realize dynamic role allocation, real-time audio acquisition, and comparative analysis functions through program drive.
[0079] Specifically, after the program is loaded into the processor, it first performs audio separation and character recognition on the target English video, generating a color-coded subtitle stream for multiple characters. When it detects that a user has selected a specific character for dubbing, the program automatically suppresses that character's original voice and activates the microphone to capture the user's dubbing audio. During playback, the program simultaneously performs phoneme matching analysis between the user's dubbing and the original voice, generating a pronunciation accuracy score and practice suggestions. The program further sets practice reminders based on a memory curve algorithm, and pushes repeated training tasks through the system notification module.
[0080] In some embodiments, the memory can be divided into a video buffer area and a user data area. The video buffer area is used to temporarily store the video clip being processed and the subtitle stream generated in real time, while the user data area is used to store the user's historical dubbing records and personalized learning progress data. When executing the program, the processor can use parallel threads to handle video decoding and audio analysis tasks to optimize the speed of real-time interactive response.
[0081] Compared to existing technologies, existing English learning terminals only support static sentence reading or fixed dialogue modes, and cannot achieve immersive learning based on dynamic switching of video characters. This solution uses programmatic control of video playback and audio processing to enable the terminal to dynamically adjust the playback mode based on the character selected by the user. This creates a personalized dubbing space while retaining the original voices of other characters, addressing the technical shortcomings of traditional learning methods such as poor interactivity and a single scenario.
[0082] Through the above technical solution, this application enables smart terminals to transform the English learning process into a role-playing interactive experience. Through real-time comparison of the original voice and dubbing and periodic practice reminders, it helps users systematically improve pronunciation accuracy while maintaining their interest in learning, overcoming the problem of distraction caused by the mechanical repetition of traditional follow-up exercises.
[0083] The present application further proposes a computer-readable storage medium. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform an English video role learning method based on a smart terminal.
[0084] A computer-readable storage medium refers to a non-temporary storage medium used to store program instructions. Specifically, this can be implemented using devices such as solid-state drives, USB flash drives, or flash memory cards. Its purpose is to ensure that program instructions are correctly read by the processor when the electronic device is powered on. Processor execution of instructions refers to the parsing and calculation of the code in the storage medium by the electronic device's computing unit. Specifically, this can be implemented using a multi-core processor or embedded chip. Its purpose is to drive the electronic device to perform video analysis, character segmentation, and voice comparison functions according to preset logic.
[0085] Specifically, once the program instructions from the storage medium are loaded into memory, the processor executes the following steps according to the code logic: First, the target English video is separated from the audio track and the roles are divided, generating color-coded subtitle data. After the user selects a dubbing role, the processor automatically mutes the original voice of the designated role and collects the user's dubbing data in real time. A speech waveform comparison algorithm is used to calculate the pronunciation accuracy score, and based on pre-set rules, improvement suggestions are generated. During this process, the processor synchronizes the operating timing of the video playback, audio acquisition, and data analysis modules.
[0086] Compared to existing technologies, traditional English learning software only enables follow-up training on fixed sentences, but is unable to replace and compare character voices in dynamic videos in real time. Programs stored in existing storage media often lack the logic for character division and dubbing interaction, making it difficult to achieve immersive learning based on video scenes. This solution optimizes the arrangement of program instructions, enabling the code carried by the storage medium to drive the device to complete the complex functions of dynamic character recognition, real-time dubbing replacement, and accurate pronunciation assessment.
[0087] Through the above technical solution, this application realizes role-playing learning based on video clips. By combining silent original sound with real-time dubbing, it effectively enhances the fun of language practice. The execution of the voice comparison algorithm enables users to obtain instant feedback on pronunciation correction. The automated processing of program instructions reduces manual switching operations and ensures that the learning process is synchronized with the rhythm of video playback. The memory reminder mechanism preset in the storage medium can regularly trigger practice tasks to strengthen the long-term memory effect of the learning content.
[0088] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, various modifications and variations of the present application are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A method for learning English video characters based on an intelligent terminal, characterized in that: include: Obtaining a target English video, and performing audio analysis on the target English video to extract audio information; The extracted audio information is divided into roles through AI audio recognition, and subtitles corresponding to the audio of the divided roles are translated and extracted, and the subtitles of the corresponding clips of the target English video are marked with different color role labels; Upon receiving an instruction to play the target English video, confirming the user's selection of a designated role for dubbing; Control the playback of the target English video according to the confirmed dubbing of the designated role. When the dubbing screen of the designated role is played, the audio of the designated role is muted and only the screen and corresponding subtitles are displayed. The dubbing audio of the current designated role is obtained in real time. At the same time, the audio of other roles is controlled to play normally. The obtained dubbing audio of the currently specified character is compared with the original dubbing audio of the corresponding character in the target English video, and a comparison score result and improvement suggestions are output.
2. The method for learning English video characters based on an intelligent terminal according to claim 1, wherein: The step of confirming the user's selection of a designated role for dubbing upon receiving the instruction to play the target English video further includes: When receiving the instruction to play the target English video, a pop-up window prompts the user whether to choose to play the role in it and perform English spoken dubbing; and displays the selectable dubbing roles; Receive the user's operation instruction and confirm the user's choice of the designated role for dubbing; According to the designated role of dubbing selected by the user, the audio of the target English video corresponding to the designated role is muted, and the audio of other roles is normal.
3. The method for learning English video characters based on an intelligent terminal according to claim 1, wherein: The step of comparing the obtained dubbing audio of the currently designated character with the original dubbing audio of the character corresponding to the target English video, and outputting a comparison score result and improvement suggestions includes: Compare the obtained dubbing audio of the currently specified character with the original dubbing audio of the corresponding character in the target English video to detect whether the dubbing audio of the currently specified character is correct and output a comparison score result and improvement suggestions for further practice; And according to the preset memory rules, users are regularly reminded to practice repeatedly.
4. The method for learning English video characters based on an intelligent terminal according to claim 3, wherein: The steps of detecting whether the dubbing audio of the currently specified character is correct and outputting a comparison score result and improvement suggestions for further practice include: When it is detected that the pronunciation of the dubbing audio of the currently specified character is not standard compared with the original dubbing audio of the corresponding character in the target English video, AI will automatically search for the pronunciation guidance practice video corresponding to the audio content with non-standard pronunciation, and output it for playback for users to practice improvement.
5. The method for learning English video characters based on an intelligent terminal according to claim 1, wherein: The step of confirming the user's selection of a designated role for dubbing upon receiving the instruction to play the target English video further includes: Categorize the target English videos for role division according to user age; When receiving the instruction to play the target English video, the currently available dubbing characters are displayed; Receive a user's operation instruction to select multiple designated roles for dubbing.
6. The method for learning English video characters based on an intelligent terminal according to claim 1, wherein: The steps of dividing the extracted audio information into roles through AI audio recognition, translating and extracting subtitles corresponding to the audio of the divided roles, and marking the subtitles of the corresponding segments of the target English video with different color role labels include: Get the extracted audio information and use AI audio recognition to identify the audio information of different people. Divide the audio information of different people into roles based on the identified audio information, and divide the audio into different roles; Translate the audio of different characters into corresponding subtitles; Subtitles corresponding to the audio of the divided roles are extracted, and subtitles corresponding to the roles are marked on the corresponding fragments of the target English video, and role labels of different colors are marked on the subtitles of the corresponding fragments.
7. The method for learning English video characters based on an intelligent terminal according to claim 1, wherein: The target English video is played according to the designated role confirmed for dubbing. When the dubbing screen of the designated role is played, the audio of the designated role is muted and only the screen and corresponding subtitles are displayed. The dubbing audio of the current designated role is obtained in real time. The steps to control the normal playback of audio for other characters include: Control the start of playing the target English video according to the designated role of the dubbing selected by the user; When the dubbing screen of a specified character is played, the dubbing screen of the specified character is controlled to be displayed, the audio of the specified character is controlled to be muted, and the corresponding subtitles marked with different color character labels are displayed; and the audio of other characters is controlled to play normally.
8. An English video role learning and processing device based on an intelligent terminal, characterized in that: The device comprises: An acquisition module is used to acquire a target English video, perform audio analysis on the target English video, and extract audio information; A role classification module is used to classify the extracted audio information into roles through AI audio recognition, translate and extract subtitles corresponding to the audio of the divided roles, and mark the subtitles of the corresponding segments of the target English video with different color role labels; A dubbing role selection module is used to confirm the user's selection of a designated dubbing role upon receiving an instruction to play a target English video; The dubbing control module is used to control the playback of the target English video according to the designated role confirmed for dubbing. When the dubbing screen of the designated role is played, the audio of the designated role is muted and only the screen and corresponding subtitles are displayed. The dubbing audio of the current designated role is obtained in real time. At the same time, the audio of other roles is controlled to play normally. The result output and suggestion module is used to compare the obtained dubbing audio of the currently specified role with the original dubbing audio of the corresponding role in the target English video, and output the comparison score results and improvement suggestions.
9. An intelligent terminal, characterized in that: The device comprises a memory and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by one or more processors, and the one or more programs include being used to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Generation method of to-be-dubbing video, computer equipment and storage medium
CN110166818A
Language learning method and system based on video dubbing and pronunciation correction training
CN111462553A
Relationship graph creation method and device, terminal and storage medium
CN117421425A
Automatic dubbing method and system based on generative AI
CN119314488A
Language learning system using video clips for role playing
TWM447567U