Personalized pronunciation practice method and system
By designing a personalized pronunciation practice system, using multiple voice processing modules to analyze user pronunciation characteristics and generate targeted exercise content, the problem that existing systems are difficult to personalize exercises is solved, and efficient and intelligent pronunciation practice effects are achieved.
Patent Information
- Application Number
- CN202510165003.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-13
AI Technical Summary
The existing pronunciation practice system is difficult to effectively practice the personalized pronunciation characteristics of different users, and lacks attention to the pronunciation situation in the daily real context of users.
A personalized pronunciation exercise system was designed to collect user voice data through intelligent terminals, combine with multiple voice processing modules on the server side (such as voice registration, speaker recognition, pronunciation quality detection, speech recognition, large language model and text to speech), analyze the user's pronunciation characteristics, generate targeted practice content, and generate reference pronunciation audio consistent with the user's tone through the text to speech module.
It realizes highly personalized pronunciation exercises, can deeply understand the user's pronunciation weaknesses and provide targeted practice content, improves the practice efficiency and effect, and has the advantages of intelligent analysis and evaluation, natural integration into daily life and timbre consistency exercises.
Smart Images

Figure CN119993133A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech processing, and in particular to a personalized pronunciation practice method and system. Background Art
[0002] Currently, there are a variety of pronunciation training methods and systems on the market, which aim to help users improve their speech expression ability. These methods and systems include traditional manual guidance, recording playback comparison, and the use of electronic devices to assist in training. However, existing pronunciation training programs still have certain limitations in terms of personalization and intelligence.
[0003] Traditional manual guidance methods rely on experience, are difficult to apply on a large scale, and have limited personalization.
[0004] The recording playback comparison method lacks objective evaluation criteria, making it difficult for users to accurately judge their own pronunciation problems.
[0005] Although existing electronic device-assisted practice systems have certain speech analysis functions, their practice content is often preset and universal, making it difficult to effectively practice according to the personalized pronunciation characteristics of different users. In addition, existing systems pay little attention to the pronunciation of users in real daily contexts, and the practice content may be out of touch with the actual needs of users. Summary of the invention
[0006] The present invention aims to solve the problem in the prior art that it is difficult to effectively practice according to the personalized pronunciation characteristics of different users, and provides a personalized pronunciation practice method and system, so as to more effectively help users discover and improve their own pronunciation weaknesses and enhance their voice expression ability.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] A personalized pronunciation practice system, comprising:
[0009] The intelligent terminal is used to collect the user's voice data, receive and present the pronunciation practice text and voice generated by the server, and receive the user's operation instructions;
[0010] The server includes a voice registration module, a voice data storage module, a voice activity detection module, a speaker recognition module, a pronunciation quality detection module, a voice recognition module, a large language model module and a text-to-speech module;
[0011] The voice registration module is used to register the user's voice and provide the user's voice features to the speaker recognition module and the text-to-speech module;
[0012] Voice data storage module, used to store user voice data uploaded by smart terminals;
[0013] The voice activity detection module is used to detect the voice activity segments in the voice data, determine the valid voice segments, and send them to the speaker recognition module;
[0014] Speaker recognition module, used to identify the speaker in the voice data and confirm whether it is the registered target user;
[0015] The pronunciation quality detection module is used to evaluate the pronunciation quality of the user's voice and obtain the weak pronunciation segments;
[0016] A speech recognition module, used to convert user speech data into text content and associate weak pronunciation segments with text content;
[0017] The large language model module is used to generate user-specific pronunciation practice texts based on the texts associated with the weak pronunciation segments;
[0018] The text-to-speech module is used to convert the pronunciation practice text generated by the large language model into reference pronunciation audio that is consistent with the user's timbre.
[0019] To optimize the above technical solutions, the specific measures taken also include:
[0020] Furthermore, the speaker recognition module is specifically implemented by a CAM++ speaker recognition model.
[0021] Furthermore, the method of evaluating the pronunciation quality of the user's voice is specifically as follows:
[0022] Detect repetition, dragging and blocking in the user's voice; output the speech feature category sequence C = [c1, c2, ..., c i ,…,c T ], where c i ∈{R,B,P,F}, and the time interval of the weak pronunciation segment, c i represents the speech feature category of the i-th frame of audio, R represents repetition, B represents obstruction, P represents dragging sound, and F represents fluency; the weak pronunciation segments include repetition segments, obstruction segments and dragging sound segments.
[0023] Furthermore, the step of associating the weakly pronounced segment with the text content specifically includes:
[0024] The speech recognition module outputs the text transcription results and the timestamp information corresponding to each word, which is expressed as the time interval at the word level as follows:
[0025]
[0026] Among them word i represents the i-th word, Indicates the start and end time of the word in the audio;
[0027] Based on the time interval of the weak segment obtained by pronunciation quality detection and the word timestamp information obtained by speech recognition, time interval matching is performed to determine the text content corresponding to the weak pronunciation segment. Specifically, for each time interval of the weak pronunciation segment Indicates the start time of the weak pronunciation segment, Indicates the end time of the weak pronunciation segment, traverses the word timestamp information, finds out the words in all word-level time intervals that overlap with the time interval of the weak pronunciation segment, and concatenates these words to obtain the text associated with the weak pronunciation segment.
[0028] Furthermore, the generation of user-specific pronunciation practice text based on the text associated with the weak pronunciation segment is specifically as follows:
[0029] Constructing a prompt word template, wherein the content of the prompt word template includes a user's weak pronunciation text WT, an instruction, a homophone and easily confused word list CWL, and a practice text type indication PTI; the user's weak pronunciation text is a text associated with a weak pronunciation segment, the instruction is used to indicate the task that the large language model needs to perform, the homophone and easily confused word list includes words that have the same or similar pronunciation as the user's weak pronunciation text WT but different semantics, and words that are easily confused with the user's weak pronunciation text WT; the practice text type indication is used to indicate the type of practice text generated by the large language model;
[0030] The constructed prompt word template is filled with specific parameter values to form a complete prompt word, wherein the parameters include the user's weak pronunciation text, a list of homophones and easily confused words, and an indication of the type of practice text; the filled complete prompt word is input into the large language model, and the large language model generates a pronunciation practice text according to the indication of the prompt word.
[0031] Furthermore, the method for constructing the list of homophones and easily confused words is specifically as follows:
[0032] Pre-build a vocabulary containing commonly used Chinese characters, words and their homophones, and easily confused characters / words. The vocabulary is organized and indexed based on dimensions such as pronunciation similarity, glyph similarity, and semantic relevance;
[0033] After obtaining the user's weakly pronounced text WT, a dynamic search is performed based on the user's weakly pronounced text to search for homophones with the same or similar pronunciation as the user's weakly pronounced text, as well as words that are easily confused with the user's weakly pronounced text, from a preset word library or online dictionary or knowledge base;
[0034] The dynamic retrieval strategy includes: searching for words whose pinyin is exactly the same or partially similar to the user's weakly pronounced text, calculating the similarity between word syllables, and determining words with a similarity higher than a set threshold as words that are easily confused with the user's weakly pronounced text.
[0035] The present invention also proposes a personalized pronunciation practice method based on the above system, comprising:
[0036] The intelligent terminal obtains pronunciation practice data from the server, including pronunciation practice text and corresponding reference audio with the same timbre as the user, and displays it to the user;
[0037] The user operates on the practice interface of the smart terminal and chooses whether to play the reference pronunciation audio, start recording personal practice voice, or switch practice data;
[0038] If the user chooses to play the reference pronunciation audio, the smart terminal plays the reference pronunciation audio, and the user listens to the demonstration pronunciation and practices;
[0039] If the user chooses not to play the reference pronunciation audio, determining whether the user chooses to start recording personal practice voice;
[0040] If the user chooses to start recording, the user recites the pronunciation practice text on the practice interface, and at the same time, the user's practice voice is recorded, and after the recording is completed, the user's practice voice data is sent to the server;
[0041] If the user chooses not to start recording, then determining whether the user chooses to switch to new practice data;
[0042] If the user chooses to switch practice data, the smart terminal obtains new pronunciation practice data from the server and starts a new round of practice;
[0043] After receiving the user practice voice data uploaded by the smart terminal, the server verifies the identity of the speaker and determines whether the practice voice comes from a registered target user; if the speaker is not a target user, the practice process ends; if the speaker is a registered target user, the pronunciation quality of the voice recorded by the user in the practice mode is evaluated to determine the weakly pronounced voice segments and texts;
[0044] The server feeds back the pronunciation quality test result to the smart terminal, and the smart terminal displays the pronunciation quality test result on the practice interface.
[0045] The beneficial effects of the present invention are:
[0046] Highly personalized: The system analyzes the user's daily conversation voice data and can deeply understand the user's personalized pronunciation characteristics and weaknesses, thereby generating practice content that is truly targeted at the user's own problems, achieving highly personalized pronunciation practice, and significantly improving practice efficiency and effectiveness.
[0047] Intelligent analysis and evaluation: The system integrates VAD, speaker recognition, pronunciation quality detection, ASR, TTS, LLM and other advanced voice processing and artificial intelligence technologies to realize intelligent analysis and evaluation of user voice data. It can objectively and accurately identify user pronunciation problems, and provide intelligent practice content generation and feedback, providing users with more professional and effective pronunciation guidance.
[0048] Naturally integrated into daily life, convenient and easy to use: The recording mode runs automatically in the background, continuously collecting the user's daily conversation voice data, without the need for additional user operation. The practice process is naturally integrated into the user's daily life. The user only needs to switch to the practice mode for targeted practice in his spare time, which greatly improves the convenience of practice and user stickiness.
[0049] Practice with consistent timbre to enhance the sense of immersion: The practice voice generated by the TTS module is consistent with the user's timbre, allowing users to gain a stronger sense of immersion during practice, making it easier to imitate and compare, thereby improving the practice effect.
[0050] Targeted practice to effectively improve expression skills: In practice mode, the system can push targeted practice content based on the weaknesses exposed by users in their daily pronunciation, so that users can practice in a targeted manner, quickly and effectively improve pronunciation problems, and improve the clarity, fluency and naturalness of voice expression, ultimately improving users' communication confidence and expression skills. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 This is a schematic diagram of the personalized pronunciation training system proposed by the present invention.
[0052] Figure 2 Flowchart for how the personalized pronunciation practice system works.
[0053] Figure 3 This is a schematic diagram of the voice registration interface.
[0054] Figure 4 This is a diagram of the pronunciation practice interface.
[0055] Figure 5 Generate flow charts for speech recording and practice data.
[0056] Figure 6 The figure is a flow chart of the personalized pronunciation training method proposed by the present invention. DETAILED DESCRIPTION
[0057] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0058] Embodiment 1
[0059] The present invention proposes a personalized pronunciation training system. The principle of the system is as follows: Figure 1 As shown, including:
[0060] The smart terminal is used to collect the user's voice data, receive and present the pronunciation practice text and voice generated by the server, and receive the user's operation instructions; the smart terminal is a portable electronic device used by the user, such as a smart phone, tablet computer, etc. Voice data collection can be achieved by calling the microphone array of the smart terminal, and noise reduction processing can be performed to improve the quality of voice data.
[0061] The server includes a voice registration module, a voice data storage module, a voice activity detection module, a speaker recognition module, a pronunciation quality detection module, a voice recognition module, a large language model module and a text-to-speech module; the server is built using a cloud computing platform or a high-performance server cluster.
[0062] The voice registration module is used to register the user's voice and provide the user's voice features to the speaker recognition module and the text-to-speech module;
[0063] Voice data storage module, used to store user voice data uploaded by smart terminals;
[0064] The voice activity detection (VAD) module is used to detect voice activity segments in voice data, determine valid voice segments, and send them to the speaker recognition module. In this implementation, silero-vad is used to implement voice activity detection. The VAD model is based on deep learning and supports dynamic adjustment of the confidence threshold of voice detection, allowing the system to automatically optimize the detection sensitivity according to the environmental noise level. It has high robustness and accuracy and can effectively detect voice activity segments in voice data. It ensures that subsequent processing (such as voice recognition) is only for valid voice data, reducing invalid calculations.
[0065] The speaker recognition module is used to identify the speaker in the voice data and confirm whether it is the registered target user. The speaker recognition module is implemented by the CAM++ speaker recognition model. The CAM++ speaker recognition model is an advanced speaker recognition model based on a densely connected time-delay neural network. Compared with traditional models such as ResNet34 and ECAPA-TDNN, the CAM++ model shows higher accuracy and faster reasoning speed in speaker recognition tasks.
[0066] The pronunciation quality detection module is used to evaluate the pronunciation quality of the user's voice and obtain the weak pronunciation segments; specifically:
[0067] Detect repetition, dragging and blocking in the user's voice; output the speech feature category sequence C = [c1, c2, ..., c i ,…,c T ], where c i ∈{R,B,P,F}, and the time interval of the weak pronunciation segment, c i represents the speech feature category of the i-th frame of audio, R represents repetition, B represents obstruction, P represents dragging sound, and F represents fluency; the weak pronunciation segments include repetition segments, obstruction segments and dragging sound segments.
[0068] In this embodiment, the pronunciation quality detection module adopts the StutterNet model. StutterNet is an advanced speech feature detection system based on deep learning, which is used in the present invention to efficiently and accurately evaluate the pronunciation quality of the user. The core architecture of the StutterNet model is the time delay neural network (TDNN), which enables it to capture the time dependency and context information in the speech signal excellently, so as to analyze the user's speech more accurately.
[0069] StutterNet performs well in speech feature detection. Its high accuracy and precise feature recognition capabilities provide strong technical support for the pronunciation quality assessment of the present invention, enabling the system to objectively and quantitatively evaluate user pronunciation and provide a key basis for subsequent personalized training content generation.
[0070] The present invention uses StutterNet to detect the following events in user speech:
[0071] Repetition (R): The repetition of a syllable, word or phrase in speech.
[0072] Pluration (P): The duration of a syllable or sound is abnormally prolonged.
[0073] Blockage (B): A pause in speech or an inability to produce sound smoothly.
[0074] Fluent speech (F): Normal speech without disfluencies.
[0075] In one embodiment of the present invention, the process of the StutterNet model for detecting the pronunciation quality of speech data is as follows: First, read the WAV file through the audio processing library to obtain the audio signal and its sampling rate. Then, the audio signal is pre-emphasized to enhance the high-frequency signal, and is framed, with each frame length of 25 milliseconds and a frame shift of 10 milliseconds. Then, the Mel-frequency cepstral coefficient (MFCC) of each frame is calculated as the input feature. After that, the pre-trained StutterNet model is loaded, and the extracted MFCC features are input into the model for detection. After the model is output, the results are analyzed to determine whether there are features such as repetition, dragging, blocking or fluent speech in the speech, and the detection results are output for further analysis or display.
[0076] The ASR module is used to convert the user's voice data into text content and associate the weakly pronounced segments with the text content. The specific steps of associating the weakly pronounced segments with the text content are as follows:
[0077] The speech recognition module outputs the text transcription results and the timestamp information corresponding to each word, which is expressed as the time interval at the word level as follows:
[0078]
[0079] Among them word i represents the i-th word, Indicates the start and end time of the word in the audio;
[0080] Based on the time interval of the weak segment obtained by pronunciation quality detection and the word timestamp information obtained by speech recognition, time interval matching is performed to determine the text content corresponding to the weak pronunciation segment. Specifically, for each time interval of the weak pronunciation segment Indicates the start time of the weak pronunciation segment, Indicates the end time of the weak pronunciation segment, traverses the word timestamp information, finds out the words in all word-level time intervals that overlap with the time interval of the weak pronunciation segment, and concatenates these words to obtain the text associated with the weak pronunciation segment.
[0081] In a specific embodiment of the present invention, the SenseVoice speech recognition engine is used to implement the speech recognition function. SenseVoice is a high-performance multi-language recognition engine. The model has been trained with large-scale data and supports more than 50 languages. It has high recognition accuracy and fast response speed, and can effectively convert user speech into text content, providing accurate text input for the subsequent LLM module.
[0082] The Large Language Model (LLM) module is used to generate user-specific pronunciation practice text based on the text associated with the weak pronunciation segments; specifically:
[0083] Constructing a prompt word template, wherein the content of the prompt word template includes a user's weak pronunciation text WT, an instruction, a homophone and easily confused word list CWL, and a practice text type indication PTI; the user's weak pronunciation text is a text associated with a weak pronunciation segment, the instruction is used to indicate the task that the large language model needs to perform, the homophone and easily confused word list includes words that have the same or similar pronunciation as the user's weak pronunciation text WT but different semantics, and words that are easily confused with the user's weak pronunciation text WT; the practice text type indication is used to indicate the type of practice text generated by the large language model;
[0084] The constructed prompt word template is filled with specific parameter values to form a complete prompt word, wherein the parameters include the user's weak pronunciation text, a list of homophones and easily confused words, and an indication of the type of practice text; the filled complete prompt word is input into the large language model, and the large language model generates a pronunciation practice text according to the indication of the prompt word.
[0085] The specific method for constructing the list of homophones and easily confused words is as follows:
[0086] Pre-build a vocabulary containing commonly used Chinese characters, words and their homophones, and easily confused characters / words. The vocabulary is organized and indexed based on dimensions such as pronunciation similarity, glyph similarity, and semantic relevance;
[0087] After obtaining the user's weakly pronounced text WT, a dynamic search is performed based on the user's weakly pronounced text to search for homophones with the same or similar pronunciation as the user's weakly pronounced text, as well as words that are easily confused with the user's weakly pronounced text, from a preset word library or online dictionary or knowledge base;
[0088] The dynamic retrieval strategy includes: searching for words whose pinyin is exactly the same or partially similar to the user's weakly pronounced text, calculating the similarity between word syllables, and determining words with a similarity higher than a set threshold as words that are easily confused with the user's weakly pronounced text.
[0089] In one embodiment of the present invention, the DeepSeek series model is used as a large language model to generate reference text for speech registration and reference text for pronunciation practice.
[0090] The text-to-speech (TTS) module is used to convert the pronunciation practice text generated by the large language model into a reference pronunciation audio that is consistent with the user's timbre. In one embodiment of the present invention, the text-to-speech (TTS) function is implemented by adopting the FishSpeech model. The model has zero-sample and few-sample speech synthesis capabilities. It only needs to provide a short user audio sample (10 to 30 seconds) to quickly clone the user's timbre and generate natural and realistic speech output based on the specified text. The model has been trained with more than 1 million hours of multilingual audio data, supports multiple languages including Chinese and English, and has powerful speech generation capabilities.
[0091] Embodiment 2
[0092] In this embodiment, the overall workflow of the personalized pronunciation training system is described. The overall workflow of the system is as follows: Figure 2 shown.
[0093] The user turns on the voice registration function of the voice terminal, and the system enters the voice registration interface, such as Figure 3 As shown in the figure, the voice registration interface mainly includes:
[0094] Voice registration reference text area: Displays reference text for users to read, such as the "Voice Registration Reference Text (Double-click to switch)" area. The reference text is dynamically generated by the server, and users can double-click the area to switch and select the appropriate reference text to read.
[0095] User current recording link: Displays the "User current recording link" area, which is used to display the link of the current recording. Users can click the link to play the current recording and judge whether the recording quality is clear and smooth.
[0096] Start / Stop Recording Button: The Start / Stop Recording Button is used to control the start and stop of recording. The user clicks the Start Recording Button to start recording, recites the reference text, and clicks the Stop Recording Button to end recording after reciting.
[0097] Save button: The "Save button" is used to save the user's registered voice.
[0098] Historical recording link area: For example, areas such as "Historical recording link 1" and "Historical recording link 2" are used to display links to historical recordings, making it easier for users to view and manage historical recordings.
[0099] In a quiet environment, the user turns on the recording function of the smart terminal and recites the voice registration reference text given by the terminal. The reference text is dynamically generated by the server, and the user can double-click the reference text area to switch and select the appropriate reference text; after the user finishes reciting, click the stop recording button, and the system generates a current recording link. The user can click the link to play the current recording and judge whether the recording is clear and smooth.
[0100] After recording, the user clicks the Save button, and the system sends all the reference text and corresponding voice data recorded by the user to the voice registration module of the server through the network. The server receives and saves the user's registered voice data for subsequent training and personalized configuration of the speaker recognition model and TTS model. In order to ensure the effectiveness of voice registration, the system requires the user to record at least 30 seconds of reference voice and ensure that the reference voice is clear and fluent to meet the needs of speaker recognition and TTS modules.
[0101] The process of voice recording and training data generation is as follows Figure 5 As shown, it mainly includes the following steps:
[0102] Smart terminals use sound receiving devices to record sounds and use VAD to detect voice signals: Smart terminals (such as smartphones and tablets) start their built-in or external sound receiving devices (such as microphones) to continuously record the sounds of the surrounding environment. At the same time, the VAD module starts working to monitor the audio signals collected by the recording device in real time to detect whether they contain voice activities.
[0103] The system uses the VAD module to determine whether a valid speech signal is detected.
[0104] If the VAD module does not detect voice activity within a certain period of time (for example, only silence or background noise), the smart terminal will continue recording and perform voice signal detection.
[0105] If the VAD module detects voice activity, the intelligent terminal will send the audio data segment containing the voice activity to the server through network transmission.
[0106] The server receives the voice data and starts the speaker recognition: After receiving the voice data from the intelligent terminal, the server immediately starts the speaker recognition module. In this embodiment, the speaker recognition module adopts the CAM++ speaker recognition model. The purpose of this model is to verify whether the received voice belongs to the pre-registered target user.
[0107] The speaker recognition module analyzes the received voice data to determine whether the speaker is the target user preset by the system; if the speaker recognition module determines that the speaker of the current voice is not the target user (for example, there are other people's voices in the environment), the current voice analysis process ends. The system will continue to monitor subsequent recording data; if the speaker recognition module confirms that the speaker is the target user, the system will conduct an in-depth analysis of the pronunciation quality.
[0108] The server starts pronunciation quality detection and speech recognition, and extracts text associated with weakly pronounced speech segments based on the pronunciation quality detection and speech recognition results: After confirming that the speaker is the target user, the server simultaneously starts the pronunciation quality detection module and the speech recognition (ASR) module.
[0109] The pronunciation quality detection module conducts a detailed analysis of the user's voice and detects pronunciation quality problems in the voice, such as repetition, dragging, blocking, etc.
[0110] The automatic speech recognition (ASR) module accurately converts user speech content into text.
[0111] After pronunciation quality detection and speech recognition are completed, by correlating the results of the two, the text content corresponding to weak pronunciation fragments such as repetition (denoted as R), dragging (denoted as P), and blocking (denoted as P) is accurately extracted for subsequent targeted practice text generation.
[0112] Input data:
[0113] The pronunciation quality detection model outputs the speech feature category sequence C = [c1, c2, ... c T ], where c i ∈{R,B,P,F}, and the time interval of the weak pronunciation segment, such as the time interval of the repeated (R) segment The time interval of the blocked fragment wait.
[0114] Speech recognition results: The ASR module outputs text transcription results and timestamp information corresponding to each word or syllable, expressed as word-level time intervals
[0115] Among them word i represents the i-th word, Indicates the start and end time of the word in the audio.
[0116] Calculate the associated pronunciation weak segment text:
[0117] Based on the time interval of the weak segment obtained by pronunciation quality detection and the word timestamp information obtained by speech recognition, time interval matching is performed to determine the text content corresponding to the weak pronunciation segment.
[0118] Specifically, for each time interval of a weak pronunciation segment The system traverses the word timestamp information in the ASR results and finds all words whose time intervals overlap with this interval. By splicing these words together, the text associated with the weak pronunciation segment can be obtained.
[0119] The system uses LLM to generate pronunciation practice text based on text associated with weakly pronounced speech fragments: The system uses the Large Language Model (LLM) module to generate targeted pronunciation practice text containing homophones / phrases and easily confused words / phrases based on the text content (words, phrases, fragments) with weak pronunciation of the user. This method is designed to help users distinguish and master easily confused pronunciations and improve the accuracy and clarity of speech expression. The specific process of generating pronunciation practice text is as follows:
[0120] a) Prepare LLM input: Construct prompt word template
[0121] In order to guide LLM to generate pronunciation practice texts that meet the requirements, the system pre-builds prompt word templates.
[0122] The core components of the prompt word template include:
[0123] User weakly pronounced text (Weak Text, denoted as WT): text content associated with user weakly pronounced speech segments, such as words, phrases or short sentences.
[0124] Instructions: Explicitly instruct the LLM on the tasks it needs to perform, such as “generate pronunciation practice texts containing homophones and easily confused words of the user’s weak pronunciation texts”.
[0125] Confusable Word List (CWL): A list of homophones and confusable words preset by the system or dynamically constructed. The list may include words that have the same or similar pronunciation as the user's weakly pronounced text (WT) but different semantics, as well as words that are easily confused with WT.
[0126] Practice Type Indicator (PTI): indicates the type of practice text generated by LLM, such as sentences, phrases, tongue twisters, short dialogues, etc.
[0127] A typical prompt word template example is as follows:
[0128] Prompt word template = "Please generate a pronunciation practice text based on the following text content with weak pronunciation of the user: '{WT}'.
[0129] The exercise text should focus on including homophones and easily confused words of '{WT}', refer to the list of homophones / easy confused words: '{CWL}'.
[0130] The exercise text type is: '{PTI}'.
[0131] b) LLM generates pronunciation practice text
[0132] Fill in the built prompt word template with specific parameter values to form a complete LLM input prompt word. The parameters include:
[0133] {WT}: User-pronounced weak text.
[0134] {CWL}: A list of homophones / confusable words dynamically retrieved from {WT} or selected from a preset list.
[0135] {PTI}: The practice text type selected according to practice needs, such as "sentences", "phrases" or "short dialogues".
[0136] The completed prompt words are input into the LLM module. Based on its powerful language understanding and generation capabilities, LLM generates pronunciation practice texts that meet the requirements according to the instructions of the prompt words.
[0137] c) Construction of Homophone / Confusable Word List (HWCL)
[0138] The construction of a homophone / confusable word list (HWCL) is a key step in generating effective pronunciation practice text. The HWCL can be constructed in the following ways:
[0139] Pre-built word library: The system pre-builds a word library containing commonly used Chinese characters, words and their homophones, and easily confused characters / words. The word library can be organized and indexed based on dimensions such as pronunciation similarity, glyph similarity, and semantic relevance.
[0140] Dynamic retrieval: After obtaining the user's weakly pronounced text [WT], the system performs a dynamic search based on [WT], searching for homophones with the same or similar pronunciation as [WT], as well as words that are easily confused with [WT] from the preset vocabulary or online dictionary or knowledge base.
[0141] Search strategies may include:
[0142] Phonetic Matching: Search for words whose phonetic spelling is exactly the same as or partially similar to [WT].
[0143] Syllable similarity calculation: Calculate the similarity between word syllables, such as similarity based on edit distance or phoneme features.
[0144] Semantic Relevance Analysis: Analyze the semantic relevance of words. For example, find words with similar semantic categories that often appear in similar contexts.
[0145] Examples of prompts for the LLM to generate pronunciation practice texts:
[0146] Please generate a pronunciation practice text based on the following text where the user has weak pronunciation: '花(huā)'.
[0147] The practice text should mainly include homophones and easily confused words of '花(huā)'. Refer to the list of homophones / easily confused words: '话(huà),画(huà),哗(huā),滑(huá),化(huà)'.
[0148] The type of the practice text is:'sentence'.
[0149] Examples of pronunciation practice texts generated by the LLM:
[0150] 1. All the flowers in the garden are in bloom in spring. The wind blows with a rustling sound, as if it's talking.
[0151] 2. This painting depicts a beautiful flower. The colors are vivid and lifelike, truly turning the ordinary into the miraculous.
[0152] 3. It took a lot of time to understand this sentence. Pay close attention to the emphasized parts. Don't try to be sensational, and don't be slippery.
[0153] 4. The beautiful flowers are like an unfolding scroll painting, reminding people of the carefree words of childhood. Time passes like sand sliding, and ultimately all turn into memories.
[0154] The server uses TTS to generate the audio corresponding to the pronunciation practice text: The server finally calls the Text-to-Speech (TTS) module. Based on the voice data registered by the user, it converts the generated pronunciation practice text into the corresponding practice audio. The TTS module strives to make the synthesized audio similar to the target user's voice in timbre, enhancing the intimacy and immersion of the practice.
[0155] Example 3
[0156] The present invention proposes a personalized pronunciation practice method based on the system of Example 1. The overall process of the method is as Figure 6 shown, including:
[0157] The intelligent terminal obtains pronunciation practice data from the server, including the pronunciation practice text and the corresponding reference audio with the same timbre as the user, and displays it to the user;
[0158] The user operates on the practice interface of the intelligent terminal to select whether to play the reference pronunciation audio, start recording personal practice voices, or switch practice data;
[0159] If the user chooses to play the reference pronunciation audio, the smart terminal plays the reference pronunciation audio, and the user listens to the demonstration pronunciation and practices;
[0160] If the user chooses not to play the reference pronunciation audio, determining whether the user chooses to start recording personal practice voice;
[0161] If the user chooses to start recording, the user recites the pronunciation practice text on the practice interface, and at the same time, the user's practice voice is recorded, and after the recording is completed, the user's practice voice data is sent to the server;
[0162] If the user chooses not to start recording, then determining whether the user chooses to switch to new practice data;
[0163] If the user chooses to switch practice data, the smart terminal obtains new pronunciation practice data from the server and starts a new round of practice;
[0164] After receiving the user practice voice data uploaded by the smart terminal, the server verifies the identity of the speaker and determines whether the practice voice comes from a registered target user; if the speaker is not a target user, the practice process ends; if the speaker is a registered target user, the pronunciation quality of the voice recorded by the user in the practice mode is evaluated to determine the weakly pronounced voice segments and texts;
[0165] The server feeds back the pronunciation quality test result to the smart terminal, and the smart terminal displays the pronunciation quality test result on the practice interface.
[0166] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0167] The above are only preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should be regarded as the protection scope of the present invention.
Claims
1. A personalized pronunciation practice system, characterized in that: include: Intelligent terminal, used to collect user's voice data, receive and present pronunciation practice text and voice generated by the server, and receive user's operation instructions; The server includes a voice registration module, a voice data storage module, a voice activity detection module, a speaker recognition module, a pronunciation quality detection module, a voice recognition module, a large language model module and a text-to-speech module; The voice registration module is used to register the user's voice and provide the user's voice features to the speaker recognition module and the text-to-speech module; Voice data storage module, used to store user voice data uploaded by smart terminals; The voice activity detection module is used to detect the voice activity segments in the voice data, determine the valid voice segments, and send them to the speaker recognition module; Speaker recognition module, used to identify the speaker in the voice data and confirm whether it is the registered target user; The pronunciation quality detection module is used to evaluate the pronunciation quality of the user's voice and obtain the weak pronunciation segments; The speech recognition module is used to convert user speech data into text content and associate weak pronunciation segments with text content; the large language model module is used to generate user-specific pronunciation practice text based on the text associated with the weak pronunciation segments; the text-to-speech module is used to convert the pronunciation practice text generated by the large language model into reference pronunciation audio consistent with the user's timbre.
2. The personalized pronunciation training system according to claim 1, characterized in that: The speaker recognition module is specifically implemented by a CAM++ speaker recognition model.
3. The personalized pronunciation training system according to claim 1, characterized in that: The method of evaluating the pronunciation quality of the user's voice is specifically as follows: Detect repetition, dragging and blocking in the user's voice; output the speech feature category sequence C = [c1, c2, ..., c i ,…,c T ], where c i ∈{R,B,P,F}, and the time interval of the weak pronunciation segment, c i represents the speech feature category of the i-th frame of audio, R represents repetition, B represents obstruction, P represents dragging sound, and F represents fluency; the weak pronunciation segments include repetition segments, obstruction segments and dragging sound segments.
4. The personalized pronunciation training system according to claim 1, characterized in that: The specific steps of associating the weak pronunciation segment with the text content are as follows: The speech recognition module outputs the text transcription results and the timestamp information corresponding to each word, which is expressed as the time interval at the word level as follows: Among them word i represents the i-th word, Indicates the start and end time of the word in the audio; Based on the time interval of the weak segment obtained by pronunciation quality detection and the word timestamp information obtained by speech recognition, time interval matching is performed to determine the text content corresponding to the weak pronunciation segment. Specifically, for each time interval of the weak pronunciation segment Indicates the start time of the weak pronunciation segment, Indicates the end time of the weak pronunciation segment, traverses the word timestamp information, finds out the words in all word-level time intervals that overlap with the time interval of the weak pronunciation segment, and concatenates these words to obtain the text associated with the weak pronunciation segment.
5. The personalized pronunciation training system according to claim 1, characterized in that: The specific method of generating a user-specific pronunciation practice text based on the text associated with the weak pronunciation segment is as follows: Constructing a prompt word template, wherein the content of the prompt word template includes a user's weak pronunciation text WT, an instruction, a homophone and easily confused word list CWL, and a practice text type indication PTI; the user's weak pronunciation text is a text associated with a weak pronunciation segment, the instruction is used to indicate the task that the large language model needs to perform, the homophone and easily confused word list includes words that have the same or similar pronunciation as the user's weak pronunciation text WT but different semantics, and words that are easily confused with the user's weak pronunciation text WT; the practice text type indication is used to indicate the type of practice text generated by the large language model; The constructed prompt word template is filled with specific parameter values to form a complete prompt word, wherein the parameters include the user's weak pronunciation text, a list of homophones and easily confused words, and an indication of the type of practice text; the filled complete prompt word is input into the large language model, and the large language model generates a pronunciation practice text according to the indication of the prompt word.
6. The personalized pronunciation training system according to claim 5, characterized in that: The method for constructing the list of homophones and easily confused words is specifically as follows: Pre-build a vocabulary containing commonly used Chinese characters, words and their homophones, and easily confused characters / words. The vocabulary is organized and indexed based on dimensions such as pronunciation similarity, glyph similarity, and semantic relevance; After obtaining the user's weakly pronounced text WT, a dynamic search is performed based on the user's weakly pronounced text to search for homophones with the same or similar pronunciation as the user's weakly pronounced text, as well as words that are easily confused with the user's weakly pronounced text, from a preset word library or online dictionary or knowledge base; The dynamic retrieval strategy includes: searching for words whose pinyin is exactly the same or partially similar to the user's weakly pronounced text, calculating the similarity between word syllables, and determining words with a similarity higher than a set threshold as words that are easily confused with the user's weakly pronounced text.
7. A personalized pronunciation practice method based on the system of claim 1, characterized in that: include: The intelligent terminal obtains pronunciation practice data from the server, including pronunciation practice text and corresponding reference audio with the same timbre as the user, and displays it to the user; The user operates on the practice interface of the smart terminal and chooses whether to play the reference pronunciation audio, start recording personal practice voice, or switch practice data; If the user chooses to play the reference pronunciation audio, the smart terminal plays the reference pronunciation audio, and the user listens to the demonstration pronunciation and practices; If the user chooses not to play the reference pronunciation audio, determining whether the user chooses to start recording personal practice voice; If the user chooses to start recording, the user recites the pronunciation practice text on the practice interface, and at the same time, the user's practice voice is recorded, and after the recording is completed, the user's practice voice data is sent to the server; If the user chooses not to start recording, then determining whether the user chooses to switch to new practice data; If the user chooses to switch practice data, the smart terminal obtains new pronunciation practice data from the server and starts a new round of practice; After receiving the user practice voice data uploaded by the smart terminal, the server verifies the speaker's identity and determines whether the practice voice comes from a registered target user; If the user is not the target user, the exercise process ends; If the speaker is a registered target user, the pronunciation quality of the speech recorded by the user in the practice mode is evaluated to determine the speech segments with weak pronunciation and the texts with weak pronunciation; The server feeds back the pronunciation quality test result to the smart terminal, and the smart terminal displays the pronunciation quality test result on the practice interface.