A method and system for comparing Chinese oral pronunciation based on speech recognition

By analyzing linguistic rules and simulating the articulatory organs to generate standard reference speech, this system solves the problem that existing systems cannot distinguish between natural speech variations and pronunciation defects, enabling more accurate pronunciation assessment and feedback, and improving the naturalness and fluency of learners' spoken expression.

CN122454977APending Publication Date: 2026-07-24HUBEI UNIV OF EDUCATION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUBEI UNIV OF EDUCATION
Filing Date
2026-04-08
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing Chinese pronunciation assessment systems cannot distinguish between natural speech variations and pronunciation defects, resulting in learners' unnatural and stiff pronunciation, and may even mislead them in the wrong learning direction.

Method used

By acquiring the text information to be pronounced, linguistic rule analysis is performed to simulate the continuous movement trajectory and acoustic adjustment rules of the speech organs, generating a standard reference speech containing speech connections and prosodic patterns, and aligning the learner's pronunciation with acoustic features to determine whether the differences are pronunciation defects.

Benefits of technology

It improves the accuracy of pronunciation assessment, provides feedback that is more in line with native speakers' perception, avoids unnatural pronunciation caused by learners correcting false errors, and enhances the efficiency and effectiveness of Chinese oral learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122454977A_ABST
    Figure CN122454977A_ABST
Patent Text Reader

Abstract

The application provides a Chinese oral pronunciation comparison method and system based on voice recognition, relates to the field of voice recognition, obtains text information to be pronounced and performs linguistic rule analysis, simulates the continuous motion track of the pronunciation organ and the acoustic adjustment law, thereby generating a standard reference voice containing voice continuity and rhythm patterns, compares the acoustic characteristics of the learner's pronunciation with the acoustic characteristics of the standard reference voice, and determines whether the difference is a pronunciation defect according to the natural voice changes reflected in the standard reference voice; the application can effectively distinguish the subtle differences in the learner's pronunciation that belong to natural voice changes from the real pronunciation defects, solve the technical problem that the existing system misjudges the natural coarticulation effect as a pronunciation error, improve the authenticity and authenticity of oral expression, overcome the deficiency that the system feedback is disconnected with the actual perception in the prior art, and significantly improve the efficiency and effect of Chinese oral learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology, and more specifically, to a method and system for comparing spoken Chinese pronunciation based on speech recognition. Background Technology

[0002] In Chinese language learning support systems, spoken pronunciation assessment and correction functions are crucial. These systems typically collect learners' pronunciation data, use speech recognition technology to convert it into analyzable acoustic information, compare it with preset standard pronunciation, and ultimately provide feedback. However, existing systems have significant problems in practical applications.

[0003] The standard pronunciation models within existing systems are built upon independent, idealized sound units, lacking an understanding and adaptability to these natural cophonic effects. When learners receive feedback from the system about these "abnormalities," they often try to "correct" these labeled "errors."

[0004] This deliberate suppression of natural cophony and excessive pursuit of a "perfect" match between individual sound units ultimately leads to learners' unnatural, stiff, and even robotic-like pronunciation. While learners may make each individual word acoustically closer to the system's set standards, the fluency, rhythm, and overall intonation of the entire word or sentence are severely disrupted. This persistent contradiction and disconnect ultimately causes learners significant confusion and frustration. They will find that despite achieving "high scores" on the learning system, their pronunciation still sounds stiff, disjointed, and even perceived as having a strange accent when communicating with native speakers. This huge discrepancy between system feedback and actual perception not only severely undermines learners' motivation and makes them doubt the system's guidance, but more importantly, it may lead learners in the wrong direction, unintentionally acquiring pronunciation habits far removed from natural spoken Chinese, thus hindering them from truly achieving fluent and natural spoken expression. Summary of the Invention

[0005] This application discloses a method and system for comparing spoken Chinese pronunciation based on speech recognition. It aims to solve the technical problem that existing spoken Chinese pronunciation assessment systems cannot distinguish between natural speech changes and pronunciation defects when processing natural speech, resulting in learners' unnatural and stiff pronunciation, and may lead them to the wrong learning direction.

[0006] The technical solution of this application is as follows: Firstly, this application discloses a method for comparing spoken Chinese pronunciation based on speech recognition, including: Obtain the text information to be pronounced; Linguistic rule analysis is performed on text information to obtain linguistic analysis results; linguistic rule analysis refers to the identification and analysis of pronunciation rules unique to Chinese. Based on the results of linguistic analysis, the continuous movement trajectory and acoustic adjustment rules of the vocal organs were simulated to obtain simulation results; Based on the simulation results, a standard reference speech containing speech cohesion and prosodic patterns is generated; Collect learners' actual pronunciation and preprocess it to obtain the preprocessed pronunciation; Extract acoustic feature sequences from the preprocessed actual pronunciation; align the acoustic feature sequences of the learner's pronunciation with the standard reference speech in time; Based on this time alignment, the acoustic features of the learner's pronunciation are compared with those of the standard reference speech to obtain the differences between the acoustic features of the learner's pronunciation and those of the standard reference speech. Based on the natural speech changes reflected in the standard reference speech, it is determined whether the difference is a pronunciation defect.

[0007] Secondly, this application also discloses a Chinese spoken pronunciation comparison system based on speech recognition, the system comprising: The text information acquisition module is used to acquire the text information to be pronounced; The linguistic rule analysis module is used to perform linguistic rule analysis on the text information to obtain linguistic analysis results; the linguistic rule analysis refers to the identification and analysis of pronunciation rules unique to Chinese. The speech organ simulation module is used to simulate the continuous movement trajectory and acoustic adjustment rules of the speech organs based on the linguistic analysis results, and obtain simulation results; A standard reference speech synthesis module is used to generate standard reference speech containing speech cohesion and prosodic patterns based on the simulation results; The actual pronunciation acquisition module is used to acquire the learner's actual pronunciation and preprocess the actual pronunciation to obtain the preprocessed actual pronunciation; The acoustic feature extraction module is used to extract acoustic feature sequences from the preprocessed actual pronunciation and to align the acoustic feature sequences of the learner's pronunciation with the standard reference speech in time. The pronunciation defect judgment module is used to compare the acoustic features of the learner's pronunciation with the acoustic features of the standard reference speech based on the time alignment, obtain the difference between the acoustic features of the learner's pronunciation and the acoustic features of the standard reference speech, and determine whether the difference is a pronunciation defect based on the natural speech changes reflected in the standard reference speech.

[0008] Beneficial effects This application discloses a Chinese spoken pronunciation comparison method based on speech recognition. By acquiring the text information to be pronounced and performing linguistic rule analysis, it simulates the continuous movement trajectory and acoustic adjustment rules of the articulatory organs, thereby generating a standard reference speech that includes speech cohesion and prosodic patterns. This standard reference speech fully considers the unique pronunciation rules of Chinese and the co-articulation effect in natural speech flow, avoiding the problems of overly idealized and independent standard pronunciation models in existing systems. When comparing learners' actual pronunciation, this method not only extracts acoustic feature sequences and performs time alignment, but more importantly, based on this time alignment, it compares the acoustic features of the learner's pronunciation with the acoustic features of the standard reference speech, and judges whether the difference is a pronunciation defect based on the natural speech changes reflected in the standard reference speech. This application can effectively distinguish between subtle differences in learners' pronunciation that belong to natural speech changes and real pronunciation defects, solving the technical problem of existing systems misjudging natural co-articulation effects as pronunciation errors. This enables the system to provide more accurate feedback that is more in line with native speaker perception, avoiding the problem of learners' pronunciation becoming unnatural and stiff due to correcting "false errors." Accordingly, learners can better understand and master the natural rhythm and fluency of spoken Chinese, thereby improving the authenticity and naturalness of their spoken expression. This overcomes the shortcomings of existing technologies where system feedback is disconnected from actual perception, and significantly improves the efficiency and effectiveness of spoken Chinese learning. Attached Figure Description

[0009] Figure 1 A schematic diagram illustrating a Chinese spoken pronunciation comparison method based on speech recognition provided in this application.

[0010] Figure 2 A schematic diagram of a Chinese spoken pronunciation comparison system based on speech recognition provided in this application. Detailed Implementation

[0011] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. The components of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0012] Reference Figure 1 The diagram illustrates an embodiment of a Chinese spoken pronunciation comparison method based on speech recognition according to the present invention, which may specifically include the following steps: S101, Obtain the text information to be pronounced; S102, perform linguistic rule analysis on the text information to obtain linguistic analysis results; the linguistic rule analysis refers to the identification and analysis of pronunciation rules unique to Chinese. S103, Based on the linguistic analysis results, the continuous movement trajectory and acoustic adjustment rules of the vocal organs are simulated to obtain simulation results; S104, Based on the simulation results, generate a standard reference speech that includes speech cohesion and prosodic patterns; S105, Collect the learner's actual pronunciation and preprocess the actual pronunciation to obtain the preprocessed actual pronunciation; S106, Extract the acoustic feature sequence from the preprocessed actual pronunciation; Align the acoustic feature sequence of the learner's pronunciation with the standard reference speech in time; S107, based on the time alignment, compare the acoustic features of the learner's pronunciation with the acoustic features of the standard reference speech to obtain the difference between the acoustic features of the learner's pronunciation and the acoustic features of the standard reference speech, and determine whether the difference is a pronunciation defect based on the natural speech changes reflected in the standard reference speech.

[0013] To facilitate a clearer understanding of the technical solutions proposed in this application, some key terms and implementation environments are first explained. The "linguistic rule analysis" referred to in this application refers to the identification and analysis of pronunciation rules unique to Chinese, such as tone changes, linking, tone sandhi, neutral tone, and retroflexion. These rules form the basis for the natural and fluent expression of spoken Chinese, and their analysis results are crucial for subsequent simulation of the movement trajectory and acoustic adjustment rules of the articulatory organs. "Continuous movement trajectory and acoustic adjustment rules of the articulatory organs" refers to how the tongue, lips, teeth, soft palate, and other articulatory organs coordinate their movements in continuous speech, and how the resulting acoustic features (such as pitch, intensity, and timbre) are dynamically adjusted. These rules reflect the coherence and adaptability of native speakers' speech in natural communication. "Phonetic cohesion and prosodic patterns" refer to the natural transitions between adjacent syllables or words in continuous speech, as well as the overall rhythm, intonation, and stress distribution of speech. These are important components of natural and fluent spoken expression. This method can be implemented on various computing devices that support speech processing, such as personal computers, smartphones, tablets, or dedicated language learning devices. These devices are typically equipped with microphones for capturing speech and have sufficient computing power to execute complex speech recognition and analysis algorithms.

[0014] The method proposed in this application achieves accurate comparison and defect judgment of spoken Chinese pronunciation through a series of steps.

[0015] First, it is necessary to obtain the text information to be pronounced. This can be achieved in various ways. For example, learners can directly input the text on the learning interface, or select the text from the preset course content, or convert the learner's speech input into text through speech recognition technology. For example, if a learner needs to practice the sentence "How are you", then "How are you" is the text information to be pronounced.

[0016] Next, perform a linguistic rule analysis on the text information to obtain the linguistic analysis result. The linguistic rule analysis is to identify and analyze the pronunciation rules unique to Chinese. Specifically, a pre-trained Chinese linguistic analysis model can be used to perform morphological analysis, syntactic analysis on the input text, and identify the unique pronunciation phenomena in Chinese such as tone sandhi, liaison, weak stress, and erhua. For example, for the text "How are you", the linguistic analysis will identify the tone sandhi rule between "How" and "are" (the rising tone + rising tone changes to the rising tone + rising tone), and the weak stress pronunciation feature of "you" as a modal particle.

[0017] Then, based on the linguistic analysis result, simulate the continuous movement trajectory of the articulatory organs and the acoustic adjustment rule to obtain the simulation result. This step aims to simulate the dynamic changes of the articulatory organs of native speakers in natural speech flow. An acoustic-articulatory model can be used to predict the continuous changes of the articulatory organs such as tongue position, lip shape, and soft palate position during pronunciation, as well as the dynamic adjustment of acoustic features such as pitch, intensity, and formants, in combination with the linguistic analysis result. For example, in the simulation of "How are you", the smooth transition between the falling tone of "How" and the rising tone of "are" will be considered, as well as the weak influence of the weak stress of "you" on the overall pitch and intensity.

[0018] According to the simulation result, generate a standard reference speech containing speech connection and prosodic pattern. This step uses speech synthesis technology to convert the simulated movement of the articulatory organs and the acoustic adjustment rule into an audible standard reference speech. Different from the traditional standard speech of independent pronunciation, the standard reference speech generated in this application will naturally include speech connection (such as liaison, tone sandhi) and the overall prosodic pattern (such as speech rate, intonation, stress). For example, the generated standard reference speech of "How are you" will be a smooth and natural whole, rather than a simple splicing of three independent syllables. The tone sandhi and liaison between "How" and "are" will be naturally reflected, and the weak stress of "you" will also be accurately synthesized.

[0019] At the same time, collect the actual pronunciation of the learner and preprocess the actual pronunciation to obtain the preprocessed actual pronunciation. The actual pronunciation of the learner can be collected in real time through devices such as microphones. The preprocessing steps usually include noise suppression, endpoint detection, gain normalization, etc., to eliminate environmental noise, determine the start and end points of the speech, and unify the loudness of the speech, so as to improve the accuracy of subsequent analysis.

[0020] Next, acoustic feature sequences are extracted from the preprocessed actual pronunciation. These acoustic feature sequences are quantized representations of the speech signal at different time points and can include Mel-frequency cepstral coefficients (MFCC), fundamental frequency (F0), energy, duration, etc. These features comprehensively reflect the acoustic characteristics of the learner's pronunciation.

[0021] Subsequently, the acoustic feature sequence of the learner's pronunciation is temporally aligned with the standard reference speech. Temporal alignment is a crucial step in speech comparison, ensuring that each phoneme or syllable of the learner's pronunciation can be accurately compared with its corresponding part in the standard reference speech. Commonly used temporal alignment algorithms include Dynamic Temporal Warping (DTW) or forced alignment based on Hidden Markov Models (HMMs).

[0022] Finally, based on time alignment, the acoustic features of the learner's pronunciation are compared with those of the standard reference speech to identify differences between the two. Then, based on the natural speech variations already reflected in the standard reference speech, the system determines whether these differences constitute pronunciation defects. This step is the core innovation of this application. During the comparison process, the system not only identifies quantitative differences in acoustic features, but more importantly, it combines these differences with natural speech variations already reflected in the standard reference speech (such as cophonation, tone sandhi, and linking) to determine whether these differences are normal natural variations or genuine pronunciation defects. For example, if a learner makes a slight pitch change between "you" and "good" in "hello ma," but this change is consistent with the natural tone sandhi trend in the standard reference speech, the system will determine that this is not a pronunciation defect, but a natural speech connection. Conversely, if a learner shows a significant deviation from the standard reference speech in the tone, duration, or timbre of a syllable that is not a natural variation, it will be judged as a pronunciation defect.

[0023] The proposed method for comparing spoken Chinese pronunciation based on speech recognition generates a standard reference speech that includes phonetic connections and prosodic patterns by incorporating in-depth analysis of Chinese-specific linguistic rules and simulating the continuous movement trajectory and acoustic adjustment patterns of the articulatory organs. This core innovation enables the method to accurately capture the natural cophonation effects and prosodic variations in spoken Chinese, thereby overcoming the limitation of existing systems that misinterpret these natural variations as pronunciation defects.

[0024] Compared to existing methods that compare independent, idealized sound units, the advantage of this application lies in its adaptability to natural speech flow. Existing systems, lacking an understanding of the consonant effect, often incorrectly label subtle acoustic differences in learners' pronunciation caused by natural speech flow as "errors," leading learners to deliberately suppress natural pronunciation and make their spoken language sound stiff and unnatural. This application, by generating a standard reference speech containing natural speech connections and prosodic patterns, can distinguish between natural speech variations and genuine pronunciation defects. For example, when learners exhibit natural linking or tone sandhi in continuous pronunciation, this method does not consider it an error but compares it with the natural variations already reflected in the standard reference speech, thus avoiding misjudgment.

[0025] Furthermore, the method described in this application can more accurately assess learners' pronunciation quality and provide more instructive feedback. By determining whether differences constitute pronunciation defects, this method helps learners identify and correct genuine pronunciation problems while encouraging them to maintain natural fluency in their spoken language. This method not only improves the accuracy of pronunciation assessment but also avoids the predicament of learners pursuing "independent perfection" pronunciation, leading to unnatural spoken language, thereby significantly improving the efficiency and effectiveness of Chinese spoken language learning.

[0026] This application further proposes to optimize the steps of simulating the continuous movement trajectory and acoustic adjustment rules of the vocal organs in order to generate simulation results that are more consistent with the characteristics of natural speech flow.

[0027] In this regard, this application further proposes the steps for obtaining simulation results based on the above-mentioned linguistic analysis results, simulating the continuous movement trajectory and acoustic adjustment rules of the vocal organs, including: Identify specific words contained in text information; Query the prosodic pattern corresponding to the specific word; the prosodic pattern includes the natural prosodic information of the specific word in real speech flow; Based on the prosodic pattern, the predicted parameters of the articulatory organ movement are adjusted to obtain the adjusted predicted parameters of the articulatory organ movement; the predicted parameters include the starting point and rate of change of the pitch curve, or the duration distribution of syllables and the shape of the energy envelope; Based on the adjusted predicted parameters of the vocal organ movement, the continuous movement trajectory and acoustic adjustment law of the vocal organ are simulated to obtain simulation results.

[0028] Specifically, identifying specific words in text information refers to performing lexical analysis on the input text to be pronounced to identify words with special pronunciation or prosodic features, such as polyphonic words, proper nouns, idioms, or words with emphasis in specific contexts. These words often exhibit prosodic characteristics different from ordinary words in real speech. Querying the prosodic pattern corresponding to specific words means that the system obtains the prosodic features that these specific words typically exhibit in actual spoken language by consulting a pre-built prosodic knowledge base or analyzing a large amount of real-world data through machine learning models. The prosodic pattern can be understood as describing the variation patterns of speech in time, frequency, and intensity, including the natural prosodic information of specific words in real-world speech, such as their stress, intonation, speech rate, and pauses in different contexts. In practical applications, adjusting the prediction parameters of articulatory organ movement based on the prosodic pattern means that before simulating articulatory organ movement, the parameters used to predict the trajectory of articulatory organ movement and acoustic output are finely adjusted based on the prosodic pattern of the queried specific words. The prediction parameters are specifically core parameters affecting speech generation. These may include, for example, the starting point and rate of change of the pitch curve, which determine the fundamental frequency trend and intonation changes of the speech; or the duration distribution of syllables and the shape of the energy envelope, which affect the rhythm, stress, and loudness of the speech. By adjusting these parameters, the simulated articulatory organ movements can more accurately reflect the prosodic features of specific words in real speech. Therefore, based on the adjusted prediction parameters of the articulatory organ movements, simulating the continuous movement trajectory and acoustic adjustment patterns of the articulatory organs to obtain simulation results means, based on the above parameter adjustments, using an articulatory organ model (e.g., a physical model or a data-driven model) to simulate the coordinated movements of the tongue, lips, vocal cords, and other articulatory organs, and the resulting acoustic signal changes, thereby generating simulation results containing more accurate prosodic information. This simulation result will serve as the basis for subsequent generation of standard reference speech.

[0029] Through the aforementioned technical solution, this application significantly improves the naturalness and accuracy of standard reference speech. Compared to methods that rely solely on general linguistic rules for simulation, this application identifies specific words and combines their prosodic patterns in real speech flow, enabling the simulation results of the articulatory organs to capture the unique prosodic variations in spoken Chinese more precisely. This refined simulation, especially in the adjustment of prediction parameters such as pitch, duration, and energy envelope, ensures that the generated standard reference speech is not only accurate at the phoneme level but also possesses a high degree of naturalness and fluency at the prosodic level. Consequently, when learners' actual pronunciation is compared with this highly natural standard reference speech, the system can more accurately distinguish between genuine pronunciation defects and natural speech variations, avoiding misjudging normal prosodic variations as errors. This provides learners with more accurate and instructive pronunciation feedback, effectively improving the learning outcome of spoken Chinese.

[0030] In some preferred embodiments, a specific example is given below for illustration. Suppose the text information to be pronounced is "I love Beijing". First, the system will perform a lexical analysis on this text to identify specific vocabulary, such as "Beijing", which may have specific prosodic manifestations in Chinese spoken language. Then, the system will query the prosodic pattern corresponding to the specific vocabulary "Beijing" in real speech flow. For example, the query result may show that when "Beijing" is used as a place name, the second syllable "Jing" usually has a slight rising tone, and the overall duration is slightly longer than that of an ordinary disyllabic word. Then, based on these queried prosodic patterns, the system will adjust the prediction parameters of the movement of the articulatory organs. Specifically, for "Beijing", the system will adjust the starting point and change rate of the pitch curve of the character "Jing" to make it show the expected rising tone, and increase the duration allocation of its syllable and the shape of the energy envelope to ensure the fluency and naturalness of its pronunciation. Finally, based on these adjusted prediction parameters of the movement of the articulatory organs, the system will simulate the continuous movement trajectory of the articulatory organs and the acoustic adjustment rules, so as to obtain a simulation result containing a more natural and accurate prosodic pronunciation of "Beijing". This simulation result will then be used to generate a standard reference speech containing speech connection and prosodic patterns, providing a high-quality reference for Chinese spoken language pronunciation for learners.

[0031] The present application further proposes the following steps on the basis of time alignment: comparing the acoustic features of the learner's pronunciation with the acoustic features of the standard reference speech, obtaining the differences between the acoustic features of the learner's pronunciation and the acoustic features of the standard reference speech, and judging whether the differences are pronunciation defects according to the natural speech variations already reflected in the standard reference speech, including: Extracting the prosodic parameters of the learner's pronunciation; Extracting the prosodic parameters of the standard reference speech; Analyzing the differences between the prosodic parameters of the learner's pronunciation and the prosodic parameters of the standard reference speech to obtain prosodic parameter differences; Identifying specific vocabulary included in the text information; Querying the first pattern corresponding to the specific vocabulary; the first pattern refers to the natural prosodic information that is generally accepted in the real native speaker community and has slight differences from the general prosodic pattern; When the prosodic parameter differences exceed the preset general tolerance range, judging whether the prosodic pattern of the learner's pronunciation conforms to the first pattern to obtain a judgment result; According to the judgment result, determining that the learner's pronunciation is a natural prosodic variant or a pronunciation defect; the natural prosodic variant refers to natural prosodic variations.

[0032] Specifically, when assessing pronunciation defects, it is first necessary to extract the prosodic parameters of the learner's pronunciation and those of the standard reference speech. These prosodic parameters can be understood as features in speech related to pitch, duration, volume, speech rate, rhythm, and stress, which together constitute the prosodic pattern of speech. By extracting and quantifying these parameters, the specific prosodic manifestations of the learner's pronunciation and the standard reference speech can be obtained. Subsequently, the extracted prosodic parameters are compared and analyzed to determine the differences between the two.

[0033] Furthermore, to more accurately determine pronunciation defects, this application also identifies specific words that may be contained in the text information to be pronounced. Specific words refer to words whose pronunciation prosodic patterns may have unique or variable characteristics in the Chinese context due to factors such as context, habits, or region; for example, some common phrases, idioms, or polyphonic words. For these specific words, the system queries their corresponding first pattern. The first pattern refers to natural prosodic information that is generally accepted in the real native speaker community and has subtle differences from the general prosodic pattern. For example, some words may have a widely accepted prosodic pattern in spoken language that deviates slightly from the standard dictionary pronunciation; this pattern is not an error but a natural language variation.

[0034] Based on this, when the differences in prosodic parameters obtained from the analysis exceed a preset general tolerance range, the system will further determine whether the learner's pronunciation prosodic pattern matches the first pattern. The general tolerance range refers to the threshold of prosodic difference that is allowed between the learner's pronunciation and the standard reference speech under normal circumstances. If the difference is within this threshold, the pronunciation is generally considered acceptable. However, when the difference exceeds this range, a deeper judgment is required. At this time, a judgment result can be obtained by comparing the learner's pronunciation prosodic pattern with the first pattern corresponding to the specific word.

[0035] Ultimately, based on the judgment results, the system can determine whether a learner's pronunciation belongs to a natural prosodic variation or a genuine pronunciation defect. Natural prosodic variations refer to prosodic changes that are generally accepted by the native speaker community, caused by factors such as context, speech rate, emotion, or personal pronunciation habits, without affecting semantic comprehension and fluency of communication. By introducing the concept of a first mode, this application can effectively distinguish these natural prosodic variations from actual pronunciation errors, thereby providing more accurate and humanized pronunciation feedback.

[0036] Through the above technical solution, the present application can significantly improve the accuracy and robustness of Chinese spoken language pronunciation comparison. Especially when dealing with the prevalent prosodic diversity and natural variants of specific words in Chinese spoken language, the present application can effectively distinguish natural prosodic changes from actual pronunciation defects, avoiding misjudgment. This not only provides learners with more accurate and targeted pronunciation feedback, helping them understand and master the subtle prosodic changes in Chinese spoken language, but also can enhance learners' learning enthusiasm, reduce the sense of frustration caused by misjudgment, and thus promote more efficient and natural Chinese spoken language learning.

[0037] In some preferred embodiments, it is assumed that the learner is practicing the pronunciation text "It doesn't matter. You go first." Among them, "It doesn't matter" is a commonly used phrase. In daily spoken language, there may be multiple natural variants of its prosodic pattern that are generally accepted by native speakers. For example, when speaking at a faster speed, the pitch or duration of the word "méi" may be slightly different from the standard pronunciation, but the overall auditory perception is still natural and fluent.

[0038] First, the system will extract the prosodic parameters of the learner's pronunciation of "It doesn't matter" and compare them with the prosodic parameters of the standard reference speech, and find that the difference in its prosodic parameters exceeds the preset general tolerance range. At this time, if judged according to the traditional method, this pronunciation may be directly determined as a defect. However, the solution of the present application will further identify that "It doesn't matter" is a specific word and query the corresponding first pattern in the community prosodic variant knowledge base. Assume that this knowledge base contains a natural prosodic variant pattern of "It doesn't matter" in fast spoken language, and this pattern allows for a slight compression of the pitch or duration of the word "méi" in a specific context.

[0039] Subsequently, the system will compare the prosodic pattern of the learner's pronunciation of "It doesn't matter" with the queried first pattern. If the comparison result shows that the two are highly consistent, that is, the learner's pronunciation pattern is consistent with this natural variant pattern generally accepted by native speakers, then even if there is a large difference from the standard reference speech, the system will determine the learner's pronunciation as a natural prosodic variant rather than a pronunciation defect. Thus, the learner will receive feedback, informing that although their pronunciation is slightly different from the standard, it belongs to an acceptable natural spoken expression and does not need to be overly corrected, thereby avoiding unnecessary corrections and helping learners better understand the flexibility of Chinese spoken language.

[0040] The present application further proposes that when the difference in the prosodic parameters exceeds the preset general tolerance range, the steps of determining whether the prosodic pattern of the learner's pronunciation matches the first pattern and obtaining a judgment result include: Calculating the matching degree between the prosodic pattern of the learner's pronunciation and the first pattern; Comparing the matching degree with a preset fuzzy matching threshold to obtain a comparison result; Based on the comparison result, determine whether the prosodic pattern of the learner's pronunciation matches the first pattern.

[0041] Specifically, calculating the matching degree between the prosodic pattern of the learner's pronunciation and the first pattern refers to evaluating the similarity or closeness between the two in a quantitative manner. For example, various algorithms such as the Dynamic Time Warping (DTW) algorithm, cosine similarity, Euclidean distance, etc. can be used to compare the prosodic pattern of the learner's pronunciation, such as the pitch curve, duration distribution, energy envelope, etc., with the first pattern, so as to obtain a value between 0 and 1, where the larger the value, the higher the matching degree.

[0042] Among them, the matching degree is compared with a preset fuzzy matching threshold to obtain a comparison result, which can be understood as introducing an elastic judgment mechanism. The fuzzy matching threshold is a preset value used to define the degree of "matching". For example, when the matching degree is higher than this threshold, it is considered that the two match; when the matching degree is lower than this threshold, it is considered that the two do not match. This threshold can be adjusted according to the actual application scenario, corpus analysis results, and expert experience to adapt to the prosodic variations in different accents or contexts.

[0043] In practical applications, determining whether the prosodic pattern of the learner's pronunciation matches the first pattern based on the comparison result means obtaining a final judgment according to the comparison result between the matching degree and the fuzzy matching threshold. For example, if the matching degree is greater than or equal to the fuzzy matching threshold, it is determined that the prosodic pattern of the learner's pronunciation matches the first pattern; otherwise, if the matching degree is less than the fuzzy matching threshold, it is determined that they do not match.

[0044] Through the above technical solution, the present application can provide a more refined and robust prosodic pattern matching and judgment method. By quantifying the matching degree and introducing a fuzzy matching threshold, the accuracy and flexibility of the system in distinguishing natural prosodic variants and pronunciation defects are significantly improved. Compared with simple binary judgment, the solution of the present application can better adapt to the rich and diverse natural prosodic variations in Chinese spoken pronunciation, reduce the misjudgment rate, and thus provide more accurate and instructive pronunciation feedback for learners, effectively improving the efficiency of speech learning and the user experience.

[0045] In some preferred embodiments, assume that the text to be pronounced contains the specific word "Hello", and its corresponding first pattern is identified as a natural variant with a slightly rising intonation in the community prosodic variant knowledge base. When the learner actually pronounces "Hello", the system first extracts its prosodic parameters and compares them with the prosodic parameters of the standard reference speech, and finds that the prosodic parameter differences exceed the preset general tolerance range. At this time, the system needs to determine whether the prosodic pattern of the learner's pronunciation matches the first pattern.

[0046] Specifically, the system calculates the degree of match between the learner's pronunciation of "hello" and the first pattern. For example, by analyzing features such as the shape of the pitch curve, changes in speech rate, and stress distribution, a matching degree value is calculated, say 0.75. Simultaneously, the system presets a fuzzy matching threshold, for example, 0.70. The calculated matching degree of 0.75 is compared with the preset fuzzy matching threshold of 0.70. Since 0.75 is greater than 0.70, the comparison result is "matched". Based on this comparison result, the system ultimately determines that the learner's pronunciation prosodic pattern matches the first pattern, thus classifying the pronunciation as a natural prosodic variation rather than a pronunciation defect.

[0047] This application further proposes to optimize the steps of querying the first pattern corresponding to the specific words mentioned above, in order to provide a more personalized and context-adaptive pronunciation comparison.

[0048] The above-mentioned method for comparing spoken Chinese pronunciation based on speech recognition includes the following step: querying the first pattern corresponding to the specific word. Receive learner-set target accent or contextual preference information; Based on the target accent or contextual preference information, a first pattern matching the target accent or contextual preference information is selected from the community prosodic variant knowledge base.

[0049] Specifically, receiving learner-set target accent or contextual preference information refers to the system acquiring the pronunciation standards or specific requirements of the communication environment that the learner expects to achieve during the learning process. Target accent can be understood as the pronunciation characteristics of a specific region or social group that the learner wishes to emulate, such as standard Mandarin pronunciation, a particular regional accent, or the pronunciation habits of a specific professional group. Contextual preference information refers to the learner's choice of pronunciation style in specific communication scenarios. For example, the rhythmic patterns of pronunciation may differ in different contexts such as formal speeches, everyday conversations, and poetry recitations. This information can be actively input by the learner through the user interface or automatically inferred by analyzing the learner's historical learning data and learning goals.

[0050] Furthermore, based on the target accent or contextual preference information, the system selects a first pattern that matches the target accent or contextual preference information from the community prosodic variant knowledge base. This means that the system uses the received personalized preference information to match and select within the pre-built community prosodic variant knowledge base. The community prosodic variant knowledge base is a collection containing a large amount of pronunciation data from real native speakers in different accents and contexts. It stores natural prosodic patterns corresponding to various specific words and their related metadata, such as accent type, contextual tags, and usage frequency. By comparing the learner's target accent or contextual preference information with the metadata in the knowledge base, the system can intelligently select the "first pattern" that best meets the learner's needs, thereby ensuring that subsequent pronunciation comparisons are more accurate and personalized.

[0051] Through the aforementioned technical solution, this application can significantly improve the accuracy and personalization of Chinese spoken pronunciation comparison. Specifically, because the selection of the "first mode" fully considers the learner's target accent and contextual preferences, the system can more accurately distinguish whether the learner's pronunciation belongs to an acceptable natural rhythmic variation or a genuine pronunciation defect. This avoids misjudging pronunciations that conform to a specific accent or context as errors, thus providing learners with more targeted and constructive feedback. Furthermore, this personalized comparison method helps learners better master pronunciation skills for specific accents or contexts, improving learning efficiency and satisfaction, and making pronunciation training closer to the needs of real-world language environments.

[0052] In some preferred embodiments, suppose a learner is learning Chinese and wants to master Mandarin pronunciation with a Beijing accent. During pronunciation practice, the learner can set their "target accent" to "Beijing accent" through the system interface. When the system needs to query the "first pattern" of a specific word (e.g., a word with many retroflex endings), it first receives the learner's set "Beijing accent" preference information. Subsequently, the system filters through a pre-built community prosodic variation knowledge base based on this preference. This knowledge base may store a large amount of pronunciation data from native speakers in the Beijing area, including their unique prosodic patterns for specific words. The system prioritizes prosodic patterns that highly match the characteristics of the "Beijing accent" as the "first pattern" for that word. For example, for words with retroflex endings, the "first pattern" of the Beijing accent might exhibit a more pronounced retroflex prosodic feature. When a learner's actual pronunciation rhythmic pattern matches the "first pattern" of the selected "Beijing accent," the system will classify it as a natural rhythmic variation rather than a pronunciation defect, even if there are slight differences from the rhythmic pattern of standard Mandarin. Conversely, if the learner's pronunciation deviates significantly from the "first pattern" of the "Beijing accent," it will be identified as a pronunciation defect.

[0053] As a specific implementation method, suppose a learner is preparing a formal Chinese speech and wants their pronunciation to sound more dignified and fluent. The learner can set their "context preference" to "formal speech" in the system. When the system compares the learner's pronunciation of a word in the speech, it will, based on the "formal speech" context preference, filter from the community's prosodic variant knowledge base to find the "first pattern" of that word in a formal context. This "first pattern" might exhibit a smoother intonation, clearer articulation, and a more moderate speaking speed. If the learner's pronunciation conforms to this prosodic pattern in a formal context, even if it differs from the prosodic pattern of everyday conversation, the system will consider it a natural variant.

[0054] This application further proposes a more refined first pattern selection method, which introduces contextual adaptability assessment and comprehensive matching score calculation to ensure that the selected first pattern can better adapt to specific pronunciation scenarios.

[0055] The steps described above, which involve selecting a first pattern from the community prosodic variant knowledge base that matches the target accent or contextual preference information, include: Receive learner-set target accent or contextual preference information; A preliminary screening of patterns in the community prosody variant knowledge base yields a pattern set. Each pattern in the pattern set is subjected to a contextual adaptability assessment, which includes analyzing the naturalness, fluency, and semantic fit of the pattern in the context of the current text to be pronounced, and obtaining the contextual adaptability assessment result. Based on the context adaptability assessment results, calculate the comprehensive matching score for each pattern in the pattern set; Select the pattern with the highest overall matching score as the first pattern.

[0056] Specifically, receiving learner-defined target accent or contextual preference information means that the system obtains the accent type (e.g., Mandarin, Sichuan dialect, etc.) or specific contextual preferences (e.g., formal occasions, daily conversations, etc.) explicitly specified by the user. This information is used to guide the subsequent first mode selection process.

[0057] The initial screening of patterns in the community prosodic variant knowledge base yields a pattern set. This can be understood as a rough filtering of the vast community prosodic variant knowledge base based on the received target accent or contextual preference information. For example, if the learner prefers "Mandarin," all prosodic patterns related to Mandarin are initially screened, forming a set of candidate patterns. The purpose is to narrow the search scope and improve the efficiency of subsequent evaluation.

[0058] In practical applications, evaluating the contextual adaptability of each pattern in the pattern set involves conducting an in-depth analysis of each initially selected candidate prosodic pattern within the context of the text to be pronounced. This evaluation specifically includes analyzing the naturalness of the pattern in the current context—whether the pattern sounds natural and consistent with the pronunciation habits of a native speaker; analyzing its fluency—whether the pattern flows smoothly without abrupt pauses or awkward transitions in speech; and analyzing its semantic fit with the text—whether the pattern accurately expresses the meaning of the text and avoids ambiguity or misunderstanding. Through these dimensions of evaluation, the contextual adaptability assessment result for each pattern can be obtained.

[0059] Furthermore, based on the context adaptability assessment results, a comprehensive matching score is calculated for each pattern in the pattern set. This comprehensive matching score is a quantitative indicator derived by weighting or combining the assessment results of naturalness, fluency, and semantic fit, and is used to measure the overall applicability of each pattern in the current specific context. For example, weights can be assigned to naturalness, fluency, and semantic fit respectively, and then the comprehensive matching score is obtained by multiplying each assessment result by its corresponding weight and summing the results.

[0060] Ultimately, the pattern with the highest overall matching score was selected as the first pattern. This means that among all candidate patterns that have undergone initial screening and context-adaptability assessment, the prosodic pattern that performs best in the context of the current text to be pronounced, best matches learner preferences, and is most natural was selected as the reference for pronunciation comparison.

[0061] Through the aforementioned technical solution, this application significantly improves the accuracy and contextual adaptability of the first mode selection. Compared to screening solely based on target accent or contextual preference, this solution assesses the contextual adaptability of the mode set and calculates a comprehensive matching score, ensuring that the selected first mode achieves optimal naturalness, fluency, and semantic fit with the text. This allows the pronunciation comparison process to more accurately identify learners' pronunciation defects, distinguishing genuine pronunciation problems from natural prosodic variations, thereby providing learners with more personalized and refined pronunciation guidance. Consequently, learners can receive pronunciation feedback that better matches their target accent and specific context, effectively avoiding misjudgments or inefficient learning caused by inappropriate reference modes, and significantly improving the efficiency and effectiveness of Chinese spoken pronunciation learning.

[0062] In some preferred embodiments, a specific example is given below. Suppose a learner wants to practice pronouncing "He likes to eat apples" and sets the target accent to "Mandarin" and the context preference to "everyday conversation".

[0063] First, the system receives this preference information.

[0064] Next, the system performs a preliminary screening of the community prosodic variant knowledge base, filtering out all prosodic patterns related to "Mandarin" and "daily conversation" to form a pattern set. For example, it may filter out pattern A ("apple" with a slightly rising intonation at the end of the sentence, indicating a question), pattern B ("apple" with a steady intonation at the end of the sentence, indicating a statement), and pattern C ("apple" with a slightly falling intonation in the middle of the sentence, indicating emphasis).

[0065] The system then performs a context-adaptability evaluation on each pattern in the pattern set. For the text to be pronounced, "He likes to eat apples," if the context is "You asked him what he likes to eat?", then Pattern A (interrogative intonation) will receive higher evaluation results in terms of naturalness, fluency, and semantic fit. If the context is "He told me he likes to eat apples," then Pattern B (declarative intonation) will receive higher evaluation results. Pattern C (emphasis intonation) may only be applicable in specific emphasis contexts.

[0066] Based on these evaluation results, the system calculates a comprehensive matching score for each pattern. For example, in the context of "You ask him what he likes to eat?", pattern A has the highest comprehensive matching score, followed by pattern B, and pattern C has the lowest.

[0067] Ultimately, the system selects the pattern with the highest overall matching score (e.g., pattern A) as the primary pattern for the specific word "apple" in the current context. In this way, the system can provide learners with a more natural reference pronunciation that not only matches their accent preferences but also closely fits the specific context, making pronunciation comparison and feedback more accurate and effective.

[0068] In some embodiments of this application, when evaluating the contextual adaptability of each pattern in the pattern set, it is necessary to analyze the naturalness, fluency, and semantic fit of the pattern within the context of the text to be pronounced. However, in practical applications, the context of the text to be pronounced may be ambiguous, leading to inaccurate and incomplete evaluation of the patterns.

[0069] To this end, this application further proposes specific steps for evaluating the contextual adaptability of each pattern in the aforementioned pattern set. This contextual adaptability evaluation includes analyzing the naturalness, fluency, and semantic fit of the pattern within the context of the current text to be pronounced, thereby obtaining the contextual adaptability evaluation result. Specifically, this step includes: Semantic analysis is performed on the context of the text to be pronounced to identify ambiguities in the context. For the points of ambiguity, multiple semantic understanding paths are generated, and a contextual representation is constructed for each semantic understanding path; Under the contextual representation corresponding to each semantic understanding path, the naturalness, fluency, and fit with the semantics of the patterns in the pattern set are evaluated to obtain multiple sets of evaluation results; Obtain learners' historical pronunciation data and learning preference information; By analyzing the historical pronunciation data and learning preference information, learners' pronunciation habits and tendencies in understanding specific semantics in similar polysemous contexts are analyzed; Based on learners' pronunciation habits and tendencies, weights are assigned to the evaluation results corresponding to each semantic understanding path; The context adaptability evaluation result of the patterns in the pattern set is calculated by weighting the evaluation results under different semantic understanding paths.

[0070] Specifically, semantic parsing of the context of the text to be pronounced aims to gain a deeper understanding of its underlying meaning and identify ambiguities that may lead to different interpretations. For example, a word or phrase may have multiple meanings in different contexts, and identifying these ambiguities is a prerequisite for accurate evaluation. For each identified ambiguity, the system generates multiple possible semantic understanding paths, each representing a reasonable interpretation of the ambiguity. Subsequently, a corresponding contextual representation is constructed for each semantic understanding path, which helps in the subsequent evaluation to perform pattern adaptability analysis for different semantic understanding perspectives.

[0071] Under the contextual representation corresponding to each semantic understanding path, the naturalness, fluency, and semantic fit of the patterns in the pattern set are evaluated, resulting in multiple sets of evaluation results. This means that the adaptability of the same pronunciation pattern may differ under different semantic understandings. To make the evaluation results more personalized and accurate, this application further obtains learners' historical pronunciation data and learning preference information. By analyzing this data, it is possible to understand learners' pronunciation habits and tendencies toward specific semantic understandings in similar polysemous contexts. For example, some learners may prefer a particular semantic interpretation and reflect this tendency in their pronunciation.

[0072] Based on learners' pronunciation habits and preferences, weights are assigned to the evaluation results corresponding to each semantic comprehension path. This means that the evaluation results of semantic comprehension paths that learners prefer or are more accustomed to will be given higher weights, thus playing a more important role in the final comprehensive evaluation. Finally, the context-adaptive evaluation results of the patterns in the pattern set are calculated by weighted averaging of the evaluation results under different semantic comprehension paths. This weighted averaging method can comprehensively consider the multiple semantic possibilities of the text as well as the learner's personalized characteristics, making the context-adaptive evaluation results more comprehensive and accurate.

[0073] The above-mentioned technical solution effectively addresses the shortcomings of traditional context-adaptive assessments, such as insufficient consideration of textual polysemy and lack of personalized adaptation. This solution not only improves the accuracy and precision of context-adaptive assessments but also provides more targeted pronunciation guidance based on learners' individual differences, thereby significantly enhancing the effectiveness of Chinese spoken pronunciation comparison and the learning experience.

[0074] In some embodiments described above, learners' pronunciation habits and tendencies toward specific semantic understandings in similar polysemous contexts are analyzed using historical pronunciation data and learning preference information. However, in practical applications, for certain specific semantic understanding paths, learners' historical pronunciation behavior data or selection frequency statistics may be sparse and insufficient to support accurate and comprehensive analysis. If this problem is not addressed, it may lead to biased judgments about learners' pronunciation habits and tendencies, thereby affecting the accuracy of pronunciation defect assessments. To address this, this application further proposes a method to optimize the analysis of learners' pronunciation habits and tendencies. Through a data fusion mechanism, it effectively addresses the data sparsity problem, thereby improving the accuracy and robustness of the analysis.

[0075] The steps described above for analyzing learners' pronunciation habits and tendencies in understanding specific semantics in similar polysemous contexts using historical pronunciation data and learning preference information include: The system identifies learners' pronunciation behaviors within the historical pronunciation data, focusing on specific semantic understanding paths within polysemous contexts. Specifically, this step aims to precisely locate and extract pronunciation segments from the learner's past pronunciation records that relate to a particular semantic understanding path within the current polysemous context to be evaluated. For example, when a word has multiple pronunciations in different contexts, the system identifies how the learner pronounced it in similar past contexts.

[0076] When the amount of pronunciation data for a specific semantic understanding path is lower than a preset minimum data amount, pronunciation segments that are similar to the specific semantic understanding path in acoustic features are extracted from the learner's overall pronunciation data to obtain similar pronunciation segments. The preset minimum data amount refers to the minimum number of data samples required to ensure the reliability of the analysis results. When pronunciation data for a specific semantic understanding path is insufficient, the system will search for pronunciation segments similar to that specific path in acoustic features such as pitch, duration, and energy envelope from all of the learner's historical pronunciation data. For example, if a learner lacks pronunciation data for a specific polysemous word, but exhibits similar intonation patterns or stress habits in other pronunciations, these similar pronunciation segments will be extracted as supplementary data.

[0077] Similar pronunciation fragments are used as supplementary pronunciation data for a specific semantic understanding path, and then fused with the pronunciation behavior data of that specific semantic understanding path to obtain fused pronunciation data. Fusion refers to integrating the original pronunciation data for a specific semantic understanding path with the supplementary pronunciation data obtained through similarity extraction to expand the dataset and improve its richness and representativeness.

[0078] Based on the fused pronunciation data, the learner's pronunciation habits under the specific semantic understanding path are analyzed. By using more comprehensive and representative fused pronunciation data, learners' pronunciation patterns, speech rate, intonation, stress, and other habits under the specific semantic understanding path can be captured more accurately.

[0079] Identify the frequency with which learners choose specific semantic comprehension paths for polysemous contexts within the historical pronunciation data. This step aims to statistically analyze learners' tendency to choose different semantic comprehension paths when faced with polysemous contexts, i.e., which semantic comprehension they employ more frequently.

[0080] When the statistical data volume of the selection frequency is lower than a preset minimum statistical data volume, the system extracts learning preferences that are semantically related to the specific semantic understanding path from the learner's overall learning preference information, thus obtaining related learning preferences. Similar to pronunciation data, when the selection frequency data for a specific semantic understanding path is insufficient, the system searches for preference information related to that specific path in semantic category from the learner's overall learning preference information. For example, if learners show a preference for a specific topic or context in reading or listening exercises, this preference information can be considered as supplementary data related to the specific semantic understanding path.

[0081] Relevance learning preferences are used as supplementary preference data for specific semantic understanding paths, and then fused with the selection frequency data of specific semantic understanding paths to obtain fused preference data. The fused preference data can more comprehensively reflect learners' selection tendencies for specific semantic understanding paths in polysemous contexts.

[0082] Based on the fused preference data, the learners' tendency towards specific semantic understanding paths is analyzed. This fused preference data allows for a more accurate assessment of which semantic understanding the learner prefers in polysemous contexts, thus providing more precise contextual information for subsequent pronunciation assessment.

[0083] Through the above technical solution, this application can effectively overcome the limitations of traditional methods in dealing with the sparsity of learners' historical data. Specifically, by intelligently extracting similar or associated supplementary data from the overall data and fusing them when the data volume is insufficient, this solution significantly improves the accuracy and robustness of the analysis of learners' pronunciation habits and semantic understanding tendencies in polysemous contexts. This enables the system to generate more accurate personalized evaluations even when the learners' historical data is insufficient, thereby more accurately judging pronunciation defects or natural prosody variants, avoiding misjudgments or missed judgments caused by insufficient data, and thus enhancing the overall effect and user experience of Chinese spoken language pronunciation comparison.

[0084] In some preferred embodiments, a specific example is described below. Suppose a learner is learning a polysemous word "行" (for example, it can be pronounced as xíng, meaning "walk"; or pronounced as háng, meaning "industry"). When conducting a pronunciation comparison for this learner, the system needs to analyze their pronunciation habits and semantic understanding tendencies for the word "行" in a similar polysemous context.

[0085] First, the system identifies the learner's pronunciation behavior for the word "行" when it means "walk" in the historical pronunciation data. If it is found that the data volume of the learner's past pronunciation of "行 (xíng)" is lower than the preset minimum data volume (for example, less than 5 times), the system will activate the data supplementation mechanism. At this time, the system will extract all pronunciation segments from the learner's overall pronunciation data that have a high degree of similarity with the pronunciation characteristics of "行 (xíng)" in terms of pitch, duration, energy envelope, etc. For example, when the learner pronounces other words with the second tone, their intonation curve and duration distribution may be similar to those of "行 (xíng)", and these segments will be extracted. Subsequently, these similar pronunciation segments are fused with the learner's existing small amount of pronunciation data of "行 (xíng)" to form a more comprehensive dataset for analyzing the learner's pronunciation habits in the semantic context of "walk". [[ID=८]]

[0086] Meanwhile, the system identifies the selection frequency of the learner in the historical learning preference information for the character "行" when it means "walking". If it is found that the statistical data volume of the learner's selection of the semantic path of "walking" is lower than the preset minimum statistical data volume (for example, less than 3 times), the system will extract the learning preferences related to the semantic category of "walking" (such as transportation, sports, etc.) from the learner's overall learning preference information. For example, if the learner frequently selects texts or topics related to transportation and sports in other learning tasks, these preference information will be regarded as supplementary preference data associated with the "walking" semantic path. These associated learning preference data will be fused with the small amount of existing "walking" semantic path selection frequency data of the learner to form a more comprehensive preference dataset for analyzing the learner's tendency for a specific semantic understanding path of "walking".

[0087] In this way, even if the historical data of the learner in a specific polysemous context is sparse, the solution of the present application can comprehensively and accurately analyze their pronunciation habits and semantic understanding tendencies by fusing supplementary data, thereby providing a solid data basis for subsequent pronunciation comparison and defect judgment.

[0088] The present application further proposes that when the amount of pronunciation behavior data of a specific semantic understanding path is lower than the preset minimum data volume, the steps of extracting pronunciation segments with similarity in acoustic features to the specific semantic understanding path from the learner's overall pronunciation data include: For each pronunciation segment in the learner's overall pronunciation data, calculate the similarity values of each pronunciation segment with the acoustic features of the specific semantic understanding path in terms of pitch, duration, and energy envelope; According to the similarity values, perform a preliminary ranking on the pronunciation segments to obtain the preliminarily ranked pronunciation segments; For the preliminarily ranked pronunciation segments, analyze the context matching degree between the context of the pronunciation segment in the current polysemous context and the context of the specific semantic understanding path. The context matching degree includes judging whether the speech rate change pattern of the pronunciation segment is consistent with the speech rate change pattern of the specific semantic understanding path, judging whether the stress distribution of the pronunciation segment is consistent with the stress distribution pattern of the specific semantic understanding path, and judging whether the intonation curve trend of the pronunciation segment is consistent with the intonation curve trend of the specific semantic understanding path; Based on the similarity values of the acoustic features and the context matching degree, select the pronunciation segment with the highest comprehensive matching degree as the similarity pronunciation segment.

[0089] Specifically, when calculating the acoustic feature similarity value between each pronunciation segment and a specific semantic understanding path, various acoustic feature parameters can be used for quantitative comparison. For example, pitch curves can be extracted using the fundamental frequency (F0) and similarity can be measured by dynamic time warping (DTW) or correlation analysis; duration can be compared by comparing duration at the phoneme or syllable level; and energy envelopes can be compared by comparing short-time energy or loudness curves. The calculation of these similarity values ​​aims to initially screen pronunciation segments that are similar to the target path at the basic acoustic level. The initially sorted pronunciation segments refer to the set of pronunciation segments arranged from high to low according to the aforementioned acoustic feature similarity values, with the purpose of providing a more promising candidate set for subsequent contextual matching analysis.

[0090] In practical applications, the key to this scheme lies in analyzing the degree of contextual matching between the pronunciation segment and the specific semantic understanding path within the current polysemous context. The assessment of contextual matching goes beyond acoustic features, taking into account the prosody and pragmatic information of the pronunciation. Specifically, the consistency of speech rate variation patterns can be judged by comparing prosodic features such as the relative duration and pause length of syllables or words in the pronunciation segment; the consistency of stress distribution patterns can be determined by analyzing the correspondence between the pitch, duration, and energy of stressed syllables or words in the pronunciation segment and the target path; and the consistency of intonation curves can be determined by comparing whether the overall intonation profile of the pronunciation segment (such as rising, falling, or level intonation) matches the intonation pattern of the target path. These judgments of contextual matching aim to ensure that the selected supplementary pronunciation segments are not only acoustically similar but also highly consistent with the target path in terms of expressed context and semantics. Therefore, by combining the similarity scores of acoustic features with the degree of contextual matching, a comprehensive evaluation can be conducted on each initially ranked pronunciation segment using methods such as weighted summation, multi-dimensional scoring, or machine learning models, thus obtaining a comprehensive matching score. Selecting the pronunciation segment with the highest comprehensive matching score as the most similar pronunciation segment ensures that the extracted supplementary data has the highest relevance to the specific semantic understanding path in both acoustic and contextual dimensions.

[0091] In some preferred embodiments, it is assumed that when a learner is learning the word "apple," the amount of pronunciation data in the two polysemous contexts of "eating apples" and "Apple Inc." is very small, making it impossible to directly analyze their pronunciation habits under a specific semantic understanding path. In this case, the system needs to supplement the learner's past pronunciation data by finding similar pronunciation segments. Specifically, firstly, the system traverses the learner's overall pronunciation data and, for each pronunciation segment, calculates the similarity value of its acoustic features (such as pitch, duration, and energy envelope) with the semantic understanding path of "eating apples." For example, a pronunciation segment of "eating pears" may be very close to the pitch, duration, and energy envelope curve of "eating apples" acoustically, thus having a high similarity value. Secondly, the system performs a preliminary sorting of the pronunciation segments based on these similarity values, prioritizing the most acoustically similar segments. Thirdly, for these preliminarily sorted pronunciation segments, the system further analyzes the degree of contextual matching between their current polysemous context and the specific semantic understanding path of "eating apples." For example, for the pronunciation segment "eat a pear," the system evaluates whether its speech rate variation pattern, stress distribution, and intonation curve are consistent with the contextual features of "eat an apple" when expressing the action of "eating." If the pronunciation of "eat a pear" reflects a similar, action-emphasizing expression in speech rate, stress, and intonation as "eat an apple," then its contextual matching degree will be high. Conversely, if a pronunciation segment is acoustically similar but expresses a question or emphasizes other information in the context, its contextual matching degree will be low. Finally, the system combines the similarity values ​​of acoustic features and the contextual matching degree to calculate a comprehensive matching score for each pronunciation segment. For example, using a weighted average, acoustic similarity is assigned a weight of 0.6, and contextual matching degree is assigned a weight of 0.4. Ultimately, the pronunciation segment with the highest comprehensive matching score, such as the pronunciation segment of "eat a pear," is selected as a supplementary similar pronunciation segment for the specific semantic understanding path of "eat an apple." In this way, it is ensured that the supplementary data are not only similar in pronunciation, but also highly relevant in semantics and context, thereby improving the accuracy of the analysis of learners' pronunciation habits.

[0092] Secondly, referring to Figure 2 This application further proposes a Chinese spoken pronunciation comparison system based on speech recognition, the system comprising: The text information acquisition module 201 is used to acquire the text information to be pronounced; The linguistic rule analysis module 202 is used to perform linguistic rule analysis on the text information to obtain linguistic analysis results; the linguistic rule analysis refers to the identification and analysis of pronunciation rules unique to Chinese. The speech organ simulation module 203 is used to simulate the continuous movement trajectory and acoustic adjustment rules of the speech organs based on the linguistic analysis results, and obtain simulation results; The standard reference speech synthesis module 204 is used to generate a standard reference speech containing speech cohesion and prosodic patterns based on the simulation results. The actual pronunciation acquisition module 205 is used to acquire the learner's actual pronunciation and preprocess the actual pronunciation to obtain the preprocessed actual pronunciation; The acoustic feature extraction module 206 is used to extract acoustic feature sequences from the preprocessed actual pronunciation and to align the acoustic feature sequences of the learner's pronunciation with the standard reference speech in time. The pronunciation defect judgment module 207 is used to compare the acoustic features of the learner's pronunciation with the acoustic features of the standard reference speech based on the time alignment, obtain the difference between the acoustic features of the learner's pronunciation and the acoustic features of the standard reference speech, and determine whether the difference is a pronunciation defect based on the natural speech changes reflected in the standard reference speech.

[0093] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for comparing spoken Chinese pronunciation based on speech recognition, characterized in that, include: Obtain the text information to be pronounced; The text information is subjected to linguistic rule analysis to obtain linguistic analysis results; The linguistic rule analysis refers to the identification and analysis of pronunciation rules unique to Chinese. Based on the linguistic analysis results, the continuous movement trajectory and acoustic adjustment rules of the vocal organs were simulated to obtain simulation results; Based on the simulation results, a standard reference speech containing speech cohesion and prosodic patterns is generated; Collect the learner's actual pronunciation and preprocess the actual pronunciation to obtain the preprocessed actual pronunciation; Extract acoustic feature sequences from the preprocessed actual pronunciation; The acoustic feature sequence of learners' pronunciation is temporally aligned with that of a standard reference speech. Based on time alignment, the acoustic features of the learner's pronunciation are compared with the acoustic features of the standard reference speech to obtain the differences between the acoustic features of the learner's pronunciation and the acoustic features of the standard reference speech. Based on the natural speech changes reflected in the standard reference speech, it is determined whether the differences are pronunciation defects.

2. The method for comparing spoken Chinese pronunciation based on speech recognition according to claim 1, characterized in that, The steps for simulating the continuous movement trajectory and acoustic adjustment rules of the vocal organs based on the linguistic analysis results to obtain the simulation results include: Identify specific words contained in text information; Query the prosodic pattern corresponding to the specific word; the prosodic pattern includes the natural prosodic information of the specific word in real speech flow; Based on the prosodic pattern, the predicted parameters of the articulatory organ movement are adjusted to obtain the adjusted predicted parameters of the articulatory organ movement; the predicted parameters include the starting point and rate of change of the pitch curve, or the duration distribution of syllables and the shape of the energy envelope; Based on the adjusted predicted parameters of the vocal organ movement, the continuous movement trajectory and acoustic adjustment law of the vocal organ are simulated to obtain simulation results.

3. The method for comparing spoken Chinese pronunciation based on speech recognition according to claim 1, characterized in that, The step of comparing the acoustic features of the learner's pronunciation with those of the standard reference speech based on time alignment, obtaining the differences between the acoustic features of the learner's pronunciation and those of the standard reference speech, and determining whether the differences constitute pronunciation defects based on the natural speech changes reflected in the standard reference speech includes: Extract prosodic parameters from learners' pronunciation; Extract prosodic parameters from the standard reference speech; The prosodic parameters of the learner's pronunciation are analyzed to obtain the prosodic parameter differences between the prosodic parameters of the standard reference speech. Identify specific words contained in text information; The query retrieves the first pattern corresponding to the specific word; the first pattern refers to the natural prosodic information of the specific word that is generally accepted in the real native speaker community and has subtle differences from the general prosodic pattern; When the difference in the prosodic parameters exceeds the preset general tolerance range, it is determined whether the prosodic pattern of the learner's pronunciation matches the first pattern, and a judgment result is obtained; Based on the judgment result, the learner's pronunciation is determined to be a natural rhythmic variation or a pronunciation defect; the natural rhythmic variation refers to natural rhythmic changes.

4. The method for comparing spoken Chinese pronunciation based on speech recognition according to claim 3, characterized in that, The step of determining whether the learner's pronunciation prosodic pattern matches the first pattern when the difference in prosodic parameters exceeds a preset general tolerance range, and obtaining the determination result, includes: Calculate the degree of matching between the learner's pronunciation prosodic pattern and the first pattern; The matching degree is compared with a preset fuzzy matching threshold to obtain a comparison result; Based on the comparison results, it is determined whether the learner's pronunciation rhythm pattern matches the first pattern.

5. The method for comparing spoken Chinese pronunciation based on speech recognition according to claim 3, characterized in that, The steps for querying the first pattern corresponding to the specific term include: Receive learner-set target accent or contextual preference information; Based on the target accent or contextual preference information, a first pattern matching the target accent or contextual preference information is selected from the community prosodic variant knowledge base.

6. The method for comparing spoken Chinese pronunciation based on speech recognition according to claim 5, characterized in that, The step of selecting a first pattern that matches the target accent or contextual preference information from the community prosodic variant knowledge base based on the target accent or contextual preference information includes: Receive learner-set target accent or contextual preference information; A preliminary screening of patterns in the community prosody variant knowledge base yields a pattern set. Each pattern in the pattern set is subjected to a contextual adaptability assessment, which includes analyzing the naturalness, fluency, and semantic fit of the pattern in the context of the current text to be pronounced, and obtaining the contextual adaptability assessment result. Based on the context adaptability assessment results, calculate the comprehensive matching score for each pattern in the pattern set; Select the pattern with the highest overall matching score as the first pattern.

7. A method for comparing spoken Chinese pronunciation based on speech recognition according to claim 6, characterized in that, The step of performing a context-adaptability assessment on each pattern in the pattern set, which includes analyzing the naturalness, fluency, and semantic fit of the pattern in the context of the current text to be pronounced, and obtaining the context-adaptability assessment result, includes: Semantic analysis is performed on the context of the text to be pronounced to identify ambiguities in the context. For the points of ambiguity, multiple semantic understanding paths are generated, and a contextual representation is constructed for each semantic understanding path; Under the contextual representation corresponding to each semantic understanding path, the naturalness, fluency, and fit with the semantics of the patterns in the pattern set are evaluated to obtain multiple sets of evaluation results; Obtain learners' historical pronunciation data and learning preference information; By analyzing the historical pronunciation data and learning preference information, learners' pronunciation habits and tendencies in understanding specific semantics in similar polysemous contexts are analyzed; Based on learners' pronunciation habits and tendencies, weights are assigned to the evaluation results corresponding to each semantic understanding path; The context adaptability evaluation result of the patterns in the pattern set is calculated by weighting the evaluation results under different semantic understanding paths.

8. The method for comparing spoken Chinese pronunciation based on speech recognition according to claim 7, characterized in that, The steps of analyzing learners' pronunciation habits and tendencies in understanding specific semantics in similar polysemous contexts using the historical pronunciation data and learning preference information include: Identify learners' pronunciation behaviors in the historical pronunciation data, specifically for the semantic understanding path of polysemous contexts; When the amount of pronunciation behavior data for a specific semantic understanding path is lower than the preset minimum amount of data, pronunciation segments that are similar to the specific semantic understanding path in terms of acoustic features are extracted from the learner's overall pronunciation data to obtain similar pronunciation segments. Similar pronunciation segments are used as supplementary pronunciation data for a specific semantic understanding path, and then fused with the pronunciation behavior data of the specific semantic understanding path to obtain fused pronunciation data; Based on the fused pronunciation data, the learner's pronunciation habits under the specific semantic understanding path are analyzed; Identify the frequency with which learners select specific semantic comprehension paths for polysemous contexts within the historical pronunciation data; When the statistical data of the selected frequency is lower than the preset minimum statistical data, the learning preferences that are related to the specific semantic understanding path in terms of semantic category are extracted from the learner's overall learning preference information to obtain the related learning preferences; The relevance learning preference is used as supplementary preference data for a specific semantic understanding path, and then fused with the selection frequency data of the specific semantic understanding path to obtain the fused preference data. Based on the fused preference data, the learner's tendency toward a specific semantic understanding path is analyzed.

9. A method for comparing spoken Chinese pronunciation based on speech recognition according to claim 8, characterized in that, The step of extracting pronunciation segments that are similar in acoustic features to the specific semantic understanding path from the learner's overall pronunciation data when the amount of pronunciation behavior data for a specific semantic understanding path is lower than a preset minimum amount of data includes: For each pronunciation segment in the learner's overall pronunciation data, calculate the similarity values ​​between each pronunciation segment and the acoustic features of the specific semantic understanding path in terms of pitch, duration, and energy envelope; Based on the similarity values, the pronunciation segments are initially sorted to obtain the initially sorted pronunciation segments; For the initially sorted pronunciation segments, the contextual matching degree between the pronunciation segments in the current polysemous context and the specific semantic understanding path is analyzed. The contextual matching degree includes determining whether the speech rate change pattern of the pronunciation segment is consistent with the speech rate change pattern of the specific semantic understanding path, determining whether the stress distribution of the pronunciation segment is consistent with the stress distribution pattern of the specific semantic understanding path, and determining whether the intonation curve of the pronunciation segment is consistent with the intonation curve of the specific semantic understanding path. Based on the similarity values ​​of the combined acoustic features and the degree of contextual matching, the pronunciation segment with the highest overall matching degree is selected as the similar pronunciation segment.

10. A Chinese spoken pronunciation comparison system based on speech recognition, characterized in that, The system includes: The text information acquisition module is used to acquire the text information to be pronounced; The linguistic rule analysis module is used to perform linguistic rule analysis on the text information to obtain linguistic analysis results; the linguistic rule analysis refers to the identification and analysis of pronunciation rules unique to Chinese. The speech organ simulation module is used to simulate the continuous movement trajectory and acoustic adjustment rules of the speech organs based on the linguistic analysis results, and obtain simulation results; A standard reference speech synthesis module is used to generate standard reference speech containing speech cohesion and prosodic patterns based on the simulation results; The actual pronunciation acquisition module is used to acquire the learner's actual pronunciation and preprocess the actual pronunciation to obtain the preprocessed actual pronunciation; The acoustic feature extraction module is used to extract acoustic feature sequences from the preprocessed actual pronunciation and to align the acoustic feature sequences of the learner's pronunciation with the standard reference speech in time. The pronunciation defect judgment module is used to compare the acoustic features of the learner's pronunciation with the acoustic features of the standard reference speech based on time alignment, obtain the differences between the acoustic features of the learner's pronunciation and the acoustic features of the standard reference speech, and determine whether the differences are pronunciation defects based on the natural speech changes reflected in the standard reference speech.