Method and system for automatic detection of tone sandhi in continuous speech in Chinese

By extracting fundamental frequency transition trajectory features in the connection region between adjacent syllables and combining them with a hierarchical tone sandhi rule knowledge base, the problems of accuracy and insufficient feedback in continuous speech tone sandhi detection in Chinese speech learning are solved, improving detection accuracy and providing intuitive learning guidance.

CN121545552BActive Publication Date: 2026-04-28SICHUAN NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SICHUAN NORMAL UNIV
Filing Date
2026-01-20
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately detect and assess tone sandhi in continuous speech flow during Chinese phonetics learning. Syllable boundary localization is inaccurate, tone sandhi feature extraction is incomplete, and there is a lack of hierarchical processing and intuitive feedback, resulting in impractical assessment results.

Method used

By extracting fundamental frequency transition trajectory features in the connection region between adjacent syllables and combining them with a hierarchical tone sandhi rule knowledge base, the system can accurately detect and evaluate tone sandhi patterns in continuous speech flow, and generate speech flow pitch curve annotation diagrams and feedback content.

Benefits of technology

It improves the accuracy of tone sandhi detection by approximately 15% to 20%, provides intuitive and effective feedback, helps learners understand and improve tone sandhi, and can be applied to Chinese phonetics teaching and Mandarin proficiency testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121545552B_ABST
    Figure CN121545552B_ABST
Patent Text Reader

Abstract

The application discloses a Chinese continuous speech tone change automatic detection method and system, and belongs to the technical field of speech signal processing.The method comprises the following steps: a speech forced alignment step is used to acquire syllable time boundaries; a fundamental frequency transition trajectory extraction step is used to extract a fundamental frequency feature vector in an adjacent syllable connection area; a tone change rule matching step is used to acquire an expected mode from a knowledge base containing necessary and variable rules; a tone change mode detection step is used to calculate a matching score and determine whether the tone change is correct, missing or excessive; and a feedback generation step is used to generate a pitch curve labeling graph and rule explanation.The application realizes accurate tone change detection by focusing on the fundamental frequency transition features of the connection area, provides reasonable evaluation by using a hierarchical rule knowledge base, and helps learners to improve pronunciation by visual feedback.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech signal processing, and particularly to a technology for automatically detecting and evaluating the continuous speech tone sandhi patterns for learners of Chinese as a second language, specifically to a method and system for automatically detecting speech tone sandhi in Chinese continuous speech. Background Art

[0002] Chinese is a typical tonal language, and tones play an important role in differentiating word meanings in Chinese. Mandarin Chinese includes four basic tones: high level, rising, falling-rising, and falling. Different tones can make the same syllable express completely different semantic meanings. In the process of learning Chinese pronunciation, learners not only need to master the four basic tones of single characters but also need to master the complex tone sandhi rules in continuous speech. Tone sandhi refers to the regular changes in the tones of adjacent syllables affected by each other in continuous speech, which is one of the important features of Chinese pronunciation. Common tone sandhi phenomena include the tone sandhi of three consecutive third tones, the tone sandhi of the character "一", the tone sandhi of the character "不", and light tone weakening. The correct implementation of these tone sandhi rules is an important guarantee for natural and fluent Chinese pronunciation.

[0003] For learners of Chinese as a second language, the acquisition of tone sandhi rules is the key point and difficulty in pronunciation learning. Research shows that although many learners can correctly pronounce each tone when reading isolated characters, they often have difficulty accurately implementing tone sandhi in continuous speech. There are many reasons for this phenomenon: Firstly, the triggering conditions of tone sandhi rules are relatively complex, involving various factors such as the tone types of adjacent syllables, word boundaries, and grammatical structures; Secondly, the implementation of tone sandhi requires switching tones within a short time, which poses high requirements for the coordinated control of the pronunciation organs; Moreover, the prosodic features of the learners' mother tongues may have a negative transfer effect on Chinese tone sandhi.

[0004] In the prior art, CN113571037A discloses a Chinese braille speech synthesis method and system. This technical solution converts a general braille text into a pinyin sequence, combines a prosody prediction model to obtain prosody labels, and finally realizes speech synthesis. In terms of tone sandhi processing, this solution mainly uses a rule-based method to convert a normal pinyin sequence into a tone sandhi pinyin sequence. For example, when processing the connection of two third tones, the first third tone changes to a rising tone, and the tone sandhi rules of the characters "一" and "不" in different tone environments are processed. This technical solution constructs a tone sandhi dictionary for function words to deal with the problem of eliminating light tone ambiguity and realizes the automatic conversion from normal pinyin to tone sandhi pinyin. However, this technical solution mainly focuses on the problem of tone sandhi generation in the field of speech synthesis, and its goal is to ensure the correctness of tone sandhi in the synthesized speech, rather than detecting and evaluating the implementation of tone sandhi in learners' speech. This solution lacks the ability to extract and analyze the tone sandhi features in actual speech signals and cannot judge whether learners have correctly implemented the expected tone sandhi.

[0005] In the field of computer-aided language learning, speech assessment technology has been widely applied. Traditional speech assessment methods mainly target the identification and scoring of tones in isolated words or monosyllables, using the matching degree between the fundamental frequency trajectory and a standard template as the evaluation criterion. These methods are effective in processing single-word tones, but they have significant shortcomings when dealing with continuous speech. Tone performance in continuous speech is affected by a variety of factors, including the coarticulation effect of adjacent syllables, tone compression caused by changes in speech rate, and pitch adjustment due to stress position. These factors make it difficult for simple template matching methods to accurately determine the implementation of tone sandhi. In addition, most existing tone assessment systems use tone classification methods based on hidden Markov models or deep neural networks, classifying each syllable independently into one of four tone categories. This method ignores the essential characteristic of tone sandhi as a cross-syllable phenomenon.

[0006] Current research on tone sandhi detection in continuous speech still faces the following technical challenges. First, the syllable boundary localization in tone sandhi detection is not precise enough. Existing methods often rely on manual annotation or fixed time window segmentation, which cannot adapt to changes in syllable boundaries under different speaking speeds and styles, leading to boundary errors affecting subsequent tone sandhi feature analysis. Second, tone sandhi feature extraction is not comprehensive enough. Existing methods mainly focus on the fundamental frequency trajectory shape within syllables, neglecting the fundamental frequency transition features in the transition region between adjacent syllables. This region is precisely the key to tone sandhi, as its acoustic manifestation is mainly reflected in the fundamental frequency change pattern in the transition interval from the end of the previous syllable to the beginning of the next syllable. Third, the processing of tone sandhi rules lacks hierarchy. Existing methods typically treat all tone sandhi rules equally, failing to distinguish between mandatory rules that must be followed and optional rules that allow for acceptable changes. This results in evaluations that are either too strict or too lenient, reducing their practical value. Fourth, the feedback mechanism is not intuitive or effective enough. Existing methods often present evaluation results in the form of numerical scores, making it difficult for learners to understand specific problems and how to improve, lacking targeted guidance and demonstration examples.

[0007] Therefore, there is an urgent need for an automatic speech tone detection method and system that can accurately detect the implementation of tone shifts in continuous speech, distinguish different types of tone shift rules, and provide intuitive and effective feedback. Summary of the Invention

[0008] To address the shortcomings of existing technologies, this invention provides an automatic detection method and system for tone sandhi in continuous Chinese speech. By extracting fundamental frequency transition trajectory features in the connection area between adjacent syllables and combining them with a hierarchical tone sandhi rule knowledge base, it achieves accurate detection and effective evaluation of tone sandhi patterns in continuous speech.

[0009] The automatic detection method for tone sandhi in Chinese continuous speech provided by the present invention includes the following steps. In the speech forced alignment step, the speech signal to be detected and the corresponding reference text are obtained, and forced alignment processing is performed to determine the time boundaries of each syllable, generating a syllable time boundary sequence. In the fundamental frequency transition trajectory extraction step, based on the syllable time boundary sequence, the connection area between adjacent syllables is defined, and the fundamental frequency transition trajectory is extracted within the connection area, calculating the starting value of the fundamental frequency, the ending value of the fundamental frequency, the slope of the fundamental frequency change, and the amplitude of the fundamental frequency change, which are combined into a fundamental frequency transition trajectory feature vector. In the tone sandhi rule matching step, the tone sandhi rule corresponding to the current adjacent syllable combination is obtained from the tone sandhi rule knowledge base, which includes two types of rules: mandatory rules and variable rules, and the expected fundamental frequency transition pattern is determined according to the tone sandhi rules. In the tone sandhi pattern detection step, the fundamental frequency transition trajectory feature vector is matched with the expected fundamental frequency transition pattern, calculating the tone sandhi matching score, and the tone sandhi detection result is determined according to the tone sandhi matching score, locating the specific syllable position with tone sandhi abnormality. In the feedback generation step, a pitch curve annotation map of the speech flow is generated according to the tone sandhi detection result, differentiating the correctly toned section, the missing toned section, and the over-toned section with different colors, and at the same time generating rule explanation text and comparative demonstration audio.

[0010] Preferably, the connection area is defined with the ending time point of the previous syllable as the center, extending forward by a first preset duration and backward by a second preset duration to form a connection area time window, where the value range of the first preset duration is from 30 milliseconds to 80 milliseconds, and the value range of the second preset duration is from 30 milliseconds to 80 milliseconds.

[0011] Preferably, the mandatory rules in the tone sandhi rule knowledge base include the three-tone consecutive tone sandhi rule, the one-character tone sandhi rule, and the not-character tone sandhi rule, and the variable rules include the weakening rule of light tone and the dialect tone sandhi rule.

[0012] Preferably, in the tone sandhi pattern detection step, for the detection of three-tone consecutive tone sandhi, a classifier based on a deep neural network is also used for auxiliary determination to improve the detection accuracy.

[0013] The automatic detection system for tone sandhi in Chinese continuous speech provided by the present invention includes a speech forced alignment module, a fundamental frequency transition trajectory extraction module, a tone sandhi rule knowledge base, a tone sandhi pattern detection engine, and a feedback generation module. The speech forced alignment module is used to perform the forced alignment of speech and text and generate a syllable time boundary sequence. The fundamental frequency transition trajectory extraction module is used to define the connection area and extract the fundamental frequency transition trajectory feature vector. The tone sandhi rule knowledge base is used to store mandatory rules and variable rules. The tone sandhi pattern detection engine is used to match the tone sandhi pattern and determine the detection result. The feedback generation module is used to generate a visual annotation map and multimedia feedback content.

[0014] The beneficial effects of this invention include the following aspects. First, by focusing on the extraction of fundamental frequency transition trajectory features in the connection region between adjacent syllables, it can more accurately capture the key acoustic cues of pitch sandhi, improving the accuracy of pitch sandhi detection by approximately 15% to 20% compared to traditional syllable internal feature analysis methods. Second, by establishing a hierarchical knowledge base of pitch sandhi rules, setting mandatory pitch sandhi rules as mandatory detection items and variable rules as reference detection items, the evaluation results are more reasonable, reducing misjudgments of learners. Third, by generating speech flow pitch curve annotation diagrams and comparative demonstration audio, it provides learners with intuitive and effective feedback, helping them understand the problems and make targeted improvements. Fourth, the method and system of this invention can be widely applied in fields such as Chinese phonetics teaching, Mandarin proficiency testing assistance, and speech therapy, and has significant application value. Attached Figure Description

[0015] Figure 1 This is a flowchart of the automatic detection method for tone sandhi in continuous Chinese speech flow according to the present invention.

[0016] Figure 2 This is an architecture diagram of the automatic detection system for tone sandhi in continuous Chinese speech flow according to the present invention. Detailed Implementation

[0017] Please refer to the attached document. Figures 1-2 The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. It should be noted that the following embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.

[0018] The overall process of the automatic detection method for tone sandhi in continuous Chinese speech provided by this invention is as follows: Figure 1 As shown, this method achieves automatic detection and evaluation of pitch shift patterns in continuous speech through five core steps: forced speech alignment, fundamental frequency transition trajectory extraction, pitch shift rule matching, pitch shift pattern detection, and feedback generation. A deeply coupled closed-loop collaborative architecture is formed among these steps, with the output of the previous step serving as the key input for the next. The detection results of subsequent steps can inversely influence the parameter adjustments of preceding steps, thereby achieving continuous optimization of detection accuracy.

[0019] The forced speech alignment module 1 is responsible for executing the forced speech alignment step, which is a fundamental step in the entire detection process. In one embodiment of the present invention, the forced speech alignment module 1 acquires the speech signal to be detected and the corresponding reference text as input data. The speech signal to be detected is typically an audio file recorded by a learner reading a specific Chinese text, with a sampling rate set to 16kHz, a quantization precision of 16 bits, and mono recording to ensure signal quality. The reference text is the standard Chinese character text corresponding to the speech signal to be detected, which can be provided in advance by the teaching system or manually input by the learner.

[0020] The specific process of the forced alignment module 1 performing forced alignment processing on the speech signal to be detected is as follows. First, the reference text is converted into a pinyin sequence and further decomposed into a phoneme sequence. In this embodiment, an acoustic model based on a deep neural network is used to align the speech with the phoneme sequence. This acoustic model can be a pre-trained Kaldi model or an end-to-end model based on connection temporal classification. Under the constraint of the given reference text, the forced alignment algorithm calculates the correspondence between each time frame in the speech signal and each phoneme in the phoneme sequence, thereby determining the start and end time points of each phoneme. Further, the time boundaries of phonemes belonging to the same syllable are merged to determine the start and end time points of each syllable in the speech signal to be detected, generating a syllable time boundary sequence.

[0021] Preferably, during the forced alignment process, for speech segments with fast speech rates or connected speech, the speech forced alignment module 1 employs a dynamic time warping algorithm to fine-tune the initial alignment results. Specifically, using the syllable energy change curve and fundamental frequency fluctuation point as auxiliary references, the initially determined syllable boundaries are fine-tuned to ensure that the syllable boundaries fall more accurately on the natural boundaries of acoustic features. In this embodiment, the time accuracy of the syllable boundaries is controlled within ±20 milliseconds, which meets the accuracy requirements for subsequent fundamental frequency transition trajectory extraction.

[0022] The fundamental frequency transition trajectory extraction module 2 is responsible for performing the fundamental frequency transition trajectory extraction step, which is one of the core innovative aspects of this invention. Traditional tone analysis methods mainly focus on the shape of the fundamental frequency trajectory within a syllable, while this invention creatively shifts the focus of analysis to the connection area between adjacent syllables, because the acoustic performance of tone shifting is mainly reflected in the fundamental frequency transition pattern between adjacent syllables.

[0023] The fundamental frequency transition trajectory extraction module 2 delineates the transition region between adjacent syllables based on the syllable time boundary sequence. The specific delineation method is as follows: Let the end time point of the i-th syllable be... The starting time of the (i+1)th syllable is Then the connecting region of the adjacent syllable pair is defined as... Centered on the first preset duration, extend forward. Extend the second preset duration backward The resulting time window. In this embodiment, the first preset duration... The value range is from 30 milliseconds to 80 milliseconds, the second preset duration. The value ranges from 30 milliseconds to 80 milliseconds. Preferably, The value is 50 milliseconds. A value of 50 milliseconds is used to create a transition window with a total duration of 100 milliseconds. This duration setting effectively covers the key intervals of pitch shifts while avoiding the introduction of excessive syllable internal information.

[0024] It is worth noting that the delineation of the connecting region needs to take into account the short pauses or overlapping coarticulation that may exist between syllables in actual speech. When When the interval exceeds the preset threshold, it indicates a significant pause between the two syllables. In this case, the second preset duration should be shortened to avoid including the pause interval in the analysis. When there is a time interval, it indicates that there is syllable overlap due to co-pronunciation. In this case, the midpoint between the two time points is used as the center of the transition area. The preset interval threshold is set to 150 milliseconds in this embodiment.

[0025] After defining the transition region, the fundamental frequency transition trajectory extraction module 2 extracts the fundamental frequency trajectory within that region from the original speech signal. Fundamental frequency extraction employs either the autocorrelation method or the PYIN algorithm, with the fundamental frequency detection frequency range set from 75Hz to 500Hz to cover the fundamental frequency range of speakers of different genders and ages. The extracted original fundamental frequency sequence may contain fundamental frequency jump points or missing values, thus requiring preprocessing. First, the original fundamental frequency sequence undergoes median filtering with a filter window length of 5 sampling points to remove fundamental frequency jump points caused by vocal cord vibration instability or the transition from voiced to unvoiced sounds. Second, missing values ​​are filled using linear interpolation to ensure the continuity of the fundamental frequency sequence.

[0026] Furthermore, the fundamental frequency transition trajectory extraction module 2 performs speaker normalization processing on the filtered fundamental frequency sequence to eliminate individual pitch differences between different speakers. In this embodiment, two normalization methods are provided for selection. The first method is the Z-score normalization method, which calculates the speaker's fundamental frequency mean and standard deviation throughout the entire speech segment, converting each fundamental frequency value into a corresponding Z-score. The calculation formula for this method is:

[0027] ,

[0028] in, This is the normalized fundamental frequency value. This is the original fundamental frequency value. The speaker's fundamental frequency mean. The standard deviation of the speaker's fundamental frequency. Indexed by time point.

[0029] The second method is the semitone conversion method, which converts the fundamental frequency value from Hertz units to a semitone value referenced to the speaker's reference pitch. The calculation formula for this method is:

[0030] ,

[0031] in, The converted semitone value. This is the original fundamental frequency value. The reference fundamental frequency of the speaker can be taken as the 5th percentile of the speaker's fundamental frequency distribution.

[0032] After normalization, the fundamental frequency transition trajectory extraction module 2 calculates multi-dimensional features based on the fundamental frequency transition trajectory within the transition region. The fundamental frequency transition trajectory feature extraction algorithm proposed in this invention includes feature calculation in the following four dimensions.

[0033] The first dimension is the starting value of the fundamental frequency. , defined as the mean fundamental frequency of the first half of the transition region, is calculated using the following formula:

[0034] ,

[0035] in, This is the sampling index corresponding to the starting time point of the connection region. This represents the number of sampling points in the first half of the connecting area.

[0036] The second dimension is the fundamental frequency termination value. , defined as the mean fundamental frequency of the latter half of the transition region, is calculated using the following formula:

[0037] ,

[0038] in, This represents the number of sampling points in the latter half of the connecting area.

[0039] The third dimension is the slope of the fundamental frequency change. The least squares method is used to linearly fit the fundamental frequency sequence within the transition region. The slope of the fitted line is taken as the slope of the fundamental frequency change. The calculation formula is as follows:

[0040] ,

[0041] in, This represents the total number of sampling points within the connecting area. For the first The time value of each sampling point The average over time. For the first Normalized fundamental frequency value of each sampling point This is the mean of the fundamental frequency.

[0042] The fourth dimension is the amplitude of fundamental frequency variation. , defined as the difference between the fundamental frequency termination value and the fundamental frequency start value, is calculated using the following formula:

[0043] ,

[0044] The fundamental frequency transition trajectory extraction module 2 combines the feature values ​​of the above four dimensions into a fundamental frequency transition trajectory feature vector. This feature vector will serve as input data for subsequent pitch shift pattern detection.

[0045] In a preferred embodiment of the present invention, the fundamental frequency transition trajectory extraction module 2 further calculates extended features to improve detection accuracy. These extended features include fundamental frequency curvature, fundamental frequency jitter, and energy change rate. Fundamental frequency curvature reflects the degree of bending of the fundamental frequency trajectory and is obtained by performing second-order differencing on the fundamental frequency sequence and calculating the mean. Fundamental frequency jitter reflects the micro-fluctuations of the fundamental frequency and is obtained by calculating the standard deviation of the fundamental frequency difference between adjacent sampling points. The energy change rate reflects the trend of speech energy change within the transition region and is obtained by calculating the slope of the short-time energy sequence. Adding these extended features to the feature vector can further improve the accuracy of pitch shift detection by approximately 5%.

[0046] The tone sandhi rule knowledge base 3 stores various tone sandhi rules in continuous Chinese speech, serving as the data foundation for the tone sandhi rule matching step. In the design of this invention, the tone sandhi rule knowledge base 3 adopts a hierarchical architecture, dividing tone sandhi rules into two levels: mandatory rules and variable rules. This design reflects the mandatory differences in tone sandhi rules, enabling the detection results to better conform to linguistic principles and practical teaching needs.

[0047] Mandatory tone sandhi rules refer to the tone sandhi rules that must be followed in standard Mandarin. Violating these rules will lead to obvious pronunciation errors, so they are set as mandatory detection items in the tone sandhi rule knowledge base 3. The mandatory tone sandhi rules in the knowledge base 3 include the following three categories.

[0048] The first category is the tone sandhi rule for three consecutive tones. This rule stipulates that when two consecutive third tone syllables are connected, the preceding syllable changes from a falling-rising tone to a rising tone, that is, the tone value changes from 214 to 35. This is the most typical tone sandhi phenomenon in Chinese and also the tone sandhi rule that has the greatest impact on learners. From a phonetic perspective, the tone sandhi of three consecutive tones occurs because two consecutive low falling tones are difficult to pronounce clearly in fast speech. The human vocal organs tend to simplify the first third tone into a rising tone to reduce the difficulty of pronunciation. In the tone sandhi rule knowledge base 3, the expected fundamental frequency transition pattern of the three consecutive tone sandhi rule is defined as: the fundamental frequency at the end of the preceding syllable should show an upward trend, and the slope of the fundamental frequency change is... It should be greater than the preset positive slope threshold. The amplitude of fundamental frequency variation It should be within the preset range of increase. Inside. In this embodiment, Set to 0.3, Set to 0.5. Set to 2.5, these values are determined based on the statistical analysis of the standard Mandarin corpus. It should be noted that the application scope of the trisyllabic tone sandhi rule is not limited to within disyllabic words, but also includes the case of three consecutive tones across word boundaries. For example, in the phrase "nǐ hǎo ma", the characters "nǐ" and "hǎo" appear consecutively, and the tone of the character "nǐ" should change.

[0049] The second category is the tone sandhi rule for the character "yī". "Yī" is one of the most frequently used numerals in Chinese, and its tone sandhi rule is relatively complex. When "yī" is followed by a fourth tone syllable, the tone of "yī" changes from the original high-level tone to a rising tone, that is, the tone value changes from 55 to 35. When "yī" is followed by a first tone, second tone or third tone syllable, the tone of "yī" changes from the high-level tone to a falling tone, that is, the tone value changes from 55 to 51. In the tone sandhi rule knowledge base 3, the tone sandhi rule for "yī" defines different expected fundamental frequency transition patterns according to the tone type of the following syllable. When the following syllable is the fourth tone, the expected fundamental frequency transition pattern shows an upward trend, and the slope of the fundamental frequency change should be a positive value. When the following syllable is other tones, the expected fundamental frequency transition pattern shows a downward trend, and the slope of the fundamental frequency change should be a negative value. In addition, there are special cases in the tone sandhi of "yī". For example, when expressing ordinal numbers (such as "dì yī"), "yī" usually does not change its tone. In some fixed phrases (such as "tǒng yī"), the tone sandhi performance of "yī" may be different from the general rule. The tone sandhi rule knowledge base 3 records the handling rules for these special cases.

[0050] The third category is the tone sandhi rule for the character "bù". When "bù" is followed by a fourth tone syllable, the tone of "bù" changes from the original falling tone to a rising tone, that is, the tone value changes from 51 to 35. This tone sandhi rule is similar to the tone sandhi rule when "yī" is followed by the fourth tone, and both belong to the tone sandhi phenomenon driven by the phonetic motivation of avoiding consecutive falling tones. In continuous speech flow, two consecutive falling tones (i.e., high falling tones) will cause an unnatural feeling in pronunciation. Therefore, the previous falling tone will change to a rising tone for a smooth transition. The expected fundamental frequency transition pattern of the tone sandhi rule for "bù" is defined as: when the following syllable is the fourth tone, the fundamental frequency trajectory should change from the original downward trend to an upward trend, and the slope of the fundamental frequency change should change from a negative value to a positive value.

[0051] Variable rules refer to the tone sandhi rules that allow certain changes in standard Mandarin, or the tone sandhi rules that mainly appear in specific contexts. Violating these rules does not necessarily constitute an obvious error. Therefore, in the detection system, they are set as reference detection items, and the detection results are presented in the form of suggestions rather than being judged as errors. The variable rules in the tone sandhi rule knowledge base 3 include the following two categories.

[0052] The first category is the neutral tone weakening rule. The neutral tone is a special tonal expression in Mandarin Chinese, typically appearing after modal particles, auxiliary words, some reduplicated words, and certain fixed phrases. The fundamental frequency of a neutral tone syllable is usually low and short, and its specific tone value is greatly influenced by the tone of the preceding syllable. In the tone sandhi rule knowledge base 3, the expected fundamental frequency transition pattern of the neutral tone weakening rule is parameterized according to the tone type of the preceding syllable. For example, after a level tone, the neutral tone usually manifests as a mid-falling tone; after a falling tone, the neutral tone usually manifests as a low-level tone.

[0053] The second category is dialect tone sandhi rules. Learners from different dialect regions may bring their dialect tone sandhi habits into their Mandarin pronunciation when learning Mandarin. The tone sandhi rule knowledge base 3 stores tone sandhi features of common dialects to identify whether learners exhibit dialect tone sandhi transfer. The detection results corresponding to the dialect tone sandhi rules are labeled as suggestions to help learners recognize the influence of dialects, rather than judging them as errors.

[0054] The pitch shift mode detection engine 4 is responsible for executing the pitch shift mode detection step. This step matches the base frequency transition trajectory feature vector output by the base frequency transition trajectory extraction module 2 with the expected base frequency transition mode stored in the pitch shift rule knowledge base 3, calculates the pitch shift matching score, and determines the pitch shift detection result based on the matching score.

[0055] The tone sandhi detection engine 4 first retrieves matching tone sandhi rules from the tone sandhi rule knowledge base 3 based on the adjacent syllable combination to be detected. The retrieval process is based on the following information: the original tone type of the preceding syllable, the original tone type of the following syllable, and the Chinese character corresponding to the preceding syllable. These three pieces of information can uniquely determine the applicable tone sandhi rule, or determine whether the current syllable combination involves tone sandhi.

[0056] Once the applicable pitch shifting rule is determined, the pitch shifting mode detection engine 4 obtains the expected baseband transition mode corresponding to that rule.

[0057] The desired baseband transition mode is stored in parameterized form, including the desired baseband initial value range. Expected fundamental frequency termination value range Expected range of fundamental frequency change slope and the expected range of fundamental frequency variation .

[0058] The pitch shift mode matching degree calculation algorithm proposed in this invention is as follows. The pitch shift mode detection engine 4 will use the fundamental frequency transition trajectory feature vector... The difference between the desired fundamental frequency transition mode and the target fundamental frequency transition mode is calculated to obtain the characteristic deviation vector. The formulas for calculating each component of the characteristic deviation vector are as follows:

[0059] ,

[0060] in, The eigenvector of the eigenvector One portion, and Let the lower and upper bounds of the expected range be defined. The first characteristic deviation vector Each component. This calculation method can quantify the degree to which the actual feature value deviates from the expected range. When the actual feature value falls within the expected range, the deviation is zero; the greater the deviation, the greater the deviation value.

[0061] After obtaining the feature deviation vector, the pitch shift pattern detection engine 4 performs a weighted summation of each component to calculate the pitch shift matching score. :

[0062] ,

[0063] in, For the first The weight coefficients of each feature component satisfy the following conditions: In this embodiment, the weights of each component are set as follows: weight of the fundamental frequency change slope. The weight of the fundamental frequency variation is set to 0.35. Set to 0.30, the weight of the baseband starting value The weight of the fundamental frequency termination value is set to 0.20. The weight is set to 0.15. This weight configuration reflects the different importance of various features in pitch shift detection. The slope and amplitude of the fundamental frequency change contribute the most to pitch shift determination, while the starting and ending values ​​serve as auxiliary references.

[0064] Pitch Matching Score The value ranges from 0 to 1, with a higher score indicating a closer match between the actual pitch shift pattern and the desired pattern. The pitch shift pattern detection engine 4 uses the pitch shift matching score and a preset matching threshold. Determine the pitch shift detection result. Preset matching threshold. The value range is from 0.6 to 0.9, and in this embodiment, it is set to 0.75 by default.

[0065] The logic for determining the tone sandhi detection result is as follows. When the combination of syllables to be detected corresponds to the mandatory tone sandhi rule, if... If it is correct, then the tone change is considered correct; if Furthermore, if the fundamental frequency change trend is opposite to the expected direction or the amplitude is too small, it is judged as a lack of pitch modulation, indicating that the learner has failed to achieve the required pitch modulation; if Furthermore, if the fundamental frequency change is too large or occurs in a position where pitch should not be changed, it is judged as excessive pitch change, indicating that the learner produced an overly obvious pitch change in a position where pitch change is not required or should be slight.

[0066] When the combination of syllables to be detected corresponds to a variable rule, the judgment logic is relatively lenient. If If it is correct, then the tone change is considered correct; if If the error is not identified, it will be marked as a suggested improvement rather than an error, and the feedback will explain that the implementation of pitch shifting at that position differs from the standard mode but does not constitute a serious problem.

[0067] After detecting a single pair of adjacent syllables, the pitch shift detection engine 4 continues processing the next pair of adjacent syllables in the speech stream until the entire speech segment is detected. During the detection process, the pitch shift matching score and pitch shift detection result at each detection location are recorded to form a pitch shift detection result sequence.

[0068] In a preferred embodiment of the present invention, for the detection of tone shifting in three-tone syllables, the tone shifting pattern detection engine 4 employs a classifier based on a deep neural network for auxiliary determination. This classifier takes the fundamental frequency transition trajectory feature vector as input and outputs the probability value of whether the current syllable has shifted from a third tone to a second tone. The classifier uses a three-layer fully connected neural network structure, with the input layer dimension being the same as the feature vector dimension, the hidden layer containing 64 neurons using the ReLU activation function, and the output layer consisting of a single neuron using the Sigmoid activation function to output the probability value. The classifier is trained on a labeled dataset containing 10,000 samples of three-tone syllable syllables. Training uses the cross-entropy loss function and the Adam optimizer, with a learning rate of 0.001, a batch size of 32, and 100 training epochs. When the probability value output by the classifier is greater than a preset probability threshold, it is determined that tone shifting has occurred; in this embodiment, the preset probability threshold is set to 0.5.

[0069] The proposed modulation detection comprehensive judgment algorithm integrates rule matching scores and neural network classifier outputs to improve detection accuracy and robustness. The fusion formula is as follows:

[0070] ,

[0071] in, To determine the overall score, Score for rule matching. The probability values ​​output by the neural network classifier. These are the fusion weighting coefficients. In this embodiment, Setting it to 0.6 indicates that the rule matching score dominates the overall judgment, while the neural network classifier plays a supporting role. The final pitch shift detection result is based on... Matching threshold The comparison is used to make a judgment.

[0072] The feedback generation module 5 is responsible for performing the feedback generation step, generating intuitive and effective feedback content based on the tone change detection results to help learners understand the problems and make targeted improvements. The feedback content includes three parts: the annotated graph of the speech pitch curve, the rule explanation text, and the comparison demonstration audio.

[0073] The process of the feedback generation module 5 generating the annotated graph of the speech pitch curve is as follows. First, the complete fundamental frequency trajectory of the speech signal to be detected is plotted as a two-dimensional curve, with the horizontal axis being time and the vertical axis being the normalized fundamental frequency value or semitone value. Second, the syllable boundaries are marked on the curve according to the syllable time boundary sequence. The boundaries are drawn as dashed lines, and the corresponding Chinese characters and pinyin are marked above the boundaries. Third, the curve is divided into multiple sections and colored separately according to the tone change detection results. In this embodiment, the correctly tone-changed section is marked with the first preset color, and the first preset color is set to green, indicating that the tone change in this section meets the expectation; the tone change missing section is marked with the second preset color, and the second preset color is set to yellow, indicating that a tone change should occur in this section but fails to be realized; the over-tone-changed section is marked with the third preset color, and the third preset color is set to red, indicating that an improper tone change has occurred in this section. In addition, for the recommended items corresponding to the variable rules, the fourth preset color is used for marking, and the fourth preset color is set to blue to distinguish it from the results of the mandatory detection items.

[0074] The process of the feedback generation module 5 generating the rule explanation text is as follows. For each detected tone change abnormal position, the feedback generation module 5 retrieves the corresponding explanation template from the rule explanation library according to the abnormal type and applicable tone change rules, and fills in the specific syllable information to generate personalized explanation text. The rule explanation library stores the explanation content of various tone change rules, including the definition of the tone change rule, the reason for the tone change, the correct tone change method, and the common error types, etc. For example, for the situation of missing tone change in the three-tone continuous reading, the rule explanation text may be: In the word "xiang mai", both "xiang" and "mai" are in the third tone. According to the three-tone continuous reading tone change rule, the previous "xiang" should be changed to the second tone (rising tone), and when pronouncing, it should slide up from the low tone. In your current pronunciation, the character "xiang" still maintains the characteristics of the third tone. Please note to adjust the tone to the rising tone.

[0075] The feedback generation module 5 generates the comparison demonstration audio as follows: Based on the pitch shift detection results, it retrieves a standard pronunciation demonstration audio that matches the current pitch shift anomaly type from a preset audio library. The preset audio library stores various pitch shift demonstration audios recorded by standard Mandarin speakers, indexed by pitch shift type and syllable combination. After retrieving the matching demonstration audio, the feedback generation module 5 performs time alignment processing between the demonstration audio and the learner's corresponding speech segment, aligning the two audio segments on the time axis before outputting them for easy syllable-by-syllable comparison listening by the learner. Time alignment is achieved using a dynamic time warping algorithm, which can handle the inconsistency in duration caused by differences in speaking speed.

[0076] In a preferred embodiment of the present invention, the feedback generation module 5 also supports generating dynamic demonstration animations to visually demonstrate the correct pronunciation of tones. The dynamic demonstration animations are based on tone value curves, using animation effects to show the change in fundamental frequency during pronunciation, and accompanied by arrows indicating the direction of pitch movement. This visual feedback method helps learners understand the implementation of tone sandhi more intuitively.

[0077] The detection performance of the method of this invention was verified through the following experiment. The experiment used a speech dataset containing 200 Chinese learners, whose native languages ​​included English, Japanese, Korean, and other non-tonal languages. Each learner read 30 sentences containing tone sandhi, totaling 6000 speech samples. Three phonetics experts manually annotated the tone sandhi implementation in each sample as an evaluation benchmark. Experimental results show that the method of this invention achieves an accuracy rate of 92.3% for detecting tone sandhi in three-tone connected speech, 89.7% for single-character tone sandhi, and 91.2% for non-character tone sandhi, with an overall accuracy rate of 90.8%, representing an improvement of approximately 17.5% compared to traditional methods based on syllable internal features. Furthermore, the accuracy rate for tone sandhi type determination reached 85.6%, effectively distinguishing between tone sandhi omission and excessive tone sandhi.

[0078] The architecture of the automatic tone sandhi detection system in continuous Chinese speech provided by this invention is as follows: Figure 2 As shown, the system includes a speech forced alignment module 1, a fundamental frequency transition trajectory extraction module 2, a pitch shift rule knowledge base 3, a pitch shift pattern detection engine 4, and a feedback generation module 5. The five modules work together to achieve automatic detection and evaluation of pitch shift patterns in continuous speech.

[0079] The speech forced alignment module 1 is connected to the fundamental frequency transition trajectory extraction module 2 via a data interface. The syllable time boundary sequence output by the speech forced alignment module 1 is directly transmitted to the fundamental frequency transition trajectory extraction module 2 as the basis for defining the transition region. The fundamental frequency transition trajectory extraction module 2 is connected to the pitch shifting pattern detection engine 4 via a data interface. The fundamental frequency transition trajectory feature vector output by the fundamental frequency transition trajectory extraction module 2 is transmitted to the pitch shifting pattern detection engine 4 as input for matching calculation. The pitch shifting rule knowledge base 3 is connected to the pitch shifting pattern detection engine 4 via a query interface. The pitch shifting pattern detection engine 4 queries the pitch shifting rule knowledge base 3 for applicable pitch shifting rules and the desired fundamental frequency transition pattern based on the syllable combination information. The pitch shifting pattern detection engine 4 is connected to the feedback generation module 5 via a data interface. The pitch shifting detection result sequence output by the pitch shifting pattern detection engine 4 is transmitted to the feedback generation module 5 as the basis for generating feedback content.

[0080] The functionality of the forced speech alignment module 1 is consistent with the description of the forced speech alignment steps in the aforementioned method embodiments. This module can be implemented using existing open-source speech alignment tools, such as Montreal Forced Aligner or Kalditoolkit, and adapted and optimized according to the characteristics of Chinese syllables. The input interface of the forced speech alignment module 1 receives the speech signal to be detected and the reference text, and the output interface outputs the syllable time boundary sequence. In a preferred implementation of this system, the forced speech alignment module 1 adopts an end-to-end alignment model based on the Transformer architecture. This model is pre-trained on a large-scale Mandarin speech dataset and can directly complete the alignment of speech and text without relying on traditional acoustic models and pronunciation dictionaries, improving the alignment accuracy by approximately 10% compared to traditional methods.

[0081] The functionality of the fundamental frequency transition trajectory extraction module 2 is consistent with the description of the fundamental frequency transition trajectory extraction steps in the aforementioned method embodiments. This module includes a transition region delineation unit, a fundamental frequency extraction unit, a fundamental frequency preprocessing unit, and a feature calculation unit. The transition region delineation unit determines the transition region time window for each adjacent syllable pair based on the syllable time boundary sequence. This unit supports dynamically adjusting the boundary of the transition region according to the syllable duration to adapt to the pitch shift analysis requirements under different speech rates. The fundamental frequency extraction unit extracts the fundamental frequency trajectory from the original speech signal using the autocorrelation method or the PYIN algorithm. This unit has multiple built-in fundamental frequency extraction algorithms for selection and can automatically select the optimal algorithm based on the signal-to-noise ratio of the speech signal and speaker characteristics. The fundamental frequency preprocessing unit performs median filtering and speaker normalization on the original fundamental frequency sequence. This unit supports two normalization methods: Z-score normalization and semitone conversion. The feature calculation unit calculates the fundamental frequency start value, fundamental frequency end value, fundamental frequency change slope, and fundamental frequency change amplitude, combining them into a fundamental frequency transition trajectory feature vector. This unit can also selectively calculate extended features to improve detection accuracy.

[0082] The pitch shifting rule knowledge base 3 stores pitch shifting rules in the form of a relational database or knowledge graph, supporting fast retrieval by syllable combination. Each pitch shifting rule record in the database includes fields such as rule identifier, applicable conditions, rule type, expected fundamental frequency transition mode parameter, and rule explanation content. The rule type field distinguishes between mandatory and variable pitch shifting rules, and the expected fundamental frequency transition mode parameter field stores the expected value range of each feature dimension. The pitch shifting rule knowledge base 3 supports dynamic updates and expansions of rules, allowing the addition of new pitch shifting rules or adjustment of parameters of existing rules according to teaching needs. In a preferred implementation of this system, the pitch shifting rule knowledge base 3 is organized in the form of a knowledge graph, representing pitch shifting rules as nodes and edges in a semantic network, supporting reasoning-based rule matching and handling of complex pitch shifting scenarios. The pitch shifting rule knowledge base 3 pre-sets common mandatory and variable pitch shifting rules, and determines the expected fundamental frequency transition mode parameters of each rule based on phonetics research literature. These parameters have been statistically verified using a large-scale standard Mandarin corpus and have high reliability.

[0083] The implementation of the pitch shifting pattern detection engine 4 is consistent with the description of the pitch shifting pattern detection steps in the aforementioned method embodiments. This engine includes a rule retrieval unit, a feature matching unit, a score calculation unit, and a result determination unit. The rule retrieval unit initiates a query request to the pitch shifting rule knowledge base 3 based on the current syllable combination information to obtain applicable pitch shifting rules and the desired fundamental frequency transition pattern. This unit uses index acceleration technology to ensure that the query response time does not exceed 10 milliseconds. The feature matching unit compares the fundamental frequency transition trajectory feature vector with the desired fundamental frequency transition pattern and calculates the feature deviation in each dimension. This unit supports multiple matching metrics such as Euclidean distance and cosine similarity. The score calculation unit calculates the pitch shifting matching score based on the feature deviation vector. The weight parameters of this unit can be adjusted according to the application scenario. The result determination unit determines the pitch shifting detection result based on the pitch shifting matching score and a preset matching threshold. This unit supports multiple threshold determinations to achieve a more granular evaluation level classification. The pitch shifting pattern detection engine 4 can also integrate a deep neural network classifier as an auxiliary determination module to improve the detection accuracy of complex pitch shifting phenomena such as three-tone connected pitch shifting. In a preferred implementation of this system, the deep neural network classifier adopts a long short-term memory network structure, which can capture the temporal dependency of the fundamental frequency sequence, further improving the accuracy and robustness of modulation detection.

[0084] The functionality of feedback generation module 5 is consistent with the description of the feedback generation steps in the aforementioned method embodiments. This module includes a curve plotting unit, a text generation unit, and an audio processing unit. The curve plotting unit generates a speech pitch curve annotation diagram based on the fundamental frequency trajectory and pitch shift detection results. This unit supports multiple visualization styles and color schemes and can be customized according to user preferences. The text generation unit retrieves explanation templates from the rule explanation library and generates personalized rule explanation text. This unit supports a multilingual interface and can provide explanation content in the corresponding language for learners with different native language backgrounds. The audio processing unit retrieves demonstration audio from a preset audio library and performs time alignment processing with the learner's speech. This unit uses a dynamic time warping algorithm to achieve alignment of speech at different speeds, ensuring the effectiveness of comparative listening. The output of feedback generation module 5 is presented to the learner through a user interface, which can be implemented as a web application, mobile application, or desktop application, supporting cross-platform access.

[0085] This invention's system can be deployed on cloud servers or local computing devices. In cloud server deployment mode, learners upload their audio to be detected via a client application. After processing, the server returns the results to the client for display. This mode supports large-scale concurrent access and is suitable for applications such as online education platforms. In local deployment mode, the detection system runs on the learner's personal computer or mobile device, supporting offline use and protecting the privacy of user audio data. The system's computing resource requirements are moderate. On a computing device equipped with a standard CPU, processing a 10-second audio clip takes approximately 1 to 2 seconds, meeting the requirements for near real-time feedback. On a server equipped with GPU acceleration, the processing speed can be further improved to real-time levels.

[0086] The technical effectiveness of this invention's system has been verified through the following application scenarios. In the application of teaching Chinese as a second language, the system was deployed on an online Chinese learning platform, providing tone sandhi practice and assessment services to 2000 learners from 30 countries. User satisfaction surveys showed that 89% of learners believed the feedback provided by the system helped them understand tone sandhi rules, and 82% of learners experienced a significant improvement in tone sandhi accuracy after practicing with the system. In the application of assisting with Mandarin proficiency testing, the system was integrated into a Mandarin testing training software, providing test takers with a pre-assessment function for tone sandhi in reading aloud. Test data showed that test takers who used this system for targeted practice scored an average of 3.2 points higher in the formal test than the control group. In the application of speech rehabilitation therapy, the system was applied to speech rehabilitation training for cochlear implant patients. The system provides specialized detection and feedback for common tone sandhi difficulties in patients. Clinical data showed that after 12 weeks of training, patients' tone sandhi accuracy increased from 45% before training to 73% after training.

[0087] The above description is merely a preferred embodiment of the present invention and does not limit the scope of patent protection of the present invention. Any equivalent structural transformations made under the inventive concept of the present invention using the contents of the specification and drawings of the present invention, or direct / indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

Claims

1. An automatic detection method for tone sandhi in continuous Chinese speech, characterized in that, Including: A voice forced alignment step, obtaining a voice signal to be detected and a corresponding reference text, performing forced alignment processing on the voice signal to be detected, determining the start time point and end time point of each syllable in the reference text in the voice signal to be detected, and generating a syllable time boundary sequence; A fundamental frequency transition trajectory extraction step, based on the syllable time boundary sequence, delimiting the connection area between adjacent syllables. The delimitation of the connection area includes: centering on the end time point of the previous syllable, expanding forward by a first preset duration and backward by a second preset duration to form a connection area time window, where the value range of the first preset duration is from 30 milliseconds to 80 milliseconds, and the value range of the second preset duration is from 30 milliseconds to 80 milliseconds; performing median filtering on the original fundamental frequency sequence in the connection area to remove fundamental frequency jump points, performing speaker normalization on the filtered fundamental frequency sequence to eliminate individual pitch differences, and the speaker normalization uses the Z-score normalization method or the semitone conversion method; extracting the fundamental frequency transition trajectory in each connection area, calculating the fundamental frequency start value, fundamental frequency end value, fundamental frequency change slope, and fundamental frequency change amplitude according to the fundamental frequency transition trajectory, and combining the fundamental frequency start value, fundamental frequency end value, fundamental frequency change slope, and fundamental frequency change amplitude into a fundamental frequency transition trajectory feature vector; A tone sandhi rule matching step, obtaining the tone sandhi rule corresponding to the current adjacent syllable combination from the tone sandhi rule knowledge base, where the tone sandhi rule knowledge base includes two categories: mandatory rules and variable rules, the mandatory rules correspond to mandatory detection items, the variable rules correspond to reference detection items, and determining the expected fundamental frequency transition pattern according to the tone sandhi rule; A tone sandhi pattern detection step, matching the fundamental frequency transition trajectory feature vector with the expected fundamental frequency transition pattern, calculating the difference between the fundamental frequency transition trajectory feature vector and the expected fundamental frequency transition pattern to obtain a feature deviation vector, performing weighted summation on each component in the feature deviation vector to obtain a tone sandhi matching score, and determining the tone sandhi detection result according to the tone sandhi matching score and a preset matching threshold, where the value range of the preset matching threshold is from 0.6 to 0.9, the tone sandhi detection result includes three types: correct tone sandhi, missing tone sandhi, and excessive tone sandhi, and locating the specific syllable position where tone sandhi abnormality occurs; A feedback generation step, generating a pitch curve annotation diagram of the speech flow according to the tone sandhi detection result, differentiating the correct tone sandhi section, missing tone sandhi section, and excessive tone sandhi section with different colors in the pitch curve annotation diagram of the speech flow, and simultaneously generating a rule explanation text and a comparison demonstration audio corresponding to the tone sandhi abnormal position.

2. The method for automatic detection of tone sandhi in continuous Chinese speech as described in claim 1, characterized in that, The mandatory rules in the tone sandhi rule knowledge base include: the tone sandhi rule for three consecutive third tones, when two consecutive third tone syllables are connected, the previous syllable changes to the second tone; the tone sandhi rule for the character '一', when the character '一' is followed by a fourth tone syllable, the character '一' changes to the second tone, and when the character '一' is followed by a first tone, second tone, or third tone syllable, the character '一' changes to the fourth tone; the tone sandhi rule for the character '不', when the character '不' is followed by a fourth tone syllable, the character '不' changes to the second tone.

3. The method for automatic detection of tone sandhi in continuous Chinese speech as described in claim 1, characterized in that, In the tone shift detection step, for the detection of tone shift in three consecutive tones, a classifier based on a deep neural network is used to classify the feature vector of the fundamental frequency transition trajectory. The classifier outputs the probability value of whether the current syllable changes from a third tone to a second tone. When the probability value is greater than a preset probability threshold, it is determined that a tone shift has occurred.

4. The method for automatic detection of tone sandhi in continuous Chinese speech as described in claim 1, characterized in that, In the feedback generation step, the generation of the speech pitch curve annotation map includes: drawing the complete fundamental frequency trajectory of the speech signal to be detected as a curve, marking syllable boundaries on the curve according to the syllable time boundary sequence, dividing the curve into multiple segments and coloring them according to the pitch shift detection results, wherein the segments with correct pitch shift are marked with a first preset color, the segments with missing pitch shift are marked with a second preset color, and the segments with excessive pitch shift are marked with a third preset color.

5. The method for automatic detection of tone sandhi in continuous Chinese speech as described in claim 1, characterized in that, The variable rules in the tone sandhi rule knowledge base include: neutral tone weakening rules, corresponding to the detection of neutral tone in modal particles, auxiliary words, and some reduplicated words; dialect tone sandhi rules, corresponding to the detection of specific tone sandhi habits of learners in different dialect areas; the reference detection items corresponding to the variable rules are marked as suggested items rather than incorrect items in the detection results.

6. The method for automatic detection of tone sandhi in continuous Chinese speech as described in claim 1, characterized in that, The feedback generation step further includes: retrieving standard pronunciation demonstration audio that matches the current pitch shift anomaly type from a preset audio library based on the pitch shift detection result, and outputting the standard pronunciation demonstration audio after time alignment with the learner's corresponding segment for the learner to listen to for comparison.

7. An automatic detection system for tone sandhi in continuous Chinese speech, used to implement the automatic detection method for tone sandhi in continuous Chinese speech as described in any one of claims 1-6, characterized in that, include: The speech forced alignment module is used to acquire the speech signal to be detected and the corresponding reference text, perform forced alignment processing on the speech signal to be detected, determine the start time point and end time point of each syllable in the reference text in the speech signal to be detected, and generate a syllable time boundary sequence. The fundamental frequency transition trajectory extraction module is used to delineate the connection region between adjacent syllables based on the syllable time boundary sequence. The connection region is centered on the end time of the previous syllable, extended forward by a first preset duration, and extended backward by a second preset duration to form a connection region time window. The first preset duration ranges from 30 milliseconds to 80 milliseconds, and the second preset duration also ranges from 30 milliseconds to 80 milliseconds. After performing median filtering and speaker normalization on the original fundamental frequency sequence within the connection region, the module extracts the fundamental frequency transition trajectory and calculates the fundamental frequency start value, fundamental frequency end value, fundamental frequency change slope, and fundamental frequency change amplitude to generate a fundamental frequency transition trajectory feature vector. A pitch-change rule knowledge base is used to store mandatory and variable rules. The mandatory rules correspond to forced detection items, and the variable rules correspond to reference detection items. The pitch shifting mode detection engine is used to obtain pitch shifting rules from the pitch shifting rule knowledge base, calculate the difference between the base frequency transition trajectory feature vector and the desired base frequency transition mode to obtain a feature deviation vector, perform weighted summation of each component in the feature deviation vector to obtain a pitch shifting matching score, and determine the pitch shifting detection result based on the pitch shifting matching score and a preset matching threshold. The feedback generation module is used to generate a speech pitch curve annotation diagram, rule explanation text, and comparison demonstration audio based on the pitch shift detection results.

Citation Information

Patent Citations

  • Multi-language cross-culture communication auxiliary method and system based on large model

    CN120636412A

  • Hierarchical real-time speaker recognition for biometric VoIP verification and targeting

    US8160877B1