A speech therapy treatment instrument abnormal voice detection method and system

By using a microphone and a main control processor for speech frame segmentation and formant analysis, combined with spectral comparison, real-time abnormal speech detection and feedback for speech therapy instruments were achieved, solving the problem of feedback lag in existing technologies and improving training effectiveness.

CN122290639APending Publication Date: 2026-06-26WOMEN & CHILDRENS MEDICAL CENTER AFFILIATED WITH GUANGZHOU MEDICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WOMEN & CHILDRENS MEDICAL CENTER AFFILIATED WITH GUANGZHOU MEDICAL UNIVERSITY
Filing Date
2026-05-21
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing speech therapy devices are unable to locate abnormal pronunciation segments and types in real time or near real time during children's training, resulting in delayed treatment feedback and affecting training effectiveness.

Method used

Employing a microphone, voice acquisition card, main control processor, and treatment feedback output module, the system identifies abnormal speech and generates treatment feedback information in real time through voice frame segmentation, formant analysis, and spectrum comparison.

Benefits of technology

It enables real-time detection and feedback of abnormal speech during children's training, improving the on-site feedback capability of the treatment instrument and the accuracy of training intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122290639A_ABST
    Figure CN122290639A_ABST
Patent Text Reader

Abstract

This invention relates to the field of pediatric therapeutic instrument technology, specifically to a method and system for detecting abnormal speech in a speech therapy instrument. The method involves acquiring voltage signals and filtering discrete speech frame sequences, extracting formant trajectories to analyze frequency change trends, comparing the articulation direction distribution of children with standard pronunciation, identifying directional deviation segments, and using formant neighborhood envelope peak comparison to pinpoint envelope offset segments. It also compares spectral peak and valley distribution characteristics, performs multidimensional temporal overlap comparison, and finally outputs the abnormal speech detection results. This invention utilizes amplitude cohesion to filter stable speech frame sequences, extracts structured formant trajectories, correlates syllable segment change directions to enhance dynamic trend discrimination, analyzes neighborhood envelope peak offsets to refine spectral characterization, cross-validates peak and valley distribution and temporal overlap segments, strengthens multidimensional consistency, achieves hierarchical identification and precise localization of pronunciation abnormalities, and improves the stability and distinguishability of detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pediatric therapeutic instrument technology, and in particular to a method and system for detecting abnormal speech in a speech therapy instrument. Background Technology

[0002] The field of pediatric therapeutic equipment technology encompasses pediatric rehabilitation equipment, pediatric speech training equipment, pediatric pronunciation correction equipment, pediatric cognitive intervention equipment, and pediatric speech assessment equipment. This technology primarily addresses issues such as delayed language development, articulation disorders, abnormal language expression, and abnormal vocalization in children. It involves pediatric speech acquisition, pronunciation training, speech feature detection, pronunciation status assessment, and rehabilitation intervention. Its core components include acquiring pediatric speech information, analyzing pediatric pronunciation features, recognizing abnormal speech, and detecting speech status during treatment and training. This is typically achieved by combining speech acquisition devices, short-time energy analysis, fundamental frequency detection, formant analysis, Mel-frequency cepstral coefficient extraction, and spectrogram analysis to complete pediatric speech detection and assessment.

[0003] Among them, the abnormal speech detection method and system for speech therapy instruments refers to the technical content for identifying and analyzing abnormal pronunciation during children's speech therapy training. It mainly covers children's speech acquisition, speech framing processing, silence segment removal, spectrum parameter extraction, fundamental frequency calculation, formant detection, abnormal pronunciation comparison, and children's pronunciation status determination. Typically, it involves acquiring children's training speech, performing pre-emphasis processing, windowing processing, and Fourier transform on the speech signal, extracting Mel frequency cepstral coefficients, short-time energy parameters, fundamental frequency parameters, and formant parameters, identifying and classifying abnormal speech based on the range of normal pronunciation characteristics of children, and combining the children's historical training speech data to complete the abnormal pronunciation status analysis.

[0004] Existing speech therapy devices for children typically only collect and record speech, or display basic speech parameters. They struggle to pinpoint specific abnormal speech segments, spectral structures, and types of abnormalities in real-time or near real-time during treatment. Doctors or therapists often need to review the recordings after training and manually assess the child's pronunciation based on their experience, leading to delayed detection of abnormal speech and hindering timely application of treatment feedback to the current training phase. Furthermore, existing speech therapy devices lack the ability to integrate speech collection, real-time detection of abnormal speech, location of abnormal segments, output of abnormal types, and treatment feedback. They cannot translate abnormal speech results into readily available segment markers, abnormal type prompts, and training adjustment guidelines for real-time use. This results in a lack of precise data for subsequent adjustments to corrective training programs, training rhythms, and pronunciation prompts, impacting the effectiveness of immediate intervention during children's speech training. Summary of the Invention

[0005] To address the technical problems existing in the prior art, embodiments of the present invention provide a method and system for detecting abnormal speech in a speech therapy instrument. The technical solution is as follows: A method for detecting abnormal speech in a speech therapy instrument, the speech therapy instrument comprising a microphone, a speech acquisition card, a main control processor, a standard pronunciation data storage unit, and a treatment feedback output module, wherein the microphone is connected to the speech acquisition card, the speech acquisition card is connected to the main control processor, and the main control processor is connected to both the standard pronunciation data storage unit and the treatment feedback output module; the method is executed by the main control processor in a child's speech therapy training setting, and includes the following steps: S1: The microphone collects the air vibration signal generated by the child's vocal training, and the voice acquisition card converts the air vibration signal into a discrete sampling point sequence corresponding to the voltage change signal. The main control processor divides the voice frames according to time, filters and sorts the voice frames with continuously changing amplitude to obtain the voice frame sequence. S2: The main control processor converts the discrete sampling point sequence corresponding to the speech frame sequence into a frequency distribution frame by frame, compares the amplitude of adjacent frequency points, filters the local peak frequency positions and checks the inter-frame continuity to obtain the formant trajectory sequence. S3: The main control processor compares the frequency position of adjacent speech frames frame by frame based on the local peak frequency position in the formant trajectory sequence, the frequency position shifts up, down and stays, divides the speech frame segments according to the training syllables, counts the distribution of the direction of change of children's pronunciation, and calls the standard pronunciation change direction distribution in the standard pronunciation data storage unit for comparison, filters the difference segments, and obtains the speech direction deviation segments. S4: The main control processor selects a child's pronunciation speech frame according to the pronunciation direction deviation segment, defines a neighborhood frequency range around the local peak frequency position corresponding to the formant, scans the envelope amplitude and compares it with the standard pronunciation envelope peak frequency position in the standard pronunciation data storage unit to obtain the envelope offset segment; S5: The main control processor reads the spectrum data of the neighboring frequency range for the envelope offset segment, identifies the peak position and valley position and compares them with the standard pronunciation peak and valley frequency positions in the standard pronunciation data storage unit, and then compares them with the articulation direction deviation segment in time to obtain the abnormal speech detection result. Based on the abnormal speech detection result, it generates treatment feedback information and sends the treatment feedback information to the treatment feedback output module.

[0006] As a further aspect of the present invention, the speech frame sequence includes a frame timestamp, frame duration, frame energy representation, and frame validity identifier; the formant trajectory sequence includes the formant center frequency, frequency variation amplitude, frequency continuity index, and formant level identifier; the articulation direction deviation segment includes the direction deviation type, deviation intensity level, segment span range, and syllable association identifier; the envelope offset segment includes the envelope peak offset, offset direction identifier, frequency band coverage range, and offset significance level; the abnormal speech detection result includes the abnormal segment number, abnormal feature category, spectral structure identifier, and syllable abnormality type; and the treatment feedback information includes abnormal pronunciation segment prompts, abnormal feature category prompts, syllable abnormality type prompts, and training adjustment prompts.

[0007] As a further aspect of the present invention, the step of obtaining S1 is as follows: S11: The microphone senses the air vibration generated by the child's vocal training and outputs a voltage change signal. The voice acquisition card reads the voltage change signal, records the voltage amplitude according to the sampling time sequence, organizes the correspondence between the sampling time and the sampling amplitude, and generates a discrete sampling point sequence. S12: The main control processor, based on the discrete sampling point sequence, continuously divides the speech frames according to the sampling time, reads the arrangement order of the sampling points frame by frame, checks the amplitude connection status of adjacent sampling points, and retains the speech frames with complete sampling point order and smooth amplitude connection to obtain a candidate speech frame set. S13: The main control processor reads the starting sampling time of each audio frame for the candidate audio frame set, arranges and retains the audio frames in chronological order, records the correspondence between the sampling point number and amplitude of each audio frame, and filters audio frames with continuously changing amplitude to obtain an audio frame sequence.

[0008] As a further aspect of the present invention, the step of obtaining S2 is as follows: S21: The main control processor reads the sampling point arrangement order frame by frame based on the discrete sampling point sequence corresponding to each speech frame in the speech frame sequence, expands the sampling amplitude according to the frequency position, forms the frequency point and amplitude correspondence relationship, compares the amplitude connection state of adjacent frequency points, and generates a frequency amplitude correspondence sequence. S22: The main control processor calls the frequency amplitude corresponding sequence, reads the amplitude of each frequency point, the left frequency point, and the right frequency point point one by one, compares the magnitude of the three amplitudes, filters the frequency positions with amplitudes higher than the amplitudes of the left and right adjacent frequency points, and arranges them in order of frequency from low to high to obtain the local peak frequency position set. S23: The main control processor selects the frequency positions corresponding to the formant distribution of the child's pronunciation according to the local peak frequency position set, arranges the selected frequency positions in the order of speech frame time, checks the continuous state of frequency position changes between adjacent speech frames, and obtains the formant trajectory sequence.

[0009] As a further aspect of the present invention, the step of obtaining S3 is as follows: S31: The main control processor reads the local peak frequency position of the formant in the formant trajectory sequence frame by frame based on the local peak frequency position of the formant in the adjacent speech frames, compares the changes in the arrangement of the frequency positions before and after, records the upward, downward and stationary frequency positions as direction marks, arranges the direction marks according to the time sequence of the speech frames, and generates the frequency direction change amount. S32: The main control processor divides the speech frame into segments according to the training syllable order based on the frequency direction change, reads the arrangement of direction marks in the segment segment by segment, counts the distribution of up-moving marks, down-moving marks, and stationary marks, sorts out the correspondence between each speech frame segment and the direction marks, and obtains the segment direction distribution. S33: The main control processor calls the segment direction distribution quantity, reads the local peak frequency position corresponding to the standard pronunciation formant of the corresponding training syllable in the standard pronunciation data storage unit, sorts out the standard pronunciation change direction distribution, performs a corresponding speech frame segment comparison between the child's pronunciation change direction distribution and the standard pronunciation change direction distribution, filters out speech frame segments with different direction distributions, and obtains the articulation direction deviation segment.

[0010] As a further aspect of the present invention, the step of obtaining S4 is as follows: S41: The main control processor selects the speech frame corresponding to the child's pronunciation based on the speech frame segment with differences in the speech direction deviation segment, extracts the local peak frequency position corresponding to the formant in the speech frame, delineates the adjacent frequency boundary around the frequency position, filters the frequency points between the boundary, records the correspondence between the speech frame number and the frequency boundary, and generates the formant neighborhood frequency range. S42: The main control processor scans the changes in the envelope amplitude within the neighborhood frequency range point by point based on the formant neighborhood frequency range, reads the envelope amplitude of each frequency point of the child's pronunciation, compares the arrangement of the envelope amplitude of adjacent frequency points, records the peak frequency position of the child's pronunciation envelope, and establishes the peak position of the child's envelope. S43: The main control processor calls the child's envelope peak position, reads the corresponding training syllable and the standard pronunciation envelope peak frequency position in the corresponding neighborhood frequency range in the standard pronunciation data storage unit, performs a position difference comparison between the child's pronunciation envelope peak frequency position and the standard pronunciation envelope peak frequency position, filters the envelope position offset speech frame segment and the corresponding neighborhood frequency range, and obtains the envelope offset segment.

[0011] As a further aspect of the present invention, the step of obtaining S5 is as follows: S51: The main control processor reads the spectrum data of the neighboring frequency range within the speech frame segment and the corresponding neighboring frequency range in the envelope offset segment, checks the amplitude arrangement of adjacent frequency points in frequency order, identifies the peak position and valley position, records the correspondence between the speech frame number and the frequency position, and generates the peak-valley frequency distribution. S52: The main control processor calls the peak-valley frequency distribution, reads the standard pronunciation spectrum data of the corresponding training syllable and the same neighborhood frequency range in the standard pronunciation data storage unit, identifies the peak position and valley position of the standard pronunciation, performs a peak-valley frequency position comparison between the child's pronunciation and the standard pronunciation, filters out speech frame segments with differences in peak-valley distribution, and obtains peak-valley difference segments. S53: The main control processor calls the speech frame segments with different directional distributions in the speech direction deviation segment according to the peak-valley difference segment, performs time comparison of the two types of speech frame segments, filters out time-overlapping segments, arranges them according to the training syllable order and records the corresponding neighborhood frequency range, and obtains the abnormal speech detection result. S54: The main control processor generates treatment feedback information based on the abnormal segment number, abnormal feature category, spectral structure identifier and syllable abnormal type in the abnormal speech detection result, and controls the treatment feedback output module to output the treatment feedback information in at least one of the following forms: time axis mark, abnormal segment highlight, abnormal type label, spectral structure prompt, syllable abnormal prompt or training task prompt. The method is applied in the context of speech therapy training for children. After a child completes a single training syllable, phrase, or sentence, the therapeutic feedback output module outputs the corresponding abnormal pronunciation segment, abnormal feature category, and syllable abnormality type in real time or near real time. This prompts the therapist to adjust subsequent training syllables, training frequency, pronunciation prompting methods, training rhythm, or repetition of training segments. As a further aspect of the present invention, the method is applied in the context of speech therapy training for children. After a child completes a single training syllable, training phrase, or training sentence, the therapeutic feedback output module outputs the corresponding abnormal pronunciation segment, abnormal feature category, and syllable abnormality type in real time or near real time to prompt the therapist to adjust subsequent training syllables, training times, pronunciation prompting methods, training rhythm, or repetition of training segments.

[0012] An abnormal speech detection system for a speech therapy instrument, the system being used to execute the above-mentioned abnormal speech detection method for a speech therapy instrument, the system comprising a microphone, a speech acquisition card, a main control processor, a standard pronunciation data storage unit, and a treatment feedback output module; The microphone is used to collect air vibration signals generated during children's vocal training. The voice acquisition card is connected to the microphone and is used to convert the air vibration signal into a discrete sampling point sequence corresponding to the voltage change signal; The standard pronunciation data storage unit is used to store the distribution of standard pronunciation change direction, the position of the peak frequency of the standard pronunciation envelope, and the position of the peak and valley frequencies of the standard pronunciation. The main control processor is connected to the voice acquisition card, the standard pronunciation data storage unit, and the treatment feedback output module, respectively. The main control processor includes: The voice acquisition module is used to acquire the discrete sampling point sequence output by the voice acquisition card, divide the voice frames according to time, filter and sort the voice frames with continuously changing amplitude to obtain the voice frame sequence. The frequency peak extraction module is used to convert the discrete sampling point sequence corresponding to the speech frame sequence into a frequency distribution frame by frame, compare the magnitude of adjacent frequency points, filter the local peak frequency positions and check the inter-frame continuity to obtain the formant trajectory sequence. The direction analysis module is used to compare the frequency position of adjacent speech frames frame by frame based on the local peak frequency position in the formant trajectory sequence, the upward and downward shift and the dwell state of the frequency position, divide the speech frame into segments according to the training syllables, count the distribution of the direction of pronunciation change of children, compare it with the distribution of the direction of pronunciation change of standard in the standard pronunciation data storage unit, filter the difference segments, and obtain the speech direction deviation segments. The envelope localization module is used to select a child's pronunciation speech frame according to the articulation direction deviation segment, define a neighborhood frequency range around the local peak frequency position corresponding to the formant, scan the envelope amplitude and compare it with the standard pronunciation envelope peak frequency position in the standard pronunciation data storage unit to obtain the envelope offset segment. The anomaly detection module is used to read the spectrum data of the neighboring frequency range for the envelope offset segment, identify the peak position and valley position and compare them with the standard pronunciation peak and valley frequency positions in the standard pronunciation data storage unit, and then compare them with the articulation direction deviation segment in time to obtain the abnormal speech detection result. The feedback output control module is used to generate treatment feedback information based on the abnormal speech detection results, and to control the treatment feedback output module to output the treatment feedback information. The therapeutic feedback output module is used to output therapeutic feedback information, including abnormal pronunciation segments, abnormal feature categories, and syllable abnormality types, at the pediatric speech therapy training site.

[0013] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: 1. In this invention, a hardware collaborative processing link is formed by a microphone, a voice acquisition card, a main control processor, a standard pronunciation data storage unit, and a treatment feedback output module, enabling the speech therapy instrument to collect children's training pronunciations at the children's speech therapy training site, analyze abnormal speech features in real time or near real time, and output treatment feedback information.

[0014] 2. This invention constructs a continuous and stable speech frame sequence by screening amplitude connection relationships, thereby enhancing the consistency of the input signal. It forms a structured formant trajectory by screening local peak frequencies and constraining cross-frame continuity, improving the coherence of frequency feature expression. It enhances the ability to discern dynamic trends in children's pronunciation by modeling the correlation between the direction of frequency position change and training syllable segments. It refines the ability to represent local spectral structure differences by analyzing the offset of envelope peak positions within the neighboring frequency band. It strengthens the multidimensional consistency of abnormal speech judgment by cross-validating peak and valley distribution features with overlapping time segments, thereby achieving hierarchical recognition and precise localization of pronunciation abnormalities and improving the stability and distinguishability of detection results.

[0015] 3. This invention can convert abnormal speech detection results into therapeutic feedback information containing abnormal pronunciation segments, abnormal feature categories, spectral structure identifiers, and syllable abnormality types. This feedback is output through a therapeutic feedback output module in the form of timeline markers, abnormal segment highlights, abnormal type labels, spectral structure prompts, syllable abnormality prompts, or training task prompts. This allows doctors or therapists to promptly identify the location and type of abnormal pronunciation during the child's current training process. Based on the detection results, they can adjust subsequent training syllables, training frequency, pronunciation prompting methods, training rhythm, or repeated training segments, forming a closed loop of speech acquisition—abnormality detection—result output—therapeutic feedback. This improves the on-site feedback capability and training intervention accuracy of speech therapy instruments. Attached Figure Description

[0016] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a system flowchart of the present invention. Detailed Implementation

[0017] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0018] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0019] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the difference between them, their intended meanings are consistent. Similarly, the terms "of," "correlation (ponding)," and "correlation (ponding)" may sometimes be used interchangeably. It should be noted that, without emphasizing the difference between them, their intended meanings are consistent.

[0020] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0021] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0022] Please see Figure 1 The present invention provides a technical solution: a method for detecting abnormal speech in a speech therapy instrument. The speech therapy instrument includes a microphone, a voice acquisition card, a main control processor, a standard pronunciation data storage unit, and a treatment feedback output module. The microphone is connected to the voice acquisition card, the voice acquisition card is connected to the main control processor, and the main control processor is connected to the standard pronunciation data storage unit and the treatment feedback output module, respectively. The method, executed by a master processor in a child speech therapy training setting, includes the following steps: S1: The air vibration signal generated by the child's pronunciation training is collected by the microphone, and the voice acquisition card converts the air vibration signal into a discrete sampling point sequence corresponding to the voltage change signal. The main control processor continuously divides the voice frame according to the time sequence of the discrete sampling points, reads the sampling amplitude change within the voice frame point by point, filters the voice frames with continuously changing amplitude according to the amplitude connection relationship of adjacent sampling points, and arranges the filtered voice frames in time sequence to obtain the voice frame sequence. S2: The main control processor converts the discrete sampling point sequence corresponding to each speech frame in the speech frame sequence into a frequency distribution frame by frame, forming a frequency-amplitude correspondence. It compares the amplitude of each frequency point with the amplitude of the left and right adjacent frequency points, filters out the local peak frequency positions with amplitudes higher than the amplitudes of the left and right adjacent frequency points, selects the local peak frequency positions corresponding to the formant distribution of the child's pronunciation in order of frequency from low to high, arranges the selected local peak frequency positions in the time sequence of the speech frames, checks the continuity of the change of local peak frequency positions between adjacent speech frames, and obtains the formant trajectory sequence. S3: The main control processor compares the changes in the local peak frequency position of the formant corresponding to the formant in adjacent speech frames frame by frame based on the local peak frequency position in the formant trajectory sequence. It records the change direction information corresponding to the upward, downward and stationary frequency position. It divides the speech frame into segments according to the order of training syllables, counts the change direction distribution within each speech frame segment, calls the local peak frequency position corresponding to the standard pronunciation formant of the training syllable in the standard pronunciation data storage unit and performs the same direction sorting to form the standard pronunciation change direction distribution. It compares the children's pronunciation change direction distribution with the standard pronunciation change direction distribution in the corresponding speech frame segments, and filters the speech frame segments with different direction distributions to obtain the articulation direction deviation segments. S4: The main control processor selects the speech frame corresponding to the child's pronunciation based on the speech frame segment with differences in the articulation direction deviation segment, extracts the local peak frequency position corresponding to the formant in the speech frame, delineates the neighborhood frequency range around the local peak frequency position corresponding to the formant, scans the change of the envelope amplitude in the neighborhood frequency range point by point, records the peak frequency position of the child's pronunciation envelope, reads the peak frequency position of the standard pronunciation envelope in the corresponding training syllable and the corresponding neighborhood frequency range in the standard pronunciation data storage unit, performs a position difference comparison between the peak frequency position of the child's pronunciation envelope and the peak frequency position of the standard pronunciation envelope, filters the speech frame segment with envelope position offset and the corresponding neighborhood frequency range, and obtains the envelope offset segment; S5: The main control processor reads the spectrum data of the neighboring frequency range within the speech frame segment and its corresponding neighboring frequency range in the envelope offset segment, identifies the peak and valley positions and records the frequency distribution information, reads the peak and valley positions of the corresponding training syllable and the standard pronunciation within the same neighboring frequency range in the standard pronunciation data storage unit, performs a peak-valley frequency position comparison between the child's pronunciation and the standard pronunciation, filters speech frame segments with different peak-valley distributions, calls speech frame segments with different directional distributions in the articulation direction deviation segment, performs a time comparison between the two types of speech frame segments, filters time overlap segments, arranges them according to the order of training syllables and records the corresponding neighboring frequency ranges, obtains the abnormal speech detection results, and generates treatment feedback information based on the abnormal speech detection results. The treatment feedback information is sent to the treatment feedback output module, which outputs the abnormal pronunciation segment, abnormal feature category and syllable abnormal type at the child speech therapy training site.

[0023] The speech frame sequence includes frame timestamp, frame duration, frame energy representation, and frame validity identifier. The formant trajectory sequence includes formant center frequency, frequency variation amplitude, frequency continuity index, and formant level identifier. The articulation direction deviation segment includes deviation type, deviation intensity level, segment span range, and syllable association identifier. The envelope offset segment includes envelope peak offset, offset direction identifier, frequency band coverage range, and offset significance level. The abnormal speech detection results include abnormal segment number, abnormal feature category, spectral structure identifier, and syllable abnormality type. The treatment feedback information includes abnormal training syllables, abnormal speech frame start and end times, abnormal segment number, abnormal feature category, frequency band coverage range, offset significance level, syllable abnormality type, and training adjustment prompts.

[0024] Specifically, the steps to obtain S1 are as follows: S11: Acquire the voltage change signal output by the microphone in the speech therapy instrument after sensing the air vibration of the child's speech, call the voice acquisition card to read the voltage change signal, record the voltage amplitude according to the sampling time sequence, organize the correspondence between sampling time and sampling amplitude, and generate a discrete sampling point sequence. The system acquires the voltage change signal output by the microphone in a speech therapy device after sensing air vibrations during a child's vocalization. This acquisition front-end uses a high-sensitivity MEMS condenser microphone as an acoustic-electric transducer, transmitting the continuous sound pressure signal to the main control chip via an I2S digital audio interface. The chip then calls upon the 24-bit Sigma-Delta A / D conversion module in the voice acquisition card's hardware circuitry to perform high-precision quantization and set the sampling frequency. for The corresponding sampling period is Set the quantization depth to This module will capture the microphone's... to The analog voltage fluctuations within the range are linearly mapped to the numerical range. Based on the sampling time Sequentially record the original quantized voltage amplitude at each moment. The data from the first 10 sampling points were extracted as follows: , , , , , , , , , Perform normalization operation ,in Take the system's preset maximum range , Take the minimum range Substituting the first point into the calculation, we get Similarly, the normalized amplitudes of the remaining points are calculated and the data from each sampling time is compiled. With the corresponding normalized amplitude The mapping relationship is used to check the logical continuity of the sampling sequence point by point, and the values ​​are stored in a bidirectional data buffer to generate a discrete sampling point sequence.

[0025] Table 1: Quantization Record Table of Discrete Sampling Points ; As shown in Table 1, the physical quantity conversion details of the first five sampling points in the initial stage of signal acquisition are shown. Through linear mapping and normalization of the original voltage, the conversion of the vocal vibration signal into a standardized digital sequence is realized.

[0026] S12: Based on the discrete sampling point sequence, the speech frames are continuously divided according to the sampling time. The order of sampling points is read frame by frame. The amplitude connection status of adjacent sampling points is checked. Speech frames with complete sampling point order and smooth amplitude connection are retained to obtain a candidate speech frame set. Based on the discrete sampling point sequence, the duration of each frame is set to... And the frame shift is The number of sampling points covered by each speech frame is calculated based on the sampling rate. Perform continuous slicing on the sequence, reading the sampling point arrangement index frame by frame. Extract the amplitude of adjacent sampling points and ; Perform subtraction to obtain the absolute difference in amplitude. The system reads the preset transition smoothness benchmark value through background noise variance analysis. The dimension is the absolute value of the difference in normalized amplitudes. For the first frame of data, the amplitudes at points 50 to 52 are extracted as follows: , , ; Calculated Determined by logical operators Record the point as a smooth state, traverse all 319 connection positions in the entire frame and count the total number of non-smooth state points. Read the maximum fault tolerance value ; If it is determined that an inconsistency in the connection is detected (i.e.) If the frame is affected by random pulse interference, then the entire frame is removed and the fault time index is recorded; if If the frame structure is complete, it is marked as usable and pushed into the second-level buffer to obtain a set of candidate speech frames.

[0027] S13: For the candidate speech frame set, read the starting sampling time of each speech frame, arrange and retain the speech frames in chronological order, record the correspondence between the sampling point number and amplitude of each speech frame, filter speech frames with continuously changing amplitude, and obtain the speech frame sequence. For the candidate speech frame set, extract the start sampling time data from the stored header file. The sorting operator is invoked to rearrange the data in ascending order along the time axis. For the rearranged data... Frame and the For each frame, a consistency check is performed to confirm the logical incrementing relationship of the sampling point numbers, and the tuple of sampling point number and amplitude within the frame is read point by point. Define the second-order difference operator To characterize the signal curvature, a continuously varying threshold is set. Calculate using real-time data; If at the 100th sampling point , , ,calculate ; determination For continuously varying amplitudes, the percentage of continuously varying points across the entire frame is counted. When this percentage exceeds a preset threshold... If the proportion is lower than the smoothing characteristic, it is determined that the smoothing characteristic is met. If a frame is found to have a risk of speech interruption or breakage, it is discarded and the linear interpolation of adjacent frames is used to fill in the data gaps. Finally, the filtered frames are physically spliced ​​together according to the time index to obtain the speech frame sequence.

[0028] Specifically, the steps to obtain S2 are as follows: S21: Based on the discrete sampling point sequence corresponding to each speech frame in the speech frame sequence, read the sampling point arrangement order frame by frame, expand the sampling amplitude according to the frequency position, form the frequency point and amplitude correspondence, compare the amplitude connection state of adjacent frequency points, and generate the frequency amplitude correspondence sequence. Based on the discrete sampling point sequence in the speech frame sequence, 320 sampling points are read frame by frame. To suppress spectral leakage, a Hamming window function is first applied to the sampling points. Then, a 512-point Fast Fourier Transform (FFT) algorithm is called to pad 192 zero points at the end of the original sequence, mapping the time-domain sequence to the frequency-domain complex plane. The amplitude of each frequency point is calculated according to the modulus formula, and the frequency resolution is set to [value missing]. The amplitude is indexed by frequency from arrive Expanding, a binary correspondence is formed, and adjacent frequency points are compared point by point. and Calculate the amplitude difference based on the amplitude change. The dimension is decibel ( dB) ), read the system's preset spectrum connection threshold Read the 20th frequency point Amplitude Point 21 Amplitude Determine the difference If the value is below the threshold, the frequency domain energy connection is confirmed to be stable. If the difference exceeds the threshold, it is judged as a calculation anomaly, and a re-windowing verification operation is performed. Finally, the full frequency band data is integrated to generate the frequency amplitude corresponding sequence.

[0029] S22: Call the frequency amplitude corresponding sequence, read the amplitude of each frequency point, the left frequency point, and the right frequency point point one by one, compare the magnitude of the three amplitudes, filter the frequency positions with amplitudes higher than the amplitudes of the left and right adjacent frequency points, arrange them in order of frequency from low to high, and obtain the set of local peak frequency positions. Call the frequency amplitude corresponding sequence, for the first For frame spectrum data, a sliding window with a length of 3 frequency points is set, and the center frequency point is extracted point by point. and the amplitude of its left and right frequency points , , Perform a three-point size comparison and judgment, when the condition is met... and At that time, mark For local maxima, read the values, such as...

[0030] , Execution judgment confirmation For the location of the local peak frequency, in Repeated filtering within the main vocal band. Reading the average energy baseline value across the entire frame. ; Set the noise filtering threshold to Remove those with amplitudes lower than The false peaks are identified. If a frame has no valid peaks, it is marked as a silent frame. The remaining valid peaks are arranged in ascending order of frequency to obtain the local peak frequency location set.

[0031] S23: Based on the local peak frequency location set, select the frequency positions corresponding to the formant distribution of the child's pronunciation, arrange the selected frequency positions according to the time sequence of the speech frames, check the continuous state of frequency position changes between adjacent speech frames, and obtain the formant trajectory sequence. Based on the local peak frequency location set and the physiological characteristics of children, the formant search interval is set, and the first formant is... Designated as Second resonance peak Designated as The third resonance peak Designated as Perform a maximum value search within the interval. If the maximum value is at the 1st position... frame Interval detected and ,because Select As this frame Feature points are extracted based on time indexes to determine the positions of formants in each frame. Calculate the frequency shift of the same order resonant peak between adjacent frames. Set a threshold for the continuity of the motion trajectory. If the 5th frame And the 6th frame ,calculate The trajectory is determined to be continuous. If the displacement exceeds the threshold, it means that the resonance peak has changed or been lost. Then, the linear prediction correction of the mean of the first three frames is performed to finally obtain the resonance peak trajectory sequence.

[0032] Specifically, the steps to obtain S3 are as follows: S31: Based on the local peak frequency position in the formant trajectory sequence, read the local peak frequency position corresponding to the formant in adjacent speech frames frame by frame, compare the changes in the arrangement of frequency positions before and after, and record the upward, downward and stationary frequency positions as directional markers. Arrange the directional markers according to the time sequence of the speech frames to generate the frequency direction change amount. Based on the formant trajectory sequence, for a selected order, the first frame is read one by one. Frame rate With the Frame rate ; The change is obtained by performing a difference operation. Dimensions are Set the dead zone coefficient for direction determination This coefficient is determined based on the frequency jitter variance caused by the microphone's ambient noise. The logical judgment is as follows: If... Marked as "1", it represents moving up; if A symbol marked "-1" indicates a downward shift; if Marked as "0" to represent stability, the frequency of the 15th frame is calculated using actual data. Frame 16 ,calculate The direction of recording is marked as "1", and so on. The state of all inter-frame intervals is defined, and these state codes are extracted in chronological order to form a state array. If more than 5 consecutive "0" marks are detected, it is defined as a holding period stable point, and the frequency direction change is generated.

[0033] S32: Based on the frequency direction change, divide the speech frame into segments according to the training syllable order, read the arrangement of the direction markers in each segment, count the distribution of the up-moving markers, down-moving markers, and stationary markers, organize the correspondence between each speech frame segment and the direction markers, and obtain the segment direction distribution. Based on the change in frequency direction, read the total duration of the current training syllable (e.g., "i"). It is divided into three sections: initiation, resistance holding, and resistance release, each accounting for a certain percentage. The corresponding frame index ranges are respectively , and Count the number of direction markers "1" segment by segment. The number of down-shifted markers "-1" And the number of times the marker "0" is left Calculate the distribution ratio of various tags within the segment. (20 frames per segment). Taking the initial segment as an example, if... Then the upward shift ratio is calculated as follows: The downward shift ratio is The percentage of people staying is A multidimensional correlation mapping is established between these proportions and their corresponding segment numbers and syllable types. If a certain proportion deviates significantly from common sense, it is marked as a collection disturbance frame in the background to obtain the segment directional distribution.

[0034] Table 2: Quantitative Table of Directional Distribution of Training Syllables in Segments ; See Table 2, which details the dynamic frequency distribution of different syllable segments during children's pronunciation, providing a precise statistical benchmark for subsequent comparison with the standard modality.

[0035] S33: Call the segment direction distribution quantity, read the local peak frequency position of the formant corresponding to the training syllable in the standard pronunciation, sort out the standard pronunciation change direction distribution, perform corresponding speech frame segment comparison for the children's pronunciation change direction distribution and the standard pronunciation change direction distribution, filter the speech frame segments with different direction distributions, and obtain the articulation direction deviation segments. The segment directional distribution is invoked, and the standard distribution parameters of the training syllable "i" are retrieved from the standard library to extract the baseline value of the standard starting segment upward shift ratio. Obtain the percentage of upward shift of the starting segment corresponding to the child's pronunciation. Perform deviation calculation Set a threshold for determining the direction of articulation deviation. This threshold was obtained by analyzing the variance of the orientation consistency distribution of a large sample of healthy children. The result was calculated by substituting the data... It was confirmed that there was a significant deviation in the articulation direction, and the frame index range covered by this starting segment was recorded. It is then marked as an abnormal point in direction. If the deviation is lower than the threshold, it is marked as "normal". All segments are traversed synchronously and the frame intervals that trigger the threshold are summarized to obtain the segments where the articulation direction deviates.

[0036] Specifically, the steps to obtain S4 are as follows: S41: Based on the speech frame segments with differences in the articulation direction deviation segment, select the speech frame corresponding to the child's pronunciation, extract the local peak frequency position corresponding to the formant in the speech frame, delineate the adjacent frequency boundary around the frequency position, filter the frequency points between the boundary, record the correspondence between the speech frame number and the frequency boundary, and generate the formant neighborhood frequency range. Based on the deviation of the articulation direction, locate the original spectral data from frame 1 to frame 20, and extract the center frequency of the formant in each frame. ,by Define a width of [value] on each side of the reference. Frequency monitoring window; Determine the left boundary And the right boundary If a certain frame Then the calculation boundary range is The process involves filtering out all discrete frequency sampling points within the closed interval, establishing a dynamic binding relationship between the speech frame number and the window, recording the total number of frequency points and the initial value of the energy integral within the window, establishing the corresponding address of each frame and its dedicated window through an index mapping table, and performing an edge truncation operation if the frequency coordinates exceed the legal range to generate the frequency range of the formant neighborhood.

[0037] S42: Based on the frequency range of the formant neighborhood, scan the changes in the envelope amplitude within the neighborhood frequency range point by point, read the envelope amplitude of each frequency point of the child's pronunciation, compare the arrangement of the envelope amplitude of adjacent frequency points, record the peak frequency position of the child's pronunciation envelope, and establish the peak position of the child's envelope. Based on the frequency range of the resonant peak neighborhood, a linear envelope extraction logic is used to scan point by point to read the frequency points. amplitude Set the exponentially weighted moving average coefficient (smoothing coefficient). Calculate the smoothed envelope value Compare its envelope amplitude with the relationship between the left and right adjacent points. When the center point is greater than the two sides, that is... and Record the current frequency point. The envelope peak value; if detected Inside, The envelope amplitude is The two sides are respectively and ,confirm The center of the envelope peak is determined. This value is stored in the feature vector. If multiple peaks exist within the window, the one with the highest amplitude is selected to establish the location of the child's envelope peak.

[0038] S43: Call the child's envelope peak position, read the corresponding training syllable and the corresponding neighboring frequency range of the envelope peak frequency position in the standard pronunciation, perform position difference comparison between the child's pronunciation envelope peak frequency position and the standard pronunciation envelope peak frequency position, filter the envelope position offset speech frame segment and the corresponding neighboring frequency range, and obtain the envelope offset segment; The system retrieves children's envelope peak position data and simultaneously calls up the standard envelope peak coordinates defined for the same syllable and frequency band in the standard model library. Extracting the peak position of children's pronunciation Perform absolute value calculation of position offset Set the allowable deviation of envelope drift. This value is set with reference to the human auditory perception limit of frequency changes, and a comparison judgment is performed. The frame was confirmed to have an envelope structure offset. A continuous sequence of such offsets was recorded. A segment merging operation was performed, and speech frames with a duration of less than two (i.e.,...) were removed. If the instantaneous jitter of the signal is highly consistent with the detection results, it is marked as "good fit". Finally, the time segments of the stable offset are summarized to obtain the envelope offset segment.

[0039] Specifically, the steps to obtain S5 are as follows: S51: For speech frame segments with envelope position offset and corresponding neighboring frequency ranges in the envelope offset segment, read the spectrum data of the neighboring frequency range within the speech frame segment, check the amplitude arrangement of adjacent frequency points in frequency order, identify the peak position and valley position, record the correspondence between the speech frame number and the frequency position, and generate peak and valley frequency distribution. For speech frame sequences with offsets in the envelope offset segment, reread them in... The fine spectral distribution within the frequency range is calculated using a first-order central difference operator to determine the gradient values ​​in the frequency direction. The coordinates where the gradient changes from positive to negative and crosses zero are identified as the peak frequency points. The coordinates where the gradient changes from negative to positive and crosses zero are identified as valley frequency points. The peak value in the 5th frame was recorded. And the valley value is located These feature points are bound to frame numbers, and the relative distance and amplitude difference between peak and valley pairs are extracted. If the waveform within the window is too flat and there are no obvious peaks and valleys, it is marked as "feature missing" and peak and valley frequency distribution is generated.

[0040] S52: Call the peak-valley frequency distribution data, read the spectrum data of the corresponding training syllable and the same neighborhood frequency range in the standard pronunciation, identify the peak position and valley position of the standard pronunciation, perform peak-valley frequency position comparison between the child's pronunciation and the standard pronunciation, filter out the speech frame segments with different peak-valley distribution, and obtain the peak-valley difference segments. Use peak-valley frequency distribution data to extract standard peak points from the standard library. Compared with standard valley point Perform a two-point offset comparison operation and calculate the peak deviation. and valley deviation Establish a comprehensive morphological difference judgment benchmark This benchmark is determined by the critical distance from pronunciation accuracy clustering analysis. Execute the discrimination logic: If... This difference is recorded as significant; in this example, due to... The segment to which the frame belongs is determined to be a difference region. All frames are traversed and their difference flags are recorded. If the difference values ​​are all below the threshold, it is determined to be "morphological fitting", and the peak-valley difference segment is obtained.

[0041] Table 3: Statistical Table of Abnormal Speech Feature Parameter Comparison ; As shown in Table 3, by quantizing the peak offset and comparing it with the threshold, the time segment in which the sound spectrum structure is distorted and its corresponding abnormal state can be accurately located.

[0042] S53: Based on the peak-valley difference segment, call the speech frame segment with different directional distribution in the speech direction deviation segment, perform time comparison of the two types of speech frame segments, filter the time overlapping segments, arrange them according to the training syllable order and record the corresponding neighborhood frequency range to obtain the abnormal speech detection result; Based on the time interval of peak-valley difference segment The time range of the articulation direction deviating from the segment. ; Perform set intersection operation The calculated overlapping section is Duration is Read the corresponding neighborhood frequency range within the overlapping time period. Set a minimum duration threshold for anomaly detection. Perform comparison judgment If the overlapping segment is confirmed as a valid piece of evidence of abnormal pronunciation, and the duration of overlap is insufficient... If the abnormal speech is detected, it is considered a physiological tremor and filtered out. These segments are then rearranged according to the syllable sequence, associated with abnormality type markers (such as "formant tongue position deviation"), and a structured diagnostic report is generated to obtain the abnormal speech detection results.

[0043] S54: The main control processor generates treatment feedback information based on the abnormal segment number, abnormal feature category, spectral structure identifier and syllable abnormal type in the abnormal speech detection results, and controls the treatment feedback output module to output the treatment feedback information in at least one of the following forms: time axis mark, abnormal segment highlight, abnormal type label, spectral structure prompt, syllable abnormal prompt or training task prompt. Specifically, after obtaining the abnormal speech detection results, the main control processor reads the abnormal segment number, abnormal feature category, spectral structure identifier, and syllable abnormality type from the abnormal speech detection results. Combining this with time overlap segments, training syllable order, and neighboring frequency range, it maps the abnormal speech detection results to the corresponding training syllables, training phrases, or training sentence fragments in the child's current training task. For the same abnormal segment, the main control processor further correlates the direction deviation type and deviation intensity level in the articulation direction deviation segment, the envelope peak deviation and deviation significance level in the envelope deviation segment, and the peak deviation and valley deviation in the peak-valley difference segment, generating treatment feedback information corresponding to the abnormal segment number.

[0044] Treatment feedback information includes abnormal training syllables, start and end times of abnormal speech frames, abnormal segment numbers, abnormal feature categories, spectral structure identifiers, frequency band coverage, envelope peak offset, offset significance level, syllable abnormality type, and training adjustment prompts. Specifically, abnormal training syllables indicate the specific syllable in the child's current training task where an abnormality occurs; abnormal speech frame start and end times indicate the position of the abnormal speech on the training timeline; abnormal feature categories indicate deviations in articulation direction, envelope peak offsets, peak-valley structure differences, or multiple feature overlap abnormalities; spectral structure identifiers and frequency band coverage indicate the neighborhood of the abnormal formant and its corresponding frequency range; syllable abnormality type indicates the type of abnormal pronunciation corresponding to the training syllable; and training adjustment prompts suggest to the therapist that they adjust subsequent training syllables, training frequency, pronunciation prompting methods, training rhythm, or repetition of training segments.

[0045] In one implementation, when the abnormal speech detection results show a time overlap anomaly in frames 5 to 7, with the abnormal feature category being articulation direction deviation and peak-valley structure difference, a frequency band coverage of [400Hz, 600Hz], and the syllable anomaly type being "formant tongue position deviation," the main control processor generates treatment feedback information including the abnormal segment number, the corresponding training syllable, the start and end times of the abnormal speech frame, the spectral structure identifier, and the syllable anomaly type. It then controls the treatment feedback output module to highlight frames 5 to 7 on the timeline of the training interface, while simultaneously displaying an anomaly type label and spectral structure prompt at the corresponding training syllable. The therapist can then use the treatment feedback information to repeatedly train the training syllable, slow down the training pace, or strengthen the articulation point prompts, thereby creating a closed loop between speech acquisition, anomaly detection, anomaly localization, result output, and treatment feedback within the speech therapy instrument at the child's training site.

[0046] Please see Figure 2 An abnormal speech detection system for speech therapy treatment instruments, the system includes a microphone, a speech acquisition card, a main control processor, a standard pronunciation data storage unit, and a treatment feedback output module; The microphone is used to collect air vibration signals generated during children's vocal training; The voice acquisition card is connected to a microphone and is used to convert air vibration signals into a discrete sampling point sequence corresponding to voltage change signals; The standard pronunciation data storage unit is used to store the distribution of the standard pronunciation change direction, the position of the peak frequency of the standard pronunciation envelope, and the position of the peak and valley frequencies of the standard pronunciation. The main control processor is connected to the voice acquisition card, the standard pronunciation data storage unit, and the therapeutic feedback output module, respectively. The main control processor includes: The voice acquisition module acquires the discrete sampling point sequence output by the voice acquisition card, divides the voice frames according to time, filters and sorts the voice frames with continuously changing amplitudes to obtain the voice frame sequence. The frequency peak extraction module converts the discrete sampling point sequence corresponding to the speech frame sequence into a frequency distribution frame by frame, compares the magnitude of adjacent frequency points, filters the local peak frequency positions and checks the inter-frame continuity to obtain the formant trajectory sequence. The direction analysis module compares the frequency position of the frequency position of adjacent speech frames frame by frame, based on the local peak frequency position in the formant trajectory sequence. It divides speech frames into segments according to training syllables, counts the distribution of the direction of pronunciation change of children, and compares it with the distribution of the direction of pronunciation change of standard pronunciation in the standard pronunciation data storage unit. It then filters out the difference segments and obtains the speech direction deviation segments. The envelope localization module selects children's speech frames based on the deviation of the articulation direction, defines the neighborhood frequency range around the local peak frequency position corresponding to the formant, and scans the envelope amplitude.

[0047] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for detecting abnormal speech in a speech therapy instrument, characterized in that, The speech therapy instrument includes a microphone, a voice acquisition card, a main control processor, a standard pronunciation data storage unit, and a treatment feedback output module. The microphone is connected to the voice acquisition card, the voice acquisition card is connected to the main control processor, and the main control processor is connected to both the standard pronunciation data storage unit and the treatment feedback output module. The method is executed by the main control processor in the context of children's speech therapy training and includes the following steps: S1: The microphone collects the air vibration signal generated by the child's vocal training, and the voice acquisition card converts the air vibration signal into a discrete sampling point sequence corresponding to the voltage change signal. The main control processor divides the voice frames according to time, filters and sorts the voice frames with continuously changing amplitude to obtain the voice frame sequence. S2: The main control processor converts the discrete sampling point sequence corresponding to the speech frame sequence into a frequency distribution frame by frame, compares the amplitude of adjacent frequency points, filters the local peak frequency positions and checks the inter-frame continuity to obtain the formant trajectory sequence. S3: The main control processor compares the frequency position of adjacent speech frames frame by frame based on the local peak frequency position in the formant trajectory sequence, the frequency position shifts up, down and stays, divides the speech frame segments according to the training syllables, counts the distribution of the direction of change of children's pronunciation, and calls the standard pronunciation change direction distribution in the standard pronunciation data storage unit for comparison, filters the difference segments, and obtains the speech direction deviation segments. S4: The main control processor selects a child's pronunciation speech frame according to the pronunciation direction deviation segment, defines a neighborhood frequency range around the local peak frequency position corresponding to the formant, scans the envelope amplitude and compares it with the standard pronunciation envelope peak frequency position in the standard pronunciation data storage unit to obtain the envelope offset segment.

2. The abnormal speech detection method for speech therapy instruments according to claim 1, characterized in that: The speech frame sequence includes frame timestamp, frame duration, frame energy characterization, and frame validity identifier. The formant trajectory sequence includes formant center frequency, frequency variation amplitude, frequency continuity index, and formant level identifier. The articulation direction deviation segment includes direction deviation type, deviation intensity level, segment span range, and syllable association identifier. The envelope offset segment includes envelope peak offset, offset direction identifier, frequency band coverage range, and offset significance level.

3. The abnormal speech detection method for speech therapy instruments according to claim 1, characterized in that: The steps for obtaining S1 are as follows: S11: The microphone senses the air vibration generated by the child's vocal training and outputs a voltage change signal. The voice acquisition card reads the voltage change signal, records the voltage amplitude according to the sampling time sequence, organizes the correspondence between the sampling time and the sampling amplitude, and generates a discrete sampling point sequence. S12: The main control processor, based on the discrete sampling point sequence, continuously divides the speech frames according to the sampling time, reads the arrangement order of the sampling points frame by frame, checks the amplitude connection status of adjacent sampling points, and retains the speech frames with complete sampling point order and smooth amplitude connection to obtain a candidate speech frame set. S13: The main control processor reads the starting sampling time of each audio frame for the candidate audio frame set, arranges and retains the audio frames in chronological order, records the correspondence between the sampling point number and amplitude of each audio frame, and filters audio frames with continuously changing amplitude to obtain an audio frame sequence.

4. The abnormal speech detection method for speech therapy instruments according to claim 1, characterized in that: The steps for obtaining S2 are as follows: S21: The main control processor reads the sampling point arrangement order frame by frame based on the discrete sampling point sequence corresponding to each speech frame in the speech frame sequence, expands the sampling amplitude according to the frequency position, forms the frequency point and amplitude correspondence relationship, compares the amplitude connection state of adjacent frequency points, and generates a frequency amplitude correspondence sequence. S22: The main control processor calls the frequency amplitude corresponding sequence, reads the amplitude of each frequency point, the left frequency point, and the right frequency point point one by one, compares the magnitude of the three amplitudes, filters the frequency positions with amplitudes higher than the amplitudes of the left and right adjacent frequency points, and arranges them in order of frequency from low to high to obtain the local peak frequency position set. S23: The main control processor selects the frequency positions corresponding to the formant distribution of the child's pronunciation according to the local peak frequency position set, arranges the selected frequency positions in the order of speech frame time, checks the continuous state of frequency position changes between adjacent speech frames, and obtains the formant trajectory sequence.

5. The abnormal speech detection method for speech therapy instruments according to claim 1, characterized in that: The steps for obtaining S3 are as follows: S31: The main control processor reads the local peak frequency position of the formant in the formant trajectory sequence frame by frame based on the local peak frequency position of the formant in the adjacent speech frames, compares the changes in the arrangement of the frequency positions before and after, records the upward, downward and stationary frequency positions as direction marks, arranges the direction marks according to the time sequence of the speech frames, and generates the frequency direction change amount. S32: The main control processor divides the speech frame into segments according to the training syllable order based on the frequency direction change, reads the arrangement of direction marks in the segment segment by segment, counts the distribution of up-moving marks, down-moving marks, and stationary marks, sorts out the correspondence between each speech frame segment and the direction marks, and obtains the segment direction distribution. S33: The main control processor calls the segment direction distribution quantity, reads the local peak frequency position corresponding to the standard pronunciation formant of the corresponding training syllable in the standard pronunciation data storage unit, sorts out the standard pronunciation change direction distribution, performs a corresponding speech frame segment comparison between the child's pronunciation change direction distribution and the standard pronunciation change direction distribution, filters out speech frame segments with different direction distributions, and obtains the articulation direction deviation segment.

6. The abnormal speech detection method for speech therapy instruments according to claim 1, characterized in that: The steps for obtaining S4 are as follows: S41: The main control processor selects the speech frame corresponding to the child's pronunciation based on the speech frame segment with differences in the speech direction deviation segment, extracts the local peak frequency position corresponding to the formant in the speech frame, delineates the adjacent frequency boundary around the frequency position, filters the frequency points between the boundary, records the correspondence between the speech frame number and the frequency boundary, and generates the formant neighborhood frequency range. S42: The main control processor scans the changes in the envelope amplitude within the neighborhood frequency range point by point based on the formant neighborhood frequency range, reads the envelope amplitude of each frequency point of the child's pronunciation, compares the arrangement of the envelope amplitude of adjacent frequency points, records the peak frequency position of the child's pronunciation envelope, and establishes the peak position of the child's envelope. S43: The main control processor calls the child's envelope peak position, reads the corresponding training syllable and the standard pronunciation envelope peak frequency position in the corresponding neighborhood frequency range in the standard pronunciation data storage unit, performs a position difference comparison between the child's pronunciation envelope peak frequency position and the standard pronunciation envelope peak frequency position, filters the envelope position offset speech frame segment and the corresponding neighborhood frequency range, and obtains the envelope offset segment.

7. The abnormal speech detection method for speech therapy instruments according to claim 1, characterized in that: The method further includes: S5: The main control processor reads the spectrum data of the neighboring frequency range for the envelope offset segment, identifies the peak position and valley position and compares them with the standard pronunciation peak and valley frequency positions in the standard pronunciation data storage unit, and then compares them with the articulation direction deviation segment in time to obtain the abnormal speech detection result. Based on the abnormal speech detection result, it generates treatment feedback information and sends the treatment feedback information to the treatment feedback output module. The abnormal speech detection results include abnormal segment number, abnormal feature category, spectral structure identifier, and syllable abnormality type; The treatment feedback information includes prompts for abnormal pronunciation segments, abnormal feature categories, abnormal syllable types, and training adjustment prompts.

8. The abnormal speech detection method for speech therapy instruments according to claim 7, characterized in that: The steps for obtaining S5 are as follows: S51: The main control processor reads the spectrum data of the neighboring frequency range within the speech frame segment and the corresponding neighboring frequency range in the envelope offset segment, checks the amplitude arrangement of adjacent frequency points in frequency order, identifies the peak position and valley position, records the correspondence between the speech frame number and the frequency position, and generates the peak-valley frequency distribution. S52: The main control processor calls the peak-valley frequency distribution, reads the standard pronunciation spectrum data of the corresponding training syllable and the same neighborhood frequency range in the standard pronunciation data storage unit, identifies the peak position and valley position of the standard pronunciation, performs a peak-valley frequency position comparison between the child's pronunciation and the standard pronunciation, filters out speech frame segments with differences in peak-valley distribution, and obtains peak-valley difference segments. S53: The main control processor calls the speech frame segments with different directional distributions in the speech direction deviation segment according to the peak-valley difference segment, performs time comparison of the two types of speech frame segments, filters out time-overlapping segments, arranges them according to the training syllable order and records the corresponding neighborhood frequency range, and obtains the abnormal speech detection result. S54: The main control processor generates treatment feedback information based on the abnormal segment number, abnormal feature category, spectral structure identifier and syllable abnormal type in the abnormal speech detection result, and controls the treatment feedback output module to output the treatment feedback information in at least one of the following forms: time axis mark, abnormal segment highlight, abnormal type label, spectral structure prompt, syllable abnormal prompt or training task prompt. The method is applied in the context of speech therapy training for children. After a child completes a single training syllable, training phrase, or training sentence, the therapeutic feedback output module outputs the corresponding abnormal pronunciation segment, abnormal feature category, and syllable abnormality type in real time or near real time to prompt the therapist to adjust the subsequent training syllables, training frequency, pronunciation prompting method, training rhythm, or repetition of training segments.

9. An abnormal speech detection system for speech therapy, characterized in that, The system is used to perform the abnormal speech detection method of the speech therapy instrument according to any one of claims 1-8, the system comprising: A microphone for collecting air vibration signals generated during children's vocal training; A voice acquisition card, connected to the microphone, is used to convert the air vibration signal into a discrete sampling point sequence corresponding to the voltage change signal; The standard pronunciation data storage unit is used to store the distribution of the standard pronunciation change direction, the peak frequency position of the standard pronunciation envelope, and the peak and valley frequency positions of the standard pronunciation. The main control processor is connected to the voice acquisition card, the standard pronunciation data storage unit, and the treatment feedback output module, respectively. The main control processor includes: The voice acquisition module is used to acquire the discrete sampling point sequence output by the voice acquisition card, divide the voice frames according to time, filter and sort the voice frames with continuously changing amplitude to obtain the voice frame sequence. The frequency peak extraction module is used to convert the discrete sampling point sequence corresponding to the speech frame sequence into a frequency distribution frame by frame, compare the magnitude of adjacent frequency points, filter the local peak frequency positions and check the inter-frame continuity to obtain the formant trajectory sequence. The direction analysis module is used to compare the frequency position of adjacent speech frames frame by frame based on the local peak frequency position in the formant trajectory sequence, the upward and downward shift and the dwell state of the frequency position, divide the speech frame into segments according to the training syllables, count the distribution of the direction of pronunciation change of children, compare it with the distribution of the direction of pronunciation change of standard in the standard pronunciation data storage unit, filter the difference segments, and obtain the speech direction deviation segments. The envelope localization module is used to select a child's pronunciation speech frame according to the articulation direction deviation segment, define a neighborhood frequency range around the local peak frequency position corresponding to the formant, scan the envelope amplitude and compare it with the standard pronunciation envelope peak frequency position in the standard pronunciation data storage unit to obtain the envelope offset segment. The anomaly detection module is used to read the spectral data of the neighborhood frequency range for the envelope offset segment, identify the peak position and valley position and compare them with the standard pronunciation peak and valley frequency positions in the standard pronunciation data storage unit, and perform a time comparison with the articulation direction deviation segment to obtain the abnormal speech detection result. The feedback output control module is used to generate treatment feedback information based on the abnormal speech detection results, and to control the treatment feedback output module to output the treatment feedback information. The therapeutic feedback output module is used to output therapeutic feedback information, including abnormal pronunciation segments, abnormal feature categories, and syllable abnormality types, in the context of children's speech therapy training.