Language disorder adjuvant therapy system based on speech recognition

By combining an adaptive speech recognition engine and a multimodal feedback device with a reinforcement learning engine, the shortcomings of existing systems in pathological speech recognition and dialect feature separation are addressed, enabling efficient neural remodeling and personalized treatment for language disorder rehabilitation.

CN120932641AInactive Publication Date: 2025-11-11JIANGXI DINGJI MINGCHUANG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511047948.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing speech recognition-based language disorder auxiliary treatment systems have shortcomings in terms of robustness of pathological speech recognition, confusion between dialect and pathological features, and disconnect between feedback mechanisms and neural remodeling, resulting in high misjudgment rates, low treatment compliance, and limited rehabilitation efficiency.

Method used

Employing an adaptive speech recognition engine, a multimodal feedback device, and a reinforcement learning engine, this system achieves accurate recognition of pathological speech and personalized rehabilitation training through dynamic threshold recognition, dialect adaptation, multimodal feedback, and personalized treatment path optimization.

Benefits of technology

It improves the efficiency of neural remodeling in the rehabilitation of language disorders, shortens the rehabilitation cycle, and enhances treatment compliance and rehabilitation outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0005522286320000011
    Figure HDA0005522286320000011
  • Figure HDA0005522286320000012
    Figure HDA0005522286320000012
  • Figure HDA0005522286320000013
    Figure HDA0005522286320000013
Patent Text Reader

Abstract

The invention belongs to the field of intelligent medical rehabilitation engineering, particularly relates to a speech recognition-based speech disorder adjuvant therapy system, and solves the problems of high pathological speech misjudgment rate, large dialect interference and single feedback in the prior art. A phoneme fault-tolerant threshold value is dynamically adjusted by constructing an adaptive speech recognition engine, regional features and pathological features are separated by using a dialect adaptation module, and three-dimensional simulation, AR real-time correction and touch graded vibration of vocal organs are realized in combination with a multi-modal feedback device; the pathological speech recognition precision is improved; and the neural remodeling efficiency is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent medical rehabilitation engineering, and more specifically, to a speech recognition-based language disorder auxiliary treatment system. Background Technology

[0002] Speech recognition-based language disorder assistance systems fall within the interdisciplinary field of intelligent medicine and rehabilitation engineering. They provide real-time intervention strategies by analyzing patients' speech characteristics. In recent years, with the widespread adoption of deep learning algorithms, traditional systems have achieved basic pronunciation assessments, but bottlenecks remain in areas such as adaptability to pathological speech, accuracy of multimodal feedback, and optimization of personalized treatment pathways, hindering substantial improvements in rehabilitation efficiency. Current technological shortcomings include:

[0003] 1. Lack of robustness in pathological speech recognition

[0004] Existing systems generally employ generic speech recognition engines, which cannot effectively distinguish between physiological pronunciation variations and pathological pronunciation errors in patients with speech disorders. When dealing with vocal cord control abnormalities in patients with articulation disorders or prosodic disorders in patients with aphasia, fixed recognition thresholds lead to increased misjudgment rates, often labeling compensatory pronunciation strategy errors as pronunciation errors, thus generating misleading training instructions. The fundamental flaw of this recognition mechanism lies in its lack of quantitative modeling capabilities for the dynamic features of pathological speech.

[0005] 2. Confusing dialect with pathological features

[0006] Traditional approaches neglect the interference of dialectal background on speech assessment, conflating regional pronunciation habits (such as dialectal differences in retroflex and alveolar consonants) with organic speech disorders. The system cannot construct a separate assessment model for dialectal and pathological features, leading to systematic biases in training programs tailored to dialectal populations. This deficiency results in numerous false positive diagnoses in cross-regional clinical applications, severely reducing treatment adherence.

[0007] 3. Disconnect between feedback mechanisms and neural remodeling

[0008] Current mainstream systems rely on single visual or auditory feedback, which is severely incompatible with the human multi-sensory collaborative learning mechanism. Visual feedback can only present a two-dimensional static diagram of pronunciation and cannot dynamically simulate the movement of deep organs such as the hyoid bone and throat; tactile feedback is limited to simple vibration cues and has not established a spatial mapping relationship between the location of incorrect pronunciation and the intensity of vibration. This fragmented feedback mode makes it difficult to activate the neuroplasticity of the brain's sensorimotor cortex and hinders the functional reconstruction of motor articulation disorders.

[0009] Therefore, a speech recognition-based language disorder auxiliary treatment system is proposed to address the above problems. Summary of the Invention

[0010] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a speech recognition-based language disorder auxiliary treatment system to solve the problems mentioned in the background art.

[0011] To achieve the above objectives, the present invention provides the following technical solution: a speech recognition-based language disorder auxiliary treatment system, comprising a user terminal, an adaptive speech recognition engine, a treatment server, and a multimodal feedback device; the user terminal is equipped with a speech acquisition module and an interactive interface for real-time acquisition of patient speech and display of treatment feedback; the adaptive speech recognition engine is connected to the user terminal and includes a noise suppression unit, a dynamic threshold recognition unit, and a dialect adaptation module, wherein the noise suppression unit performs environmental noise reduction processing on the input speech, the dynamic threshold recognition unit dynamically adjusts the phoneme recognition tolerance threshold according to the patient's language disorder type, and the dialect adaptation module loads a regional pronunciation feature library to suit patients with dialect disorders; the treatment server stores a personalized treatment strategy database and generates real-time multimodal feedback instructions based on the speech recognition results; the multimodal feedback device includes a visual speech organ movement simulation unit, a tactile vibration feedback unit, and an AR augmented reality correction guidance module, respectively providing three-dimensional movement simulation of speech organs, tactile vibration feedback graded by syllable accuracy, and virtual pronunciation guidance superimposed in a real environment.

[0012] Furthermore, the dynamic threshold recognition unit automatically classifies articulation disorders, aphasia, or fluency disorders by analyzing the patient's historical data, and dynamically adjusts the phoneme similarity judgment threshold according to the degree of improvement in pronunciation during the treatment process. When the pronunciation stability is improved by more than 15% for three consecutive times, the error tolerance threshold is automatically relaxed by 5% to 10%. When the error rate of a specific phoneme is detected to be higher than 50% for an extended period, the corresponding phoneme threshold is tightened by 8% to 12%.

[0013] Furthermore, the dialect adaptation module establishes a benchmark pronunciation model by collecting speech samples from healthy individuals in the target dialect area, extracts a set of dialect-specific phoneme replacement rules, and automatically marks the patient's pronunciation as acceptable when it deviates from the standard phonemes but conforms to the preset dialect replacement rules during the speech recognition process. At the same time, it generates a separation assessment report of dialect features and pathological pronunciation.

[0014] Furthermore, the visualized speech organ motion simulation unit synchronously generates a three-dimensional dynamic tongue trajectory projection, a dynamic heat map of lip shape changes, and a visualized model of vocal cord vibration frequency. The tongue trajectory projection displays the distance parameters between the tongue tip and back and the hard palate in real time. The lip shape heat map uses color depth to indicate the contraction intensity of the orbicularis oris muscle. The vocal cord vibration model displays the fundamental frequency stability through waveform diagrams.

[0015] Furthermore, the treatment server has a built-in reinforcement learning engine. By recording the recognition results and feedback data of each pronunciation training session, it uses the Q-learning algorithm to optimize the difficulty sequence of practice items in the treatment strategy database and dynamically generates a personalized treatment path that includes the syllable complexity increment rate and training duration allocation factor. When the accuracy of a specific phoneme increases by more than 20% in five consecutive training sessions, a high-difficulty compound training module is automatically inserted.

[0016] Furthermore, the tactile vibration feedback unit switches between three working modes based on the real-time syllable recognition accuracy: when the accuracy is greater than or equal to 90%, it outputs a continuous and stable vibration at a frequency of 100 Hz; when the accuracy is between 70% and 89%, it outputs a pulse vibration with an interval of 0.5 seconds; and when the accuracy is less than 70%, it outputs a high-frequency warning vibration at a frequency of 200 Hz. The vibration intensity automatically increases by 30% to 50% depending on the position of the incorrect syllable.

[0017] Furthermore, it also includes a remote diagnosis and treatment interface module, which converts speech recognition results into standardized articulation disorder assessment scale data, automatically generates diagnosis and treatment reports containing heatmaps of the spatial distribution of error syllables and temporal spectra of error frequencies, and supports multilingual therapists to collaboratively annotate abnormal pronunciation segments through a speech annotation system and generate annotation confidence scores.

[0018] A speech disorder treatment method based on a speech recognition-based auxiliary treatment system includes: establishing a patient-specific speech recognition model through a dynamic threshold recognition unit and optimizing phoneme error tolerance parameters in real time; using an AR augmented reality correction guidance module to overlay virtual speech organ standard positions dynamically indicating deviation angles on the patient's real-time image; adjusting the intensity of lip and tongue muscle exertion in real time according to the vibration pattern changes of a tactile vibration feedback unit; automatically activating a high-difficulty training module when the accuracy rate improves by more than 15% after three consecutive training sessions; and triggering decomposed muscle training animations when the error rate of a specific phoneme remains above 50%.

[0019] The technical effects and advantages of this invention are as follows:

[0020] Compared with existing technologies, this invention constructs an adaptive speech recognition engine and uses a dynamic threshold recognition unit to adjust phoneme tolerance parameters in real time to accurately distinguish between pathological and compensatory pronunciation variations. It utilizes a dialect adaptation module to load a regional feature database and establish a dialect-pathology separation model, eliminating the interference of regional pronunciation habits on diagnosis. Combined with a multimodal feedback device, a visualized speech organ motion simulation unit generates a three-dimensional dynamic tongue trajectory and vocal cord vibration model. An AR augmented reality correction guidance module overlays virtual organ standard positions onto the patient's real-time image, providing concrete guidance for pronunciation errors. A synchronous tactile vibration feedback unit outputs differentiated vibration patterns based on syllable accuracy, establishing a spatial mapping between incorrect pronunciation locations and tactile feedback. The treatment server incorporates a reinforcement learning engine to continuously optimize personalized treatment paths and dynamically adjusts the training difficulty sequence based on the Q-learning algorithm. These technologies work synergistically to achieve simultaneous optimization of environmental noise suppression and physiological parameter acquisition at the hardware level, and to complete closed-loop control of pathological feature extraction and rehabilitation strategy generation at the software level, ultimately improving the efficiency of neural remodeling in language disorder rehabilitation. Attached Figure Description

[0021] Figure 1 This is a system framework diagram of the present invention.

[0022] Figure 2 This is a flowchart of the dynamic threshold adjustment process of the present invention.

[0023] Figure 3 This is a diagram of the multimodal feedback collaboration mechanism of the present invention.

[0024] Figure 4 This is a diagram illustrating the dialect-pathology separation process of the present invention. Detailed Implementation

[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Implementation Process 1: Dynamic Threshold Adaptive Therapy (Focusing on Articulation Dysfunction Rehabilitation)

[0026] Scenario: Rehabilitation training for patients with dysarthria after stroke

[0027] Core Innovation: Real-time Adjustment of Phoneme Recognition Error Tolerance Threshold

[0028] Step 1: Initial parameter configuration

[0029] The patient wears AR glasses with an integrated microphone (user terminal) and selects the "Articulation Disorder - Apical Dysfunction" category.

[0030] The system retrieves historical medical record data: the basic phoneme error rate is 42%, and the main errors are concentrated in the / z / and / c / sounds.

[0031] Step 2: Dynamic threshold generation

[0032] When the algorithm executes, the real-time parameter adjustment sub-module is started by the dynamic threshold recognition unit. Based on the patient's historical error rate data (0.42), the error rate weight coefficient 0.336 is calculated (formula: historical error rate × 0.8). Then, through the dynamic threshold formula [threshold base = healthy person's standard threshold × (1 + error rate weight coefficient)], the initial threshold for the / z / sound is set to 66.8 milliseconds (i.e., 30 milliseconds × (1 + 0.336)), and at the same time, a real-time elastic tolerance interval of ±8.5 milliseconds is generated to continuously adapt to the fluctuations in the patient's pronunciation state.

[0033] Step 3: Real-time pronunciation training

[0034] In the real-time evaluation process of a language disorder assisted treatment system based on speech recognition, the patient reads the target word "zongzi" (focusing on the pronunciation of the / z / phoneme). The system collects the original audio through a high-sensitivity microphone array. The noise suppression unit eliminates 60 dB environmental noise based on the FFT spectral subtraction method (mainly filtering out environmental interference in the 100 - 4000 Hz frequency band). The preprocessed audio signal is transmitted to the deep speech recognition engine for millisecond-level parsing: the measured duration of the / z / phoneme is 58 milliseconds (this value exceeds the standard range of 40 - 45 milliseconds for the healthy population, but is lower than the personalized threshold of 66.8 milliseconds dynamically calculated based on the patient's historical error rate - this threshold is generated in real time by the dynamic threshold recognition module through the formula [healthy person's standard threshold 30 milliseconds × (1 + historical average error rate 0.42 × weight coefficient 0.8)]). Based on this, the system determines that the current pronunciation is at the "acceptable" level (judgment logic: when the measured duration ≤ dynamic threshold, activate the green passage instruction). At the same time, the adaptive algorithm generates an elastic tolerance band of ±8.5 milliseconds (calculation formula: threshold base × 13%, update period 0.5 seconds) for real-time adaptation of subsequent continuous speech streams, and outputs a visual feedback signal (green progress bar + 87 - point score) through the HMI interface.

[0035] Step 4: Dynamic threshold adjustment

[0036] In the adaptive training cycle of a speech recognition-based language disorder auxiliary treatment system, the real-time parameter tuning submodule strictly follows the algorithm logic of "performing threshold update every 3 training sessions": when the error rate record of the patient in three consecutive training sessions shows a gradient decreasing trend of 42%→35%→28%, the system calls the progress magnitude calculation formula ΔP=[(current error rate-previous error rate) / previous error rate] for dynamic analysis (first iteration ΔP1=(35%-42%) / 42%≈-0.167, second iteration ΔP2=(28%-35%) / 35%≈-0.200), and calculates the composite progress magnitude ΔP=-0.183 (satisfying the rigid judgment condition of ΔP≤-0.15). The system then activates the threshold compression mechanism, and based on the formula [new threshold = current threshold × (1 - progress reward coefficient)] (fixed reward coefficient 0.05), precisely updates the / z / phoneme dynamic threshold from 66.8 milliseconds to 63.46 milliseconds (mathematical verification: 66.8 × (1 - 0.05) = 63.46 ± 0.03 milliseconds). Simultaneously, four-fold collaborative optimization is triggered:

[0037] ① Tolerance band compression: The adaptive elasticity range shrinks from ±8.5 milliseconds to ±7.2 milliseconds (a shrinkage rate of 15.3%, reducing the scoring tolerance).

[0038] ② Historical weight decay: The patient error rate weight coefficient was iterated from 0.336 to 0.319 (reducing the constraint of historical performance on the standard).

[0039] ③ Scoring benchmark reconstruction: The original 58-millisecond pronunciation score under the new threshold increased from 87 to 92 points (the slope of the activation function increased by 37%).

[0040] ④ Backtracking defense mechanism: When the preset error rate rebound is greater than 5%, it automatically rolls back to the previous threshold (to build a safe training boundary).

[0041] This algorithm, through millisecond-level threshold iteration (error control <0.1%), ensures the continuity of patient training (threshold change delay <200ms) while gradually bringing rehabilitation standards closer to the level of healthy individuals (40-45 milliseconds). Its innovation lies in the ΔP composite calculation model eliminating the interference of single fluctuations, and the fixed reward coefficient design avoiding the risk of human intervention. Ultimately, in clinical trials, it shortened the patient's target achievement period by 23% ± 4% (p < 0.01).

[0042] Step 5: Multimodal feedback linkage

[0043] In an augmented reality rehabilitation training scenario, the system uses the millimeter-wave radar array of the AR glasses to capture the movement trajectory of the patient's tongue in real time (sampling rate 120Hz). It calculates that the distance between the tip of the tongue and the palate is 5.3mm (exceeding the healthy standard value of 3.0mm), immediately triggering a red holographic projection warning box (continuous flashing frequency 2Hz) and superimposing the dynamic error parameter "+2.3mm" in the center of the field of view. The tactile feedback system is synchronously activated: according to the preset response logic tree (the current pronunciation accuracy rate of 35% falls within the medium warning range of 30%-70%), the tactile wristband performs pulsed vibration intervention. The vibration intensity is calculated via a dynamic enhancement formula:

[0044] Intensity = base intensity (15mN·s) × (1 + deviation coefficient)

[0045] (Deviation coefficient = actual distance / standard distance = 5.3 / 3.0 ≈ 0.77)

[0046] The output intensity is upgraded to 26.55mN·s (an increase of 77%), and the pulse pattern uses a 10Hz intermittent square wave (continuous for 200ms → interval of 300ms in a cycle). At the same time, the wrist displacement is monitored through the IMU sensor. If a resistant swing (acceleration > 3g) is detected, the intensity is automatically reduced by 20% to prevent muscle stress.

[0047] Implementation process two: Dialect integration therapy (for dialectal speech confusion)

[0048] Scenario: Rehabilitation of Cantonese-speaking patients with confusion between / n / and / l / sounds

[0049] Core innovation: Separation of dialect features and pathological pronunciation

[0050] Step 1: Construction of a dialect feature library

[0051] In the multi-dialect adaptation module of a language disorder assisted therapy system based on speech recognition, when the patient selects the Cantonese training mode, the system automatically retrieves the Cantonese pronunciation feature database. For the / n / nasal consonant feature of the target word "你", the system identifies that the median nasal cavity resonance duration of Cantonese pronunciation is 110 milliseconds (standard deviation σ = ±5 milliseconds), which is 22.2% longer than the standard Mandarin benchmark value of 90 milliseconds. Based on the statistical normal distribution model (95% confidence interval), the system establishes a dynamic acceptable interval of [100, 120] milliseconds (calculation logic: median ± 2σ), and simultaneously generates a triple Cantonese feature adaptation strategy:

[0052] 1. Judgment interval expansion: The upper limit is relaxed by 33% compared to Mandarin (120ms vs 90ms) to match the characteristic of the longer air flow path in Cantonese pronunciation.

[0053] 2. Formant calibration: Add a center compensation frequency of 1120Hz for Cantonese features (the Mandarin benchmark is 950Hz).

[0054] 3. Real-time auditory weighting: When the unique laryngeal vibration of Cantonese is detected (VOT = -25ms), the scoring weighting coefficient of nasal cavity duration is automatically increased to 0.55 (default 0.3).

[0055] Step 2: Isolation of pathological features

[0056] In cross-dialect intelligent diagnosis scenarios, when a patient reads the target word "milk," the system uses a multi-dimensional sensor array to collect speech and tongue position dynamics in parallel:

[0057] / n / sound feature analysis: Nasal resonance duration is 115 milliseconds (real-time spectrum analysis accuracy ±1.2ms), which is higher than the standard Cantonese range [100, 120 milliseconds] in the dialect feature library. This triggers the dialect recognition protocol (confidence level 97%), generates a blue "dialect feature" label, and dynamically relaxes the rehabilitation standard of this phoneme to the upper limit of 120 milliseconds.

[0058] / l / sound pathology discrimination: the tongue pressure sensor grid detected that the contact area of ​​the tongue side was only 30% (health standard ≥60%), and ultrasound imaging showed that the amplitude of tongue muscle activity decreased by 63%. Dual-modal cross-validation triggered a red warning for "pathological pronunciation" (the system automatically reduced the dialect feature weight to 0.3 and increased the pathology discrimination weight to 0.9).

[0059] Step 3: Dynamic rule generation

[0060] In a speech recognition-based language disorder assisted treatment system, the system generates a dialect-pathology separation report based on multi-round training data. The core process achieves accurate diagnosis and attribution through the confusion index (CI) algorithm: First, the system counts the number of dialect feature tags (1 time) and the number of pathology errors (2 times) in the current training unit. The patented algorithm CI = number of pathology errors / (number of dialect feature tags + number of pathology errors) is applied to calculate CI = 2 / (1 + 2) = 0.67 (the numerator and denominator have a preset non-zero protection protocol).

[0061] Step 4: Targeted Training

[0062] Based on a real-time calculated confusion index (CI) of 0.67 (moderate risk of confusion), the treatment server activates a Cantonese-adaptive enhancement program and performs precise intervention through a multimodal interactive system: AR glasses immediately overlay a dynamic tongue-side elevation holographic animation (60 frames / second) in the patient's field of vision, highlighting the target area (60% of the standard tongue-palatate contact area) with a green highlight area (the animation loops at a frequency of 5Hz to demonstrate the elevation trajectory), and simultaneously renders a semi-transparent red mask at the actual contact position of the tongue (transparency = 1 - measured area / standard area × 100%, in this case, 30% area corresponds to 70% transparency); when the tongue pressure sensor grid detects a real-time contact area < 45% critical value... When the contact area reaches 75% (the lower limit of the health standard), the haptic wristband immediately outputs a high-frequency vibration of 200±5Hz (intensity algorithm: I = basic intensity 12mN·s × (1 + (60% - S) / 15%), where S = 30%, substituting into I = 12 × (1 + 2) = 36mN·s). The vibration mode switches to a square wave sequence of 50ms pulse / 50ms interval, and at the same time, a three-level safety linkage control is activated. If the contact area is less than 35% within 3 seconds, the vibration frequency increases stepwise to 250Hz (intensity limited to 40mN·s). Conversely, if the contact area is greater than 40%, a positive feedback mechanism is triggered (AR animation switches to a green checkmark + vibration intensity drops sharply by 70%).

[0063] Step 5: Real-time dialect adaptation

[0064] This diagnostic model breaks through the limitations of traditional single standards, enabling Cantonese-speaking patients to pass the training with 50% contact area (saving 33% of rehabilitation time compared to the general standard). Clinical validation shows that the achievement rate of dialect mixed disorder has increased to 91.3% (Δ=+27.5%, p<0.01).

[0065] Implementation Process Three: Strengthening Learning Optimization Path (Children's Language Delay Rehabilitation)

[0066] Scenario: A 5-year-old child with delayed language development

[0067] Core Innovation: Q-learning algorithm dynamically adjusts the training path

[0068] Step 1: Capability Baseline Assessment

[0069] Treatment server initialization of Q-table:

[0070] State (phoneme complexity) Action (training module) Low Q value (monosyllabic) Module A (lip training) 0.0

[0071] Middle (Disyllabic) Module B (Tongue Training) 0.0

[0072] Step 2: Q-learning decision loop

[0073] In a speech recognition-based language disorder assistive therapy system, the system performs an algorithm update every 5 training sessions: based on the current state S_t = "low complexity" (the child's pronunciation ability is limited to / b / and / m / sounds), the system selects action A_t = "module A" through an ε-greedy strategy (exploration probability ε = 0.3) to drive the child's training target word "daddy" (focusing on the / b / sound). The immediate reward R is calculated after training: 1.5 points are obtained by multiplying the accuracy increment (40% → 55%, Δ = 0.15) by a coefficient of 10. After deducting the fatigue coefficient (training time 2 minutes / preset upper limit 5 minutes = 0.4), the final R = 1.5 - 0.4 × 5 = -0.5 (formula: R = accuracy increment × 10 - fatigue coefficient × 5). The Q-value is then updated: applying the temporal difference algorithm [Q(S_t,A_t)=(1-α)×original Q-value+α×(R+γ×maxQ(S_t+1))], substituting the learning rate α=0.2, discount factor γ=0.9, and new state maxQ=0 (because subsequent states were not explored), we calculate Q("low","module A")=(1-0.2)×0+0.2×[-0.5+0.9×0]=-0.1. This negative value indicates that module A's performance in low complexity is lower than expected (a decrease of 0.1 from the base Q-value of 0), triggering a policy adjustment—the ε policy will increase the probability of exploring module B / C to 42% in the next training session.

[0074] Step 3: Dynamic Path Optimization

[0075] Automatic activation after 10 training sessions: When the system detects that the accuracy rate of the / b / phoneme is consistently and stably >60% (average of 67% ± 3.2% over 10 training sessions, with a dispersion coefficient ≤ 5%), Q-table autonomously performs a state upgrade operation—migrating the original state S_t = "low complexity" (limited training on / b / and / m / sounds) to S_t+1 = "medium complexity" (adding training permissions for / p / and / t / sounds). Simultaneously, a specific action module C (palatopharyngeal closure biofeedback training) is dynamically generated, with its Q value initialized to 1.8 (calculation logic: historical average benefit 1.5 × progress coefficient 1.2), while simultaneously increasing the exploration rate ε to 0.45 to enhance the sampling probability of the new module. This upgrade mechanism ensures safe migration through three layers of verification: three consecutive qualifying tests (accuracy rates of 63% / 68% / 70%), fatigue coefficient maintained ≤1.0 (average of 0.8), and dispersion fluctuation < threshold, ultimately achieving an intelligent leap in the rehabilitation pathway (training dimension expansion +38%).

[0076] Step 4: Multimodal Cooperative Enhancement

[0077] In the pronunciation enhancement training for the target word "grape," the system precisely controls lip movements through multimodal feedback: AR glasses project a lip shape heatmap in real time (the target area is marked in dark red, corresponding to a lip closure pressure of ≥14kPa), simultaneously overlaid with a dynamic blast pressure dashboard (standard value 15cmH2O); when the pressure sensor detects that the measured pressure is only 12cmH2O (deviation -3cmH2O), the system executes dual-track intervention:

[0078] Visual guidance: Projection -3cmH2O deviation animation (red difference bar flashes at a frequency of 5Hz, arrows dynamically indicate the direction of lip closure tightening);

[0079] Tactile error correction: When the error rate falls into the medium-risk range [30%, 70%], the tactile wristband activates graded pulse vibration (intensity = 10mN·s × [1 + |Δp| / 3] = 20mN·s, mode: 120ms pulse / 180ms intermittent square wave, frequency 150Hz), and the vibration intensity increases linearly with the air pressure loss value;

[0080] Dynamic upgrade mechanism: If the air pressure does not rise to 13cmH2O within 3 seconds, the vibration intensity will be automatically increased by 30% (to 26mN·s) and the airflow simulation animation will be triggered (blue airflow jet with a flow rate of 6L / min) to form a closed-loop correction of biomechanical-neural feedback (error control ±0.2cmH2O).

[0081] Step 5: Cross-module jump rules

[0082] When the system detects that the accuracy improvement ΔP ≥ 0.2 over 5 consecutive training cycles (ΔP = average accuracy of this cycle / average accuracy of the previous cycle - 1) and the current module's Q value > preset threshold θ (θ = 1.5), it triggers the intelligent path jump protocol: pausing the original disyllabic training module (Q = 1.8) and activating the higher-order trisyllabic mixed training program (e.g., the target word "tractor" requires coordinated control of the / t / , / l / , and / j / consonant groups). The jump process performs triple verification:

[0083] 1. Competency verification: Pass the / p / phono blast pressure test (mean 18cmH2O > threshold 15) and tongue coordination test (transition delay <100ms).

[0084] 2. Dynamic parameter reset:

[0085] The initial Q value of the new module = the original Q × the upgrade coefficient (1.8 × 1.3 = 2.34).

[0086] The exploration rate ε is reset to 0.4 (33% higher than the original state).

[0087] 3. Cross-level protection mechanism: If the accuracy rate of three-syllable training is less than 45%, it will automatically revert to two-syllable reinforcement training (activation delay not exceeding 2 cycles).

[0088] Finally, the following points should be noted: First, in the description of this application, it should be noted that, unless otherwise specified and limited, the terms "installation", "connection", and "linkage" should be interpreted broadly, and can be mechanical or electrical connections, or internal connections between two components, or direct connections. "Up", "down", "left", "right", etc. are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may change.

[0089] Secondly: The accompanying drawings of the embodiments disclosed in this invention only involve the structures involved in the embodiments disclosed in this invention. Other structures can refer to the general design. In the absence of conflict, the same embodiment and different embodiments of this invention can be combined with each other.

[0090] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A speech recognition-based language disorder auxiliary treatment system, characterized in that, The system includes a user terminal, an adaptive speech recognition engine, a treatment server, and a multimodal feedback device. The user terminal is equipped with a speech acquisition module and an interactive interface for real-time acquisition of patient speech and display of treatment feedback. The adaptive speech recognition engine is connected to the user terminal and includes a noise suppression unit, a dynamic threshold recognition unit, and a dialect adaptation module. The noise suppression unit performs environmental noise reduction processing on the input speech, the dynamic threshold recognition unit dynamically adjusts the phoneme recognition tolerance threshold according to the patient's language impairment type, and the dialect adaptation module loads a regional pronunciation feature library to suit patients with dialect impairments. The treatment server stores a personalized treatment strategy database and generates real-time multimodal feedback instructions based on the speech recognition results. The multimodal feedback device includes a visual speech organ movement simulation unit, a tactile vibration feedback unit, and an AR augmented reality correction guidance module, which respectively provide three-dimensional movement simulation of speech organs, tactile vibration feedback graded by syllable accuracy, and virtual pronunciation guidance superimposed in a real environment.

2. The speech recognition-based language disorder auxiliary treatment system according to claim 1, characterized in that, The dynamic threshold recognition unit automatically classifies articulation disorders, aphasia, or fluency disorders by analyzing the patient's historical data. During treatment, it dynamically adjusts the phoneme similarity judgment threshold based on the degree of improvement in pronunciation. When three consecutive improvements in pronunciation stability are detected to exceed 15%, the error tolerance threshold is automatically relaxed by 5% to 10%. When a specific phoneme error rate is detected to be consistently higher than 50%, the corresponding phoneme threshold is tightened by 8% to 12%.

3. The speech recognition-based language disorder auxiliary treatment system according to claim 1, characterized in that, The dialect adaptation module establishes a baseline pronunciation model by collecting speech samples from healthy people in the target dialect area, extracts a set of dialect-specific phoneme replacement rules, and automatically marks the patient's pronunciation as acceptable when the patient's pronunciation deviates from the standard phonemes but conforms to the preset dialect replacement rules during the speech recognition process. At the same time, it generates a separation assessment report of dialect features and pathological pronunciation.

4. The speech recognition-based language disorder auxiliary treatment system according to claim 1, characterized in that, The visualized vocal organ motion simulation unit synchronously generates a three-dimensional dynamic tongue trajectory projection, a dynamic heat map of lip shape changes, and a visualized model of vocal cord vibration frequency. The tongue trajectory projection displays the distance parameters between the tongue tip and back and the hard palate in real time. The lip shape heat map uses color depth to indicate the contraction intensity of the orbicularis oris muscle. The vocal cord vibration model displays the fundamental frequency stability through waveform diagrams.

5. The speech recognition-based language disorder auxiliary treatment system according to claim 1, characterized in that, The treatment server has a built-in reinforcement learning engine. By recording the recognition results and feedback data of each pronunciation training session, it uses the Q-learning algorithm to optimize the difficulty sequence of practice items in the treatment strategy database and dynamically generates a personalized treatment path that includes the syllable complexity increment rate and training time allocation factor. When the accuracy of a specific phoneme improves by more than 20% in five consecutive training sessions, a high-difficulty compound training module is automatically inserted.

6. The speech recognition-based language disorder auxiliary treatment system according to claim 1, characterized in that, The tactile vibration feedback unit switches between three working modes based on the real-time syllable recognition accuracy: when the accuracy is greater than or equal to 90%, it outputs a continuous and stable vibration at a frequency of 100 Hz; when the accuracy is between 70% and 89%, it outputs a pulse vibration at an interval of 0.5 seconds; and when the accuracy is less than 70%, it outputs a high-frequency warning vibration at a frequency of 200 Hz. The vibration intensity automatically increases by 30% to 50% depending on the position of the incorrect syllable.

7. The speech recognition-based language disorder auxiliary treatment system according to claim 1, characterized in that, It also includes a remote diagnosis and treatment interface module, which converts speech recognition results into standardized articulation disorder assessment scale data, automatically generates diagnosis and treatment reports containing heatmaps of the spatial distribution of error syllables and temporal spectra of error frequencies, and supports multilingual therapists to collaboratively annotate abnormal pronunciation segments through a speech annotation system and generate annotation confidence scores.

8. A language disorder treatment method based on the system according to any one of claims 1 to 7, characterized in that, include: A patient-specific speech recognition model is established through a dynamic threshold recognition unit, and phoneme error tolerance parameters are optimized in real time. An AR augmented reality correction guidance module is used to overlay the standard position of virtual speech organs onto the patient's real-time image to dynamically indicate the deviation angle. The intensity of lip and tongue muscle exertion is adjusted in real time according to the vibration pattern changes of the tactile vibration feedback unit. When the accuracy rate improves by more than 15% in three consecutive training sessions, a high-difficulty training module is automatically activated. When the error rate of a specific phoneme is consistently higher than 50%, a decomposed muscle training animation is triggered.