Personalized language rehabilitation system and method based on neuroplasticity and voice recognition
By extracting multi-dimensional pronunciation defect parameters and dynamically generating personalized rehabilitation tasks, the problem of existing systems being unable to quantify pronunciation deviations has been solved, improving the efficiency and adaptability of language rehabilitation training and activating neuroplasticity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANLI PEOPLES HOSPITAL
- Filing Date
- 2026-01-12
- Publication Date
- 2026-04-21
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing speech rehabilitation training systems cannot quantitatively analyze the acoustic details of the pronunciation process, making it difficult to identify pronunciation deviations at the phoneme level. This results in training tasks not being able to be dynamically adjusted, leading to user frustration and low training efficiency.
By extracting multi-dimensional pronunciation defect parameters such as phoneme error rate, fundamental frequency trajectory deviation, and formant Euclidean offset, and combining them with users' historical data to assess language ability, personalized rehabilitation tasks are dynamically generated, and multi-channel feedback is provided to adjust the training difficulty, forming a closed-loop mechanism.
It enables dynamic adjustment of training difficulty based on the user's actual pronunciation performance, maintaining the repetitive and targeted stimulation required for neural plasticity activation, and improving the efficiency and adaptability of rehabilitation training.
Smart Images

Figure CN121905151A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interdisciplinary technology of artificial intelligence and speech rehabilitation, and more specifically, to a personalized language rehabilitation system and method based on neuroplasticity and speech recognition. Background Technology
[0002] Existing language rehabilitation training systems mostly rely on manual assessment or simple speech recognition based on keyword matching. They can only determine whether the user pronounces the correct words, but cannot quantitatively analyze the acoustic details of the pronunciation process. Even if some systems introduce automatic speech recognition (ASR) technology, their output is limited to text transcription results, lacking the ability to extract and compare key pronunciation features such as fundamental frequency, formants, and speech rate. Therefore, the system struggles to identify pronunciation deviations at the phoneme level (such as vowel distortion, abnormal intonation, or non-semantic pauses), resulting in training tasks that cannot be dynamically adjusted according to the user's actual pronunciation deficiencies. Long-term use of content with fixed difficulty not only fails to effectively stimulate the "moderate challenge" required for neuroplasticity but also easily leads to user frustration or low training efficiency.
[0003] Therefore, how to objectively quantify the user's pronunciation defects based on multi-dimensional acoustic features and use this to drive the generation of personalized and adaptive rehabilitation tasks has become a core technical problem that urgently needs to be solved. A personalized language rehabilitation system and method based on neuroplasticity and speech recognition is proposed. Summary of the Invention
[0004] To overcome the aforementioned shortcomings of existing technologies, this invention extracts multi-dimensional pronunciation defect parameters, including phoneme error rate, fundamental frequency trajectory deviation, and formant Euclidean offset, through a speech recognition and defect analysis module. These parameters, along with the user's historical data, are input into a language ability assessment module to calculate a comprehensive score reflecting the current rehabilitation stage. This score is further used by a personalized rehabilitation task generation module to match training templates and dynamically adjust execution parameters such as target phoneme density and background noise intensity. Simultaneously, a multi-channel feedback output module provides voice demonstrations and visual cues based on the defect type, while a rehabilitation database module continuously records data throughout the process and supports progress trend analysis, thus forming a closed-loop mechanism of "defect quantification—ability assessment—task generation—multimodal feedback—data storage—strategy optimization." Because the training difficulty is dynamically adjusted based on the user's actual pronunciation performance, and the feedback content is directly related to defect characteristics, each training session is conducted within the user's "zone of proximal development" of language ability, effectively maintaining the repetitive, targeted, and progressive stimulation required for neuroplasticity activation, thereby improving the efficiency and adaptability of rehabilitation training.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a personalized language rehabilitation system based on neuroplasticity and speech recognition, comprising:
[0006] Voice input module: used to collect the user's spoken voice signal;
[0007] Speech recognition and defect analysis module: connected to the speech input module, used to recognize the speech signal and extract pronunciation defect parameters;
[0008] User language ability assessment module: connected to the speech recognition and defect analysis module, used to assess the user's language ability level based on the user's historical speech data and current pronunciation defect parameters;
[0009] Personalized rehabilitation task generation module: connected to the user language ability assessment module, used to dynamically generate personalized language rehabilitation training tasks based on the user's language ability level;
[0010] Multi-channel feedback output module: connected to the personalized rehabilitation task generation module, used to output multi-channel feedback information to the user, including voice demonstrations, visual cues, or text guidance;
[0011] The rehabilitation database module stores the user's historical voice data, language ability assessment results, training records, and rehabilitation progress information, and supports data access from the user's language ability assessment module and the personalized rehabilitation task generation module.
[0012] In a preferred embodiment: the voice input module includes an audio acquisition unit and a voice preprocessing unit; the audio acquisition unit is a directional microphone, a microphone array, or a microphone integrated into a mobile terminal, used to acquire the user's spoken voice signal; the voice preprocessing unit is used to sequentially perform bandpass filtering, noise reduction processing based on spectral subtraction or deep neural networks, and voice activity detection on the voice signal to extract effective voice segments, and to divide the effective voice segments into frames according to a preset frame length and frame shift, outputting a standardized voice frame sequence for subsequent recognition.
[0013] In a preferred embodiment: the speech recognition and defect analysis module includes a speech transcription submodule and a pronunciation defect quantification submodule; the speech transcription submodule uses an end-to-end automatic speech recognition model based on Transformer or Conformer architecture to convert the speech signal into corresponding text and output phoneme-level time alignment results; the pronunciation defect quantification submodule compares the user's actual pronunciation with the target phonemes in the standard pronunciation library based on the phoneme-level time alignment results, and calculates multi-dimensional pronunciation defect parameters, including: phoneme error rate (PER), fundamental frequency trajectory deviation, Euclidean distance offset between the first formant (F1) and the second formant (F2), percentage deviation of speech rate relative to standard speech rate, and number of non-semantic pauses per unit time.
[0014] In a preferred embodiment: the user language ability assessment module includes a historical data retrieval unit, an ability index calculation unit, and a rehabilitation stage determination unit; the historical data retrieval unit extracts the pronunciation defect parameters and task completion records of the user's most recent N training sessions from the rehabilitation database module; the ability index calculation unit calculates a comprehensive language ability score based on the current pronunciation defect parameters and historical data, the comprehensive score being obtained by weighted summation of phoneme accuracy weight, sentence fluency weight, and task completion weight; the rehabilitation stage determination unit classifies the user's current language ability level into a primary, intermediate, or advanced rehabilitation stage according to the preset threshold range of the comprehensive score, and outputs an ability level identifier.
[0015] In a preferred embodiment: the personalized rehabilitation task generation module includes a task matching unit and a difficulty adjustment unit; the task matching unit retrieves matching training templates from a preset task knowledge base according to the language proficiency level, the training templates include one or more combinations of monosyllabic discrimination exercises, disyllabic word repetition, declarative sentence retelling, and question-and-answer dialogue simulation; the difficulty adjustment unit dynamically adjusts the execution parameters of the training templates according to the user's recent ability change trends, the execution parameters include: the occurrence density of target phonemes in sentences, background noise signal-to-noise ratio, the maximum allowable response delay time, and whether to enable real-time pronunciation waveform comparison prompts.
[0016] In a preferred embodiment: the multi-channel feedback output module includes a voice feedback submodule, a visual feedback submodule, and an interactive control submodule; the voice feedback submodule calls the text-to-speech synthesis engine to play the standard pronunciation audio corresponding to the user's incorrect pronunciation, and superimposes difference prompts at the phoneme level; the visual feedback submodule synchronously displays lip-shape animation, glottal opening and closing diagrams, or real-time spectrograms on the graphical user interface, and highlights the formant deviation areas between the user's pronunciation and the standard pronunciation; the interactive control submodule receives the user's confirmation or skip instruction for the feedback content, and records the feedback interaction log and sends it back to the rehabilitation database module.
[0017] In a preferred embodiment: the rehabilitation database module adopts a relational or time-series database structure, using the user's unique identifier as the primary key, and establishes a storage system containing the following data tables: the original speech table, storing the encrypted speech file path and sampling rate information; the recognition result table, storing the corresponding text, phoneme alignment timestamps, and confidence scores; the defect parameter table, storing multi-dimensional pronunciation defect parameters for each training session; the task execution table, storing the pushed task ID, difficulty parameters, and user completion status; and the ability assessment table, storing the comprehensive language ability score and rehabilitation stage identifier for each session. The rehabilitation database module also provides an API interface for the user's language ability assessment module and the personalized rehabilitation task generation module to query and update data in real time, and supports generating rehabilitation progress trend charts at a weekly / monthly granularity.
[0018] A personalized language rehabilitation method based on neuroplasticity and speech recognition, applying the personalized language rehabilitation system based on neuroplasticity and speech recognition as described in any one of claims 1-7, includes the following steps: S1, acquiring the user's pronunciation speech signal; S2, performing speech recognition on the speech signal and extracting pronunciation defect parameters; S3, assessing the user's language ability level based on the user's historical speech data and current pronunciation defect parameters; S4, dynamically generating personalized language rehabilitation training tasks based on the language ability level; S5, outputting multi-channel feedback information to the user including speech demonstrations, visual cues, or text guidance; S6, storing the user's speech data, assessment results, training records, and rehabilitation progress in a rehabilitation database; S7, continuously optimizing the difficulty and content of subsequent training tasks based on historical data in the database to achieve closed-loop adaptive rehabilitation.
[0019] The technical effects and advantages of this invention are as follows: This invention extracts multi-dimensional pronunciation defect parameters, including phoneme error rate, fundamental frequency trajectory deviation, and formant Euclidean offset, through a speech recognition and defect analysis module. These parameters, along with the user's historical data, are input into a language ability assessment module to calculate a comprehensive score reflecting the current rehabilitation stage. This score is further used by a personalized rehabilitation task generation module to match training templates and dynamically adjust execution parameters such as target phoneme density and background noise intensity. Simultaneously, a multi-channel feedback output module provides voice demonstrations and visual cues based on the defect type, while a rehabilitation database module continuously records data throughout the process and supports progress trend analysis, thus forming a closed-loop mechanism of "defect quantification—ability assessment—task generation—multimodal feedback—data storage—strategy optimization." Because the training difficulty is dynamically adjusted based on the user's actual pronunciation performance, and the feedback content is directly related to the defect characteristics, each training session is conducted within the user's "zone of proximal development" of language ability, effectively maintaining the repetitive, targeted, and progressive stimulation required for neuroplasticity activation, thereby improving the efficiency and adaptability of rehabilitation training. Attached Figure Description
[0020] Figure 1 This is a schematic diagram of the working system modules of the present invention. Figure 2 This is a schematic diagram of the working steps of the present invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] An exemplary embodiment will now be described more fully with reference to the accompanying drawings. A personalized language rehabilitation system based on neuroplasticity and speech recognition includes: a speech input module for acquiring a user's pronunciation speech signal; a speech recognition and defect analysis module connected to the speech input module for recognizing the speech signal and extracting pronunciation defect parameters; a user language ability assessment module connected to the speech recognition and defect analysis module for assessing the user's language ability level based on the user's historical speech data and current pronunciation defect parameters; a personalized rehabilitation task generation module connected to the user language ability assessment module for dynamically generating personalized language rehabilitation training tasks based on the language ability level; a multi-channel feedback output module connected to the personalized rehabilitation task generation module for outputting multi-channel feedback information to the user, including speech demonstrations, visual cues, or text guidance; and a rehabilitation database module for storing the user's historical speech data, language ability assessment results, training records, and rehabilitation progress information, and supporting data retrieval by the user language ability assessment module and the personalized rehabilitation task generation module.
[0023] The raw speech signal acquired by the voice input module is recorded as a continuous-time signal. Sampling rate The signal first passes through a digital bandpass filter with a cutoff frequency of [80Hz, 8kHz] to remove the DC component and high-frequency noise. Subsequently, the system employs an improved spectral subtraction method for noise reduction; let the... The short-time Fourier transform of the frame speech is The noise power spectrum is estimated as follows: The enhanced spectrum is then:
[0024]
[0025] in, For over-subtraction factor, This is the lower limit protection factor. To compensate for residual noise gain, this design effectively suppresses background interference while preserving speech intelligibility;
[0026] Voice Activity Detection (VAD) employs a joint energy-zero-crossing rate criterion.
[0027] Definition of the first Frame Short Time Energy With zero crossing rate for:
[0028]
[0029] System dynamically updates threshold , ,in The mean and standard deviation of the silent segment are statistically analyzed.
[0030] Only when and When a valid audio frame is selected, invalid segments are prevented from being included in the recognition process.
[0031] The speech recognition and defect analysis module includes a speech transcription submodule and a pronunciation defect quantification submodule. The speech transcription submodule uses an end-to-end automatic speech recognition model based on the Transformer or Conformer architecture to convert the speech signal into corresponding text and output phoneme-level time alignment results. The pronunciation defect quantification submodule compares the user's actual pronunciation with the target phonemes in the standard pronunciation library based on the phoneme-level time alignment results to calculate multi-dimensional pronunciation defect parameters. The multi-dimensional pronunciation defect parameters include: phoneme error rate (PER), fundamental frequency trajectory deviation, Euclidean distance offset between the first formant (F1) and the second formant (F2), percentage deviation of speech rate relative to the standard speech rate, and number of non-semantic pauses per unit time.
[0032] The aforementioned end-to-end ASR model based on the Conformer architecture integrates convolution and self-attention mechanisms in its encoder, outputting a phoneme sequence and temporal alignment results; given an input speech frame sequence... The model outputs a sequence of phoneme labels. and its start and end timestamps The pronunciation defect quantification submodule constructs a multi-dimensional acoustic bias vector for each target phoneme. The system retrieves its reference features from the standard pronunciation database. and the user's actual pronunciation characteristics Calculate the following parameters:
[0033] Phoneme Error Rate (PER):
[0034]
[0035] in These represent the number of errors for deletion, insertion, and replacement, respectively. The total number of standard phonemes reflects the accuracy of the phoneme segment;
[0036] Fundamental frequency trajectory deviation ( ):
[0037]
[0038] It measures the stability of pitch control and is suitable for assessing intonation abnormalities.
[0039] Euclidean shift of resonance peaks ( ):
[0040]
[0041] It reflects deviations in vocal tract configuration and is used to determine the articulation position of vowels / consonants.
[0042] speech rate deviation rate ( ):
[0043]
[0044] Identify abnormal speech rhythm;
[0045] Non-semantic pause frequency: The number of silent gaps lasting more than 300ms per unit time is counted and used to assess speech fluency; the above parameters together constitute a quantitative profile of pronunciation defects, providing a refined basis for subsequent personalized intervention.
[0046] The user language ability assessment module includes a historical data retrieval unit, an ability indicator calculation unit, and a rehabilitation stage determination unit.
[0047] The historical data retrieval unit extracts the pronunciation defect parameters and task completion records of the user's most recent N training sessions from the rehabilitation database module; the ability index calculation unit calculates the comprehensive language ability score based on the current pronunciation defect parameters and historical data, and the comprehensive score is obtained by weighted summation of phoneme accuracy weight, sentence fluency weight and task completion weight.
[0048] The rehabilitation stage determination unit classifies the user's current language ability level into a primary, intermediate, or advanced rehabilitation stage based on the preset threshold range of the comprehensive score, and outputs an ability level identifier.
[0049] The user language ability assessment module dynamically assesses the user's rehabilitation level by integrating current performance with historical trends;
[0050] The historical data retrieval unit extracts the user's most recent data from the rehabilitation database. The pronunciation defect parameter vector of the second training And task completion records. The capability indicator calculation unit first calculates the current phoneme accuracy rate. PER / 100), Sentence fluency (in (Maximum pause frequency preset), task completion rate (Based on a comprehensive assessment of task completion rate and response timeout), then, calculate the overall language proficiency score:
[0051]
[0052] Among them, weight satisfy The default value is The treatment can also be adjusted by the rehabilitation therapist according to the patient's type.
[0053] The rehabilitation stage assessment unit will Compared with the preset threshold:
[0054] like It is classified as the initial stage;
[0055] like This is the intermediate stage;
[0056] like This represents the advanced stage and outputs the corresponding capability level identifier.
[0057] The personalized rehabilitation task generation module includes a task matching unit and a difficulty adjustment unit; the task matching unit retrieves matching training templates from a preset task knowledge base according to the language proficiency level, and the training templates include one or more combinations of monosyllabic discrimination exercises, disyllabic word repetition, declarative sentence retelling, and question-and-answer dialogue simulation;
[0058] The difficulty adjustment unit dynamically adjusts the execution parameters of the training template based on the user's recent ability change trend. The execution parameters include: the occurrence density of the target phoneme in the sentence, the background noise signal-to-noise ratio, the maximum allowed response delay time, and whether to enable real-time pronunciation waveform comparison prompts.
[0059] The personalized rehabilitation task generation module matches training templates from a preset task knowledge base based on the user's current rehabilitation stage.
[0060] The task knowledge base stores the following template types according to difficulty level:
[0061] Beginner level: Distinguishing between single syllables (e.g., / s / vs / sh / ),
[0062] Intermediate: Reading aloud disyllabic words and repeating short declarative sentences.
[0063] Advanced: Question-and-answer dialogue simulation, instruction execution with background noise;
[0064] The difficulty adjustment unit further dynamically adjusts the execution parameters based on the user's recent ability trends. Let's assume the recent... The overall score sequence for this training session is as follows: The slope of the system's fitted linear regression:
[0065]
[0066] like This is judged as rapid progress, with an increase in difficulty.
[0067] like If so, then reduce the difficulty;
[0068] The specific parameter adjustment rules are as follows:
[0069] Target phoneme density: benchmark ;
[0070] Background noise signal-to-noise ratio: ;
[0071] Maximum response latency: Second( s);
[0072] Real-time waveform comparison prompt: When Forced to enable at any time.
[0073] The multi-channel feedback output module includes a voice feedback submodule, a visual feedback submodule, and an interactive control submodule;
[0074] The voice feedback submodule calls the text-to-speech synthesis engine to play the standard pronunciation audio corresponding to the user's incorrect pronunciation, and superimposes the difference prompt sound at the phoneme level;
[0075] The visual feedback submodule synchronously displays lip-shape animation, glottal opening and closing diagram or real-time spectrogram on the graphical user interface, and highlights the formant deviation area between the user's pronunciation and the standard pronunciation with a bright color; the interactive control submodule receives the user's confirmation or skip instruction for the feedback content, and records the feedback interaction log and sends it back to the rehabilitation database module.
[0076] The multi-channel feedback output module provides correction guidance simultaneously through auditory, visual, and interactive channels; the voice feedback submodule calls a high-quality text-to-speech (TTS) engine to play standard pronunciation audio and overlays difference prompts at incorrect phonemes.
[0077] The audio frequency offset is proportional to the fundamental frequency deviation.
[0078]
[0079] For example, if the user's fundamental frequency is 20Hz lower, the prompt tone will be raised by 1kHz to form an auditory anchor point;
[0080] The visual feedback submodule is displayed in real time in the graphical user interface:
[0081] Lip animation (phoneme-driven 3D mouthmodel);
[0082] A diagram illustrating the opening and closing of the glottis (reflecting the phonation state);
[0083] Real-time spectrogram or vowel formant distribution diagram;
[0084] Specifically, the system calculates the user resonance peak value ( , ) and standard point ( , The vector difference is used to indicate the direction of correction, which is highlighted with a red arrow to intuitively guide tongue position adjustment.
[0085] The rehabilitation database module adopts a relational or time-series database structure, using the user's unique identifier as the primary key, and establishes a storage system containing the following data tables:
[0086] The original speech table stores the encrypted speech file path and sampling rate information; the recognition result table stores the corresponding text, phoneme alignment timestamp, and confidence score.
[0087] The defect parameter table stores multi-dimensional pronunciation defect parameters for each training session;
[0088] The task execution table stores the pushed task ID, difficulty parameters, and user completion status.
[0089] Ability assessment form, storing the comprehensive scores of language ability in previous tests and the indicators of rehabilitation stages;
[0090] The rehabilitation database module also provides an API interface for the user language ability assessment module and the personalized rehabilitation task generation module to query and update data in real time, and supports the generation of rehabilitation progress trend charts at a weekly / monthly granularity.
[0091] The rehabilitation database module is built using a relational database (such as MySQL or PostgreSQL) or a time-series database (such as InfluxDB). It uses the user's unique identifier (UserID) as the primary key to establish a structured data table system, which is used to persistently store multi-source heterogeneous data generated throughout the training process and to provide efficient data services for other functional modules.
[0092] Specifically, the rehabilitation database module includes the following core data tables:
[0093] Raw Audio Table:
[0094] The encrypted audio file path, sampling rate (e.g., 16kHz), number of channels, recording timestamp, and associated task ID are stored. The audio file is encrypted using AES-256 to ensure user privacy and security.
[0095] ASRResultTable:
[0096] Store the recognized text, phoneme sequence, phoneme-level time alignment results (start and end timestamps), recognition confidence (0-1), and ASR model version used for each training session;
[0097] Defect Param Table:
[0098] Store the multi-dimensional pronunciation defect parameters obtained from each training calculation, including: PER value, ΔF0, ΔF12, δv, number of non-semantic pauses, etc., and each parameter is associated with a corresponding phoneme or sentence ID;
[0099] Task Log Table:
[0100] Record the ID of the pushed task, the task type (e.g., "monosyllabic discrimination"), the difficulty parameters (e.g., target phoneme density ρ, background noise SNR), the user response time, whether it has timed out, and the completion status (success / failure / abandonment).
[0101] Competency Assessment Table
[0102] Store the comprehensive language ability score S, rehabilitation stage indicator (beginner / intermediate / advanced), assessment time, and the historical training number N on which the assessment was based;
[0103] All tables are linked by UserID and SessionID, supporting multi-dimensional queries by user, time period, task type, and other criteria.
[0104] The rehabilitation database module provides a standardized RESTful API interface for the user's language ability assessment module and personalized rehabilitation task generation module to call in real time; a typical interface includes: retrieving training records for the last 7 days.
[0105] Submit the results of this assessment and update your competency score;
[0106] Returns rehabilitation progress trend data at a weekly or monthly granularity;
[0107] To support closed-loop adaptive training, the system periodically calculates a rehabilitation progress index based on time-series data in the capability assessment table:
[0108]
[0109] in, This is the current overall score. for The score from days ago (usually) );
[0110] If continuous Training (e.g.) Satisfying RPI (For example If the condition is not met, it is determined that rehabilitation has stalled, and the system will automatically trigger an early warning mechanism to notify a remote rehabilitation therapist to intervene and conduct an assessment via push notification or email.
[0111] In addition, the database supports generating rehabilitation progress trend charts on a weekly and monthly basis, including: PER change curves, overall score trends, task difficulty evolution charts, etc., which make it easier for users and doctors to intuitively grasp the rehabilitation effect;
[0112] Through the above design, the rehabilitation database module not only achieves secure data storage and efficient access, but also becomes the core hub driving the "assessment-intervention-optimization" closed loop, ensuring that the entire system has the ability to continuously learn and evolve in a personalized manner.
[0113] A personalized language rehabilitation method based on neuroplasticity and speech recognition includes the following steps:
[0114] S1. Acquiring the user's spoken voice signal: The system uses a microphone or directional microphone array integrated into the mobile terminal to acquire the user's voice signal when reading a specified training text in real time. The sampling rate was set to 16kHz. To improve the signal-to-noise ratio, bandpass filtering (80Hz-8kHz) and spectral subtraction-based noise reduction were performed simultaneously during the acquisition process. Effective speech segments were extracted using voice activity detection (VAD), and a standardized speech frame sequence was output. For further processing.
[0115] S2. Perform speech recognition on the speech signal and extract pronunciation defect parameters:
[0116] The system employs an end-to-end Automatic Speech Recognition (ASR) model based on the Conformer architecture, converting speech frame sequences into text and outputting phoneme-level time-aligned results. ;
[0117] Subsequently, the pronunciation defect quantification submodule compares the user's actual pronunciation characteristics with the target phonemes in the standard pronunciation library to calculate multi-dimensional pronunciation defect parameters, including:
[0118] Phoneme Error Rate: PER ;
[0119] Fundamental frequency trajectory deviation: ;
[0120] Euclidean shift of resonance peaks: ;
[0121] Speech rate deviation rate: ;
[0122] The number of non-semantic pauses per unit time (count of silent segments > 300ms).
[0123] This multidimensional quantization mechanism breaks through the limitations of traditional methods that rely solely on text accuracy, enabling refined modeling of the physiological processes of pronunciation.
[0124] S3. Assess the user's language proficiency level based on the user's historical speech data and current pronunciation defect parameters:
[0125] The system retrieves the user's most recent [database / resources] from the rehabilitation database. Based on the defect parameters and task completion records of the previous training, the current phoneme accuracy is calculated. PER / 100, Sentence fluency Task completion rate The overall language proficiency score is obtained by weighted summation:
[0126]
[0127] The default weight , , ;
[0128] according to The interval (e.g.) 0.4 is the initial level. Intermediate level (For advanced settings), output the corresponding rehabilitation stage identifier.
[0129] S4. Based on the stated language proficiency level, dynamically generate personalized language rehabilitation training tasks: The system retrieves matching templates from a preset task knowledge base.
[0130] Beginner users are given monosyllabic discrimination exercises, intermediate users practice reading aloud disyllabic words or repeating short sentences, and advanced users participate in question-and-answer dialogue simulations.
[0131] Furthermore, the difficulty adjustment unit dynamically adjusts the execution parameters based on the user's recent ability change trend (quantified by the linear regression slope α), including:
[0132] Target phoneme density The background noise signal-to-noise ratio (SNR), maximum response delay, and whether real-time vocal waveform comparison cues are enabled are all considered to ensure that training remains within the "zone of proximal development." .
[0133] S5. Output multi-channel feedback information to the user, including voice demonstrations, visual cues, or text guidance: The system provides this synchronously through the multi-channel feedback output module.
[0134] Voice feedback: Plays standard pronunciation audio and overlays frequency offset prompts at incorrect phonemes. ;
[0135] Visual feedback: Display lip-shape animation, glottal opening and closing diagram or real-time spectrogram in the graphical interface, and mark the formant correction direction with a highlighted arrow;
[0136] Interactive control: Receives user "confirm" or "skip" commands and records feedback interaction logs; this multimodal feedback mechanism collaboratively activates the auditory, visual and motor cortexes, accelerating the reconstruction of sensory-motor mapping.
[0137] S6. Store the user's voice data, assessment results, training records, and rehabilitation progress in the rehabilitation database: All original voice files (encrypted storage), recognition results, defect parameters, task logs, ability scores, and other data are written to a relational or time-series database using the user ID as the primary key, forming a structured rehabilitation profile. This supports generating progress trend charts at a weekly / monthly granular level.
[0138] S7. Based on historical data in the database, continuously optimize the difficulty and content of subsequent training tasks to achieve closed-loop adaptive rehabilitation: The system periodically calculates the Rehabilitation Progress Index (RPI). If the RPI falls below the threshold multiple times consecutively... If the condition is not met, it is determined that rehabilitation has stalled, an early warning is automatically triggered, and the training strategy is adjusted. At the same time, historical successful case data is used to optimize the task recommendation algorithm, enabling the system to have continuous learning and personalized presentation capabilities.
[0139] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, and other arbitrary combinations. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions and computer programs. When the computer instructions and computer programs are loaded and executed on a computer, the processes and functions described in the embodiments of this application are generated, in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, and data center to another website, computer, server, and data center via wired (e.g., infrared, wireless, microwave) means. The computer-readable storage medium can be any available medium that a computer can access or a server or data center data storage device containing one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0140] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0141] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0142] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0143] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0144] If the aforementioned functions are implemented as software functional units and sold and used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device) to execute all and part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and various media capable of storing program code.
[0145] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations and substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A personalized language rehabilitation system based on neuroplasticity and speech recognition, characterized in that, This includes: Voice input module: used to collect the user's spoken voice signal; Speech recognition and defect analysis module: connected to the speech input module, used to recognize the speech signal and extract pronunciation defect parameters; User language ability assessment module: connected to the speech recognition and defect analysis module, used to assess the user's language ability level based on the user's historical speech data and current pronunciation defect parameters; Personalized rehabilitation task generation module: connected to the user language ability assessment module, used to dynamically generate personalized language rehabilitation training tasks based on the user's language ability level; Multi-channel feedback output module: connected to the personalized rehabilitation task generation module, used to output multi-channel feedback information to the user, including voice demonstrations, visual cues, or text guidance; The rehabilitation database module stores the user's historical voice data, language ability assessment results, training records, and rehabilitation progress information, and supports data access from the user's language ability assessment module and the personalized rehabilitation task generation module.
2. The personalized language rehabilitation system based on neuroplasticity and speech recognition according to claim 1, characterized in that: The voice input module includes an audio acquisition unit and a voice preprocessing unit. The audio acquisition unit is a directional microphone, a microphone array, or a microphone integrated into a mobile terminal, used to acquire the user's spoken voice signal. The voice preprocessing unit is used to sequentially perform bandpass filtering, noise reduction processing based on spectral subtraction or deep neural networks, and voice activity detection on the voice signal to extract effective voice segments. The effective voice segments are then divided into frames according to a preset frame length and frame shift, and a standardized voice frame sequence is output for subsequent recognition.
3. The personalized language rehabilitation system based on neuroplasticity and speech recognition according to claim 1, characterized in that: The speech recognition and defect analysis module includes a speech transcription submodule and a pronunciation defect quantification submodule. The speech transcription submodule uses an end-to-end automatic speech recognition model based on the Transformer or Conformer architecture to convert the speech signal into corresponding text and output phoneme-level time alignment results. The pronunciation defect quantification submodule compares the user's actual pronunciation with the target phonemes in the standard pronunciation library based on the phoneme-level time alignment results to calculate multi-dimensional pronunciation defect parameters. The multi-dimensional pronunciation defect parameters include: phoneme error rate (PER), fundamental frequency trajectory deviation, Euclidean distance offset between the first formant (F1) and the second formant (F2), percentage deviation of speech rate relative to the standard speech rate, and the number of non-semantic pauses per unit time.
4. The personalized language rehabilitation system based on neuroplasticity and speech recognition according to claim 1, characterized in that: The user language ability assessment module includes a historical data retrieval unit, an ability indicator calculation unit, and a rehabilitation stage determination unit. The historical data retrieval unit extracts the pronunciation defect parameters and task completion records of the user's most recent N training sessions from the rehabilitation database module; the ability index calculation unit calculates the comprehensive language ability score based on the current pronunciation defect parameters and historical data, and the comprehensive score is obtained by weighted summation of phoneme accuracy weight, sentence fluency weight and task completion weight. The rehabilitation stage determination unit classifies the user's current language ability level into primary, intermediate, or advanced rehabilitation stages based on the preset threshold range of the comprehensive score, and outputs the ability level identifier.
5. A personalized language rehabilitation system based on neuroplasticity and speech recognition according to claim 1, characterized in that: The personalized rehabilitation task generation module includes a task matching unit and a difficulty adjustment unit; the task matching unit retrieves matching training templates from a preset task knowledge base according to the language proficiency level, and the training templates include one or more combinations of monosyllabic discrimination exercises, disyllabic word repetition, declarative sentence retelling, and question-and-answer dialogue simulation; The difficulty adjustment unit dynamically adjusts the execution parameters of the training template based on the user's recent ability change trend. The execution parameters include: the occurrence density of the target phoneme in the sentence, the background noise signal-to-noise ratio, the maximum allowed response delay time, and whether to enable real-time pronunciation waveform comparison prompts.
6. The personalized language rehabilitation system based on neuroplasticity and speech recognition according to claim 1, characterized in that: The multi-channel feedback output module includes a voice feedback submodule, a visual feedback submodule, and an interactive control submodule; The voice feedback submodule calls the text-to-speech synthesis engine to play the standard pronunciation audio corresponding to the user's incorrect pronunciation, and superimposes the difference prompt sound at the phoneme level; The visual feedback submodule synchronously displays lip-shape animation, glottal opening and closing diagram or real-time spectrogram on the graphical user interface, and highlights the formant deviation area between the user's pronunciation and the standard pronunciation with a bright color; the interactive control submodule receives the user's confirmation or skip instruction for the feedback content, and records the feedback interaction log and sends it back to the rehabilitation database module.
7. A personalized language rehabilitation system based on neuroplasticity and speech recognition according to claim 1, characterized in that: The rehabilitation database module adopts a relational or time-series database structure, using the user's unique identifier as the primary key, and establishes a storage system containing the following data tables: The original audio table stores the encrypted audio file path and sampling rate information; The recognition results table stores the corresponding text, phoneme alignment timestamp, and confidence level. The defect parameter table stores multi-dimensional pronunciation defect parameters for each training session; The task execution table stores the pushed task ID, difficulty parameters, and user completion status. Ability assessment form, storing the comprehensive scores of language ability in previous tests and the indicators of rehabilitation stages; The rehabilitation database module also provides an API interface for the user language ability assessment module and the personalized rehabilitation task generation module to query and update data in real time, and supports the generation of rehabilitation progress trend charts at a weekly / monthly granularity.
8. A personalized language rehabilitation method based on neuroplasticity and speech recognition, employing the personalized language rehabilitation system based on neuroplasticity and speech recognition as described in any one of claims 1-7, characterized in that, Includes the following steps: S1. Collect the user's spoken voice signal; S2. Perform speech recognition on the voice signal and extract pronunciation defect parameters; S3. Assess the user's language proficiency level based on the user's historical speech data and current pronunciation defect parameters; S4. Dynamically generate personalized language rehabilitation training tasks based on the language proficiency level; S5. Output multi-channel feedback information to the user, including voice demonstrations, visual cues, or text instructions; S6. Store the user's voice data, assessment results, training records, and rehabilitation progress in the rehabilitation database. S7. Based on historical data in the database, continuously optimize the difficulty and content of subsequent training tasks to achieve closed-loop adaptive rehabilitation.