Auditory speech rehabilitation training method fusing semantic noise control and difficulty self-adaption

By using reinforcement learning models and semantic noise control, the training difficulty and background noise are dynamically adjusted, which solves the problems of modular fragmentation and insufficient feedback in existing language rehabilitation systems and improves users' auditory and semantic abilities in noisy environments.

CN121506175APending Publication Date: 2026-02-10HANGZHOU HUIER HEARING INSTR & TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511487513.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing language rehabilitation systems suffer from fragmented modular training content, primitive logic for adjusting training difficulty, lack of intelligent noise processing and semantic coordination capabilities, and a lack of multimodal interactive support for training feedback, resulting in limited user self-awareness and self-correction abilities.

Method used

By acquiring user training data, the training difficulty is adjusted using a reinforcement learning model. Combined with semantic noise control, the task difficulty and background noise are dynamically adjusted to achieve multimodal feedback, including pronunciation scoring graphs, speech comparison playback, and semantic error localization prompts.

Benefits of technology

It achieves personalized and precise matching of training content, improves users' auditory separation and semantic reconstruction capabilities in noisy environments, and enhances users' self-awareness and self-correction capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121506175A_ABST
    Figure CN121506175A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an auditory speech rehabilitation training method fusing semantic noise control and difficulty self-adaption, and relates to the technical field of auditory speech rehabilitation and artificial intelligence auxiliary language training. The auditory speech rehabilitation training method comprises the steps of obtaining a plurality of subdivision indexes of a user in recent rounds of training in auditory speech rehabilitation training, generating a comprehensive training score of the same round of training, and determining a training state of the user; the training state and the comprehensive training score of the current round of training serve as state variables, a reinforcement learning model is started, an award signal is generated according to the subdivision indexes and the comprehensive training score of the user in the current round of training, the task difficulty of the next round of training is decided according to the award signal, and the task difficulty comprises the number of keywords and the noise signal-to-noise ratio; corpus content of the next round of training is generated according to the number of the keywords, and environment noise matched with the training scene is added at the positions of the keywords according to the noise signal-to-noise ratio. According to the embodiment, the training difficulty and the noise adding position can be adaptively adjusted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of auditory-speech rehabilitation and artificial intelligence-assisted language training, and in particular to an auditory-speech rehabilitation training method that integrates semantic noise control and difficulty adaptation. Background Technology

[0002] In recent years, with the rapid development of speech recognition, speech assessment, artificial intelligence, and big data analysis technologies, the digitalization process in the field of hearing and speech rehabilitation has continued to accelerate. Multiple assistive training systems have made continuous progress in areas such as auditory and speech training content, user ability assessment, training path recommendation, difficulty adjustment mechanisms, and interactive feedback methods. Personalized rehabilitation training platforms based on artificial intelligence are gradually replacing traditional manual training methods, driving the evolution of speech rehabilitation systems from "static template-driven" to "adaptive intelligent feedback-driven."

[0003] While existing language rehabilitation systems have achieved preliminary structured designs in modules such as speech recognition, scoring mechanisms, training task recommendation, corpus playback, and interactive feedback, they still generally suffer from the following technical limitations:

[0004] The training content is modular and fragmented: articulation training, vocabulary recognition, and situational dialogue often exist independently, lacking a unified training path scheduling mechanism based on ability profiles, and cannot form a training closed loop from phonemes to semantics.

[0005] The logic for adjusting training difficulty is relatively primitive: it generally sets levels or tiered recommendations based on static scoring results, without establishing a dynamic behavior modeling mechanism and real-time state perception capability.

[0006] Noise processing lacks intelligence and semantic collaboration capabilities: In the existing system, noise is played in the background and cannot be intelligently matched and dynamically intervened in terms of noise interference.

[0007] Training feedback lacks multimodal interactive support: feedback methods mostly rely on numerical scoring and text prompts, lacking multi-channel auxiliary error correction methods such as audio comparison playback, pronunciation trajectory visualization, and semantic error location prompts, which limits the user's self-awareness and self-correction capabilities. Summary of the Invention

[0008] This invention provides an auditory-verbal rehabilitation training method that integrates semantic noise control and difficulty adaptation to solve at least one of the above-mentioned problems.

[0009] In a first aspect, embodiments of the present invention provide an auditory-verbal rehabilitation training method that integrates semantic noise control and difficulty adaptation, comprising:

[0010] The system obtains multiple sub-indicators from the user's most recent rounds of auditory-verbal rehabilitation training, including answer accuracy, voice score, number of repeated playbacks, and answer duration.

[0011] Based on the various sub-indicators of the same training round, a comprehensive training score for the same training round is generated;

[0012] The user's training status is determined based on the various sub-indicators and the overall training score of the multiple rounds of training.

[0013] Using the training state and the comprehensive training score of this round of training as state variables, the reinforcement learning model is started. The reinforcement learning model generates a reward signal based on the user's sub-indicators and comprehensive training score in this round of training, and determines the task difficulty of the next round of training based on the reward signal. The task difficulty includes the number of keywords and the noise signal-to-noise ratio.

[0014] The corpus content for the next round of training is generated based on the number of keywords, and environmental noise matching the training scenario is added at the keyword positions based on the noise signal-to-noise ratio.

[0015] In a second aspect, embodiments of the present invention provide an electronic device, the electronic device comprising:

[0016] One or more processors;

[0017] Memory, used to store one or more programs.

[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement the auditory-verbal rehabilitation training method that integrates semantic noise control and difficulty adaptation as described in any embodiment.

[0019] Thirdly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the auditory-verbal rehabilitation training method that integrates semantic noise control and difficulty adaptation as described in any embodiment.

[0020] In summary, this invention provides an auditory-verbal rehabilitation training method that integrates semantic noise control and difficulty adaptation. During training, the system collects data in real time, including user answer accuracy, speech score, response duration, and number of repetitions. Based on an ability trend model, it dynamically adjusts parameters such as speech rate, semantic complexity, and background interference intensity of the training content to achieve task difficulty adaptation. Simultaneously, it extracts keywords or core semantic information from the corpus using a semantic recognition algorithm and dynamically adjusts the background noise level before and after these keywords to simulate real attentional interference scenarios. Specifically, this embodiment has the following technical advantages:

[0021] First, considering the significant individual differences and fluctuating abilities among adult hearing-impaired users, this embodiment designs an intelligent adaptive training difficulty mechanism. This mechanism not only collects user performance data in each round of training, including answer accuracy, speech score, number of re-listening attempts, and answer duration, but also incorporates historical ability trends for multi-dimensional modeling. Through a reinforcement learning model, it adjusts the complexity of training tasks, interaction frequency, and training path direction in real time. This approach effectively avoids the "overestimation" or "underestimation" problems caused by static grading, achieving a dynamic and accurate match between user ability levels and task challenge.

[0022] In terms of environmental simulation, a semantically linked, real-scene noise intervention mechanism is proposed for the first time. Unlike existing technologies that only use background noise of fixed intensity as generalized interference, this embodiment introduces dynamic noise sources that are related to or adversarial to the semantic logic during key nodes of the user's semantic task (such as information extraction, dialogue response, and digit recognition), thereby improving the user's auditory separation and semantic reconstruction capabilities in noisy environments. This semantic-noise coupling design is more in line with the auditory challenges under real-world communication conditions and helps to promote the transition of rehabilitation training for hearing-impaired users from "static indoor environments" to "realistic environments." Attached Figure Description

[0023] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0024] Figure 1 This is a schematic diagram of an auditory-verbal rehabilitation training system that integrates semantic noise control and difficulty adaptation, provided in an embodiment of the present invention.

[0025] Figure 2 This is a flowchart of an auditory-verbal rehabilitation training method that integrates semantic noise control and difficulty adaptation, provided by an embodiment of the present invention;

[0026] Figure 3 This is a scene example diagram of a contextual dialogue training module provided in an embodiment of the present invention;

[0027] Figure 4 This is an architecture diagram of user interface information provided in an embodiment of the present invention;

[0028] Figure 5 This is a flowchart of another auditory-verbal rehabilitation training method that integrates semantic noise control and difficulty adaptation provided in an embodiment of the present invention;

[0029] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0031] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0032] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0033] This embodiment provides an auditory-speech rehabilitation training method that integrates semantic noise control and difficulty adaptation. To illustrate this method, an auditory-speech rehabilitation training system that integrates semantic noise control and difficulty adaptation supporting the application of this method is introduced first. Figure 1 This is a schematic diagram of an auditory-speech rehabilitation training system that integrates semantic noise control and difficulty adaptation, provided by an embodiment of the present invention. Figure 1 As shown, the system includes a user information collection module, a user ability profile modeling module, a multi-level training module, a difficulty adaptation module, a semantic noise dynamic control module, a data recording and feedback module, and a user interaction interface module. The modules are interconnected through data flow or control flow to jointly complete the entire process of auditory and speech rehabilitation training for adults with hearing impairment.

[0034] The user information collection module gathers basic user information and initial speech ability test data via mobile or web platforms. This information includes: basic attributes (age, gender, type of hearing aid, rehabilitation stage, etc.); and initial hearing and speech ability indicators (including but not limited to: phoneme recognition accuracy, high-frequency word recognition accuracy, simple sentence listening and repetition ability, and information extraction efficiency in noisy environments). This data will serve as the input basis for subsequent training paths and parameter adjustments.

[0035] The user capability profile modeling module is used to build user capability vector models, generate comprehensive scores of user capabilities, and provide support for intelligent recommendation and dynamic adjustment of training content.

[0036] The multi-level training module divides the training process into three core modules:

[0037] Articulation training module: Based on the comparison between acoustic models (such as Mel frequency cepstral coefficients MFCC + dynamic time warping DTW) and standard pronunciation templates, users receive phoneme-based scoring feedback after recording audio.

[0038] Semantic word module: Users receive a task to select high-frequency words, with support for slow playback and key phoneme prompts;

[0039] Contextual Dialogue Module: Introduces multi-turn interactive tasks with clear semantic goals from everyday dialogue scenarios (such as banks, hospitals, restaurants, and transportation), and constructs a complete pragmatic chain by combining the "task-instruction-response" structure.

[0040] The difficulty adaptive adjustment module is used to access performance data during training in real time, including answer accuracy, pronunciation score, frequency of repeated listening and answer time, historical trend fluctuation curves, etc.; and uses Long Short-Term Memory (LSTM) modeling to judge the user's current state; and then dynamically adjusts the language complexity (sentence length, vocabulary density), interaction density (number of dialogue turns, speed of situation change), noise intensity, etc. of the training task according to the user's state to achieve dynamic adjustment of task granularity.

[0041] The semantic noise dynamic control module is used to dynamically insert "semantic adversarial noise" based on the semantic content of the dialogue and key nodes of the task. This module makes noise no longer a fixed background, but a "semantic perturbation factor" that interacts with the training semantics, thereby enhancing the user's language parsing and reaction capabilities in real high-noise environments.

[0042] The data recording and feedback module is used to record data and provide multimodal feedback. The multimodal feedback is aimed at enhancing the user's ability and self-correction after training, and provides the following functions: pronunciation scoring map (initial consonant-final vowel-tone sub-items), comparison and playback of user recordings and standard audio, corresponding lip shape videos, airflow diagram animation demonstration, training process score curve and ability change trend graph, personalized prompts and suggestions, support for user review, annotation and personal note accumulation.

[0043] Based on the above systems, Figure 2 This is a flowchart of an auditory-speech rehabilitation training method integrating semantic noise control and difficulty adaptation, provided by an embodiment of the present invention. This method is applicable to auditory-speech rehabilitation for adult hearing-impaired users and can be executed by the various modules in the above system, or by other electronic devices. For ease of understanding and description, the following embodiments all use the modules in the above system as the execution subject; when the execution subject changes, the models in the system are replaced with the changed execution subject. Figure 2 As shown, the method specifically includes the following steps:

[0044] S110, User Information Collection.

[0045] This step involves the user information collection module collecting the user's personal information and initial capability data.

[0046] In one specific implementation, the user first completes registration and login via a mobile terminal, and the system records basic information such as the user's name, age, gender, hearing impairment type, hearing aid type, and rehabilitation stage.

[0047] Then, the user information collection module conducts an initial capability test on the user, including:

[0048] Single-sound recognition accuracy test: The system plays single sounds of initials, finals or tones in sequence. The user selects or imitates the pronunciation, and the system records the correct answer rate.

[0049] Word recognition accuracy test: The system plays commonly used words in daily life, and the user selects or imitates the pronunciation, recording the correct answer rate;

[0050] Simple sentence listening and retelling test: The system plays short sentences, and the user needs to retell them or select the correct meaning;

[0051] Information capture capability test in noisy environments: The system plays speech materials with different signal-to-noise ratios, and the user completes the listening identification or retelling.

[0052] S120, Competency Assessment and Document Generation.

[0053] This step is performed by the user capability profile module. In one specific implementation, after the test results are processed, the user capability profile module first generates an initial user capability profile based on the collected data, denoted as:

[0054]

[0055] in, Represents user capability data, Indicates the accuracy rate of single-tone recognition. Indicates word recognition accuracy. This indicates the accuracy rate of listening comprehension and retelling of simple sentences. This indicates the accuracy of information capture in noisy environments.

[0056] Then, the user capability data can be standardized using the following formula:

[0057]

[0058] in, This represents the raw score of a test item. This represents the historical average score for this project. This represents the historical standard deviation of the project score. This represents the standardized score of the project.

[0059] Finally, based on the standardized data, a comprehensive user ability score is generated. :

[0060]

[0061] in, , , and These represent the standardized accuracy rates for single-sound recognition, word recognition, simple sentence listening and repetition, and information capture in noisy environments, respectively. , , and These represent the corresponding weights. This score is used to guide subsequent personalized training path recommendations. (Weight parameters are listed below.) , , and It can be pre-set based on the experience of rehabilitation experts, or it can be obtained through statistical learning (such as regression analysis, factor analysis, or machine learning model training) on ​​historical user data.

[0062] S130, Personalized Training Path Recommendation.

[0063] This step can be performed by a multi-level trained model, based on the user's overall rating. The system recommends training paths and requirements (such as training objectives and scenario preferences). Through a rule base and adaptive recommendation engine, it generates dynamic task paths that include modules such as basic articulation training, high-frequency vocabulary training, semantic sentence training, and situational dialogue simulation training, and automatically distinguishes between beginner, intermediate, and advanced levels. Different levels correspond to different training content and complexity levels.

[0064] Optionally, the system first determines whether each of the user's ability indicators is below the corresponding threshold. If any indicator is below the threshold, the system prioritizes recommending the corresponding specialized training modules: articulation training module, vocabulary training module, and situational dialogue training module. If the user's ability is high, the system recommends an advanced training path.

[0065] S140. Perform multi-level training according to the training path.

[0066] This step can be performed by a multi-layered training model, training the corresponding modules according to the personalized path. Combined with... Figure 1 The training task execution module specifically includes an articulation training module, a vocabulary training module, and a situational dialogue training module.

[0067] In the articulation training module, an acoustic standard template is first loaded, including a three-dimensional acoustic model of initials, finals, and tones. Then, the user records the pronunciation of the target syllable, and the acoustic features of the user's recording are extracted and compared with the standard recording using MFCC (Multi-Functional Sound Comparison). The similarity score is then calculated. The calculation formula is:

[0068]

[0069] in, The MFCC feature representing the user's recording, MFCC characteristics representing standard recordings; This represents the similarity function, which can be DTW, cosine similarity, or other similarity methods. Optionally, in the scoring of initials, finals, and tones, if any score is lower than a preset threshold... This module will guide users into specialized articulation training.

[0070] In the vocabulary training module, standard word audio is first played, supporting both normal and slow playback speeds. Then, users record their pronunciations and submit them to the system, which then performs a multi-dimensional scoring based on the following formula:

[0071]

[0072] in, The word score is represented by PronAcc, which treats words as a continuous syllable sequence, extracts the overall MFCC features and compares them with a standard template, and uses DTW to eliminate duration differences, or uses cosine similarity and other measurement methods to quantify pronunciation accuracy. ToneAcc is evaluated based on the fundamental frequency (F0) contour features of the tone. Specific implementation methods include: concatenating MFCC and F0 features as input for comprehensive comparison, or extracting the F0 curve separately and quantifying word tone accuracy through DTW and similarity measurement. , This represents the weighting parameter.

[0073] Optionally, if a user's overall score is lower than the threshold for three consecutive times... This module will guide users into specialized articulation training to solidify their basic pronunciation skills.

[0074] In the contextual dialogue training module, preset training scenarios are first loaded, such as coffee shops, supermarkets, asking for directions while traveling, and hospital consultations. Background noise simulation is then introduced, including noise types such as: clinking of cups and plates, running water, conversations, traffic noise outside the window, and music playing in a coffee shop environment; broadcast announcements, crowd noise, and the sound of goods being stored in a supermarket environment; and announcements and conversations in a hospital environment. The noise signal-to-noise ratio (SNR) is controlled within a certain range. Then, the user interface module presents users with multi-turn dialogue tasks, including listening and speaking questions. The noise signal-to-noise ratio (SNR) refers to the ratio of the power of the speech signal to the power of the noise after noise is added; a higher SNR indicates a stronger speech signal and weaker noise, while a lower SNR indicates that the noise is relatively stronger than the speech signal.

[0075] Figure 3 This is a scenario example diagram for the contextual dialogue training module, demonstrating the dialogue training process and screen design in a coffee shop setting. It includes NPC dialogue, user answer options, background noise settings, and training prompts to simulate a real-world social communication environment. Table 1 schematically illustrates a difficulty-level corpus design for the coffee shop scenario. Easy difficulty involves word or phrase responses, medium difficulty involves complete short sentences, and high difficulty involves more complex sentence structures and polite expressions, all while incorporating noise interference, reflecting the multi-layered, progressive training strategy of this invention. Optionally, if the user answers incorrectly consecutively, the system reduces the difficulty or recommends lower-level specialized training.

[0076] Table 1

[0077]

[0078] S150, Judging Training Performance Status.

[0079] Combination Figure 1The data recording and feedback module records multiple performance metrics for each training session, including the accuracy of answering questions. Voice score S2= Number of times to repeat playback and response time The difficulty adaptation module will use multiple sub-indicators from recent training rounds to predict the user's current training status, specifically including several status types such as well-adapted, mildly fatigued, capable adjacency, and consistently efficient.

[0080] In one specific implementation, the above training state prediction process includes the following steps:

[0081] First, based on the various sub-indicators of the same training round, a comprehensive training score for the same training round is generated. Thus, the historical training performance curve is obtained. :

[0082]

[0083] in, , , , Weights derived from experience or learning; t represents the current time. based on , , and It is obtained through weighted calculation, smoothing, or trend fitting, and is used to characterize the changing trend of user capabilities over time.

[0084] Then, based on the various sub-indicators and the overall training score from the multiple rounds of training, the user's training status is determined. Specifically, the aforementioned sub-indicators and overall training score, after standardization, are input into an LSTM-based time series learning model to learn the user's current training status, which is a comprehensive reflection of the user's fatigue level, cognitive load, and ability stability.

[0085] Optionally, the user's sub-metrics and overall training score in the same training round can be concatenated into a training performance vector for that round; the training performance vectors from multiple rounds can be arranged into a time series and input into an LSTM model, which then outputs the user's training state. Assume the input to the LSTM model is the training round window. The feature sequence within the data is the data input to the LSTM model for this prediction. include:

[0086]

[0087] The output of the LSTM model is the user's current training state label. The system will use the label ∈{well adapted, mildly fatigued, at the critical point of ability, consistently high efficiency} to determine whether the training difficulty needs to be adjusted.

[0088] S160, difficulty adaptive adjustment.

[0089] This step is also performed by the difficulty adaptive module, which controls the task difficulty at a higher level. The specific control objects include language complexity (sentence length, vocabulary density, semantic nesting level), interaction density (number of dialogue turns, prompt interval, context switching), and rhythm control (playback speed, task interval time).

[0090] In one specific implementation, this adjustment can be performed using a reinforcement learning algorithm. After obtaining the training state, the reinforcement learning model is started using the training state and the combined training score of the current training session as state variables. The reinforcement learning model generates a reward signal based on the user's various sub-indicators and the combined training score in the current training session, and determines the task difficulty for the next round of training based on the reward signal.

[0091] Specifically, the state variables of this reinforcement learning model include two parts: training state. This reflects the user's current state; the overall training score. This reflects the dynamic changes in user behavior. Therefore, the state vector of the reinforcement learning model is represented as: .

[0092] The action variables of the reinforcement learning model Corresponding to fine-grained adjustments to task parameters, this specifically includes the following four categories of parameters:

[0093] Language complexity parameters include sentence length (number of tokens). Lexical density (proportion of low-frequency words) Semantic nesting hierarchy (number of complex sentence structures such as subject-verb-object, time condition, etc.) Information content (number of named entities + number of keywords) ;

[0094] Interaction density parameter: includes the total number of dialogue rounds Context switching frequency ; Minimum interval prompt ;

[0095] Rhythm parameters: including playback speed Task interval time ;

[0096] Noise parameters: including background noise signal-to-noise ratio .

[0097] Finally, the system's action vector is: , The elements in the table represent the adjustments made to sentence length, lexical density, semantic nesting level, information content, total number of dialogue turns, context switching frequency, minimum cue interval, playback speed, task interval time, and signal-to-noise ratio.

[0098] The reward function of this reinforcement learning model It is still based on the user's training performance and is used to guide policy optimization:

[0099]

[0100] in, , , , and These represent the weighting coefficients.

[0101] In one specific implementation, the policy optimization process of the reinforcement learning model includes two stages: offline and online. The offline stage utilizes historical training log data and employs offline reinforcement learning methods (such as implicit Q-learning or conservative Q-learning) to pre-train the initial policy, ensuring safety and availability during cold starts. The online stage refers to updating the policy during actual training using a proximal policy optimization algorithm with a recurrent neural network structure.

[0102] For discrete parameters , , The policy network outputs a multinomial distribution; for continuous parameters... The policy network outputs a Gaussian distribution (mean and variance); the system samples actions from the distribution and introduces constraint mechanisms (such as Lagrange penalty or constraint policy optimization) to ensure that parameter adjustments do not exceed a preset threshold.

[0103] The sampled action vector This will be mapped to the increment of training parameters. : ,in, This represents the mapping function. The update rule for the training parameters is:

[0104]

[0105] in: and These represent the current parameter vector and the updated parameter vector, respectively. and Indicates the upper and lower limits of each dimension parameter (such as the upper and lower limits determined by the range of sentence length, the range of speech rate ratio, etc.). It is a range control function used to ensure that the parameter does not exceed the acceptable range.

[0106] Finally, based on the updated parameter vector Generate the task instruction set for the next training round. Based on this, language materials (controlling sentence length, vocabulary density, semantic level), interaction flow (number of dialogue turns, prompt interval, scene switching) and playback rhythm (speech speed, task interval) are dynamically constructed to achieve fine-grained scheduling of task difficulty.

[0107] S170, Semantic noise dynamic control.

[0108] This step is executed by the semantic noise dynamic control module, which is responsible for "external interference control of the task environment" and determines the complexity of the background noise. Specifically, based on the semantic difficulty obtained in S6, the corpus content for the next round of training can be generated. For example, training corpus with the specified number of keywords can be generated, and environmental noise matching the training scenario can be added at the keyword positions according to the noise signal-to-noise ratio obtained in S6.

[0109] In one specific implementation, firstly, matched ambient noise is retrieved based on the training scenario. Optionally, based on the scenario category of the current training task (e.g., medical consultation, shopping, transportation, ordering food), semantically equivalent or conflicting scene sounds are matched from a background noise library (e.g., hospital corridor prompts and conversations, supermarket broadcasts and shopping conversations, etc.). Matching uses semantic tag alignment.

[0110]

[0111] in, For semantic topics, A predefined scene noise label library; Match represents the matching function, which can be implemented based on BERT (Bidirectional Encoder Representations from Transformers) vectors or TF-IDF (Term Frequency – Inverse Document Frequency) similarity. This indicates the type of ambient noise that was matched.

[0112] Then, based on the noise signal-to-noise ratio, environmental noise matching the training scenario is added at the keyword location. Optionally, key semantic information (such as time, location, numbers, verb commands, etc.) is identified in the training corpus using NLP algorithms, and dynamic background noise is loaded within a 2-second range before and after the keyword or core semantic meaning. The noise types include human conversation, broadcast prompts, environmental noise, etc., to construct a "semantic-guided" interference scenario. For example, when the user needs to identify key nouns (such as drug names, amounts, addresses), interference words with similar syllables or similar semantics are superimposed; background human voices or dynamic environmental noise (such as supermarket broadcasts announcing different discount information) are inserted into transitional sentences and interrogative sentences.

[0113] Furthermore, to more closely approximate the effects of realistic environmental noise, the location of the noise source in the training scene can be determined based on the type of environmental noise. For example, in a coffee shop scene, different types of environmental noise include the sound of cups and plates clattering, the sound of running water, the sound of conversation, the sound of traffic outside the window, and the sound of music playing, etc., and their corresponding noise source locations are the dining area, the operation area, the multi-person dining area, outside the window, and the sound system area, respectively.

[0114] Then, the signal-to-noise ratio (SNR) of the noise given by the reinforcement learning model is used as the SNR at each noise source location. Based on the distance from the user's dialogue location in the training scene to each noise source location, the SNR of different types of environmental noise heard by the user can be calculated. Here, the SNR of different types of environmental noise refers to the ratio of the speech signal power to the power of different types of environmental noise. When the speech signal power remains constant, the higher the power (or intensity) of the environmental noise, the lower the SNR. For example, if the user's current dialogue location is the ordering area, then according to the inverse relationship between noise intensity (i.e., noise power) and the square of the distance, the actual intensity of each type of environmental noise heard by the user can be calculated from the distance from the ordering area to the aforementioned sound source locations. The system assumes that the power (or intensity) of the effective speech signal heard by the user remains constant at all locations (even if it changes, it is due to the user's volume adjustment of all sounds). Therefore, the SNR of different types of environmental noise heard by the user at different locations can be obtained.

[0115] Furthermore, the user interface module can also provide users with the function of selecting the dialogue position. When users feel uncomfortable in a noisy environment, they can choose to adjust the dialogue position. This embodiment provides a way to roughly judge the user's ability to identify mixed environmental noise based on the user's adjustment of the dialogue position, and adjust the corresponding training content accordingly. In a specific implementation, this process may include the following steps:

[0116] Step 1: In the same training round, if the user chooses to adjust their dialogue position when performing poorly, the signal-to-noise ratio (SNR) of each type of environmental noise heard by the user can be dynamically adjusted based on the spatial distance between the user's dialogue position and the noise sources of different types of environmental noise. Specifically, when the user's position changes, the distance between the new user position and each environmental noise source also changes. The SNR of each type of environmental noise heard by the user can be re-determined based on the new distance, and the user will continue to the next round of training under the new environmental noise intensity.

[0117] Step Two: If the user repeatedly adjusts their dialogue position and demonstrates good training performance at the final dialogue position, the type of environmental noise that has the greatest impact on the user can be identified based on the signal-to-noise ratio (SNR) changes before and after each adjustment. Optionally, the user continues auditory-verbal training after each position adjustment. If the user continues to change dialogue positions after each training session until finally settling at a certain position, and maintains a good level of auditory-verbal training at that position, it indicates that the user has found a suitable noise environment for training by repeatedly adjusting their dialogue position. At this point, the intensity changes of each type of environmental noise at the original and final dialogue positions can be calculated separately. Combined with the semantic signal power, the noise sources with the largest SNR changes are identified. These noise sources have the greatest impact on the user, i.e., the noise sources the user avoids by continuously adjusting their position.

[0118] Step 3: Determine the shortest path between the noise source of the most impactful environmental noise type and the final dialogue location; and assess the user's ability to identify mixed environmental noise based on the similarity between the user's movement trajectory and the shortest path. Optionally, Frechet distance can be used to calculate the similarity between the two paths; if the similarity between the two paths is lower than a set threshold, it indicates that the user repeatedly wandered around before finding a suitable location, rather than quickly identifying the most impactful noise source and directly reaching the effective location. This may indicate that the user's ability to identify mixed environmental noise sources is lacking.

[0119] Step 4: Based on the above predictions, if the similarity between the two paths is too low, environmental sound recognition training can be added in subsequent training. Research shows that environmental sound recognition training helps improve speech recognition ability. Therefore, it is advisable to introduce environmental sound recognition training into auditory-speech rehabilitation training, including identification of single environmental sound types, mixed identification of multiple environmental sound types, identification of environmental sound source location, identification of environmental sound source intensity, and identification of environmental sound source distance, etc., to train users to accurately identify the type, location, intensity, and distance of environmental sounds, thereby improving auditory-speech ability.

[0120] This embodiment provides an assessment strategy for users' ability to recognize ambient sounds and adjusts subsequent training content accordingly. Compared to completely disregarding users' ability to recognize ambient sounds or not assessing this ability at all, this approach improves the relevance of auditory-verbal training tasks and provides a suitable time to add ambient sound training.

[0121] S180, Data Recording and Feedback.

[0122] As described above, the data recording and feedback module records the user's training data for each session to the database, including answer results, pronunciation score details, number of replays, and answer duration. It then generates detailed reports based on the training data, including: scores for each training stage, error messages and improvement suggestions, the user's skill improvement curve, and a function to compare the user's recording with the standard pronunciation. Users can view these feedback reports at any time through the interactive interface.

[0123] In summary, the user interface architecture of the entire method is as follows: Figure 4 As shown, this is the software interface information architecture layout for the user side, including the registration and login interface, the ability test entry, the training mode selection interface, the training process interface, and the training report viewing interface, which aims to improve the intuitiveness and ease of use of user operations.

[0124] In summary, this embodiment provides an auditory-verbal rehabilitation training method that integrates semantic noise control and difficulty adaptation. First, it collects basic user information and initial speech ability data to generate a personalized ability profile. Then, it recommends training paths based on the profile results, executing articulation training, vocabulary training, and multi-turn situational dialogue tasks. During training, the system collects data in real time, including user answer accuracy, speech score, response time, and number of repetitions. Based on an ability trend model, it dynamically adjusts parameters such as speech rate, semantic complexity, and background interference intensity of the training content to achieve task difficulty adaptation. Simultaneously, it extracts keywords or core semantic information from the corpus using a semantic recognition algorithm and dynamically adjusts the background noise level before and after these keywords to simulate real attentional interference scenarios. After training, the system outputs multimodal feedback, including audio comparison, pronunciation lip-syncing videos, pronunciation animations, and ability trend graphs, improving adult users' auditory comprehension and language expression abilities in noisy environments. This embodiment improves the personalization accuracy, contextual adaptability, and training efficiency of rehabilitation training. Specifically, this embodiment has the following technical advantages:

[0125] Firstly, regarding the construction of training tasks, this embodiment proposes a four-level progressive training system consisting of articulation training, word training, sentence training, and situational dialogue training. This structure not only covers all ability dimensions in language rehabilitation, from basic pronunciation to semantic expression, but also features scenario-based design to meet the language needs of adult users in different contexts such as work, life, and social situations. This significantly improves the matching degree between training tasks and actual communication needs, avoiding the problems of simplistic language data and situational mismatch caused by existing systems that focus on children as the core audience.

[0126] Secondly, considering the significant individual differences and fluctuating abilities among adult hearing-impaired users, this embodiment designs an intelligent adaptive training difficulty mechanism. This mechanism not only collects user performance data in each round of training, including accuracy, speech score, number of re-listening attempts, and response time, but also incorporates historical ability trends for multi-dimensional modeling. Through a reinforcement learning model, it adjusts the complexity of training tasks, interaction frequency, and training path direction in real time. This approach effectively avoids the "overestimation" or "underestimation" problems caused by static grading, achieving a dynamic and accurate match between user ability levels and task challenge.

[0127] In terms of environmental simulation, this embodiment proposes for the first time a semantically linked, real-scene noise intervention mechanism. Unlike existing technologies that only use background noise of fixed intensity as generalized interference, this embodiment introduces dynamic noise sources that are related to or adversarial to the semantic logic during key nodes of the user's semantic task (such as information extraction, dialogue response, and digit recognition), thereby improving the user's auditory separation and semantic reconstruction capabilities in noisy environments. This semantic-noise coupling design is more in line with the auditory challenges under real-world communication conditions and helps to promote the transition of rehabilitation training for hearing-impaired users from "static indoor environments" to "realistic environments."

[0128] Furthermore, to enhance user engagement and perception of training effectiveness, this embodiment constructs a multimodal feedback mechanism. The system can automatically generate structured feedback content after training, including pronunciation scoring maps, error phoneme localization, comparison playback of standard pronunciation and user audio, synchronized lip-sync video, and airflow trajectory animation. Users can not only intuitively understand their strengths and weaknesses in training, but also perform self-correction and proactive learning through visualization, enhancing their sense of accomplishment and sustained motivation in rehabilitation.

[0129] Figure 5 This is a flowchart of another auditory-verbal rehabilitation training method integrating semantic noise control and difficulty adaptation provided by an embodiment of the present invention. The method is executed by an electronic device and specifically includes the following steps:

[0130] S210. Obtain multiple sub-indicators from the user's most recent rounds of auditory-verbal rehabilitation training, wherein the multiple sub-indicators include answer accuracy, voice score, number of repeated playbacks, and answer duration.

[0131] S220. Generate a comprehensive training score for the same training round based on the various sub-indicators of the same training round.

[0132] Optionally, a weighted average of the sub-indicators in the same round of training can be calculated to obtain the comprehensive training score for the same round of training.

[0133] S230. Determine the user's training status based on the sub-indicators and comprehensive training score of the multi-round training.

[0134] Optionally, the user's various sub-indicators and comprehensive training scores in the same round of training can be concatenated into a training performance vector for the same round of training; the training performance vectors of the multiple rounds of training can be arranged into a time sequence and input into an LSTM model to obtain the user's training state, wherein the training state includes well-adapted, slightly fatigued, capable adjacency, and continuously efficient.

[0135] S240. Using the training state and the comprehensive training score of this round of training as state variables, start the reinforcement learning model. The reinforcement learning model generates a reward signal based on the user's sub-indicators and comprehensive training score in this round of training, and determines the task difficulty of the next round of training based on the reward signal. The task difficulty includes the number of keywords and the noise signal-to-noise ratio.

[0136] Optionally, the task difficulty includes language complexity parameters, interaction density parameters, rhythm parameters, and noise parameters; wherein, the language complexity parameters include sentence length, vocabulary density, semantic nesting level, and number of keywords; the interaction density parameters include the total number of dialogue rounds, scene switching frequency, and minimum prompt interval; the rhythm parameters include playback speed and task interval time; and the noise parameters include noise signal-to-noise ratio.

[0137] S250. Generate the corpus content for the next round of training based on the number of keywords, and add environmental noise matching the training scene at the keyword positions according to the noise signal-to-noise ratio.

[0138] Optionally, based on the training scenario, matching environmental noise can be invoked; and based on the noise signal-to-noise ratio, the environmental noise can be added to the positions of keywords and key tone.

[0139] Optionally, the location of the noise source in the training scene can be determined according to the type of environmental noise; the noise signal-to-noise ratio can be used as the signal-to-noise ratio at the location of the noise source; the signal-to-noise ratio heard by the user can be calculated based on the distance from the user's dialogue location in the training scene to the location of the noise source; and environmental noise can be added at the keyword location based on the signal-to-noise ratio heard by the user.

[0140] Furthermore, the environmental noise can be of various types; correspondingly, the method further includes:

[0141] Step 1: In the same round of training, if the user chooses to adjust the dialogue position when the performance is poor, the signal-to-noise ratio of each type of environmental noise will be dynamically adjusted according to the spatial distance between the user's dialogue position and the noise source of different types of environmental noise.

[0142] Step 2: If the user adjusts the dialogue position multiple times and has good training performance at the final dialogue position, identify the type of environmental noise that has the greatest impact on the user based on the change in signal-to-noise ratio before and after the adjustment of each type of environmental noise.

[0143] Step 3: Determine the shortest path between the noise source of the most influential environmental noise type and the final dialogue location, and judge the user's ability to identify mixed environmental noise based on the similarity between the user's movement trajectory and the shortest path.

[0144] Step 4: Based on the recognition ability, add environmental sound recognition training in subsequent training.

[0145] Optionally, if the recognition ability is too low, at least one of the following environmental sound recognition training methods may be added in subsequent training: single environmental sound type recognition, mixed recognition of multiple environmental sound types, environmental sound source location recognition, environmental sound source intensity recognition, and environmental sound source distance recognition.

[0146] It is worth mentioning that the method of this embodiment is based on the same inventive concept as the method of any embodiment constituted by S110-S180 above, and the limitations in any of the above embodiments are applicable to this embodiment and can achieve the same beneficial effects.

[0147] It should be noted that the user data involved in this application (including facial data, facial expression data, body movement data, high-definition video images of learners, etc.) are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0148] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention, such as... Figure 6As shown, the device includes a processor 60, a memory 61, an input device 62, and an output device 63; the number of processors 60 in the device can be one or more. Figure 6 Taking a processor 60 as an example; the processor 60, memory 61, input device 62, and output device 63 in the device can be connected via a bus or other means. Figure 6 Taking the example of a connection between China and Israel via a bus.

[0149] The memory 61, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the auditory-speech rehabilitation training method integrating semantic noise control and difficulty adaptation in this embodiment of the invention. The processor 60 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 61, thereby realizing the aforementioned auditory-speech rehabilitation training method integrating semantic noise control and difficulty adaptation.

[0150] The memory 61 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function; the data storage area may store data created based on terminal usage. Furthermore, the memory 61 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory, or other non-volatile solid-state storage device. In some instances, the memory 61 may further include memory remotely located relative to the processor 60, which can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0151] Input device 62 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the device. Output device 63 may include display devices such as a display screen.

[0152] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the auditory-verbal rehabilitation training method of any embodiment that integrates semantic noise control and difficulty adaptation.

[0153] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0154] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.

[0155] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0156] Computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, and C++—as well as conventional procedural programming languages—such as C or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0157] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.

Claims

1. A method for auditory-verbal rehabilitation training that integrates semantic noise control and difficulty adaptation, characterized in that, include: The system obtains multiple sub-indicators from the user's most recent rounds of auditory-verbal rehabilitation training, including answer accuracy, voice score, number of repeated playbacks, and answer duration. Based on the various sub-indicators of the same training round, a comprehensive training score for the same training round is generated; The user's training status is determined based on the various sub-indicators and the overall training score of the multiple rounds of training. Using the training state and the comprehensive training score of this round of training as state variables, the reinforcement learning model is started. The reinforcement learning model generates a reward signal based on the user's sub-indicators and comprehensive training score in this round of training, and determines the task difficulty of the next round of training based on the reward signal. The task difficulty includes the number of keywords and the noise signal-to-noise ratio. The corpus content for the next round of training is generated based on the number of keywords, and environmental noise matching the training scenario is added at the keyword positions based on the noise signal-to-noise ratio.

2. The method according to claim 1, characterized in that, The process of generating a comprehensive training score for the same training round based on various sub-indicators includes: The weighted average of each sub-indicator in the same round of training is used to obtain the comprehensive training score for that round of training.

3. The method according to claim 1, characterized in that, The step of determining the user's training status based on the various sub-indicators and the overall training score of the multiple rounds of training includes: The user's various sub-metrics and overall training score in the same round of training are concatenated into a training performance vector for the same round of training. The training performance vectors from the multiple rounds of training are arranged into a time sequence and input into the LSTM model to obtain the user's training state, which includes well-adapted, slightly fatigued, capable adjacency, and continuously efficient.

4. The method according to claim 1, characterized in that, The task difficulty includes parameters such as language complexity, interaction density, rhythm, and noise. The language complexity parameters include sentence length, lexical density, semantic nesting level, and number of keywords; The interaction density parameters include the total number of dialogue rounds, scene switching frequency, and minimum prompt interval; The rhythm parameters include playback speed and task interval time. The noise parameters include the noise signal-to-noise ratio.

5. The method according to claim 1, characterized in that, The step of adding environmental noise matching the training scene at the keyword location based on the noise signal-to-noise ratio includes: Based on the training scenario, call up the matching environmental noise; Based on the noise signal-to-noise ratio, the environmental noise is added to the positions of keywords and key tones.

6. The method according to claim 1, characterized in that, The step of adding environmental noise matching the training scene at the keyword location based on the noise signal-to-noise ratio includes: Based on the type of environmental noise, determine the location of the noise source in the training scenario; The noise signal-to-noise ratio is used as the signal-to-noise ratio at the location of the noise source. The signal-to-noise ratio heard by the user is calculated based on the distance from the user's dialogue location in the training scenario to the location of the noise source. Based on the signal-to-noise ratio heard by the user, ambient noise is added at the keyword location.

7. The method according to claim 1, characterized in that, There are many types of environmental noise; Accordingly, the method further includes: In the same round of training, if a user chooses to adjust their dialogue position when performing poorly, the signal-to-noise ratio of each type of environmental noise heard by the user is dynamically adjusted based on the spatial distance between the user's dialogue position and the noise sources of different types of environmental noise. If a user adjusts the dialogue position multiple times in succession and has good training performance at the final dialogue position, the type of environmental noise that has the greatest impact on the user can be identified based on the change in signal-to-noise ratio before and after the adjustment of various types of environmental noise heard by the user. Determine the shortest path between the noise source of the most influential environmental noise type and the final dialogue location, and judge the user's ability to identify mixed environmental noise based on the similarity between the user's movement trajectory and the shortest path. Based on the aforementioned recognition ability, environmental sound recognition training will be added in subsequent training.

8. The method according to claim 7, characterized in that, The step of adding environmental sound recognition training in subsequent training based on the recognition ability includes: If the recognition ability is too low, at least one of the following environmental sound recognition training methods should be added in subsequent training: single environmental sound type recognition, mixed recognition of multiple environmental sound types, environmental sound source location recognition, environmental sound source intensity recognition, and environmental sound source distance recognition.

9. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the auditory-verbal rehabilitation training method that integrates semantic noise control and difficulty adaptation as described in any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the auditory-speech rehabilitation training method that integrates semantic noise control and difficulty adaptation as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Hearing and speech rehabilitation device and method, electronic device and storage medium

    CN110782962A

  • Reinforced learning hearing aid adaptation method combined with hearing loss speech evaluation model

    CN120302226A

  • AI-based preschool special child language rehabilitation training guiding method

    CN120432093A

  • Voice training system and data training method after complete denture repair

    CN120496383A

  • Sound source localization rehabilitation training method, system, equipment and medium

    CN120539676A