Methods, systems, electronic devices and storage media for generating auditory training tasks
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-25
- Publication Date
- 2026-08-14
AI Technical Summary
[0005]本发明提供一种听觉训练任务生成方法、系统、电子设备及存储介质,用以解决现有技术中听觉训练的训练模式僵化、场景适应性差以及缺乏情感维度考量的缺陷
Smart Images

Figure CN122575627A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of healthcare technology, and in particular to a method, system, electronic device, and storage medium for generating auditory training tasks. Background Technology
[0002] Hearing is a vital sensory channel for humans to acquire external information, engage in language communication, and interact socially. For individuals with hearing impairments, auditory neuropathy, children with language delays, and those requiring speech rehabilitation, restoring or enhancing auditory perception and comprehension is crucial for returning to normal life and integrating into society. Auditory training aims to reshape the function of the brain's auditory cortex through systematic sound stimulation and task practice, improving an individual's ability to perceive, distinguish, recognize, and understand sounds. It not only helps users better adapt to the sound signals from hearing aids, cochlear implants, and other hearing devices but also effectively enhances their verbal communication abilities in complex acoustic environments, thereby improving their quality of life and reducing the psychological burden caused by hearing impairment. Therefore, scientific and efficient methods for generating auditory training tasks play a crucial role in daily assistive training.
[0003] Existing auditory training methods typically rely on manual guidance from therapists or basic computer-assisted training software. Traditional rehabilitation training often follows a fixed process, from simple pure-tone audiometry to basic speech recognition and then to short sentence comprehension. The training content is mostly based on pre-set static corpora, lacking real-time dynamic adjustment mechanisms tailored to individual user differences. Furthermore, existing technologies are limited in the diversity of training samples, often lacking effective simulation of different speaker identities, speech rate variations, and complex acoustic environments, making it difficult to transfer trained auditory abilities to the varied scenarios of real life. More importantly, existing training programs focus primarily on improving cognitive skills while neglecting the cultivation of users' emotional cognitive abilities during training. The lack of systematic planning for the emotional training phase makes the training process tedious and difficult to maintain long-term user adherence.
[0004] In conclusion, how to dynamically adjust training methods based on user performance data has become a pressing technical problem to be solved in this field. Summary of the Invention
[0005] This invention provides a method, system, electronic device, and storage medium for generating auditory training tasks, in order to solve the shortcomings of existing auditory training technologies, such as rigid training modes, poor scene adaptability, and lack of consideration for emotional dimensions.
[0006] This application provides a method for generating auditory training tasks, including the following steps: Based on user performance data, the core training parameters are dynamically adjusted, including the dominant weight of the training mode, task complexity, stimulus variability, and emotional training stage. Based on the aforementioned core training parameters, a training task is generated; The dominant weight of the training mode is the ratio of analytical training to comprehensive training. Analytical training is a training mode for sub-vocabularies, while comprehensive training is a holistic training mode in real-world contexts. The task complexity reflects the difficulty of the statement length and syntactic structure. The stimulus variability reflects the level of diversity in speaker identity, speech rate, and acoustic environment in the training samples. The emotional training phase reflects the progression path of a user's emotional cognitive abilities.
[0007] According to the auditory training task generation method provided in this application, the training steps of the analytical training are as follows: While maintaining the playback of a specific phoneme audio stream in the audio channel, a video stream with similar pronunciation locations but different acoustic characteristics is superimposed and displayed in the visual channel, thus merging the audio stream and the video stream into a single multimedia stimulus stream; Play the multimedia stimulus stream and obtain the user's identification results; If the identification result indicates that the user has correctly identified the user, then it is determined that the user's auditory channel weight has been increased. If the identification result indicates that the user has made an incorrect identification, a correction process is executed, which is used to guide the user to focus on the auditory channel.
[0008] According to the auditory training task generation method provided in this application, the dominant weight of the training mode is dynamically adjusted based on the following steps: Based on the user's real-time response data and historical error records in the sub-vocabulary recognition task, identify the user's phoneme confusion pattern; After generating adversarial training samples based on the phoneme confusion pattern, training is performed based on the adversarial training samples. If the user's recognition accuracy in the sub-vocabulary recognition task reaches a preset threshold, the proportion of analytical training will be reduced.
[0009] According to the auditory training task generation method provided in this application, the auditory training task generation method further includes: During training, the video stream played to the user contains visual information about lip shape, teeth, visible parts of the tongue, cheek movements, jaw movements, facial expressions, and head posture.
[0010] According to the auditory training task generation method provided in this application, the task complexity is dynamically adjusted based on the following steps: Based on the task complexity, multimodal contextual cues are provided and corresponding multimedia stimulus streams are played. Based on the user's performance data in response to the multimedia stimulus stream, the user's verbal prediction error and emotional attitude prediction error are calculated. The task complexity is adjusted based on the verbal prediction error and the affective attitude prediction error.
[0011] According to the auditory training task generation method provided in this application, the auditory training task generation method further includes: After receiving the user's audio and video input, the audio and video input is deconstructed into sublexical unit dimension and sentiment dimension; Identify the confusing phonemes in the sublexical unit dimension and the misjudged emotion categories in the emotion dimension; Based on the confused phonemes, generate instructions for the confused phonemes and compare and play them; Based on the misjudged emotion category, emotional visual cues are generated and the corresponding video clips are played back.
[0012] This application also provides an auditory training device, including the following modules: The parameter adjustment module is used to dynamically adjust the core training parameters based on the user's performance data. The core training parameters include the dominant weight of the training mode, task complexity, stimulus variability, and emotional training stage. The task generation module is used to: generate training tasks based on the core training parameters; The dominant weight of the training mode is the ratio of analytical training to comprehensive training. Analytical training is a training mode for sub-vocabularies, while comprehensive training is a holistic training mode in real-world contexts. The task complexity reflects the difficulty of the statement length and syntactic structure. The stimulus variability reflects the level of diversity in speaker identity, speech rate, and acoustic environment in the training samples. The emotional training phase reflects the progression path of a user's emotional cognitive abilities.
[0013] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the auditory training task generation method described above.
[0014] This application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the auditory training task generation method as described above.
[0015] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the auditory training task generation method as described above.
[0016] The auditory training task generation method, system, electronic device, and storage medium provided in this application can dynamically adjust the ratio of analytical training to comprehensive training based on the user's real-time performance data through the training mode's dominant weight. That is, when the system detects a high error rate at the sub-vocabulary level, it increases the weight of analytical training to strengthen basic auditory discrimination ability; conversely, when the user's basic ability is solid but their performance in long sentence comprehension is poor, the system increases the weight of comprehensive training to guide the user into realistic contexts. This dynamic balancing mechanism ensures that the user can both clearly hear details and understand meaning, thus achieving a seamless transition from micro-level speech perception to macro-level language comprehension, significantly improving the overall effectiveness of rehabilitation training. Furthermore, the system dynamically adjusts the weight of analytical training based on the task complexity parameter. By adjusting sentence length and grammatical structure, the system progressively trains users' short-term memory and syntactic processing abilities. Simultaneously, by introducing diversity in speaker identity, speech rate, and acoustic environment, the system simulates varied soundscapes in real life, thereby improving speech recognition rates under different signal-to-noise ratios and speaker conditions. This enhances the transferability of auditory skills to everyday life scenarios, solving the problem of disconnect between training and practical application in existing technologies. Furthermore, by incorporating emotional cognition tasks into training, the system enriches its content, avoiding the metalanguage trap and emotional neglect issues of traditional training. By providing dual feedback on speech and emotion within context, it provides targeted and systematic improvement in speech comprehension and social communication abilities in complex environments. In summary, this application achieves a precise balance between analytical and comprehensive training by dynamically adjusting core training parameters, constructs a highly adaptable multi-dimensional difficulty control system, and integrates the emotional cognition dimension, thus significantly improving the personalization, scenario transferability, and user compliance of auditory training. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating the auditory training task generation method provided in this application; Figure 2 This is a schematic diagram of the auditory training task generation system provided in this application; Figure 3 This is a schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] The following is combined with Figures 1 to 3 This application describes the auditory training task generation method, system, electronic device, and storage medium.
[0021] Figure 1 This is a flowchart illustrating the auditory training task generation method provided in this application, as shown below. Figure 1 As shown, the method includes the following: S110, dynamically adjust the core training parameters based on user performance data, including the dominant weight of the training mode, task complexity, stimulus variability, and emotional training stage. S120, Based on the core training parameters, generate a training task.
[0022] In this embodiment, the dominant weight of the training mode is the ratio of analytical training to comprehensive training. Analytical training is a training mode targeting sub-vocabularies, while comprehensive training is a holistic training mode in real-world contexts. The task complexity reflects the difficulty of sentence length and grammatical structure. The stimulus variability reflects the diversity level of speaker identity, speech rate, and acoustic environment in the training samples. The emotion training stage reflects the progression path of the user's emotional cognitive ability.
[0023] It should be noted that the execution subject of the auditory training task generation method provided in this application embodiment can be a server or computer device, such as a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. The execution subject of the auditory training task generation method provided in this application embodiment can also be an auditory training task generation system. For ease of understanding, the following embodiments use an auditory training task generation system as the execution subject to describe the auditory training task generation method provided in this application.
[0024] In S110, performance data refers to the quantitative indicators generated by users during the performance of auditory training tasks. These may include, but are not limited to, task accuracy, reaction time, error type, number of repeated requests, etc., and may also include physiological feedback data, such as eye tracking, natural facial expressions, EEG attention index, etc.
[0025] In the actual implementation process, the system collects user performance data in real time, analyzes the performance data in real time, and updates the core parameters of the training through the preset algorithm model based on the analysis results.
[0026] Here, the dominant weight of the training mode refers to the allocation ratio between the micro-analysis and macro-comprehension modes in a single training task or a set of training tasks. Analytical training focuses on bottom-up processing, performing detailed listening and discrimination on sub-lexical units (such as initials, finals, tones, and syllables); comprehensive training focuses on top-down processing, using context, grammar, and logical clues to understand the overall meaning, such as sentence paraphrasing and story comprehension.
[0027] Here, task complexity refers to the degree of cognitive load required by the training task. Sentence length increases from easy to difficult, from single words and two-word phrases to phrases, long sentences, and paragraphs; grammatical structure increases from easy to difficult, from simple subject-verb-object structures to passive sentences, inverted sentences, double negative sentences, or complex clauses.
[0028] Here, stimulus variability refers to the acoustic feature diversity of the training samples, which aims to prevent users from memorizing specific sounds and promote the generalization of auditory skills. Speaker identity includes different age groups, genders, and dialect accents; speech rate can include slow, normal, fast, and variable speed; acoustic environment can include quiet, steady-state noise (such as fan noise), and non-steady-state noise (such as traffic flow and noisy crowds).
[0029] Here, the emotion training stage refers to the progressive settings for users' emotional cognitive abilities. For example, three stages can be set: beginner, intermediate, and advanced. From beginner to advanced, the learning progresses from recognizing basic and obvious emotions to distinguishing complex and mixed emotions.
[0030] In some embodiments, a predetermined threshold table is set, and the core training parameters are adjusted according to the threshold table. For example, if the accuracy of the last 5 tasks is >90%, the task complexity level is increased by one level.
[0031] In other embodiments, reinforcement learning algorithms are used to adjust the core training parameters. For example, the user is regarded as the environment, the system as the agent, and the user's accuracy and reaction time are used as reward functions. Reinforcement learning algorithms are used to continuously adjust the core training parameters to maximize the user's long-term learning benefits.
[0032] In S120, based on the updated parameters, the system retrieves or synthesizes specific training materials from the cloud or local database for subsequent training.
[0033] In some embodiments, a large amount of video material is pre-recorded and labeled with corresponding phonemes, complexity, stimulus variability, and emotion tags. Based on the core training parameters, video segments with corresponding tags are extracted from the library and spliced together to form a training task.
[0034] In other embodiments, after generating corresponding system instructions based on the training core parameters, the system instructions are input into the large language model, and the large language model is used to generate the corresponding video segments in real time.
[0035] The auditory training task generation method provided in this application dynamically adjusts the ratio of analytical training to comprehensive training based on the user's real-time performance data through the training mode-dominant weighting. That is, when the system detects a high error rate at the sub-vocabulary level, it increases the weight of analytical training to strengthen basic auditory discrimination ability. Conversely, when the user has a solid foundation but performs poorly in long sentence comprehension, the system increases the weight of comprehensive training to guide the user into realistic contexts. This dynamic balancing mechanism ensures that the user can both clearly hear details and understand the meaning, thus achieving a seamless transition from micro-level speech perception to macro-level language comprehension, significantly improving the overall effectiveness of rehabilitation training. Furthermore, it dynamically adjusts the sentence length through task complexity parameters. In addition to grammatical structure, the system can progressively train users' short-term memory and syntactic processing abilities. Simultaneously, by introducing diversity in speaker identity, speech rate, and acoustic environment, the system simulates varied soundscapes in real life, thereby improving users' speech recognition rate under different signal-to-noise ratios and speaker conditions. This enhances the transferability of auditory skills to everyday life scenarios, solving the problem of disconnect between training and practical application in existing technologies. Furthermore, by incorporating emotional cognition tasks into training, the system enriches its content, avoiding the metalanguage trap and emotional neglect issues of traditional training. By providing dual feedback of speech and emotion within context, it achieves targeted and systematic improvement in speech comprehension and social communication abilities in complex environments. In summary, this application achieves a precise balance between analytical and comprehensive training by dynamically adjusting core training parameters, constructs a highly adaptable multi-dimensional difficulty control system, and integrates the emotional cognition dimension, thereby significantly improving the personalization, scenario transferability, and user compliance of auditory training.
[0036] In an optional embodiment, the training steps of the analytical training are as follows: While maintaining the playback of a specific phoneme audio stream in the audio channel, a video stream with similar pronunciation locations but different acoustic characteristics is superimposed and displayed in the visual channel, thus merging the audio stream and the video stream into a single multimedia stimulus stream; Play the multimedia stimulus stream and obtain the user's identification results; If the identification result indicates that the user has correctly identified the user, then it is determined that the user's auditory channel weight has been increased. If the identification result indicates that the user has made an incorrect identification, a correction process is executed, which is used to guide the user to focus on the auditory channel.
[0037] This application's embodiments construct a perceptual loop that deeply integrates auditory and visual senses. Unlike simple parallel multimodal stimulation, this approach uses an audiovisual conflict task (such as playing the phoneme / ba / while simultaneously displaying the lip movements of / pa / ) to force the brain to perform cross-modal calibration, actively breaking the patient's visual dependence and promoting the reconstruction and strengthening of their auditory pathways.
[0038] Here, a specific phoneme audio stream refers to a pure audio signal recorded for a specific speech unit (such as the initial consonant / b / , the final vowel / a / , or a specific syllable). This audio stream has been standardized to ensure the consistency of intensity and duration, serving as the basic stimulus source for auditory training.
[0039] Here, video streams with similar pronunciation points but different acoustic characteristics refer to video content where the visual and auditory content conflict or differ. For example, the audio might play / b / while the video displays the lip movements for / p / , or the audio might use a front vowel while the video displays the lip movements for a back vowel. This design aims to create audiovisual conflict to test and train users' reliance on auditory information.
[0040] In practice, the system constructs training samples in the background, selecting an audio sample of a target phoneme, such as / pa / , and simultaneously matching it with a video of similar articulation points but different auditory perceptions, such as a video of lip movements for / ba / or / ma / . The system merges these two samples to generate a multimedia stimulus stream with a mismatch between sound and image. The system plays this multimedia stimulus stream to the user, who watches the video and listens to the audio. The user is then asked to identify what they hear. For example, the screen may display three options: "fear," "dad," and "mom." At this point, the user faces cognitive conflict: their eyes see the lip movements for "dad," while their ears hear the sound for "fear." If the user identifies correctly, it indicates that they have successfully suppressed visual interference and accurately extracted auditory features. The system determines the task is successful and increases the weight of the auditory channel. In subsequent training, the system will further increase the difficulty of the auditory task or reduce the salience of visual assistance. If a user makes a mistake, it indicates that the user has been deceived by visual cues and is overly reliant on vision. The system determines that the task has failed, triggers a correction process, and guides the user to focus on the auditory channel. For example, targeted training can be conducted for the two confused options, multimedia stimulus streams can be played back, and the differences in auditory features can be highlighted to guide the user to refocus on the auditory channel.
[0041] The auditory training task generation method provided in this application embodiment deliberately creates a conflict between sound and image, forcing users to suppress excessive reliance on vision, readjust and strengthen auditory processing capabilities, and improve the auditory channel weight by continuously confirming correct recognition, thereby improving the user's real hearing level in an environment without visual assistance.
[0042] In an optional embodiment, the dominant weights of the training mode are dynamically adjusted based on the following steps: Based on the user's real-time response data and historical error records in the sub-vocabulary recognition task, identify the user's phoneme confusion pattern; After generating adversarial training samples based on the phoneme confusion pattern, training is performed based on the adversarial training samples. If the user's recognition accuracy in the sub-vocabulary recognition task reaches a preset threshold, the proportion of analytical training will be reduced.
[0043] According to the inverse hierarchy theory, effective perceptual learning requires external feedback to guide learners to focus on and utilize sublexical stimuli that exist at the early sensory level but have not been used for higher-level classification tasks. Therefore, this application's embodiments create a progressive training path from sublexical to vocabulary, and then to sentences. Simultaneously, addressing the phoneme confusion issues unique to cochlear implant users, such as / s / and / ʃ / , / p / and / b / , a confusion model is constructed. Based on the user's real-time and historical error data, adversarial training samples are dynamically generated for minimal contrast training, inducing effective perceptual learning through consonant level feedback.
[0044] Here, sublexical recognition tasks refer to recognition exercises targeting speech units smaller than the word level, specifically including recognition tasks for initials, finals, tones, or syllables. This is the most basic micro-level training in auditory training.
[0045] Here, phoneme confusion patterns refer to specific and recurring errors exhibited by users in auditory recognition. For example, users may consistently mishear aspirated sounds (such as "pa") as unaspirated sounds (such as "ba"), or they may consistently confuse alveolar consonants (such as / z / ) with retroflex consonants (such as / zh / ). This pattern reflects a deficiency in the user's auditory system's ability to distinguish specific acoustic features.
[0046] Here, adversarial training samples refer to highly deceptive training materials specifically generated to target users' phoneme confusion patterns. These samples are extremely similar in acoustic features or are easily misjudged in context, and are used to target users' weak auditory areas, forcing them to make more refined distinctions.
[0047] In this embodiment, the system monitors the user's performance in sub-vocabulary tasks in real time, records the user's current reaction and reaction time, retrieves the user's historical error records, identifies recurring error pairs, and uses statistical algorithms to determine the user's specific phoneme confusion patterns. Once a phoneme confusion pattern is identified, targeted adversarial samples are generated or retrieved and inserted into the current training stream. If the user performs consistently in adversarial training, the proportion of analytical training is automatically reduced, while the proportion of comprehensive training is increased. The training path transitions from single phoneme discrimination to syllable combinations, and then extends to words and sentence contexts containing the phoneme. Through this bottom-up, progressive path, the system ensures that the underlying perceptual repair can stably support high-level language comprehension capabilities.
[0048] The auditory training task generation method provided in this application accurately locates the user's auditory weaknesses by identifying phoneme confusion patterns. The subsequently generated adversarial training samples focus on these weaknesses, significantly shortening the training cycle. When the user's identification accuracy reaches a preset threshold, more complex comprehensive training is introduced to ensure that the training difficulty is always dynamically matched with the user's ability, achieving a dynamic and smooth transition in training difficulty, preventing the occurrence of training plateaus, further shortening the training cycle, and improving training efficiency.
[0049] In an optional embodiment, the auditory training task generation method further includes: During training, the video stream played to the user contains visual information about lip shape, teeth, visible parts of the tongue, cheek movements, jaw movements, facial expressions, and head posture.
[0050] This application's embodiments change the traditional lip-reading training model that isolates close-ups of the mouth, fully utilizing the value of the face as a whole in speech perception and communication. Based on the fact that the brain naturally integrates various visual cues such as lip movements, facial expressions and their dynamic changes when processing speech information, it provides the following dynamic facial context integration training: (1) Utilization of full-face dynamic information: All visual and speech stimuli are presented as complete, dynamic video of the speaker's face, allowing users to obtain comprehensive visual and speech information, including lip shape, teeth, visible parts of the tongue, cheek and jaw movements, rather than just isolated lips; (2) Hyperlinguistic information fusion: The system deliberately preserves and utilizes the speaker's natural facial expressions (such as smiling and puzzled) and subtle changes in head posture. These hyperlinguistic information are an indispensable part of real communication and help users learn in training how to use these cues to aid understanding and interaction in actual communication.
[0051] (3) Contextualized and emotional communication scenarios: The training tasks are set in video scenarios that simulate real life, such as pleasant parties, serious discussions, and caring greetings, so that facial information and emotional expression are always in a meaningful context. This overcomes the drawback of traditional training that artificially separates visual cues, emotional expression and communication context, and promotes the effective transfer of audiovisual integration and emotional understanding to the real world.
[0052] (4) Specialized emotion perception training: The system includes specialized tasks that require users to identify or match the speaker's emotional state. The training material library is labeled with multi-dimensional emotion tags (such as emotion category, intensity, and valence). The tasks progress from identifying basic and obvious emotions to distinguishing complex and mixed emotions.
[0053] The auditory training task generation method provided in this application provides users with visual counterparts of acoustic features by displaying details such as teeth, tongue, and jaw. This full-dimensional visual input strengthens the connection between the auditory cortex and the visual cortex of the brain. When users encounter hearing difficulties in real life, they can more automatically and accurately call these visual cues to assist in understanding, significantly improving the speech recognition rate in noisy environments.
[0054] In an optional embodiment, the task complexity is dynamically adjusted based on the following steps: Based on the task complexity, multimodal contextual cues are provided and corresponding multimedia stimulus streams are played. Based on the user's performance data in response to the multimedia stimulus stream, the user's verbal prediction error and emotional attitude prediction error are calculated. The task complexity is adjusted based on the verbal prediction error and the affective attitude prediction error.
[0055] Here, multimodal contextual cues refer to auxiliary information presented to users before or during the playback of core training videos, which aims to activate users' background knowledge and expectations. For example, in the visual modality, images and short videos related to the topic are displayed, while in the text modality, background information or question guidance for the dialogue is provided.
[0056] Here, speech prediction error refers to the difference between what a user actually hears and what is predicted based on context, as well as the user's ability to process this difference. The computational dimensions include whether the user correctly predicted words using contextual cues, or whether the user can overcome contextual interference and accurately identify the words actually heard when the context is misleading.
[0057] Here, the emotion / attitude prediction error refers to the difference between the speaker's expected emotion / attitude and the emotion / attitude actually expressed in the audio.
[0058] In this embodiment, to transform training from a tedious metalinguistic task into vivid, realistic communication, a design based on predictive coding theory is employed, integrating an emotion prediction dimension. Before each training round, the system provides multimodal contextual cues, such as images, scene descriptions, and the first half of a dialogue, guiding users to actively predict subsequent verbal content and potential accompanying emotional attitudes. The system dynamically adjusts the difficulty and complexity of subsequent stimuli based on the user's verbal and emotional prediction errors. This design simulates the anticipation, understanding, and empathy processes in real communication, effectively promoting the transfer of learning outcomes to everyday situations and avoiding the drawbacks of traditional reflective learning.
[0059] In practice, the system determines the intensity of contextual cues to provide based on the current task complexity level. The system plays a stream of target multimedia stimuli, and the user identifies or answers questions. The system records the user's reactions and reaction times, and calculates verbal prediction error and affective attitude prediction error. If the user's error is too large, the task complexity level is reduced; if the error is moderate or small, the task complexity level is further increased.
[0060] The auditory training task generation method provided in this application calculates speech prediction error and trains users to find a balance between utilizing context and verifying auditory perception, enabling users to quickly complete ambiguous speech signals using context in real life, thus significantly improving communication efficiency.
[0061] In an optional embodiment, the auditory training task generation method further includes: After receiving the user's audio and video input, the audio and video input is deconstructed into sublexical unit dimension and sentiment dimension; Identify the confusing phonemes in the sublexical unit dimension and the misjudged emotion categories in the emotion dimension; Based on the confused phonemes, generate instructions for the confused phonemes and compare and play them; Based on the misjudged emotion category, emotional visual cues are generated and the corresponding video clips are played back.
[0062] Here, audio and video input refers to real-time interactive data collected by the user during training through microphones, cameras, or sensors. This includes not only the user's voice responses but also their facial expressions, body movements, or clicks on on-screen options.
[0063] Here, the sublexical unit dimension refers to the set of acoustic features in a speech signal below the word level, specifically including phonemes of initials and finals, syllable structure, tone, and prosodic features such as stress and duration.
[0064] Here, confused phonemes refer to phoneme pairs that are misidentified or generated in a user's auditory perception or pronunciation repetition due to similar acoustic features. For example, mishearing the voiceless consonant / p / as the voiced consonant / b / , or mishearing the alveolar consonant / z / as the retroflex consonant / zh / .
[0065] Here, the emotional dimension refers to the non-verbal information carried in the speech signal, reflecting the speaker's emotional state, attitude, or intention.
[0066] Here, misclassification of emotion category refers to a user incorrectly identifying the actual emotion expressed in the audio as another emotion, or being unable to identify the emotion category label.
[0067] Here, the confusing phoneme guidance is a text or voice prompt generated by the system for a specific confusing phoneme pair, which includes a description of the place of articulation, a comparison of acoustic features, or a mnemonic.
[0068] Here, emotional visual cue analysis refers to the digital interpretation of key facial features of a speaker in a video, used to explain why a particular expression corresponds to a specific emotion.
[0069] Traditional training systems typically employ simple binary feedback (correct / incorrect) or provide overall confirmation only for completely correct words. This feedback lacks specificity and fails to provide effective guidance for partially correct responses. This application's embodiments introduce a partially correct feedback mechanism at the consonant level and an analytical feedback mechanism based on emotional cues. In sentence or dialogue context training, the system, through real-time phoneme-phoneme alignment analysis and emotional feature analysis, can accurately identify correct and incorrect phoneme components in user responses, as well as the correctness and deviation of emotional judgments. For example, when a user misidentifies the target word / ba / as / pa / , the system not only marks the error but also explicitly points out that the initial phoneme / b / was misheard as / p / , the middle vowel / a / was correctly identified, and provides targeted comparative training. When a user misjudges the speaker's emotion as anger, the system provides feedback: "You judged it as anger, but more accurately, surprise. Please observe the speaker's raised eyebrows and wide eyes, while anger is usually characterized by furrowed brows and downturned lips," and replays key segments. This refined, multi-dimensional feedback mechanism aligns perfectly with the reverse hierarchy theory. By guiding users to focus on sublexical features and emotional cues that have been received by the sensory system but have not yet been fully utilized through external feedback, it achieves precise positioning and efficient error correction.
[0070] In the specific implementation process, the system receives audio and video input from users and decomposes it into phoneme sequences and emotional feature vectors. In analytical training, the acoustic distance between the user's pronunciation and the standard pronunciation is directly calculated. In comprehensive training, the system calculates whether the user has misunderstood the system's question based on the user's answer, thereby identifying confusing phonemes at the sub-lexical unit level. The system analyzes the user's intonation and facial expressions to identify misjudged emotion categories. For the identified confusing phonemes, comparative teaching materials are generated. For the identified misjudged emotions, key visual cues are retrospectively highlighted.
[0071] The auditory training task generation method provided in this application provides a way to visualize abstract auditory errors by generating confusing phoneme instructions and analyzing emotional visual cues. This helps users establish a clear cognitive model of errors, causes, and corrections, enabling precise location and efficient error correction, and significantly shortening the trial-and-error time.
[0072] The auditory training task generation system provided in this application is described below. The auditory training task generation system described below can be referred to in correspondence with the auditory training task generation method described above.
[0073] Figure 2 This is a schematic diagram of the auditory training task generation system provided in this application, such as... Figure 2 As shown, the auditory training task generation system may include, but is not limited to: The parameter adjustment module 210 is used to: dynamically adjust the core training parameters based on the user's performance data, wherein the core training parameters include the dominant weight of the training mode, task complexity, stimulus variability, and emotional training stage. Task generation module 220 is used to: generate training tasks based on the training core parameters; The dominant weight of the training mode is the ratio of analytical training to comprehensive training. Analytical training is a training mode for sub-vocabularies, while comprehensive training is a holistic training mode in real-world contexts. The task complexity reflects the difficulty of the statement length and syntactic structure. The stimulus variability reflects the level of diversity in speaker identity, speech rate, and acoustic environment in the training samples. The emotional training phase reflects the progression path of a user's emotional cognitive abilities.
[0074] The auditory training task generation system provided in this application dynamically adjusts the weighting of analytical and comprehensive training based on the user's real-time performance data through a training mode-dominant weighting mechanism. Specifically, when the system detects a high error rate at the sub-vocabulary level, it increases the weighting of analytical training to strengthen basic auditory discrimination abilities. Conversely, when the user has a solid foundation but performs poorly in long sentence comprehension, the system increases the weighting of comprehensive training to guide the user into realistic contexts. This dynamic balancing mechanism ensures that the user can clearly hear details and understand meaning, thus achieving a seamless transition from micro-level speech perception to macro-level language comprehension, significantly improving the overall effectiveness of rehabilitation training. Furthermore, the system dynamically adjusts the sentence length and... The system progressively trains users' short-term memory and syntactic processing abilities through its grammatical structure. Simultaneously, by introducing diversity in speaker identity, speech rate, and acoustic environment, the system simulates varied soundscapes in real life, thereby improving speech recognition rates under different signal-to-noise ratios and speaker conditions. This enhances the transferability of auditory skills to everyday life scenarios, solving the problem of disconnect between training and practical application in existing technologies. Furthermore, by incorporating emotional cognition tasks into the training, the system enriches its content, avoiding the metalanguage trap and emotional neglect issues of traditional training. By providing dual feedback on speech and emotion within context, it achieves targeted and systematic improvement in speech comprehension and social communication abilities in complex environments. In summary, this application achieves a precise balance between analytical and comprehensive training by dynamically adjusting core training parameters, constructs a highly adaptable multi-dimensional difficulty control system, and integrates the emotional cognition dimension, thus significantly improving the personalization, scenario transferability, and user compliance of auditory training.
[0075] In an optional embodiment, the training steps of the analytical training are as follows: While maintaining the playback of a specific phoneme audio stream in the audio channel, a video stream with similar pronunciation locations but different acoustic characteristics is superimposed and displayed in the visual channel, thus merging the audio stream and the video stream into a single multimedia stimulus stream; Play the multimedia stimulus stream and obtain the user's identification results; If the identification result indicates that the user has correctly identified the user, then it is determined that the user's auditory channel weight has been increased. If the identification result indicates that the user has made an incorrect identification, a correction process is executed, which is used to guide the user to focus on the auditory channel.
[0076] In an optional embodiment, the parameter adjustment module 210 is specifically used for: Based on the user's real-time response data and historical error records in the sub-vocabulary recognition task, identify the user's phoneme confusion pattern; After generating adversarial training samples based on the phoneme confusion pattern, training is performed based on the adversarial training samples. If the user's recognition accuracy in the sub-vocabulary recognition task reaches a preset threshold, the proportion of analytical training will be reduced.
[0077] In an optional embodiment, the auditory training task generation system further includes a display module for: During training, the video stream played to the user contains visual information about lip shape, teeth, visible parts of the tongue, cheek movements, jaw movements, facial expressions, and head posture.
[0078] In an optional embodiment, the parameter adjustment module 210 is further specifically used for: Based on the task complexity, multimodal contextual cues are provided and corresponding multimedia stimulus streams are played. Based on the user's performance data in response to the multimedia stimulus stream, the user's verbal prediction error and emotional attitude prediction error are calculated. The task complexity is adjusted based on the verbal prediction error and the affective attitude prediction error.
[0079] In an optional embodiment, the auditory training task generation system further includes a feedback module for: After receiving the user's audio and video input, the audio and video input is deconstructed into sublexical unit dimension and sentiment dimension; Identify the confusing phonemes in the sublexical unit dimension and the misjudged emotion categories in the emotion dimension; Based on the confused phonemes, generate instructions for the confused phonemes and compare and play them; Based on the misjudged emotion category, emotional visual cues are generated and the corresponding video clips are played back.
[0080] It should be noted that the auditory training task generation system provided in this embodiment of the invention can execute the auditory training task generation method described in any of the above embodiments during specific operation, which will not be elaborated in this embodiment.
[0081] Figure 3 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 3 As shown, the electronic device may include: a processor 310, a communications interface 320, a memory 330, and a communication bus 340, wherein the processor 310, the communications interface 320, and the memory 330 communicate with each other via the communication bus 340. The processor 310 can call logical instructions in the memory 330 to execute an auditory training task generation method, which includes: Based on user performance data, the core training parameters are dynamically adjusted, including the dominant weight of the training mode, task complexity, stimulus variability, and emotional training stage. Based on the aforementioned core training parameters, a training task is generated; The dominant weight of the training mode is the ratio of analytical training to comprehensive training. Analytical training is a training mode for sub-vocabularies, while comprehensive training is a holistic training mode in real-world contexts. The task complexity reflects the difficulty of the statement length and syntactic structure. The stimulus variability reflects the level of diversity in speaker identity, speech rate, and acoustic environment in the training samples. The emotional training phase reflects the progression path of a user's emotional cognitive abilities.
[0082] Furthermore, the logical instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0083] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the auditory training task generation method provided by the above methods, the method including: Based on user performance data, the core training parameters are dynamically adjusted, including the dominant weight of the training mode, task complexity, stimulus variability, and emotional training stage. Based on the aforementioned core training parameters, a training task is generated; The dominant weight of the training mode is the ratio of analytical training to comprehensive training. Analytical training is a training mode for sub-vocabularies, while comprehensive training is a holistic training mode in real-world contexts. The task complexity reflects the difficulty of the statement length and syntactic structure. The stimulus variability reflects the level of diversity in speaker identity, speech rate, and acoustic environment in the training samples. The emotional training phase reflects the progression path of a user's emotional cognitive abilities.
[0084] Furthermore, this application also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, is implemented to perform the auditory training task generation method provided by the methods described above, the method comprising: Based on user performance data, the core training parameters are dynamically adjusted, including the dominant weight of the training mode, task complexity, stimulus variability, and emotional training stage. Based on the aforementioned core training parameters, a training task is generated; The dominant weight of the training mode is the ratio of analytical training to comprehensive training. Analytical training is a training mode for sub-vocabularies, while comprehensive training is a holistic training mode in real-world contexts. The task complexity reflects the difficulty of the statement length and syntactic structure. The stimulus variability reflects the level of diversity in speaker identity, speech rate, and acoustic environment in the training samples. The emotional training phase reflects the progression path of a user's emotional cognitive abilities.
[0085] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0086] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for generating auditory training tasks, characterized in that, include: Based on user performance data, the core training parameters are dynamically adjusted, including the dominant weight of the training mode, task complexity, stimulus variability, and emotional training stage. Based on the aforementioned core training parameters, a training task is generated; The dominant weight of the training mode is the ratio of analytical training to comprehensive training. Analytical training is a training mode for sub-vocabularies, while comprehensive training is a holistic training mode in real-world contexts. The task complexity reflects the difficulty of the statement length and syntactic structure. The stimulus variability reflects the level of diversity in speaker identity, speech rate, and acoustic environment in the training samples. The emotional training phase reflects the progression path of a user's emotional cognitive abilities.
2. The auditory training task generation method according to claim 1, characterized in that, The training steps for the analytical training are as follows: While maintaining the playback of a specific phoneme audio stream in the audio channel, a video stream with similar pronunciation locations but different acoustic characteristics is superimposed and displayed in the visual channel, thus merging the audio stream and the video stream into a single multimedia stimulus stream; Play the multimedia stimulus stream and obtain the user's identification results; If the identification result indicates that the user has correctly identified the user, then it is determined that the user's auditory channel weight has been increased. If the identification result indicates that the user has made an incorrect identification, a correction process is executed, which is used to guide the user to focus on the auditory channel.
3. The auditory training task generation method according to claim 2, characterized in that, The auditory training task generation method also includes: During training, the video stream played to the user contains visual information about lip shape, teeth, visible parts of the tongue, cheek movements, jaw movements, facial expressions, and head posture.
4. The auditory training task generation method according to claim 2, characterized in that, The task complexity is dynamically adjusted based on the following steps: Based on the task complexity, multimodal contextual cues are provided and corresponding multimedia stimulus streams are played. Based on the user's performance data in response to the multimedia stimulus stream, the user's verbal prediction error and emotional attitude prediction error are calculated. The task complexity is adjusted based on the verbal prediction error and the affective attitude prediction error.
5. The auditory training task generation method according to claim 1, characterized in that, The dominant weights of the training mode are dynamically adjusted based on the following steps: Based on the user's real-time response data and historical error records in the sub-vocabulary recognition task, identify the user's phoneme confusion pattern; After generating adversarial training samples based on the phoneme confusion pattern, training is performed based on the adversarial training samples. If the user's recognition accuracy in the sub-vocabulary recognition task reaches a preset threshold, the proportion of analytical training will be reduced.
6. The auditory training task generation method according to any one of claims 1-5, characterized in that, The auditory training task generation method also includes: After receiving the user's audio and video input, the audio and video input is deconstructed into sublexical unit dimension and sentiment dimension; Identify the confusing phonemes in the sublexical unit dimension and the misjudged emotion categories in the emotion dimension; Based on the confused phonemes, generate instructions for the confused phonemes and compare and play them; Based on the misjudged emotion category, emotional visual cues are generated and the corresponding video clips are played back.
7. An auditory training task generation system, characterized in that, The parameter adjustment module is used to dynamically adjust the core training parameters based on the user's performance data. The core training parameters include the dominant weight of the training mode, task complexity, stimulus variability, and emotional training stage. The task generation module is used to: generate training tasks based on the core training parameters; The dominant weight of the training mode is the ratio of analytical training to comprehensive training. Analytical training is a training mode for sub-vocabularies, while comprehensive training is a holistic training mode in real-world contexts. The task complexity reflects the difficulty of the statement length and syntactic structure. The stimulus variability reflects the level of diversity in speaker identity, speech rate, and acoustic environment in the training samples. The emotional training phase reflects the progression path of a user's emotional cognitive abilities.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the auditory training task generation method as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the auditory training task generation method as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the auditory training task generation method as described in any one of claims 1 to 6.