Rehabilitation training system for hearing disorder patient based on virtual reality
Through multi-stage cross-modal attention-enhanced audio-visual representation method, voiceprint dynamic mapping and cross-modal contrast learning, combined with virtual reality technology, the shortcomings of the existing system in multi-sensory data integration, personalized adaptation and intelligent support are solved, and efficient rehabilitation training for patients with hearing impairments are achieved.
Patent Information
- Application Number
- CN202510461909.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing rehabilitation training system for patients with hearing impairments cannot effectively integrate multi-sensory data, resulting in patients being unable to fully perceive voice information, insufficient data fusion accuracy and real-time performance, lack of personalized adaptability, unable to dynamically adjust training content, and lack of intelligent support in lip training and contextual understanding.
Multi-stage cross-modal attention-enhanced audio-visual representation method is used to fusion of multimodal data, combined with the audio generation dynamic adjustment method of voiceprint dynamic mapping for basic auditory reconstruction, cross-modal comparison learning method is used for lip synchronous training, and dynamic generation training scheme is combined with virtual reality technology.
It realizes high-precision, real-time multimodal data fusion, dynamically generates personalized sound-scene mapping, improves lip-voice alignment and targeted training of language content, and improves the overall effect of rehabilitation training.
Smart Images

Figure CN120015235A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of auditory rehabilitation training, and in particular relates to a rehabilitation training system for hearing-impaired patients based on virtual reality. Background Art
[0002] The virtual reality-based rehabilitation training system for hearing-impaired patients is an intelligent rehabilitation tool that uses virtual reality (VR) technology, combined with multimodal data fusion and artificial intelligence algorithms, to help hearing-impaired patients rebuild their auditory perception ability, improve their speech comprehension ability and enhance their environmental adaptability. The system builds an immersive virtual scene, converts sound signals into visual light waves and tactile vibrations, helps patients perceive sound information through vision and touch, and helps patients better understand and communicate in complex environments. Its role is to improve patients' auditory perception, language comprehension and social adaptability through high-precision, personalized rehabilitation training, and ultimately improve their quality of life.
[0003] However, in the existing rehabilitation training systems for hearing-impaired patients, there is a technical problem that the existing systems cannot effectively integrate multi-sensory data. During the rehabilitation process, it is easy for patients to be unable to fully perceive sound information through alternative senses. At the same time, the accuracy and real-time performance of data fusion are insufficient, and it is impossible to provide patients with an immersive rehabilitation training experience, and it is difficult to effectively combine virtual reality technology for integrated application. In the existing auditory perception reconstruction process, there is a technical problem that the existing systems lack personalized adaptation capabilities in basic auditory perception reconstruction, and cannot dynamically adjust the training content according to the patient's hearing loss degree and rehabilitation progress. At the same time, the accuracy of sound feature extraction and mapping is also insufficient, resulting in There is a technical problem that patients' ability to perceive sound frequency, direction and intensity is limited. In the process of integrating existing language content, the existing system lacks intelligent support in lip reading training and contextual understanding training, and is unable to effectively identify and complete the phonemes, context and lip movement information that are easily confused by patients, affecting the patient's learning effect on the speech-lip shape association. In the existing rehabilitation training system for hearing-impaired patients, there is a technical problem that the existing rehabilitation training system lacks interactivity and intelligence, and the scene construction of the existing rehabilitation system supported by virtual reality relies on preset content, and is unable to dynamically generate training plans according to patient needs, resulting in poor overall rehabilitation effects. Summary of the invention
[0004] In view of the above situation, in order to overcome the defects of the prior art, the present invention provides a rehabilitation training system for hearing-impaired patients based on virtual reality. In view of the technical problems that the existing rehabilitation training system for hearing-impaired patients cannot effectively integrate multi-sensory data, and during the rehabilitation process, it is easy for patients to be unable to fully perceive sound information through alternative senses. At the same time, the accuracy and real-time performance of data fusion are insufficient, and it is impossible to provide patients with an immersive rehabilitation training experience, and it is difficult to effectively combine virtual reality technology for integrated application, this solution creatively adopts a multi-stage cross-modal attention enhanced speech and audio-visual representation method, and converts sound signals into visualized light waves and tactile vibration data through a deep learning model to achieve high-precision and real-time multi-modal data fusion; in view of the existing auditory perception reconstruction process, the existing system lacks personalized adaptation capabilities in basic auditory perception reconstruction, and cannot dynamically adjust the training content according to the patient's hearing loss degree and rehabilitation progress. At the same time, the accuracy of sound feature extraction and mapping is also insufficient, resulting in limited improvement in the patient's perception of sound frequency, direction and intensity. This solution creatively adopts a dynamic adjustment method for audio generation combined with dynamic mapping of voiceprints to perform basic auditory reconstruction and effectively dynamically generate personalized soundscapes. Mapping helps patients rebuild their basic perception of sound; in view of the technical problem that in the process of integrating existing language content, the existing system lacks intelligent support in lip reading training and context understanding training, and cannot effectively identify and complete the phonemes, context and lip movement information that are easily confused by patients, which affects the patient's learning effect on speech-lip shape association, this solution creatively adopts cross-modal comparative learning method to assist lip reading synchronization training, and provides human voice context background prompts during rehabilitation training to assist patients in human voice and speech understanding, thus achieving high-precision lip reading-speech alignment, easily confused phonemes and targeted training of language content; In the existing rehabilitation training systems for hearing-impaired patients, there is a technical problem that the existing rehabilitation training systems lack interactivity and intelligence, and the scene construction of the existing rehabilitation systems supported by virtual reality relies on preset content and cannot dynamically generate training plans according to patient needs, resulting in poor overall rehabilitation effects. This solution creatively adopts the integration of virtual reality technology, auditory reconstruction and language content understanding, and integrates the rehabilitation training system for hearing-impaired patients based on virtual reality, which improves the overall effect of rehabilitation training and provides strong support and exploration experience for the intelligent and automated rehabilitation of hearing-impaired patients.
[0005] The technical solution adopted by the present invention is as follows: the virtual reality-based rehabilitation training system for hearing-impaired patients provided by the present invention includes a perception fusion module, a basic perception reconstruction module, a language content integration module, a rehabilitation training module and a virtual reality auxiliary module;
[0006] The perception fusion module is used for multimodal data collection and processing, obtains a hearing impairment rehabilitation training fusion data set through multimodal data collection and processing, and sends the hearing impairment rehabilitation training fusion data set to the basic perception reconstruction module and the language content integration module;
[0007] The basic perception reconstruction module is used for basic auditory perception reconstruction, obtains basic sound feature perception optimization data through basic auditory perception reconstruction, and sends the basic sound feature perception optimization data to the rehabilitation training module;
[0008] The language content integration module is used for lip synchronization training and context comprehension training, obtains language content comprehension rehabilitation optimization data through lip synchronization training and context comprehension training, and sends the language content comprehension rehabilitation optimization data to the rehabilitation training module;
[0009] The rehabilitation training module is used for comprehensive rehabilitation training. Through comprehensive rehabilitation training, comprehensive rehabilitation training reference evaluation data is obtained, and combined with the virtual reality auxiliary module, rehabilitation training for hearing-impaired patients is carried out;
[0010] The virtual reality auxiliary module is used to construct a virtual reality scene and assist rehabilitation training, and conduct rehabilitation training with the assistance of virtual reality technology.
[0011] Furthermore, the multimodal data collection and processing is used to collect the original data set required for virtual reality rehabilitation training, specifically, to obtain the original data set for hearing impairment rehabilitation training through multimodal data collection, and to obtain the fused data set for hearing impairment rehabilitation training through data fusion;
[0012] The multimodal data collection collects environmental sound, user line of sight direction and body posture data in real time through microphone arrays, spatial audio sensors, eye tracking devices and tactile feedback devices;
[0013] The environmental sound includes audio data and sound spectrogram data; the audio data includes human voice data, human voice video data and environmental noise data;
[0014] The data fusion extracts sound signal features by combining a multi-stage cross-modal attention enhanced speech and audio-visual representation method, performs feature cross-modal mapping based on the environmental sound, obtains sound signal feature data and maps them to obtain visual spectrogram data and tactile vibration data, and obtains a hearing impairment rehabilitation training fusion data set through feature cross-modal mapping, including the following steps: multi-stage cross-modal feature extraction, multi-scale feature fusion optimization, cross-modal attention optimization, cross-modal feature update, cross-modal feature loss construction, and feature cross-modal mapping;
[0015] The multi-stage cross-modal feature extraction is used to construct a basic problem model of the speech audio-visual representation method, specifically defining audio-visual instance feature vector pairs, and setting each pair of audio-visual instance feature vector pairs as a one-hot category vector, performing multi-stage cross-modal feature extraction problem modeling, and obtaining audio-visual instance feature vector data;
[0016] The multi-scale feature fusion optimization is used for extracting sound and visual features by linear projection and convolution, specifically, the sound feature data and the visual feature data are subjected to feature extraction and feature fusion respectively through one-dimensional convolution operation to obtain multi-scale feature fusion optimization data;
[0017] The cross-modal attention optimization is used to optimize the attention information of the sound features and the visual features and calculate the attention weights, specifically, performing cross-modal attention optimization based on the multi-scale feature fusion optimization data to obtain cross-modal attention weight data;
[0018] The cross-modal feature updating is specifically to perform cross-modal feature updating on the multi-scale feature fusion optimization data according to the cross-modal attention weight data to obtain updated attention feature data;
[0019] The cross-modal feature loss construction is used to optimize the model training process, specifically constructing the discriminant loss, the intra-modal loss and the cross-modal loss in sequence, and performing loss integration to obtain the cross-modal feature loss function;
[0020] The calculation formula of the cross-modal feature loss function is:
[0021] L=L dis +a·L intra +b·L cross ;
[0022] Where L is the cross-modal feature loss function, L dis is the discrimination loss, L intra is the intra-modal loss, L cross is the cross-modal loss, a is the intra-modal loss weight, and b is the cross-modal loss weight;
[0023] The feature cross-modal mapping is specifically to perform feature cross-modal mapping model training through the multi-stage cross-modal feature extraction, the multi-scale feature fusion optimization, the cross-modal attention optimization, the cross-modal feature update and the cross-modal feature loss construction to obtain a sound-visual feature cross-modal mapping model, and to perform feature cross-modal mapping based on the environmental sound by using the sound-visual feature cross-modal mapping model to obtain sound signal feature data and visualized spectrum data, and to generate tactile vibration simulation data based on the visualized spectrum data to obtain a hearing impairment rehabilitation training fusion data set;
[0024] The hearing impairment rehabilitation training fusion data set specifically includes human voice feature data, human voice video feature data, environmental noise feature data, visualized spectrum data and tactile vibration simulation data;
[0025] The hearing impairment rehabilitation training fusion data set is combined with a virtual reality scene to form virtual reality auxiliary data, and the virtual reality auxiliary data is used for comprehensive rehabilitation training; the virtual reality auxiliary data includes sound source location, sound intensity, sound visualization auxiliary information and sound tactile auxiliary information.
[0026] Furthermore, the basic auditory perception reconstruction is used to assist hearing-impaired patients in reconstructing their basic perception ability of sound. Specifically, based on the hearing-impaired rehabilitation training fusion data set, a method for dynamically adjusting the audio frequency generation combined with dynamic voiceprint mapping is used to perform basic auditory reconstruction to obtain basic sound feature perception optimization data, including the following steps: dynamic voiceprint mapping, eye tracking auxiliary guidance, adaptive sound field training optimization and basic auditory perception reconstruction;
[0027] The dynamic voiceprint mapping is specifically to use an audio adversarial generation network combined with sound visualization auxiliary information to generate dynamic light waves of sound signals to obtain visible light wave data of generated sound signals;
[0028] The audio adversarial generative network combined with the sound visualization auxiliary information adopts standard generator training and discriminator training to perform generative adversarial training, and constructs a comprehensive loss function of adversarial loss and spectrum reconstruction loss to perform generative adversarial training to obtain the generated sound signal visible light wave data;
[0029] The calculation formula of the comprehensive loss function is:
[0030] ;
[0031] Where, L total is the comprehensive loss function, L GAN is the adversarial loss function, G is the generator identifier, D is the discriminator identifier, r is the spectral reconstruction loss weight, S REAL is the visible light wave data corresponding to the real sound spectrum, S GEN It is the visible light wave data corresponding to the generated sound spectrum;
[0032] The eye tracking auxiliary guidance specifically collects the user's eye tracking data as auxiliary data, and dynamically adjusts the brightness of the sound source light wave of the generated sound signal visual light wave data through the spatial attention mechanism, performs sight line guidance enhancement training, and obtains sight line guidance optimization data;
[0033] The adaptive sound field training optimization is specifically to use a standard reinforcement learning method to perform the rehabilitation training difficulty of the basic auditory perception reconstruction training stage according to the generated sound signal visible light wave data and the sight guidance optimization data, and to perform the adaptive sound field training optimization by constructing state parameters, action parameters and reward functions to obtain adaptive training adjustment reference strategy data;
[0034] The basic auditory perception reconstruction is specifically to perform basic auditory perception reconstruction through the voiceprint dynamic mapping, eye tracking auxiliary guidance and adaptive sound field training optimization, and to perform basic auditory perception reconstruction rehabilitation training in combination with the virtual reality auxiliary data to obtain basic sound feature perception optimization data.
[0035] Furthermore, the lip synchronization training is used to assist hearing-impaired patients in training their ability to recognize the association between lip shape and speech, specifically, based on the hearing-impaired rehabilitation training fusion data set, a cross-modal contrastive learning method is used to assist in lip synchronization training to obtain language comprehension rehabilitation optimization data;
[0036] The cross-modal contrastive learning method specifically uses the human voice feature data and human voice video feature data in the hearing impairment rehabilitation training fusion data set as input data, and constructs a standard three-dimensional convolutional network as a lip reading feature encoder, and constructs a standard Wav2Vec2.0 model as a speech feature encoder, and performs cross-modal contrastive learning by extracting lip movement features and speech features, and constructs a contrast loss function to maximize the similarity of lip movement features and speech features to obtain lip movement speech mapping data, and combines the virtual reality auxiliary data to perform lip reading synchronization rehabilitation training to obtain language comprehension rehabilitation optimization data;
[0037] The calculation formula of the contrast loss function is:
[0038] ;
[0039] Where, L contrastive is the contrast loss function, exp(·) is the natural base function, and f lip It is the lip movement feature. is a positive sample in the speech feature, which is used to represent the speech feature that matches the lip movement feature. is the nth speech feature, m is the total number of feature pairs of lip movement features and speech features, n is the feature pair index, is the characteristic pair correction factor.
[0040] Furthermore, the contextual comprehension training is used to assist hearing-impaired patients in understanding contextual information in virtual scenes. Specifically, the virtual reality auxiliary data is combined to provide vocal context background prompts during rehabilitation training to assist patients in understanding vocal speech and obtain language content comprehension rehabilitation optimization data.
[0041] Furthermore, the comprehensive rehabilitation training is used to carry out rehabilitation training in combination with the results of virtual visual environment, auditory perception reconstruction and language content comprehension training. Specifically, virtual reality scene simulation is carried out based on the virtual reality auxiliary data, and real-time gesture and posture monitoring of the patient is combined. The patient's auditory perception during rehabilitation training is enhanced by the basic sound feature perception optimization data, and the patient's language content perception during rehabilitation training is enhanced by the language content comprehension rehabilitation optimization data, so as to obtain comprehensive rehabilitation training reference evaluation data.
[0042] The beneficial effects achieved by the present invention using the above scheme are as follows:
[0043] (1) In view of the technical problems that the existing rehabilitation training system for hearing-impaired patients cannot effectively integrate multi-sensory data, which easily leads to patients being unable to fully perceive sound information through alternative senses during the rehabilitation process. At the same time, the accuracy and real-time performance of data fusion are insufficient, which makes it impossible to provide patients with an immersive rehabilitation training experience and difficult to effectively combine with virtual reality technology for integrated application, this solution creatively adopts a multi-stage cross-modal attention enhanced audio-visual representation method, and converts sound signals into visualized light waves and tactile vibration data through a deep learning model, thus achieving high-precision and real-time multi-modal data fusion;
[0044] (2) In view of the technical problems that the existing system lacks personalized adaptation capabilities in basic auditory perception reconstruction and cannot dynamically adjust the training content according to the patient's hearing loss degree and rehabilitation progress, and the accuracy of sound feature extraction and mapping is also insufficient, resulting in limited improvement in patients' perception of sound frequency, direction and intensity, this solution creatively adopts a dynamic adjustment method of audio generation combined with dynamic voiceprint mapping to perform basic auditory reconstruction, effectively dynamically generate personalized soundscape mapping, and help patients rebuild their basic perception of sound;
[0045] (3) In view of the technical problem that the existing system lacks intelligent support in lip reading training and context understanding training in the process of integrating existing language content, and cannot effectively identify and complete the phonemes, context and lip movement information that are easily confused by patients, which affects the patient's learning effect on the speech-lip shape association, this program creatively adopts a cross-modal comparative learning method to assist in lip reading synchronization training, and provides human voice context background prompts during rehabilitation training to assist patients in human voice and speech understanding, thus achieving high-precision lip reading-speech alignment, and targeted training of easily confused phonemes and language content;
[0046] (4) In view of the technical problem that the existing rehabilitation training system for hearing-impaired patients lacks interactivity and intelligence, and the scene construction of the existing rehabilitation system supported by virtual reality relies on preset content and cannot dynamically generate training plans according to patient needs, resulting in poor overall rehabilitation effects, this solution creatively adopts the integration of virtual reality technology, auditory reconstruction and language content understanding, and integrates the rehabilitation training system for hearing-impaired patients based on virtual reality, which improves the overall effect of rehabilitation training and provides strong support and exploration experience for the intelligent and automated rehabilitation of hearing-impaired patients. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 A schematic diagram of the structure of a virtual reality-based rehabilitation training system for hearing-impaired patients provided by the present invention;
[0048] Figure 2 A flowchart of the steps performed by the system is provided for the present invention;
[0049] Figure 3 A flowchart of the steps performed for data fusion in the perception fusion module;
[0050] Figure 4 Flowchart of the steps performed by the base perception reconstruction module.
[0051] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention. DETAILED DESCRIPTION
[0052] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0053] Example 1, see Figure 1 , the technical solution adopted by the present invention is as follows: the rehabilitation training system for hearing-impaired patients based on virtual reality provided by the present invention includes a perception fusion module, a basic perception reconstruction module, a language content integration module, a rehabilitation training module and a virtual reality auxiliary module;
[0054] The perception fusion module is used for multimodal data collection and processing, obtains a hearing impairment rehabilitation training fusion data set through multimodal data collection and processing, and sends the hearing impairment rehabilitation training fusion data set to the basic perception reconstruction module and the language content integration module;
[0055] The basic perception reconstruction module is used for basic auditory perception reconstruction, obtains basic sound feature perception optimization data through basic auditory perception reconstruction, and sends the basic sound feature perception optimization data to the rehabilitation training module;
[0056] The language content integration module is used for lip synchronization training and context comprehension training, obtains language content comprehension rehabilitation optimization data through lip synchronization training and context comprehension training, and sends the language content comprehension rehabilitation optimization data to the rehabilitation training module;
[0057] The rehabilitation training module is used for comprehensive rehabilitation training. Through comprehensive rehabilitation training, comprehensive rehabilitation training reference evaluation data is obtained, and combined with the virtual reality auxiliary module, rehabilitation training for hearing-impaired patients is carried out;
[0058] The virtual reality auxiliary module is used to construct a virtual reality scene and assist rehabilitation training, and conduct rehabilitation training with the assistance of virtual reality technology.
[0059] By performing the above operations, the existing rehabilitation training system for hearing-impaired patients lacks interactivity and intelligence, and the scene construction of the existing rehabilitation system supported by virtual reality relies on preset content, and cannot dynamically generate training plans according to patient needs, resulting in poor overall rehabilitation effects. This solution creatively adopts the integration of virtual reality technology, auditory reconstruction and language content understanding, and integrates the rehabilitation training system for hearing-impaired patients based on virtual reality, which improves the overall effect of rehabilitation training and provides strong support and exploration experience for the intelligent and automated rehabilitation of hearing-impaired patients.
[0060] Embodiment 2, this embodiment is based on the above embodiment, see Figure 1 , Figure 2 and Figure 3The multimodal data collection and processing is used to collect the original data set required for virtual reality rehabilitation training, specifically, to obtain the original data set for hearing impairment rehabilitation training through multimodal data collection, and to obtain the fused data set for hearing impairment rehabilitation training through data fusion;
[0061] The multimodal data collection collects environmental sound, user line of sight direction and body posture data in real time through microphone arrays, spatial audio sensors, eye tracking devices and tactile feedback devices;
[0062] The environmental sound includes audio data and sound spectrogram data; the audio data includes human voice data, human voice video data and environmental noise data;
[0063] The data fusion extracts sound signal features by combining a multi-stage cross-modal attention enhanced speech and audio-visual representation method, performs feature cross-modal mapping based on the environmental sound, obtains sound signal feature data and maps them to obtain visual spectrogram data and tactile vibration data, and obtains a hearing impairment rehabilitation training fusion data set through feature cross-modal mapping, including the following steps: multi-stage cross-modal feature extraction, multi-scale feature fusion optimization, cross-modal attention optimization, cross-modal feature update, cross-modal feature loss construction, and feature cross-modal mapping;
[0064] The multi-stage cross-modal feature extraction is used to construct a basic problem model of the speech audio-visual representation method, specifically defining audio-visual instance feature vector pairs, and setting each pair of audio-visual instance feature vector pairs as a unique hot category vector, performing multi-stage cross-modal feature extraction problem modeling, and obtaining audio-visual instance feature vector data. The calculation formula is:
[0065] ;
[0066] Where D is the feature vector data of the audio-visual instance, is sound feature vector data, wherein the sound feature is used to represent the characteristics of the original sound data, is visual feature vector data, the visual feature is used to represent the waveform visualization data corresponding to the original sound, N is the total number of audio-visual instances, and i is the audio-visual instance index;
[0067] The multi-scale feature fusion optimization is used for extracting sound and visual features by linear projection and convolution. Specifically, the sound feature data and the visual feature data are subjected to feature extraction and feature fusion respectively through one-dimensional convolution operation to obtain multi-scale feature fusion optimization data. The calculation formula is:
[0068] ;
[0069] In the formula, is the convolution sound feature data, k is the convolution kernel size, Conv1D(·) is the one-dimensional convolution operation function, F a is the sound feature data, is the sound feature extraction parameter, is the convolutional visual feature data, F v is the visual feature data, is the visual feature extraction parameter, M is the multi-scale feature fusion optimization data, W M is the feature fusion weight;
[0070] Preferably, the value range of the convolution kernel size is specifically {3, 5, 7};
[0071] The cross-modal attention optimization is used to optimize the attention information of the sound features and the visual features and calculate the attention weight. Specifically, the cross-modal attention optimization is performed based on the multi-scale feature fusion optimization data to obtain the cross-modal attention weight data. The calculation formula is:
[0072] ;
[0073] Where A is the cross-modal attention weight data, including sound attention weight and visual attention weight, is the sound attention weight, k is the convolution kernel size, is the visual attention weight, softmax(·) is the classifier function, Q a is the sound query matrix, K a is the sound key matrix, V a is the sound value matrix, M is the multi-scale feature fusion optimization data, T is the transposition operator, is the dimension value of the sound key matrix, Q v is the visual query matrix, K v is the visual key matrix, V v is the visual value matrix, is the dimension value of the visual key matrix;
[0074] The cross-modal feature updating is specifically to perform cross-modal feature updating on the multi-scale feature fusion optimization data according to the cross-modal attention weight data to obtain updated attention feature data;
[0075] The calculation formula for the cross-modal feature update is:
[0076] ;
[0077] In the formula, is to update the sound feature data, is to update the visual feature data;
[0078] The calculation formula for updating the attention feature data is:
[0079] ;
[0080] Where F is the updated attention feature data, including the updated attention integrated sound feature data and the updated attention integrated visual feature data. is to update the attention and integrate the sound feature data, is to update the attention and integrate the visual feature data, It is the updated sound feature data obtained by performing feature fusion optimization using a one-dimensional convolution kernel with a convolution kernel size of 3. It is the updated visual feature data obtained by performing feature fusion optimization using a one-dimensional convolution kernel with a convolution kernel size of 3; It is the updated sound feature data obtained by performing feature fusion optimization using a one-dimensional convolution kernel with a convolution kernel size of 5. It is the updated visual feature data obtained by performing feature fusion optimization using a one-dimensional convolution kernel with a convolution kernel size of 5; It is the updated sound feature data obtained by performing feature fusion optimization using a one-dimensional convolution kernel with a convolution kernel size of 7. It is the updated visual feature data obtained by performing feature fusion optimization using a one-dimensional convolution kernel with a convolution kernel size of 7;
[0081] The cross-modal feature loss construction is used to optimize the model training process, specifically constructing the discriminant loss, the intra-modal loss and the cross-modal loss in sequence, and performing loss integration to obtain the cross-modal feature loss function;
[0082] The calculation formula of the discrimination loss is:
[0083] ;
[0084] Where, L dis is the discrimination loss, is the sound feature discrimination loss, specifically using the cross entropy loss function. It is the visual feature discrimination loss, specifically using the cross entropy loss function;
[0085] The calculation formula of the intra-modal loss is:
[0086] ;
[0087] Where, L intra is the intra-modal loss, N is the total number of audio-visual instances, i is the audio-visual instance index, is the sound feature value corresponding to the i-th audiovisual instance, is the central value of the sound feature corresponding to the i-th audiovisual instance, is the visual feature value corresponding to the i-th audiovisual instance, is the central value of the visual feature corresponding to the ith audiovisual instance;
[0088] The calculation formula of the cross-modal loss is:
[0089] ;
[0090] Where, L cross is the cross-modal loss, N is the total number of audio-visual instances, i is the audio-visual instance index, e is the natural base, is the cosine similarity value of the positive sample, is the cosine similarity value of the negative sample;
[0091] The calculation formula of the cross-modal feature loss function is:
[0092] L=L dis +a·L intra +b·L cross ;
[0093] Where L is the cross-modal feature loss function, L dis is the discrimination loss, L intra is the intra-modal loss, L cross is the cross-modal loss, a is the intra-modal loss weight, and b is the cross-modal loss weight;
[0094] The feature cross-modal mapping is specifically to perform feature cross-modal mapping model training through the multi-stage cross-modal feature extraction, the multi-scale feature fusion optimization, the cross-modal attention optimization, the cross-modal feature update and the cross-modal feature loss construction to obtain a sound-visual feature cross-modal mapping model, and to perform feature cross-modal mapping based on the environmental sound by using the sound-visual feature cross-modal mapping model to obtain sound signal feature data and visualized spectrum data, and to generate tactile vibration simulation data based on the visualized spectrum data to obtain a hearing impairment rehabilitation training fusion data set;
[0095] The hearing impairment rehabilitation training fusion data set specifically includes human voice feature data, human voice video feature data, environmental noise feature data, visualized spectrum data and tactile vibration simulation data;
[0096] The hearing impairment rehabilitation training fusion data set is combined with a virtual reality scene to form virtual reality auxiliary data, and the virtual reality auxiliary data is used for comprehensive rehabilitation training; the virtual reality auxiliary data includes sound source location, sound intensity, sound visualization auxiliary information and sound tactile auxiliary information.
[0097] By performing the above operations, the existing rehabilitation training system for hearing-impaired patients cannot effectively integrate multi-sensory data. During the rehabilitation process, it is easy for patients to be unable to fully perceive sound information through alternative senses. At the same time, the accuracy and real-time performance of data fusion are insufficient, and it is impossible to provide patients with an immersive rehabilitation training experience. It is difficult to effectively combine virtual reality technology for integrated application. This solution creatively adopts a multi-stage cross-modal attention enhanced audio-visual representation method for speech, and converts sound signals into visualized light waves and tactile vibration data through a deep learning model to achieve high-precision and real-time multi-modal data fusion.
[0098] Embodiment 3, this embodiment is based on the above embodiment, see Figure 1 , Figure 2 and Figure 4 The basic auditory perception reconstruction is used to assist hearing-impaired patients in reconstructing their basic perception ability of sound. Specifically, based on the hearing-impaired rehabilitation training fusion data set, a method for dynamically adjusting the audio frequency generation combined with dynamic voiceprint mapping is used to perform basic auditory reconstruction to obtain basic sound feature perception optimization data, including the following steps: dynamic voiceprint mapping, eye tracking auxiliary guidance, adaptive sound field training optimization and basic auditory perception reconstruction;
[0099] The dynamic voiceprint mapping is specifically to use an audio adversarial generation network combined with sound visualization auxiliary information to generate dynamic light waves of sound signals to obtain visible light wave data of generated sound signals;
[0100] The audio adversarial generative network combined with the sound visualization auxiliary information adopts standard generator training and discriminator training to perform generative adversarial training, and constructs a comprehensive loss function of adversarial loss and spectrum reconstruction loss to perform generative adversarial training to obtain the generated sound signal visible light wave data;
[0101] The calculation formula of the comprehensive loss function is:
[0102] ;
[0103] Where, L total is the comprehensive loss function, L GAN is the adversarial loss function, G is the generator identifier, D is the discriminator identifier, r is the spectral reconstruction loss weight, S REAL is the visible light wave data corresponding to the real sound spectrum, S GEN It is the visible light wave data corresponding to the generated sound spectrum;
[0104] The eye tracking auxiliary guidance specifically collects the user's eye tracking data as auxiliary data, and dynamically adjusts the brightness of the sound source light wave of the generated sound signal visual light wave data through the spatial attention mechanism, performs sight line guidance enhancement training, and obtains sight line guidance optimization data;
[0105] The adaptive sound field training optimization is specifically to use a standard reinforcement learning method to perform the rehabilitation training difficulty of the basic auditory perception reconstruction training stage according to the generated sound signal visible light wave data and the sight guidance optimization data, and to perform the adaptive sound field training optimization by constructing state parameters, action parameters and reward functions to obtain adaptive training adjustment reference strategy data;
[0106] The state parameters include the user's reaction time and accuracy to the sound source positioning;
[0107] The action parameters include increasing the interfering sound source, reducing the interfering sound source, and adjusting the distance between the sound source and the patient;
[0108] The reward function specifically refers to the positive feedback reward for the user's positioning accuracy;
[0109] The basic auditory perception reconstruction is specifically to perform basic auditory perception reconstruction through the voiceprint dynamic mapping, eye tracking auxiliary guidance and adaptive sound field training optimization, and to perform basic auditory perception reconstruction rehabilitation training in combination with the virtual reality auxiliary data to obtain basic sound feature perception optimization data.
[0110] By performing the above operations, in view of the technical problems that in the existing auditory perception reconstruction process, the existing system lacks personalized adaptation capabilities in basic auditory perception reconstruction, and cannot dynamically adjust the training content according to the patient's hearing loss degree and rehabilitation progress. At the same time, the accuracy of sound feature extraction and mapping is also insufficient, resulting in limited improvement in patients' perception of sound frequency, direction and intensity. This solution creatively adopts a dynamic adjustment method of audio generation combined with dynamic mapping of voiceprints to perform basic auditory reconstruction, effectively and dynamically generate personalized soundscape mapping, and help patients rebuild their basic perception of sound.
[0111] Embodiment 4, this embodiment is based on the above embodiment, see Figure 1 and Figure 2 The lip synchronization training is used to assist hearing-impaired patients in training their ability to recognize the association between lip shape and speech. Specifically, based on the hearing-impaired rehabilitation training fusion data set, a cross-modal contrastive learning method is used to assist in lip synchronization training to obtain language comprehension rehabilitation optimization data;
[0112] The cross-modal contrastive learning method specifically uses the human voice feature data and human voice video feature data in the hearing impairment rehabilitation training fusion data set as input data, and constructs a standard three-dimensional convolutional network as a lip reading feature encoder, and constructs a standard Wav2Vec2.0 model as a speech feature encoder, and performs cross-modal contrastive learning by extracting lip movement features and speech features, and constructs a contrast loss function to maximize the similarity of lip movement features and speech features to obtain lip movement speech mapping data, and combines the virtual reality auxiliary data to perform lip reading synchronization rehabilitation training to obtain language comprehension rehabilitation optimization data;
[0113] The calculation formula of the contrast loss function is:
[0114] ;
[0115] Where, L contrastive is the contrast loss function, exp(·) is the natural base function, and f lip It is the lip movement feature. is a positive sample in the speech feature, which is used to represent the speech feature that matches the lip movement feature. is the nth speech feature, m is the total number of feature pairs of lip movement features and speech features, n is the feature pair index, is the characteristic pair correction factor.
[0116] Embodiment 5, this embodiment is based on the above embodiment, see Figure 1 and Figure 2 The context comprehension training is used to assist hearing-impaired patients in understanding context information in virtual scenes. Specifically, it combines the virtual reality auxiliary data to provide vocal context background prompts during rehabilitation training to assist patients in understanding vocal speech and obtain language content comprehension rehabilitation optimization data.
[0117] By performing the above operations, in order to address the technical problem that in the process of integrating existing language content, the existing system lacks intelligent support in lip reading training and context understanding training, and is unable to effectively identify and complete the patient's easily confused phonemes, context and lip movement information, which affects the patient's learning effect on speech-lip shape association, this solution creatively adopts a cross-modal comparative learning method to assist in lip reading synchronization training, and provides vocal context background prompts during rehabilitation training to assist patients in vocal speech understanding, thereby achieving high-precision lip reading-speech alignment, easily confused phonemes and targeted training of language content.
[0118] Embodiment 6, this embodiment is based on the above embodiment, see Figure 1The comprehensive rehabilitation training is used to carry out rehabilitation training in combination with the results of virtual visual environment, auditory perception reconstruction and language content comprehension training. Specifically, virtual reality scene simulation is carried out based on the virtual reality auxiliary data, and real-time gesture and posture monitoring of the patient is combined. The auditory perception of the patient during the rehabilitation training is enhanced by the basic sound feature perception optimization data, and the language content perception of the patient during the rehabilitation training is enhanced by the language content comprehension rehabilitation optimization data, so as to obtain comprehensive rehabilitation training reference evaluation data.
[0119] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process and method including a series of elements include not only those elements, but also other elements not explicitly listed, or also include elements inherent to such process and method.
[0120] While the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that many changes, modifications, substitutions and variations can be made to the embodiments without departing from the principles and spirit of the invention.
[0121] The present invention and its embodiments are described above, and such description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if ordinary technicians in the field are inspired by it, without departing from the purpose of the invention, they can design a structure and embodiment similar to the technical solution without creativity, which should belong to the protection scope of the present invention.
Claims
1. A rehabilitation training system for hearing-impaired patients based on virtual reality, characterized by: It includes perception fusion module, basic perception reconstruction module, language content integration module, rehabilitation training module and virtual reality assistance module; The perception fusion module is used for multimodal data collection and processing, and obtains a hearing impairment rehabilitation training fusion data set through multimodal data collection and processing and data fusion, and sends the hearing impairment rehabilitation training fusion data set to the basic perception reconstruction module and the language content integration module; The data fusion comprises the following steps: multi-stage cross-modal feature extraction, multi-scale feature fusion optimization, cross-modal attention optimization, cross-modal feature update, cross-modal feature loss construction and feature cross-modal mapping; The basic perception reconstruction module is used for basic auditory perception reconstruction, obtains basic sound feature perception optimization data through basic auditory perception reconstruction, and sends the basic sound feature perception optimization data to the rehabilitation training module; The basic auditory perception reconstruction includes the following steps: voiceprint dynamic mapping, eye tracking auxiliary guidance, adaptive sound field training optimization and basic auditory perception reconstruction; The language content integration module is used for lip synchronization training and context comprehension training, obtains language content comprehension rehabilitation optimization data through lip synchronization training and context comprehension training, and sends the language content comprehension rehabilitation optimization data to the rehabilitation training module; The rehabilitation training module is used for comprehensive rehabilitation training. Through comprehensive rehabilitation training, comprehensive rehabilitation training reference evaluation data is obtained, and combined with the virtual reality auxiliary module, rehabilitation training for hearing-impaired patients is carried out; The virtual reality auxiliary module is used to construct a virtual reality scene and assist rehabilitation training, and conduct rehabilitation training with the assistance of virtual reality technology.
2. The virtual reality-based rehabilitation training system for hearing-impaired patients according to claim 1, characterized in that: The multimodal data collection and processing is used to collect the original data set required for virtual reality rehabilitation training, specifically, to obtain the hearing impairment rehabilitation training original data set through multimodal data collection, and to obtain the hearing impairment rehabilitation training fusion data set through data fusion; The multimodal data collection collects environmental sound, user line of sight direction and body posture data in real time through microphone arrays, spatial audio sensors, eye tracking devices and tactile feedback devices; The environmental sound includes audio data and sound spectrogram data; the audio data includes human voice data, human voice video data and environmental noise data; The data fusion extracts sound signal features by combining a multi-stage cross-modal attention enhanced speech and audio-visual representation method, performs feature cross-modal mapping based on the environmental sound, obtains sound signal feature data and maps them to obtain visual spectrogram data and tactile vibration data, and obtains a hearing impairment rehabilitation training fusion data set through feature cross-modal mapping, including the following steps: multi-stage cross-modal feature extraction, multi-scale feature fusion optimization, cross-modal attention optimization, cross-modal feature update, cross-modal feature loss construction, and feature cross-modal mapping; The hearing impairment rehabilitation training fusion data set is combined with a virtual reality scene to form virtual reality auxiliary data, and the virtual reality auxiliary data is used for comprehensive rehabilitation training; the virtual reality auxiliary data includes sound source location, sound intensity, sound visualization auxiliary information and sound tactile auxiliary information.
3. The virtual reality-based rehabilitation training system for hearing-impaired patients according to claim 2, characterized in that: The multi-stage cross-modal feature extraction is used to construct a basic problem model of the speech audio-visual representation method, specifically defining audio-visual instance feature vector pairs, and setting each pair of audio-visual instance feature vector pairs as a one-hot category vector, performing multi-stage cross-modal feature extraction problem modeling, and obtaining audio-visual instance feature vector data; The multi-scale feature fusion optimization is used for extracting sound and visual features by linear projection and convolution, specifically, the sound feature data and the visual feature data are subjected to feature extraction and feature fusion respectively through one-dimensional convolution operation to obtain multi-scale feature fusion optimization data; The cross-modal attention optimization is used to optimize the attention information of the sound features and the visual features and calculate the attention weights, specifically, performing cross-modal attention optimization based on the multi-scale feature fusion optimization data to obtain cross-modal attention weight data; The cross-modal feature updating is specifically to perform cross-modal feature updating on the multi-scale feature fusion optimization data according to the cross-modal attention weight data to obtain updated attention feature data; The cross-modal feature loss construction is used to optimize the model training process, specifically constructing the discriminant loss, the intra-modal loss and the cross-modal loss in sequence, and performing loss integration to obtain the cross-modal feature loss function; The calculation formula of the cross-modal feature loss function is: L=L dis +a·L intra +b·L cross ; Where L is the cross-modal feature loss function, L dis is the discrimination loss, L intra is the intra-modal loss, L cross is the cross-modal loss, a is the intra-modal loss weight, and b is the cross-modal loss weight; The feature cross-modal mapping is specifically to perform feature cross-modal mapping model training through the multi-stage cross-modal feature extraction, the multi-scale feature fusion optimization, the cross-modal attention optimization, the cross-modal feature update and the cross-modal feature loss construction to obtain a sound-visual feature cross-modal mapping model, and to perform feature cross-modal mapping based on the environmental sound by using the sound-visual feature cross-modal mapping model to obtain sound signal feature data and visualized spectrum data, and to generate tactile vibration simulation data based on the visualized spectrum data to obtain a hearing impairment rehabilitation training fusion data set; The hearing impairment rehabilitation training fusion data set specifically includes human voice feature data, human voice video feature data, environmental noise feature data, visualized spectrum data and tactile vibration simulation data.
4. The virtual reality-based rehabilitation training system for hearing-impaired patients according to claim 3, characterized in that: The basic auditory perception reconstruction is used to assist hearing-impaired patients in reconstructing their basic perception ability of sound. Specifically, based on the hearing-impairment rehabilitation training fusion data set, a dynamic adjustment method of audio generation combined with voiceprint dynamic mapping is used to perform basic auditory reconstruction to obtain basic sound feature perception optimization data, including the following steps: voiceprint dynamic mapping, eye tracking auxiliary guidance, adaptive sound field training optimization and basic auditory perception reconstruction.
5. The virtual reality-based rehabilitation training system for hearing-impaired patients according to claim 4, characterized in that: The dynamic voiceprint mapping is specifically to use an audio adversarial generation network combined with sound visualization auxiliary information to generate dynamic light waves of sound signals to obtain visible light wave data of generated sound signals; The audio adversarial generative network combined with the sound visualization auxiliary information adopts standard generator training and discriminator training to perform generative adversarial training, and constructs a comprehensive loss function of adversarial loss and spectrum reconstruction loss to perform generative adversarial training to obtain the generated sound signal visible light wave data; The calculation formula of the comprehensive loss function is: ; Where, L total is the comprehensive loss function, L GAN is the adversarial loss function, G is the generator identifier, D is the discriminator identifier, r is the spectral reconstruction loss weight, S REAL is the visible light wave data corresponding to the real sound spectrum, S GEN It is the visible light wave data corresponding to the generated sound spectrum; The eye tracking auxiliary guidance specifically collects the user's eye tracking data as auxiliary data, and dynamically adjusts the brightness of the sound source light wave of the generated sound signal visual light wave data through the spatial attention mechanism, performs sight line guidance enhancement training, and obtains sight line guidance optimization data; The adaptive sound field training optimization is specifically to use a standard reinforcement learning method to perform the rehabilitation training difficulty of the basic auditory perception reconstruction training stage according to the generated sound signal visible light wave data and the sight guidance optimization data, and to perform the adaptive sound field training optimization by constructing state parameters, action parameters and reward functions to obtain adaptive training adjustment reference strategy data; The basic auditory perception reconstruction is specifically to perform basic auditory perception reconstruction through the voiceprint dynamic mapping, eye tracking auxiliary guidance and adaptive sound field training optimization, and to perform basic auditory perception reconstruction rehabilitation training in combination with the virtual reality auxiliary data to obtain basic sound feature perception optimization data.
6. The virtual reality-based rehabilitation training system for hearing-impaired patients according to claim 5, characterized in that: The lip synchronization training is used to assist hearing-impaired patients in training their ability to associate lip shape and speech. Specifically, based on the hearing-impaired rehabilitation training fusion data set, a cross-modal contrastive learning method is used to assist in lip synchronization training to obtain language comprehension rehabilitation optimization data; The cross-modal contrastive learning method specifically uses the human voice feature data and human voice video feature data in the hearing impairment rehabilitation training fusion data set as input data, and constructs a standard three-dimensional convolutional network as a lip reading feature encoder, and constructs a standard Wav2Vec2.0 model as a speech feature encoder, and performs cross-modal contrastive learning by extracting lip movement features and speech features, and constructs a contrast loss function to maximize the similarity of lip movement features and speech features to obtain lip movement speech mapping data, and combines the virtual reality auxiliary data to perform lip reading synchronization rehabilitation training to obtain language comprehension rehabilitation optimization data; The calculation formula of the contrast loss function is: ; Where, L contrastive is the contrast loss function, exp(·) is the natural base function, and f lip It is the lip movement feature. is a positive sample in the speech feature, which is used to represent the speech feature that matches the lip movement feature. is the nth speech feature, m is the total number of feature pairs of lip movement features and speech features, n is the feature pair index, is the characteristic pair correction factor.
7. The virtual reality-based rehabilitation training system for hearing-impaired patients according to claim 6, characterized in that: The contextual comprehension training is used to assist hearing-impaired patients in understanding contextual information in virtual scenes. Specifically, the virtual reality auxiliary data is combined to provide vocal context background prompts during the rehabilitation training process to assist patients in understanding vocal speech and obtain language content comprehension rehabilitation optimization data.
8. The virtual reality-based rehabilitation training system for hearing-impaired patients according to claim 7, characterized in that: The comprehensive rehabilitation training is used to carry out rehabilitation training by combining the results of virtual visual environment, auditory perception reconstruction and language content comprehension training. Specifically, virtual reality scene simulation is carried out based on the virtual reality auxiliary data, and real-time gesture and posture monitoring of the patient is combined. The patient's auditory perception during the rehabilitation training is enhanced by the basic sound feature perception optimization data, and the patient's language content perception during the rehabilitation training is enhanced by the language content comprehension rehabilitation optimization data, so as to obtain comprehensive rehabilitation training reference evaluation data.
Citation Information
Patent Citations
Auditory cognitive dysfunction evaluation and rehabilitation training device
CN107591196A
VR-based audio-video-touch sense multi-mode hand function rehabilitation training method
CN108417249A
Multi-modal speech rehabilitation training system based on virtual reality
CN117789982A
Environment simulation equipment based on VR and psychology
CN118629285A
Adjuvant Method for the Interface of Psychosomatic Approaches and Technology for Improving Medical Outcomes
US20150174362A1
Cited By
Voice disorder detection method, device and equipment and readable storage medium
CN120318639A
Cross-modal fusion intelligent joint training method and system
CN121237323A