Virtual Reality-based Rehabilitation Training System for Hearing-impaired Patients

Through multi-stage cross-modal attention-enhanced audio-visual representation method, voiceprint dynamic mapping and cross-modal comparison learning method, the shortcomings of the existing system in multi-sensory data integration, personalized adaptation and intelligent support are solved, and high-precision multi-modal data fusion and personalized rehabilitation training are achieved, which improves the rehabilitation effect of patients with hearing impairments.

CN120015235BActive Publication Date: 2025-06-17川北医学院附属医院
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510461909.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-06-17
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The existing rehabilitation training system for patients with hearing impairments cannot effectively integrate multi-sensory data, resulting in patients being unable to fully perceive voice information, insufficient data fusion accuracy and real-time performance, lack of personalized adaptability, unable to dynamically adjust training content, and lack of intelligent support in lip training and contextual understanding training.

Method used

Multi-stage cross-modal attention-enhanced audio-visual representation method is used to fusion of multimodal data, combined with the audio generation dynamic adjustment method of vocal print dynamic mapping for basic auditory reconstruction, and lip synchronous training is used for cross-modal contrast learning method, and vocal context background prompts are provided in rehabilitation training.

Benefits of technology

It realizes high-precision, real-time multimodal data fusion, dynamically generates personalized sound-scene mapping, improves patients' basic perception of sound, realizes high-precision lip-voice alignment and targeted training of language content, and improves the overall effect of rehabilitation training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015235B_ABST
    Figure CN120015235B_ABST
Patent Text Reader

Abstract

The present invention discloses a rehabilitation training system for hearing-impaired patients based on virtual reality, including a perception fusion module, a basic perception reconstruction module, a language content integration module, a rehabilitation training module, and a virtual reality assistance module. The present invention belongs to the technical field of auditory rehabilitation training. Specifically, it is a rehabilitation training system for hearing-impaired patients based on virtual reality. It adopts a multi-stage cross-modal attention-enhanced speech audio-visual representation method to convert sound signals into visual light waves and tactile vibration data through a deep learning model, realizing high-precision and real-time multi-modal data fusion; adopts a sound frequency generation dynamic adjustment method combined with voiceprint dynamic mapping for basic auditory reconstruction, effectively generating personalized soundscape mappings dynamically to help patients reconstruct the basic perception ability of sounds; adopts a cross-modal contrast learning method for lip-reading synchronization training assistance and provides human voice context background cues during the rehabilitation training process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of auditory rehabilitation training, and specifically relates to a rehabilitation training system for hearing-impaired patients based on virtual reality. Background Art

[0002] A rehabilitation training system for hearing-impaired patients based on virtual reality is an intelligent rehabilitation tool that uses virtual reality (VR) technology, combines multi-modal data fusion and artificial intelligence algorithms to help hearing-impaired patients reconstruct auditory perception ability, improve speech understanding ability and enhance environmental adaptation ability. The system constructs an immersive virtual scene, converts sound signals into visual light waves and tactile vibrations, helps patients perceive sound information through vision and touch, and helps patients better understand and communicate in complex environments. Its function is to improve the patient's auditory perception ability, language understanding ability and social adaptation ability through high-precision and personalized rehabilitation training, and ultimately improve their quality of life.

[0003] However, in the existing rehabilitation training systems for hearing-impaired patients, there are technical problems that the existing systems cannot effectively integrate multi-sensory data, and it is easy for patients to fail to fully perceive sound information through alternative senses during the rehabilitation process. At the same time, the accuracy and real-time performance of data fusion are insufficient, and an immersive rehabilitation training experience cannot be provided for patients, making it difficult to effectively integrate and apply virtual reality technology; in the existing process of auditory perception reconstruction, there are technical problems that the existing systems lack personalized adaptation ability in basic auditory perception reconstruction, cannot dynamically adjust training content according to the degree of hearing loss and rehabilitation progress of patients, and at the same time, the accuracy of sound feature extraction and mapping is also insufficient, resulting in limited improvement in patients' perception ability of sound frequency, direction and intensity; in the existing process of language content integration, there are technical problems that the existing systems also lack intelligent support in lip reading training and context understanding training, cannot effectively identify and complement phonemes, contexts and lip movement information that patients are prone to confuse, and affect the learning effect of patients' speech-lip shape association; in the existing rehabilitation training systems for hearing-impaired patients, there are technical problems that the existing rehabilitation training systems lack interactivity and intelligence, and the scene construction of the existing rehabilitation systems supported by virtual reality depends on preset content and cannot dynamically generate training plans according to patients' needs, resulting in poor overall rehabilitation effects. Summary of the Invention

[0004] In view of the above situation, to overcome the defects of the existing technology, the present invention provides a rehabilitation training system for hearing-impaired patients based on virtual reality. In the existing rehabilitation training system for hearing-impaired patients, there are technical problems that the existing system cannot effectively integrate multi-sensory data, which is likely to cause patients to be unable to fully perceive sound information through alternative senses during the rehabilitation process. At the same time, the accuracy and real-time performance of data fusion are insufficient, unable to provide patients with an immersive rehabilitation training experience, and it is difficult to effectively integrate and apply virtual reality technology. This solution creatively adopts a multi-stage cross-modal attention-enhanced speech-audio-visual representation method to convert sound signals into visual light waves and tactile vibration data through a deep learning model, realizing high-precision and real-time multi-modal data fusion; in view of the technical problems existing in the existing auditory perception reconstruction process, that is, the existing system lacks personalized adaptation ability in basic auditory perception reconstruction, unable to dynamically adjust training content according to the degree of hearing loss and rehabilitation progress of patients. At the same time, the accuracy of sound feature extraction and mapping is also insufficient, resulting in limited improvement in patients' perception ability of sound frequency, direction, and intensity. This solution creatively adopts a sound frequency generation dynamic adjustment method combined with voiceprint dynamic mapping for basic auditory reconstruction, effectively generating personalized soundscape mappings dynamically to help patients reconstruct the basic perception ability of sound; in view of the technical problems existing in the existing language content integration process, that is, the existing system also lacks intelligent support in lip-reading training and context understanding training, unable to effectively identify and complete the phonemes, contexts, and lip movement information that patients are prone to confuse, affecting the learning effect of patients' speech-lip shape association. This solution creatively adopts a cross-modal contrast learning method for lip-reading synchronization training assistance and provides human voice context background prompts during the rehabilitation training process to assist patients in understanding human voice speech, realizing high-precision lip-reading-speech alignment, targeted training of easily confused phonemes, and language content; in view of the technical problems existing in the existing rehabilitation training system for hearing-impaired patients, that is, the existing rehabilitation training system lacks interactivity and intelligence, and the scene construction of the existing rehabilitation system supported by virtual reality depends on preset content and cannot dynamically generate training plans according to patients' needs, resulting in poor overall rehabilitation effects. This solution creatively adopts the integration of three technologies, namely virtual reality technology, auditory reconstruction, and language content understanding, and integratively realizes a rehabilitation training system for hearing-impaired patients based on virtual reality, improving the overall effect of rehabilitation training and providing strong support and exploration experience for the intelligent and automated rehabilitation of hearing-impaired patients.

[0005] The technical solution adopted by the present invention is as follows: The rehabilitation training system for hearing-impaired patients based on virtual reality provided by the present invention includes a perception fusion module, a basic perception reconstruction module, a language content integration module, a rehabilitation training module, and a virtual reality assistance module;

[0006] The perception fusion module is used for multi-modal data collection and processing. Through multi-modal data collection and processing, a fusion dataset for hearing impairment rehabilitation training is obtained, and the fusion dataset for hearing impairment rehabilitation training is sent to the basic perception reconstruction module and the language content integration module;

[0007] The basic perception reconstruction module is used for basic auditory perception reconstruction. Through basic auditory perception reconstruction, optimized data of basic sound feature perception is obtained, and the optimized data of basic sound feature perception is sent to the rehabilitation training module;

[0008] The language content integration module is used for lip-reading synchronization training and context understanding training. Through lip-reading synchronization training and context understanding training, optimized data of language content understanding for rehabilitation is obtained, and the optimized data of language content understanding for rehabilitation is sent to the rehabilitation training module;

[0009] The rehabilitation training module is used for comprehensive rehabilitation training. Through comprehensive rehabilitation training, reference evaluation data for comprehensive rehabilitation training is obtained, and in combination with the virtual reality assistance module, rehabilitation training for hearing-impaired patients is carried out;

[0010] The virtual reality assistance module is used for constructing a virtual reality scene and assisting in rehabilitation training. Through virtual reality technology assistance, rehabilitation training is carried out.

[0011] Furthermore, the multi-modal data collection and processing is used for collecting the original dataset required for virtual reality rehabilitation training. Specifically, through multi-modal data collection, an original dataset for hearing impairment rehabilitation training is obtained, and through data fusion, a fusion dataset for hearing impairment rehabilitation training is obtained;

[0012] The multi-modal data acquisition is used to collect environmental sound, user's line-of-sight direction, and body posture data in real time through a microphone array, a spatial audio sensor, an eye-tracking device, and a tactile feedback device;

[0013] The environmental sound includes audio data and sound spectrogram data; the audio data includes human voice data, human voice video data, and environmental noise data;

[0014] The data fusion is used to extract sound signal features through a speech audio-visual representation method combined with multi-stage cross-modal attention enhancement, and based on the environmental sound, perform feature cross-modal mapping to obtain sound signal feature data and map to obtain visual spectrogram data and tactile vibration data. Through feature cross-modal mapping, a fusion dataset for hearing impairment rehabilitation training is obtained, including the following steps: multi-stage cross-modal feature extraction, multi-scale feature fusion optimization, cross-modal attention optimization, cross-modal feature update, cross-modal feature loss construction, and feature cross-modal mapping;

[0015] The multi-stage cross-modal feature extraction is used to construct a basic problem model of the speech-audio-visual representation method. Specifically, an audio-visual instance feature vector pair is defined, and each pair of audio-visual instance feature vectors is set as a one-hot category vector, and the multi-stage cross-modal feature extraction problem is modeled to obtain audio-visual instance feature vector data;

[0016] The multi-scale feature fusion optimization is used for linear projection and convolution to extract sound and visual features. Specifically, the sound feature data and the visual feature data are respectively subjected to feature extraction and feature fusion through one-dimensional convolution operations to obtain multi-scale feature fusion optimization data;

[0017] The cross-modal attention optimization is used to optimize the attention information of the sound feature and the visual feature and calculate the attention weight. Specifically, based on the multi-scale feature fusion optimization data, cross-modal attention optimization is performed to obtain cross-modal attention weight data;

[0018] The cross-modal feature update is specifically to perform cross-modal feature update on the multi-scale feature fusion optimization data according to the cross-modal attention weight data to obtain updated attention feature data;

[0019] The cross-modal feature loss construction is used to optimize the process of model training. Specifically, a discriminant loss, an intra-modal loss, and a cross-modal loss are constructed in sequence, and loss integration is performed to obtain a cross-modal feature loss function;

[0020] The calculation formula of the cross-modal feature loss function is:

[0021] L = L dis + a·L intra + b·L cross ;

[0022] In the formula, L is the cross-modal feature loss function, L dis is the discriminant loss, L intra is the intra-modal loss, L cross is the cross-modal loss, a is the intra-modal loss weight, and b is the cross-modal loss weight;

[0023] The feature cross-modal mapping is specifically to perform feature cross-modal mapping model training through the multi-stage cross-modal feature extraction, the multi-scale feature fusion optimization, the cross-modal attention optimization, the cross-modal feature update, and the cross-modal feature loss construction to obtain a sound-visual feature cross-modal mapping model, and by using the sound-visual feature cross-modal mapping model, based on the environmental sound, perform feature cross-modal mapping to obtain sound signal feature data and visual spectrum diagram data, and generate tactile vibration simulation data based on the visual spectrum diagram data to obtain an auditory impairment rehabilitation training fusion dataset;

[0024] The hearing impairment rehabilitation training fusion data set specifically includes human voice feature data, human voice video feature data, environmental noise feature data, visualized spectrum data and tactile vibration simulation data;

[0025] The hearing impairment rehabilitation training fusion data set is combined with a virtual reality scene to form virtual reality auxiliary data, and the virtual reality auxiliary data is used for comprehensive rehabilitation training; the virtual reality auxiliary data includes sound source location, sound intensity, sound visualization auxiliary information and sound tactile auxiliary information.

[0026] Furthermore, the basic auditory perception reconstruction is used to assist hearing-impaired patients in reconstructing their basic perception ability of sound. Specifically, based on the hearing-impaired rehabilitation training fusion data set, a method for dynamically adjusting the audio frequency generation combined with dynamic voiceprint mapping is used to perform basic auditory reconstruction to obtain basic sound feature perception optimization data, including the following steps: dynamic voiceprint mapping, eye tracking auxiliary guidance, adaptive sound field training optimization and basic auditory perception reconstruction;

[0027] The dynamic voiceprint mapping is specifically to use an audio adversarial generation network combined with sound visualization auxiliary information to generate dynamic light waves of sound signals to obtain visible light wave data of generated sound signals;

[0028] The audio adversarial generative network combined with the sound visualization auxiliary information adopts standard generator training and discriminator training to perform generative adversarial training, and constructs a comprehensive loss function of adversarial loss and spectrum reconstruction loss to perform generative adversarial training to obtain the generated sound signal visible light wave data;

[0029] The calculation formula of the comprehensive loss function is:

[0030] ;

[0031] Where, L total is the comprehensive loss function, L GAN is the adversarial loss function, G is the generator identifier, D is the discriminator identifier, r is the spectral reconstruction loss weight, S REAL is the visible light wave data corresponding to the real sound spectrum, S GEN It is the visible light wave data corresponding to the generated sound spectrum;

[0032] The eye tracking auxiliary guidance specifically collects the user's eye tracking data as auxiliary data, and dynamically adjusts the brightness of the sound source light wave of the generated sound signal visual light wave data through the spatial attention mechanism, performs sight line guidance enhancement training, and obtains sight line guidance optimization data;

[0033] The adaptive sound field training optimization specifically refers to, based on the generated visual light wave data of the sound signal and the line-of-sight guidance optimization data, using the standard reinforcement learning method to adjust the rehabilitation training difficulty in the basic auditory perception reconstruction training stage, and by constructing state parameters, action parameters, and a reward function, performing adaptive sound field training optimization to obtain adaptive training adjustment reference policy data;

[0034] The basic auditory perception reconstruction specifically refers to, through the voiceprint dynamic mapping, eye tracking assisted guidance, and adaptive sound field training optimization, performing basic auditory perception reconstruction, and combining the virtual reality assisted data to conduct basic auditory perception reconstruction rehabilitation training to obtain basic sound feature perception optimization data.

[0035] Furthermore, the lip-reading synchronization training is used to assist hearing-impaired patients in training their ability to recognize the association between lip shapes and speech. Specifically, based on the auditory disorder rehabilitation training fusion dataset, the cross-modal contrast learning method is adopted to perform lip-reading synchronization training assistance to obtain language understanding rehabilitation optimization data;

[0036] The cross-modal contrast learning method specifically takes the vocal feature data and vocal video feature data in the auditory disorder rehabilitation training fusion dataset as input data, constructs a standard three-dimensional convolutional network as the lip-reading feature encoder, constructs a standard Wav2Vec2.0 model as the speech feature encoder, performs cross-modal contrast learning by extracting lip movement features and speech features, and constructs a contrast loss function to maximize the similarity between lip movement features and speech features to obtain lip movement-speech mapping data, and combines the virtual reality assisted data to conduct lip-reading synchronization rehabilitation training to obtain language understanding rehabilitation optimization data;

[0037] The calculation formula of the contrast loss function is:

[0038] ;

[0039] In the formula, L contrastive is the contrast loss function, exp(·) is the natural exponential function, f lip is the lip movement feature, is the positive sample in the speech feature, used to represent the speech feature that matches the lip movement feature, is the nth speech feature, m is the total number of feature pairs of lip movement features and speech features, n is the feature pair index, is the feature pair correction coefficient.

[0040] Furthermore, the context understanding training is used to assist hearing-impaired patients in understanding context information in a virtual scenario. Specifically, in combination with the virtual reality assistance data, it provides vocal context background cues during the rehabilitation training process to assist the patients in understanding vocal speech and obtain rehabilitation optimization data for language content understanding.

[0041] Furthermore, the comprehensive rehabilitation training is used to conduct rehabilitation training by combining the results of virtual visual environment, auditory perception reconstruction, and language content understanding training. Specifically, based on the virtual reality assistance data, it simulates virtual reality scenarios and combines real-time monitoring of the patients' gesture postures. It enhances the patients' auditory perception during the rehabilitation training process through the basic sound feature perception optimization data and enhances the patients' language content perception during the rehabilitation training process through the language content understanding rehabilitation optimization data to obtain comprehensive rehabilitation training reference evaluation data.

[0042] The beneficial effects achieved by the present invention using the above solution are as follows:

[0043] (1) Aiming at the technical problems existing in the existing rehabilitation training system for hearing-impaired patients, where the existing system cannot effectively integrate multi-sensory data, and during the rehabilitation process, it is easy for patients to fail to fully perceive sound information through alternative senses. At the same time, the accuracy and real-time performance of data fusion are insufficient, and it is impossible to provide an immersive rehabilitation training experience for patients and it is difficult to effectively integrate and apply virtual reality technology. This solution creatively adopts a multi-stage cross-modal attention-enhanced speech-audio-visual representation method to convert sound signals into visual light waves and tactile vibration data through a deep learning model, realizing high-precision and real-time multi-modal data fusion;

[0044] (2) Aiming at the technical problems existing in the existing auditory perception reconstruction process, where the existing system lacks personalized adaptation ability in basic auditory perception reconstruction and cannot dynamically adjust the training content according to the patients' hearing loss degree and rehabilitation progress. At the same time, the accuracy of sound feature extraction and mapping is also insufficient, resulting in limited improvement in the patients' perception ability of sound frequency, direction, and intensity. This solution creatively adopts a sound frequency generation dynamic adjustment method combined with voiceprint dynamic mapping for basic auditory reconstruction, effectively generating personalized soundscape mappings dynamically to help patients reconstruct the basic perception ability of sound;

[0045] (3) In view of the technical problem that in the existing language content integration process, the existing system lacks intelligent support in lip-reading training and context understanding training, and cannot effectively identify and complete the phonemes, contexts, and lip movement information that patients are prone to confuse, affecting the learning effect of patients on the speech-lip shape association, this solution creatively adopts a cross-modal contrast learning method to assist in lip-reading synchronization training, and provides a vocal context background prompt during the rehabilitation training process to assist patients in understanding vocal speech, achieving high-precision lip-reading-speech alignment, targeted training for easily confused phonemes, and language content;

[0046] (4) In view of the technical problem that in the existing rehabilitation training system for hearing-impaired patients, the existing rehabilitation training system lacks interactivity and intelligence, and the scene construction of the existing rehabilitation system supported by virtual reality depends on preset content and cannot dynamically generate training plans according to the needs of patients, resulting in poor overall rehabilitation effects, this solution creatively adopts the integration of three technologies: virtual reality technology, auditory reconstruction, and language content understanding, and integrally realizes a rehabilitation training system for hearing-impaired patients based on virtual reality, improving the overall effect of rehabilitation training, and also providing strong support and exploration experience for the intelligent and automated rehabilitation of hearing-impaired patients. Brief Description of the Drawings

[0047] Figure 1 It is a schematic structural diagram of a rehabilitation training system for hearing-impaired patients based on virtual reality provided by the present invention;

[0048] Figure 2 It is a schematic flow diagram of the steps executed by the system provided by the present invention;

[0049] Figure 3 It is a schematic flow diagram of the steps executed by data fusion in the perception fusion module;

[0050] Figure 4 It is a schematic flow diagram of the steps executed by the basic perception reconstruction module.

[0051] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention. Detailed Embodiments

[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0053] Embodiment 1, Refer toFigure 1 , the technical solution adopted by the present invention is as follows: The rehabilitation training system for hearing-impaired patients based on virtual reality provided by the present invention includes a perception fusion module, a basic perception reconstruction module, a language content integration module, a rehabilitation training module, and a virtual reality assistance module;

[0054] The perception fusion module is used for collecting and processing multi-modal data. Through the collection and processing of multi-modal data, a fusion data set for hearing impairment rehabilitation training is obtained, and the fusion data set for hearing impairment rehabilitation training is sent to the basic perception reconstruction module and the language content integration module;

[0055] The basic perception reconstruction module is used for basic auditory perception reconstruction. Through basic auditory perception reconstruction, optimized data for basic sound feature perception is obtained, and the optimized data for basic sound feature perception is sent to the rehabilitation training module;

[0056] The language content integration module is used for lip-reading synchronization training and context understanding training. Through lip-reading synchronization training and context understanding training, optimized data for language content understanding rehabilitation is obtained, and the optimized data for language content understanding rehabilitation is sent to the rehabilitation training module;

[0057] The rehabilitation training module is used for comprehensive rehabilitation training. Through comprehensive rehabilitation training, reference evaluation data for comprehensive rehabilitation training is obtained, and combined with the virtual reality assistance module, rehabilitation training for hearing-impaired patients is carried out;

[0058] The virtual reality assistance module is used for constructing a virtual reality scene and assisting in rehabilitation training. Through virtual reality technology assistance, rehabilitation training is carried out.

[0059] By performing the above operations, in the existing rehabilitation training system for hearing-impaired patients, there are problems that the existing rehabilitation training system lacks interactivity and intelligence, and the scene construction of the existing rehabilitation system supported by virtual reality depends on preset content and cannot dynamically generate training plans according to the needs of patients, resulting in poor overall rehabilitation effects. This solution creatively combines three technologies: virtual reality technology, auditory reconstruction, and language content understanding, and integrally realizes a rehabilitation training system for hearing-impaired patients based on virtual reality, improving the overall effect of rehabilitation training, and also providing strong support and exploration experience for the intelligent and automated rehabilitation of hearing-impaired patients.

[0060] Embodiment 2, this embodiment is based on the above embodiment, refer to Figure 1 , Figure 2 and Figure 3, the multimodal data collection and processing is used to collect the original data set required for virtual reality rehabilitation training. Specifically, through multimodal data collection, the original data set for auditory impairment rehabilitation training is obtained, and through data fusion, the fused data set for auditory impairment rehabilitation training is obtained;

[0061] The multimodal data acquisition is to collect environmental sound, user's line-of-sight direction and body posture data in real time through a microphone array, a spatial audio sensor, an eye-tracking device and a tactile feedback device;

[0062] The environmental sound includes audio data and sound spectrogram data; the audio data includes human voice data, human voice video data and environmental noise data;

[0063] The data fusion is to perform sound signal feature extraction through a speech audio-visual representation method enhanced by multi-stage cross-modal attention, and based on the environmental sound, perform feature cross-modal mapping to obtain sound signal feature data and map to obtain visual spectrogram data and tactile vibration data. Through feature cross-modal mapping, the fused data set for auditory impairment rehabilitation training is obtained, including the following steps: multi-stage cross-modal feature extraction, multi-scale feature fusion optimization, cross-modal attention optimization, cross-modal feature update, cross-modal feature loss construction and feature cross-modal mapping;

[0064] The multi-stage cross-modal feature extraction is used to construct the basic problem model of the speech audio-visual representation method. Specifically, the audio-visual instance feature vector pair is defined, and each pair of audio-visual instance feature vector pairs is set as a one-hot category vector to perform multi-stage cross-modal feature extraction problem modeling to obtain audio-visual instance feature vector data. The calculation formula is:

[0065] ;

[0066] In the formula, D is the audio-visual instance feature vector data, is the sound feature vector data, and the sound feature is used to represent the features of the original sound data, is the visual feature vector data, and the visual feature is used to represent the waveform visualization data corresponding to the original sound. N is the total number of audio-visual instances, and i is the audio-visual instance index;

[0067] The multi-scale feature fusion optimization is used for linear projection and convolution to extract sound and visual features. Specifically, the sound feature data and the visual feature data are respectively subjected to feature extraction and feature fusion through one-dimensional convolution operation to obtain multi-scale feature fusion optimization data. The calculation formula is:

[0068] ;

[0069] In the formula, is the convolutional sound feature data, k is the convolutional kernel size, Conv1D(·) is the one-dimensional convolutional operation function, F a is the sound feature data, is the sound feature extraction parameter, is the convolutional visual feature data, F v is the visual feature data, is the visual feature extraction parameter, M is the multi-scale feature fusion optimization data, W M is the feature fusion weight;

[0070] Preferably, the value range of the convolutional kernel size is specifically {3, 5, 7};

[0071] The cross-modal attention optimization is used to optimize the attention information of the sound feature and the visual feature and calculate the attention weight. Specifically, according to the multi-scale feature fusion optimization data, cross-modal attention optimization is performed to obtain the cross-modal attention weight data. The calculation formula is:

[0072] ;

[0073] In the formula, A is the cross-modal attention weight data, including the sound attention weight and the visual attention weight, is the sound attention weight, k is the convolutional kernel size, is the visual attention weight, softmax(·) is the classifier function, Q a is the sound query matrix, K a is the sound key matrix, V a is the sound value matrix, M is the multi-scale feature fusion optimization data, T is the transpose operator, is the dimension value of the sound key matrix, Q v is the visual query matrix, K v is the visual key matrix, V v is the visual value matrix, is the dimension value of the visual key matrix;

[0074] The cross-modal feature update is specifically to perform cross-modal feature update on the multi-scale feature fusion optimization data according to the cross-modal attention weight data to obtain the updated attention feature data;

[0075] The calculation formula of the cross-modal feature update is:

[0076] ;

[0077] In the formula, is the updated sound feature data, is the updated visual feature data;

[0078] The calculation formula for updating the attention feature data is as follows:

[0079] ;

[0080] In the formula, F is the updated attention feature data, including the updated attention integrated sound feature data and the updated attention integrated visual feature data. is the updated attention integrated sound feature data. is the updated attention integrated visual feature data. is the updated sound feature data obtained by performing feature fusion optimization using a one-dimensional convolutional kernel with a kernel size of 3. is the updated visual feature data obtained by performing feature fusion optimization using a one-dimensional convolutional kernel with a kernel size of 3. is the updated sound feature data obtained by performing feature fusion optimization using a one-dimensional convolutional kernel with a kernel size of 5. is the updated visual feature data obtained by performing feature fusion optimization using a one-dimensional convolutional kernel with a kernel size of 5. is the updated sound feature data obtained by performing feature fusion optimization using a one-dimensional convolutional kernel with a kernel size of 7. is the updated visual feature data obtained by performing feature fusion optimization using a one-dimensional convolutional kernel with a kernel size of 7.

[0081] The cross-modal feature loss construction is used to optimize the model training process. Specifically, the discriminant loss, intra-modal loss, and cross-modal loss are constructed in sequence, and loss integration is performed to obtain the cross-modal feature loss function.

[0082] The calculation formula for the discriminant loss is as follows:

[0083] ;

[0084] In the formula, L dis is the discriminant loss. is the sound feature discriminant loss, specifically using the cross-entropy loss function. is the visual feature discriminant loss, specifically using the cross-entropy loss function.

[0085] The calculation formula for the intra-modal loss is as follows:

[0086] ;

[0087] In the formula, L intra is the intra-modal loss, N is the total number of audio-visual instances, i is the audio-visual instance index. is the sound feature value corresponding to the i-th audio-visual instance. is the sound feature center value corresponding to the i-th audio-visual instance. is the visual feature value corresponding to the i-th audio-visual instance, is the visual feature center value corresponding to the i-th audio-visual instance;

[0088] The calculation formula of the cross-modal loss is:

[0089] ;

[0090] In the formula, L cross is the cross-modal loss, N is the total number of audio-visual instances, i is the audio-visual instance index, e is the natural base, is the cosine similarity value of the positive sample, is the cosine similarity value of the negative sample;

[0091] The calculation formula of the cross-modal feature loss function is:

[0092] L = L dis + a·L intra + b·L cross ;

[0093] In the formula, L is the cross-modal feature loss function, L dis is the discriminant loss, L intra is the intra-modal loss, L cross is the cross-modal loss, a is the intra-modal loss weight, b is the cross-modal loss weight;

[0094] The feature cross-modal mapping is specifically carried out by constructing the multi-stage cross-modal feature extraction, the multi-scale feature fusion optimization, the cross-modal attention optimization, the cross-modal feature update and the cross-modal feature loss, training the feature cross-modal mapping model, obtaining the sound-visual feature cross-modal mapping model, and using the sound-visual feature cross-modal mapping model, according to the environmental sound, performing feature cross-modal mapping to obtain sound signal feature data and visualization spectrogram data, and generating tactile vibration simulation data according to the visualization spectrogram data to obtain an auditory impairment rehabilitation training fusion dataset;

[0095] The auditory impairment rehabilitation training fusion dataset specifically includes human voice feature data, human voice video feature data, environmental noise feature data, visualization spectrum data and tactile vibration simulation data;

[0096] The auditory impairment rehabilitation training fusion dataset is combined with a virtual reality scene to form virtual reality auxiliary data, and the virtual reality auxiliary data is used for comprehensive rehabilitation training; the virtual reality auxiliary data includes sound source position, sound intensity, sound visualization auxiliary information and sound tactileization auxiliary information.

[0097] By performing the above operations, in the existing hearing impairment patient rehabilitation training system, there are technical problems that the existing system cannot effectively integrate multi-sensory data. During the rehabilitation process, it is easy for patients to be unable to fully perceive sound information through alternative senses. At the same time, the accuracy and real-time performance of data fusion are insufficient, and it is impossible to provide patients with an immersive rehabilitation training experience, and it is difficult to effectively integrate and apply virtual reality technology. This solution creatively adopts a multi-stage cross-modal attention-enhanced speech-audio-visual representation method, which converts sound signals into visual light waves and tactile vibration data through a deep learning model to achieve high-precision and real-time multi-modal data fusion.

[0098] Embodiment 3. This embodiment is based on the above embodiment. Refer to Figure 1 、 Figure 2 and Figure 4 , the basic auditory perception reconstruction is used to assist hearing-impaired patients in reconstructing the basic perception ability of sound. Specifically, according to the auditory impairment rehabilitation training fusion data set, a method for dynamically adjusting the generation of sound frequency by combining voiceprint dynamic mapping is adopted for basic auditory reconstruction to obtain optimized data for basic sound feature perception, including the following steps: voiceprint dynamic mapping, eye tracking-assisted guidance, adaptive sound field training optimization, and basic auditory perception reconstruction;

[0099] The voiceprint dynamic mapping is specifically to adopt an audio frequency adversarial generation network combined with voice visualization auxiliary information to generate dynamic light waves of sound signals to obtain visual light wave data of generated sound signals;

[0100] The audio frequency adversarial generation network combined with voice visualization auxiliary information adopts standard generator training and discriminator training for generative adversarial training, and constructs a comprehensive loss function of adversarial loss and spectrum reconstruction loss for generative adversarial training to obtain visual light wave data of generated sound signals;

[0101] The calculation formula of the comprehensive loss function is:

[0102] ;

[0103] In the formula, L total is the comprehensive loss function, L GAN is the adversarial loss function, G is the generator identifier, D is the discriminator identifier, r is the spectrum reconstruction loss weight, S REAL is the visual light wave data corresponding to the real sound spectrum, and S GEN is the visual light wave data corresponding to the generated sound spectrum;

[0104] The eye-tracking assisted guidance specifically involves collecting the user's eye-tracking data as auxiliary data, and through a spatial attention mechanism, dynamically adjusting the brightness of the sound source light wave of the generated sound signal visual light wave data to perform gaze guidance enhancement training, and obtaining gaze guidance optimized data;

[0105] The adaptive sound field training optimization specifically involves, based on the generated sound signal visual light wave data and the gaze guidance optimized data, using a standard reinforcement learning method to perform rehabilitation training difficulty in the basic auditory perception reconstruction training stage, and through constructing state parameters, action parameters, and a reward function, performing adaptive sound field training optimization to obtain adaptive training adjustment reference policy data;

[0106] The state parameters include the reaction time and accuracy rate of the user's sound source localization;

[0107] The action parameters include adding interfering sound sources, reducing interfering sound sources, and adjusting the distance between the sound source and the patient;

[0108] The reward function specifically refers to the positive feedback reward for the user's localization correct rate;

[0109] The basic auditory perception reconstruction specifically involves, through the voiceprint dynamic mapping, eye-tracking assisted guidance, and adaptive sound field training optimization, performing basic auditory perception reconstruction, and combining the virtual reality assisted data to perform basic auditory perception reconstruction rehabilitation training to obtain basic sound feature perception optimized data.

[0110] By performing the above operations, in the existing auditory perception reconstruction process, there are technical problems that the existing system lacks personalized adaptation ability in basic auditory perception reconstruction, cannot dynamically adjust the training content according to the patient's hearing loss degree and rehabilitation progress, and at the same time, the accuracy of sound feature extraction and mapping is also insufficient, resulting in limited improvement in the patient's perception ability of sound frequency, direction, and intensity. This solution creatively adopts an audio frequency generation dynamic adjustment method combined with voiceprint dynamic mapping to perform basic auditory reconstruction, effectively generating personalized soundscape mapping dynamically to help patients reconstruct the basic perception ability of sound.

[0111] Example 4, this example is based on the above example, refer to Figure 1 and Figure 2 The lip-reading synchronization training is used to assist hearing-impaired patients in training the association recognition ability of lip shapes and voices. Specifically, based on the auditory disorder rehabilitation training fusion dataset, a cross-modal contrast learning method is used to perform lip-reading synchronization training assistance to obtain language understanding rehabilitation optimized data;

[0112] The cross-modal contrastive learning method specifically uses the vocal feature data and vocal video feature data in the auditory impairment rehabilitation training fusion dataset as input data, constructs a standard three-dimensional convolutional network as the lip language feature encoder, constructs a standard Wav2Vec2.0 model as the speech feature encoder, performs cross-modal contrastive learning by extracting lip movement features and speech features, constructs a contrast loss function to maximize the similarity between the lip movement features and speech features, obtains lip movement-speech mapping data, and combines the virtual reality assistance data to perform lip language synchronization rehabilitation training to obtain language understanding rehabilitation optimization data;

[0113] The calculation formula of the contrast loss function is:

[0114] ;

[0115] In the formula, L contrastive is the contrast loss function, exp(·) is the natural base function, f lip is the lip movement feature, is the positive sample in the speech feature, used to represent the speech feature that matches the lip movement feature, is the nth speech feature, m is the total number of feature pairs of the lip movement feature and the speech feature, n is the feature pair index, is the feature pair correction coefficient.

[0116] Example 5. This example is based on the above example. Refer to Figure 1 and Figure 2 The context understanding training is used to assist hearing-impaired patients in understanding context information in a virtual scene. Specifically, it combines the virtual reality assistance data to provide a vocal context background prompt during the rehabilitation training to assist the patient's vocal speech understanding and obtain language content understanding rehabilitation optimization data.

[0117] By performing the above operations, aiming at the technical problem that in the existing language content integration process, the existing system also lacks intelligent support in lip language training and context understanding training, and cannot effectively identify and complete the phonemes, contexts, and lip movement information that patients are prone to confuse, affecting the learning effect of the patient's speech-lip shape association, this solution creatively adopts a cross-modal contrastive learning method to assist in lip language synchronization training, and provides a vocal context background prompt during the rehabilitation training to assist the patient's vocal speech understanding, realizing high-precision lip language-speech alignment, targeted training of easily confused phonemes, and language content.

[0118] Example 6. This example is based on the above example. Refer to Figure 1, the comprehensive rehabilitation training is used to conduct rehabilitation training by combining the results of virtual visual environment, auditory perception reconstruction, and language content understanding training. Specifically, based on the virtual reality assistance data, virtual reality scene simulation is performed, and the gesture postures of the patient are monitored in real time. During the rehabilitation training, the auditory perception of the patient is enhanced through the optimized data of the basic sound feature perception, and the language content perception of the patient is enhanced through the optimized data of the language content understanding rehabilitation, so as to obtain the comprehensive rehabilitation training reference evaluation data.

[0119] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process or method comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such a process or method.

[0120] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present invention.

[0121] The above describes the present invention and its implementation manners. Such a description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments without creative efforts without departing from the purpose of the present invention, they shall fall within the protection scope of the present invention.

Claims

1. A rehabilitation training system for hearing-impaired patients based on virtual reality, characterized by: It includes perception fusion module, basic perception reconstruction module, language content integration module, rehabilitation training module and virtual reality assistance module; The perception fusion module is used for multimodal data collection and processing, and obtains a hearing impairment rehabilitation training fusion data set through multimodal data collection and processing and data fusion, and sends the hearing impairment rehabilitation training fusion data set to the basic perception reconstruction module and the language content integration module; The data fusion comprises the following steps: multi-stage cross-modal feature extraction, multi-scale feature fusion optimization, cross-modal attention optimization, cross-modal feature update, cross-modal feature loss construction and feature cross-modal mapping; The basic perception reconstruction module is used for basic auditory perception reconstruction, obtains basic sound feature perception optimization data through basic auditory perception reconstruction, and sends the basic sound feature perception optimization data to the rehabilitation training module; The basic auditory perception reconstruction includes the following steps: voiceprint dynamic mapping, eye tracking auxiliary guidance, adaptive sound field training optimization and basic auditory perception reconstruction; The language content integration module is used for lip synchronization training and context comprehension training, obtains language content comprehension rehabilitation optimization data through lip synchronization training and context comprehension training, and sends the language content comprehension rehabilitation optimization data to the rehabilitation training module; The rehabilitation training module is used for comprehensive rehabilitation training. Through comprehensive rehabilitation training, comprehensive rehabilitation training reference evaluation data is obtained, and combined with the virtual reality auxiliary module, rehabilitation training for hearing-impaired patients is carried out; The virtual reality auxiliary module is used to construct a virtual reality scene and assist rehabilitation training, and conduct rehabilitation training with the assistance of virtual reality technology.

2. The virtual reality-based rehabilitation training system for hearing-impaired patients according to claim 1, characterized in that: The multimodal data collection and processing is used to collect the original data set required for virtual reality rehabilitation training, specifically, to obtain the hearing impairment rehabilitation training original data set through multimodal data collection, and to obtain the hearing impairment rehabilitation training fusion data set through data fusion; The multimodal data collection collects environmental sound, user line of sight direction and body posture data in real time through microphone arrays, spatial audio sensors, eye tracking devices and tactile feedback devices; The environmental sound includes audio data and sound spectrogram data; the audio data includes human voice data, human voice video data and environmental noise data; The data fusion extracts sound signal features by combining a multi-stage cross-modal attention enhanced speech and audio-visual representation method, performs feature cross-modal mapping based on the environmental sound, obtains sound signal feature data and maps them to obtain visual spectrogram data and tactile vibration data, and obtains a hearing impairment rehabilitation training fusion data set through feature cross-modal mapping, including the following steps: multi-stage cross-modal feature extraction, multi-scale feature fusion optimization, cross-modal attention optimization, cross-modal feature update, cross-modal feature loss construction, and feature cross-modal mapping; The hearing impairment rehabilitation training fusion data set is combined with a virtual reality scene to form virtual reality auxiliary data, and the virtual reality auxiliary data is used for comprehensive rehabilitation training; the virtual reality auxiliary data includes sound source location, sound intensity, sound visualization auxiliary information and sound tactile auxiliary information.

3. The virtual reality-based rehabilitation training system for hearing-impaired patients according to claim 2, characterized in that: The multi-stage cross-modal feature extraction is used to construct a basic problem model of the speech audio-visual representation method, specifically defining audio-visual instance feature vector pairs, and setting each pair of audio-visual instance feature vector pairs as a one-hot category vector, performing multi-stage cross-modal feature extraction problem modeling, and obtaining audio-visual instance feature vector data; The multi-scale feature fusion optimization is used for extracting sound and visual features by linear projection and convolution, specifically, the sound feature data and the visual feature data are subjected to feature extraction and feature fusion respectively through one-dimensional convolution operation to obtain multi-scale feature fusion optimization data; The cross-modal attention optimization is used to optimize the attention information of the sound features and the visual features and calculate the attention weights, specifically, performing cross-modal attention optimization based on the multi-scale feature fusion optimization data to obtain cross-modal attention weight data; The cross-modal feature updating is specifically to perform cross-modal feature updating on the multi-scale feature fusion optimization data according to the cross-modal attention weight data to obtain updated attention feature data; The cross-modal feature loss construction is used to optimize the model training process, specifically constructing the discriminant loss, the intra-modal loss and the cross-modal loss in sequence, and performing loss integration to obtain the cross-modal feature loss function; The calculation formula of the cross-modal feature loss function is: L=L dis +a·L intra +b·L cross ; Where L is the cross-modal feature loss function, L dis is the discrimination loss, L intra is the intra-modal loss, L cross is the cross-modal loss, a is the intra-modal loss weight, and b is the cross-modal loss weight; The feature cross-modal mapping is specifically to perform feature cross-modal mapping model training through the multi-stage cross-modal feature extraction, the multi-scale feature fusion optimization, the cross-modal attention optimization, the cross-modal feature update and the cross-modal feature loss construction to obtain a sound-visual feature cross-modal mapping model, and to perform feature cross-modal mapping based on the environmental sound by using the sound-visual feature cross-modal mapping model to obtain sound signal feature data and visualized spectrum data, and to generate tactile vibration simulation data based on the visualized spectrum data to obtain a hearing impairment rehabilitation training fusion data set; The hearing impairment rehabilitation training fusion data set specifically includes human voice feature data, human voice video feature data, environmental noise feature data, visualized spectrum data and tactile vibration simulation data.

4. The virtual reality-based rehabilitation training system for hearing-impaired patients according to claim 3, characterized in that: The basic auditory perception reconstruction is used to assist hearing-impaired patients in reconstructing their basic perception ability of sound. Specifically, based on the hearing-impairment rehabilitation training fusion data set, a dynamic adjustment method of audio generation combined with voiceprint dynamic mapping is used to perform basic auditory reconstruction to obtain basic sound feature perception optimization data, including the following steps: voiceprint dynamic mapping, eye tracking auxiliary guidance, adaptive sound field training optimization and basic auditory perception reconstruction.

5. The virtual reality-based rehabilitation training system for hearing-impaired patients according to claim 4, characterized in that: The dynamic voiceprint mapping is specifically to use an audio adversarial generation network combined with sound visualization auxiliary information to generate dynamic light waves of sound signals to obtain visible light wave data of generated sound signals; The audio adversarial generative network combined with the sound visualization auxiliary information adopts standard generator training and discriminator training to perform generative adversarial training, and constructs a comprehensive loss function of adversarial loss and spectrum reconstruction loss to perform generative adversarial training to obtain the generated sound signal visible light wave data; The calculation formula of the comprehensive loss function is: ; Where, L total is the comprehensive loss function, L GAN is the adversarial loss function, G is the generator identifier, D is the discriminator identifier, r is the spectral reconstruction loss weight, S REAL is the visible light wave data corresponding to the real sound spectrum, S GEN It is the visible light wave data corresponding to the generated sound spectrum; The eye tracking auxiliary guidance specifically collects the user's eye tracking data as auxiliary data, and dynamically adjusts the brightness of the sound source light wave of the generated sound signal visual light wave data through the spatial attention mechanism, performs sight line guidance enhancement training, and obtains sight line guidance optimization data; The adaptive sound field training optimization is specifically to use a standard reinforcement learning method to perform the rehabilitation training difficulty of the basic auditory perception reconstruction training stage according to the generated sound signal visible light wave data and the sight guidance optimization data, and to perform the adaptive sound field training optimization by constructing state parameters, action parameters and reward functions to obtain adaptive training adjustment reference strategy data; The basic auditory perception reconstruction is specifically to perform basic auditory perception reconstruction through the voiceprint dynamic mapping, eye tracking auxiliary guidance and adaptive sound field training optimization, and to perform basic auditory perception reconstruction rehabilitation training in combination with the virtual reality auxiliary data to obtain basic sound feature perception optimization data.

6. The virtual reality-based rehabilitation training system for hearing-impaired patients according to claim 5, characterized in that: The lip synchronization training is used to assist hearing-impaired patients in training their ability to associate lip shape and speech. Specifically, based on the hearing-impaired rehabilitation training fusion data set, a cross-modal contrastive learning method is used to assist in lip synchronization training to obtain language comprehension rehabilitation optimization data; The cross-modal contrastive learning method specifically uses the human voice feature data and human voice video feature data in the hearing impairment rehabilitation training fusion data set as input data, and constructs a standard three-dimensional convolutional network as a lip reading feature encoder, and constructs a standard Wav2Vec2.0 model as a speech feature encoder, and performs cross-modal contrastive learning by extracting lip movement features and speech features, and constructs a contrast loss function to maximize the similarity of lip movement features and speech features to obtain lip movement speech mapping data, and combines the virtual reality auxiliary data to perform lip reading synchronization rehabilitation training to obtain language comprehension rehabilitation optimization data; The calculation formula of the contrast loss function is: ; Where, L contrastive is the contrast loss function, exp(·) is the natural base function, and f lip It is the lip movement feature. is a positive sample in the speech feature, which is used to represent the speech feature that matches the lip movement feature. is the nth speech feature, m is the total number of feature pairs of lip movement features and speech features, n is the feature pair index, is the characteristic pair correction factor.

7. The virtual reality-based rehabilitation training system for hearing-impaired patients according to claim 6, characterized in that: The contextual comprehension training is used to assist hearing-impaired patients in understanding contextual information in virtual scenes. Specifically, the virtual reality auxiliary data is combined to provide vocal context background prompts during the rehabilitation training process to assist patients in understanding vocal speech and obtain language content comprehension rehabilitation optimization data.

8. The virtual reality-based rehabilitation training system for hearing-impaired patients according to claim 7, characterized in that: The comprehensive rehabilitation training is used to carry out rehabilitation training by combining the results of virtual visual environment, auditory perception reconstruction and language content comprehension training. Specifically, virtual reality scene simulation is carried out based on the virtual reality auxiliary data, and real-time gesture and posture monitoring of the patient is combined. The patient's auditory perception during the rehabilitation training is enhanced by the basic sound feature perception optimization data, and the patient's language content perception during the rehabilitation training is enhanced by the language content comprehension rehabilitation optimization data, so as to obtain comprehensive rehabilitation training reference evaluation data.

Citation Information

Patent Citations

  • Auditory cognitive dysfunction evaluation and rehabilitation training device

    CN107591196A

  • Systems and methods for presenting visual, audible, and tactile CUES within an augmented reality, virtual reality, or mixed reality game environment

    WO2024028759A1