Rail transit hearing impairment voice auxiliary system based on multi-mode perception
By adopting multimodal perception technology in rail transit environments, combining deep noise reduction networks, spectrum correction algorithms and spatiotemporal convolutional networks, speech enhancement, vibration feature extraction and lip recognition are achieved, solving the problem of hearing-impaired people obtaining information in complex noise environments, and significantly improving the reliability and richness of information acquisition.
Patent Information
- Application Number
- CN202510535859.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-27
AI Technical Summary
In rail transit environments, hearing-impaired people find it difficult to obtain key information from complex background noise, such as train arrival prompts and transfer instructions. Traditional auxiliary equipment cannot synchronize voice content in real time and accurately, and lip recognition technology has low accuracy in complex environments.
The multimodal perception-based speech assistance system for hearing impaired rail transit is adopted to obtain speech signals, environmental vibration signals and visual lip information through the multimodal data acquisition module, and combine deep noise reduction networks, spectrum correction algorithms, spatiotemporal convolution networks and dynamic weighted fusion models to realize speech enhancement, vibration feature extraction, lip recognition and multimodal information fusion.
It significantly improves the recognizableness of voice information, provides information acquisition channels other than voice and lip, enhances the reliability and richness of information acquisition, and improves the information acquisition ability and travel experience of hearing impaired people in the rail transit environment.
Smart Images

Figure CN120071950A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of rail transit hearing aids, and particularly to a voice assistance system for hearing-impaired people in rail transit based on multi-modal perception. Background Art
[0002] In modern society, rail transit has become one of the important ways for people to travel. It has the advantages of high efficiency, convenience, large passenger capacity, etc., and greatly meets people's daily travel needs. However, for hearing-impaired people, obtaining voice information in the rail transit environment faces many challenges.
[0003] During the train operation, the environmental noise is very complex. The noise generated by the friction between the wheels and the track during train travel, the operation sound of the ventilation system in the carriage, the conversations of passengers, etc. are intertwined, forming a strong background noise. These noises seriously interfere with the transmission of voice information such as train announcements. Even for people with normal hearing, it is difficult to clearly distinguish the announcement content during some noisy periods, and hearing-impaired people are more severely affected. They often cannot obtain key information from these noisy sounds, such as train arrival announcements, transfer instructions, etc., which brings great inconvenience to their travel and may even lead to problems such as missing stations and unable to transfer smoothly.
[0004] Traditional auxiliary devices have obvious limitations when dealing with the rail transit environment. Taking a simple subtitle display device as an example, it usually can only provide fixed-format text information and cannot synchronize with the voice content in real time and accurately. In actual applications, the speech rate and content of train announcements vary greatly, and subtitle display may have delays, incomplete information, etc., making it difficult to meet the needs of hearing-impaired people to obtain information in a timely manner. Moreover, such devices only rely on single visual information and do not fully consider other information clues in the environment, and cannot comprehensively improve the understanding and reception effect of hearing-impaired people on voice information.
[0005] In addition, lip-reading recognition technology also faces difficulties in the complex rail transit environment. On the one hand, the light conditions in the carriage are complex and changeable. At different times and different positions in the carriage, the light intensity and angle are different, which may lead to a decrease in the quality of lip-reading images captured by the camera and affect the accuracy of lip-reading recognition. On the other hand, the positions and postures of passengers are not fixed. Sometimes the mouth of the speaker may be partially blocked, or in the shooting blind spot of the camera, making it difficult to effectively perform lip-reading recognition. Summary of the Invention
[0006] The purpose of the present invention is to provide a voice assistance system for hearing-impaired people in rail transit based on multi-modal perception to solve the problems raised in the above background art.
[0007] To achieve the above object, the present invention provides the following technical solution: A rail transit hearing-impaired voice assistance system based on multi-modal perception, the system includes: Multi-modal data acquisition module: configured to obtain a multi-source perception data set in the rail transit environment in real time, the multi-source perception data includes voice signals, environmental vibration signals and visual lip language information; Voice enhancement processing module: configured to separate the target voice and environmental noise based on the voice signal through a deep noise reduction network to generate a denoised voice stream; the deep noise reduction network adopts a dual-channel spectrum masking algorithm and a time-domain waveform reconstruction technique; Vibration signal analysis module: configured to extract vibration feature vectors synchronized with the voice signal according to the environmental vibration signal through a spectrum correction algorithm, and construct a vibration-voice mapping relationship model; Lip language recognition module: configured to perform dynamic semantic segmentation processing on the visual lip language information, and use a spatio-temporal convolutional network to extract a sequence of key mouth shape frames to generate a lip language text prediction result; Multi-modal fusion module: input the denoised voice stream, vibration feature vectors and lip language text prediction results into a dynamic weighted fusion model, and output an enhanced voice assistance information stream; the dynamic weighted fusion model includes an adaptive weight allocation layer and a cross-modal alignment mechanism.
[0008] Preferably, the dual-channel spectrum masking algorithm of the deep noise reduction network includes: Perform short-time Fourier transform on the input voice signal to generate a spectrogram, and divide it into high-frequency and low-frequency sub-bands; Use parallel convolutional gated units to process the high-frequency and low-frequency sub-bands respectively to generate a spectrum masking coefficient matrix; Filter the original spectrum through the masking coefficient matrix to reconstruct the denoised time-domain voice waveform.
[0009] Preferably, the steps of the spectrum correction algorithm for extracting vibration feature vectors synchronized with the voice signal include: Perform time-frequency transformation on the vibration signal to generate a time-frequency spectrum matrix, detect the position of the formant synchronized with the voice signal, calculate the vibration energy distribution characteristics according to the formant position, and eliminate the interference components in the non-voice related frequency bands; Adopt a sliding window difference method to extract the transient features in the vibration signal and construct a multi-dimensional vibration feature vector.
[0010] Preferably, the method for the spatio-temporal convolutional network to extract a sequence of key mouth shape frames includes: Perform key frame sampling on the continuous lip language video stream to generate a group of time series images, and extract the spatial features and time correlation of the mouth shape movement through a three-dimensional convolutional kernel; The semantic weights of key lip movement change frames are strengthened by using the attention mechanism to generate a lip movement semantic coding sequence.
[0011] Preferably, the adaptive weight allocation layer of the dynamic weighted fusion model includes: Calculating the signal-to-noise ratio index for the denoised speech stream to generate a speech confidence weight; Calculating a vibration assistance weight according to the mapping relationship between the vibration feature vector and the speech signal; Generating a lip language correction weight based on the semantic integrity of the lip language text prediction result; Performing dynamic normalization processing on the three types of weights through a differentiable optimization algorithm to generate a fusion weight matrix.
[0012] Preferably, the implementation method of the cross-modal alignment mechanism includes: Mapping the speech stream, vibration feature vector, and lip language text to a unified time coordinate system; Calculating the time offset of multi-modal data through the dynamic time warping algorithm and compensating for the alignment error; Using a gated recurrent unit to perform sequence fusion on the aligned multi-modal data to generate a synchronized enhanced speech stream.
[0013] Preferably, the optimization method of the time-domain waveform reconstruction technology includes: Performing phase recovery processing on the filtered spectrum, optimizing the continuity of the time-domain waveform by using an iterative projection algorithm, evaluating the auditory quality of the reconstructed waveform through a generative adversarial network, and iteratively adjusting the spectrum mask parameters; Compensating the loudness of the reconstructed speech in combination with the human ear auditory characteristic curve.
[0014] Preferably, the construction method of the vibration-speech mapping relationship model includes: Collecting multiple groups of synchronous vibration signals and pure speech samples to construct a training data set; Learning the non-linear mapping relationship between vibration features and speech spectra through a bidirectional long short-term memory network, and optimizing the network parameters by using a composite loss function of mean square error and spectral correlation.
[0015] Preferably, the implementation steps of the attention mechanism include: Calculating the semantic correlation degree between each lip movement frame and the front and back frames to generate a time attention score; Extracting multi-scale features of the lip movement area through spatial pyramid pooling to generate a spatial attention score; Multiplying the time and spatial attention scores to generate a comprehensive attention weight matrix; Performing weighted fusion on the key frame features according to the weight matrix and outputting the strengthened lip movement semantic coding.
[0016] Preferably, the sequence fusion method of the gated recurrent unit includes: Initializing the hidden state as the mean vector of multimodal features, calculating the activation values of the input gate, forget gate, and output gate at each time step, selectively updating the hidden state through the gating mechanism, and outputting the fused feature vector at the current step; Concatenating the fused feature vectors of all time steps into the final enhanced speech-assisted information stream.
[0017] Compared with the prior art, the beneficial effects of the present invention are: The rail transit hearing-impaired speech assistance system based on multimodal perception of the present invention has significant beneficial effects in many aspects.
[0018] In terms of information collection and processing, the multimodal data collection module can real-time obtain voice signals, environmental vibration signals, and visual lip-reading information in the rail transit environment, achieving a comprehensive perception of environmental information. The speech enhancement processing module uses a deep denoising network to separate the target speech from environmental noise, and adopts a dual-channel spectral mask algorithm and a time-domain waveform reconstruction technology to effectively improve the speech quality. For example, when the train is running, even if there are various noisy noises in the carriage, this module can clearly separate key voices such as train announcements from complex noises, making the speech content clearer and more distinguishable, greatly improving the recognizability of speech information, and laying a solid foundation for hearing-impaired people to accurately receive information subsequently.
[0019] The vibration signal analysis module extracts the vibration feature vector synchronized with the speech signal through a spectral correction algorithm and constructs a vibration-speech mapping relationship model. This function fully explores the potential connection between the vibration signal and the speech signal. In the actual rail transit scenario, the vibration generated by the train running often has certain rules and is related to the speech information. Through the processing of this module, the system can use the vibration signal to assist in understanding the speech content, providing another important information acquisition channel for hearing-impaired people in addition to speech and lip-reading, further enriching the information source and increasing the reliability of information acquisition.
[0020] The lip-reading recognition module performs dynamic semantic segmentation processing on the visual lip-reading information, uses a spatio-temporal convolutional network to extract the key frame sequence of mouth shapes, and generates the lip-reading text prediction result. In the complex carriage environment, this module can overcome adverse factors such as light and personnel occlusion and identify lip-reading as accurately as possible. For example, when the train announcement sound is small or some voices are drowned out by noise, the lip-reading recognition module can be used as a supplement to help hearing-impaired people obtain key information, enabling them to more comprehensively understand the communication content of people around them and the train announcement information in various situations.
[0021] The multi-modal fusion module inputs the denoised speech stream, vibration feature vectors, and lip text prediction results into the dynamic weighted fusion model, and outputs an enhanced speech auxiliary information stream. Among them, the adaptive weight allocation layer dynamically adjusts the weights of each modal information according to the characteristics of each modal information, such as the signal-to-noise ratio index of the speech, the mapping relationship between vibration and speech, and the semantic integrity of the lip text, realizing the intelligent fusion of multi-modal information. The cross-modal alignment mechanism ensures the temporal synchronization of multi-modal information, making the output enhanced speech auxiliary information stream more accurate and coherent. This multi-modal fusion method gives full play to the advantages of different information sources, overcomes the limitations of a single information source, greatly improves the adaptability of the system to complex environments, provides more accurate and rich speech auxiliary information for hearing-impaired people, significantly enhances their information acquisition ability and travel experience in the rail transit environment, and enhances the safety and convenience of their independent travel by rail transit. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 FIG. is a working principle diagram of the rail transit hearing-impaired speech assistance system based on multi-modal perception according to the present invention; Figure 2 FIG. is a schematic diagram of the dual-channel spectral mask algorithm of the deep noise reduction network; Figure 3 FIG. is a schematic diagram of extracting vibration feature vectors by the spectral correction algorithm; Figure 4 FIG. is a schematic diagram of the adaptive weight allocation layer of the dynamic weighted fusion model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0024] Please refer to Figures 1 - 4 , the present invention provides a rail transit hearing-impaired speech assistance system based on multi-modal perception, and its specific implementation solution is described in detail below.
[0025] The system mainly includes a multi-modal data acquisition module, a speech enhancement processing module, a vibration signal analysis module, a lip reading recognition module, and a multi-modal fusion module.
[0026] The multi-modal data acquisition module is responsible for real-time acquisition of a multi-source perception data set in the rail transit environment. In the rail transit scenario, such as at subway platforms, inside carriages, etc., this module uses specific acquisition devices, such as microphone arrays, to collect voice signals. The microphones can be arranged at different positions to ensure comprehensive acquisition of surrounding voice information; vibration sensors are used to collect environmental vibration signals, and these sensors are installed at parts where vibration may occur, such as the platform floor and carriage seats; visual lip-reading information is collected through cameras, and the shooting angles of the cameras are carefully adjusted to clearly capture the lip movements of the speaker.
[0027] The voice enhancement processing module processes the voice signals obtained from the multi-modal data acquisition module. This module uses a deep noise reduction network, through a dual-channel spectral masking algorithm and time-domain waveform reconstruction technology, to separate the target voice from environmental noise and generate a denoised voice stream.
[0028] The vibration signal analysis module works based on the collected environmental vibration signals. It uses a spectral correction algorithm to extract vibration feature vectors synchronized with the voice signals and constructs a vibration-voice mapping relationship model. Its purpose is to extract voice-related features from complex vibration signals to provide effective data for subsequent fusion.
[0029] The lip-reading recognition module processes the collected visual lip-reading information. Through dynamic semantic segmentation processing, it uses a spatio-temporal convolutional network to extract the key frame sequence of mouth shapes and generate a lip-reading text prediction result. In actual operation, this module analyzes the received lip-reading video stream.
[0030] The multi-modal fusion module fuses the denoised voice stream, vibration feature vectors, and lip-reading text prediction results obtained after processing by the above modules. This module uses a dynamic weighted fusion model, which includes an adaptive weight allocation layer and a cross-modal alignment mechanism, and finally outputs an enhanced voice-assisted information stream.
[0031] The following further elaborates on the specific implementation methods of each module of the present invention through specific embodiments.
[0032] Embodiment 1: This embodiment mainly focuses on the dual-channel spectral masking algorithm of the deep noise reduction network in the voice enhancement processing module. In the actual rail transit environment, voice signals are often interfered by various noises, such as the roar of train operation, the noise of the crowd, etc. To effectively separate the target voice from environmental noise, the dual-channel spectral masking algorithm of the deep noise reduction network operates according to the following steps.
[0033] For the input voice signal Perform a short-time Fourier transform (STFT) to generate a spectrogram. The formula for the short-time Fourier transform is:
[0034] where represents the frame index, represents the frequency index, is the window length, is the frame shift, is the window function, is the number of points for the Fourier transform. Through this transform, the speech signal in the time domain is converted to the frequency domain to obtain the spectrogram. Then, the spectrogram is divided into high-frequency and low-frequency subbands, so that targeted processing can be carried out according to the characteristics of different frequency bands.
[0035] The parallel convolutional gated unit is used to process the high-frequency and low-frequency subbands respectively. The parallel convolutional gated unit can perform convolutional operations on data in different frequency bands simultaneously, improving the processing efficiency. During the processing, a spectral mask coefficient matrix will be generated. Assume that the spectral mask coefficient matrix for the high-frequency band is and the spectral mask coefficient matrix for the low-frequency band is .
[0036] The original spectrum is filtered by the mask coefficient matrix to reconstruct the denoised speech waveform in the time domain. Specifically, for the high-frequency band spectrum , the spectrum after filtering is ; for the low-frequency band spectrum , the filtered spectrum is . Then, the filtered frequency-domain signal is converted back to the time domain through the inverse short-time Fourier transform (ISTFT) to obtain the denoised speech waveform , and the formula for the inverse short-time Fourier transform is:
[0037] where is the total number of frames.
[0038] In an actual application scenario, taking the subway platform as an example, when the train enters the station, the surrounding environmental noise is very large. Through this dual-channel spectral mask algorithm, the noise can be effectively removed, making the speech signal clearer and providing a good data basis for subsequent processing. At the same time, during the processing, due to the use of the parallel convolutional gated unit, the processing speed is greatly improved, meeting the real-time requirement.
[0039] Example 2: This example focuses on elaborating the specific steps of extracting the vibration feature vector synchronized with the speech signal by the spectral correction algorithm in the vibration signal analysis module. In the rail transit environment, the vibration signal source is complex and contains a lot of interference information unrelated to speech. Therefore, it is crucial to accurately extract the vibration feature vector synchronized with the speech signal.
[0040] Perform time-frequency transformation on the vibration signal to generate a time-frequency spectrum matrix. Through this transformation, the vibration signal in the time domain is converted to the frequency domain, facilitating subsequent analysis. Then, detect the formant positions synchronized with the speech signal. Formants are important features of speech signals and are closely related to speech pronunciation. Calculate the vibration energy distribution characteristics based on the formant positions. The specific calculation method is as follows:
[0041] where represents the vibration energy at frequency . Through this formula, the vibration energy distribution at different frequencies can be calculated, and then the interference components in the non-speech related frequency bands can be eliminated.
[0042] Adopt the sliding window difference method to extract the transient features in the vibration signal. Assume the size of the sliding window is . Within each window, calculate the difference between the vibration signals at adjacent times. In this way, the transient changes in the vibration signal can be highlighted, and these transient changes are often related to the changes in speech. Finally, combine the calculated vibration energy distribution characteristics and transient features to construct a multi-dimensional vibration feature vector. For example, the vibration energy and transient features at different frequencies can be arranged in a certain order to form a multi-dimensional vector.
[0043] In an actual scenario, such as in a train carriage, when a passenger speaks, the seat and the ground will generate weak vibrations. Through the above spectrum correction algorithm, the vibration feature vector synchronized with the speech can be accurately extracted from the complex vibration signal. These feature vectors can provide important information for subsequent multi-modal fusion, helping hearing-impaired people better understand the speech content.
[0044] Example 3: This example details the method of extracting the key frame sequence of lip movements by the spatio-temporal convolutional network in the lip reading module. In the rail transit scenario, with frequent personnel flow, the lip movements of the speaker are an important source of information. For hearing-impaired people, lip reading is one of the key ways to obtain information.
[0045] Perform key frame sampling on the continuous lip reading video stream to generate a group of time series images. Assume the lip reading video stream is . During the sampling process, select key frames at a certain time interval to obtain the group of time series images , . These key frames contain important information about the lip movement changes. Then, extract the spatial features and temporal correlation of the lip movement through a three-dimensional convolution kernel. The three-dimensional convolution kernel performs convolution operations simultaneously in the spatial and temporal dimensions, and its convolution formula is:
[0046] Among them is the output feature after convolution, is a three-dimensional convolution kernel, is the input image feature, are respectively the sizes of the convolution kernel in the directions. In this way, the spatial changes of lip shapes at different moments and the continuity features in time can be extracted.
[0047] The attention mechanism is adopted to strengthen the semantic weights of key lip shape change frames. Calculate the semantic correlation degrees between each lip shape frame and its front and back frames to generate time attention scores. Extract multi-scale features of the lip shape region through spatial pyramid pooling to generate spatial attention scores. Spatial pyramid pooling divides the lip shape region into sub-regions of different scales, performs pooling operations on each sub-region to obtain features of different scales. Assuming that the pooling results at different scales are respectively , the spatial attention score can be obtained by weighted calculation of these pooling results. For example , where are weights. Multiply the time and spatial attention scores to generate a comprehensive attention weight matrix . Weightedly fuse the key frame features according to the weight matrix, and output the strengthened lip shape semantic encoding. Let the key frame feature be , and the strengthened lip shape semantic encoding be .
[0048] In practical applications, such as in the broadcast area of a subway station, when staff convey information through the broadcast, the spatio-temporal convolutional network combined with the attention mechanism can accurately extract the key frame sequence of lip shapes, generate more accurate lip text prediction results, and provide effective information support for hearing-impaired people.
[0049] Example 4: This example details the adaptive weight assignment layer of the dynamic weighted fusion model. In the multi-modal fusion process, in order to give full play to the advantages of different modal data, reasonable weights need to be assigned to the denoised speech stream, vibration feature vectors, and lip text prediction results.
[0050] Calculate the signal-to-noise ratio index for the denoised speech stream to generate the speech confidence weight. Let the denoised speech stream be , and its noise signal be , the signal-to-noise ratio index . According to the signal-to-noise ratio index, the speech confidence weight can be calculated through a function. For example , where is a constant used to adjust the range of the weight. In this way, the higher the signal-to-noise ratio, the greater the speech confidence weight, indicating that the reliability of this speech stream is higher.
[0051] Calculate the vibration assistance weight according to the mapping relationship between the vibration feature vector and the speech signal. Assume the vibration feature vector is , and the feature vector of the speech signal is . The mapping relationship between them can be represented by a function , that is . The vibration assistance weight can be determined by calculating the similarity between the two. For example , where represents the dot product of vectors, and represents the norm of the vector. The higher the similarity, the greater the vibration assistance weight, indicating that the vibration feature vector has a stronger auxiliary effect on the speech signal.
[0052] Generate the lip-reading correction weight based on the semantic integrity of the lip-reading text prediction result. Let the lip-reading text prediction result be , and the semantic integrity can be measured by the similarity with the reference text . For example, use the edit distance algorithm to calculate the similarity . The lip-reading correction weight , where is a constant. The higher the semantic integrity, the greater the lip-reading correction weight, indicating that the lip-reading text prediction result is more important for the correction of speech assistance.
[0053] Perform dynamic normalization on the three types of weights through a differentiable optimization algorithm to generate a fusion weight matrix. Let the speech confidence weight, vibration assistance weight, and lip-reading correction weight obtained through the above calculations be , , respectively. The fusion weight matrix can be calculated by the following formula: . In this way, the sum of the three types of weights is 1, achieving dynamic normalization and ensuring that multi-modal data can be reasonably fused in different rail transit scenarios.
[0054] In an actual scenario, such as in a subway carriage, when the surrounding environmental noise is large, the speech confidence weight may decrease, while the vibration assistance weight and lip-reading correction weight will increase accordingly. Through the adaptive weight allocation layer, the weights can be dynamically adjusted to improve the accuracy of the speech assistance information flow.
[0055] Example 5: This example elaborates on the implementation method of the cross-modal alignment mechanism and the sequence fusion method of the gated recurrent unit. During the multi-modal fusion process, due to the differences in the acquisition and processing times of the speech stream, vibration feature vector, and lip-reading text, cross-modal alignment is required to ensure data synchronization.
[0056] Map the speech stream, vibration feature vectors, and lip text to a unified time coordinate system. Assume the time series of the speech stream is , the time series of the vibration feature vectors is , and the time series of the lip text is . Through a time calibration algorithm, map them to a common time coordinate system . For example, the method of linear interpolation can be used to align the time points of different modality data on the common time axis.
[0057] Calculate the time offset of the multimodal data through the dynamic time warping algorithm and compensate for the alignment error. The core idea of the dynamic time warping algorithm is to find the optimal matching path between two time series so that they are as aligned as possible in time. Let the speech stream feature sequence be , the vibration feature vector sequence be , and the lip text feature sequence be . The time offsets calculated through the dynamic time warping algorithm are and respectively. According to these offsets, perform time compensation on the vibration feature vectors and lip text features. For example, translate the vibration feature vector in time by , and translate the lip text feature in time by .
[0058] Use gated recurrent units to perform sequence fusion on the aligned multimodal data. Concatenate the fusion feature vectors at all time steps into the final enhanced speech-aided information stream.
[0059] In practical applications, such as in the transfer passage of a subway station where there are dense crowds and complex information, through the cross-modal alignment mechanism and the sequence fusion method of gated recurrent units, the speech stream, vibration feature vectors, and lip text can be accurately fused to provide clear and accurate speech-aided information for hearing-impaired people.
[0060] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device.
[0061] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will appreciate that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A rail transit hearing-impaired voice assistance system based on multimodal perception, characterized in that: include: Multimodal data acquisition module: used to obtain a multi-source perception data set in the rail transit environment in real time, wherein the multi-source perception data includes speech signals, environmental vibration signals and visual lip reading information; A speech enhancement processing module: configured to separate the target speech from the environmental noise based on the speech signal through a deep noise reduction network to generate a denoised speech stream; the deep noise reduction network adopts a dual-channel spectrum mask algorithm and a time domain waveform reconstruction technology; Vibration signal analysis module: used to extract the vibration feature vector synchronized with the voice signal according to the environmental vibration signal through a spectrum correction algorithm, and to construct a vibration-voice mapping relationship model; Lip reading recognition module: used to perform dynamic semantic segmentation processing on the visual lip reading information, extract lip shape key frame sequence using spatiotemporal convolutional network, and generate lip reading text prediction results; Multimodal fusion module: input the denoised speech stream, vibration feature vector and lip reading text prediction results into a dynamic weighted fusion model, and output an enhanced speech auxiliary information stream; the dynamic weighted fusion model includes an adaptive weight allocation layer and a cross-modal alignment mechanism.
2. The rail transit hearing-impaired voice assistance system based on multimodal perception according to claim 1 is characterized in that: The dual-channel spectral mask algorithm of the deep denoising network includes: Perform short-time Fourier transform on the input speech signal to generate a spectrum diagram, and divide it into high-frequency and low-frequency sub-bands; Parallel convolutional gating units are used to process high-frequency and low-frequency sub-bands separately to generate spectrum mask coefficient matrices; The original spectrum is filtered through the mask coefficient matrix to reconstruct the denoised time domain speech waveform.
3. The rail transit hearing-impaired voice assistance system based on multimodal perception according to claim 1 is characterized in that: The step of extracting the vibration feature vector synchronized with the speech signal by the spectrum correction algorithm comprises: Perform time-frequency transformation on the vibration signal to generate a time-frequency spectrum matrix, detect the position of the resonance peak synchronized with the speech signal, calculate the vibration energy distribution characteristics based on the resonance peak position, and eliminate the interference components of the non-speech related frequency band; The sliding window difference method is used to extract the transient features in the vibration signal and construct a multi-dimensional vibration feature vector.
4. The rail transit hearing-impaired voice assistance system based on multimodal perception according to claim 1 is characterized in that: The method for extracting a lip-sync key frame sequence using a spatiotemporal convolutional network comprises: The continuous lip reading video stream is sampled by key frames to generate a time series image group, and the spatial characteristics and temporal correlation of lip movement are extracted through a three-dimensional convolution kernel; The attention mechanism is used to strengthen the semantic weight of key lip shape change frames and generate lip shape semantic coding sequences.
5. The rail transit hearing-impaired voice assistance system based on multimodal perception according to claim 1 is characterized in that: The adaptive weight allocation layer of the dynamic weighted fusion model includes: Calculate the signal-to-noise ratio index for the denoised speech stream and generate the speech confidence weight; Calculating the vibration auxiliary weight according to the mapping relationship between the vibration feature vector and the speech signal; Generate lip reading correction weights based on the semantic completeness of the lip reading text prediction results; The three types of weights are dynamically normalized through a differentiable optimization algorithm to generate a fusion weight matrix.
6. The rail transit hearing-impaired voice assistance system based on multimodal perception according to claim 5 is characterized in that: The implementation method of the cross-modal alignment mechanism includes: Mapping speech stream, vibration feature vector and lip reading text to a unified time coordinate system; The time offset of multimodal data is calculated through the dynamic time warping algorithm, and the alignment error is compensated; The gated recurrent unit is used to perform sequence fusion on the aligned multimodal data to generate a synchronized enhanced speech stream.
7. The rail transit hearing-impaired voice assistance system based on multimodal perception according to claim 2 is characterized in that: The optimization method of the time domain waveform reconstruction technology includes: The filtered spectrum is subjected to phase recovery processing, the iterative projection algorithm is used to optimize the continuity of the time domain waveform, the auditory quality of the reconstructed waveform is evaluated through the adversarial generative network, and the spectrum mask parameters are iteratively adjusted; The loudness of the reconstructed speech is compensated in combination with the human hearing characteristic curve.
8. The rail transit hearing-impaired voice assistance system based on multimodal perception according to claim 3 is characterized in that: The method for constructing the vibration-speech mapping relationship model comprises: Collect multiple sets of synchronous vibration signals and pure voice samples to build a training data set; The nonlinear mapping relationship between vibration features and speech spectrum is learned through a bidirectional long short-term memory network, and the network parameters are optimized using the composite loss function of mean square error and spectrum correlation.
9. The rail transit hearing-impaired voice assistance system based on multimodal perception according to claim 4 is characterized in that: The implementation steps of the attention mechanism include: Calculate the semantic relevance of each lip frame with the previous and next frames to generate a temporal attention score; The multi-scale features of the mouth shape area are extracted through spatial pyramid pooling to generate a spatial attention score; Multiply the temporal and spatial attention scores to generate a comprehensive attention weight matrix; The key frame features are weightedly fused according to the weight matrix, and the enhanced lip semantic coding is output.
10. The rail transit hearing-impaired voice assistance system based on multimodal perception according to claim 6, characterized in that: The sequence fusion method of the gated recurrent unit includes: Initialize the hidden state to the mean vector of multimodal features, calculate the activation values of the input gate, forget gate, and output gate at each time step, selectively update the hidden state through the gating mechanism, and output the fused feature vector of the current step; The fused feature vectors of all time steps are concatenated into the final enhanced speech auxiliary information stream.
Citation Information
Patent Citations
Method and device for speech recognition
CN106157956A
Sound-vibration fusion signal identification method and system, computer equipment and medium
CN117292494A
Multi-mode adaptive pickup method and system, earphone and storage medium
CN118764765A
Multi-mode synchronous fusion speech recognition system
CN119296523A
Voice detection method based on millimeter wave radar
CN119780912A
Cited By
Video micro-vibration signal correction method
CN120705481A
Intelligent social assistance system and method for tinnitus patient based on multi-modal fusion
CN120954436A
Intelligent lock identity authentication method and system for multi-mode voiceprint verification
CN121075020A
Method for sensing track state based on multi-modal monitoring data fusion of ballastless track
CN121527721A
Virtual anchor real-time driving system based on facial motion capture
CN121842342A