Rail Transit Hearing-Impaired Speech Assistance System Based on Multimodal Perception
Through the multimodal perception system, combined with voice enhancement, vibration signal analysis and lip recognition, the problem of inconvenient information acquisition for hearing-impaired people in the rail transit environment is solved, and clear voice assistance is achieved in complex noise and light environments, improving the information acquisition ability and safety of hearing-impaired people in the rail transit.
Patent Information
- Application Number
- CN202510535859.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-04-27
AI Technical Summary
In the rail transit environment, it is difficult for people with hearing impairments to clearly obtain voice information such as train radio, and existing auxiliary equipment cannot synchronize voice content in real time and accurately, and lip recognition is affected by light and occlusion, resulting in inconvenient information acquisition.
The multimodal perception system is adopted, integrating speech enhancement processing, vibration signal analysis and lip recognition modules, and synchronous and intelligent fusion of multimodal information is achieved through deep noise reduction network, spectrum correction algorithm and spatiotemporal convolution network, combined with dynamic weighted fusion model.
Clearly separate voice content in complex noise environments, use vibration signals to assist in understanding voice, overcome the influence of light and occlusion, provide accurate and rich voice assist information, and improve the information acquisition ability and travel safety of hearing impaired people.
Smart Images

Figure CN120071950B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of rail transit hearing assistance, and in particular to a rail transit hearing-impaired voice assistance system based on multimodal perception. Background Art
[0002] In modern society, rail transit has become an important means of transportation for people. Its advantages, such as efficiency, convenience, and large capacity, greatly meet people's daily travel needs. However, for people with hearing impairments, obtaining voice information in rail transit environments faces many challenges.
[0003] During train operation, the ambient noise is complex. The noise from the wheels rubbing against the tracks, the hum of the ventilation system, and the chatter of passengers all intertwine to create a formidable background din. This noise severely interferes with the transmission of voice information, such as train announcements. Even for people with normal hearing, it can be difficult to clearly discern the content during certain noisy periods. People with hearing impairments are even more severely affected, often unable to discern crucial information, such as arrival notices and transfer instructions, from the cacophony. This creates significant inconvenience for their journeys and can even lead to missing stops and being unable to transfer.
[0004] Traditional assistive devices have significant limitations when dealing with rail transit environments. For example, simple subtitle display devices typically only provide text information in a fixed format and cannot accurately synchronize with the voice content in real time. In practice, the speed and content of train announcements vary widely, and subtitle displays may experience delays and incomplete information, making it difficult to meet the needs of hearing-impaired people for timely information. Moreover, such devices rely solely on single visual information and fail to fully consider other information cues in the environment, failing to fully enhance the hearing-impaired's understanding and reception of voice information.
[0005] Furthermore, lip reading recognition technology faces challenges in the complex environment of rail transit. For one thing, the lighting conditions within train cars are complex and variable. Light intensity and angle vary at different times and locations within the car, which can degrade the quality of lip reading images captured by the camera and affect lip reading recognition accuracy. Furthermore, passengers' positions and postures vary. Sometimes the speaker's mouth may be partially obscured or in a blind spot within the camera, making effective lip reading recognition difficult. Summary of the Invention
[0006] The purpose of the present invention is to provide a rail transit hearing-impaired voice assistance system based on multimodal perception to solve the problems raised in the above background technology.
[0007] To achieve the above objectives, the present invention provides the following technical solution: a rail transit hearing-impaired voice assistance system based on multimodal perception, the system comprising:
[0008] Multimodal data acquisition module: used to acquire multi-source perception data sets in the rail transit environment in real time, wherein the multi-source perception data includes voice signals, environmental vibration signals and visual lip reading information;
[0009] A speech enhancement processing module is configured to separate the target speech from the ambient noise based on the speech signal through a deep noise reduction network, thereby generating a denoised speech stream; the deep noise reduction network adopts a dual-channel spectrum masking algorithm and time-domain waveform reconstruction technology;
[0010] Vibration signal analysis module: used to extract the vibration feature vector synchronized with the voice signal based on the environmental vibration signal through a spectrum correction algorithm, and construct a vibration-voice mapping relationship model;
[0011] Lip reading recognition module: used to perform dynamic semantic segmentation processing on the visual lip reading information, extract the lip shape key frame sequence using a spatiotemporal convolutional network, and generate lip reading text prediction results;
[0012] Multimodal fusion module: inputs the denoised speech stream, vibration feature vector and lip reading text prediction results into a dynamic weighted fusion model, and outputs an enhanced speech auxiliary information stream; the dynamic weighted fusion model includes an adaptive weight allocation layer and a cross-modal alignment mechanism.
[0013] Preferably, the dual-channel spectral mask algorithm of the deep noise reduction network includes:
[0014] Perform short-time Fourier transform on the input speech signal to generate a spectrogram, and divide it into high-frequency and low-frequency sub-bands;
[0015] Parallel convolutional gating units are used to process high-frequency and low-frequency sub-bands separately to generate spectrum mask coefficient matrices;
[0016] The original spectrum is filtered through the mask coefficient matrix to reconstruct the denoised time domain speech waveform.
[0017] Preferably, the step of extracting the vibration feature vector synchronized with the speech signal by the spectrum correction algorithm includes:
[0018] Perform time-frequency transformation on the vibration signal to generate a time-frequency spectrum matrix, detect the position of the formant synchronized with the speech signal, calculate the vibration energy distribution characteristics based on the formant position, and eliminate interference components in non-speech related frequency bands;
[0019] The sliding window difference method is used to extract transient features from vibration signals and construct multi-dimensional vibration feature vectors.
[0020] Preferably, the method for extracting a lip-sync keyframe sequence using a spatiotemporal convolutional network includes:
[0021] Keyframe sampling is performed on the continuous lip movement video stream to generate a time series image group, and the spatial characteristics and temporal correlation of lip movement are extracted through a three-dimensional convolution kernel;
[0022] The attention mechanism is used to enhance the semantic weight of key lip-changing frames and generate lip-shaped semantic coding sequences.
[0023] Preferably, the adaptive weight allocation layer of the dynamic weighted fusion model includes:
[0024] Calculate the signal-to-noise ratio index for the denoised speech stream and generate the speech confidence weight;
[0025] Calculating the vibration auxiliary weight according to the mapping relationship between the vibration feature vector and the speech signal;
[0026] Generate lip reading correction weights based on the semantic completeness of the lip reading text prediction results;
[0027] The three types of weights are dynamically normalized through a differentiable optimization algorithm to generate a fusion weight matrix.
[0028] Preferably, the implementation method of the cross-modal alignment mechanism includes:
[0029] Mapping the speech stream, vibration feature vector and lip reading text to a unified time coordinate system;
[0030] Calculate the time offset of multimodal data and compensate for alignment errors through dynamic time warping algorithm;
[0031] The gated recurrent unit is used to perform sequence fusion on the aligned multimodal data to generate a synchronized enhanced speech stream.
[0032] Preferably, the optimization method of the time domain waveform reconstruction technology includes:
[0033] The filtered spectrum is subjected to phase recovery processing, and the continuity of the time domain waveform is optimized using an iterative projection algorithm. The auditory quality of the reconstructed waveform is evaluated using a generative adversarial network, and the spectrum mask parameters are iteratively adjusted.
[0034] The loudness of the reconstructed speech is compensated in combination with the human ear auditory characteristic curve.
[0035] Preferably, the method for constructing the vibration-speech mapping relationship model includes:
[0036] Collect multiple sets of synchronized vibration signals and pure speech samples to build a training data set;
[0037] The nonlinear mapping relationship between vibration features and speech spectrum is learned through a bidirectional long short-term memory network, and the network parameters are optimized using a composite loss function of mean square error and spectrum correlation.
[0038] Preferably, the steps for implementing the attention mechanism include:
[0039] Calculate the semantic relevance of each lip frame with the previous and next frames to generate a temporal attention score;
[0040] Extract multi-scale features of the mouth shape area through spatial pyramid pooling and generate spatial attention scores;
[0041] Multiply the temporal and spatial attention scores to generate a comprehensive attention weight matrix;
[0042] The key frame features are weightedly fused according to the weight matrix and the enhanced lip semantic coding is output.
[0043] Preferably, the sequence fusion method of the gated recurrent unit includes:
[0044] Initialize the hidden state to the mean vector of multimodal features, calculate the activation values of the input gate, forget gate, and output gate at each time step, selectively update the hidden state through the gating mechanism, and output the fused feature vector of the current step;
[0045] The fused feature vectors of all time steps are concatenated into the final enhanced speech auxiliary information stream.
[0046] Compared with the prior art, the present invention has the following beneficial effects:
[0047] The rail transit hearing-impaired voice assistance system based on multimodal perception of the present invention has many significant beneficial effects.
[0048] In terms of information acquisition and processing, the multimodal data acquisition module can acquire voice signals, ambient vibration signals, and visual lip readings in the rail transit environment in real time, achieving comprehensive environmental awareness. The speech enhancement processing module utilizes a deep noise reduction network to separate target speech from ambient noise, employing a dual-channel spectral masking algorithm and time-domain waveform reconstruction technology to effectively improve speech quality. For example, while the train is running, even in the presence of various noisy noises within the carriage, this module can clearly separate key speech, such as train announcements, from the complex noise, making the speech content more legible and significantly improving the recognizability of speech information, laying a solid foundation for hearing-impaired individuals to accurately receive information.
[0049] The vibration signal analysis module uses a spectrum correction algorithm to extract vibration feature vectors synchronized with the speech signal and constructs a vibration-speech mapping model. This function fully exploits the potential connection between vibration and speech signals. In actual rail transit scenarios, the vibrations generated by train movement often exhibit certain patterns and are correlated with speech information. Through processing by this module, the system can use vibration signals to assist in understanding speech content, providing hearing-impaired individuals with an important alternative to speech and lip reading to obtain information, further enriching information sources and increasing the reliability of information acquisition.
[0050] The lip reading recognition module performs dynamic semantic segmentation on visual lip reading information, employing a spatiotemporal convolutional network to extract lip shape keyframe sequences and generate lip reading text predictions. This module can accurately recognize lip readings in complex train environments, overcoming obstacles such as lighting and occlusion. For example, when train announcements are low or partially drowned out by noise, the lip reading recognition module can supplement this information, helping hearing-impaired individuals obtain key information. This allows them to more fully understand the conversations of those around them and train announcements in all situations.
[0051] The multimodal fusion module inputs the denoised speech stream, vibration feature vectors and lip reading text prediction results into a dynamic weighted fusion model and outputs an enhanced speech auxiliary information stream. The adaptive weight allocation layer dynamically adjusts the weight of each modal information according to the characteristics of each modal information, such as the signal-to-noise ratio index of speech, the mapping relationship between vibration and speech, and the semantic completeness of lip reading text, thereby realizing the intelligent fusion of multimodal information. The cross-modal alignment mechanism ensures the temporal synchronization of multimodal information, making the output enhanced speech auxiliary information stream more accurate and coherent. This multimodal fusion method fully utilizes the advantages of different information sources, overcomes the limitations of a single information source, greatly improves the system's adaptability to complex environments, provides more accurate and rich speech auxiliary information for people with hearing impairments, significantly improves their information acquisition ability and travel experience in the rail transit environment, and enhances their safety and convenience when traveling alone on the rail transit. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 This is a working principle diagram of the rail transit hearing-impaired voice assistance system based on multimodal perception according to the present invention;
[0053] Figure 2 Schematic diagram of the dual-channel spectral mask algorithm for deep noise reduction network;
[0054] Figure 3 Schematic diagram of extracting vibration eigenvectors for spectrum correction algorithm;
[0055] Figure 4Schematic diagram of the adaptive weight assignment layer of the dynamic weighted fusion model. DETAILED DESCRIPTION
[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0057] See also Figures 1-4 The present invention provides a rail transit hearing-impaired voice assistance system based on multimodal perception, and its specific implementation scheme is elaborated in detail below.
[0058] The system mainly includes a multimodal data acquisition module, a speech enhancement processing module, a vibration signal analysis module, a lip reading recognition module and a multimodal fusion module.
[0059] The multimodal data acquisition module is responsible for acquiring multi-source sensory data sets in real time within the rail transit environment. In rail transit scenarios, such as subway platforms and train interiors, the module utilizes specialized acquisition equipment, such as microphone arrays, to collect voice signals. Microphones can be positioned in various locations to ensure comprehensive coverage of surrounding voice information. Environmental vibration signals are collected using vibration sensors installed in areas likely to vibrate, such as platform floors and train seats. Furthermore, visual lip reading information is captured through cameras, whose angles are carefully adjusted to clearly capture the speaker's lip movements.
[0060] The speech enhancement processing module processes the speech signals acquired by the multimodal data acquisition module. It uses a deep noise reduction network, a dual-channel spectral masking algorithm, and time-domain waveform reconstruction technology to separate the target speech from the ambient noise, generating a denoised speech stream.
[0061] The vibration signal analysis module operates based on the collected ambient vibration signals. It uses a spectrum correction algorithm to extract vibration feature vectors synchronized with the speech signal and constructs a vibration-speech mapping model. Its goal is to extract speech-related features from complex vibration signals, providing valid data for subsequent fusion.
[0062] The lip reading recognition module processes the collected visual lip reading information. Through dynamic semantic segmentation, it uses a spatiotemporal convolutional network to extract lip keyframe sequences and generate lip reading text predictions. In practice, this module analyzes the received lip reading video stream.
[0063] The multimodal fusion module combines the denoised speech stream, vibration feature vectors, and lip-reading text predictions generated by the above modules. This module uses a dynamic weighted fusion model, including an adaptive weight allocation layer and a cross-modal alignment mechanism, to ultimately output an enhanced speech auxiliary information stream.
[0064] The specific implementation of each module of the present invention is further described below through specific embodiments.
[0065] Example 1: This example focuses on the dual-channel spectral masking algorithm of the deep noise reduction network in the speech enhancement processing module. In actual rail transit environments, speech signals are often interfered with by various noises, such as the roar of running trains and the clamor of crowds. To effectively separate the target speech from the ambient noise, the dual-channel spectral masking algorithm of the deep noise reduction network operates according to the following steps.
[0066] For input speech signal Perform short-time Fourier transform (STFT) to generate the spectrum. The formula for short-time Fourier transform is:
[0067]
[0068] in Indicates the frame index, represents the frequency index, is the window length, It's frame shift. is the window function, is the number of Fourier transform points. This transform converts the time-domain speech signal into the frequency domain, generating a spectrogram. The spectrogram is then divided into high-frequency and low-frequency subbands, allowing for targeted processing based on the characteristics of each frequency band.
[0069] The parallel convolution gating unit is used to process the high frequency band and low frequency band sub-bands separately. The parallel convolution gating unit can perform convolution operations on data of different frequency bands at the same time, improving processing efficiency. During the processing, a spectrum mask coefficient matrix is generated. Assume that the spectrum mask coefficient matrix of the high frequency band is , the spectrum mask coefficient matrix of the low frequency band is .
[0070] The original spectrum is filtered through the mask coefficient matrix to reconstruct the denoised time domain speech waveform. Specifically, the high frequency spectrum , the spectrum after filtering is ; For low frequency spectrum , the filtered spectrum is The filtered frequency domain signal is then converted back to the time domain through the inverse short-time Fourier transform (ISTFT) to obtain the denoised speech waveform. , the inverse short-time Fourier transform formula is:
[0071]
[0072] in is the total number of frames.
[0073] In practical applications, such as subway platforms, when a train arrives, the surrounding noise is very high. This dual-channel spectral masking algorithm can effectively remove the noise, making the voice signal clearer and providing a good data foundation for subsequent processing. Furthermore, the use of parallel convolutional gating units during processing greatly improves processing speed, meeting real-time requirements.
[0074] Example 2: This example focuses on the specific steps used by the spectrum correction algorithm in the vibration signal analysis module to extract vibration eigenvectors synchronized with the speech signal. In rail transit environments, vibration signals come from complex sources and contain a lot of interference information unrelated to speech. Therefore, accurately extracting vibration eigenvectors synchronized with the speech signal is crucial.
[0075] The vibration signal is transformed into a time-frequency spectrum matrix. This transform converts the vibration signal from the time domain into the frequency domain, facilitating subsequent analysis. Next, the position of the formant synchronized with the speech signal is detected. Formants are important features of speech signals and are closely related to speech pronunciation. The vibration energy distribution characteristics are calculated based on the formant positions. The specific calculation method is as follows:
[0076]
[0077] in Indicates the frequency The vibration energy at the location can be calculated by this formula to calculate the distribution of vibration energy at different frequencies, thereby eliminating the interference components in the non-speech related frequency bands.
[0078] The sliding window difference method is used to extract the transient features in the vibration signal. Assume that the sliding window size is Within each window, the difference between vibration signals at adjacent moments is calculated. This method highlights transient changes in the vibration signal, which are often correlated with changes in speech. Finally, the calculated vibration energy distribution features and transient features are combined to construct a multidimensional vibration feature vector. For example, the vibration energy and transient features at different frequencies can be arranged in a specific order to form a multidimensional vector.
[0079] In real-world scenarios, such as on a train, when passengers speak, the seats and floor generate subtle vibrations. The spectral correction algorithm described above accurately extracts vibration feature vectors synchronized with speech from complex vibration signals. These feature vectors provide crucial information for subsequent multimodal fusion, helping hearing-impaired individuals better understand speech content.
[0080] Example 3: This example details the method for extracting lip-reading keyframe sequences using a spatiotemporal convolutional network in the lip reading recognition module. In rail transit scenarios, with frequent personnel movement, the speaker's lip movements are an important source of information. For the hearing-impaired, lip reading recognition is one of the key ways to obtain this information.
[0081] Perform key frame sampling on the continuous lip reading video stream to generate a time series image group. Assume that the lip reading video stream is , during the sampling process, at a certain time interval Select key frames to obtain time series image groups , These key frames contain important information about the changes in lip shape. Then, the spatial characteristics and temporal correlation of lip shape movement are extracted through the three-dimensional convolution kernel. The three-dimensional convolution kernel performs convolution operations in both spatial and temporal dimensions. The convolution formula is:
[0082]
[0083] in is the output feature after convolution, is the three-dimensional convolution kernel, is the input image feature, The convolution kernel is In this way, the spatial variation of the mouth shape at different moments and the temporal continuity features can be extracted.
[0084] The attention mechanism is used to strengthen the semantic weight of the key lip shape change frames. The semantic correlation between each lip shape frame and the previous and next frames is calculated to generate a temporal attention score. The multi-scale features of the lip shape area are extracted through spatial pyramid pooling to generate a spatial attention score. Spatial pyramid pooling divides the lip shape area into sub-areas of different scales, and performs pooling operations on each sub-area to obtain features of different scales. Assume that the pooling results at different scales are , spatial attention score It can be obtained by weighting these pooling results, for example ,in Is the weight. Multiply the temporal and spatial attention scores to generate a comprehensive attention weight matrix The key frame features are weighted and fused according to the weight matrix, and the enhanced lip semantic coding is output. Assume that the key frame features are , the enhanced lip semantic encoding is .
[0085] In practical applications, such as in the broadcasting area of a subway station, when staff convey information through broadcasting, the spatiotemporal convolutional network combined with the attention mechanism can accurately extract the lip key frame sequence and generate more accurate lip reading text prediction results, providing effective information support for the hearing impaired.
[0086] Example 4: This example introduces the adaptive weight allocation layer of the dynamic weighted fusion model in detail. In the multimodal fusion process, in order to give full play to the advantages of different modal data, it is necessary to assign reasonable weights to the denoised speech stream, vibration feature vectors and lip reading text prediction results.
[0087] Calculate the signal-to-noise ratio index for the denoised speech stream and generate the speech confidence weight. Suppose the denoised speech stream is , and its noise signal is , signal-to-noise ratio index According to the signal-to-noise ratio index, the speech confidence weight This can be calculated using a function, such as ,in is a constant used to adjust the weight range. Thus, the higher the signal-to-noise ratio, the greater the voice confidence weight, indicating a higher reliability of the voice stream.
[0088] The vibration auxiliary weight is calculated based on the mapping relationship between the vibration feature vector and the speech signal. Assume that the vibration feature vector is , the eigenvector of the speech signal is , the mapping relationship between them can be achieved through a function Indicates that . Vibration Assisted Weights This can be determined by calculating the similarity between the two, e.g. ,in represents the dot product of vectors, Represents the norm of the vector. The higher the similarity, the greater the vibration auxiliary weight, indicating that the vibration feature vector has a stronger auxiliary effect on the speech signal.
[0089] The lip reading correction weight is generated based on the semantic completeness of the lip reading text prediction result. Assume that the lip reading text prediction result is , semantic completeness can be measured by comparing it with the reference text For example, the edit distance algorithm is used to calculate the similarity Lip reading correction weight ,in is a constant. The higher the semantic completeness, the greater the lip reading correction weight, indicating that the lip reading text prediction result is more important for the correction of speech assistance.
[0090] The three types of weights are dynamically normalized by the differentiable optimization algorithm to generate a fusion weight matrix. Assume that the speech confidence weight, vibration assistance weight and lip reading correction weight obtained by the above calculation are 、 、 , fusion weight matrix It can be calculated by the following formula: ,In this way, the sum of the three types of weights is made 1,,dynamic normalization is achieved, and multimodal data can be reasonably,fused in different rail transit scenarios.
[0091] In actual scenarios, such as in a subway car, when the surrounding noise is large, the voice confidence weight may decrease, while the vibration assistance weight and lip reading correction weight will increase accordingly. Through the adaptive weight allocation layer, the weights can be dynamically adjusted to improve the accuracy of the voice assistance information flow.
[0092] Example 5: This example describes a method for implementing a cross-modal alignment mechanism and a sequence fusion method for gated recurrent units. During the multimodal fusion process, due to differences in the acquisition and processing time of the speech stream, vibration feature vectors, and lip reading text, cross-modal alignment is required to ensure data synchronization.
[0093] Map the speech stream, vibration feature vector and lip reading text to a unified time coordinate system. Assume that the time series of the speech stream is , the time series of vibration eigenvectors is , the time series of lip reading text is Through the time calibration algorithm, they are uniformly mapped to a common time coordinate system For example, linear interpolation can be used to align the time points of different modal data onto a common time axis.
[0094] The dynamic time warping algorithm is used to calculate the time offset of multimodal data and compensate for the alignment error. The core idea of the dynamic time warping algorithm is to find the optimal matching path between two time series so that they are aligned as much as possible in time. Suppose the speech stream feature sequence is , the vibration eigenvector sequence is , the lip reading text feature sequence is , the time offsets calculated by the dynamic time warping algorithm are and According to these offsets, the vibration feature vector and the lip reading text feature are time-compensated. For example, the vibration feature vector Shift in time , the lip reading text features Shift in time .
[0095] The gated recurrent unit is used to perform sequence fusion on the aligned multimodal data, and the fused feature vectors of all time steps are concatenated into the final enhanced speech auxiliary information stream.
[0096] In practical applications, such as in transfer passages at subway stations, where there are dense crowds and complex information, the cross-modal alignment mechanism and the sequence fusion method of gated recurrent units can accurately fuse speech streams, vibration feature vectors, and lip reading text to provide clear and accurate speech auxiliary information for the hearing impaired.
[0097] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0098] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A rail transit hearing-impaired voice assistance system based on multimodal perception, characterized in that: include: Multimodal data acquisition module: used to acquire multi-source perception data sets in the rail transit environment in real time, wherein the multi-source perception data includes voice signals, environmental vibration signals and visual lip reading information; A speech enhancement processing module is configured to separate the target speech from the ambient noise based on the speech signal through a deep noise reduction network, thereby generating a denoised speech stream; the deep noise reduction network adopts a dual-channel spectrum masking algorithm and time-domain waveform reconstruction technology; Vibration signal analysis module: used to extract the vibration feature vector synchronized with the voice signal based on the environmental vibration signal through a spectrum correction algorithm, and construct a vibration-voice mapping relationship model; Lip reading recognition module: used to perform dynamic semantic segmentation processing on the visual lip reading information, extract the lip shape key frame sequence using a spatiotemporal convolutional network, and generate lip reading text prediction results; Multimodal fusion module: inputs the denoised speech stream, vibration feature vector and lip reading text prediction results into a dynamic weighted fusion model, and outputs an enhanced speech auxiliary information stream; The dynamic weighted fusion model includes an adaptive weight distribution layer and a cross-modal alignment mechanism; The adaptive weight allocation layer of the dynamic weighted fusion model includes: Calculate the signal-to-noise ratio index for the denoised speech stream and generate the speech confidence weight; Calculating the vibration auxiliary weight according to the mapping relationship between the vibration feature vector and the speech signal; Generate lip reading correction weights based on the semantic completeness of the lip reading text prediction results; The three types of weights are dynamically normalized through a differentiable optimization algorithm to generate a fusion weight matrix.
2. The rail transit hearing-impaired voice assistance system based on multimodal perception according to claim 1 is characterized in that: The dual-channel spectral mask algorithm of the deep noise reduction network includes: Perform short-time Fourier transform on the input speech signal to generate a spectrogram, and divide it into high-frequency and low-frequency sub-bands; Parallel convolutional gating units are used to process high-frequency and low-frequency sub-bands separately to generate spectrum mask coefficient matrices; The original spectrum is filtered through the mask coefficient matrix to reconstruct the denoised time domain speech waveform.
3. The rail transit hearing-impaired voice assistance system based on multimodal perception according to claim 1 is characterized in that: The step of extracting the vibration feature vector synchronized with the speech signal by the spectrum correction algorithm comprises: Perform time-frequency transformation on the vibration signal to generate a time-frequency spectrum matrix, detect the position of the formant synchronized with the speech signal, calculate the vibration energy distribution characteristics based on the formant position, and eliminate interference components in non-speech related frequency bands; The sliding window difference method is used to extract transient features from vibration signals and construct multi-dimensional vibration feature vectors.
4. The rail transit hearing-impaired voice assistance system based on multimodal perception according to claim 1 is characterized in that: The method for extracting a lip-sync key frame sequence using a spatiotemporal convolutional network includes: Keyframe sampling is performed on the continuous lip reading video stream to generate a time series image group, and the spatial characteristics and temporal correlation of lip movement are extracted through a three-dimensional convolution kernel; The attention mechanism is used to enhance the semantic weight of key lip-changing frames and generate lip-shaped semantic coding sequences.
5. The rail transit hearing-impaired voice assistance system based on multimodal perception according to claim 1 is characterized in that: The implementation method of the cross-modal alignment mechanism includes: Mapping the speech stream, vibration feature vector and lip reading text to a unified time coordinate system; Calculate the time offset of multimodal data and compensate for alignment errors through dynamic time warping algorithm; The gated recurrent unit is used to perform sequence fusion on the aligned multimodal data to generate a synchronized enhanced speech stream.
6. The rail transit hearing-impaired voice assistance system based on multimodal perception according to claim 2 is characterized in that: The optimization method of the time domain waveform reconstruction technology includes: The filtered spectrum is subjected to phase recovery processing, and the continuity of the time domain waveform is optimized using an iterative projection algorithm. The auditory quality of the reconstructed waveform is evaluated using a generative adversarial network, and the spectrum mask parameters are iteratively adjusted. The loudness of the reconstructed speech is compensated in combination with the human ear auditory characteristic curve.
7. The rail transit hearing-impaired voice assistance system based on multimodal perception according to claim 3 is characterized in that: The method for constructing the vibration-speech mapping relationship model includes: Collect multiple sets of synchronized vibration signals and pure speech samples to construct a training dataset; The nonlinear mapping relationship between vibration features and speech spectrum is learned through a bidirectional long short-term memory network, and the network parameters are optimized using a composite loss function of mean square error and spectrum correlation.
8. The rail transit hearing-impaired voice assistance system based on multimodal perception according to claim 4 is characterized in that: The implementation steps of the attention mechanism include: Calculate the semantic relevance of each lip frame with the previous and next frames to generate a temporal attention score; Extract multi-scale features of the mouth shape area through spatial pyramid pooling and generate spatial attention scores; Multiply the temporal and spatial attention scores to generate a comprehensive attention weight matrix; The key frame features are weightedly fused according to the weight matrix and the enhanced lip semantic coding is output.
9. The rail transit hearing-impaired voice assistance system based on multimodal perception according to claim 5 is characterized in that: The sequence fusion method of the gated recurrent unit includes: Initialize the hidden state to the mean vector of multimodal features, calculate the activation values of the input gate, forget gate, and output gate at each time step, selectively update the hidden state through the gating mechanism, and output the fused feature vector of the current step; The fused feature vectors of all time steps are concatenated into the final enhanced speech auxiliary information stream.
Citation Information
Patent Citations
Method and device for speech recognition
CN106157956A
Voice detection method based on millimeter wave radar
CN119780912A