Voice Sensing System for Rail Transit Hearing Impairment Assistance

Through CEEMDAN and deep learning technology, audio signals are decomposed and feature extraction, noise is filtered out, and key information is captured, which solves the problem of difficulty in obtaining sound information for hearing impaired people in rail transit environments, and realizes clear voice output.

CN119316785BActive Publication Date: 2025-07-04SHANGHAI BOZHIWEI ELECTRONIC SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411428208.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-14
Publication Date
2025-07-04
Estimated Expiration
2044-10-14

AI Technical Summary

Technical Problem

In rail transit environments, hearing impaired people find it difficult to effectively obtain sound information, and traditional hearing assistive devices are severely disturbed by background noise, making it difficult to achieve ideal results.

Method used

The audio signal is decomposed by CEEMDAN technology, and the frequency domain characteristics of IMF components are extracted by deep learning signal frequency domain analysis technology. Through semantic contribution characteristic filtering and full-frequency domain feature aggregation, noise components are filtered out, key semantic information is captured, and the output is amplified after signal enhancement processing is performed.

Benefits of technology

It improves the hearing experience of hearing impaired people in the rail transit environment, enhances the ability to obtain useful voice information, and provides clearer voice information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119316785B_ABST
    Figure CN119316785B_ABST
Patent Text Reader

Abstract

The present application discloses a voice sensing system for rail transit hearing impairment assistance, which includes an audio acquisition module, an audio signal processing module, a power amplifier module, and an output module. Specifically, the audio acquisition module is used to receive input audio; the audio signal processing module is used to perform signal enhancement processing on the input audio to obtain an enhanced audio signal; the power amplifier module is used to perform amplification processing on the enhanced audio signal to obtain an amplified audio signal; and the output module is used to output the amplified audio signal. In this way, clearer voice information is provided for hearing-impaired persons through audio signal processing, so as to effectively improve the auditory experience of hearing-impaired persons in the rail transit environment and enhance their ability to obtain useful sound information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent assistance, and more specifically, to a voice sensing system for assisting the hearing-impaired in rail transit. Background Art

[0002] With the accelerated development of urbanization, rail transit has become one of the indispensable public transportation tools in modern cities, greatly facilitating people's daily travel. However, for the hearing-impaired group, it is often difficult to effectively obtain sound information in the rail transit environment (such as station announcements, emergency broadcasts, etc.), which not only affects their travel experience, but also may bring safety hazards. Therefore, how to provide effective auxiliary hearing services for the hearing-impaired has become one of the current problems that need to be solved urgently.

[0003] Traditional hearing aids mainly rely on physical amplification to improve the audibility of sound. However, in rail transit environments, background noise interference is particularly serious, such as vehicle operation noise, passenger conversation noise, etc., making it difficult for traditional hearing aids to achieve the ideal hearing aid effect.

[0004] Therefore, an optimized voice sensing system for assisting the hearing-impaired in rail transit is desired. Summary of the invention

[0005] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides a voice sensing system for assisting the hearing-impaired in rail transit, which decomposes the input audio by adopting CEEMDAN technology, and uses the signal frequency domain analysis technology based on deep learning to perform frequency domain analysis on each IMF component after decomposition, so as to extract the frequency domain feature representation of each IMF component, and then filter out the noise components in the audio signal by performing feature filtering based on semantic contribution and full-frequency domain feature aggregation on each IMF component, and capture key semantic information to achieve enhanced processing of the audio signal. Finally, the enhanced audio signal is amplified and output, thereby providing clearer voice information for the hearing-impaired. In this way, the auditory experience of the hearing-impaired in the rail transit environment can be effectively improved, and their ability to obtain useful sound information can be enhanced.

[0006] According to one aspect of the present application, there is provided a voice sensing system for assisting hearing impairment in rail transit, which includes: an audio acquisition module, an audio signal processing module, a power amplifier module and an output module;

[0007] The audio acquisition module is used to receive input audio;

[0008] The audio signal processing module is used to perform signal enhancement processing on the input audio to obtain an enhanced audio signal;

[0009] The power amplifier module is used to amplify the enhanced audio signal to obtain an amplified audio signal;

[0010] The output module is used to output the amplified audio signal;

[0011] Wherein, the audio signal processing module includes: a signal decomposition unit, configured to perform CEEMDAN decomposition on the input audio to obtain a set of audio signal IMF components; a frequency-domain feature extraction unit, configured to extract the frequency-domain features of each audio signal IMF component in the set of audio signal IMF components respectively to obtain a set of audio signal IMF component frequency-domain feature vectors; a feature filtering unit, configured to perform feature filtering based on semantic contribution degree on the set of audio signal IMF component frequency-domain feature vectors to obtain a set of filtered audio signal IMF component frequency-domain feature vectors; and a signal restoration unit, configured to perform signal restoration based on the set of filtered audio signal IMF component frequency-domain feature vectors to obtain the enhanced audio signal.

[0012] Compared with the prior art, a voice sensing system for rail transit hearing impairment assistance provided by the present application decomposes the input audio by adopting the CEEMDAN technology, and performs frequency-domain analysis on each decomposed IMF component by using the signal frequency-domain analysis technology based on deep learning to extract the frequency-domain feature representations of each IMF component. Furthermore, by performing feature filtering based on semantic contribution degree and full-frequency domain feature aggregation on each IMF component, the noise components in the audio signal are filtered out, and the key semantic information is captured to implement the enhancement processing of the audio signal. Finally, the enhanced audio signal is amplified and output, so as to provide clearer voice information for the hearing-impaired. In this way, the auditory experience of the hearing-impaired in the rail transit environment can be effectively improved, and their ability to obtain useful sound information can be enhanced. Description of the Drawings

[0013] By describing the embodiments of the present application in more detail in conjunction with the drawings, the above and other objects, features and advantages of the present application will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0014] Figure 1 It is a block diagram of a voice sensing system for rail transit hearing impairment assistance according to an embodiment of the present application;

[0015] Figure 2 It is a schematic diagram of data flow of a voice sensing system for rail transit hearing impairment assistance according to an embodiment of the present application;

[0016] Figure 3 is a flowchart of the working principle of the MCU according to an embodiment of the present application;

[0017] Figure 4 is a block diagram of the audio signal processing module in the voice sensing system for rail transit hearing impairment assistance according to an embodiment of the present application. Detailed implementation manners

[0018] Next, example embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments of the present application. It should be understood that the present application is not limited by the example embodiments described herein.

[0019] As shown in the present application and the claims, unless the context clearly indicates otherwise, words such as "a", "an", "one" and / or "the" are not specifically singular and may also include plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.

[0020] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, however, any number of different modules can be used and run on the user terminal and / or the server. The modules are only illustrative, and different aspects of the system and method can use different modules.

[0021] Flowcharts are used in the present application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the operations before or below do not necessarily need to be executed precisely in sequence. On the contrary, various steps can be processed in reverse order or simultaneously as needed. At the same time, other operations can also be added to these processes, or one or several steps can be removed from these processes.

[0022] Next, example embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments of the present application. It should be understood that the present application is not limited by the example embodiments described herein.

[0023] Traditional hearing aids mainly rely on physical amplification to improve the audibility of sounds. However, in the rail transit environment, the interference of background noise is particularly serious, such as vehicle running noise, passenger conversation noise, etc., making it difficult for traditional hearing aids to achieve an ideal hearing aid effect. Therefore, an optimized voice sensing system for rail transit hearing impairment assistance is expected.

[0024] In the technical solution of the present application, a voice sensing system for rail transit hearing impairment assistance is proposed. Figure 1 It is a block diagram of a voice sensing system for rail transit hearing impairment assistance according to an embodiment of the present application. Figure 2 It is a schematic diagram of data flow of a voice sensing system for rail transit hearing impairment assistance according to an embodiment of the present application. As Figure 1 and Figure 2 shown, a voice sensing system 300 for rail transit hearing impairment assistance according to an embodiment of the present application includes: an audio acquisition module 310 for receiving input audio; an audio signal processing module 320 for performing signal enhancement processing on the input audio to obtain an enhanced audio signal; a power amplifier module 330 for performing amplification processing on the enhanced audio signal to obtain an amplified audio signal; and an output module 340 for outputting the amplified audio signal.

[0025] Those of ordinary skill in the art should know that according to the principle of electromagnetism, when there is a current passing through a long straight wire, magnetic lines of force will be generated around it. According to the right-hand rule, the magnetic lines of force are concentric circles in the same plane and perpendicular to the wire. The magnetic lines of force are centered on the wire, from the inside to the outside, from dense to sparse, and the magnetic field changes from strong to weak accordingly. When the current in a closed loop changes, the magnetic field generated by the current also changes accordingly. After the surrounding coil generates an induced magnetic field in the carriage space, the passengers wearing hearing aids in the carriage can switch to the "T file" of the hearing aid ("T" represents the magnetoelectric induction input mode), and then use the induction coil inside the hearing aid to cut the magnetic lines of force in the magnetic field to generate a current, thereby picking up electromagnetic signals with specific frequencies around and converting them into electrical signals - the basic principle relied on in this process is the law of electromagnetic induction, which is expressed as: the magnitude of the induced electromotive force in any closed circuit is equal to the rate of change of the magnetic flux through this circuit. The process of the electric field changing to the magnetic field and the magnetic field changing to the electric field is the basic principle of the induction coil system. Therefore, the process of signal transmission and reception of the audio induction loop system is that the host filters and amplifies the received audio signal and outputs it into the long straight wire, and then an alternating magnetic field that changes with the audio can be formed in the carriage; at this time, the coil of the hearing aid in the magnetic field induces an audio current, which is appropriately amplified, and the current signal is converted into a voice signal, so that the hearing-impaired person can receive clearer voice content.

[0026] Specifically, the audio acquisition module 310 is used to receive input audio. Correspondingly, after the audio acquisition module 310 receives the input audio, the input audio is transmitted to the audio signal processing module 320 for signal processing.

[0027] Correspondingly, in a feasible implementation manner, after receiving the input audio, the audio acquisition module 310 performs A / D conversion on it, converting the analog quantity into a digital quantity, and then inputs it to the audio signal processing module 320 for further signal enhancement processing. It is worth mentioning that the audio acquisition module is mainly responsible for receiving at most two channels of external audio inputs, filtering the noise in the external audio through an isolation transformer, and then outputting it to the digital voice processing module.

[0028] In an example, after the audio acquisition module 320 receives the input audio, it gives it to the MCU. The MCU performs A / D conversion, converting the analog quantity into a digital quantity, and performs FFT transformation for spectrum analysis, processes and digitally filters the frequency bands with problems in the frequency response of the audio signal. After the signal processed by the MCU is output to the hardware filter for signal filtering and smoothing, it is output to the external loop circuit after power amplification (that is, after the power amplification processing of the power amplification module 330).

[0029] As Figure 3 shown, the MCU is mainly composed of two parts: digital voice processing and a controller; digital voice processing is mainly responsible for performing analog quantity analysis, A / D conversion, digital quantity analysis, FFT transformation, spectrum analysis, gain amplification, anti-aliasing filtering, etc. on the audio transmitted by the voice acquisition module. The voice processing uses STM32F407, with a maximum main frequency of up to 168 MHZ, providing 3 channels of ADC, with a maximum configuration of 12-bit resolution, and using the DMA mode for continuous high-speed AD conversion to convert the analog quantity into a data quantity, and then using the floating-point operation unit FPU to perform FFT operations to achieve the conversion from the time domain to the frequency domain; the controller mainly feeds back the control signals collected by DI and the gain control for device debugging to the MCU together; the MCU outputs the signals of the digital voice processing and the controller to the power amplification module. Among them, DI acquisition is mainly responsible for collecting the external switch quantity inputs corresponding to two channels of audio signals, and at the same time providing the collected signals to the controller in the MCU; in addition, in the case of only one channel of audio input and output, the corresponding control switch can be internally short-circuited.

[0030] Particularly, due to audio signals, especially those from the actual environment, often containing complex non-linear components and transient changes, this characteristic makes it difficult for traditional signal processing methods to effectively extract their useful information.

[0031] Therefore, in the preferred embodiment of the present application, the audio signal processing module 320 adopts another method to perform signal enhancement processing to obtain an enhanced audio signal. Specifically, as Figure 4As shown, the audio signal processing module 320 includes: a signal decomposition unit 321 for performing CEEMDAN decomposition on the input audio to obtain a set of audio signal IMF components; a frequency domain feature extraction unit 322 for respectively extracting the frequency domain features of each audio signal IMF component in the set of audio signal IMF components to obtain a set of audio signal IMF component frequency domain feature vectors; a feature filtering unit 323 for performing feature filtering based on semantic contribution degree on the set of audio signal IMF component frequency domain feature vectors to obtain a set of filtered audio signal IMF component frequency domain feature vectors; and a signal restoration unit 324 for performing signal restoration based on the set of filtered audio signal IMF component frequency domain feature vectors to obtain the enhanced audio signal.

[0032] Specifically, the signal decomposition unit 321 is used to perform CEEMDAN decomposition on the input audio to obtain a set of audio signal IMF components. It should be understood that since audio signals, especially those from the actual environment, often contain complex non-linear components and transient changes, this characteristic makes it difficult for traditional signal processing methods to effectively extract their useful information. CEEMDAN is an adaptive decomposition method for non-linear and non-stationary signals, which can effectively process audio signals containing multiple frequency components and complex noise. During the decomposition process, CEEMDAN can decompose the signal into a series of IMF components with different frequency scales by adding white noise multiple times and applying the EMD (Empirical Mode Decomposition) algorithm, thereby separating noise and useful signals in different IMF components. For example, high-frequency IMF components usually contain more instantaneous changes and noise, while low-frequency IMF components may carry more stable and useful information. In this way, it helps to process each IMF component separately, effectively suppressing noise interference while retaining the important features of the signal.

[0033] Specifically, the frequency domain feature extraction unit 322 is used to respectively extract the frequency domain features of each audio signal IMF component in the set of audio signal IMF components to obtain a set of audio signal IMF component frequency domain feature vectors. In a specific example of the present application, each audio signal IMF component in the set of audio signal IMF components is input into an audio signal component frequency domain feature extractor based on a 1D-CNN model to obtain a set of audio signal IMF component frequency domain feature vectors. That is, in order to further extract key feature information from each audio signal IMF component, the present application adopts a 1D-CNN model as an audio signal component frequency domain feature extractor to perform frequency domain feature extraction on each audio signal IMF component in the set of audio signal IMF components. By stacking multiple convolutional layers, activation layers (such as ReLU) and pooling layers (such as maximum pooling), the 1D-CNN model can make full use of the local correlation of audio signals on the time axis, and gradually extract the deep features of audio signals at different frequency scales through layer-by-layer convolution operations, so as to mine the frequency domain characteristics of each audio signal IMF component, thereby obtaining a set of frequency domain feature vectors of the audio signal IMF components, providing strong data support for subsequent feature filtering.

[0034] Specifically, the feature filtering unit 323 is configured to perform feature filtering based on semantic contribution degree on the set of frequency-domain feature vectors of the IMF components of the audio signal to obtain a set of frequency-domain feature vectors of the IMF components of the filtered audio signal. In order to remove redundant information and noise components in the audio signal and achieve enhanced processing of the audio signal, the present application further performs feature filtering on the set of frequency-domain feature vectors of the IMF components of the audio signal. It is worth mentioning that, in order to construct a more refined feature filtering mechanism, the present application introduces a feature filtering method based on semantic contribution degree, which, by introducing the concept of pseudo-anchoring center and combining with the semantic contribution degree query mechanism, can intelligently evaluate and screen out the most critical feature information for the overall audio signal in each frequency-domain feature vector of the IMF components. In a specific example of the present application, first, perform clustering analysis on the set of frequency-domain feature vectors of the IMF components of the audio signal to obtain a frequency-domain feature clustering center representation vector of the IMF components of the audio signal as the pseudo-anchoring center; capture the overall internal structure and distribution of the data by performing clustering analysis on the set of frequency-domain feature vectors of the IMF components of the audio signal, and extract the frequency-domain feature clustering center representation vector of the IMF components of the audio signal as the pseudo-anchoring center, so as to provide a reference for subsequent feature evaluation. Specifically, calculate the position-wise mean vector of the set of frequency-domain feature vectors of the IMF components of the audio signal as the frequency-domain feature clustering center representation vector of the IMF components of the audio signal. Then, based on the semantic contribution of each frequency-domain feature vector of the IMF components of the audio signal in the set of frequency-domain feature vectors of the IMF components of the audio signal relative to the pseudo-anchoring center, perform feature filtering on the set of frequency-domain feature vectors of the IMF components of the audio signal to obtain the set of frequency-domain feature vectors of the IMF components of the filtered audio signal. Specifically, calculate the semantic contribution degree representation matrix of each frequency-domain feature vector of the IMF components of the audio signal in the set of frequency-domain feature vectors of the IMF components of the audio signal relative to the pseudo-anchoring center to obtain a set of semantic contribution degree representation matrices of the frequency-domain features of the IMF components of the audio signal; and based on the set of semantic contribution degree representation matrices of the frequency-domain features of the IMF components of the audio signal, perform feature modulation on the set of frequency-domain feature vectors of the IMF components of the audio signal to obtain the set of frequency-domain feature vectors of the IMF components of the filtered audio signal. Here, by calculating the semantic contribution of each frequency-domain feature vector of the IMF components of the audio signal relative to the pseudo-anchoring center, the correlation and relative importance between each frequency-domain feature vector of the IMF components of the audio signal and the central tendency of the audio signal are quantitatively represented, and a sequence of semantic contribution degree representation matrices is generated. Furthermore, by performing dropout masking on each semantic contribution degree representation matrix, the attention to key features is strengthened, the generalization ability of the model is enhanced, and the number of parameters is reduced.Finally, using each semantic contribution degree representation matrix after masking as the modulation matrix, filter out unimportant or noise features by weighting the frequency domain feature vectors of each original audio signal IMF component, so as to obtain a set of frequency domain feature vectors of the filtered audio signal IMF components.

[0035] Among them, the process of calculating the semantic contribution degree representation matrix of each audio signal IMF component frequency domain feature vector in the set of audio signal IMF component frequency domain feature vectors relative to the pseudo-anchoring center to obtain a set of audio signal IMF component frequency domain feature semantic contribution degree representation matrices includes: multiplying the audio signal IMF component frequency domain feature vector by the transposed vector of the pseudo-anchoring center and then dividing by the two-norm of the pseudo-anchoring center to obtain the audio signal IMF component frequency domain feature semantic contribution degree representation matrix.

[0036] More specifically, based on the set of audio signal IMF component frequency domain feature semantic contribution degree representation matrices, the process of performing feature modulation on the set of audio signal IMF component frequency domain feature vectors to obtain the set of filtered audio signal IMF component frequency domain feature vectors includes: performing contribution degree mask sparsification processing based on the dropout mechanism on each audio signal IMF component frequency domain feature semantic contribution degree representation matrix in the set of audio signal IMF component frequency domain feature semantic contribution degree representation matrices to obtain a set of masked sparse audio signal IMF component frequency domain feature semantic contribution degree representation matrices; using each masked sparse audio signal IMF component frequency domain feature semantic contribution degree representation matrix in the set of masked sparse audio signal IMF component frequency domain feature semantic contribution degree representation matrices as the modulation matrix, and respectively calculating the matrix product between it and each audio signal IMF component frequency domain feature vector to obtain the set of filtered audio signal IMF component frequency domain feature vectors.

[0037] In summary, in the above embodiments, performing semantic contribution degree-based feature filtering on the set of audio signal IMF component frequency domain feature vectors to obtain a set of filtered audio signal IMF component frequency domain feature vectors includes: processing the set of audio signal IMF component frequency domain feature vectors with the following feature filtering formula to obtain the set of filtered audio signal IMF component frequency domain feature vectors, where the feature filtering formula is:

[0038] X = {x1, x2,..., x m}

[0039]

[0040]

[0041] {x i’} = {WF i *x i}

[0042] Wherein, X represents the set of frequency-domain feature vectors of the IMF components of the audio signal, x1, x2, x i and x m are respectively the first, second, i-th, and m-th frequency-domain feature vectors of the IMF components of the audio signal in the set of frequency-domain feature vectors of the IMF components of the audio signal. The value of m is the number of frequency-domain feature vectors of the IMF components of the audio signal. x c represents the pseudo-anchoring center, ||·||2 represents the second norm of the feature vector, (·) T represents the transpose of the vector, softmax is the normalized exponential function, dropout[·] represents the random inactivation process, WF i represents the semantic contribution degree representation matrix of the i-th masked sparse audio signal IMF component frequency-domain feature, * represents matrix multiplication, x i ’ represents the i-th filtered frequency-domain feature vector of the IMF component of the audio signal.

[0043] It is worth mentioning that in other signal processing solutions of this application, the input audio can also be enhanced by other means to obtain an enhanced audio signal. For example: input the input audio; use an audio processing library (such as Librosa, PyDub) to read the input audio file; convert the audio file into a digital signal form; extract the basic parameters of the audio, such as the sampling rate, number of channels, etc.; use a noise estimation algorithm (such as minimum statistical noise estimation) to estimate the characteristics of the background noise; use the short-time Fourier transform (STFT) to convert the audio signal from the time domain to the frequency domain; divide the audio signal into multiple frames and perform FFT transformation on each frame; suppress the noise by subtracting the estimated noise spectrum from the spectrum; use a Wiener filter to weight the spectrum and reduce the noise impact; use a deep learning model (such as CNN, RNN, Transformer) for noise suppression; make the audio more uniform and clear by compressing the dynamic range of the audio signal; use a compressor to adjust the gain of the audio signal and reduce the dynamic range; use spectral enhancement techniques (such as MMSE-STSA) to further improve the clarity of the audio signal; enhance the speech signal by adjusting certain frequency bands in the spectrum; use an equalizer to adjust the frequency response of the audio signal to make the audio more balanced and natural; use the inverse short-time Fourier transform (ISTFT) to convert the processed frequency-domain signal back to the time domain; convert the frequency-domain signal of each frame back to the time-domain signal and perform overlap-and-add (OLA) to reconstruct the original audio signal to obtain the enhanced audio signal. This is not limited to this application.

[0044] Further, the signal restoration unit 324 is configured to perform signal restoration based on the set of frequency-domain feature vectors of the IMF components of the filtered audio signal to obtain the enhanced audio signal. In a specific example of the present application, the set of frequency-domain feature vectors of the IMF components of the filtered audio signal is first concatenated to obtain a full-frequency-domain concatenated representation vector of the audio signal; the set of frequency-domain feature vectors of the IMF components of the filtered audio signal is concatenated to integrate the key features of the audio signal in different frequency bands, forming a full-frequency-domain feature representation of the audio signal, that is, a full-frequency-domain concatenated representation vector of the audio signal, so as to comprehensively reflect the overall characteristics of the audio signal. Subsequently, the full-frequency-domain concatenated representation vector of the audio signal is input into an audio signal restorer based on a decoder to obtain the enhanced audio signal. That is, in the technical solution of the present application, in order to restore the full-frequency-domain concatenated representation vector of the audio signal to the original audio signal form, the present application uses an audio signal restorer based on a decoder to perform signal reconstruction on the full-frequency-domain concatenated representation vector of the audio signal to obtain an enhanced audio signal. Through multi-layer deconvolution operations, the decoder can gradually map the deep features in the full-frequency-domain concatenated representation vector of the audio signal back to the audio signal in the original time domain, thereby realizing the enhanced reconstruction of the audio signal.

[0045] Here, in the preferred example, each frequency-domain feature vector of the IMF components of the audio signal in the set of frequency-domain feature vectors of the IMF components of the audio signal respectively represents the frequency-domain feature of each IMF component of the input audio. Thus, when the set of frequency-domain feature vectors of the IMF components of the audio signal is input into the feature filtering network based on pseudo-anchored center contribution degree semantic query, the feature filtering network based on pseudo-anchored center contribution degree semantic query performs feature filtering based on the pseudo-anchored center contribution degree of the frequency-domain feature vector of the IMF components of the audio signal. However, considering that the pseudo-anchored center is not a real feature filtering anchor point, this will cause a contribution degree mismatch when calculating the pseudo-anchored center contribution degree, which will lead to a full-frequency-domain aggregation difference in the full-frequency-domain concatenated representation vector of the audio signal obtained by concatenation due to the node contribution degree mismatch. Therefore, it is desired to improve the detailed semantic aggregation expression effect of the full-frequency-domain local frequency-domain feature aggregation difference of the full-frequency-domain concatenated representation vector of the audio signal.

[0046] Preferably, inputting the full-frequency-domain concatenated representation vector of the audio signal into an audio signal restorer based on a decoder to obtain an enhanced audio signal includes:

[0047] Calculating the sum of the absolute values of the respective eigenvalues of the full-frequency-domain concatenated representation vector of the audio signal to obtain a first full-frequency-domain concatenated sum and modulation value of the audio signal, and calculating the square root of the sum of the squares of the respective eigenvalues of the full-frequency-domain concatenated representation vector of the audio signal to obtain a second full-frequency-domain concatenated sum and modulation value of the audio signal;

[0048] After subtracting the full-frequency domain concatenated representation vector of the audio signal from the full-frequency domain concatenated and modulation value of the second audio signal, perform dot multiplications with the number of eigenvalues of the full-frequency domain concatenated representation vector of the audio signal and the reciprocal of the full-frequency domain concatenated and modulation value of the first audio signal respectively, and take the reciprocal of each eigenvalue to obtain the full-frequency domain concatenated phase conversion vector of the first audio signal;

[0049] After subtracting the full-frequency domain concatenated representation vector of the audio signal from the full-frequency domain concatenated and modulation value of the first audio signal, perform dot multiplications with the square root of the number of eigenvalues of the full-frequency domain concatenated representation vector of the audio signal and the reciprocal of the full-frequency domain concatenated and modulation value of the second audio signal respectively, and take the reciprocal of each eigenvalue to obtain the full-frequency domain concatenated phase conversion vector of the second audio signal; and

[0050] Calculate the dot product vector of the full-frequency domain concatenated phase conversion vector of the second audio signal and the weighted hyperparameter, and subtract the dot product vector from the full-frequency domain concatenated phase conversion vector of the first audio signal to obtain the optimized full-frequency domain concatenated representation vector of the audio signal; and

[0051] Input the optimized full-frequency domain concatenated representation vector of the audio signal into a decoder-based audio signal restorer to obtain an enhanced audio signal.

[0052] Here, the optimized representation of the full-frequency domain concatenated representation vector of the audio signal, denoted as V, is:

[0053]

[0054]

[0055] v i ∈V∈R n

[0056] where V is the full-frequency domain concatenated representation vector of the audio signal, v i are the respective eigenvalues of the full-frequency domain concatenated representation vector of the audio signal, α is the full-frequency domain concatenated and modulation value of the first audio signal, and β is the full-frequency domain concatenated and modulation value of the second audio signal, n is the number of eigenvalues of the full-frequency domain concatenated representation vector of the audio signal, V1 is the full-frequency domain concatenated phase conversion vector of the first audio signal, V2 is the full-frequency domain concatenated phase conversion vector of the second audio signal, ω is the weighted hyperparameter, θ is subtraction, ⊙ is dot multiplication, and V′ is the optimized full-frequency domain concatenated representation vector of the audio signal.

[0057] Based on this, the difference between the eigenvalues of the full-frequency domain concatenated representation vector of the audio signal and the differential of the different and modulation representations of the feature set of the vector as a whole of the full-frequency domain concatenated representation vector of the audio signal is used as the semantic change intensity information, and a class-phase conversion corresponding to the position-based intensity modulation is performed through different and modulation representation forms, so as to perform a space translation operation based on alternating stacking under the scale balance of the vector set of the full-frequency domain concatenated representation vector of the audio signal, so that the aggregation enhancement of semantic change phase perception can improve the axial aggregation receptive field along the feature aggregation direction, thereby improving the perception effect of the aggregation semantics of the full-frequency domain concatenated representation vector of the audio signal on the detailed semantic changes, so as to improve the expression effect of the full-frequency domain concatenated representation vector of the audio signal, and improve the signal quality and accuracy of the enhanced audio signal obtained by inputting it into the audio signal restorer based on the decoder.

[0058] Specifically, the power amplifier module 330 and the output module 340 are used to amplify the enhanced audio signal to obtain an amplified audio signal; and output the amplified audio signal. That is, after obtaining the enhanced audio signal, the power amplifier module will further perform power amplification processing on it to enhance the signal strength and ensure clear transmission to the hearing-impaired in the rail transit environment. It is worth mentioning that the power amplifier module is responsible for amplifying the voice signal processed by the MCU, converting the voice signal into current, and outputting it to the loop.

[0059] Specifically, the output module 340 is used to output the amplified audio signal. That is, the amplified audio signal is then output to the hearing-impaired through an output module, such as headphones, so that the hearing-impaired can clearly hear and understand important sound information even in a noisy rail transit environment. In one example, the feedback circuit is responsible for collecting the output of the power amplifier and feeding it back to the MCU. The MCU compares the collected output of the power amplifier with the original output and makes corresponding feedback and records on the output.

[0060] It is worth mentioning that in the technical solution of this application, the voice sensing system 300 for rail transit hearing impairment assistance further includes other necessary functional modules, such as a power supply module, etc. For this, since it is not the focus of this application, it will not be elaborated here.

[0061] As described above, the voice sensing system 300 for rail transit hearing impairment assistance according to the embodiments of the present application can be implemented in various wireless terminals, such as a server having a voice sensing algorithm for rail transit hearing impairment assistance, etc. In a possible implementation manner, the voice sensing system 300 for rail transit hearing impairment assistance according to the embodiments of the present application can be integrated into the wireless terminal as a software module and / or a hardware module. For example, the voice sensing system 300 for rail transit hearing impairment assistance can be a software module in the operating system of the wireless terminal, or can be an application developed for the wireless terminal; of course, the voice sensing system 300 for rail transit hearing impairment assistance can also be one of the numerous hardware modules of the wireless terminal.

[0062] Alternatively, in another example, the voice sensing system 300 for rail transit hearing impairment assistance and the wireless terminal can also be separate devices, and the voice sensing system 300 for rail transit hearing impairment assistance can be connected to the wireless terminal through a wired and / or wireless network, and transmit interaction information according to a predefined data format.

[0063] The embodiments of the present disclosure have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technologies in the market, or to enable other ordinary skill in the art in the technical field to understand the embodiments disclosed herein.

Claims

1. A voice sensing system for rail transit hearing impairment assistance, characterized in that, Including: An audio acquisition module, an audio signal processing module, a power amplifier module, and an output module; The audio acquisition module is used to receive input audio; The audio signal processing module is used to perform signal enhancement processing on the input audio to obtain an enhanced audio signal; The power amplifier module is used to perform amplification processing on the enhanced audio signal to obtain an amplified audio signal; The output module is used to output the amplified audio signal; Among them, the audio signal processing module includes: a signal decomposition unit, which is used to perform CEEMDAN decomposition on the input audio to obtain a set of audio signal IMF components; a frequency domain feature extraction unit, which is used to extract the frequency domain features of each audio signal IMF component in the set of audio signal IMF components to obtain a set of audio signal IMF component frequency domain feature vectors; a feature filtering unit, which is used to perform feature filtering based on semantic contribution degree on the set of audio signal IMF component frequency domain feature vectors to obtain a filtered set of audio signal IMF component frequency domain feature vectors; a signal restoration unit, which is used to perform signal restoration based on the filtered set of audio signal IMF component frequency domain feature vectors to obtain the enhanced audio signal; Among them, the feature filtering unit includes: A clustering analysis sub-unit, which is used to perform clustering analysis on the set of audio signal IMF component frequency domain feature vectors to obtain an audio signal IMF component frequency domain feature clustering center representation vector as a pseudo-anchoring center; A centralized feature filtering sub-unit, which is used to perform feature filtering on the set of audio signal IMF component frequency domain feature vectors based on the semantic contribution of each audio signal IMF component frequency domain feature vector in the set of audio signal IMF component frequency domain feature vectors relative to the pseudo-anchoring center to obtain the filtered set of audio signal IMF component frequency domain feature vectors; The signal restoration unit includes: A full frequency domain feature aggregation sub-unit, which is used to cascade the filtered set of audio signal IMF component frequency domain feature vectors to obtain an audio signal full frequency domain cascade representation vector; A decoding sub-unit, which is used to input the audio signal full frequency domain cascade representation vector into an audio signal restorer based on a decoder to obtain the enhanced audio signal.

2. The voice sensing system for rail transit hearing impairment assistance according to claim 1, characterized in that, The frequency domain feature extraction unit is used to: Input each audio signal IMF component in the set of audio signal IMF components into an audio signal component frequency domain feature extractor based on a 1D-CNN model to obtain a set of audio signal IMF component frequency domain feature vectors.

3. The voice sensing system for rail transit hearing impairment assistance according to claim 2, wherein The audio signal component frequency domain feature extractor based on the 1D-CNN model includes an input layer, a one-dimensional convolutional layer, an activation layer based on the ReLU activation function, a pooling layer, and an output layer.

4. The voice sensing system for rail transit hearing impairment assistance according to claim 3, characterized in that, The clustering analysis sub-unit is used to: Calculate the position-wise mean vector of the set of audio signal IMF component frequency domain feature vectors as the audio signal IMF component frequency domain feature clustering center representation vector.

5. The voice sensing system for rail transit hearing impairment assistance according to claim 4, wherein The centralized feature filtering sub-unit includes: A semantic contribution degree query secondary subunit, configured to calculate the semantic contribution degree representation matrix of each audio signal IMF component frequency domain feature vector in the set of audio signal IMF component frequency domain feature vectors with respect to the pseudo-anchoring center to obtain a set of audio signal IMF component frequency domain feature semantic contribution degree representation matrices; A feature modulation secondary subunit, configured to perform feature modulation on the set of audio signal IMF component frequency domain feature vectors based on the set of audio signal IMF component frequency domain feature semantic contribution degree representation matrices to obtain the set of filtered audio signal IMF component frequency domain feature vectors.

6. The voice sensing system for rail transit hearing impairment assistance according to claim 5, wherein The semantic contribution degree query secondary subunit is configured to: Multiply the audio signal IMF component frequency domain feature vector by the transposed vector of the pseudo-anchoring center and then divide by the two-norm of the pseudo-anchoring center to obtain the audio signal IMF component frequency domain feature semantic contribution degree representation matrix.

7. The voice sensing system for rail transit hearing impairment assistance according to claim 6, characterized in that, The feature modulation secondary subunit is configured to: Perform contribution degree mask sparsification processing based on a random inactivation mechanism on each audio signal IMF component frequency domain feature semantic contribution degree representation matrix in the set of audio signal IMF component frequency domain feature semantic contribution degree representation matrices to obtain a set of masked sparse audio signal IMF component frequency domain feature semantic contribution degree representation matrices; Using each masked sparse audio signal IMF component frequency domain feature semantic contribution degree representation matrix in the set of masked sparse audio signal IMF component frequency domain feature semantic contribution degree representation matrices as a modulation matrix, respectively calculate the matrix product between it and each audio signal IMF component frequency domain feature vector to obtain the set of filtered audio signal IMF component frequency domain feature vectors.

Citation Information

Patent Citations

  • Digital hearing aid noise reduction method and system and special DSP

    CN110602621A

  • Method and device for speech enhancement based on improved CEEMDAN algorithm

    CN115798494A