A voice acquisition and extraction system and method based on metasurface

By using a metasurface-based voice acquisition system, electromagnetic wave reflection and phase unwinding techniques, combined with voice recognition algorithms, deaf and mute individuals can directly "see" voice signals, solving their difficulties in obtaining information and improving their voice information processing capabilities.

CN116312530BActive Publication Date: 2025-10-28FOURTH MILITARY MEDICAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310058519.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-17
Publication Date
2025-10-28
Estimated Expiration
2043-01-17

Smart Images

  • Figure CN116312530B_ABST
    Figure CN116312530B_ABST
Patent Text Reader

Abstract

The present invention discloses a voice collection and extraction system and method based on a metasurface, including a phase gradient metasurface, an eye tracker, a wearable display, and a vector network analyzer; the phase gradient metasurface is connected to the output end of a controller, and the wearable display is connected to the eye tracker; the phase gradient metasurface is used to emit electromagnetic waves of a set band; the vector network analyzer is used to receive reflected electromagnetic waves from a target object, and the reflected electromagnetic waves carry surface vibration information of the target object; the phase gradient metasurface includes multiple units, the length and width of each unit are more than four times the wavelength corresponding to the working center frequency, and each unit is provided with an electrically adjustable device, which is a diode or a varactor diode; the human ability to process voice information is improved; the structure is simple and easy to implement, and the efficiency is significantly improved; a voice perception framework is provided for the deaf and dumb, enabling the deaf and dumb to "see" voice signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of electronic technology, specifically relating to a voice acquisition and extraction system and method based on metasurfaces. Background Technology

[0002] Speech is a vital tool for information sharing between people. Deaf and mute individuals face significant challenges in accessing information due to inherent difficulties in processing spoken language, leading to reduced social interaction and contributing to increased morbidity and mortality. Therefore, improving speech perception abilities in deaf and mute individuals is of paramount importance.

[0003] Metasurfaces are man-made planar two-dimensional structures, typically composed of periodic or aperiodic arrangements of metallic unit cells. Their unique electromagnetic properties are not found in naturally occurring materials. For example, they can alter the polarization direction, phase, and propagation mode of transmitted or reflected electromagnetic waves. Different types of metasurfaces exhibit different modulation effects on electromagnetic waves. Leveraging these superior properties, vision-driven metasurface platforms can provide healthcare services to the deaf and mute, overcoming communication barriers and breaking down the barriers between visual and auditory information, thus helping them access speech information without obstacles. Summary of the Invention

[0004] This invention provides a solution for realizing a metasurface-based voice acquisition system, which can provide healthcare services for the deaf and mute and overcome communication barriers.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: a voice acquisition and extraction system based on a metasurface, comprising a phase gradient metasurface, an eye tracker, a wearable display, and a vector network analyzer; the phase gradient metasurface is connected to the output terminal of a controller, and the wearable display is connected to the eye tracker; the phase gradient metasurface is used to emit electromagnetic waves of a set band; the vector network analyzer is used to receive reflected electromagnetic waves from a target object, wherein the reflected electromagnetic waves carry surface vibration information of the target object; the phase gradient metasurface comprises multiple units, each unit having a length and width more than four times the wavelength corresponding to the working center frequency, and each unit is equipped with an electrically tunable device, wherein the electrically tunable device is a diode or a varactor diode.

[0006] The controller uses an FPGA module or a microcontroller.

[0007] The speech acquisition and extraction method based on metasurface-based speech acquisition and extraction system includes the following steps:

[0008] The eye tracker selects the sound source to be collected based on the gaze state of the human eye;

[0009] The electromagnetic wave reflection mode of the phase gradient metasurface is changed so that electromagnetic waves illuminate the selected sound source.

[0010] The vector network analyzer receives vibration signals reflected from the sound source. First, it uses the phase unwinding method to extract noisy speech signals from the electromagnetic wave echo. Second, it uses the short-time amplitude spectrum minimum mean square error method to enhance the speech signal. Finally, it uses a speech recognition module to obtain the final speech-to-text recognition result.

[0011] Eye trackers record eye movements, including fixation points, fixation duration, and saccade information. The eye trackers track and determine the user's observation area and points of interest in real time.

[0012] By changing the feed voltage of the electrically tunable device on the phase gradient metasurface, the phase gradient and operating mode of the control electromagnetic metasurface are controlled, thereby adjusting the electromagnetic wave radiation region.

[0013] Once the irradiation direction is determined, the phase gradient of the metasurface is determined according to the generalized Snell's law.

[0014] The in-phase baseband signal and the quadrature baseband signal are obtained from the time-domain signal of the target acoustic reflection irradiated by the metasurface and the vibration signal received by the vector network analyzer. The time-domain signal of the target acoustic reflection irradiated by the metasurface is s. T (t), the received signal collected by the receiving antenna is s R (t), represented as follows:

[0015] s T (t)=A t cos(ωt) (1)

[0016]

[0017] In the formula, A t and A r Let ω be the amplitude of the transmitted signal and λ be the wavelength of the carrier frequency, c be the speed of light, and d0 be the distance between the metasurface and the acoustic diaphragm. x(t) represents the instantaneous displacement of the acoustic diaphragm when the acoustic diaphragm emits sound, as follows:

[0018]

[0019] Among them, A s and These represent the amplitude and phase shift of the diaphragm displacement caused by the sound emission of the speaker, respectively.

[0020] Demodulating the received echo signal in formula (2) yields the in-phase baseband signal s. I (t) and orthogonal baseband signal s Q (t),

[0021] s I (t)=s T (t)×sR (t) (4)

[0022]

[0023] The phase unwinding method for extracting speech signals is as follows: First, the inverted baseband signal is divided by the in-phase baseband signal, and the result is arctangent, thus extracting the phase signal that changes with time, i.e., the displacement signal of the target acoustic diaphragm; Second, phase unwinding is performed by subtracting 2π from the phase difference when the phase difference between consecutive phase values ​​is greater than or less than ±π; Finally, the phase difference operation is performed on the unwinded phase signal to obtain the noisy speech signal.

[0024] The speech enhancement method based on the minimum mean square error of short-time amplitude spectrum is as follows.

[0025] The amplitude estimator for the speech enhancement algorithm based on the minimum mean square error of the short-time amplitude spectrum is:

[0026]

[0027] Estimate A from the noisy signal y(t) k The optimal solution to the above equation is:

[0028]

[0029] Where Γ is the gamma function, I0 and I1 are the modified zeroth and first-order Bessel functions, and v k For the prior signal-to-noise ratio and the posterior signal-to-noise ratio γ k Defined gain value.

[0030] Compared with existing technologies, the present invention has at least the following beneficial effects: the system breaks through the barriers of visual and auditory information, realizing the "visualization" of speech signals; the speech acquisition system and method based on metasurface described in the present invention improves human beings' ability to process speech information; it endows human hearing with the ability to extract specific target sound sources from the mixture of multiple sound sources and background; the structure is simple and easy to implement, significantly improving efficiency; and it provides a speech perception framework for deaf and mute people, enabling them to "see" speech signals. Attached Figure Description

[0031] Figure 1 This is a schematic diagram of the metasurface-based voice acquisition system used.

[0032] Figure 2 This is a schematic diagram of the electromagnetic response simulation under different phase arrangements of the metasurface.

[0033] Figure 3 These are the waveforms of (b) the original speech, (c) the speech acquired by the system, and (d) the enhanced speech signal in the embodiments of the present invention, as well as (e) the speech recognition result. Detailed Implementation

[0034] This invention proposes a metasurface-based voice acquisition system based on a vision-driven metasurface platform. This system aims to provide healthcare services to the deaf and mute and overcome communication barriers. The system includes... Figure 1 As shown, a barrier-free speech acquisition and enhancement system based on a vision-driven metasurface platform is proposed. The deaf person first observes their environment and selects a target of interest that can emit sound, such as a loudspeaker. Next, an eye tracker collects the deaf person's eye movement signals in real time. The eye movement signals are analyzed, and specific eye movements, such as repeated blinking or looking at the target for more than a time threshold, are used to determine the area where the deaf person wants to perceive speech. Then, the metasurface guides electromagnetic waves to the target for electromagnetic irradiation, and the returned electromagnetic wave signals are received in real time. The echo signals are phase-demodulated to obtain the speech signal of the acoustic diaphragm based on metasurface measurements. The speech signal is further denoised to improve the signal-to-noise ratio. Then, a speech recognition algorithm is used to process the denoised speech signal to obtain the corresponding speech recognition result. The speech recognition result is displayed on the deaf person's eyes using a wearable display, showing the corresponding text information. In this way, sound wave information is converted into microwave information, and then into a visual signal directly visible to the human eye.

[0035] A voice acquisition and extraction system based on a metasurface includes a phase gradient metasurface, an eye tracker, a wearable display, and a vector network analyzer. The phase gradient metasurface is connected to the output of a controller, and the wearable display is connected to the eye tracker. The phase gradient metasurface is used to emit electromagnetic waves of a set wavelength band. The vector network analyzer is used to receive reflected electromagnetic waves from a target object, wherein the reflected electromagnetic waves carry surface vibration information of the target object. The phase gradient metasurface includes multiple units, each unit having a length and width more than four times the wavelength corresponding to the working center frequency. Each unit is equipped with an electrically tunable device, wherein the electrically tunable device is a diode or a varactor diode.

[0036] To explain in detail the technical content, structural features, objectives, and effects of the invention, the invention will be further described below with reference to the accompanying drawings. The invention includes the following steps:

[0037] The system referred to in this invention includes an eye tracker, a vector network analyzer, a phase gradient metasurface, and a wearable display, etc.; the phase gradient metasurface is used to emit electromagnetic waves of a set band to a specific angle and region; the vector network analyzer is used to analyze the reflected electromagnetic waves of a target object, wherein the reflected electromagnetic waves carry the surface vibration information of the target object.

[0038] First, the deaf-mute person observes a target sound source of interest, and an eye tracker records their eye movements. Through an eye-tracking analysis program, the program reads the coordinates of the direction the deaf-mute person is looking at when they blink multiple times, and it can also read the coordinates of a region after the person has been looking at it for more than a time threshold.

[0039] Next, after processing by the eye-tracking analysis program, the required direction of electromagnetic wave irradiation can be determined. After the irradiation direction is determined, the phase gradient of the metasurface is determined according to the generalized Snell's law, thereby changing the direction of electromagnetic wave reflection.

[0040] The electromagnetic metasurface used consists of multiple units arranged in a grid, with a length and width more than four times the wavelength corresponding to the operating center frequency. Each unit has a diode or varactor diode welded onto it. In a preferred embodiment of the invention, a varactor diode, model SMV1430-040LF, is selected. By adjusting the voltage across the varactor diode in real time, the reflection phase of each metasurface unit can be changed in real time. Figure 2 (a) and (b) show the reflection amplitude and phase of the metasurface unit under different voltages, respectively. When four different voltages (2V, 6V, 9V, 23V) are applied across the varactor diode, the metasurface unit has different reflection phases, which are labeled "00", "01", "10", and "11", respectively. By arranging the metasurface units with different reflection phases, different phase gradients can be obtained, and thus different far-field electromagnetic radiation modes can be obtained. Figure 2 (c)-(e) show the metasurface unit arrangements with different reflection phases and their corresponding far-field radiation. Figure 2 (c)(d) are arranged in sequence as “00”, “01”, “10” and “11”, and the electromagnetic waves radiate in the ±36° direction. Figure 2 (e) When all metasurface unit states are "00", the electromagnetic wave radiates in the vertical direction. It should be noted that the voltage selected here is an example, and the voltage value is not limited to the four values ​​of 2V, 6V, 9V, and 23V. With more precise voltage selection, the phase gradient can be adjusted to the desired value, thereby arbitrarily adjusting the radiation direction of the electromagnetic wave. In this case, the electromagnetic beam is pointed at the audio source to detect surface vibrations driven by the speech signal; finally, the reflected electromagnetic wave signal containing surface vibration information is used with a vector network analyzer, and the micro-displacement information of the audio signal is contained in the phase of this electromagnetic wave signal.

[0041] As an example, the target of the test was a loudspeaker. One loudspeaker was set as the target sound source and connected to a laptop via cable, repeatedly playing the phrase, "Good Morning. How are you doing?" Another loudspeaker was connected via Bluetooth and played ambient white noise to interfere with the target sound source's audio signal. A deaf person, using an eye tracker, adjusted the beam of a metasurface to point towards the target sound source, detecting the faint vibration signal of the speaker's eardrum.

[0042] The time-domain signal of the sound emitted by the metasurface and reflected onto the target is denoted as s. T (t), the received signal collected by the receiving antenna is denoted as s. R (t), represented as follows:

[0043] s T (t)=A t cos(ωt) (1)

[0044]

[0045] In the formula, A t and A r Let be the amplitudes of the transmitted and received signals, respectively; ω be the angular frequency of the transmitted signal; λ be the carrier wavelength; c be the speed of light; d0 be the distance between the metasurface and the acoustic diaphragm; and x(t) represent the instantaneous displacement of the acoustic diaphragm when the acoustic diaphragm emits sound, which can be expressed as follows:

[0046]

[0047] Among them, A s and These represent the amplitude and phase shift of the diaphragm displacement caused by the sound emission of the speaker, respectively.

[0048] Demodulating the received echo signal in formula (2) yields the in-phase baseband signal s. I (t) and orthogonal baseband signal s Q (t),

[0049] s I (t)=s T (t)×s R (t) (4)

[0050]

[0051] The above s were obtained I (t) and s Q(t) After the baseband signal, this invention employs the most common phase unwinding method to extract the speech signal. The specific implementation steps of this algorithm are as follows: First, the inverted baseband signal is divided by the in-phase baseband signal, and the result is arctangent, thus extracting the time-varying phase signal, i.e., the displacement signal of the target speaker diaphragm. Second, phase unwinding is performed. Since the phase value is between [-π, π], phase unwinding is required to obtain the actual displacement of the diaphragm. This is achieved by subtracting 2π from the phase difference when the phase difference between consecutive phase values ​​is greater than or less than ±π. Finally, phase difference operation is performed on the unwound phase signal to enhance the vibration signal and eliminate any potential phase shift. Through the above operations, a noisy speech signal can be obtained.

[0052] Because the original extracted microwave speech information contains strong noise, it affects the accurate recognition of the speech information. This invention uses a speech enhancement algorithm based on minimum mean square error of short-time amplitude spectrum (MMSE-STSA) to suppress noise.

[0053] Assuming the real speech signal x(t) and the environmental noise signal n(t) are independent, which holds true for this scenario, the signal model can be obtained as follows:

[0054] y(t)=x(t)+n(t) (6)

[0055] In the formula, y(t) is the noisy microwave voice signal after phase unwinding.

[0056] Let X k =A k exp(jα k ), N k , and Y k =R k exp(jαθ k Let x(t), n(t), and y(t) represent the k-th spectral components of the signal x(t), noise n(t), and noise observation y(t), respectively. It is also assumed that all spectral components follow a Gaussian statistical model.

[0057] The MMSE STSA amplitude estimator can be written as:

[0058]

[0059] The objective of this invention is to estimate A from a noisy signal y(t). k The optimal solution of equation (7) can be derived as:

[0060]

[0061] Where Γ is the gamma function, and I0 and I1 are the modified zeroth and first-order Bessel functions. k For the prior signal-to-noise ratio and the posterior signal-to-noise ratio γ k Defined gain value.

[0062] This algorithm can effectively enhance audio information from noisy microwave speech signals. After obtaining a clean speech signal, it uses Microsoft Azure, a speech recognition module developed by Microsoft, to obtain the final speech recognition result, such as... Figure 3 As shown.

[0063] Unlike existing technologies, this invention starts with electromagnetic metasurfaces and obtains complete speech signals by analyzing reflected electromagnetic waves, rather than a series of simple commands. Simultaneously, the system can eliminate the influence of environmental noise and multiple sound sources, giving human hearing the ability to extract specific target sound sources from a mixture of various sound sources and background.

[0064] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the implementation methods of the present invention, and should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of the present invention.

Claims

1. A speech acquisition and extraction system based on metasurfaces, characterized in that, The system includes a phase gradient metasurface, an eye tracker, a wearable display, and a vector network analyzer. The phase gradient metasurface is connected to the output of the controller, and the wearable display is connected to the eye tracker. The phase gradient metasurface emits electromagnetic waves of a set wavelength. The vector network analyzer receives reflected electromagnetic waves from the target object, which carry surface vibration information of the target object. The phase gradient metasurface comprises multiple units, each unit having a length and width at least four times the wavelength corresponding to the working center frequency. Each unit is equipped with an electrically tunable device, which is a diode or a varactor diode. During operation, the eye tracker selects the sound source to be collected based on the gaze state of the human eye. The electromagnetic wave reflection mode of the phase gradient metasurface is changed so that electromagnetic waves illuminate the selected sound source. The vector network analyzer receives vibration signals reflected from the sound source and first uses the phase unwinding method to extract noisy speech signals from the electromagnetic wave echo. Secondly, the speech signal is enhanced using the short-time amplitude spectrum minimum mean square error method; finally, the speech recognition module is used to obtain the final speech-to-text recognition result.

2. The speech acquisition and extraction system based on metasurfaces according to claim 1, characterized in that, The controller uses an FPGA module or a microcontroller.

3. A speech acquisition and extraction method based on the metasurface-based speech acquisition and extraction system as described in claim 1 or 2, characterized in that, Includes the following steps: The eye tracker selects the sound source to be collected based on the gaze state of the human eye; The electromagnetic wave reflection mode of the phase gradient metasurface is changed so that electromagnetic waves illuminate the selected sound source. The vector network analyzer receives vibration signals reflected from the sound source and first uses the phase unwinding method to extract noisy speech signals from the electromagnetic wave echo. Secondly, the speech signal is enhanced using the short-time amplitude spectrum minimum mean square error method; finally, the speech recognition module is used to obtain the final speech-to-text recognition result.

4. The speech acquisition and extraction method according to claim 3, characterized in that, Eye trackers record eye movements, including fixation points, fixation duration, and saccade information. The eye trackers track and determine the user's observation area and points of interest in real time.

5. The speech acquisition and extraction method according to claim 3, characterized in that, By changing the feed voltage of the electrically tunable device on the phase gradient metasurface, the phase gradient and operating mode of the control electromagnetic metasurface are controlled, thereby adjusting the electromagnetic wave radiation region.

6. The speech acquisition and extraction method according to claim 3, characterized in that, Once the irradiation direction is determined, the phase gradient of the metasurface is determined according to the generalized Snell's law.

7. The speech acquisition and extraction method according to claim 3, characterized in that, The in-phase baseband signal and the quadrature baseband signal are obtained from the time-domain signal of the target acoustic reflection irradiated by the metasurface and the vibration signal received by the vector network analyzer. The time-domain signal of the target acoustic reflection irradiated by the metasurface is s. T (t), the received signal collected by the receiving antenna is s R (t), represented as follows: s T (t)=A t cos(ωt) (1) In the formula, A t and A r Let ω be the amplitude of the transmitted signal and λ be the wavelength of the carrier frequency, c be the speed of light, and d0 be the distance between the metasurface and the acoustic diaphragm. x(t) represents the instantaneous displacement of the acoustic diaphragm when the acoustic diaphragm emits sound, as follows: Among them, A s and These represent the amplitude and phase shift of the diaphragm displacement caused by the sound emission of the speaker, respectively. Demodulating the received echo signal in formula (2) yields the in-phase baseband signal s. I (t) and orthogonal baseband signal s Q (t), s I (t)=s T (t)×s R (t) (4) 8. The speech acquisition and extraction method according to claim 7, characterized in that, The phase unwinding method for extracting speech signals is as follows: First, the inverted baseband signal is divided by the in-phase baseband signal, and the result is arctangent, thus extracting the phase signal that changes with time, i.e., the displacement signal of the target acoustic diaphragm; Second, phase unwinding is performed by subtracting 2π from the phase difference when the phase difference between consecutive phase values ​​is greater than or less than ±π; Finally, the phase difference operation is performed on the unwinded phase signal to obtain the noisy speech signal.

9. The method for voice acquisition and extraction based on metasurfaces according to claim 3, characterized in that, The speech enhancement method based on the minimum mean square error of short-time amplitude spectrum is as follows. The amplitude estimator for the speech enhancement algorithm based on the minimum mean square error of the short-time amplitude spectrum is: Estimate A from the noisy signal y(t) k The optimal solution to the above equation is: Where Γ is the gamma function, I0 and I1 are the modified zeroth and first-order Bessel functions, and v k For the prior signal-to-noise ratio and the posterior signal-to-noise ratio γ k Defined gain value.

Citation Information

Patent Citations

  • Super surface structure with adjustable sound wave focusing

    CN107492370A

  • Metasurface with light redirecting structure comprising multiple materials and method of manufacture

    CN114641713A