Biological radar-based speech processing method and apparatus
By combining the throat vibration signal collected by the bio-radar sensor with the microphone, the problem of noise removal in the existing voice signal is solved, thereby improving the quality of voice calls and reducing the amount of data transmitted.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-07
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies have failed to effectively utilize bio-radar sensors to process human speech signals, resulting in the inability to effectively remove noise from speech signals.
By collecting vibration signals from the human throat using a bio-radar sensor and combining them with voice signals collected by a microphone, the voice signals of non-target individuals are removed and noise in the voice signals are eliminated using feature parameter comparison and noise filtering algorithms.
This technology enables noise cancellation in the audio signal captured by the microphone, improving the quality of voice calls and reducing the amount of audio signal transmitted.
Smart Images

Figure CN116778941B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of speech processing, and in particular to a speech processing method and device based on biological radar. BACKGROUND
[0002] The non-contact signal detection technology based on biological radar uses electromagnetic wave as the detection medium. When the electromagnetic wave reaches the human body, the body surface micro-motion caused by physiological activities modulates the electromagnetic wave phase and frequency, and the physiological signals of the human body can be obtained by demodulating the received radar echo signal.
[0003] Therefore, the application research of biological radar detection technology is mainly to measure the movement of the human vocal organ, and the radar sensor has not been directly applied to the detection of human speech signals. SUMMARY
[0004] Therefore, the application research of biological radar detection technology is mainly to measure the movement of the human vocal organ, and the radar sensor has not been directly applied to the detection of human speech signals.
[0005] An embodiment of the present application provides a speech processing method based on biological radar, which comprises the following steps: acquiring a speech signal collected by a microphone; acquiring a vibration signal collected by a biological radar sensor; reserving an effective speech signal in the speech signal according to the energy distribution of the speech signal;
[0006] reserving an effective vibration signal in the vibration signal according to the energy distribution of the vibration signal; removing a speech signal of a non-target person and reserving a speech signal of a target person according to the effective speech signal and the effective vibration signal; extracting a characteristic parameter of the speech signal of the target person and a characteristic parameter of the effective vibration signal respectively, comparing waveform similarity according to the characteristic parameter of the speech signal of the target person and the characteristic parameter of the effective vibration signal; judging whether the waveform similarity is less than a similarity threshold value; and when it is judged that the waveform similarity is less than the similarity threshold value, removing noise in the speech signal of the target person according to the envelope and the spectrum distribution of the effective vibration signal, and outputting the speech signal of the target person.
[0007] The embodiment of the present application also provides a voice processing device based on biological radar, characterized by comprising: a microphone for collecting voice signals; a biological radar sensor for collecting vibration signals; a computer readable storage medium for storing at least one instruction; and a processor for executing the at least one instruction to realize the following steps: obtaining the voice signals; obtaining the vibration signals; reserving effective voice signals in the voice signals according to the energy distribution of the voice signals; reserving effective vibration signals in the vibration signals according to the energy distribution of the vibration signals; removing non-target person voice signals and reserving target person voice signals according to the effective voice signals and the effective vibration signals; extracting feature parameters of the target person voice signals and feature parameters of the effective vibration signals respectively, and comparing waveform similarity according to the feature parameters of the target person voice signals and the feature parameters of the effective vibration signals; judging whether the waveform similarity is less than a similarity threshold; and when judging that the waveform similarity is less than the similarity threshold, removing noise in the target person voice signals according to the envelope and the spectrum distribution of the effective vibration signals, and outputting the target person voice signals.
[0008] Compared with the prior art, the voice processing method and device based on biological radar provided by the present application can eliminate noise in the voice signals collected by the microphone according to the vibration signals collected by the biological radar. BRIEF DESCRIPTION OF DRAWINGS
[0009] Figure 1 A block diagram of the voice processing device according to the embodiment of the present application.
[0010] Figure 2 A flowchart of the voice processing method according to the embodiment of the present application.
[0011] Figure 3 A flowchart of extracting effective signals according to the embodiment of the present application.
[0012] Figure 4 A flowchart of removing non-target person voice signals according to the embodiment of the present application.
[0013] Figure 5 A flowchart of signal similarity comparison according to the embodiment of the present application.
[0014] Figure 6 A flowchart of removing noise in target person voice signals according to the embodiment of the present application.
[0015] MAIN ELEMENT SYMBOL EXPLANATION
[0016]
[0017]
[0018] The following detailed description of the application will further explain the application with reference to the accompanying drawings. DETAILED DESCRIPTION
[0019] For the purpose of clarity, detailed descriptions of well-known tools, circuits, and techniques of the application will not be discussed in detail herein. The foregoing detailed description of the application has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the application to the precise form disclosed. Many modifications and variations are possible in light of the above teaching. It is intended that the scope of the application be limited not with this detailed description. Rather, it is intended to cover this application broadly, within the scope of the appended claims and their equivalents. While the application has been described above with particularity, and reference has been made to embodiments of the application, this description is illustrative only, and not restrictive. Many other embodiments will become apparent to those of ordinary skill in the art, having the benefit of this disclosure, upon reading and interpreting this description and the associated drawings. Certain features that are, for clarity, described above and illustrated in the drawings can optionally be omitted and / or combined. Additionally, any suitable size, shape or type of materials or processes could be used as appropriate, and are intended to be within the scope of the claims. The claims should not be limited to the embodiments set forth in the detailed description, which are intended as being exemplary, but rather can include any other embodiments that fall within the scope of the claims.
[0020] The description of the application provided in this specification is provided for the purpose of illustrating and describing the application and its best mode. The detailed description is not intended to limit the application to the exact embodiment disclosed. Various modifications and variations are possible in light of the above teachings provided they come within the scope of the claims. It is therefore to be understood that within the scope of the claims the application can be practiced otherwise than as specifically described. While the application has been described above with particularity, and reference has been made to embodiments of the application, this description is illustrative only, and not restrictive. Many other embodiments will become apparent to those of ordinary skill in the art, having the benefit of this disclosure, upon reading and interpreting this description and the associated drawings. Certain features that are, for clarity, described above and illustrated in the drawings can optionally be omitted and / or combined. Additionally, any suitable size, shape or type of materials or processes could be used as appropriate, and are intended to be within the scope of the claims. The claims should not be limited to the embodiments set forth in the detailed description, which are intended as being exemplary, but rather can include any other embodiments that fall within the scope of the claims.
[0021] Moreover, in describing the present application, some embodiments are described in a specific order. However, the order of the steps described in the specification is not necessarily the order in which the steps are performed. The steps can be performed in any order, unless otherwise specified. Other sequences of steps can be possible, and are intended to be within the scope of the claims. The specific order of steps described in the specification is intended to illustrate the method and process of the present application, and is not intended to limit the scope of the claims. The claims should not be limited to the specific order of steps described in the specification. Moreover, the application is not limited to the specific steps described in the specification, and other steps can be included within the scope of the claims.
[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description of the application herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. As used in this description of the application, the following terms have the following meanings. The use of "and / or" means that one or all possible combinations are included. Some embodiments of the application are described in detail below with reference to the attached drawing figures and detailed description. It should be noted that the description set forth herein describes specific embodiments of the application and that the scope of the application is not limited to these specific embodiments.
[0023] Reference will now be made to Figure 1, as shown in a block diagram of a voice processing device 100 in an embodiment of the present application. The device 100 includes a processor 102, a computer readable storage medium 104, a bio-radar sensor 106, and a microphone 108. The bio-radar sensor 106 and the microphone 108 are coupled to the processor 102. Those skilled in the art should understand that, Figure 1 The components of the device 100 shown do not constitute a limitation of the embodiments of the present application, Figure 4 The device 100 shown is simplified for ease of description, and in different embodiments, can include fewer or more components than shown.
[0024] In an embodiment, the processor 102 can be composed of integrated circuits, for example, can be composed of a single packaged integrated circuit, or can be composed of multiple packaged integrated circuits of the same function or different functions, including one or more combinations of central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 102 is the control core of the device 100, and connects various components of the entire device 100 through various interfaces and lines, executes computer programs or modules stored in the computer readable storage medium 104, and calls data stored in the computer readable storage medium 104, to perform various functions of the device 100 and process data, such as a voice processing method.
[0025] In an embodiment, the computer readable storage medium 104 is configured to store computer program codes and various data, such as the voice processing method, and to realize high-speed and automatic access of programs or data during the operation of the device 100. The computer readable storage medium 104 includes a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically-erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, a magnetic disk storage, a magnetic tape storage, or any other computer readable storage medium capable of carrying or storing data.
[0026] The biological radar sensor 106 includes a biological radar signal transmitter 1061 and a biological radar signal receiver 1062. The biological radar signal transmitter 1061 is configured to generate and transmit a biological radar monitoring signal, and the biological radar signal receiver 1062 is configured to collect a reflected biological radar echo signal.
[0027] The microphone 108 is configured to collect a voice signal.
[0028] In the embodiment, the microphone 108 is configured to collect a voice signal in the environment, and the biological radar sensor 106 is configured to collect a throat vibration signal via a directional millimeter wave radar, such as a 94 GHz directional millimeter wave radar with less interference. Then, an algorithm is used to filter out clutter from the voice signal and optimize the voice signal.
[0029] Specifically, it is assumed that the single-frequency signal transmitted by the biological radar signal transmitter 1061 is P T (t) = A cos(2pft + 0i), where A is the vibration amplitude of the transmitted signal, f0is the frequency of the transmitted signal, and 0iis the initial phase. When the transmitted signal reaches the throat of a human body at a distance d0, the phase change of the echo signal caused by d0is 0i2, and the phase change of the echo signal caused by the vibration x(t) of the human throat is 4p x(t) / l0, then the echo signal received by the biological radar signal receiver 1062 can be represented as: Wherein, λ0=c / f0, c is the speed of light, K is the attenuation coefficient of the vibration amplitude of the transmitted signal. The radar echo signal is mixed with the transmitted signal: After low-pass filtering and removing DC, the baseband signal is: Wherein, Δθ is the phase offset caused by the distance d0 between the transmitted signal and the throat. When the displacement x(t) caused by the throat micro-motion of the human body is much smaller than the radar wavelength and Δθ / 2 is an odd multiple, the baseband signal can be represented as: At this time, the information of the throat vibration of the human body is contained in the demodulated baseband signal, and the vibration signal can be obtained by processing.
[0030] Please refer to Figure 2 , which is a flow chart of the speech processing method in an embodiment of the present application. The method can be executed by the processor 102 of the device 100 in Figure 1 , and mainly includes the following steps:
[0031] Step S201, detecting the speech signal and determining whether there is sound. When it is determined that there is sound, step S202 is executed, and when it is determined that there is no sound, the detection is continued.
[0032] Step S202, simultaneously collecting the speech signal and the vibration signal.
[0033] Specifically, the speech signal is collected through the microphone 108, and the vibration signal is collected through the biometric radar sensor 106. The simultaneously collected speech signal and vibration signal are sampled and stored in the computer readable storage medium 104 for further processing by the processor 102.
[0034] Step S203, retaining the valid speech signal in the speech signal according to the energy distribution of the speech signal.
[0035] Specifically, the energy distribution of the speech signal is analyzed, and the invalid speech signal with an energy value less than the energy threshold value is removed, and the remaining is the valid speech signal.
[0036] Step S204, retaining the valid vibration signal in the vibration signal according to the energy distribution of the vibration signal.
[0037] Specifically, the energy distribution of the vibration signal is analyzed, and the invalid vibration signal with an energy value less than the energy threshold value is removed, and the remaining is the valid vibration signal.
[0038] Wherein, the energy threshold value is determined according to the analysis result of the energy distribution of the speech signal and the vibration signal.
[0039] Step S205, removing the speech signal of the non-target person and retaining the speech signal of the target person according to the valid speech signal and the valid vibration signal.
[0040] Step S206, respectively extracting the feature parameters of the speech signal of the target person and the feature parameters of the effective vibration signal, and performing waveform similarity comparison according to the feature parameters of the speech signal of the target person and the feature parameters of the effective vibration signal.
[0041] In an example, the feature parameters include but are not limited to spectral coefficients.
[0042] In an example, the spectral coefficients include but are not limited to mel-frequency cepstral coefficients.
[0043] Step S207, judging whether the waveform similarity is less than a similarity threshold value. When it is judged that the waveform similarity is less than the similarity threshold value, step S208 is executed; when it is judged that the waveform similarity is greater than or equal to the similarity threshold value, step S209 is executed.
[0044] In an example, the similarity threshold value is 0.8.
[0045] Step S208, removing noise in the speech signal of the target person according to the envelope and spectral distribution of the effective vibration signal.
[0046] Step S209, outputting the speech signal of the target person.
[0047] In an embodiment, the specific flow steps of steps S203 and S204 are as shown in Figure 3 .
[0048] Step S301, filtering out the high-frequency part and the low-frequency part of the speech signal and the vibration signal.
[0049] Step S302, analyzing the energy distribution of the speech signal and the vibration signal.
[0050] Step S303, removing invalid signals in the speech signal and the vibration signal with energy values less than an energy threshold value.
[0051] Step S304, removing invalid signals in the speech signal and the vibration signal with energy values greater than or equal to the energy threshold value but with time durations less than a time length threshold value.
[0052] In an example, the energy threshold value is 0.1, and the time length threshold value is 0.2 seconds.
[0053] Step S305, extracting and storing the processed effective speech signal and the effective vibration signal.
[0054] In an embodiment, the specific flow steps of step S205 are as shown in Figure 4 .
[0055] Step S401, synchronizing the effective speech signal and the effective vibration signal according to the energy values and time indexes of the effective speech signal and the effective vibration signal.
[0056] Step S402, according to the valid vibration signal, remove the part of the valid speech signal without corresponding valid vibration signal.
[0057] Step S403, extract and save the valid speech signal after processing.
[0058] At this time, the saved valid speech signal is the speech signal of the target person.
[0059] In an embodiment, the specific flow steps of step S206 and step S207 are as shown in Figure 5
[0060] Step S501, calculate the mel-frequency cepstral coefficients of the valid speech signal and the synchronized valid vibration signal respectively.
[0061] Step S502, perform waveform similarity comparison on the mel-frequency cepstral coefficient curve of the valid speech signal and the mel-frequency cepstral coefficient curve of the valid vibration signal.
[0062] Step S503, judge whether the waveform similarity is less than the similarity threshold value. When it is judged that the similarity is less than the similarity threshold value, execute step S504; when it is judged that the waveform similarity is greater than or equal to the similarity threshold value, end the flow step.
[0063] Step S504, remove the noise in the speech signal of the target person.
[0064] In an embodiment, the specific flow steps of step S208 are as shown in Figure 6
[0065] Step S601, synchronize the valid vibration signal and the speech signal of the target person according to the energy value and the time index.
[0066] Step S602, obtain the envelope amplitude of the valid vibration signal.
[0067] Step S603, obtain the spectral amplitude of the valid vibration signal.
[0068] Step S604, filter out the noise in the speech signal of the target person according to the envelope amplitude and the spectral amplitude.
[0069] Step S605, save the processed speech signal of the target person.
[0070] In summary, the above-mentioned speech processing method and device based on biological radar use biological radar sensors to detect human throat vibration, filter and process the speech signal collected by the microphone, which can completely filter out the noise in the speech signal, thereby improving the quality of voice communication and reducing the transmission amount of the speech signal.
[0071] It is to be noted that the above examples are only used to illustrate the technical solutions of the present application, rather than limit the present application. Although the present application is described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or equivalent replaced, without departing from the spirit and scope of the technical solutions of the present application.
Claims
1. A speech processing method based on bio-radar, characterized in that, The method includes the following steps: Acquire the voice signal captured by the microphone; Acquire throat vibration signals collected by a bio-radar sensor; Based on the energy distribution of the speech signal, retain the valid speech signal in the speech signal; Based on the energy distribution of the throat vibration signal, the effective vibration signal in the throat vibration signal is retained; Synchronize the effective voice signal and the effective vibration signal according to the energy value and time label of the effective voice signal and the effective vibration signal; Based on the effective vibration signal, the portion of the effective voice signal that does not correspond to the effective vibration signal is removed, and the processed effective voice signal is the voice signal of the target person; The feature parameters of the voice signal and the feature parameters of the effective vibration signal of the target person are extracted respectively, and the waveform similarity is compared according to the feature parameters of the voice signal and the feature parameters of the effective vibration signal of the target person. Determine whether the waveform similarity is less than a similarity threshold; as well as When the waveform similarity is determined to be less than the similarity threshold, noise in the target person's voice signal is removed based on the envelope and spectral distribution of the effective vibration signal, and the target person's voice signal is output.
2. The speech processing method as described in claim 1, characterized in that, The method also includes the following steps: When the waveform similarity is determined to be greater than or equal to the similarity threshold, the voice signal of the target person is output.
3. The speech processing method as described in claim 1, characterized in that, The step of retaining the valid speech signal in the speech signal based on the energy distribution of the speech signal further includes: Filter out the high-frequency and low-frequency components of the speech signal; Analyze the energy distribution of the speech signal and remove invalid speech signals with energy values less than the energy threshold; as well as Remove invalid speech signals whose energy value is greater than or equal to the energy threshold but whose duration is less than the duration threshold.
4. The speech processing method as described in claim 1, characterized in that, The step of retaining the effective vibration signal from the throat vibration signal based on the energy distribution of the throat vibration signal further includes: Filter out the high-frequency and low-frequency components of the throat vibration signal; Analyze the energy distribution of the throat vibration signal and remove invalid vibration signals with energy values less than the energy threshold; as well as Remove invalid vibration signals whose energy value is greater than or equal to the energy threshold but whose duration is less than the duration threshold.
5. The speech processing method as described in claim 1, characterized in that, The step of extracting feature parameters of the target person's voice signal and feature parameters of the effective vibration signal, and comparing waveform similarity based on the feature parameters of the target person's voice signal and the feature parameters of the effective vibration signal, further includes: Calculate the Mel-frequency cepstral coefficients of the effective speech signal and the effective vibration signal respectively; The waveform similarity of the Mel-frequency cepstral coefficient curves of the effective speech signal and the effective vibration signal is compared.
6. The speech processing method as described in claim 1, characterized in that, The step of removing noise from the target person's speech signal based on the envelope and spectral distribution of the effective vibration signal further includes: The effective vibration signal and the voice signal of the target person are synchronized according to the energy value and time label; Obtain the envelope amplitude of the effective vibration signal; Obtain the spectral amplitude of the effective vibration signal; Noise in the speech signal of the target person is filtered out based on the envelope amplitude and the spectral amplitude.
7. The speech processing method as described in claim 4, characterized in that, The method also includes the following steps: The energy threshold is determined based on the analysis results of the energy distribution of the speech signal and the throat vibration signal.
8. A voice processing device based on bio-radar, characterized in that, The device includes: A microphone is used to collect voice signals; Bio-radar sensor; used to collect throat vibration signals; A computer-readable storage medium for storing at least one instruction; as well as A processor for executing the at least one instruction to perform the following steps: Acquire the voice signal; Acquire the throat vibration signal; Based on the energy distribution of the speech signal, retain the valid speech signal in the speech signal; Based on the energy distribution of the throat vibration signal, the effective vibration signal in the throat vibration signal is retained; Synchronize the effective voice signal and the effective vibration signal according to the energy value and time label of the effective voice signal and the effective vibration signal; Based on the effective vibration signal, the portion of the effective voice signal that does not correspond to the effective vibration signal is removed, and the processed effective voice signal is the voice signal of the target person; The feature parameters of the voice signal and the feature parameters of the effective vibration signal of the target person are extracted respectively, and the waveform similarity is compared according to the feature parameters of the voice signal and the feature parameters of the effective vibration signal of the target person. Determine whether the waveform similarity is less than a similarity threshold; as well as When the waveform similarity is determined to be less than the similarity threshold, noise in the target person's voice signal is removed based on the envelope and spectral distribution of the effective vibration signal, and the target person's voice signal is output.
Citation Information
Patent Citations
Speech enhancement method based on fusion of radar speech and microphone speech
CN109584894A
Method and device for acquiring sounding motion characteristic waveforms of multiple sounders, and electronic equipment
CN113257271A