A method for correcting bone-conducted voice signals
By performing peak processing, weighted filtering and phase recovery on the bone conduction signal, the problem of large calculation volume, insufficient anti-interference ability and stability of the bone conduction voice signal correction method is solved, and high recognition and stable bone conduction signal correction is achieved.
Patent Information
- Application Number
- CN202210408067.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-19
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-04-19
AI Technical Summary
The existing bone-guided voice signal correction methods have large calculation volume, insufficient anti-interference ability and stability.
By collecting and processing the bone conduction signal, peak processing is performed to extract characteristic frequencies, weighted filtering is performed, and the signal is phase restored to obtain the modified bone conduction signal.
It realizes high recognition and stability of bone conduction signals, reduces the calculation amount of the correction process, and improves the anti-interference ability.
Smart Images

Figure CN114842865B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of acoustic-electric signal conversion and digital signal processing, and specifically to a method for correcting bone-conducted voice signals. Background Art
[0002] The existence of environmental noise causes many troubles and inconveniences to voice communication. Compared with traditional air-conduction processing methods, bone-conducted voice mainly collects the vibration signals of the speaker's vocalization part, can shield the strong noise in the speaker's surrounding environment, and has a good noise reduction effect.
[0003] However, bone-conducted voice has problems such as low clarity and poor recognition, and it is necessary to correct bone-conducted voice. The existing bone-conducted signal correction methods are mainly divided into two categories: machine learning and analog circuit methods. Each method has its own advantages and disadvantages. The machine learning method has good effects, but requires a large amount of sample data and high computing power; the analog circuit method is simple and low-cost, but has poor anti-interference ability and unstable circuits. Summary of the Invention
[0004] The problem solved by the present invention is how to reduce the computational amount of the bone-conducted signal correction method and improve the anti-interference ability and stability.
[0005] To solve the above problems, the present invention provides a method for correcting bone-conducted voice signals, including the following steps:
[0006] Step 1: Collect and process to obtain a bone-conducted signal S;
[0007] Step 2: Perform peak processing on the bone-conducted signal S to obtain a characteristic frequency ω i ;
[0008] Step 3: Perform weighted filtering processing on the bone-conducted signal S to obtain a signal S L ;
[0009] Step 4: Perform phase recovery on each corresponding characteristic frequency ω in the signal S L to obtain a corrected bone-conducted signal Y as an output signal. i
[0010] The beneficial effects of the present invention are: by using digital signal processing technology to correct bone-conducted signals, the corrected signals have clear recognition, strong stability, and can also reduce the computational amount of bone-conducted signal correction.
[0011] Preferably, the specific steps of Step 1 include:
[0012] Step 101: Collect vibration signals at at least 1 position on the human head through a sensor;
[0013] Step 102: Determine whether the number of vibration signals is equal to 1. If so, this vibration signal is directly used as the vibration signal x(t) with the maximum signal-to-noise ratio. If not, select the vibration signal x(t) with the maximum signal-to-noise ratio from the vibration signals through a discriminator.
[0014] Step 103: After overall amplification of the vibration signal x(t), perform short-time Fourier transform to obtain the bone conduction signal S.
[0015] Since the collected vibration signal is relatively weak, it is necessary to perform overall amplification on the vibration signal and then perform short-time Fourier transform on the amplified weak vibration signal without distortion.
[0016] Preferably, the characteristic frequency ω i includes the fundamental frequency ω 0 and the harmonic frequency ω j ; The specific steps of step 2 are as follows:
[0017] Step 201: Based on the lower and upper limits of the human vocal frequency band, detect the fundamental frequency ω 0 ;
[0018] Step 202: Based on the fundamental frequency ω 0 , find the amplitude and frequency [A , ω n , ω n of the maximum peak between
[0019] Step 203: Preset the maximum receiving frequency value ω fmax of the human ear. Based on ω n , set the detection area Judge whether it is less than or equal to ω fmax . If not, go to step 3. If so, judge whether the maximum peak can be found between . If the maximum peak can be found, mark the amplitude and frequency [A n+1 , ω n+1 and go to step 204. If the maximum peak cannot be found, select the amplitude and frequency [A n+1 , ω n+1 of the center point in the retrieval area and go to step 204;
[0020] Step 204: n = n + 1, and return to 203.
[0021] Preferably, the formula for weighted filtering processing of the bone conduction signal S in step 3 is:
[0022]
[0023] wherein, a i is the weight coefficient obtained by comparing the amplitudes of the fundamental frequency and harmonic frequencies of the bone conduction signal and the traditional air conduction voice signal; is the band-pass filtering process, and the lower limit of the bandwidth of the band-pass filtering process is ω i (1-α), and the upper limit of the bandwidth of the band-pass filtering process is ω i (1 + α); thereby overcoming the problem of large differences in the spectral characteristics between the bone conduction signal S and the corresponding voice signal;
[0024] Preferably, step 4 specifically includes:
[0025] Step 401, based on the short-time Fourier transform in step 103 Take the logarithm to obtain the phase information
[0026] Step 402, based on the phase information Perform phase recovery on S L to obtain the corrected bone conduction signal Y, and a good voice correction effect can be achieved after phase recovery. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is the flowchart of the present invention;
[0028] Figure 2 is the spectrum comparison diagram of the bone conduction signal S and the voice signal of the present invention;
[0029] Figure 3 is the spectrum comparison diagram of the bone conduction signal Y and the voice signal after voice correction processing of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] In order to make the above objects, features and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be given with reference to the accompanying drawings.
[0031] As Figure 1 shown, a method for correcting a bone conduction voice signal includes the following steps:
[0032] Step 1, collect and process to obtain the bone conduction signal S; specifically including
[0033] Step 101, collect the vibration signals at at least one position of the human head through a sensor. In this specific embodiment, the vibration signals at the positions of the human larynx, forehead and behind the ear are collected by means of piezoelectric sheets or vibration sensors;
[0034] Step 102: Determine whether the number of vibration signals is equal to 1. If so, this vibration signal is directly used as the vibration signal x(t) with the maximum signal-to-noise ratio. If not, the vibration signal x(t) with the maximum signal-to-noise ratio is selected from the vibration signals through a discriminator. In this specific embodiment, the discriminator selects the vibration signal x(t) with the maximum signal-to-noise ratio, which is a prior art and will not be elaborated here.
[0035] Step 103: Since the collected vibration signal is relatively weak, it is necessary to amplify the vibration signal as a whole. On the premise of no distortion, the vibration signal x(t) is amplified as a whole and then subjected to short-time Fourier transform to obtain the bone conduction signal S, as Figure 2 shown;
[0036] Step 2: Perform peak processing on the bone conduction signal S to obtain the characteristic frequency ω i ; The characteristic frequency ω i includes the fundamental frequency ω 0 and the harmonic frequency ω j ; Specifically, it includes:
[0037] Step 201: Based on the lower and upper limits of the human vocal frequency band, detect the fundamental frequency ω 0 ;
[0038] Step 202: Based on the fundamental frequency ω 0 , find the amplitude and frequency [A , ω n , ω n of the maximum peak between
[0039] Step 203: Preset the maximum receiving frequency value ω fmax of the human ear. Based on ω n , set the detection area Judge whether it is less than or equal to ω fmax . If not, go to Step 3. If so, judge whether the maximum peak can be found between . If the maximum peak can be found, mark the amplitude and frequency [A n+1 , ω n+1 , and go to Step 204. If the maximum peak cannot be found, select the amplitude and frequency [A n+1 , ω n+1 of the short-time Fourier transform at the center point within the retrieval area, and go to Step 204;
[0040] Step 204: n = n + 1, and return to 203;
[0041] Step 3. To overcome the problem that the spectral characteristics of the bone-conducted signal S and the corresponding speech signal are quite different, the bone-conducted signal S is subjected to weighted filtering to obtain signal S L ; The formula for weighted filtering is:
[0042]
[0043] In the formula, a i is the weight coefficient obtained by comparing the amplitudes of the fundamental frequency and harmonic frequencies of the bone-conducted signal and the traditional air-conducted speech signal; is the band-pass filtering process, and the lower limit of the bandwidth of the band-pass filtering process is ω i (1-α), and the upper limit of the bandwidth of the band-pass filtering process is ω i (1 + α);
[0044] Step 4. For each characteristic frequency ω L corresponding in signal S i phase recovery is performed to obtain the corrected bone-conducted signal Y as the output signal, specifically including:
[0045] Step 401. Based on the short-time Fourier transform in step 103 take the logarithm to obtain the phase information
[0046] Step 402. Based on the phase information perform phase recovery on S L to obtain the corrected bone-conducted signal Y, as Figure 3 shown, after phase recovery, a good speech correction effect can be achieved.
[0047] In this specific embodiment, the fundamental frequency and harmonic frequencies can be obtained through peak detection first. However, since the fundamental frequency component of the bone-conducted signal S accounts for too much proportion relative to the speech signal spectrum and the harmonic peaks decay too much, the bone-conducted signal S sounds very dull. Therefore, weighted filtering processing is adopted. It is necessary to use weighted filtering processing to correct the fundamental frequency and harmonic frequencies of the bone-conducted signal S, so as to restore it to the same degree as the speech signal frequency distribution and make it sound clearer; in this specific implementation, the center frequency of the weighted filter is the fundamental frequency and harmonic frequencies obtained by peak detection;
[0048] The bone-conducted signal after weighted filtering will lose the original phase information, resulting in reduced clarity and generating a certain degree of noise. Therefore, in this specific embodiment, the weighted filtered bone-conducted signal is corrected according to the phase information of the original bone-conducted signal, so as to obtain a clear bone-conducted signal Y output.
[0049] Although the present disclosure is disclosed as above, the scope of protection of the present disclosure is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present disclosure, and these changes and modifications will all fall within the scope of protection of the present invention.
Claims
1. A method for correcting bone-conducted voice signals, characterized in that, it includes the following steps: Step 1: Collect and process to obtain a bone-conducted signal S; Step 2: Perform peak processing on the bone conduction signal S to obtain the characteristic frequency ; Specifically, it includes: Step 201: Detect the fundamental frequency based on the lower and upper limits of the human vocal frequency band ; Step 202: Based on the fundamental frequency , search for the amplitude and frequency of the maximum peak ; Step 203: Preset the maximum receiving frequency value of the human ear , based on set the detection area , and judge whether it is less than or equal to . If not, go to step 3; if so, judge whether the maximum peak can be found between . If the maximum peak can be found, mark the amplitude and frequency of the maximum peak , and go to step 204. If the maximum peak cannot be found, select the short-time Fourier transform amplitude and frequency of the center point within the retrieval area , and go to step 204; Step 204, and return 203; Step 3: Perform weighted filtering on the bone conduction signal S to obtain the signal ; The formula for performing weighted filtering on the bone conduction signal S is: ; In the formula, is the weight coefficient obtained by comparing the amplitude ratios of the fundamental frequencies and harmonic frequencies of the bone conduction signal and the traditional air conduction voice signal; is the band-pass filtering process, and the lower limit of the bandwidth of the band-pass filtering process is and the upper limit of the bandwidth of the band-pass filtering process is ; Step 4. For each characteristic frequency corresponding in the signal perform phase recovery to obtain the corrected bone conduction signal Y as the output signal.
2. The method for correcting bone-conducted voice signals according to claim 1, characterized in that, the specific content of said Step 1 includes: Step 101: Collect vibration signals at at least 1 position on the human head through a sensor; Step 102: Determine whether the number of vibration signals is equal to 1. If so, this vibration signal is directly used as the vibration signal with the maximum signal-to-noise ratio. If not, the decision maker is used to select the vibration signal with the maximum signal-to-noise ratio from the vibration signals. ; Step 103: After overall amplification of the vibration signal perform short-time Fourier transform to obtain the bone conduction signal S.
3. The method for correcting bone-conducted voice signals according to claim 2, characterized in that, the specific content of said Step 4 includes: Step 401: Based on the short-time Fourier transform in Step 103 , take the logarithm , to obtain phase information ; Step 402: Based on the phase information Perform phase recovery to obtain the corrected bone-conducted signal Y.
Citation Information
Patent Citations
Voice signal processing method and apparatus
CN105845146A
Bone conduction hearing aid automatic gain control method based on electroencephalogram EEG
CN110366086A