Noise reduction electronic artificial larynx and noise reduction method based on ultrasonic voice

By employing an ultrasonic pitch and spectrum shifting module within the electronic artificial larynx, the problem of noise leakage was solved, resulting in clearer voice output and an improved user experience.

CN116327430BActive Publication Date: 2026-04-24HANGZHOU DIANZI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU DIANZI UNIV
Filing Date
2023-03-01
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing electronic artificial larynxes suffer from noise leakage when vibrating, leading to decreased speech clarity and impacting user experience.

Method used

The traditional pitch is replaced by an ultrasonic pitch, and a receiver and spectrum shifting module are added to the electronic artificial larynx. Ultrasonic waves are generated by an ultrasonic generator, and the echo signal is received by a receiver. Noise cancellation is achieved by combining a bandpass filter, a spectrum conversion module and a modulator.

Benefits of technology

It almost completely eliminates noise, improves speech clarity, reduces the user's vocal strain, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116327430B_ABST
    Figure CN116327430B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on ultrasonic fundamental tone's noise reduction electronic artificial throat and noise reduction method, wherein, transmitting receiving end includes ultrasonic generator and ultrasonic receiver;The PC end includes band-pass filter and spectrum transform module;The power module is powered for transmitting receiving end and PC end, and ultrasonic generator sends ultrasonic wave to human throat, and ultrasonic receiver receives echo of human throat, and transmitting receiving end sends ultrasonic signal to PC end, and the band-pass filter of PC end filters ultrasonic signal and sends to spectrum transform module to carry out noise reduction, conversion and output voice signal.The application replaces fundamental tone with ultrasonic wave, reduces frequency to sound fundamental tone frequency range in the ultrasonic fundamental tone in sound signal, realizes spectrum shift, and finally forms the voice signal audible to human ear.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biomedical acoustics, and specifically relates to a noise-reducing electronic artificial larynx based on ultrasonic fundamental tone and a noise reduction method. Background Technology

[0002] Human speech involves sound production, vibration, resonance, and amplification. Sound production is generated by the movement of airflow during exhalation in the lungs. Vibration occurs when the vocal cords vibrate to produce a fundamental tone. This fundamental tone is then amplified through resonance in the pharynx, oral cavity, and nasal cavity, as well as through amplification by the tongue, teeth, lips, and palate, ultimately becoming a recognizable sound. Therefore, for patients who have undergone laryngectomy, as long as a vibration source to replace the vocal cords can be found, other organs can be coordinated to resume speaking. Electronic artificial larynxes utilize this principle.

[0003] An electronic larynx is a battery-powered speech restoration device. Its basic principle is to use a mechanical device to replace the vibration of the vocal cords to generate a fundamental tone. The sound signal is transmitted through the patient's neck tissue to the pharynx. The signal is modulated (resonance and anti-resonance) in the upper vocal tract and radiated through the oral cavity, forming a near-human voice at the tip of the lips. Current electronic larynxes not only achieve clear and stable speech production but also allow for frequency and amplitude modulation. They are also compact, lightweight, inexpensive, and hygienic to use. Through continuous improvement and innovation, its indications are constantly expanding. It is now not only a speech rehabilitation tool for those without a larynx but also a treatment tool in otolaryngology, becoming a popular choice for many patients.

[0004] However, because electronic larynxes are placed on the outside of the neck, the vibration occurs outside the body when the patient presses the switch to speak. Since the contact between the artificial larynx and neck tissue is not tight enough, vibrational energy is easily leaked. This leaked signal (generally between 50-500Hz) is directly radiated by the electronic larynx itself, becoming radiated noise. This noise is heard simultaneously with the artificial larynx's speech, reducing the clarity of the speech and affecting the listening and speaking experience. Currently, there are approximately 600,000 laryngectomy patients worldwide, and the incidence of laryngeal cancer is constantly rising. As a product with an expanding target population and requiring long-term use, user experience is a crucial indicator. A poor user experience can significantly impact the user's daily and mental well-being, causing considerable suffering for the patient.

[0005] As the number of people using and the scope of applications for artificial larynx continue to expand, compact, convenient, relatively inexpensive, and hygienic electronic artificial larynxes will become the choice of more people. How to reduce noise and enhance the user experience is a question worth exploring.

[0006] Traditional electronic larynxes consist of four parts: an oscillator, a power amplifier, a transducer, and a power supply. The oscillator generates pulse waves, the power amplifier amplifies the power intensity of the pulse waves, and the transducer converts electrical energy into sound, which is then transmitted to the neck tissues. The transmitted fundamental tone is modulated and amplified by the pharynx, oral cavity, and nostrils, ultimately forming a sound that is discernible to the human ear. The user simply places the vibrating end of the electronic larynx near the glottis, and the signal generated by the vibration can replace the sound source and enter the vocal tract, causing airflow within the vocal tract and modulation through the oral cavity to form speech. However, because the fundamental tone is generated externally, some noise is inevitably generated, reducing the clarity of speech from the artificial larynx. Long-term use can cause many inconveniences in the patient's daily life and can easily lead to psychological stress. Summary of the Invention

[0007] In view of this, this invention addresses the noise problem caused by oscillations in electronic artificial larynxes by proposing the use of ultrasound instead of the fundamental tone. A receiver and a spectrum shifting module are added within the electronic artificial larynx to achieve noise cancellation. A noise-reducing electronic artificial larynx based on ultrasonic fundamental tone is provided, comprising a power module and a transmitter / receiver and a PC connected to the power module. The transmitter / receiver is connected to the PC.

[0008] The transmitting and receiving end includes an ultrasonic generator and an ultrasonic receiver; the PC end includes a bandpass filter and a spectrum conversion module; the power module supplies power to the transmitting and receiving end and the PC end. The ultrasonic generator sends ultrasonic waves to the human throat, the ultrasonic receiver receives the echo from the human throat, the transmitting and receiving end sends ultrasonic signals to the PC end, and the bandpass filter on the PC end filters the ultrasonic signals and sends them to the spectrum conversion module for noise reduction, conversion and output of voice signals.

[0009] Preferably, the spectrum conversion module includes a resampler and a duration warping unit. The resampler is connected to a bandpass filter to downsample the ultrasonic signal. The resampler is connected to the duration warping unit, which uses an overlap superposition algorithm to warp the downsampled signal.

[0010] Preferably, the spectrum conversion module includes a synchronous detector demodulator and a DSB modulator. The synchronous detector demodulator is connected to a bandpass filter to remove the fundamental frequency signal from the ultrasonic signal. The synchronous detector demodulator is connected to the DSB modulator, which modulates the received signal to form a voice signal and plays it.

[0011] To achieve the above objectives, the present invention also provides a noise reduction method for an electronic artificial larynx based on ultrasonic pitch, comprising the following steps:

[0012] S10 uses an ultrasonic generator to generate ultrasonic waves and transmit them into the human throat. It uses an ultrasonic receiver head to receive the transmitted ultrasonic signals outside the mouth and transmit them to the PC.

[0013] S21, Input high-frequency signal data: Import ultrasonic signals and perform digital processing;

[0014] S22, Bandpass filtering: The bandpass filter removes noise from the input signal;

[0015] S20 changes the base frequency of the signal and processes the data to output a voice signal.

[0016] Preferably, step S20 specifically includes the following steps:

[0017] S23, resample the signal to change the base frequency;

[0018] S24 uses an overlap superposition algorithm to perform duration normalization on the signal and generate a speech signal.

[0019] Preferably, step S24 specifically includes the following steps:

[0020] S241, extract W data points from the input sequence and store them directly into the output sequence;

[0021] S242, take W more data points at interval Sa points from the input sequence;

[0022] S243, take the W data points taken previously as the front sequence and the W data points taken later as the back sequence; compare the consistency between the last Wov data points of the front sequence and the first Wov data points of the back sequence, record the analysis results, and denote the number of comparisons as k (k is initially 0), k = k + 1;

[0023] S244, determine if k is equal to Kmax; if not, move the analysis window (i.e., capture a window of W points) and return to S243 to continue recording and comparing; if yes, proceed to S245.

[0024] S245, select the case with the most consistent results, superimpose the Wov data before and after, and then store the remaining Ss data from the W data into the output sequence;

[0025] S246, determine whether data processing is complete. If not, return to S242 to continue the value comparison process until signal processing is complete; if yes, end.

[0026] Where W is the window length, representing the minimum length of speech processing; Sa is the analysis delay, representing the interval between the first addresses of the speech segments that are sequentially extracted and processed; Ss is the synthesis delay, representing the interval between the first addresses of the speech segments that are sequentially output; Kmax is the lookup delay, which is the delay that must occur for the analysis window to be consistent with the end of the output signal; Wov is the length of the superposition of the previous and next speech segments, where the previous / next speech segments are the speech segments extracted first / last.

[0027] Preferably, step S20 specifically includes the following steps:

[0028] S33 performs synchronous detection and demodulation on the signal; generates a sine wave with the same frequency as the ultrasonic fundamental tone, multiplies it with the signal, and performs amplitude modulation demodulation through a low-pass filter to obtain a sound signal with the ultrasonic fundamental tone removed.

[0029] S34 modulates the signal; generates a low-frequency pitch signal, and modulates it with the audio signal on both sides to form the final speech signal and play it.

[0030] Preferably, the carrier signal used for modulation in S34 is in the range of 50-500Hz.

[0031] Beneficial effects: Compared with traditional electronic artificial larynx, this invention can almost completely eliminate the noise of electronic artificial larynx, and adds the functions of receiving end and filtering amplification, which can reduce the user's vocal pressure, bring a clearer and more relaxed vocal experience, bring more convenience to the communication of "laryngeal people", improve their quality of life, bring more comfort to their hearts, and also contribute a small part to the development of medicine. Attached Figure Description

[0032] Figure 1 This is a structural block diagram of a noise-reducing electronic artificial larynx based on ultrasonic pitch according to an embodiment of the present invention;

[0033] Figure 2 This is a structural block diagram of a noise-reducing electronic artificial larynx based on ultrasonic pitch, according to another embodiment of the present invention;

[0034] Figure 3 This is a flowchart illustrating the steps of a noise reduction method for an electronic artificial larynx based on ultrasonic pitch, according to an embodiment of the present invention.

[0035] Figure 4 This is a flowchart illustrating the specific steps of step S24 of a noise reduction method for an electronic artificial larynx based on ultrasonic pitch, according to an embodiment of the present invention.

[0036] Figure 5 This is a diagram showing the parameter relationships in S24 of a noise reduction method for an electronic artificial larynx based on ultrasonic pitch, according to an embodiment of the present invention.

[0037] Figure 6 This is a flowchart illustrating the steps of a noise reduction method for an electronic artificial larynx based on ultrasonic pitch, according to another embodiment of the present invention. Detailed Implementation

[0038] The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0039] See one embodiment of the electronic artificial larynx Figure 1 An ultrasonic pitch-based noise-reducing electronic artificial larynx includes a power module 40 and a transmitter / receiver 20 and a PC 30 connected to the power module 40. The transmitter / receiver 20 includes an ultrasonic generator 21 and an ultrasonic receiver 22. The PC 30 includes a bandpass filter 31 and a spectrum conversion module. The power module 40 supplies power to the transmitter / receiver 20 and the PC 30. The ultrasonic generator 21 sends ultrasonic waves to the human larynx 10, and the ultrasonic receiver 22 receives the echoes from the human larynx 10. The transmitter / receiver 20 sends ultrasonic signals to the PC 30. The bandpass filter 31 of the PC 30 filters the ultrasonic signals and sends them to the spectrum conversion module for noise reduction, conversion, and output of speech signals. The spectrum conversion module includes a resampler 32 and a duration warping unit 33. The resampler 32 is connected to a bandpass filter 31 to downsample the ultrasonic signal. The resampler 32 is connected to the duration warping unit 33, which uses an overlap superposition algorithm to warp the downsampled signal.

[0040] See another embodiment of the electronic artificial larynx Figure 2 The spectrum conversion module includes a synchronous detector demodulator 34 and a DSB modulator 35. The synchronous detector demodulator 34 is connected to a bandpass filter 31 to remove the fundamental frequency signal from the ultrasonic signal. The synchronous detector demodulator 34 is connected to the DSB modulator 35, which modulates the received signal to form a voice signal and plays it.

[0041] To achieve the above objectives, the present invention also provides a noise reduction method for an electronic artificial larynx based on ultrasonic pitch, employing the aforementioned electronic artificial larynx, see [link to relevant documentation]. Figure 3 This includes the following steps:

[0042] S10 uses an ultrasonic generator to generate ultrasonic waves and transmit them into the human throat. It uses an ultrasonic receiver head to receive the transmitted ultrasonic signals outside the mouth and transmit them to the PC.

[0043] S21, Input high-frequency signal data: Import ultrasonic signals and perform digital processing;

[0044] S22, Bandpass filtering: The bandpass filter removes noise from the input signal;

[0045] S20 changes the base frequency of the signal and processes the data to output a voice signal.

[0046] S20 specifically includes the following steps:

[0047] S23, resample the signal to change the base frequency;

[0048] S24 utilizes an overlap-overlay algorithm to perform duration normalization on the signal, generating a speech signal. Speech duration normalization is divided into two stages: decomposition and synthesis. In the decomposition stage, frames are divided with a frame length of N and a frame interval of Sa. In the synthesis stage, the signals are synthesized using the frame interval Ss. The ratio of Sa to Ss determines the normalization factor F. A Hamming window is added to ensure that the amplitude of the overlapping region remains unchanged.

[0049] See Figure 4 S24 specifically includes the following steps:

[0050] S241, extract W data points from the input sequence and store them directly into the output sequence;

[0051] S242, take W more data points at interval Sa points from the input sequence;

[0052] S243, take the W data points taken previously as the front sequence and the W data points taken later as the back sequence; compare the consistency between the last Wov data points of the front sequence and the first Wov data points of the back sequence, record the analysis results, and denote the number of comparisons as k (k is initially 0), k = k + 1;

[0053] S244, determine if k is equal to Kmax; if not, move the analysis window (i.e., capture a window of W points) and return to S243 to continue recording and comparing; if yes, proceed to S245.

[0054] S245, select the case with the most consistent results, superimpose the Wov data before and after, and then store the remaining Ss data from the W data into the output sequence;

[0055] S246, determine whether data processing is complete. If not, return to S242 to continue the value comparison process until signal processing is complete; if yes, end.

[0056] See Figure 5 W is the window length, representing the minimum length of speech processing; Sa is the analysis delay, representing the interval between the first addresses of the speech segments that are sequentially extracted and processed; Ss is the synthesis delay, representing the interval between the first addresses of the speech segments that are sequentially output; Kmax is the lookup delay, which is the delay that must occur for the analysis window to be consistent with the end of the output signal; Wov is the length of the superposition of the previous and next speech segments.

[0057] See Figure 6 S20 specifically includes the following steps:

[0058] S33 performs synchronous detection and demodulation on the signal; generates a sine wave with the same frequency as the ultrasonic fundamental tone, multiplies it with the signal, and performs amplitude modulation demodulation through a low-pass filter to obtain a sound signal with the ultrasonic fundamental tone removed.

[0059] S34 modulates the signal; generates a low-frequency pitch signal, and modulates it with the audio signal on both sides to form the final speech signal and play it.

[0060] The carrier signal used for modulation in S34 is between 50-500Hz.

[0061] This invention addresses the noise problem caused by the oscillation of electronic artificial larynxes by proposing the use of ultrasound to replace the fundamental tone. A receiver and a spectrum shifting module are added to the electronic artificial larynx to eliminate noise. In the oscillation chamber, the oscillation frequency is increased to the ultrasonic range, and an ultrasonic generator produces the ultrasonic fundamental tone. Furthermore, this invention adds a receiver and a spectrum shifting module to the traditional electronic artificial larynx. The receiver uses an ultrasonic receiver head to receive the ultrasonic signal, which is then uploaded to a PC. Noise is removed by bandpass filtering of the ultrasonic signal, and the signal spectrum is then transformed to down-convert the ultrasonic fundamental tone in the sound signal to the range of the original sound fundamental tone, thus achieving spectrum shifting and ultimately forming an audible speech signal.

[0062] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A noise reduction method for an electronic artificial larynx based on ultrasonic pitch, characterized in that, The electronic artificial larynx used includes a power module and a transmitter / receiver connected to the power module and a PC. The transmitter / receiver is connected to the PC. The transmitting and receiving end includes an ultrasonic generator and an ultrasonic receiver; the PC end includes a bandpass filter and a spectrum conversion module; the power module supplies power to the transmitting and receiving end and the PC end. The ultrasonic generator sends ultrasonic waves to the human throat, the ultrasonic receiver receives the echo from the human throat, the transmitting and receiving end sends ultrasonic signals to the PC end, and the bandpass filter on the PC end filters the ultrasonic signals and sends them to the spectrum conversion module for noise reduction, conversion and output of voice signals. The spectrum conversion module includes a resampler and a duration warping unit. The resampler is connected to a bandpass filter to downsample the ultrasonic signal. The resampler is connected to the duration warping unit, which uses an overlap superposition algorithm to warp the downsampled signal. The noise reduction method includes the following steps: S10 uses an ultrasonic generator to generate ultrasonic waves and transmit them into the human throat. It uses an ultrasonic receiver head to receive the transmitted ultrasonic signals outside the mouth and transmit them to the PC. S21, Input high-frequency signal data: Import ultrasonic signals and perform digital processing; S22, Bandpass filtering: The bandpass filter removes noise from the input signal; S20 changes the base frequency of the signal and processes the data to output a voice signal; S20 specifically includes the following steps: S23, resample the signal to change the base frequency; S24, the signal is time-normalized using an overlap superposition algorithm to generate a speech signal; S24 specifically includes the following steps: S241, extract W data points from the input sequence and store them directly into the output sequence; S242, take W more data points at interval Sa points from the input sequence; S243, the W data points taken previously are used as the front sequence, and the W data points taken later are used as the back sequence; compare the consistency between the last Wov data points of the front sequence and the first Wov data points of the back sequence, record the analysis results, and denote the number of comparisons as k, k=k+1, with an initial value of 0 for k; S244, determine if k is equal to Kmax; if not, move the analysis window, i.e., select a window of W points, and return to S243 to continue recording and comparing; if yes, proceed to S245. S245, select the case with the most consistent results, superimpose the Wov data before and after, and then store the remaining Ss data from the W data into the output sequence; S246, determine whether data processing is complete. If not, return to S242 to continue the value comparison process until signal processing is complete; if yes, end. Where W is the window length, representing the minimum length of speech processing; Sa is the analysis delay, representing the interval between the first addresses of the speech segments that are sequentially extracted and processed; Ss is the synthesis delay, representing the interval between the first addresses of the speech segments that are sequentially output; Kmax is the lookup delay, which is the delay that must occur for the analysis window to be consistent with the end of the output signal; Wov is the length of the superposition of the previous and next speech segments, where the previous / next speech segments are the speech segments extracted first / last.

2. A noise reduction method for an electronic artificial larynx based on ultrasonic pitch, characterized in that, The electronic artificial larynx used includes a power module and a transmitter / receiver connected to the power module and a PC. The transmitter / receiver is connected to the PC. The transmitting and receiving end includes an ultrasonic generator and an ultrasonic receiver; the PC end includes a bandpass filter and a spectrum conversion module; the power module supplies power to the transmitting and receiving end and the PC end. The ultrasonic generator sends ultrasonic waves to the human throat, the ultrasonic receiver receives the echo from the human throat, the transmitting and receiving end sends ultrasonic signals to the PC end, and the bandpass filter on the PC end filters the ultrasonic signals and sends them to the spectrum conversion module for noise reduction, conversion and output of voice signals. The spectrum conversion module includes a synchronous detector demodulator and a DSB modulator. The synchronous detector demodulator is connected to a bandpass filter to remove the fundamental frequency signal from the ultrasonic signal. The synchronous detector demodulator is connected to the DSB modulator, which modulates the received signal to form a voice signal and plays it. The noise reduction method includes the following steps: S10 uses an ultrasonic generator to generate ultrasonic waves and transmit them into the human throat. It uses an ultrasonic receiver head to receive the transmitted ultrasonic signals outside the mouth and transmit them to the PC. S21, Input high-frequency signal data: Import ultrasonic signals and perform digital processing; S22, Bandpass filtering: The bandpass filter removes noise from the input signal; S20 changes the base frequency of the signal and processes the data to output a voice signal; S20 specifically includes the following steps: S33 performs synchronous detection and demodulation on the signal; generates a sine wave with the same frequency as the ultrasonic fundamental tone, multiplies it with the signal, and performs amplitude modulation demodulation through a low-pass filter to obtain a sound signal with the ultrasonic fundamental tone removed. S34 modulates the signal; generates a low-frequency pitch signal, and modulates it with the audio signal on both sides to form the final speech signal and play it. The carrier signal used for modulation in S34 is between 50-500Hz.

Citation Information

Patent Citations

  • Method and device for Doppler ultrasound pickup analysis processing

    CN103169505A

  • Audio comprehensive training aid based on digital signal processing system

    CN104606762A

  • Split type electronic throat

    CN214908659U