Sound production device and method based on ultrasonic recognition and storage medium
Through the sound generating device based on ultrasonic recognition, the ultrasonic transducer and speech recognition algorithm are used to generate a speech signal of natural tone, which solves the problems of monotonous sound effects and inconvenient use in the prior art, and achieves a more natural sound and a higher usage experience.
Patent Information
- Application Number
- CN202510211986.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing pneumatic artificial throat and handheld electronic throat have monotonous and unnatural tone in terms of sound effects, and there are risks of infection and inconvenience.
Using a sound generating device based on ultrasonic recognition, a transmitting and receiving ultrasonic transducer, an analog-to-digital converter, an encoding chip and a processing module, combined with a speech recognition algorithm, a speech signal of natural tone is generated and a sound is emitted through the speaker.
It achieves a more natural sounding effect, reduces noise interference, improves communication efficiency and user experience, and reduces the risk of infection and inconvenience of use.
Smart Images

Figure CN120164448A_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of medical devices, and particularly to a voice generating device, method, and storage medium based on ultrasonic recognition. Background Art
[0002] Laryngectomees lose their voice due to laryngeal surgery or diseases, which brings great obstacles to communication.
[0003] Existing technologies mainly use auxiliary voice generating devices, such as pneumatic artificial larynxes or handheld electronic larynxes, to help patients generate voices. Among them, pneumatic artificial larynxes need to be directly in contact with the lesion opening, and some even need to be implanted to help patients generate voices; handheld electronic larynxes need to be continuously held by users and pressed against the larynx to help patients generate voices.
[0004] In terms of usage effects, the existing pneumatic artificial larynxes in direct contact with the lesion opening will increase the risk of infection, and due to the harsh usage environment, the implanted components need to be frequently replaced; while handheld electronic larynxes bring inconvenience to patients' lives. In terms of voice generation effects, both existing pneumatic artificial larynxes and handheld electronic larynxes use external sound sources to replace the vocal cords, and then cooperate with the movement of the muscles in the user's oral cavity to modulate the sound to form speech. This voice generation method makes the timbre monotonous and unnatural, and there will also be sound leakage, affecting the user experience. Summary of the Invention
[0005] In view of the above solutions, the present application aims to propose a voice generating device, method, and storage medium based on ultrasonic recognition to solve at least one of the above technical problems.
[0006] In a first aspect, one or more embodiments of this specification provide a voice generating device based on ultrasonic recognition, including:
[0007] A transmitting ultrasonic transducer for emitting ultrasonic signals; the transmitting ultrasonic transducer is installed in the user's oral cavity or nasal cavity, or installed in the user's neck or jaw;
[0008] A receiving ultrasonic transducer for collecting ultrasonic signals;
[0009] An analog-to-digital converter for converting the ultrasonic signals into digital signals;
[0010] An encoding chip for encoding the digital signals to obtain encoded signals;
[0011] A processing module for performing voice recognition on the encoded signals based on a voice recognition algorithm to obtain synthesized voice signals; and
[0012] A voice emitting module for driving a speaker to emit sounds based on the synthesized voice signals.
[0013] Further, it further includes:
[0014] A power-on control module, configured to collect sound signals in real time by using a microphone;
[0015] Convert the sound signals into electrical signals;
[0016] Based on a speech recognition algorithm, determine whether the electrical signals are preset sound signals;
[0017] If so, trigger a power-on signal to turn on the transmit ultrasonic transducer, the receive ultrasonic transducer, the analog-to-digital converter, and the encoding chip.
[0018] Further, it further includes:
[0019] A frequency band selection module, configured to randomly select an ultrasonic frequency band for the transmit ultrasonic transducer after power-on.
[0020] Further, the voice emitting module is configured to:
[0021] Decode the synthesized voice signals by using a decoding chip; and
[0022] Drive a speaker to emit the voice signals by using a drive circuit.
[0023] Further, it further includes:
[0024] A sound adjustment module, configured to adjust the encoded signals based on a pitch-shifting without changing speed algorithm to obtain an audible frequency band.
[0025] In a second aspect, an embodiment of the present application provides a voice emitting method based on ultrasonic recognition, including:
[0026] Emit ultrasonic signals by using a transmit ultrasonic transducer, where the transmit ultrasonic transducer is installed in the oral cavity or nasal cavity of a user, or installed on the neck or jaw of the user;
[0027] Collect the ultrasonic signals by using a receive ultrasonic transducer;
[0028] Convert the ultrasonic signals into digital signals by using an analog-to-digital converter;
[0029] Encode the digital signals based on an encoding chip to obtain encoded signals;
[0030] Perform speech recognition on the encoded signals based on a speech recognition algorithm to obtain synthesized voice signals; and
[0031] Drive a speaker to emit sounds based on the synthesized voice signals.
[0032] Further, before the ultrasonic signal is emitted by the emission ultrasonic transducer, the method includes:
[0033] Using a microphone to collect sound signals in real time;
[0034] Converting the sound signals into electrical signals;
[0035] Based on a speech recognition algorithm, determining whether the electrical signals are preset sound signals; and
[0036] If so, triggering a power-on signal to turn on the emission ultrasonic transducer, the reception ultrasonic transducer, the analog-to-digital converter, and the encoding chip.
[0037] Further, it further includes:
[0038] After power-on, randomly selecting an ultrasonic frequency band for the emission ultrasonic transducer.
[0039] Further, after the synthesized speech signal is obtained, the method further includes:
[0040] Using a decoding chip to decode the synthesized speech signal; and
[0041] Using a drive circuit to drive a speaker to emit the speech signal.
[0042] In a third aspect, an embodiment of the present application provides a storage medium for storing computer-executable instructions, characterized in that the computer-executable instructions, when executed, implement the steps of the ultrasonic recognition and voice generation method described in any one of the second aspects.
[0043] Compared with the prior art, the present application can at least achieve the following technical effects:
[0044] The present application can use an emission ultrasonic transducer to help a user emit sounds that are inaudible to the human ear, and can reduce noise interference; then use a reception transducer, an analog-to-digital converter, an encoding chip, and a processing module to process the inaudible sounds, and synthesize and customize a sound that meets the user's needs, making the timbre more natural; finally, use a speaker to emit the sound, thereby improving the communication efficiency of the user and enhancing the user experience. Description of the Drawings
[0045] To more clearly illustrate the technical solutions in one or more embodiments of the present specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present specification. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0046] Figure 1 Schematic diagram of a sound - generating device based on ultrasonic recognition provided for one or more embodiments of this specification;
[0047] Figure 2 Flowchart of a sound - generating method based on ultrasonic recognition provided for one or more embodiments of this specification;
[0048] Figure 3 Schematic diagram of an artificial larynx based on ultrasonic and voice recognition provided for one or more embodiments of this specification;
[0049] Figure 4 Schematic diagram for realizing remote communication in a silent or reduced - ambient - noise situation provided for one or more embodiments of this specification. Detailed implementation manners
[0050] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the following will clearly and completely describe the technical solutions in one or more embodiments of this specification in conjunction with the accompanying drawings in one or more embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this document.
[0051] For laryngectomees, the existing technologies mainly use auxiliary sound - generating devices for sound production, which are divided into pneumatic artificial larynxes and handheld artificial larynxes. Both of these devices have deficiencies in terms of comfort. The pneumatic artificial larynx needs to be in direct contact with the stoma, and some even need to be implanted, which not only increases the risk of infection, but also due to the harsh usage environment, the components need to be frequently replaced. The electronic larynx requires the user to continuously hold it and press it against the larynx. In terms of the sound - generating effect, both are not satisfactory. The timbre is monotonous and unnatural, lacking pitch variation, and the electronic larynx also has the problem of electronic sound leakage, thus affecting the usage experience of laryngectomees.
[0052] Embodiment 1
[0053] To address the above - mentioned technical problems, this application proposes a sound - generating device based on ultrasonic recognition, as Figure 1 shown, specifically including:
[0054] In an embodiment of the present application, a transmitting ultrasonic transducer 101 is configured to emit ultrasonic signals. The transmitting ultrasonic transducer is installed inside the oral cavity or nasal cavity of a user, or outside the user's body. The generated ultrasonic waves pass through the skin and muscles of the neck or jaw and other parts and are transmitted into the oral cavity. A receiving ultrasonic transducer 102 is configured to collect ultrasonic signals. An analog-to-digital converter 103 is configured to convert the ultrasonic signals into digital signals. An encoding chip 104 is configured to encode the digital signals to obtain encoded signals. A processing module 105 is configured to perform speech recognition on the encoded signals based on a speech recognition algorithm to obtain synthesized speech signals. A voice output module 106 is configured to drive a speaker to emit sounds based on the synthesized speech signals.
[0055] Specifically, first, an ultrasonic signal inaudible to the human ear is emitted by the transmitting ultrasonic transducer and introduced into the oral cavity. Under the joint modulation of multiple organs such as the internal muscles of the oral cavity and the tongue, an ultrasonic frequency band speech signal with different wave peaks and exceeding 20 kHz is generated. Among them, the transmitting ultrasonic transducer is installed inside the oral cavity or nasal cavity, or the transmitting ultrasonic transducer can be held by hand, and the ultrasonic transducer can be fixed to the larynx through a neckband. The ultrasonic waves will pass through the skin and muscles of the neck or jaw and be transmitted into the oral cavity. Second, the receiving ultrasonic transducer is used to collect the ultrasonic speech signal in this frequency band after being jointly modulated by multiple organs such as the internal muscles of the oral cavity and the tongue. The collected ultrasonic speech signal is the original analog signal. Among them, the receiving ultrasonic transducer is integrated on the sound generating device, such as being installed together with a speaker or a microphone. Then, after the analog signal is amplified and noise-reduced by an analog front end (AFE), the analog-to-digital converter (ADC) in the analog front end converts the analog signal into a digital signal. Next, various encoding techniques (such as pulse code modulation, differential pulse code modulation, etc.) in the encoding chip are used to encode the digital signal to obtain an encoded signal. Encoding the digital signal by the encoding chip can improve the anti-interference ability and transmission efficiency of the signal, making it easier for subsequent signal processing and decoding. Then, a speech recognition algorithm (such as Gaussian mixture model GMM, deep learning algorithm, Transformer and Conformer architectures, etc.) is used to analyze the encoded signal to identify the speech content of the speaker and obtain a synthesized speech signal. Finally, the synthesized speech signal is transmitted to the speaker to drive the speaker to emit the words the user wants to say.
[0056] In an embodiment of the present application, a power-on control module is configured to: collect a sound signal in real time using a microphone; convert the sound signal into an electrical signal; determine, based on a speech recognition algorithm, whether the electrical signal is a preset sound signal; if so, trigger a power-on signal to turn on the transmit ultrasonic transducer, the receive ultrasonic transducer, the analog-to-digital converter, and the encoding chip.
[0057] Specifically, in order to extend the battery life and service life of the device, three working states are set for the sound-emitting device: off, standby, and use. In the standby state, the transmit ultrasonic transducer, the receive ultrasonic transducer, the analog-to-digital converter, and the encoding chip do not work, and the microphone collects sound signals in real time. Once it is detected that the trigger signal sent by the user is the same as the preset sound signal, all devices immediately switch to the use state, and each module starts to operate normally. For example, the trigger signal is set in advance as an impact signal generated by the collision of the upper and lower teeth. If this signal is monitored, it indicates that the user needs to speak at this time and needs to start receiving ultrasonic signals; or the user tells the sound-emitting device that the user needs to speak at this time through a handheld switch.
[0058] In an embodiment of the present application, a frequency band selection module is configured to randomly select an ultrasonic frequency band for the transmit ultrasonic transducer after power-on.
[0059] Specifically, first, the frequency division multiplexing technology is adopted to randomly select a channel in advance. For example, the range of 20 kHz - 23 kHz is divided into 100 frequency bands, and a channel is randomly selected each time the device is powered on, so as to reduce the probability of channel repetition. Second, the frequency hopping technology (Frequency-Hopping Spread Spectrum, FHSS) can also be adopted. When it is detected that the signal-to-noise ratio is low, or a strong ultrasonic signal with the same frequency as the frequency currently selected by the user is detected in the standby state, the system will actively switch the frequency to avoid signal conflicts.
[0060] In the present application, when there are multiple users using the sound-emitting device, the problems of signal crosstalk can be avoided through the frequency division multiplexing technology and the frequency hopping technology, thus ensuring the stability and reliability of the product in a multi-user environment.
[0061] In an embodiment of the present application, a voice emission module is configured to: decode the synthesized voice signal using a decoding chip; drive a speaker to emit the voice signal using a driving circuit.
[0062] Specifically, the synthesized voice signal is input into the decoding chip. At this time, the signal is usually stored or transmitted in digital form. During the decoding process, these digital signals are converted back into analog signals to ensure that the converted analog signal is as close as possible to the user's own voice. Finally, an analog voice signal is output. Then, a driving circuit is used to amplify the analog voice signal output by the decoding chip, and the driving circuit is used to ensure that the output impedance of the decoding chip matches the input impedance of the speaker. Once the signal is amplified and impedance-matched, the driving circuit will send these signals to the speaker to drive the speaker to vibrate and emit sound.
[0063] In this application, impedance matching in the driving circuit can reduce the reflection of signals on the transmission line, thereby ensuring that more signal energy can be transmitted to the speaker, achieving the effect of improving the sound quality.
[0064] In the embodiment of this application, a sound adjustment module is used to adjust the encoded signal based on the pitch-shifting without changing the speed algorithm to obtain an audible frequency band.
[0065] Specifically, the pitch-shifting without changing the speed algorithm includes the time-domain method, the frequency-domain method, and the parametric method. Among them, the time-domain method realizes pitch-shifting without changing the speed by operating on the voice signal in the time domain. The usage methods include the cut-and-paste method, the synchronized overlap-add (SOLA) method, the waveform similarity overlap-and-add (WSOLA) method, the time-domain pitch synchronized overlap-add (TD-PSOLA) method, etc. For example: The TD-PSOLA algorithm performs frame-by-frame processing on the voice signal in the time domain and divides the voice units by detecting the pitch period. By repeating or omitting these voice units, the control of the speech rate can be achieved. At the same time, by adjusting the overlapping length or resampling rate of adjacent voice units, the fundamental frequency of the voice can be changed, thereby achieving a decrease in pitch. The specific implementation method is as follows: First, preprocess the encoded signal, such as denoising and filtering; secondly, use the TD-PSOLA algorithm to perform frame-by-frame and pitch period detection on the signal; then adjust the overlapping length or resampling rate of the voice units according to the target pitch; finally, synthesize the adjusted voice signal and perform post-processing to improve the sound quality, and finally obtain a signal in the audible frequency band.
[0066] The frequency domain method realizes pitch change without speed change by operating on the spectrum of the speech signal. The methods used include the LSEE-MSTFTM algorithm and the spectrum interpolation / decimation method. Among them, the LSEE-MSTFTM algorithm is based on the short-time Fourier transform and uses the least mean square error principle to find the short-time Fourier amplitude spectrum of a time-domain signal that approximates the spectrum of the ideal signal, thereby realizing pitch change without speed change. The spectrum interpolation / decimation method directly interpolates or decimates the signal spectrum to expand or compress each frequency component. For example: perform a short-time Fourier transform on the encoded signal to obtain a spectrum representation; then scale the spectrum in the frequency domain to reduce the frequency of the signal; then perform an inverse short-time Fourier transform (ISTFT) on the scaled spectrum to obtain a frequency-reduced signal in the time domain, that is, a signal in the audible frequency band.
[0067] The parametric method realizes pitch change without speed change by establishing a model for the speech signal and modifying the model parameters. The methods used include the phase vocoder, the sine model, etc. For example: extract the parameters of its sine model by performing parameter extraction on the encoded signal; then adjust the frequency parameters of the model according to the target pitch; finally, resynthesize the speech signal using the adjusted parameters to obtain a signal in the audible frequency band.
[0068] Embodiment 2
[0069] The embodiment of the present application provides a voice generation method based on ultrasonic recognition, as Figure 2 shown, including the following steps:
[0070] Step S1, use a transmitting ultrasonic transducer to emit an ultrasonic signal. The transmitting ultrasonic transducer is installed in the user's oral cavity or nasal cavity, or installed outside the user's body. The generated ultrasonic waves pass through the skin and muscles in places such as the neck or jaw and are transmitted into the oral cavity.
[0071] In the embodiment of the present application, by emitting an ultrasonic signal through a transmitting ultrasonic transducer, then introducing the ultrasonic sound source into the oral cavity, and then modulating the ultrasonic signal formed by the modulation of the internal muscles and tongue in the oral cavity is very close to the speech signal, and thus can carry the most useful information.
[0072] And there are three ways to introduce the ultrasonic sound source into the oral cavity:
[0073] Method 1, fix components such as a transmitting ultrasonic transducer, a driving circuit, and a battery on a dental appliance. When the user wears the dental appliance, the sound source can be directly introduced into the oral cavity. The transmitting ultrasonic transducer can be an electromagnetic voice coil motor (VCM) or a piezoelectric ceramic-based ultrasonic probe. A rechargeable battery is used for power supply.
[0074] In Method 2, the outside of the transmitting ultrasonic transducer is wrapped with elastic silicone material. Its size and structure are similar to those of in-ear Bluetooth headsets and can be inserted into the nostrils to transmit ultrasonic waves into the oral cavity through the nasal cavity. The transmitting ultrasonic transducer can be an electromagnetic voice coil motor (VCM) or an ultrasonic probe based on piezoelectric ceramics. The power supply uses a rechargeable battery.
[0075] In Method 3, the transmitting ultrasonic transducer is directly placed in contact with the lower jaw to transmit ultrasonic waves into the oral cavity through the skin. The usage method can be handheld or, through a neckband, the transmitting ultrasonic transducer can be bound to the throat. The transmitting ultrasonic transducer can be an electromagnetic voice coil motor (VCM) or an ultrasonic probe based on piezoelectric ceramics.
[0076] Step S2: Use the receiving ultrasonic transducer to collect the ultrasonic signal.
[0077] In the embodiment of the present application, the ultrasonic transducer for collecting ultrasonic waves is placed near the mouth, and the collected ultrasonic voice signal is initially an analog signal.
[0078] Step S3: Use an analog-to-digital converter to convert the ultrasonic signal into a digital signal.
[0079] In the embodiment of the present application, after the analog signal is amplified and noise-reduced in the analog front end (AFE), it is converted into a digital signal by an analog-to-digital converter.
[0080] Step S4: Based on the encoding chip, encode the digital signal to obtain an encoded signal.
[0081] In the embodiment of the present application, first, preprocess the original digital signal, including signal amplification, filtering, noise reduction, etc., to improve the efficiency and accuracy of encoding; then encode the preprocessed digital signal according to the pre-designed encoding rules to obtain an encoded signal.
[0082] Step S5: Based on the speech recognition algorithm, perform speech recognition on the encoded signal to obtain a synthesized speech signal.
[0083] In the embodiments of the present application, the speech recognition algorithm in the edge-side hardware acquisition / processing unit is used to perform speech recognition on the encoded signal. The edge-side hardware acquisition / processing unit has two working modes: Mode 1, in the network-connected state, the encoded signal is uploaded to the cloud. The cloud server completes speech recognition and synthesis, and then transmits the synthesized speech signal back to the local. The advantage of this method lies in the powerful computing power of the server, which can provide a higher recognition accuracy and more natural synthesized speech. And in the network-connected state, it can provide a larger model and stronger computing power, thereby improving the operation speed and enhancing the accuracy of speech recognition. Among them, computing power refers to the number of operations that a CPU or GPU can perform per second. The stronger the computing power, the faster the operation speed. In the case of the same model, the faster the recognition speed. In the same amount of time, a larger model can be run, and the recognition accuracy is higher. For example: in the case of having a network, use large models such as deep learning or Transformer / Conformer for speech recognition, as Figure 3 shown.
[0084] Mode 2, in the non-network-connected state, both speech recognition and synthesis are completed locally. For example: perform speech recognition by deploying a Gaussian mixture model (GMM). The speech recognition model deployed in the non-network state can be carried with the user and can be used in areas without network such as tunnels, underground or remote areas, providing users with an instant and efficient usage experience.
[0085] The output result of speech recognition is text. Subsequently, the text is synthesized into speech through the TTS (Text-to-Speech) algorithm, and models such as F5-TTS and CosyVoice can be used for voice cloning, so as to complete the customization of the synthesized speech tone. For example, using the user's voice before undergoing a laryngectomy to make the user's voice before and after the operation consistent, providing a good experience for the user.
[0086] Step S6, based on the synthesized speech signal, drive the speaker to emit sound.
[0087] In the embodiments of the present application, after receiving the synthesized speech signal, the signal will pass through a decoding chip and a driving circuit, and finally drive the speaker to emit sound.
[0088] Further, in the embodiments of the present application, a microphone is used to collect sound signals in real time; the sound signals are converted into electrical signals; based on the speech recognition algorithm, it is determined whether the electrical signals are preset sound signals; if so, a power-on signal is triggered to turn on the transmit ultrasonic transducer, the receive ultrasonic transducer, the analog-to-digital converter, and the encoding chip.
[0089] Specifically, taking the microphone as the sound acquisition device, it can capture the sounds in the surrounding environment in real time. There is a vibrating diaphragm inside the microphone. When sound waves act on the vibrating diaphragm, it will cause the diaphragm to vibrate. This vibration is then converted into an analog electrical signal. This analog signal represents the waveform and intensity of the sound. Then, the collected analog electrical signal is converted into a digital signal and processed through a speech recognition algorithm. The speech recognition algorithm will analyze the characteristics of the digital signal, such as frequency, amplitude, and waveform pattern, to determine whether these signals match the preset sound signals (such as waking up the mobile phone through specific voice commands). If the algorithm recognizes an electrical signal that matches the preset sound signal, the system will send a power-on signal. This signal will be used to start a series of devices, including but not limited to a transmitting ultrasonic transducer, a receiving ultrasonic transducer, an analog-to-digital converter, and a coding chip. When not in use, a voice signal can also be set to turn off a series of devices.
[0090] Triggering device startup through sound recognition not only improves the convenience of the system but also increases the flexibility of user interaction.
[0091] Further, in the embodiment of the present application, after power-on, an ultrasonic frequency band is randomly selected for the transmitting ultrasonic transducer.
[0092] Specifically, Frequency Division Multiplexing (FDM) is a multiplexing method that divides the channel according to frequency. In an FDM system, the bandwidth of the channel is divided into multiple non-overlapping sub-channel frequency bands. Each signal occupies one of the sub-channels, and there are unused frequency bands left between each channel as guard bands to prevent signal overlap and interference. In the actual use process, there will be multiple laryngectomee users communicating. If the same ultrasonic frequency band is used, there may be a problem of signal crosstalk. Therefore, the FDM technology is used to randomly select an ultrasonic frequency band for the user and monitor the working state of the transmitting ultrasonic transducer in real time to ensure that it can work stably within the selected frequency band. If signal interference or performance degradation is detected, the Frequency Hopping Spread Spectrum (FHSS) technology is used to re-select the ultrasonic frequency band.
[0093] The FDM technology and the FHSS technology can improve the flexibility and reliability of the system and reduce the risk of signal interference at the same time.
[0094] Further, in the embodiment of the present application, a decoding chip is used to decode the synthesized voice signal; and a driving circuit is used to drive the speaker to emit the voice signal.
[0095] Specifically, after the sounding device receives the synthesized voice signal, the signal will pass through the decoding chip and the driving circuit, and finally drive the speaker to emit sound.
[0096] This application can help laryngectomy patients vocalize in a safer, healthier and infection-free manner, and can customize a more natural voice according to the user's own needs, thereby improving the user's experience and making communication smoother.
[0097] Example 3
[0098] This solution can also serve the healthy population, and the application scenarios include but are not limited to the following situations:
[0099] Situation 1: Making a call when the surrounding environment needs to be kept quiet, such as answering a call in a library or participating in a remote meeting. In this case, the user's need is to convey voice information or text information to the other party without making a sound. Applying the present invention, the user only needs to make mouth shapes without using the vocal cords to make a sound, and then ultrasonic frequency band speech recognition can be achieved, and then the recognized text information or the voice information with a natural tone synthesized according to the text can be transmitted to the other party. Since there is no audible sound frequency band sound emitted, it will not affect others on the scene.
[0100] Situation 2: The surrounding environment has a lot of noise, such as in industrial sites, construction sites, high-noise environments such as airplanes and high-speed rails. At this time, when communicating with the outside world, it is necessary to shield the noise and clearly convey the content spoken to the other party. Because the carrier wave used in the present invention is in the ultrasonic frequency band, and the noise in nature is concentrated in the audible sound frequency band, the influence on the ultrasonic wave receiving end is very small, and the noise can be effectively shielded, and only the ultrasonic signal modulated by the muscles in the user's oral cavity is picked up. And perform speech recognition on this ultrasonic signal, and then transmit the restored voice signal to the other party.
[0101] In the embodiment of this application, for silent communication of consumer electronics and communication in a high-noise environment, the structure of the system is slightly different, as Figure 4 shown. Since there is no need to play the synthesized sound locally, there is no need for a speaker. The recognized voice signal will be in the form of text or further synthesized into a voice signal with a natural tone, and transmitted to a remote computer or mobile phone through the network, and played by the remote device to achieve silent communication or noise reduction communication.
[0102] This application can provide a smooth and clear communication device for users in a noisy environment, thereby improving the communication efficiency of users, improving work efficiency, and also making users no longer need to shout loudly or frequently ask the other party to repeat, enhancing the user experience.
[0103] The embodiment of this application provides a storage medium for storing computer-executable instructions, characterized in that the computer-executable instructions, when executed, implement the steps of the ultrasonic recognition vocalization method described in any one of the above embodiments.
[0104] It should be noted that the embodiments of the storage medium in this specification and the embodiments of the blockchain-based service providing method in this specification are based on the same inventive concept. Therefore, for the specific implementation of this embodiment, reference may be made to the corresponding implementation of the blockchain-based service providing method described above, and repeated parts will not be elaborated.
[0105] The specific embodiments of this specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0106] In the 1930s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (e.g., a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). There is not just one type of HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0107] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.
[0108] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0109] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing the embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0110] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0111] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general purpose computers, special purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0112] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0113] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0114] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0115] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0116] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0117] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising said element.
[0118] One or more embodiments of this specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification may also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media including storage devices.
[0119] Each embodiment in this specification is described in a progressive manner, and the same or similar parts among the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments.
[0120] The above are only examples of this document and are not intended to limit this document. For those skilled in the art, this document may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this document shall be included within the scope of the claims of this document.
Claims
1. A sound-generating device based on ultrasonic recognition, characterized in that include: A transmitting ultrasonic transducer, used for emitting ultrasonic signals; the transmitting ultrasonic transducer is installed in the oral cavity or nasal cavity of the user, or installed on the neck or jaw of the user; A receiving ultrasonic transducer, used for collecting ultrasonic signals; An analog-to-digital converter, used for converting the ultrasonic signal into a digital signal; A coding chip, used for encoding the digital signal to obtain a coded signal; A processing module, configured to perform speech recognition on the coded signal based on a speech recognition algorithm to obtain a synthesized speech signal; as well as The voice emitting module is used to drive the speaker to emit sound based on the synthesized voice signal.
2. The device according to claim 1, characterized in that Also includes: A power-on control module, used to collect sound signals in real time using a microphone; converting the sound signal into an electrical signal; Based on a speech recognition algorithm, determining whether the electrical signal is a preset sound signal; If so, a power-on signal is triggered to turn on the transmitting ultrasonic transducer, the receiving ultrasonic transducer, the analog-to-digital converter and the encoding chip.
3. The device according to claim 2, characterized in that Also includes: The frequency band selection module is used to randomly select an ultrasonic frequency band for the transmitting ultrasonic transducer after powering on.
4. The device according to claim 1, characterized in that The voice issuing module is configured as follows: Decoding the synthesized speech signal using a decoding chip; and The driving circuit is used to drive the speaker to emit the voice signal.
5. The device according to claim 1, characterized in that Also includes: The sound adjustment module is used to adjust the encoded signal based on a pitch-changing but speed-invariant algorithm to obtain an audible sound frequency band.
6. A method for producing sound based on ultrasonic recognition, characterized in that include: An ultrasonic signal is emitted by a transmitting ultrasonic transducer, wherein the transmitting ultrasonic transducer is installed in the oral cavity or nasal cavity of the user, or on the neck or jaw of the user; collecting the ultrasonic signal using a receiving ultrasonic transducer; Converting the ultrasonic signal into a digital signal using an analog-to-digital converter; Based on the encoding chip, the digital signal is encoded to obtain an encoded signal; Based on a speech recognition algorithm, performing speech recognition on the coded signal to obtain a synthesized speech signal; as well as Based on the synthesized voice signal, a speaker is driven to produce sound.
7. The method according to claim 6, characterized in that Before using the transmitting ultrasonic transducer to emit an ultrasonic signal, the method comprises: Use microphone to collect sound signals in real time; converting the sound signal into an electrical signal; Based on a speech recognition algorithm, determining whether the electrical signal is a preset sound signal; and If so, a power-on signal is triggered to turn on the transmitting ultrasonic transducer, the receiving ultrasonic transducer, the analog-to-digital converter and the encoding chip.
8. The method according to claim 7, characterized in that The method further comprises: After powering on, an ultrasonic frequency band is randomly selected for the transmitting ultrasonic transducer.
9. The method according to claim 6, characterized in that After obtaining the synthesized speech signal, the method further comprises: Decoding the synthesized speech signal using a decoding chip; and The driving circuit is used to drive the speaker to emit the voice signal.
10. A storage medium for storing computer executable instructions, characterized in that: When the computer executable instructions are executed, the steps of the method for utterance of ultrasonic recognition according to any one of claims 6 to 9 are implemented.
Citation Information
Patent Citations
Wireless headset based on supersonic wave, and control system of the same
CN107396228A
Silent communication method, device and equipment and storage medium
CN111899713A
Voice synthesis method and device based on timbre cloning and related equipment
CN113160794A
Speech enhancement method and system fusing ultrasonic signal features
CN114067824A
Identity verification method and system based on machine learning
CN115424608A
Cited By
Voice generation method and system based on throat vibration signal analysis
CN120544536A