Voice output device
The audio output device generates pseudo-voice by analyzing lip-syncing vibrations, addressing the limitations of existing devices by enhancing accuracy through machine learning and diverse data inputs, enabling voice output for those without laryngeal sounds.
Patent Information
- Application Number
- JP2024071946
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-25
- Publication Date
- 2025-11-07
AI Technical Summary
Existing voice output devices cannot generate pseudo-voices for individuals who cannot produce laryngeal primary sounds, requiring them to produce inaudible sounds or rely on skin surface electrodes.
An audio output device that generates pseudo-voice by analyzing vibrations caused by lip-syncing, utilizing a vibration measurement unit, audio information generation unit, and audio output unit, optionally incorporating audio measurement, muscle signal measurement, and imaging units to enhance accuracy.
Enables pseudo-voice generation for those unable to produce laryngeal sounds, improving accuracy through machine learning and incorporating multiple data sources like vibrations, audio, and muscle signals, and visual cues.
Smart Images

Figure 2025167394000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a voice output device, and more particularly to a voice output device that can generate pseudo-voice even for a person who has had their vocal cords removed and is unable to generate laryngeal primary sounds. [Background technology]
[0002] Japanese Patent Application Laid-Open Publication No. 202-0124444 describes a speech aid that amplifies sounds input to the nasal cavity. This device cannot generate pseudo-voices for people who cannot produce laryngeal primary sounds. Furthermore, in order to generate pseudo-voices with this device, the person must forcefully produce inaudible sounds.
[0003] Japanese Patent Publication No. 3455921 describes a voice output device that uses machine learning. This device uses multiple skin surface electrodes to detect skin movement. Therefore, this device cannot output voice without skin surface electrodes. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 202-0124444 [Patent Document 2] Patent No. 3455921 Summary of the Invention [Problem to be solved by the invention]
[0005] This specification provides a voice output device that can generate pseudo-voice even for a person who cannot generate laryngeal original sounds. [Means for solving the problem]
[0006] This invention is based on the finding that pseudo-voice can be generated by analyzing vibrations caused by lip-syncing, without forcing the subject to produce voice.
[0007] The present invention relates to an audio output device 1. The audio output device 1 includes a vibration measuring unit 3, an audio information generating unit 5, and an audio output unit . The vibration measurement unit 3 is an element for measuring vibrations originating from a living organism. An example of vibrations originating from a living organism is vibrations generated when a living organism tries to speak, which are air vibrations that cannot be heard as sound by anyone other than the living organism. This vibration may be vibrations generated during so-called lip-syncing. The audio information generation unit 5 is an element for generating audio information based on audio source information. The audio source information includes vibration data, which is data related to vibrations measured by the vibration measurement unit 3. The audio information generation unit 5 has a machine learning unit 9, and obtains audio information by inputting the audio source information into a trained model of the machine learning unit 9. The audio output unit 7 is an element for outputting audio based on the audio information generated by the audio information generation unit 5.
[0008] One aspect of the audio output device 1 further includes an audio measurement unit 11 for measuring audio. The audio source information used by the audio information generation unit 5 to generate audio information further includes audio data measured by the audio measurement unit 11 in addition to the vibration data.
[0009] In one embodiment of the audio output device 1, the audio source information further includes a muscle signal measurement unit 13 that measures muscle signals derived from the movement of muscles of a living organism.
[0010] In one aspect of the audio output device 1, the audio output device 1 further includes an imaging unit 15 that captures an image of an area including the mouth of the living body and obtains moving image data of the living body. In this aspect, the audio source information further includes moving image data of the living body. [Effects of the Invention]
[0011] This specification can provide a voice output device that can generate pseudo-voice even for those who cannot generate laryngeal original sounds. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a block diagram illustrating an example of the basic configuration of an audio output device. DETAILED DESCRIPTION OF THE INVENTION
[0013] The following describes embodiments of the present invention with reference to the drawings. The present invention is not limited to the embodiments described below, and also includes appropriate modifications of the embodiments below within the scope obvious to those skilled in the art.
[0014] Fig. 1 is a block diagram illustrating an example of the basic configuration of an audio output device. As shown in Fig. 1, this audio output device 1 has a vibration measurement unit 3, an audio information generation unit 5, and an audio output unit 7. The audio output device 1 may further have one or more of an audio measurement unit 11, a muscle signal measurement unit 13, and an imaging unit 15.
[0015] Audio output device 1 The audio output device 1 is a device for outputting pseudo-voice without actually speaking. For example, even a subject (patient) who has had their vocal cords removed and lost their voice can use this audio output device 1 to output pseudo-voice. Even healthy individuals can suffer from adverse effects on their vocal cords if they continue to lecture or speak for long periods of time. In this case, if the subject uses the audio output device 1, pseudo-voice can be output without continuously speaking. The pseudo-voice output in this manner is usually air vibration based on digital information. The audio output device 1 may include a computer. The audio output device 1 may also be an application for a mobile device. Examples of mobile devices include mobile phones, smartphones, wristwatch-type devices (digital watches), earphone-type devices, headsets, helmet-type devices, mask-type devices, neck-sticker-type devices, Magneban-type devices, VR goggles, and eyeglass-type devices.
[0016] A computer has an input unit, an output unit, a control unit, a calculation unit, and a memory unit, and each element is connected by a bus or the like to enable the exchange of information. For example, the memory unit may store a program or various information. When predetermined information is input from the input unit, the control unit reads the program stored in the memory unit. The control unit then reads the information stored in the memory unit as appropriate and transmits it to the calculation unit. The control unit also transmits the input information to the calculation unit as appropriate. The calculation unit performs calculation processing using the various received information and stores it in the memory unit. The control unit reads the calculation results stored in the memory unit and outputs them from the output unit. In this way, various processes and steps are executed. Each unit or means executes these various processes. A computer may have a processor, and the processor may realize various functions and steps. A computer may be standalone. A computer may have some of its functions distributed between a server and a terminal. In this case, it is preferable that the server and the terminal can exchange information via a network such as the Internet or an intranet. The computer may include a processor and a memory coupled to the processor. The memory may store instructions that, when executed by the processor, cause the computer to perform various processes or function as various elements. The computer may be provided with various training data to construct a learning model and perform various calculations through machine learning. In this case, the computer may perform various analyses using a learning model created through machine learning and deep learning in AI (artificial intelligence). This improves the accuracy of machine learning.
[0017] The audio output device 1 has a vibration measurement unit 3, an audio information generation unit 5, and an audio output unit 7. The audio output device 1 may further have one or more of an audio measurement unit 11, a muscle signal measurement unit 13, and an imaging unit 15. Each element will be described below.
[0018] The vibration measurement unit 3 is an element for measuring vibrations originating from a living organism. An example of a vibration originating from a living organism is vibrations generated when a living organism attempts to speak. An example of a vibration originating from a living organism is air vibrations that cannot be heard as sound by those other than the living organism. However, this vibration may also be vibrations originating from a living organism. This vibration may also be vibrations generated during lip-syncing. In other words, when a living organism attempts to speak, it may be when the living organism moves its lips (or a part of the organism) as if to make a sound, even though it does not actually make a sound. Examples of the vibration measurement unit 3 are a microphone and a vibration meter. When the audio output device 1 is a mobile terminal, for example, the vibration measurement unit 3 can measure air vibrations sensed by the vibration measurement unit 3. When the audio output device 1 is a sticker-type device or a Magnebang-type device attached to the neck, it can measure neck vibrations and output pseudo-sound based on the neck vibrations. The vibration measurement unit 3 may be capable of measuring air vibrations and biological vibrations. Such vibration measurement units 3 are well known. The vibration measurement unit 3 converts the measured vibration information into digital data (vibration data) that can be handled by a computer. The vibration measuring unit 3 may then store the vibration data in a storage unit as appropriate. The vibration measuring unit 3 also transmits the vibration data to the voice information generating unit 5. In this way, the digitized information is transmitted to the voice information generating unit 5 in a processable form.
[0019] The audio information generation unit 5 is an element for generating audio information based on audio source information. The audio source information includes vibration data, which is data related to vibrations measured by the vibration measurement unit 3. The audio information generation unit 5 has a machine learning unit 9, and obtains audio information by inputting the audio source information into a trained model of the machine learning unit 9.
[0020] A trained model can be created by inputting multiple training data and training data into a training model (neural network) and performing machine learning. The accuracy of the trained model can also be improved by inputting certain training data into the trained model and feeding back information on the correctness of the obtained answers. In this way, the accuracy of the trained model can be improved.
[0021] The vibration data, which is data related to vibrations measured by the vibration measurement unit 3, relates to vibrations when the subject lip-syncs, for example. Therefore, there is a correlation between the vibration data and the original voice. For this reason, vibration data observed when the subject actually speaks and the voice related to that vibration data are input as training data into a learning model (neural network) and machine learning is performed, thereby constructing a trained model. In addition, the accuracy of the trained model can be improved by collecting the voice when the subject actually speaks using a microphone or the like, inputting the measured vibration data into the trained model, obtaining a pseudo-voice, and comparing the collected voice with the pseudo-voice, and feeding back the comparison results.
[0022] As described above, the trained model is obtained by performing machine learning using the voice source information as training data or teacher data. Therefore, the voice output device 1 obtains the voice source information and inputs the voice source information to the trained model of the machine learning unit 9, thereby obtaining voice information. This voice information is digital information that is output as voice via the voice output unit 7.
[0023] One aspect of the audio output device 1 further includes an audio measurement unit 11 for measuring audio. The audio source information used by the audio information generation unit 5 to generate audio information includes not only vibration data but also audio data measured by the audio measurement unit 11. For example, the audio measurement unit 11 is a sound collector such as a microphone. When a person who has had their vocal cords removed speaks or when lip-syncing, normal audio is not output. On the other hand, even in these cases, audio correlated with normal audio output is still output. This aspect measures such audio, converts the measured audio into digital information, and performs machine learning using the digital information as one element of training data or learning data to construct a trained model. In this way, audio information can be obtained with higher accuracy.
[0024] One aspect of the audio output device 1 further includes a muscle signal measurement unit 13 that measures muscle signals derived from the movement of muscles in a living body. In this aspect, the audio source information further includes muscle signals measured by the muscle signal measurement unit 13. This aspect is similar to that described in Japanese Patent No. 3455921. This aspect is preferable when the audio output device 1 is a sticker-type device or a Magneban-type device that is attached to the neck, because it can measure the movement of the muscles in the neck. There is a correlation between muscle signals derived from the movement of muscles in a living body and the audio that should actually be emitted. For this reason, in this aspect, the muscle signals are measured, and the obtained digital data is used as one element of the audio source information. In this aspect, too, the muscle signals can be used as one element of the training data or teacher data when constructing a trained model. In this way, audio information can be obtained with greater accuracy.
[0025] In one embodiment of the audio output device 1, the audio output device 1 further includes an imaging unit 15 that captures an image of an area including the mouth of a living body and obtains video data of the living body. In this embodiment, the audio source information further includes video data of the living body. For example, if the audio output device 1 is a helmet-type device or a mask-type device, the audio output device 1 includes an imaging unit 15 such as a digital camera, which can capture an image of an area including the mouth of a living body and obtain video data of the living body. In this case, the area including the mouth of a living body is captured, and machine learning is performed using the video data of the living body as training data or teacher data (data that also includes actual audio) to construct a trained model. Mouth movement and audio are correlated. Therefore, by using video data of an area including the mouth of a living body as one element of audio source information, audio information can be obtained with greater accuracy.
[0026] The audio output unit 7 is an element for outputting audio based on the audio information generated by the audio information generation unit 5. [Industrial Applicability]
[0027] The present invention can be used in the medical device and audio related industries. [Explanation of symbols]
[0028] 1. Audio output device 3. Vibration measurement section 5. Audio information generation unit 7 Audio output section 11 Audio measurement section 13 Muscle signal measurement unit 15 Imaging unit
Claims
1. a vibration measurement unit (3) that measures vibrations originating from a living body; a voice information generating unit (5) that generates voice information based on voice source information including vibration data that is data related to vibrations measured by the vibration measuring unit (3); an audio output unit (7) that outputs audio based on the audio information generated by the audio information generation unit (5), The voice information generation unit (5) has a machine learning unit (9), and obtains the voice information by inputting the voice source information into a trained model of the machine learning unit (9).
2. 2. An audio output device (1) according to claim 1, The vibrations originating from the living body are vibrations of air that occur when the living body tries to make a sound, and cannot be heard as sound by anyone other than the living body.
3. 2. An audio output device (1) according to claim 1, It further has a voice measurement unit (11) that measures voice, The audio output device (1), wherein the audio source information further includes audio data measured by the audio measurement unit (11).
4. 2. An audio output device (1) according to claim 1, The device further includes a muscle signal measuring unit (13) for measuring a muscle signal derived from the movement of the muscle of the living body, The audio source information further includes the muscle signal measured by the muscle signal measurement unit (13).
5. 2. An audio output device (1) according to claim 1, The method further includes an imaging unit (15) that captures an image of a region including the mouth of the living body and obtains video data of the living body, The audio output device (1), wherein the audio source information further includes video data of the living body.
Citation Information
Patent Citations
JP202-0124444A
Voice substitute device
JP3455921B2