Voice call system

The voice call system addresses noise issues by converting in-ear sounds to text and back to voice, enhancing clarity and call quality through STT and TTS technologies, along with noise cancellation and screen display.

WO2026048641A1PCT designated stage Publication Date: 2026-03-05FOSTER ELECTRIC CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Existing voice communication systems fail to adequately suppress ambient noise, wind noise, and microphone noise caused by body movement, leading to unclear audio transmission.

Method used

A voice call system that converts in-ear sounds into text using Speech To Text (STT) technology, then converts the text into voice using Text To Speech (TTS) technology, with additional features like speaker identification, side tone generation, and noise cancellation to enhance clarity.

Benefits of technology

The system effectively reduces ambient and body movement noise, ensuring clearer voice transmission by displaying text on a screen and using voice synthesis to mimic the wearer's voice, thereby improving call quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025029271_05032026_PF_FP_ABST
    Figure JP2025029271_05032026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention comprises a text conversion unit that converts an in-ear voice signal to text, the in-ear voice signal being in-ear sound collected by an in-ear microphone installed on a voice call system headset, the in-ear sound including spoken voice from a speaker that is wearing the headset, a voice data conversion unit that uses prescribed voice synthesis data to convert the text to voice data in the language spoken by the speaker, and a communication endpoint that transmits the voice data to a call destination.
Need to check novelty before this filing date? Find Prior Art

Description

Voice Call System

[0001] The present invention relates to a voice communication system.

[0002] Japanese Patent Application Laid-Open Publication No. 2013-135334 states, "The communication terminal device of this embodiment is a communication terminal 1 capable of communicating with a voice terminal 2 via a telephone network, and includes an input unit 11 for inputting text data, a voice synthesis engine unit 13 for converting the text data into a voice signal by voice synthesis, a communication interface 15 for transmitting the voice signal to the voice terminal 2 and receiving the voice signal from the voice terminal 2, a voice recognition engine unit 14 for converting the voice signal from the voice terminal 2 into text data by voice recognition, and a display unit 12 for displaying the converted text data."

[0003] JP 2023-040244 A states that "an audio input / output device 200 which is a headphone, an earphone or a headset has an internal microphone 201 as a main audio acquisition unit, an external microphone 202 as a noise acquisition unit, a speaker 203 as an audio output unit, and an audio processing unit 290. The audio processing unit 290 has a noise cancellation unit 204 and an echo cancellation unit 205. The internal microphone 201 captures mixed audio which is a mixture of external noise, output audio 231 and the main audio, and outputs mixed audio signal 212. The external microphone 202 is arranged facing outside the body of the user 270, and captures external noise arriving from outside the user. A received signal 240 received by a communication unit 260 is converted into an output audio signal 232 and input to the speaker 203."

[0004] Japanese Patent Laid-Open Publication No. 08-340590 states, "A microphone 31 is provided inside the shell 27 for detecting a sound pressure signal output from the speaker 8 and converting it into an electrical signal c. A subtraction circuit 34 is inserted in the signal path of the amplified transmission signal a3 output from the vibration pickup 7, for subtracting the electrical signal c output from the microphone 31 from the transmission signal a3. Furthermore, an equalizer circuit 32 is inserted in the signal path of the electrical signal c output from the microphone 31 to the subtraction circuit 34, for correcting the frequency characteristics of the electrical signal output from the microphone 31 to the frequency characteristics of the amplified transmission signal a3 output from the vibration pickup 7."

[0005] JP 2022-506788 A states, "The system and method for providing application services using a noise-shielding earset according to the present invention includes a wireless earset including a left earphone including a left speaker driver unit, a left microphone, and a left wireless communication module, and a right earphone including a right speaker driver unit, a right microphone, and a right wireless communication module, and a terminal that processes and controls acoustic signals and voice signals for each of the left earphone and the right earphone, and provides services corresponding to the execution of an application, wherein the wireless earset is a noise-shielding earset, and the noise-shielding earset is characterized in that back holes of the left speaker driver unit and the right speaker driver unit are connected to fine holes that shield noise."

[0006] While clear conversations are desirable for headsets, no system has been established that can adequately suppress ambient noise, wind noise, and microphone noise caused by body movement, while still providing clear audio to the other party under all conditions.

[0007] For example, even when multiple microphones are used to collect speech using beamforming technology, wind noise and other ambient noises are still picked up. Even when algorithms to reduce ambient noise are applied, depending on the situation, speech can be reduced along with the surrounding noise, making it difficult to hear. While methods that are known to be resistant to noise include collecting speech using in-ear sound or contact pickup, the speech remains muffled and difficult to hear even after signal processing. It is also unavoidable that noise such as body movement is picked up.

[0008] The present disclosure has been made in consideration of the above circumstances, and aims to provide a voice call system that can transmit clearer voice to the other party when collecting spoken voice using in-ear sound, compared to when the voice collected using in-ear sound is transmitted to the other party as is.

[0009] A voice call system according to a first aspect of the present disclosure includes a text conversion unit that converts in-ear sounds, including the speech of a wearer wearing a headset, collected by an in-ear microphone mounted on the headset into text, a voice data conversion unit that uses predetermined voice synthesis data to convert the text into voice data in the language spoken by the wearer, and a call endpoint that transmits the voice data to the other party.

[0010] A voice call system according to a second aspect of the present disclosure is the voice call system according to the first aspect, wherein the voice data conversion unit converts the text into the voice data using the voice synthesis data generated from the wearer's spoken voice.

[0011] A voice call system according to a third aspect of the present disclosure is the voice call system according to the second aspect, further comprising a speaker identification unit that identifies a speaker based on at least one of the spoken voice of the wearer and the text.

[0012] A voice call system according to a fourth aspect of the present disclosure is the voice call system according to any one of the first to third aspects, further comprising a text display unit that displays the text on a screen.

[0013] A voice call system according to a fifth aspect of the present disclosure is the voice call system according to the fourth aspect, wherein the text conversion unit converts the call destination voice, including the speech of the call destination, into text and displays the text on a screen so that the wearer's speech and the speech of the call destination are distinguishable.

[0014] A voice call system according to a sixth aspect of the present disclosure is the voice call system according to any one of the first to fifth aspects, further comprising a side tone unit that adjusts external sounds including speech sounds of the wearer collected by an extra-aural microphone mounted on the headset and generates a side tone that is fed back to the wearer.

[0015] A voice call system according to a seventh aspect of the present disclosure is the voice call system according to the sixth aspect, further comprising a cancellation unit that cancels the speech of the call destination and the side tone from the in-ear sound.

[0016] The voice call system according to an eighth aspect of the present disclosure is the voice call system according to any one of the first to seventh aspects, further comprising a bypass unit that bypasses the functional unit that automatically controls the signal level of the in-ear sound.

[0017] A voice call system according to a ninth aspect of the present disclosure is the voice call system according to the eighth aspect, wherein the bypass unit bypasses the functional unit in a first mode in which the text conversion unit and the voice data conversion unit are activated, and does not bypass the functional unit in a second mode in which the text conversion unit and the voice data conversion unit are deactivated.

[0018] A voice call system according to a tenth aspect of the present disclosure is the voice call system according to any one of the first to ninth aspects, further comprising: a detection unit that detects a predetermined action by the wearer based on the in-ear sound; and a remote control unit that outputs a remote control signal that controls a controlled object in accordance with the action.

[0019] An eleventh aspect of the present disclosure relates to a voice call system in which, in the tenth aspect of the voice call system, the controlled object includes a voice assistant, and the remote control signal includes a wake-up command that causes the voice assistant to start a service.

[0020] A voice call system according to a twelfth aspect of the present disclosure is a voice call system according to any one of the first to eleventh aspects, including the headset and a terminal device communicatively connected to the headset, and the call endpoint is provided in the terminal device.

[0021] A voice call system according to a thirteenth aspect of the present disclosure is the voice call system according to the twelfth aspect, wherein the text conversion unit and the voice data conversion unit are provided in the terminal device.

[0022] A voice call system according to a fourteenth aspect of the present disclosure is the voice call system according to the twelfth aspect, wherein the text conversion unit is provided in the headset and the voice data conversion unit is provided in the terminal device.

[0023] A voice call system according to a fifteenth aspect of the present disclosure is the voice call system according to the twelfth aspect, wherein the text conversion unit and the voice data conversion unit are provided in the headset.

[0024] According to the voice call system of the present disclosure, when the spoken voice is collected by the ear, clearer voice can be transmitted to the call destination compared to when the voice collected by the ear is sent directly to the call destination.

[0025] FIG. 1 is a diagram showing an example of a schematic configuration of a voice call system 10 according to the present embodiment; FIG. 2 is a diagram showing an example of the internal structure of a headset 100 according to the present embodiment; FIG. 3 is a diagram showing an example of a functional configuration of the voice call system 10 according to the present embodiment; and FIG. 4 is a diagram showing an example of a functional configuration of a conventional voice call system 10'.

[0026] An example of an embodiment of the present disclosure will be described below with reference to the drawings. In each drawing, the same or equivalent components and parts are denoted by the same reference numerals. Furthermore, the dimensional proportions in the drawings are exaggerated for the sake of explanation and may differ from the actual proportions.

[0027] Prior to describing the voice call system 10 according to this embodiment, a conventional voice call system 10' will be described. Fig. 4 is a diagram showing an example of the functional configuration of the conventional voice call system 10'. The conventional voice call system 10' includes a headset 100' and a terminal device 500' communicably connected to the headset 100'.

[0028] The headset 100′ is a typical earphone for making calls. Typical earphones for making calls include inner-ear earphones, bone conduction earphones, and ear-hook earphones. The headset 100′ includes an extra-ear microphone 110′, a driver 120′, a side tone unit 130′, a cancellation unit 140′, an automatic gain control unit 150′, a communication endpoint 160′, and a remote control unit 170′.

[0029] The extra-aural microphone 110' is a microphone placed outside the ear, and collects the wearer's speech transmitted from outside the ear, as well as external sounds including speech of others and other noises.

[0030] The driver 120' is a speaker that outputs audio, and outputs audio received from the other end of the call (also called the "far side") and side tones.

[0031] The side tone unit 130' adjusts the wearer's speech transmitted from outside the ears, external sounds including other people's speech and other noises, and generates a side tone to be fed back to the wearer.

[0032] The cancellation unit 140' cancels the speech of the call destination output from the driver 120' from the microphone sound, thereby reducing the echo of the speech of the call destination returning to the call destination.

[0033] The automatic gain control unit 150' adjusts the volume of the call voice to an appropriate level.

[0034] The communication endpoint 160' communicates with the terminal device 500' via, for example, Bluetooth (registered trademark).

[0035] The remote control unit 170' controls the start and end of a call in response to an input operation.

[0036] The terminal device 500′ is a device capable of communicating with a call destination. Examples of the terminal device 500′ include a smartphone, a tablet terminal, and a notebook computer. The terminal device 500′ includes a communication endpoint 510′ and a call endpoint 520′.

[0037] The communication endpoint 510' communicates with the headset 100' via, for example, Bluetooth.

[0038] The call endpoint 520' conducts the call with the callee.

[0039] When making a call using open-type earphones, as in the conventional voice communication system 10', the headset 100' picks up voices from people other than the wearer. To prevent wind noise and other ambient noise from being picked up, the adoption of a boom-type microphone structure, in which a microphone is placed near the mouth, has been considered. However, a boom-type microphone structure imposes significant restrictions on the shape of the product, and is not easily implemented in the small, independent earphones that are currently widely used.

[0040] Therefore, it is considered to make a call using an in-ear microphone in a closed earphone. However, when making a call using an in-ear microphone in a closed earphone, although outside sounds are not picked up, the voice of the wearer that travels inside the head is collected by the closed ear canal, which can cause degradation such as muffled voice.

[0041] The voice call system 10 according to this embodiment has been made in consideration of the above circumstances, and aims to transmit clearer voice to the call destination when collecting the spoken voice as in-ear sound, compared to when the voice collected as in-ear sound is sent to the call destination as is.

[0042] 1 is a diagram showing an example of a schematic configuration of a voice call system 10 according to this embodiment. The voice call system 10 according to this embodiment includes a headset 100 and a terminal device 500 communicably connected to the headset 100. The headset 100 is a closed-type earphone, which differs from the headset 100'. The terminal device 500, like the terminal device 500', is a device capable of communicating with a call destination.

[0043] In this diagram, the headset 100 is shown as an in-ear type wireless earphone with independent left and right earphones as an example. Also, in this diagram, the terminal device 500 is shown as a smartphone as an example. The terminal device 500 may be communicably connected to the headset 100 using a short-range wireless communication standard such as Bluetooth.

[0044] 2 is a diagram showing an example of the internal structure of the headset 100 according to this embodiment. The headset 100 further includes an in-ear microphone 210 in addition to the extra-aural microphone 110 and the driver 120.

[0045] Canal-type earphones are used by inserting an ear tip (also called an earpiece, ear pad, or ear cap) 300 made of an elastic material such as silicone rubber, which is fixed to the outer periphery of the tip that is inserted into the ear canal, deep into the ear like an earplug. The earphones form an internal space (also called "in-ear") connected to the inner ear and ear canal, and an external space (also called "out-of-ear") that is theoretically separated from the internal space. The out-of-ear microphone 110 may be disposed in the external space, and the in-ear microphone 210 may be disposed in the internal space. The voice call system 10 according to this embodiment uses the in-ear microphone 210 disposed in the ear as a microphone for collecting spoken voice.

[0046] 3 is a diagram showing an example of the functional configuration of the voice call system 10 according to this embodiment. The headset 100 includes an extra-ear microphone 110, a driver 120, a side tone unit 130, a cancellation unit 140, an automatic gain control unit 150, a communication endpoint 160, and a remote control unit 170, as well as an in-ear microphone 210, a bypass unit 220, and a detection unit 230. The extra-ear microphone 110 to the remote control unit 170 correspond to the extra-ear microphone 110' to the remote control unit 170' in a conventional headset 100', respectively.

[0047] The extra-aural microphone 110 is a microphone placed outside the ear, and collects the wearer's speech transmitted from outside the ear, as well as external sounds including other people's speech and other noises.

[0048] The driver 120 is a speaker that outputs audio, and outputs audio received from the other end of the call and side tones.

[0049] The side tone unit 130 collects and adjusts the wearer's speech transmitted from outside the ear, as well as external sounds including speech of others and other noises, using the extra-ear microphone 110, and generates a side tone that is fed back to the wearer.

[0050] The cancellation unit 140 cancels the speech of the call destination and the side tone from the in-ear sound, thereby reducing input of sounds other than the wearer's speech to the text conversion unit 610, which will be described later.

[0051] The automatic gain control unit 150 adjusts the volume of the call audio to an appropriate level. That is, the automatic gain control unit 150 is a functional unit that automatically controls the signal level of the in-ear audio. Note that, as will be described later, this functional unit can be bypassed in this embodiment.

[0052] The communication endpoint 160 communicates with the terminal device 500 via Bluetooth, for example.

[0053] The remote control unit 170 outputs a remote control signal for controlling a controlled object in response to an operation detected by a detection unit 230 (to be described later).

[0054] The in-ear microphone 210 is a microphone placed in the ear, and collects in-ear sounds including the wearer's speech.

[0055] The bypass unit 220 bypasses the functional unit that automatically controls the signal level of the in-ear sound. More specifically, the bypass unit 220 bypasses this functional unit in a first mode in which the text conversion unit 610 and the voice data conversion unit 630, which will be described later, are activated, and does not bypass this functional unit in a second mode in which the text conversion unit 610 and the voice data conversion unit 630 are deactivated.

[0056] Here, the automatic gain control unit 150 is a functional unit provided for conventional voice processing for telephone calls. However, it has been found that performing this processing reduces the accuracy of voice recognition in the text conversion unit 610, which will be described later. Therefore, in this embodiment, the bypass unit 220 can bypass the automatic gain control unit 150 when performing text conversion and voice conversion.

[0057] The detector 230 detects a predetermined action by the wearer based on the in-ear sound. The remote controller 170 then outputs a remote control signal to control a control target in response to the action. This allows control of starting and ending a call, etc.

[0058] The controlled object may include a voice assistant. In this case, the remote control signal may include a wake-up command that causes the voice assistant to start a service.

[0059] Here, examples of the predetermined action include throat clearing and jaw up and down movement. As an example, by pre-registering the sound of the wearer clearing their throat, the remote control unit 170 may output a wake-up command when the detection unit 230 detects the throat clearing. Also, by pre-registering pressure fluctuations in the ear canal when the wearer moves their jaw up and down, the remote control unit 170 may output a control command when the detection unit 230 detects the jaw up and down movement. In this case, the wearer can end a call by, for example, clearing their throat twice and then moving their jaw up and down three times.

[0060] Conventionally, a technology for remote control by speaking pre-registered words is known. However, it is difficult to use in environments where there are other people around, such as on a train. Another technology for remote control by tapping earphones is also known, but remote control cannot be performed if your hands are full or you are wearing equipment that covers your ears, such as a helmet. In contrast, according to this embodiment, remote control can be performed without being noticed by others, even if your hands are full or you are wearing equipment that covers your ears.

[0061] The terminal device 500 includes a communication endpoint 510 and a call endpoint 520 , as well as a text conversion unit 610 , a speaker identification unit 620 , a voice data conversion unit 630 , and a text display unit 640 .

[0062] The communication endpoint 510 communicates with the headset 100 via, for example, Bluetooth.

[0063] The call endpoint 520 communicates with the call recipient. In this embodiment, the call endpoint 520 does not transmit collected voice data to the call recipient as is, but rather, as will be described later, transmits voice data that has been converted into text and then converted back into voice.

[0064] The text conversion unit 610 converts into text the in-ear sound, including the speech of the wearer wearing the headset 100, collected by the in-ear microphone 210 mounted on the headset 100. In this case, the text conversion unit 610 may use a technology such as STT (Speech To Text).

[0065] The text conversion unit 610 may perform the text conversion process using an internal STT engine of the terminal device 500. Alternatively, the text conversion unit 610 may perform the text conversion process by accessing an external STT engine provided as an external service. In this case, for example, if the external service is unavailable for some reason, the terminal device 500 may control the text conversion unit 610 and the voice data conversion unit 630 to be inactive and notify the headset 100 of an instruction to switch from the first mode to the second mode.

[0066] In the above description, the text conversion unit 610 converts into text only the speech of the wearer of the headset 100. However, the text conversion unit 610 may also convert into text the speech of the call destination received via the call endpoint 520.

[0067] The speaker identification unit 620 identifies the speaker based on at least one of the wearer's speech and the text. When identifying the speaker based on the wearer's speech, the speaker identification unit 620 may identify the speaker by, for example, identifying a voiceprint. When identifying the speaker based on the converted text, the speaker identification unit 620 may identify the speaker by, for example, the pronoun used in the text or the wording used.

[0068] The voice data conversion unit 630 uses predetermined voice synthesis data to convert text into voice data in the language spoken by the wearer. That is, the voice data conversion unit 630 converts text into voice data in the language spoken by the wearer without going through a translation process. In this case, the voice data conversion unit 630 may use a technology such as TTS (Text To Speech).

[0069] The voice data conversion unit 630 may perform voice data conversion using an internal TTS engine of the terminal device 500. Alternatively, the voice data conversion unit 630 may perform voice data conversion processing by accessing an external TTS engine provided as an external service. In this case, if the external service is unavailable for some reason, the terminal device 500 may control the text conversion unit 610 and the voice data conversion unit 630 to be inactive and notify the headset 100 of an instruction to switch from the first mode to the second mode.

[0070] When executing the voice data conversion process, if the speaker has been identified, the voice data conversion unit 630 may convert text into voice data using voice synthesis data generated from the wearer's spoken voice. This allows the voice data conversion unit 630 to restore voice data with the wearer's own voice tone intact. However, this is not limited to this. The voice data conversion unit 630 may also convert text into voice data using voice synthesis data generated from the spoken voice of any speaker. This also allows the voice data conversion unit 630 to function as a voice avatar.

[0071] The text display unit 640 displays the converted text on the screen. In the past, it was difficult to understand how what was said sounded to the other party. To address this, while sidetone has been used to provide some feedback of the spoken voice, it has sometimes been insufficient. This problem is particularly pronounced when making a call in a noisy environment. In contrast, according to this embodiment, the text conversion result is displayed as text on the screen of the terminal device 500, allowing the speaker wearing the headset 100 to understand what is being conveyed to the other party. This allows the wearer to make a call while checking the call quality in real time.

[0072] Furthermore, when the text conversion unit 610 converts the speech of the call destination into text in addition to the speech of the wearer, the text display unit 640 may display the text converted from the speech of the call destination on the screen in addition to the text converted from the speech of the wearer. In this case, the text display unit 640 may display the text on the screen so that the speech of the wearer and the speech of the call destination can be distinguished from each other by, for example, adding an identifier that identifies which text was spoken or by displaying it in a different color.

[0073] As described above, the voice call system 10 according to the present embodiment uses an approach in which in-ear sounds are converted into text using STT or the like, and then converted back into audio using TTS or the like. This allows the voice call system 10 according to the present embodiment to reduce the transmission of sounds other than conversation to the call recipient as the call voice. Therefore, the voice call system 10 according to the present embodiment can transmit clear audio to the call recipient, with ambient noise, wind noise, body movement noise, and the like reduced. In this case, the voice call system 10 according to the present embodiment uses an in-ear microphone 210 mounted in a closed-type earphone or the like, so that ambient sounds do not reach the text conversion engine in principle, thereby minimizing the transmission of ambient sounds to the call recipient as the call voice.

[0074] Although one possible embodiment has been described above as an example, the technology according to this embodiment can be modified or applied in various ways.

[0075] For example, in the above description, the case where the text conversion unit 610 and the voice data conversion unit 630 are provided in the terminal device 500 has been described as an example, but the present invention is not limited to this.

[0076] The text conversion unit 610 may be provided in the headset 100, and the voice data conversion unit 630 may be provided in the terminal device 500. In this case, the data exchanged in communication between the headset 100 and the terminal device 500 is text data, so the amount of data can be reduced compared to exchanging voice data. Therefore, for example, in communication that allows collisions using frequency hopping, such as Bluetooth, the probability of collisions can be reduced.

[0077] The text conversion unit 610 and the voice data conversion unit 630 may also be provided in the headset 100. This allows processing to be completed within the headset 100 without relying on processing by the terminal device 500. In this case, the voice call system 10 may include only the headset 100, without including the terminal device 500. In other words, the term "system" in the voice call system 10 may be interpreted as a concept that also includes a system configured from a single device such as an earphone.

[0078] Furthermore, although the above description has been given with reference to an example in which the system is used for a two-party call, the present invention is not limited to this. The voice call system 10 according to the present embodiment can also be used for conversations between multiple people. Therefore, by participating in a web conference using the voice call system 10 according to the present embodiment, the quality of the conversation in the conference can be improved and the converted text can be recorded as minutes of the conference.

[0079] In addition to the above, it goes without saying that the present disclosure can be implemented in various modifications within the scope of the gist thereof.

[0080] The disclosure of Japanese Patent Application No. 2024-151008, filed on September 2, 2024, is incorporated herein by reference in its entirety. In addition, all documents, patent applications, and technical standards described herein are incorporated herein by reference to the same extent as if each individual document, patent application, and technical standard was specifically and individually indicated to be incorporated by reference.

Claims

1. A voice call system comprising: a text conversion unit that converts in-ear sounds, including the speech of a wearer wearing a headset, collected by an in-ear microphone mounted on the headset into text; a voice data conversion unit that uses predetermined voice synthesis data to convert the text into voice data in the language spoken by the wearer; and a call endpoint that transmits the voice data to the other party.

2. The voice communication system according to claim 1, wherein the voice data conversion unit converts the text into the voice data using the voice synthesis data generated from the wearer's speech.

3. The voice communication system according to claim 2, further comprising a speaker identification unit that identifies a speaker based on at least one of the spoken voice of the wearer and the text.

4. The voice communication system according to claim 1, further comprising a text display unit that displays the text on a screen.

5. The voice call system according to claim 4, wherein the text conversion unit converts the callee's voice, including the callee's speech, into text, and displays the text on a screen so that the wearer's speech and the callee's speech can be distinguished from each other.

6. The voice call system according to claim 1, further comprising a side tone unit that adjusts external sounds, including the wearer's speech, collected by an extra-aural microphone mounted on the headset and generates a side tone that is fed back to the wearer.

7. The voice communication system according to claim 6, further comprising a cancellation unit that cancels the speech of the call destination and the side tone from the in-ear sound.

8. The voice communication system according to claim 1, further comprising a bypass unit that bypasses the function unit that automatically controls the signal level of the in-ear sound.

9. The voice call system according to claim 8, wherein the bypass unit bypasses the functional unit in a first mode in which the text conversion unit and the voice data conversion unit are activated, and does not bypass the functional unit in a second mode in which the text conversion unit and the voice data conversion unit are deactivated.

10. The voice communication system according to claim 1, further comprising: a detection unit that detects a predetermined action by the wearer based on the in-ear sound; and a remote control unit that outputs a remote control signal that controls a controlled object in accordance with the action.

11. The voice call system according to claim 10, wherein the controlled object includes a voice assistant, and the remote control signal includes a wake-up command that causes the voice assistant to start a service.

12. The voice communication system according to any one of claims 1 to 11, comprising: the headset; and a terminal device communicatively connected to the headset; wherein the call endpoint is provided in the terminal device.

13. The voice communication system according to claim 12, wherein the text conversion unit and the voice data conversion unit are provided in the terminal device.

14. The voice communication system according to claim 12, wherein the text conversion unit is provided in the headset, and the voice data conversion unit is provided in the terminal device.

15. The voice communication system according to claim 12, wherein the text conversion unit and the voice data conversion unit are provided in the headset.

Citation Information

Patent Citations

  • Mouth covering gesture recognition method based on single-earphone voice conversation process

    CN112133313A

  • Utterance recognition device and computer program

    JP2019208138A

  • KeywordsSmart earphones with wake-up function

    JP2022506787A

  • Earphones, audio processing method, and audio processing program

    JP2023040244A

  • Acoustic control system and acoustic control method

    JP2023091448A