Voice messaging device and conference system

The voice communication terminal addresses echo issues by controlling speaker and microphone operations based on speech detection, ensuring users can convey their intentions to the communication partner.

JP2026019293APending Publication Date: 2026-02-05SHARP KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024120766
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Conventional communication terminals face issues with echo (feedback) when audio output from a speaker is input to a microphone, especially when they are built into the same device or located close to each other, preventing simultaneous speaking by the user and the communication destination.

Method used

A voice communication terminal with a speaker, microphone, camera, and control unit that determines user speech from captured images, enabling or disabling voice output and input based on speech detection, and optionally sending a notification signal to the communication destination.

Benefits of technology

Enables users to express their intentions to the communication partner effectively by controlling echo and crosstalk, allowing for natural conversation even with echo suppressors enabled.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019293000001_ABST
    Figure 2026019293000001_ABST
Patent Text Reader

Abstract

To provide a voice messaging apparatus capable of indicating intention to a communication destination when a user speaks.SOLUTION: The voice messaging apparatus 10 includes a speaker 2 that transmits and receives a voice via a network and outputs a voice received from a communication destination, a microphone 3 that receives an input of a voice to be transmitted to the communication destination, a camera 4 that captures an image of a user, a control unit 5 that executes first control of enabling a voice output of the speaker 2 and disabling or suppressing a voice input of the microphone 3 when a voice is received from the communication destination, and an utterance determination unit 6 that determines whether or not the user is uttering on the basis of an image captured by the camera 4. When the utterance determination unit 6 determines that the user is uttering during execution of the first control, the control unit 5 executes the second control of disabling or suppressing the sound output of the speaker 2 and enabling the sound input of the microphone 3.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an audio communication terminal and a conference system that transmit and receive audio via a network. [Background technology]

[0002] In recent years, with the spread of communication networks, conference systems that hold conferences among multiple terminals connected to a communication network have become widely used. In conference systems, participants' voices are transmitted and received via speakers and microphones, and images of participants are transmitted and received via cameras. A method for controlling the orientation of a camera with respect to images transmitted and received in a conference system has been proposed (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2009-135634 Summary of the Invention [Problem to be solved by the invention]

[0004] A conventional communication terminal includes multiple cameras whose imaging direction can be adjusted, an image data extraction unit that extracts and outputs image data captured by the multiple cameras, an imaging direction instruction unit that instructs the imaging direction, and a control unit that controls the extraction of image data from a second camera facing in another direction when imaging in another direction is instructed while image data from a first camera is being extracted.

[0005] Regarding audio input and output in a conference system, there is a concern about echo (feedback) that occurs when the audio output from the speaker is input to the microphone. This becomes a problem when the speaker and microphone are built into the same device or are located close to each other, making it impossible for you and the other party (communication destination) to speak at the same time.

[0006] The present disclosure has been made to solve the above-mentioned problems, and aims to provide a voice communication terminal and a conference system that allow a user to express their intentions to a communication partner when they speak. [Means for solving the problem]

[0007] The voice communication terminal of the present disclosure is a voice communication terminal that transmits and receives voice via a network, and includes a speaker that outputs voice received from a communication destination, a microphone that accepts voice input to be sent to the communication destination, a camera that images a user, a control unit that executes a first control that enables the voice output of the speaker and disables or suppresses the voice input of the microphone when voice is received from the communication destination, and a speech determination unit that determines whether the user is speaking based on an image captured by the camera, and is characterized in that when the speech determination unit determines that the user is speaking while executing the first control, the control unit executes a second control that disables or suppresses the voice output of the speaker and enables the voice input of the microphone.

[0008] The voice communication terminal of the present disclosure is a voice communication terminal that transmits and receives voice via a network, and includes a speaker that outputs voice received from a communication destination, a microphone that accepts voice input to be sent to the communication destination, a camera that images a user, a control unit that executes a first control that enables the voice output of the speaker and disables or suppresses the voice input of the microphone when voice is received from the communication destination, and a speech determination unit that determines whether the user is speaking based on an image captured by the camera, and is characterized in that the control unit executes a third control that sends a notification signal to the communication destination when the speech determination unit determines that the user is speaking while executing the first control.

[0009] In the voice communication terminal according to the present disclosure, the notification signal may be a pre-stored voice.

[0010] In the voice communication terminal according to the present disclosure, the notification signal may be a notification indicating that the user is speaking.

[0011] A conference system according to the present disclosure is characterized by including an audio communication terminal according to the present disclosure. [Effects of the Invention]

[0012] According to the present disclosure, even if the echo suppressor function is enabled, when the user speaks, appropriate control is executed to allow the user to express his / her intention to the communication destination. [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a schematic diagram illustrating a conference system according to a first embodiment of the present disclosure. [Figure 2] FIG. 2 is a diagram illustrating the configuration of a voice communication terminal. [Figure 3] FIG. 4 is a schematic diagram showing a state in which a first control is being performed. [Figure 4] FIG. 4 is a schematic diagram showing a state in which second control is being performed. [Figure 5] FIG. 10 is a schematic diagram showing a state in which a third control is being performed. DETAILED DESCRIPTION OF THE INVENTION

[0014] (First embodiment) Hereinafter, an audio communication terminal and a conference system according to a first embodiment of the present disclosure will be described with reference to the drawings.

[0015] FIG. 1 is a schematic diagram illustrating a conference system according to a first embodiment of the present disclosure.

[0016] In a conference system 1 according to a first embodiment of the present disclosure, a plurality of voice communication terminals 10 transmit and receive voice to and from each other via a network. Each voice communication terminal 10 is, for example, a mobile phone terminal, a tablet, or a personal computer, and includes a speaker 2 that outputs voice received from a communication destination, a microphone 3 that accepts voice input to be transmitted to the communication destination, and a camera 4 that captures an image of the user. While FIG. 1 illustrates a case in which two users (a first user U1 and a second user U2) are conferencing, this is not limiting, and three or more users (voice communication terminals 10) may confer in the conference system 1. In the following description, the first user U1 is used as the reference and the second user U2 is used as the communication destination.

[0017] FIG. 2 is a diagram showing the configuration of a voice communication terminal.

[0018] The voice communication terminal 10 includes a speaker 2, a microphone 3, and a camera 4, as well as a control unit 5, an utterance determination unit 6, a storage unit 7, a display unit 8, and a communication unit 9.

[0019] The control unit 5 is a CPU provided in the voice communication terminal 10, and controls various operations of the voice communication terminal 10. The control performed by the control unit 5 will be described later with reference to Figs. 3 and 4.

[0020] The speech determination unit 6 determines whether the user is speaking based on the image captured by the camera 4. In determining whether the user is speaking, the speech determination unit 6 pays particular attention to the movement of the user's mouth among other things on the user's face, and recognizes whether the user is speaking. The content of the user's speech may also be inferred from the movement of the user's mouth.

[0021] The storage unit 7 is a storage medium such as a memory or a HDD, and stores various types of information. The storage unit 7 may also store the user's voice, which has been recorded in advance using the microphone 3 or the like.

[0022] The display unit 8 is a display that displays various images. Although not shown in Fig. 1, a camera 4 may also be provided on the second user U2 side, and an image of the second user U2 may be displayed on the display unit 8 on the first user U1 side. The communication unit 9 communicates with various terminals via a network or short-range wireless communication.

[0023] FIG. 3 is a schematic diagram showing a state in which the first control is being executed.

[0024] When receiving voice from the communication destination, the control unit 5 executes a first control (echo suppressor function) that enables the voice output of the speaker 2 and disables or suppresses the voice input of the microphone 3.

[0025] 3, the second user U2 is speaking and audio is being output from the speaker 2 on the first user U1 side. As a result of the control unit 5 executing the first control, audio input is disabled from the microphone 3 on the first user U1 side, so that sound emitted from the first user U1 side is not output from the speaker 2 on the second user U2 side. Note that, in the first control, instead of disabling audio input from the microphone 3, the volume of the audio input may be reduced to a level that does not cause echo (feedback).

[0026] 3, the second user U2, who is the communication destination, continues to speak, and while the first control is being executed, even if the first user U1 speaks, the voice does not reach the second user U2, preventing natural actions such as interrupting the conversation. In contrast, in this embodiment, even when the echo suppressor function is enabled, the second control described below is executed so that a user can interrupt the communication destination's speech (crosstalk).

[0027] FIG. 4 is a schematic diagram showing a state in which the second control is being executed.

[0028] When the speech determination unit 6 determines that the user is speaking while the control unit 5 is executing the first control, the control unit 5 executes the second control, which disables or suppresses the audio output from the speaker 2 and enables the audio input from the microphone 3.

[0029] Specifically, FIG. 4 corresponds to a state in which the first user U1 speaks while the first control is being executed on the first user U1 side as shown in FIG. 3. When it is determined that the first user U1 is speaking based on an image captured by the camera 4, the first control is interrupted and the second control is executed. In the second control, the output of the speaker 2 on the first user U1 side is disabled or the volume is lowered. Then, the microphone 3 on the first user U1 side is enabled for voice input, so that the first user U1's remarks are delivered to the second user U2. In this way, even if the echo suppressor function is enabled, when a user speaks, appropriate control can be executed to express the user's intention to the communication destination.

[0030] (Second embodiment) Next, a conference system according to a second embodiment of the present disclosure will be described with reference to the drawings. The second embodiment differs from the first embodiment in that a third control, which will be described later, is executed instead of the second control. Note that the second embodiment has substantially the same configuration as the first embodiment shown in Figures 1 to 4, and therefore a description thereof will be omitted, and only the differences will be described.

[0031] FIG. 5 is a schematic diagram showing a state in which the third control is being executed.

[0032] When the utterance determination unit 6 determines that the user is speaking while the control unit 5 is executing the first control, the control unit 5 executes the third control of transmitting a notification signal TS to the communication destination.

[0033] Specifically, Fig. 5 corresponds to a state in which the first user U1 is speaking while the first control is being executed on the first user U1 side as shown in Fig. 3. When it is determined that the first user U1 is speaking and the third control is executed, the voice input to the microphone 3 on the first user U1 side remains disabled, but a notification signal TS is transmitted to the second user U2 side in place of the first user U1's speech.

[0034] The notification signal TS is a pre-stored voice, and the substitute voice transmitted as the notification signal TS is a simple call such as "excuse me" or "hey." By transmitting a pre-prepared substitute voice in this way, crosstalk can be achieved without suppressing the voice from the other party. Furthermore, since the voice reaches the other party, it feels as if you are having a natural conversation.

[0035] Alternatively, instead of the substitute voice, a notification signal indicating that the user is speaking may be transmitted as a notification signal. The communication destination may display light or text based on the received notification.

[0036] It should be noted that the embodiments disclosed herein are illustrative in all respects and are not intended to be limiting. Therefore, the technical scope of the present disclosure should not be interpreted solely by the above-described embodiments, but should be defined based on the claims. Furthermore, all modifications within the scope and meaning equivalent to the claims are included. [Explanation of symbols]

[0037] 1. Conference System 2 speakers 3. Microphone 4. Camera 5. Control section 6. Speech determination unit 7 Memory section 8 Display 9. Communications Department 10 Voice communication terminals TS notification signal

Claims

1. A voice communication terminal that transmits and receives voice via a network, a speaker that outputs audio received from the communication destination; a microphone for receiving input of voice to be transmitted to the communication destination; a camera for capturing an image of a user; a control unit that executes a first control of enabling audio output from the speaker and disabling or suppressing audio input from the microphone when audio is received from the communication destination; an utterance determination unit that determines whether the user is speaking based on the image captured by the camera, When the speech determination unit determines that the user is speaking while the control unit is executing the first control, the control unit executes a second control of disabling or suppressing audio output from the speaker and enabling audio input from the microphone. A voice communication terminal characterized by:

2. A voice communication terminal that transmits and receives voice via a network, a speaker that outputs audio received from the communication destination; a microphone for receiving input of voice to be transmitted to the communication destination; a camera for capturing an image of a user; a control unit that executes a first control of enabling audio output from the speaker and disabling or suppressing audio input from the microphone when audio is received from the communication destination; an utterance determination unit that determines whether the user is speaking based on the image captured by the camera, When the utterance determination unit determines that the user is uttering during execution of the first control, the control unit executes a third control of transmitting a notification signal to the communication destination. A voice communication terminal characterized by:

3. 3. The voice communication terminal according to claim 2, The notification signal is a pre-stored voice signal. A voice communication terminal characterized by:

4. 3. The voice communication terminal according to claim 2, The notification signal is a notification indicating that the user is speaking. A voice communication terminal characterized by:

5. A conference system comprising the audio communication terminal according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Camera apparatus, communication terminal and image data output method

    JP2009135634A