Audio output determination device, electronic conference terminal device, electronic conference server device, and information processing method

The audio output determination device and electronic conference server system address the issue of direct voice hearing and delayed outputs by correlating ambient sound with received audio signals, enhancing voice clarity in electronic conferences.

JP2025146690APending Publication Date: 2025-10-03JVC KENWOOD CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025020055
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-21
Filing Date
2025-02-10
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

In conventional electronic conferences, participants wearing headsets in close proximity experience direct voice hearing and delayed voice output, making it difficult to distinguish and hear nearby participants clearly.

Method used

An audio output determination device that collects ambient sound and correlates it with received audio signals, determining if the output should be muted or reduced based on a threshold, and an electronic conference server that synthesizes and controls audio streams to exclude or mute specific terminals.

Benefits of technology

Enhances the clarity of hearing nearby participants' voices by reducing interference and delayed outputs, improving the overall audio experience in electronic conferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025146690000001_ABST
    Figure 2025146690000001_ABST
Patent Text Reader

Abstract

To make it easier to hear the voices of other nearby participants participating in the same electronic conference.SOLUTION: An audio output determination device includes a first communication unit that transmits and receives an audio stream, a sound collection unit that collects audio, an audio output unit that outputs the audio of the audio stream, and a determination unit that determines which audio streams should not be output from the audio output unit. The determination unit collects a second audio signal, which is ambient sound, with the sound collection unit, acquires a first audio signal included in the audio stream received by the first communication unit, and determines that the first audio signal should not be output from the audio output unit or that the output of the first audio signal should be reduced if the correlation value between the first audio signal and the second audio signal is greater than or equal to a predetermined threshold.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a voice output determination device, an electronic conference terminal device, an electronic conference server device, and an information processing method. [Background technology]

[0002] There is known an electronic conference system that connects remote locations via a communication network to hold a conference. Patent Document 1 discloses a technology that identifies the terminals of conference participants who are in the same room (close distance) via short-distance wireless, determines a representative terminal from among the terminals in the same room, and enables audio input and output only from the representative terminal. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2014-165888 Summary of the Invention [Problem to be solved by the invention]

[0004] In conventional electronic conferences, when each participant wears a headset and multiple participants are in the same conference room in close proximity with no partitions, the voices emitted by the participants can be heard directly by other participants, and the same voices are output with a delay from the headset via the network, making them difficult to hear.

[0005] The present invention aims to provide a voice output determination device, a teleconference terminal device, a teleconference server device, and an information processing method that make it easier to hear the voices of other nearby participants participating in the same teleconference. [Means for solving the problem]

[0006] The audio output determination device of the present invention comprises a first communication unit that transmits and receives an audio stream, a sound collection unit that collects audio, an audio output unit that outputs the audio of the audio stream, and a determination unit that determines which of the audio streams should not be output from the audio output unit, wherein the determination unit collects a second audio signal that is ambient sound using the sound collection unit, acquires a first audio signal included in the audio stream received by the first communication unit, and determines that the first audio signal should not be output from the audio output unit or that the output of the first audio signal should be reduced if the correlation value between the first audio signal and the second audio signal is greater than or equal to a predetermined threshold.

[0007] The electronic conference terminal device of the present invention is an electronic conference terminal device equipped with an audio output determination device, and further includes a second communication unit that wirelessly transmits and receives a second identifier, which is identification information related to other electronic conference terminal devices.When the second communication unit receives the second identifier related to at least one other electronic conference terminal, if the first identifier related to the other electronic conference terminal that transmitted the first audio signal matches the second identifier, and the correlation value between the first audio signal and the second audio signal is greater than or equal to the predetermined threshold, the determination unit does not output the first audio signal from the audio output unit.

[0008] The electronic conference server device of the present invention includes an information acquisition unit that acquires a mute list that indicates audio streams to be muted from among audio streams of an electronic conference received from a plurality of terminal devices, and a stream control unit that controls transmission of a synthesized audio stream to the first terminal device, the synthesized audio stream being obtained by synthesizing the audio streams from a plurality of second terminal devices that are different from the first terminal device to which the audio is to be transmitted.The stream control unit excludes the audio streams of the second terminal devices indicated by the mute list corresponding to the first terminal device from among the plurality of second terminal devices, and transmits the synthesized audio stream being obtained by synthesizing the audio streams of the remaining second terminal devices to the first terminal device.

[0009] The information processing method of the present invention is an information processing method executed by an audio output determination device having a first communication unit that transmits and receives an audio stream, an audio collection unit that collects audio, and an audio output unit that outputs the audio of the audio stream, and includes the steps of collecting a second audio signal, which is ambient sound, by the audio collection unit, acquiring a first audio signal included in the audio stream received by the first communication unit, and, if the correlation value between the first audio signal and the second audio signal is greater than or equal to a predetermined threshold, not outputting the first audio signal from the audio output unit or reducing the output of the first audio signal. [Effects of the Invention]

[0010] According to the present invention, it is possible to make it easier to hear the voices of other nearby participants who are participating in the same electronic conference. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of an electronic conference system according to the first embodiment. [Figure 2] FIG. 2 is a diagram showing the flow of an audio signal in the electronic conference system according to the first embodiment. [Figure 3] FIG. 3 is a block diagram showing an example of the configuration of a terminal device according to the first embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of a sequence of the electronic conference system according to the first embodiment. [Figure 5] FIG. 5 is a flowchart showing a processing procedure of the terminal device according to the first embodiment. [Figure 6] FIG. 6 is a diagram showing the flow of an audio signal and an identifier in the electronic conference system according to the second embodiment. [Figure 7] FIG. 7 is a block diagram showing an example of the configuration of a terminal device according to the second embodiment. [Figure 8] FIG. 8 is a diagram showing an example of a sequence of the electronic conference system according to the second embodiment. [Figure 9]FIG. 9 is a flowchart showing a processing procedure when transmitting an audio signal in a terminal device according to the second embodiment. [Figure 10] FIG. 10 is a flowchart showing a processing procedure when a terminal device according to the second embodiment receives an audio signal. [Figure 11] FIG. 11 is a diagram showing the flow of an audio signal in the electronic conference system according to the third embodiment. [Figure 12] FIG. 12 is a diagram illustrating an example of the configuration of a server device according to the third embodiment. [Figure 13] FIG. 13 is a flowchart showing a processing procedure of the server device according to the third embodiment. [Figure 14] FIG. 14 is a flowchart showing a processing procedure of the terminal device according to the fourth embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, an embodiment of the present invention will be described in detail with reference to the accompanying drawings. Note that the present invention is not limited to this embodiment, and in the following embodiments, the same components are designated by the same reference numerals, and redundant explanations will be omitted.

[0013] [First embodiment] (Electronic conference system) The electronic conference system according to the first embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram showing an example of the configuration of the electronic conference system according to the first embodiment.

[0014] The electronic conference system 1 includes a plurality of terminal devices 10 and a server device 100. The terminal device 10 is an example of an electronic conference terminal device equipped with an audio output determination device and is a device used by a user to participate in an electronic conference (online conference). The terminal device 10 is, for example, but not limited to, a notebook PC (Personal Computer), a desktop PC, a tablet terminal, a smartphone, an HMD (Head Mounted Display), or an HUD (Head Up Display). The server device 100 is an example of an electronic conference server device that transfers audio and video received from the terminal device 10 to another terminal device 10. The terminal devices 10 and the server device 100 are communicatively connected via a wired or wireless network N. The electronic conference system 1 may include any number of terminal devices 10 and server devices 100. The electronic conference system 1 is an online conference system in which a plurality of users can hold an online conference using their respective terminal devices 10.

[0015] FIG. 2 is a diagram showing the flow of audio signals in the electronic conference system according to the first embodiment. The electronic conference system 1 shown in FIG. 2 is a system that holds an electronic conference using an SFU (Selective Forwarding Unit) method. In the example shown in FIG. 2, the electronic conference system 1 holds an electronic conference between adjacently installed terminal devices 10A and 10B and a remotely installed terminal device 10C. The conference participants using the terminal devices 10A, 10B, and 10C participate in the electronic conference using headsets. A headset is a device that combines earphones or headphones with a microphone. In the following description, when there is no need to distinguish between the terminal devices 10A, 10B, and 10C, they will simply be referred to as the terminal device 10.

[0016] The electronic conference system 1 treats a series of signals, such as data, audio, and images, as a single stream. In the example shown in FIG. 2, the terminal device 10A adds its own identifier and other information to a series of audio signals acquired by its microphone and transmits (uploads) them to the server device 100 as an audio stream STA. Specifically, the terminal device 10A packets the series of audio signals for each predetermined data amount, and adds identifiers, serial numbers, timestamps, and other information to identify the packets as a series of audio signals before transmitting them. Packetization allows the signals to be time-division multiplexed with other data, image, and other packets for transmission. On the receiving side, the original series of audio signals is obtained by concatenating packets with the same identifier in serial number order. A stream is a flow of packets with the same identifier communicated over the network N. The terminal device 10B adds its own identifier and other information to a series of audio signals acquired by its microphone and transmits them to the server device 100 as an audio stream STB. The terminal device 10C adds its own identifier and other information to a series of audio signals acquired by its microphone and transmits them to the server device 100 as an audio stream STC. The server device 100 transmits (transfers) to the terminal device 10A the audio stream STB received from the terminal device 10B and the audio stream STC received from the terminal device 10C. The server device 100 transmits (transfers) to the terminal device 10B the audio stream STA received from the terminal device 10A and the audio stream STC received from the terminal device 10C. The server device 100 transmits (transfers) to the terminal device 10C the audio stream STA received from the terminal device 10A and the audio stream STB received from the terminal device 10B. As a result, the terminal device 10A realizes a teleconference by outputting the audio signals of the audio stream STB and the audio stream STC received from the server device 100 through earphones or headphones. The terminal device 10B realizes a teleconference by outputting the audio signals of the audio stream STA and the audio stream STC received from the server device 100 through earphones or headphones. The terminal device 10C realizes a teleconference by outputting the audio signals of the audio stream STA and the audio stream STB received from the server device 100 through earphones or headphones.

[0017] The electronic conference system 1 uses a method in which audio streams of devices other than the server device 100 are sent independently from each other, and connects the terminal device 10A, the terminal device 10B, and the terminal device 10C via a communication network to hold an electronic conference.

[0018] (Terminal Device) An example of the configuration of a terminal device according to the first embodiment will be described with reference to Fig. 3. Fig. 3 is a block diagram showing an example of the configuration of a terminal device according to the first embodiment.

[0019] 3, the terminal device 10 includes a camera 12, an operation unit 14, a sound collection unit 16, a display unit 18, an audio output unit 20, a first communication unit 22, a second communication unit 24, a storage unit 26, and a control unit 28. Note that the terminal device 10 may be configured to include only the first communication unit 22 of the first communication unit 22 and the second communication unit 24.

[0020] The camera 12 is a camera that captures an image of a subject, etc. The camera 12 captures, for example, the face, body, facial expression, and movements of a user who is using the terminal device 10.

[0021] The operation unit 14 accepts various operations for the terminal device 10. The operation unit 14 is realized by various input devices such as a keyboard, a mouse, a switch, a button, and a touch panel.

[0022] The sound collection unit 16 detects voice, ambient sound, etc. For example, the sound collection unit 16 detects voice emitted by a user using the terminal device 10. The sound collection unit 16 converts the detected voice into an audio signal. The sound collection unit 16 is realized by a microphone.

[0023] The display unit 18 has a display screen that displays various types of information. The display unit 18 is realized by a display including, for example, a liquid crystal display or an organic EL (Electro-Luminescence) display. The display unit 18 displays, for example, information related to an online conference on the display screen. The display unit 18 displays, for example, a list of participants in the online conference, shared materials, and faces of other users photographed by the camera 12 of another terminal device. If the operation unit 14 is a touch panel, the operation unit 14 and the display unit 18 are configured integrally. Here, if the terminal device 10 is a device that displays image information by projecting or projecting it, such as a HUD, the display unit 18 does not need to have a display screen.

[0024] The audio output unit 20 outputs an audio signal of an audio stream. The audio output unit 20 has a connection terminal such as a general-purpose terminal like an earphone microphone connector or a USB (Universal Serial Bus) terminal, and outputs the audio signal of the audio stream to an external output device connected to the connection terminal, causing the external output device to output the audio. The external output device includes, for example, a headset, earphones, an external speaker, etc. The audio output unit 20 has a speaker that outputs various types of audio. For example, the audio output unit 20 outputs the audio signal of the audio stream to output the audio of users of other terminal devices 10 participating in the online conference. Output from the audio output unit 20 includes outputting audio from an external output device connected to the connection terminal. The audio output unit 20 may be connected to the external output device via Bluetooth (registered trademark).

[0025] The first communication unit 22 is a communication interface that performs communication between the terminal device 10 and the server device 100. The first communication unit 22 performs communication between, for example, the terminal device 10 and the server device 100. The communication method of the first communication unit 22 is realized by, for example, a wireless LAN (Local Area Network), a wired LAN, or the like.

[0026] The second communication unit 24 is a communication interface that directly communicates with other terminal devices 10 via short-range wireless communication. The second communication unit 24 executes, for example, communication between its own device and other terminal devices 10. The communication method of the second communication unit 24 is realized by, for example, Wi-Fi (registered trademark), Bluetooth (registered trademark), or the like.

[0027] The storage unit 26 stores various types of information. The storage unit 26 stores information such as the contents of calculations performed by the control unit 28 and programs. The storage unit 26 includes at least one of a RAM (Random Access Memory), a main storage device such as a ROM (Read Only Memory), and an external storage device such as an HDD (Hard Disk Drive).

[0028] The storage unit 26 stores, for example, identification information 26A. The identification information 26A includes information for identifying various objects, such as an identifier of the device itself or another device, an identifier of a user who uses the device itself or another device, and an identifier of a stream to be communicated. The identifier includes, for example, an ID uniquely assigned to the device itself.

[0029] The control unit 28 controls each unit of the terminal device 10. The control unit 28 has, for example, an information processing device such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit), and a storage device such as a RAM or a ROM. The control unit 28 executes a program that controls the operation of the terminal device 10 according to the present invention. The control unit 28 may be realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). The control unit 28 may be realized by a combination of hardware and software.

[0030] The control unit 28 acquires various pieces of input information input by the user to the operation unit 14. The control unit 28 performs control to transmit a video stream including a video signal captured by the camera 12 to the server device 100 via the first communication unit 22. The control unit 28 performs control to display the video stream received from the server device 100 on the display unit 18 via the first communication unit 22. The control unit 28 performs control to transmit an audio stream including an audio signal collected by the audio collection unit to the server device 100 via the first communication unit 22. The control unit 28 performs control to output an audio signal of the audio stream received from the server device 100 from the audio output unit 20 via the first communication unit 22. The control unit 28 transmits a stream including the video stream, the audio stream, etc., and an identifier of the terminal device 10 to the server device 100. In this embodiment, the terminal device 10 will be described by omitting the video stream and focusing on the audio stream.

[0031] The control unit 28 includes a determination unit 28A as a functional block realized by processing by the control unit 28. The determination unit 28A determines an audio stream not to be output from the audio output unit 20. The determination unit 28A collects a second audio signal, which is an ambient sound, using the sound collection unit 16, acquires a first audio signal included in the audio stream received by the first communication unit 22, and determines whether the correlation value between the first audio signal and the second audio signal is equal to or greater than a predetermined threshold. If the correlation value between the first audio signal and the second audio signal is equal to or greater than the predetermined threshold, i.e., if there is a correlation greater than the predetermined threshold, the determination unit 28A determines not to output the first audio signal from the audio output unit 20 or to reduce the output volume of the first audio signal. If the first communication unit 22 receives multiple audio streams, the determination unit 28A calculates the correlation value between each of the first audio signal and the second audio signal and determines the first audio signal of the audio stream not to be output from the audio output unit 20.

[0032] The determination unit 28A uses a known autocorrelation method that takes into account the time lag between signals and differences in volume to determine the correlation between audio signals. In this case, the audio signal (first audio signal) of the audio stream transmitted from the other terminal device is the audio of the user (speaker) of the other terminal device speaking close to the audio pickup unit 16, so the user's (speaker's) voice is loud and the surrounding environmental sound (noise) is low. On the other hand, the second audio signal picked up by the audio pickup unit 16 of the terminal device 10, which is close to the other terminal device, is the audio of the user (speaker) of the other terminal device is quiet and the environmental sound is loud. The terminal device 10 may be configured to calculate a correlation value while amplifying the second audio signal picked up by the audio pickup unit 16 of the terminal device 10 and adopt the maximum correlation value. In this case, if the correlation value is high, the audio of the user (speaker) of the other terminal device is heard directly by the user of the terminal device 10 and is also heard as a delayed output from the audio output unit 20 of the terminal device 10. In this case, it is difficult to hear the voice of the user (speaker) of the other nearby terminal device, so it is better to mute the first audio signal and not output it. On the other hand, because the volume of the environmental sound included in the second audio signal is high, the correlation value may not be high even if the voice of the user (speaker) of the other terminal device is included. In this case, it is considered that the voice of the user (speaker) of the other terminal device is buried in the noise due to the high ambient noise, and the user of the own device cannot directly hear the voice of the user (speaker) of the other terminal device. In this case, it is better not to mute the first audio signal. In other words, in a noisy room with high ambient noise (a lot of noise), a high correlation value cannot be obtained even if the terminal devices are installed close to each other, so it is better not to mute the first audio signal. Taking such situations into consideration, a predetermined threshold is set and stored in the storage unit 26. The predetermined threshold may be set by machine learning using an evaluation questionnaire regarding ease of listening by users. Note that the correlation of audio signals may include the degree of agreement between audio signals.

[0033] After determining that the audio signal (first audio signal) received from the server device 100 should not be output from the audio output unit, if the correlation value between the audio signal (first audio signal) and the second audio signal becomes less than a predetermined threshold, the determination unit 28A outputs the audio signal from the audio output unit 20. As a result, when the volume of the speech of a nearby user decreases, the terminal device 10 can output the audio signal from the server device 100 from the audio output unit 20, thereby maintaining ease of hearing of the voices of nearby participants in the electronic conference.

[0034] (Sequence of electronic conference system) An example of a sequence of an audio signal in the SFU system of the electronic conference system according to the first embodiment will be described with reference to Fig. 4. Fig. 4 is a diagram showing an example of a sequence of the electronic conference system according to the first embodiment. In Fig. 4, in the electronic conference system 1, the terminal device 10A and the terminal device 10B are placed near each other in the same room, and the terminal device 10C is placed in a remote location. Each user of the terminal device 10A, the terminal device 10B, and the terminal device 10C is wearing a headset and participating in the electronic conference.

[0035] First, when the user of the terminal device 10C speaks, the terminal device 10C collects the voice with the sound collection unit 16, converts the collected voice signal MC into a voice stream, and starts transmitting it to the server device 100 (step S101).

[0036] The server device 100 transmits the audio signal MC received from the terminal device 10C as an audio stream to the terminal devices 10A and 10B (step S102). In this case, the server device 100 controls the destination of the audio signal MC.

[0037] Each of the terminal device 10A and the terminal device 10B outputs the audio signal MC received via the first communication unit 22 from the audio output unit 20. This allows the users of the terminal device 10A and the terminal device 10B to hear the speech of the user of the terminal device 10C from the audio output unit 20.

[0038] Next, when the user of terminal device 10B speaks, terminal device 10B collects the voice with sound collection unit 16, converts the collected voice signal MB into a voice stream, and starts transmitting it to server device 100 (step S103). In this case, the speech sound waves WB uttered by the user of terminal device 10B are directly heard by the user of terminal device 10A located near terminal device 10B.

[0039] The server device 100 transmits the audio signal MB received from the terminal device 10B as an audio stream to the terminal device 10A and the terminal device 10C (step S104). In this case, the server device 100 controls the destinations of the audio signal MB and the audio signal MC, respectively.

[0040] The terminal device 10C outputs the voice signal MB received via the first communication unit 22 from the voice output unit 20. This allows the user of the terminal device 10C to hear from the voice output unit 20 the speech of the user of the terminal device 10B.

[0041] Terminal device 10A receives the audio signal MB via first communication unit 22, collects a second audio signal including speech sound waves WB using audio collection unit 16, and determines whether the correlation value between audio signal MB and the second audio signal is greater than or equal to a predetermined threshold. If terminal device 10A determines that the correlation value between audio signal MB and the second audio signal is greater than or equal to the predetermined threshold, i.e., that the user of terminal device 10A can hear the speech sound waves WB, terminal device 10A mutes the output of audio signal MB to prevent the audio signal MB from being output from audio output unit 20 (step S105). As a result, the user of terminal device 10A can hear the speech sound waves WB uttered by the user of terminal device 10B without interference between the audio signal MB and the speech sound waves WB, since the audio of audio signal MB is not output from audio output unit 20. At this time, if the user of terminal device 10C is speaking, audio output of the user of terminal device 10C from audio output unit 20 continues.

[0042] When the user of terminal device 10B finishes speaking, terminal device 10B ends transmission of the audio signal MB to server device 100 as an audio stream (step S106). As a result, when server device 100 finishes receiving the audio signal MB from terminal device 10B, it ends transmission of the audio signal MB to terminal device 10A and terminal device 10C (step S107). Then, when terminal device 10C finishes receiving the audio signal MB via the first communication unit 22, it ends output of the audio of the user of terminal device 10B from the audio output unit 20. Furthermore, when terminal device 10A finishes receiving the audio signal MB via the first communication unit 22, it mutes the output of the audio signal MB and ends the audio output.

[0043] When the user of terminal device 10C finishes speaking, terminal device 10C finishes transmitting the audio signal MC to server device 100 as an audio stream (step S108). As a result, when server device 100 finishes receiving the audio signal MC from terminal device 10C, it finishes transmitting the audio signal MC to terminal device 10A and terminal device 10B (step S109). Then, when each of terminal device 10A and terminal device 10B finishes receiving the audio signal MC via first communication unit 22, it finishes outputting the audio of the user of terminal device 10C from audio output unit 20.

[0044] Thereafter, in the electronic conference system 1, the terminal device 10 transmits a voice signal uttered by the user to the server device 100 in the same manner, and the server device 100 distributes the received voice signal to the other terminal devices 10.

[0045] Also, in step S105, terminal device 10A receives a voice signal MB via first communication unit 22, collects a second voice signal including speech sound waves WB with sound collection unit 16, and, if it determines that the correlation value between voice signal MB and the second voice signal is less than a predetermined threshold and there is no correlation, outputs the received voice signal MB from voice output unit 20. As a result, the user of terminal device 10A can hear the voice signal MB spoken by the user of terminal device 10B without interference between the voice of voice signal MB and the voice of speech sound waves WB, since the voice of voice signal MB is output from voice output unit 20. At this time, if the user of terminal device 10C is speaking, the voices of both users of terminal device 10B and terminal device 10C are output from voice output unit 20.

[0046] (Terminal device processing procedure) An example of a processing procedure of the terminal device according to the first embodiment will be described with reference to Fig. 5. Fig. 5 is a flowchart showing the processing procedure of the terminal device according to the first embodiment. The processing procedure shown in Fig. 5 is realized by the control unit 28 of the terminal device 10 executing a program.

[0047] The control unit 28 of the terminal device 10 starts receiving an audio signal from the server device 100 (step S1001). In detail, the control unit 28 starts acquiring an audio signal (first audio signal) of an audio stream received from the server device 100 via the first communication unit 22. When the control unit 28 starts acquiring the first audio signal in step S1001, the process proceeds to S1002.

[0048] The control unit 28 determines whether or not a speech sound wave has been collected (step S1002). For example, the control unit 28 determines that a speech sound wave has been collected when the sound collection unit 16 is collecting the speech sound wave. If the control unit 28 determines that a speech sound wave has not been collected (step S1002; No), the control unit 28 proceeds to step S1004, which will be described later. On the other hand, if the control unit 28 determines that a speech sound wave has been collected (step S1002; Yes), the control unit 28 proceeds to step S1003.

[0049] The control unit 28 determines the correlation between the first audio signal received from the server device 100 and the second audio signal collected by the audio collection unit 16 (step S1003). For example, the control unit 28 determines the correlation between the first audio signal and the second audio signal, and stores the determined correlation value in the storage unit 26 in association with the identifier of the first audio signal. A higher correlation value indicates a higher degree of match. When the process of step S1003 is completed, the control unit 28 advances the process to S1004.

[0050] The control unit 28 determines whether the first audio signal has a correlation (step S1004). For example, if the correlation value of the first audio signal stored in the storage unit 26 in step S1003 is equal to or greater than a predetermined threshold, the control unit 28 determines that the first audio signal has a correlation. The predetermined threshold is a value set for determining the correlation value and is stored in the storage unit 26. Note that if it is determined in step S1002 that no speech sound waves are being collected (step S1002; No), the storage unit 26 does not store a correlation value, and therefore the first audio signal is not determined to have a correlation. In this case, for example, different thresholds may be set for a first threshold used for determining when the first audio signal is not muted and a second threshold used for determining when the first audio signal is already muted. By setting different thresholds, frequent switching between determining to mute and unmute the first audio signal can be prevented when the correlation value is close to the threshold.

[0051] When the control unit 28 determines that the first audio signal has a correlation (step S1004; Yes), the control unit 28 advances the process to step S1005. The control unit 28 does not output the first audio signal having a correlation with the speech sound wave from the audio output unit 20 (step S1005). That is, the control unit 28 mutes the output of the first audio signal having a correlation with the speech sound wave. When the process of step S1005 ends, the control unit 28 advances the process to step S1007, which will be described later.

[0052] Furthermore, if the control unit 28 determines that there is no correlation with the first audio signal (step S1004; No), the control unit 28 proceeds to step S1006. The control unit 28 outputs the first audio signal from the audio output unit 20 (step S1006). For example, if the first audio signal has already been muted, the control unit 28 unmutes the first audio signal and causes it to be output from the audio output unit 20. When the process of step S1006 ends, the control unit 28 proceeds to step S1007.

[0053] The control unit 28 determines whether reception of the first audio signal has ended (step S1007). Specifically, the control unit 28 determines that reception of the first audio signal has ended when reception of the audio stream from the server device 100 has ended via the first communication unit 22. If reception of the first audio signal has not ended (step S1007; No), the control unit 28 continues receiving the audio stream from the server device 100, and returns the process to step S1002, which has already been described, and continues the process. This allows the control unit 28 to monitor changes in the state of the second audio signal during reception of the first audio signal, for example, when the noise level in the room changes, and control the output of the first audio signal. Furthermore, if the control unit 28 determines that reception of the audio signal has ended (step S1007; Yes), it ends the processing procedure shown in FIG. 5. If multiple audio streams are received simultaneously, the process of FIG. 5 is performed for each of the multiple audio streams, and it is determined whether or not to output each first audio signal from the audio output unit 20.

[0054] 5, the method of muting the output of the first audio signal correlated with the speech sound wave (second audio signal) in step S1005 has been described, but the present invention is not limited to this. For example, step S1005 in the processing procedure may be replaced with a process of lowering (suppressing) the output volume of the first audio signal correlated with the speech sound wave (second audio signal).

[0055] In the processing procedure shown in FIG. 5, a determination process may be added immediately before step S1002 to determine whether the audio stream received from the server device 100 is an audio stream from a nearby terminal device 10. For example, identifiers of terminal devices 10 installed near the server device 100 may be stored in the storage unit 26, and the identifier of the terminal device 10 that is the sender, included in the audio stream received from the server device 100, may be compared to determine whether the audio stream is from a nearby terminal device 10. Then, in the processing procedure shown in FIG. 5, if it is determined that the received audio stream is an audio stream from a nearby terminal device 10, the process from step S1002 onwards is executed. Furthermore, in the processing procedure shown in FIG. 5, if it is determined that the received audio stream is not an audio stream from a nearby terminal device 10, the process from step S1002 onwards is not executed. The first audio signal of the audio stream for which the process from step S1002 onwards is not executed is output from the audio output unit 20. By adding such a determination process, the processing procedure shown in FIG. 5 can improve the accuracy of the determination.

[0056] In the terminal device 10 according to the first embodiment, when another terminal device 10 is located near the terminal device 10 and a user is participating in a teleconference using a headset, the terminal device 10 receives a voice signal of the user using the other terminal device 10 from the server device 100. When the correlation value between the received voice signal (first voice signal) and the second voice signal is equal to or greater than a predetermined threshold, the terminal device 10 does not output the received voice signal from the voice output unit 20. This prevents the terminal device 10 from outputting a voice signal of a nearby participant participating in the teleconference within a distance where the voice signal can be directly heard, thereby making the voice of the nearby participant easier to hear. Using the correlation value of the voice signal makes it easier to hear the voice of the nearby participant than when controlling the output from the voice output unit 20 based solely on location information indicating whether the nearby participant is installed nearby. Furthermore, even if the voice output unit 20 is a speaker, the terminal device 10 can make the voice of the nearby participant easier to hear by not outputting the voice signal of the nearby participant. As a result, even if the terminal device 10 is used in the same electronic conference as a different terminal device 10 located nearby, the voices of other participants can be more easily heard.

[0057] [Second embodiment] (Electronic conference system) The electronic conference system according to the second embodiment will be described with reference to Fig. 6. Fig. 6 is a diagram showing the flow of audio signals and identifiers in the electronic conference system according to the second embodiment.

[0058] The electronic conference system 1 is a system for holding an electronic conference using the SFU system, and includes a plurality of terminal devices 10-1 and a server device 100. In the example shown in Fig. 6, the electronic conference system 1 holds an electronic conference between adjacent terminal devices 10-1A and 10-1B and a remote terminal device 10-1C. The conference participants using the terminal devices 10-1A, 10-1B, and 10-1C participate in the electronic conference using headsets. In the following description, when there is no need to distinguish between the terminal devices 10-1A, 10-1B, and 10-1C, they will be simply referred to as the terminal device 10-1.

[0059] The electronic conference system 1 handles a series of signals such as data, audio, and images as a single flowing stream. In the example shown in Fig. 6, the terminal device 10-1A adds its own identifier to a series of audio signals acquired by its own microphone and transmits (uploads) the signals as an audio stream STA to the server device 100. The terminal device 10-1B adds its own identifier to a series of audio signals acquired by its own microphone and transmits the signals as an audio stream STB to the server device 100. The terminal device 10-1C adds its own identifier to a series of audio signals acquired by its own microphone and transmits the signals as an audio stream STC to the server device 100.

[0060] The server apparatus 100 transmits (transfers) to the terminal apparatus 10-1A an audio stream STB including the identifier received from the terminal apparatus 10-1B, and an audio stream STC including the identifier received from the terminal apparatus 10-1C. The server apparatus 100 transmits (transfers) to the terminal apparatus 10-1B an audio stream STA including the identifier received from the terminal apparatus 10-1A, and an audio stream STC received from the terminal apparatus 10-1C. The server apparatus 100 transmits (transfers) to the terminal apparatus 10-1C an audio stream STA including the identifier received from the terminal apparatus 10-1A, and an audio stream STB including the identifier received from the terminal apparatus 10-1B. As a result, the terminal apparatus 10-1A realizes a teleconference by outputting the audio signals of the audio stream STB and audio stream STC received from the server apparatus 100. The terminal apparatus 10-1B realizes a teleconference by outputting the audio signals of the audio stream STA and audio stream STC received from the server apparatus 100. The terminal device 10-1C outputs the audio signals of the audio streams STA and STB received from the server device 100, thereby realizing a teleconference.

[0061] Each of the terminal device 10-1A, terminal device 10-1B, and terminal device 10-1C has a function of transmitting its own identifier to another terminal device 10-1 installed near the own device by short-range wireless communication. In the example shown in Fig. 6, the terminal device 10-1A transmits its own identifier IDA to the terminal device 10-1B installed near the own device by short-range wireless communication. The terminal device 10-1B transmits its own identifier IDB to the terminal device 10-1A installed near the own device by short-range wireless communication.

[0062] (Terminal Device) An example of the configuration of the terminal device 10-1 according to the second embodiment will be described with reference to Fig. 7. Fig. 7 is a block diagram showing an example of the configuration of the terminal device 10-1 according to the second embodiment.

[0063] The terminal device 10-1 includes a camera 12, an operation unit 14, a sound collection unit 16, a display unit 18, an audio output unit 20, a first communication unit 22, a second communication unit 24, a memory unit 26, and a control unit 28.

[0064] The storage unit 26 stores, for example, identification information 26A, proximity information 26B, etc. The proximity information 26B includes information indicating a list of second identifiers related to other terminal devices 10-1 installed near the user's device (in the same room). The proximity information 26B indicates other terminal devices 10-1 installed nearby from which the user of the user's device may directly hear the audio emitted by the user of the terminal device 10-1, and indicates other terminal devices 10-1 that will mute or lower the volume of the audio output from the audio output unit 20 of the user's device. The proximity information 26B includes information indicating second identifiers related to other terminal devices 10 received via the second communication unit 24, second identifiers set by the administrator or user of the electronic conference, etc.

[0065] The control unit 28 includes, as functional blocks realized by the processing performed by the control unit 28, a determination unit 28A and a transmission / reception unit 28B.

[0066] When the second identifier is received by the second communication unit 24 from another terminal device 10-1 and stored in the proximity information 26B, the determination unit 28A of the control unit 28 determines whether the second identifier matches the first identifier related to the other terminal device 10-1 that transmitted the audio stream included in the audio stream received from the server device 100. If the first identifier matches the second identifier and the second identifier and the correlation value between the first audio signal and the second audio signal is equal to or greater than a predetermined threshold, the determination unit 28A determines not to output the audio signal included in the audio stream received from the server device 100 from the audio output unit 20, or to reduce the output volume of the audio signal.

[0067] The transmitter / receiver 28B transmits the identifier of its own device via the second communication unit 24 based on the identification information 26A of the storage unit 26. The transmitter / receiver 28B transmits the identifier of its own device to the outside, for example, at the timing when an audio signal of the voice uttered by the user of its own device and picked up by the sound pickup unit 16 is transmitted to the server device 100, or at the timing when an electronic conference is started. The transmitter / receiver 28B stores the identifier indicating the other terminal device 10 received via the second communication unit 24 in the proximity information 26B of the storage unit 26 as a second identifier.

[0068] (Sequence of electronic conference system) An example of a sequence of an audio signal in the SFU system of the electronic conference system 1 according to the second embodiment will be described with reference to Fig. 8. Fig. 8 is a diagram showing an example of a sequence of the electronic conference system 1 according to the second embodiment. In Fig. 8, in the electronic conference system 1, the terminal device 10-1A and the terminal device 10-1B are placed nearby in the same room, and the terminal device 10-1C is placed in a remote location. Each user of the terminal device 10-1A, the terminal device 10-1B, and the terminal device 10-1C is wearing a headset and participating in the electronic conference.

[0069] First, when the user of terminal device 10-1C speaks, terminal device 10-1C picks up the voice with sound collection section 16, converts the picked-up voice signal MC into a voice stream, and starts transmitting it to server device 100 (step S101).

[0070] Server apparatus 100 transmits audio signal MC received from terminal apparatus 10-1C as an audio stream to terminal apparatus 10-1A and terminal apparatus 10-1B (step S102). In this case, server apparatus 100 controls the destination of audio signal MC.

[0071] Each of terminal device 10-1A and terminal device 10-1B outputs the audio signal MC received via first communication unit 22 from audio output unit 20. This allows the users of terminal device 10-1A and terminal device 10-1B to hear the speech of the user of terminal device 10C from audio output unit 20.

[0072] Next, when the user of terminal device 10-1B speaks while the user of terminal device 10-1C is speaking, terminal device 10-1B picks up the voice with sound collection unit 16, converts the picked-up voice signal MB into a voice stream, and starts transmitting it to server device 100 (step S103). In this case, the speech sound waves WB uttered by the user of terminal device 10-1B can be directly heard by the user of terminal device 10-1A located near terminal device 10-1B. For this reason, terminal device 10-1B transmits its own identifier IDB from second communication unit 24 (step S110).

[0073] Server apparatus 100 transmits audio signal MB received from terminal apparatus 10-1B as an audio stream to terminal apparatus 10-1A and terminal apparatus 10-1C (step S104). In this case, server apparatus 100 controls the destinations of audio signal MB and audio signal MC, respectively.

[0074] The terminal device 10-1C outputs the voice signal MB received via the first communication unit 22 from the voice output unit 20. This allows the user of the terminal device 10-1C to hear from the voice output unit 20 the speech of the user of the terminal device 10B.

[0075] Furthermore, the terminal device 10-1A receives an identifier IDB (second identifier) ​​from the terminal device 10-1B via the second communication unit 24 (step S110) and stores it in the proximity information 26B of the storage unit 26. Then, the terminal device 10-1A compares the identifier IDB (first identifier) ​​indicating the terminal device 10-1B included in the audio stream of the audio signal MB received from the server device 100 via the first communication unit 22 with the second identifier. If it is determined that the first identifier and the second identifier are both identifier IDB and match, the audio signal MB received from the server device 100 is not output from the audio output unit 20 (step S111). As a result, the user of the terminal device 10-1A can hear the speech sound waves WB uttered by the user of the terminal device 10-1B without interference between the audio of the audio signal MB and the speech sound waves WB, since the audio of the audio signal MB is not output from the audio output unit 20. At this time, if the user of terminal device 10-1C is speaking, the voice output unit 20 continues to output the voice of the user of terminal device 10-1C.

[0076] When the user of terminal device 10-1B finishes speaking, terminal device 10-1B ends transmission of the audio signal MB to server device 100 as an audio stream (step S106). As a result, when server device 100 finishes receiving the audio signal MB from terminal device 10-1B, it ends transmission of the audio signal MB to terminal device 10-1A and terminal device 10-1C (step S107). Then, when terminal device 10-1C finishes receiving the audio signal MB via first communication unit 22, it ends output of the audio of the user of terminal device 10B from audio output unit 20. Also, when terminal device 10-1A finishes receiving the audio signal MB via first communication unit 22, it mutes the output of the audio signal MB and ends the audio output.

[0077] When the user of terminal device 10-1C finishes speaking, terminal device 10-1C ends transmission of voice signal MC to server device 100 in the voice stream (step S108). As a result, when server device 100 finishes receiving voice signal MC from terminal device 10-1C, it ends transmission of voice signal MC to terminal device 10-1A and terminal device 10-1B (step S109). Then, when terminal device 10-1A and terminal device 10-1B finish receiving voice signal MC via first communication unit 22, they each stop outputting the voice of the user of terminal device 10-1C from voice output unit 20.

[0078] Thereafter, in the electronic conference system 1, the terminal device 10-1 transmits a voice signal uttered by the user to the server device 100 in the same manner, and the server device 100 distributes the received voice signal to the other terminal devices 10-1.

[0079] Furthermore, in step S111, terminal device 10-1A receives identifier IDB (second identifier) ​​from terminal device 10-1B via second communication unit 24, and if the identifier IDB (first identifier) ​​indicating terminal device 10-1B included in the audio stream of audio signal MB received from server device 100 via first communication unit 22 does not match, terminal device 10-1A outputs audio signal MB received from server device 100 from audio output unit 20. This allows the user of terminal device 10-1A to hear the voice of the user of terminal device 10-1B output from audio output unit 20.

[0080] (Procedure for adjacent terminal devices) An example of the processing procedure of adjacently installed terminal devices 10-1 according to the second embodiment will be described using Fig. 9 and Fig. 10. Fig. 9 is a flowchart showing the processing procedure when the terminal device 10-1 according to the second embodiment transmits an audio signal. Fig. 10 is a flowchart showing the processing procedure when the terminal device according to the second embodiment receives an audio signal. The processing procedures shown in Fig. 9 and Fig. 10 are realized by the control unit 28 of the terminal device 10-1 executing a program.

[0081] 9, the control unit 28 of the terminal device 10-1 that transmits the audio signal determines whether or not the audio signal is currently being transmitted by the first communication unit 22 (step S1101). For example, the control unit 28 of the terminal device 10-1B determines that the audio signal is currently being transmitted by the first communication unit 22 when the audio stream of the audio signal MB is being transmitted to the server device 100 via the first communication unit 22. If the control unit 28 determines that the audio signal is not currently being transmitted by the first communication unit 22 (step S1101; No), the control unit 28 proceeds to step S1103, which will be described later. If the control unit 28 determines that the audio signal is currently being transmitted by the first communication unit 22 (step S1101; Yes), the control unit 28 proceeds to step S1102.

[0082] Control unit 28 transmits the identifier of its own device through second communication unit 24 (step S1102). For example, control unit 28 generates a transmission signal including identifier IDB indicating its own device based on identification information 26A in storage unit 26, and transmits the transmission signal from second communication unit 24. When control unit 28 completes the process of step S1102, it advances the process to step S1103.

[0083] The control unit 28 determines whether or not to terminate (step S1103). For example, the control unit 28 determines to terminate when a termination condition is met, such as when the transmission signal transmitted by the control unit in step S1102 is received by another terminal device 10-1 and a response to the reception of the identifier is returned from the other terminal device 10-1, or when the identifier has been repeatedly transmitted a predetermined number of times. If the control unit 28 determines not to terminate (step S1103; No), the control unit 28 returns the process to step S1101 already described and continues the process. On the other hand, if the control unit 28 determines to terminate (step S1103; Yes), the control unit 28 terminates the processing procedure shown in FIG. 9.

[0084] 10, the control unit 28 of the terminal device 10-1 receiving the voice signal determines whether or not the second communication unit 24 has received an identifier (step S1201). For example, the control unit 28 determines that the second communication unit 24 has received an identifier when a signal including the identifier is received via the second communication unit 24, or when the identifier is registered in the proximity information 26B of the storage unit 26. When the control unit 28 determines that the second communication unit 24 has not received an identifier (step S1201; No), the control unit 28 proceeds to step S1205, which will be described later. When the control unit 28 determines that the second communication unit 24 has received an identifier (step S1201; Yes), the control unit 28 proceeds to step S1202.

[0085] The control unit 28 determines whether or not speech sound waves are being collected (step S1202). For example, when sound waves are being collected by the sound collection unit 16, the control unit 28 determines that speech sound waves are being collected. When the control unit 28 determines that speech sound waves are not being collected (step S1202; No), the control unit 28 proceeds to the process at step S1205, which will be described later. When the control unit 28 determines that speech sound waves are being collected (step S1202; Yes), the control unit 28 proceeds to the process at step S1203.

[0086] The control unit 28 determines whether the identifiers match and whether there is a correlation between the audio signals (step S1203). For example, the control unit 28 determines that the identifiers match when the identifier of the other terminal device 10-1 received by the second communication unit 24 matches an identifier included in the audio stream of the audio signal received by the first communication unit 22, or when the identifier included in the audio stream of the audio signal received by the first communication unit 22 matches an identifier registered in the proximity information 26B of the storage unit 26. For example, the control unit 28 calculates a correlation value between the first audio signal received by the first communication unit 22 and the second audio signal collected by the audio collection unit 16, and determines that there is a correlation between the audio signals if the correlation value is equal to or greater than a predetermined threshold. If the control unit 28 determines that the identifiers match and there is a correlation between the audio signals (step S1203; Yes), the control unit 28 proceeds to step S1204.

[0087] The control unit 28 mutes the audio signal of the audio stream so that it is not output from the audio output unit 20 (step S1204). After completing the process of step S1204, the control unit 28 advances the process to step S1206, which will be described later.

[0088] Furthermore, if the control unit 28 determines in step S1203 that the identifiers match and that there is no correlation between the audio signals (step S1203; No), the control unit 28 proceeds to step S1205. The control unit 28 outputs the audio signal of the audio stream from the audio output unit 20 (step S1205). For example, if the output of the audio signal has been muted, the control unit 28 cancels the mute and causes the audio signal to be output from the audio output unit 20. When the process of step S1205 ends, the control unit 28 proceeds to step S1206. As a result, if there is a change in the state of the second audio signal while the first audio signal is being received, for example, if the noise level in the room changes, the control unit 28 can monitor the change in state as needed and control the output of the first audio signal.

[0089] The control unit 28 determines whether to terminate (step S1206). For example, the control unit 28 determines to terminate when a termination condition is met, such as when reception of the target audio stream is terminated or when a termination request from the user is received. When the control unit 28 determines not to terminate (step S1206; No), the control unit 28 returns the process to step S1201 already described and continues the process. When the control unit 28 determines to terminate (step S1206; Yes), the control unit 28 terminates the processing procedure shown in FIG. 10. When multiple audio streams are received simultaneously, the process of FIG. 10 is performed for each of the multiple audio streams, and it is determined whether or not to output each first audio signal from the audio output unit 20.

[0090] 10, the order of steps S1201 and S1202 may be reversed. That is, the processing procedure may be changed to a processing procedure in which first, "collection of speech sound waves" in the vicinity of the own device is determined, and then "reception of an identifier by the second communication unit" is determined.

[0091] In the terminal device 10-1 according to the second embodiment, when another terminal device 10-1 is located near the terminal device 10-1 and a user is participating in a teleconference using a headset, the terminal device 10-1 receives a voice signal of the user using the other terminal device 10-1 from the server device 100. When the correlation value between the received voice signal (first voice signal) and the second voice signal is equal to or greater than a predetermined threshold and the identifiers received by the second communication unit 24 match, the terminal device 10-1 does not output the received voice signal from the voice output unit 20. This prevents the terminal device 10-1 from outputting the voice signal of a nearby participant participating in the teleconference within a distance where the voice wave can be directly heard from the voice output unit 20, making it easier to hear the voice of the nearby participant. Furthermore, by receiving the identifier by the second communication unit 24, the terminal device 10-1 can determine the timing for collecting ambient sound and comparing it with the voice signal received by the first communication unit 22. When the first communication unit 22 simultaneously receives multiple audio streams, the terminal device 10-1 can easily identify the audio stream to be compared with the audio signal that captures the ambient sound to find a correlation value by receiving an identifier via the second communication unit 24. Furthermore, even if the audio output unit 20 is a speaker, the terminal device 10-1 can make the audio of nearby participants easier to hear by not outputting the audio signal of the nearby participants. As a result, even when the terminal device 10-1 is used in the same electronic conference as a different terminal device 10-1 located nearby, the terminal device 10-1 can make the audio of other participants easier to hear.

[0092] [Third embodiment] (Electronic conference system) The electronic conference system according to the third embodiment will be described with reference to Fig. 11. Fig. 11 is a diagram showing the flow of an audio signal in the electronic conference system according to the third embodiment.

[0093] The electronic conference system 1 shown in Fig. 11 is a system that holds an electronic conference using an MCU (Multi-point Control Unit) system. The electronic conference system 1 includes a plurality of terminal devices 10-1 and a server device 100-1. The MCU system is a system in which audio signals from devices other than the server device 100-1 are synthesized and sent to the terminal device 10-1 as a single audio stream.

[0094] The electronic conference system 1 treats a series of signals such as data, audio, and images as a single stream. In the example shown in FIG. 11, the terminal device 10-1A adds its own identifier and the like to a series of audio signals acquired by its own microphone and transmits (uploads) the result as an audio stream STA to the server device 100-1. The terminal device 10-1B adds its own identifier and the like to a series of audio signals acquired by its own microphone and transmits the result as an audio stream STB to the server device 100-1. The terminal device 10-1C adds its own identifier and the like to a series of audio signals acquired by its own microphone and transmits the result as an audio stream STC to the server device 100-1.

[0095] The server device 100-1 transmits to the terminal device 10-1A an audio signal obtained by combining the audio signal of the audio stream STB received from the terminal device 10-1B and the audio signal of the audio stream STC received from the terminal device 10-1C as an audio stream STBC. The server device 100-1 transmits to the terminal device 10-1B an audio signal obtained by combining the audio signal of the audio stream STA received from the terminal device 10-1A and the audio signal of the audio stream STC received from the terminal device 10-1C as an audio stream STAC. The server device 100-1 transmits to the terminal device 10-1C an audio signal obtained by combining the audio signal of the audio stream STA received from the terminal device 10-1A and the audio signal of the audio stream STB received from the terminal device 10-1B as an audio stream STAB. In this way, the terminal device 10-1A realizes a teleconference by outputting the audio signal of the audio stream STBC received from the server device 100. Terminal device 10-1B realizes the electronic conference by outputting the audio signal of audio stream STAC received from server device 100. Terminal device 10-1C realizes the electronic conference by outputting the audio signal of audio stream STAB received from server device 100.

[0096] The server device 100-1 further has a function of synthesizing the audio of the stream to be transmitted to each terminal device 10-1 based on the mute list received from each terminal device 10-1. For example, in the example shown in FIG. 11, the server device 100-1 stores the mute list received from each terminal device 10-1 in the mute information. The mute list is information that describes, for example, the identifiers of the terminal devices 10-1 for which audio signal synthesis is to be suppressed (no synthesis or synthesis with reduced volume). For example, if the mute list received from the terminal device 10-1A includes the identifier of the terminal device 10-1B, the server device 100-1 does not synthesize (mute) the audio signal of the audio stream STB from the terminal device 10-1B, and transmits to the terminal device 10-1A an audio stream STBC containing only the audio signal of the audio stream STC from the terminal device 10-1C. For example, if the mute list received from terminal device 10-1B contains the identifier of terminal device 10-1A, server device 100-1 does not synthesize (mute) the audio signal of audio stream STA from terminal device 10-1A to terminal device 10-1B, but transmits audio stream STAC containing only the audio signal of audio stream STC from terminal device 10-1C.

[0097] (Server device) An example of the configuration of the server device 100-1 according to the third embodiment will be described with reference to Fig. 12. Fig. 12 is a diagram showing an example of the configuration of the server device 100-1 according to the third embodiment.

[0098] The server device 100-1 acquires a series of video, audio, and other streams from multiple terminal devices 10-1 and distributes them to participants of an electronic conference. For example, in an electronic conference, the server device 100 may be realized by the terminal device 10-1 of the organizer of the electronic conference.

[0099] The server device 100-1 includes a communication unit 102, a storage unit 104, and a control unit .

[0100] The communication unit 102 is a communication interface that performs communication between the server device 100-1 and an external device. The communication unit 102 executes communication between the server device 100-1 and the terminal device 10-1, for example. The communication unit 102 is realized by, for example, a wireless LAN or a wired LAN.

[0101] The storage unit 104 stores various types of information. The storage unit 104 stores information such as the contents of calculations performed by the control unit 106 and programs. The storage unit 104 includes at least one of, for example, a RAM, a main storage device such as a ROM, and an external storage device such as an HDD.

[0102] The storage unit 104 stores mute information 200 including a mute list. The mute information 200 includes information that enables identification of other terminal devices 10-1 to be muted in the audio signals transmitted to each terminal device 10-1. In the example shown in FIG. 12, the mute information 200 includes information such as mute lists received from multiple terminal devices 10-1 and the identifiers of the terminal devices 10-1 that transmitted the mute lists. The mute information 200 associates each mute list with the identifiers of the terminal devices 10-1 that transmitted the mute lists. For example, the mute list contains the identifiers of the terminal devices 10-1 that are located near one or more terminal devices 10-1 that transmitted the mute lists to be muted.

[0103] The control unit 106 controls each unit of the server device 100. The control unit 106 has, for example, an information processing device such as a CPU or an MPU, and a storage device such as a RAM or a ROM. The control unit 106 executes a program that controls the operation of the server device 100 according to the present invention. The control unit 106 may be realized by, for example, an integrated circuit such as an ASIC or an FPGA. The control unit 106 may also be realized by a combination of hardware and software.

[0104] The control unit 106 includes an information acquisition unit 110 and a stream control unit 112 as functional blocks realized by processing by the control unit 106. Each of these functional blocks may be realized by the control unit 28 of the terminal device 10, or may be realized by sharing the processing between the control unit 106 and the control unit 28 of the terminal device 10-1.

[0105] The information acquisition unit 110 acquires a mute list indicating which audio streams of the electronic conference are to be muted from the multiple terminal devices 10-1. The information acquisition unit 110 creates and updates the mute information 200 based on the acquired mute list. For example, the information acquisition unit 110 acquires a mute list from the terminal device 10-1 via the communication unit 102, creates the mute information 200 based on the acquired mute list, and stores the mute information 200 in the storage unit 104. Note that the information acquisition unit 110 may also acquire the mute information 200 corresponding to the electronic conference from a database or the like and store it in the storage unit 104.

[0106] The stream control unit 112 controls transmission of an audio stream of an audio signal obtained by synthesizing audio signals of audio streams from multiple terminal devices 10-1 (second terminal devices) different from the terminal device 10-1 (first terminal device) that is the transmission target to the terminal device 10-1 (first terminal device). The stream control unit 112 excludes audio streams of second terminal devices indicated in the mute list corresponding to the terminal device 10-1 (first terminal device) from among the multiple terminal devices 10-1 (second terminal devices) and transmits a synthesized audio stream obtained by synthesizing audio streams of the remaining second terminal devices to the first terminal device. In other words, when receiving audio streams from one or more terminal devices 10-1, the stream control unit 112 mutes the audio signals of the audio streams from the terminal device 10-1 indicated in the mute list of the terminal device 10-1 that is the transmission target, and transmits an audio stream (synthesized audio stream) obtained by synthesizing audio signals of audio streams from terminal devices 10-1 other than the transmission target to the terminal device 10-1 that is the transmission target. Furthermore, if the stream control unit 112 receives audio streams from multiple terminal devices 10-1 but does not receive an audio stream from the terminal device 10-1 indicated in the mute list of the terminal device 10-1 to which the audio stream is to be sent, or if the mute list of the terminal device 10-1 to which the audio stream is to be sent is not registered in the mute information 200, the stream control unit 112 transmits an audio stream that combines the audio signals from all terminal devices 10-1 other than the terminal device 10-1 to which the audio stream is to be sent to the terminal device 10-1 to which the audio stream is to be sent.

[0107] (Terminal Device) Similar to the terminal device 10-1 shown in FIG. 7, the terminal device 10-1 includes a camera 12, an operation unit 14, a sound collection unit 16, a display unit 18, an audio output unit 20, a first communication unit 22, a second communication unit 24, a memory unit 26, and a control unit 28.

[0108] The storage unit 26 stores, in the proximity information 26B, a mute list including identifiers of audio streams of other terminal devices 10-1 for which the determination unit 28A has determined that audio should not be output from the audio output unit 20. The mute list may include identifiers of nearby terminal devices 10-1 received by the second communication unit 24.

[0109] The control unit 28 transmits the mute list data to the teleconference server device 100-1 via the first communication unit 22 and controls the server device 100-1 to request that the server device 100-1 not transmit audio streams containing audio signals of audio streams with identifiers included in the mute list. If the teleconference system 1 uses the MCU system, the control unit 28 controls the transmission of the mute list in the proximity information 26B to the server device 100-1 via the first communication unit 22. The timing of transmitting the mute list includes, for example, the start of the teleconference, when an audio stream is transmitted or received, and when another terminal device 10-1 is detected that the determination unit 28A has determined not to output audio from the audio output unit 20 or to output audio.

[0110] (Server device processing procedure) An example of the processing procedure of the server device 100-1 according to the third embodiment will be described with reference to Fig. 13. Fig. 13 is a flowchart showing the processing procedure of the server device 100-1 according to the third embodiment. The processing procedure shown in Fig. 13 is realized by the control unit 106 of the server device 100-1 executing a program.

[0111] The control unit 106 of the server device 100-1 determines whether or not a mute list has been acquired (step S2001). For example, the control unit 106 determines that a mute list has been acquired when the mute list of the terminal device 10-1 has been acquired from the terminal device 10-1, a database, or the like. If the control unit 106 determines that a mute list has not been acquired (step S2001; No), the control unit 106 proceeds to step S2003, which will be described later. If the control unit 106 determines that a mute list has been acquired (step S2001; Yes), the control unit 106 proceeds to step S2002.

[0112] The control unit 106 sets the acquired mute list in the mute information 200 (step S2002). For example, if the acquired mute list is not registered, the control unit 106 newly sets the mute list in the mute information 200 in the storage unit 104, and if the acquired mute list is already registered, the control unit 106 updates the mute information 200 in the storage unit 104. After completing the process of step S2002, the control unit 106 proceeds to step S2003.

[0113] The control unit 106 determines whether to transmit an audio stream (step S2003). For example, the control unit 106 determines whether to transmit the audio signal of the audio stream received via the communication unit 102 to the terminal device 10-1 as an audio stream. If the control unit 106 determines not to transmit an audio stream (step S2003; No), the control unit 106 proceeds to step S2008, which will be described later. On the other hand, if the control unit 106 determines to transmit an audio stream (step S2003; Yes), the control unit 106 proceeds to step S2004.

[0114] The control unit 106 determines whether or not a mute list to be transmitted exists (step S2004). For example, the control unit 106 determines that a mute list to be transmitted exists when it extracts a mute list corresponding to the identifier of the terminal device 10-1 to be transmitted from the mute information 200 in the storage unit 104. If the control unit 106 determines that a mute list to be transmitted exists (step S2004; Yes), the control unit 106 proceeds to step S2005.

[0115] The control unit 106 excludes the audio stream received from the terminal device 10-1 indicated in the mute list and combines the audio signals of the audio streams received from the other terminal devices 10-1 (step S2005). For example, when the control unit 106 is receiving multiple audio streams, it excludes the audio stream from the terminal device 10-1 indicated in the mute list from the multiple audio streams and combines the audio signals of the audio streams from the other terminal devices 10-1 using publicly known voice synthesis software or the like. For example, when the control unit 106 is receiving multiple audio streams but the audio stream from the terminal device 10-1 indicated in the mute list is not included in the multiple audio streams being received, it combines the audio signals of all the audio streams from the terminal devices 10-1 being received using publicly known voice synthesis software or the like. After completing the process of step S2005, the control unit 106 proceeds to step S2007, which will be described later.

[0116] If the control unit 106 determines that there is no mute list to be transmitted (step S2004; No), the control unit 106 proceeds to step S2006. The control unit 106 synthesizes the audio signals of all the received audio streams (step S2006). For example, the control unit 106 synthesizes the audio signals of the audio streams received from all the terminal devices 10-1 using known audio synthesis software or the like. After completing step S2006, the control unit 106 proceeds to step S2007, which will be described later.

[0117] The control unit 106 transmits the synthesized audio stream to the terminal device 10-1 that is the transmission target (step S2007). For example, the control unit 106 controls the communication unit 102 to transmit the synthesized audio stream to the terminal device 10-1 that is the transmission target. When the process of step S2007 ends, the control unit 106 advances the process to step S2008.

[0118] The control unit 106 determines whether to terminate (step S2008). For example, the control unit 106 determines to terminate when a termination condition is met, such as the termination of the electronic conference or the receipt of a termination request from a user. If the control unit 106 determines not to terminate (step S2008; No), the control unit 106 returns the process to step S2001, which has already been described, and continues the process. If the control unit 106 determines to terminate (step S2008; Yes), the control unit 106 terminates the processing procedure shown in FIG. 13.

[0119] The server device 100-1 according to the third embodiment can exclude audio streams received from the terminal device 10-1 indicated in the mute list corresponding to the terminal device 10-1 that is the transmission target among the multiple terminal devices 10-1, and transmit a synthesized audio stream obtained by synthesizing audio signals of the audio streams of the other terminal devices 10-1 to the terminal device 10-1 that is the transmission target. By acquiring the mute list, the server device 100-1 can prevent the audio signal of a nearby participant from being output from the audio output unit 20 of a terminal device 10-1 that is located within a distance where the speech sound waves of the nearby participant participating in the same electronic conference can be directly heard, thereby making the audio of the nearby participant easier to hear. Furthermore, even if the audio output unit 20 of the terminal device 10-1 is a speaker, the server device 100-1 can prevent the audio signal of the nearby participant from being output, making the audio of the nearby participant easier to hear.

[0120] [Fourth embodiment] (Terminal device processing procedure) An example of a processing procedure of a terminal device according to the fourth embodiment will be described with reference to Fig. 14. Fig. 14 is a flowchart showing the processing procedure of the terminal device according to the fourth embodiment. The configuration of the terminal device according to the fourth embodiment is the same as the configuration of the terminal device 10, and therefore the description will be omitted.

[0121] The processing shown in FIG. 14 is the same as the processing shown in FIG. 5 except for the processing in steps S1003A, S1004A, and S1010, and therefore a description thereof will be omitted.

[0122] In step S1003A, the control unit 28 determines the correlation between the first audio signal received from the server device 100 and the second audio signal collected by the audio collection unit 16 (step S1003A). For example, the control unit 28 determines the correlation between the first audio signal and the second audio signal, and stores the determined correlation value in the storage unit 26 in association with the identifier of the first audio signal. At this time, the control unit 28 calculates the correlation value for a predetermined time interval (correlation period), and stores the time at which the correlation value was calculated in association with the identifier of the first audio signal in the storage unit 26. For example, if the correlation period is 5 seconds, the control unit 28 calculates the correlation value of the audio signal for 5 seconds, and stores the time in association with the identifier of the first audio signal in the storage unit 26. A known autocorrelation function may be used to calculate the correlation value. For example, the autocorrelation function may be, but is not limited to, a power spectrum calculated by Fourier transforming the audio signal. When the process of step S1003A ends, the control unit 28 advances the process to S1010.

[0123] In step S1010, the control unit 28 determines whether the first audio signal is silent (step S1010). Specifically, the control unit 28 detects whether the first audio signal is silent at the current time. Silence refers to, for example, the sound pressure in the frequency range of human voice being equal to or lower than a predetermined sound pressure. For example, the control unit 28 detects the first audio signal as silent when the sound pressure in the frequency range of human voice is equal to or lower than a predetermined sound pressure continuously for a predetermined period (silence period). For example, the control unit 28 may detect the first audio signal as silent when a state in which human voice cannot be detected using known voice recognition technology continues for a silent period. If the first audio signal is determined to be silent (step S1010; Yes), the control unit 28 proceeds to step S1004A. If the first audio signal is not determined to be silent (step S1010; No), the control unit 28 returns to step S1002 and continues the correlation determination process. As a result, the output state of the first audio signal is not changed, and the state in which the first audio signal is output or not output from the audio output unit continues.

[0124] In step S1004A, the control unit 28 determines whether or not the first audio signal has a correlation (step S1004A). Specifically, the control unit 28 refers to a correlation value associated with a time prior to the silent period from the current time, and determines that the first audio signal has a correlation if the correlation value is equal to or greater than a predetermined value. If it is determined that the first audio signal has a correlation (step S1004A; Yes), the control unit 28 proceeds to step S1005. If it is not determined that the first audio signal has a correlation (step S1004A; No), the control unit 28 proceeds to step S1006.

[0125] In the fourth embodiment, the terminal device 10 switches the audio output of the first audio signal between on and off depending on whether the correlation value between the first audio signal and the second audio signal is equal to or greater than a predetermined value, starting from the point in time when it detects that the first audio signal is silent. As a result, in the fourth embodiment, the audio output is not switched on and off while the user is speaking, and therefore it is possible to prevent the audio of the first audio signal from being output or not being output while the first audio signal is being spoken, and thus to prevent the user from finding the audio of the first audio signal unpleasant to the ear.

[0126] [Other embodiments] The terminal device 10 and the terminal device 10-1 may be configured to determine whether to mute the audio stream based only on the identifier. That is, the terminal device 10 and the terminal device 10-1 may include a first communication unit 22 that transmits and receives an audio stream of a teleconference, a second communication unit 24 that wirelessly transmits and receives identifiers of other teleconference terminal devices, a sound collection unit 16 that collects audio, an audio output unit 20 that outputs the audio of the audio stream, and a determination unit 28A that determines which audio streams should not be output from the audio output unit 20. The determination unit 28A may receive a second identifier from another teleconference terminal device via the second communication unit 24, acquire a first audio signal included in the audio stream received by the first communication unit 22, and, if it determines that the second identifier matches a preset first identifier, not output the first audio signal from the audio output unit 20 or reduce the volume of the first audio signal. Furthermore, the electronic conference terminal device may be equipped with a user interface that displays icons of participants in the electronic conference on the display unit 18 based on the identifier of the electronic conference terminal device that is transmitting the first audio signal received by the first communication unit 22, and accepts operations to prevent the selected first audio signal from being output from the audio output unit 20 or to lower the volume of the first audio signal by selecting an icon with the mouse of the operation unit 14.

[0127] The components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads and usage conditions. This distribution and integration configuration may also be performed dynamically.

[0128] Although the embodiments of the present invention have been described above, the present invention is not limited to the contents of these embodiments. Furthermore, the above-described components include those that can be easily imagined by a person skilled in the art, those that are substantially the same, and those that are within the so-called equivalent range. Furthermore, the above-described components can be combined as appropriate. Furthermore, various omissions, substitutions, or modifications of the components can be made without departing from the spirit of the above-described embodiments. [Explanation of symbols]

[0129] 1. Electronic conference system 10,10-1 Terminal equipment 12 Camera 14 Control section 16 Sound pickup section 18 Display 20 Audio output section 22 First Communications Department 24 Second Communications Department 26 Memory section 26A Identification Information 26B Neighborhood Information 28 Control Unit 28A Judgment section 28B Transmitter / Receiver 100,100-1 Server device 102 Communications Department 104 Storage section 106 Control Unit 110 Information Acquisition Department 112 Stream control section 200 Mute Information

Claims

1. a first communication unit that transmits and receives an audio stream; a sound pickup unit that picks up sound; an audio output unit that outputs the audio of the audio stream; a determination unit that determines which of the audio streams should not be output from the audio output unit; Equipped with The determination unit collects a second audio signal, which is ambient sound, using the audio collection unit, acquires a first audio signal included in the audio stream received by the first communication unit, and, if the correlation value between the first audio signal and the second audio signal is greater than or equal to a predetermined threshold, determines not to output the first audio signal from the audio output unit or to reduce the output of the first audio signal.

2. the determination unit determines whether the first audio signal is silent, and when the first audio signal is silent and the correlation value between the first audio signal and the second audio signal is equal to or greater than a predetermined threshold, determines not to output the first audio signal from the audio output unit or to reduce the output of the first audio signal. The audio output determination device according to claim 1 .

3. An electronic conference terminal device comprising the audio output determination device according to claim 1 or 2, a second communication unit that wirelessly transmits and receives a second identifier that is identification information related to the other electronic conference terminal device; When the second identifier for at least one other electronic conference terminal is received by the second communication unit, if the first identifier for the other electronic conference terminal that transmitted the first audio signal matches the second identifier, and if the correlation value between the first audio signal and the second audio signal is greater than or equal to the predetermined threshold, the determination unit does not output the first audio signal from the audio output unit.

4. Even when the determination unit has determined that the audio of the first audio signal should not be output from the audio output unit, when a correlation value between the first audio signal and the second audio signal becomes less than the predetermined threshold, the determination unit changes to a determination that the audio of the first audio signal should be output from the audio output unit. The electronic conference terminal device according to claim 3.

5. a storage unit configured to store a mute list capable of identifying the audio streams for which the determination unit has determined that audio should not be output from the audio output unit; the first communication unit transmits data of the mute list to a server device of the electronic conference and requests the server device not to transmit the audio stream indicated by the mute list; The electronic conference terminal device according to claim 3.

6. an information acquisition unit that acquires a mute list indicating audio streams to be muted from among audio streams of the electronic conference received from a plurality of terminal devices; a stream control unit that controls transmission to the first terminal device of a synthesized audio stream obtained by synthesizing the audio streams from a plurality of second terminal devices different from the first terminal device to which the audio streams are to be transmitted; The stream control unit excludes the audio streams of the second terminal devices indicated in the mute list corresponding to the first terminal device among the plurality of second terminal devices, and transmits the synthesized audio stream obtained by synthesizing the audio streams of the remaining second terminal devices to the first terminal device.

7. a first communication unit that transmits and receives an audio stream; a sound pickup unit that picks up sound; an audio output unit that outputs the audio of the audio stream; An information processing method executed by an audio output determination device comprising: collecting a second audio signal, which is an ambient sound, by the sound collection unit; acquiring a first audio signal included in the audio stream received by the first communication unit; When the correlation value between the first audio signal and the second audio signal is equal to or greater than a predetermined threshold, the first audio signal is not output from the audio output unit or an output of the first audio signal is reduced. An information processing method, including:

Citation Information

Patent Citations

  • Conference terminal, conference server, conference system and program

    JP2014165888A