Voice output determination device, electronic conference terminal device, electronic conference server device, and information processing method

The audio output determination device in electronic conference systems addresses the challenge of hearing nearby participants by correlating incoming audio with ambient noise to mute or reduce overlapping voices, improving audio clarity in conferences.

WO2025197952A1PCT designated stage Publication Date: 2025-09-25JVC KENWOOD CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/010646
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-10
Filing Date
2025-03-19
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

In conventional electronic conferences, voices emitted by participants in close proximity can be directly heard by others and output with a delay, making it difficult to hear the voices of nearby participants clearly.

Method used

An audio output determination device that collects ambient sound and determines the correlation between incoming audio signals and ambient noise to mute or reduce audio output when the correlation exceeds a threshold, using a combination of audio collection, processing, and output control units in electronic conference systems.

Benefits of technology

Enhances the clarity of hearing nearby participants' voices by reducing interference from direct voice overlap and ambient noise, ensuring easier audio perception in electronic conferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025010646_25092025_PF_FP_ABST
    Figure JP2025010646_25092025_PF_FP_ABST
Patent Text Reader

Abstract

This audio output determination device comprises: a first communication unit that transmits and receives an audio stream; a sound collection unit that collects audio; an audio output unit that outputs audio of the audio stream; and a determination unit that determines an audio stream that is not to be output from the audio output unit. The determination unit: collects, by means of a sound collection unit, a second audio signal that is an ambient sound; acquires a first audio signal included in the audio stream received by the first communication unit; and makes a determination not to output the first audio signal from the audio output unit, or to reduce the output of the first audio signal, if a correlation value between the first audio signal and the second audio signal is greater than or equal to a prescribed threshold.
Need to check novelty before this filing date? Find Prior Art

Description

Audio output determination device, electronic conference terminal device, electronic conference server device, and information processing method

[0001] The present invention relates to a voice output determination device, an electronic conference terminal device, an electronic conference server device, and an information processing method.

[0002] There are known electronic conference systems that connect remote locations via a communication network to hold a conference. Patent Document 1 discloses a technology that identifies terminals of conference participants who are in the same room (close distance) via short-distance wireless communication, determines a representative terminal from among the terminals in the same room, and enables audio input and output only from the representative terminal.

[0003] JP 2014-165888 A

[0004] In conventional electronic conferences, when each participant wears a headset and multiple participants are in the same conference room in close proximity with no partitions, the voices emitted by the participants can be heard directly by other participants, and the same voices are output with a delay from the headset via the network, making them difficult to hear.

[0005] The present invention aims to provide a voice output determination device, a teleconference terminal device, a teleconference server device, and an information processing method that make it easier to hear the voices of other nearby participants participating in the same teleconference.

[0006] The audio output determination device of the present invention comprises a first communication unit that transmits and receives an audio stream, a sound collection unit that collects audio, an audio output unit that outputs the audio of the audio stream, and a determination unit that determines which of the audio streams should not be output from the audio output unit, wherein the determination unit collects a second audio signal that is ambient sound using the sound collection unit, acquires a first audio signal included in the audio stream received by the first communication unit, and determines that the first audio signal should not be output from the audio output unit or that the output of the first audio signal should be reduced if the correlation value between the first audio signal and the second audio signal is greater than or equal to a predetermined threshold.

[0007] The electronic conference terminal device of the present invention is an electronic conference terminal device equipped with an audio output determination device, and further includes a second communication unit that wirelessly transmits and receives a second identifier, which is identification information related to other electronic conference terminal devices, and when the second communication unit receives the second identifier related to at least one other electronic conference terminal, if the first identifier related to the other electronic conference terminal that transmitted the first audio signal matches the second identifier and the correlation value between the first audio signal and the second audio signal is greater than or equal to the predetermined threshold, the determination unit does not output the first audio signal from the audio output unit.

[0008] The electronic conference server device of the present invention comprises an information acquisition unit that acquires a mute list indicating audio streams to be muted from among audio streams of an electronic conference received from a plurality of terminal devices, and a stream control unit that controls the transmission of a synthesized audio stream to the first terminal device, which is a synthesized audio stream obtained by synthesizing the audio streams from a plurality of second terminal devices that are different from the first terminal device to which the audio is to be transmitted.The stream control unit excludes the audio streams of the second terminal devices indicated by the mute list corresponding to the first terminal device among the plurality of second terminal devices, and transmits to the first terminal device the synthesized audio stream obtained by synthesizing the audio streams of the remaining second terminal devices.

[0009] The information processing method of the present invention is an information processing method executed by an audio output determination device having a first communication unit that transmits and receives an audio stream, a sound collection unit that collects audio, and an audio output unit that outputs the audio of the audio stream, and includes the steps of collecting a second audio signal, which is ambient sound, by the sound collection unit, acquiring a first audio signal included in the audio stream received by the first communication unit, and, if the correlation value between the first audio signal and the second audio signal is greater than or equal to a predetermined threshold, not outputting the first audio signal from the audio output unit or reducing the output of the first audio signal.

[0010] According to the present invention, it is possible to make it easier to hear the voices of other nearby participants who are participating in the same electronic conference.

[0011] FIG. 1 is a diagram illustrating an example of the configuration of an electronic conference system according to a first embodiment. FIG. 2 is a diagram illustrating the flow of audio signals in the electronic conference system according to the first embodiment. FIG. 3 is a block diagram illustrating an example of the configuration of a terminal device according to the first embodiment. FIG. 4 is a diagram illustrating an example of a sequence of the electronic conference system according to the first embodiment. FIG. 5 is a flowchart illustrating a processing procedure of the terminal device according to the first embodiment. FIG. 6 is a diagram illustrating the flow of audio signals and identifiers in an electronic conference system according to a second embodiment. FIG. 7 is a block diagram illustrating an example of the configuration of a terminal device according to the second embodiment. FIG. 8 is a diagram illustrating an example of a sequence of the electronic conference system according to the second embodiment. FIG. 9 is a flowchart illustrating a processing procedure of the terminal device according to the second embodiment when transmitting an audio signal. FIG. 10 is a flowchart illustrating a processing procedure of the terminal device according to the second embodiment when receiving an audio signal. FIG. 11 is a diagram illustrating the flow of audio signals in an electronic conference system according to a third embodiment. FIG. 12 is a diagram illustrating an example of the configuration of a server device according to the third embodiment. FIG. 13 is a flowchart illustrating a processing procedure of the server device according to the third embodiment. FIG. 14 is a flowchart illustrating a processing procedure of the terminal device according to the fourth embodiment.

[0012] Hereinafter, an embodiment of the present invention will be described in detail with reference to the accompanying drawings. Note that the present invention is not limited to this embodiment, and in the following embodiments, the same components are designated by the same reference numerals, and redundant explanations will be omitted.

[0013] [First embodiment] (Electronic conference system) An electronic conference system according to a first embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram showing an example of the configuration of the electronic conference system according to the first embodiment.

[0014] The electronic conference system 1 includes multiple terminal devices 10 and a server device 100. The terminal device 10 is an example of an electronic conference terminal device equipped with an audio output determination device and is used by users to participate in an electronic conference (online conference). The terminal device 10 may be, for example, a laptop personal computer (PC), a desktop PC, a tablet terminal, a smartphone, a head-mounted display (HMD), or a head-up display (HUD), but is not limited to these. The server device 100 is an example of an electronic conference server device that transfers audio and video received from the terminal device 10 to other terminal devices 10. The terminal devices 10 and the server device 100 are communicatively connected via a wired or wireless network N. The electronic conference system 1 may include any number of terminal devices 10 and server devices 100. The electronic conference system 1 is an online conference system in which a plurality of users can hold an online conference using their respective terminal devices 10 .

[0015] FIG. 2 is a diagram illustrating the flow of audio signals in the electronic conference system according to the first embodiment. The electronic conference system 1 illustrated in FIG. 2 is a system that conducts an electronic conference using a selective forwarding unit (SFU) method. In the example illustrated in FIG. 2, the electronic conference system 1 conducts an electronic conference between adjacently installed terminal devices 10A and 10B and a remotely installed terminal device 10C. Conference participants using the terminal devices 10A, 10B, and 10C participate in the electronic conference using headsets. A headset is a device that combines earphones or headphones with a microphone. In the following description, when it is not necessary to distinguish between the terminal devices 10A, 10B, and 10C, they will simply be referred to as the terminal device 10.

[0016] The electronic conference system 1 treats a series of signals, such as data, audio, and images, as a single stream. In the example shown in FIG. 2 , the terminal device 10A adds its own device identifier and other information to a series of audio signals acquired by its microphone and transmits (uploads) them to the server device 100 as an audio stream STA. Specifically, the series of audio signals are packetized into packets of a predetermined data amount, and are transmitted with identifiers, serial numbers, timestamps, and other information to identify them as packets of a series of audio signals. Packetization allows the signals to be transmitted by time division multiplexing with other data, image, and other packets. On the receiving side, the original series of audio signals is obtained by concatenating packets with the same identifier in serial number order. A stream is a flow of packets with the same identifier communicated over the network N. The terminal device 10B adds its own device identifier and other information to a series of audio signals acquired by its microphone and transmits them to the server device 100 as an audio stream STB. The terminal device 10C adds its own device identifier and other information to a series of audio signals acquired by its microphone and transmits them to the server device 100 as an audio stream STC. The server device 100 transmits (transfers) to the terminal device 10A the audio stream STB received from the terminal device 10B and the audio stream STC received from the terminal device 10C. The server device 100 transmits (transfers) to the terminal device 10B the audio stream STA received from the terminal device 10A and the audio stream STC received from the terminal device 10C. The server device 100 transmits (transfers) to the terminal device 10C the audio stream STA received from the terminal device 10A and the audio stream STB received from the terminal device 10B. As a result, the terminal device 10A realizes a teleconference by outputting the audio signals of the audio stream STB and the audio stream STC received from the server device 100 from earphones or headphones. The terminal device 10B realizes a teleconference by outputting the audio signals of the audio stream STA and the audio stream STC received from the server device 100 from earphones or headphones. The terminal device 10C outputs the audio signals of the audio stream STA and the audio stream STB received from the server device 100 through earphones or headphones, thereby realizing an electronic conference.

[0017] The electronic conference system 1 uses a method in which audio streams from devices other than the server device 100 are sent independently, and connects the terminal devices 10A, 10B, and 10C via a communication network to hold an electronic conference.

[0018] (Terminal Device) An example of the configuration of a terminal device according to the first embodiment will be described with reference to Fig. 3. Fig. 3 is a block diagram showing an example of the configuration of a terminal device according to the first embodiment.

[0019] 3, the terminal device 10 includes a camera 12, an operation unit 14, a sound collection unit 16, a display unit 18, an audio output unit 20, a first communication unit 22, a second communication unit 24, a storage unit 26, and a control unit 28. Note that the terminal device 10 may be configured to include only the first communication unit 22 out of the first communication unit 22 and the second communication unit 24.

[0020] The camera 12 is a camera that captures an image of a target, etc. The camera 12 captures, for example, the face, body, facial expression, and movements of a user who is using the terminal device 10.

[0021] The operation unit 14 accepts various operations for the terminal device 10. The operation unit 14 is realized by various input devices such as a keyboard, a mouse, a switch, a button, and a touch panel.

[0022] The sound collection unit 16 detects voice, ambient sound, etc. For example, the sound collection unit 16 detects voice emitted by a user using the terminal device 10. The sound collection unit 16 converts the detected voice into an audio signal. The sound collection unit 16 is realized by a microphone.

[0023] The display unit 18 has a display screen that displays various information. The display unit 18 is realized by a display including, for example, a liquid crystal display or an organic EL (Electro-Luminescence) display. The display unit 18 displays, for example, information related to an online conference on the display screen. The display unit 18 displays, for example, a list of participants in the online conference, shared materials, and the faces of other users captured by the camera 12 of another terminal device. If the operation unit 14 is a touch panel, the operation unit 14 and the display unit 18 are configured as an integrated unit. Here, if the terminal device 10 is a device that displays image information by projecting or projecting it, such as a HUD, the display unit 18 does not need to have a display screen.

[0024] The audio output unit 20 outputs an audio signal of an audio stream. The audio output unit 20 has a connection terminal such as a general-purpose terminal like an earphone microphone connector or a USB (Universal Serial Bus) terminal, and outputs the audio signal of the audio stream to an external output device connected to the connection terminal, causing the external output device to output the audio as audio. The external output device includes, for example, a headset, earphones, an external speaker, etc. The audio output unit 20 has a speaker that outputs various types of audio. For example, the audio output unit 20 outputs the audio signal of the audio stream to output the audio of users of other terminal devices 10 participating in the online conference. Output from the audio output unit 20 includes outputting audio from an external output device connected to the connection terminal. The audio output unit 20 may be connected to the external output device via Bluetooth (registered trademark).

[0025] The first communication unit 22 is a communication interface that performs communication between the terminal device 10 and the server device 100. The first communication unit 22 performs communication between, for example, the terminal device 10 and the server device 100. The communication method of the first communication unit 22 is realized by, for example, a wireless LAN (Local Area Network), a wired LAN, or the like.

[0026] The second communication unit 24 is a communication interface that communicates directly with other terminal devices 10 via short-range wireless communication. The second communication unit 24 executes, for example, communication between the device itself and other terminal devices 10. The communication method of the second communication unit 24 is realized by, for example, Wi-Fi (registered trademark), Bluetooth (registered trademark), or the like.

[0027] The storage unit 26 stores various types of information, such as the contents of calculations performed by the control unit 28 and information such as programs. The storage unit 26 includes at least one of a main storage device such as a random access memory (RAM) and a read-only memory (ROM), and an external storage device such as a hard disk drive (HDD).

[0028] The storage unit 26 stores, for example, identification information 26A. The identification information 26A includes information for identifying various objects, such as identifiers of the device itself and other devices, identifiers of users who use the device itself and other devices, and identifiers of streams to be communicated. The identifiers include, for example, an ID uniquely assigned to the device itself.

[0029] The control unit 28 controls each unit of the terminal device 10. The control unit 28 has, for example, an information processing device such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit), and a storage device such as a RAM or a ROM. The control unit 28 executes a program that controls the operation of the terminal device 10 according to the present invention. The control unit 28 may be realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). The control unit 28 may be realized by a combination of hardware and software.

[0030] The control unit 28 acquires various pieces of input information input by the user to the operation unit 14. The control unit 28 controls the transmission of a video stream including a video signal captured by the camera 12 to the server device 100 via the first communication unit 22. The control unit 28 controls the display unit 18 to display the video stream received from the server device 100 via the first communication unit 22. The control unit 28 controls the transmission of an audio stream including an audio signal collected by the audio collection unit to the server device 100 via the first communication unit 22. The control unit 28 controls the output of an audio signal of the audio stream received from the server device 100 from the audio output unit 20 via the first communication unit 22. The control unit 28 transmits a stream including the video stream, the audio stream, etc., and an identifier of the terminal device 10 to the server device 100. In this embodiment, the terminal device 10 will be described focusing on the audio stream, omitting the video stream.

[0031] The control unit 28 includes a determination unit 28A as a functional block realized by processing by the control unit 28. The determination unit 28A determines an audio stream not to be output from the audio output unit 20. The determination unit 28A collects a second audio signal, which is ambient sound, using the sound collection unit 16, acquires a first audio signal included in the audio stream received by the first communication unit 22, and determines whether a correlation value between the first audio signal and the second audio signal is equal to or greater than a predetermined threshold. If the correlation value between the first audio signal and the second audio signal is equal to or greater than the predetermined threshold, i.e., if there is a correlation greater than the predetermined threshold, the determination unit 28A determines not to output the first audio signal from the audio output unit 20 or to reduce the output volume of the first audio signal. If the first communication unit 22 receives multiple audio streams, the determination unit 28A calculates the correlation value between each of the first audio signal and the second audio signal and determines the first audio signal of the audio stream not to be output from the audio output unit 20.

[0032] The determination unit 28A uses a known autocorrelation technique that takes into account the time lag between signals and differences in volume to determine the correlation between audio signals. In this case, the audio signal (first audio signal) of the audio stream transmitted by the other terminal device is the audio of the user (speaker) of the other terminal device speaking close to the audio pickup unit 16, so the user's (speaker's) voice is loud and the surrounding environmental sound (noise) is low. On the other hand, the second audio signal picked up by the audio pickup unit 16 of the terminal device 10, which is close to the other terminal device, is the audio of the user (speaker) of the other terminal device being quiet and the environmental sound being loud. The terminal device 10 may be configured to calculate a correlation value while amplifying the second audio signal picked up by the audio pickup unit 16 of the terminal device 10 and adopt the maximum correlation value. In this case, a high correlation value indicates that the audio of the user (speaker) of the other terminal device is heard directly by the user of the terminal device 10 and is also heard as a delayed output from the audio output unit 20 of the terminal device 10. In this case, it is difficult to hear the voice of the user (speaker) of the other nearby terminal device, so it is better to mute the first audio signal and not output it. On the other hand, because the volume of the environmental sound included in the second audio signal is high, the correlation value may not be high even if the voice of the user (speaker) of the other terminal device is included. In this case, it is considered that the voice of the user (speaker) of the other terminal device is buried in the noise due to the high ambient noise, and the user of the own device cannot directly hear it. In this case, it is better not to mute the first audio signal. In other words, in a noisy room with high ambient noise (a lot of noise), a high correlation value cannot be obtained even if the terminal devices are installed close to each other, so it is better not to mute the first audio signal. Taking such situations into consideration, a predetermined threshold is set and stored in the storage unit 26. The predetermined threshold may be set by machine learning using a user evaluation questionnaire regarding ease of listening. Note that the correlation between audio signals may include the degree of agreement between audio signals.

[0033] After determining that the audio signal (first audio signal) received from the server device 100 should not be output from the audio output unit, if the correlation value between the audio signal (first audio signal) and the second audio signal is less than a predetermined threshold, the determination unit 28A outputs the audio signal from the audio output unit 20. As a result, if the volume of the speech of a nearby user decreases, the terminal device 10 can output the audio signal from the server device 100 from the audio output unit 20, thereby maintaining ease of hearing of the speech of nearby participants in the electronic conference.

[0034] (Sequence of electronic conference system) An example of a sequence of an audio signal in the SFU method of the electronic conference system according to the first embodiment will be described with reference to Fig. 4. Fig. 4 is a diagram showing an example of a sequence of the electronic conference system according to the first embodiment. In Fig. 4, in the electronic conference system 1, the terminal device 10A and the terminal device 10B are placed near each other in the same room, and the terminal device 10C is placed in a remote location. Each user of the terminal device 10A, the terminal device 10B, and the terminal device 10C is wearing a headset and participating in the electronic conference.

[0035] First, when the user of the terminal device 10C speaks, the terminal device 10C picks up the voice with the sound pickup unit 16, converts the picked-up voice signal MC into an audio stream, and starts transmitting it to the server device 100 (step S101).

[0036] The server device 100 transmits the audio signal MC received from the terminal device 10C as an audio stream to the terminal device 10A and the terminal device 10B (step S102). In this case, the server device 100 controls the destination of the audio signal MC.

[0037] Each of the terminal device 10A and the terminal device 10B outputs the audio signal MC received via the first communication unit 22 from the audio output unit 20. This allows the users of the terminal device 10A and the terminal device 10B to hear the speech of the user of the terminal device 10C from the audio output unit 20.

[0038] Next, when the user of terminal device 10B speaks, terminal device 10B picks up the voice with sound pickup unit 16, converts the picked-up voice signal MB into an audio stream, and starts transmitting it to server device 100 (step S103). In this case, the speech sound waves WB uttered by the user of terminal device 10B can be directly heard by the user of terminal device 10A located near terminal device 10B.

[0039] The server device 100 transmits the audio signal MB received from the terminal device 10B as an audio stream to the terminal device 10A and the terminal device 10C (step S104). In this case, the server device 100 controls the destinations of the audio signal MB and the audio signal MC, respectively.

[0040] The terminal device 10C outputs the audio signal MB received via the first communication unit 22 from the audio output unit 20. This allows the user of the terminal device 10C to hear the speech of the user of the terminal device 10B from the audio output unit 20.

[0041] Terminal device 10A receives a voice signal MB via first communication unit 22, collects a second voice signal including speech sound waves WB using sound collection unit 16, and determines whether the correlation value between voice signal MB and the second voice signal is greater than or equal to a predetermined threshold. If terminal device 10A determines that the correlation value between voice signal MB and the second voice signal is greater than or equal to the predetermined threshold, i.e., that the user of terminal device 10A can hear the speech sound waves WB, terminal device 10A mutes the output of voice signal MB to prevent it from being output from audio output unit 20 (step S105). As a result, the user of terminal device 10A can hear the speech sound waves WB spoken by the user of terminal device 10B without interference between the voice of voice signal MB and the speech sound waves WB, since the voice of voice signal MB is not output from audio output unit 20. If the user of terminal device 10C is speaking, audio output from audio output unit 20 of the user of terminal device 10C continues.

[0042] When the user of terminal device 10B finishes speaking, terminal device 10B ends transmission of the audio signal MB to server device 100 as an audio stream (step S106). As a result, when server device 100 finishes receiving the audio signal MB from terminal device 10B, it ends transmission of the audio signal MB to terminal device 10A and terminal device 10C (step S107). Then, when terminal device 10C finishes receiving the audio signal MB via first communication unit 22, it ends output of the audio of the user of terminal device 10B from audio output unit 20. Furthermore, when terminal device 10A finishes receiving the audio signal MB via first communication unit 22, it mutes the output of audio signal MB and ends audio output.

[0043] When the user of terminal device 10C finishes speaking, terminal device 10C ends transmission of the audio signal MC to server device 100 via an audio stream (step S108). As a result, when server device 100 finishes receiving the audio signal MC from terminal device 10C, it ends transmission of the audio signal MC to terminal device 10A and terminal device 10B (step S109). Then, when each of terminal device 10A and terminal device 10B finishes receiving the audio signal MC via first communication unit 22, it ends output of the voice of the user of terminal device 10C from audio output unit 20.

[0044] Thereafter, the electronic conference system 1 similarly transmits the voice signal uttered by the user from the terminal device 10 to the server device 100 , and the server device 100 distributes the received voice signal to the other terminal devices 10 .

[0045] Also, in step S105, terminal device 10A receives audio signal MB via first communication unit 22, collects a second audio signal including speech sound waves WB with sound collection unit 16, and, if it determines that the correlation value between audio signal MB and the second audio signal is less than a predetermined threshold and there is no correlation, outputs the received audio signal MB from audio output unit 20. As a result, the user of terminal device 10A hears audio signal MB spoken by the user of terminal device 10B without interference between the audio of audio signal MB and the audio of speech sound waves WB, because the audio of audio signal MB is output from audio output unit 20. At this time, if the user of terminal device 10C is speaking, the audio of both users of terminal device 10B and terminal device 10C is output from audio output unit 20.

[0046] (Processing Procedure of Terminal Device) An example of the processing procedure of the terminal device according to the first embodiment will be described with reference to Fig. 5. Fig. 5 is a flowchart showing the processing procedure of the terminal device according to the first embodiment. The processing procedure shown in Fig. 5 is realized by the control unit 28 of the terminal device 10 executing a program.

[0047] The control unit 28 of the terminal device 10 starts receiving an audio signal from the server device 100 (step S1001). Specifically, the control unit 28 starts acquiring an audio signal (first audio signal) of the audio stream received from the server device 100 via the first communication unit 22. After starting acquisition of the first audio signal in step S1001, the control unit 28 proceeds to S1002.

[0048] The control unit 28 determines whether or not speech sound waves have been collected (step S1002). For example, the control unit 28 determines that speech sound waves have been collected when speech sound waves are being collected by the sound collection unit 16. If the control unit 28 determines that speech sound waves have not been collected (step S1002; No), the control unit 28 proceeds to step S1004, which will be described later. On the other hand, if the control unit 28 determines that speech sound waves have been collected (step S1002; Yes), the control unit 28 proceeds to step S1003.

[0049] The control unit 28 determines the correlation between the first audio signal received from the server device 100 and the second audio signal collected by the audio collection unit 16 (step S1003). For example, the control unit 28 determines the correlation between the first audio signal and the second audio signal, and stores the determined correlation value in the storage unit 26 in association with the identifier of the first audio signal. A higher correlation value indicates a higher degree of match. When the process of step S1003 is completed, the control unit 28 proceeds to the process of S1004.

[0050] The control unit 28 determines whether the first audio signal has a correlation (step S1004). For example, if the correlation value of the first audio signal stored in the storage unit 26 in step S1003 is equal to or greater than a predetermined threshold, the control unit 28 determines that the first audio signal has a correlation. The predetermined threshold is a value set for determining the correlation value and is stored in the storage unit 26. Note that if it is determined in step S1002 that no speech sound waves are being collected (step S1002; No), the storage unit 26 does not store a correlation value, and therefore the first audio signal is not determined to have a correlation. In this case, for example, different thresholds may be set for a first threshold used to determine if the first audio signal is not muted and a second threshold used to determine if the first audio signal is already muted. By setting different thresholds, frequent switching between determining whether to mute and unmute the first audio signal can be prevented when the correlation value is near the threshold.

[0051] If the control unit 28 determines that the first audio signal has a correlation (step S1004; Yes), the control unit 28 proceeds to step S1005. The control unit 28 does not output the first audio signal that has a correlation with the speech sound waves from the audio output unit 20 (step S1005). That is, the control unit 28 mutes the output of the first audio signal that has a correlation with the speech sound waves. When the process of step S1005 ends, the control unit 28 proceeds to step S1007, which will be described later.

[0052] Furthermore, if the control unit 28 determines that there is no correlation with the first audio signal (step S1004; No), the control unit 28 proceeds to step S1006. The control unit 28 outputs the first audio signal from the audio output unit 20 (step S1006). For example, if the first audio signal has already been muted, the control unit 28 unmutes the first audio signal and causes it to be output from the audio output unit 20. When the process of step S1006 ends, the control unit 28 proceeds to step S1007.

[0053] The control unit 28 determines whether reception of the first audio signal has ended (step S1007). Specifically, the control unit 28 determines that reception of the first audio signal has ended if reception of the audio stream from the server device 100 has ended via the first communication unit 22. If reception of the first audio signal has not ended (step S1007; No), the control unit 28 continues receiving the audio stream from the server device 100, and returns to step S1002, as already described, and continues processing. This allows the control unit 28 to monitor changes in the state of the second audio signal during reception of the first audio signal, such as a change in the noise level in the room, and control the output of the first audio signal. Furthermore, if the control unit 28 determines that reception of the audio signal has ended (step S1007; Yes), it terminates the processing procedure shown in FIG. 5. When multiple audio streams are received simultaneously, the process of FIG. 5 is performed for each of the multiple audio streams, and it is determined whether or not to output each first audio signal from the audio output unit 20.

[0054] 5, the method of muting the output of the first audio signal correlated with the speech sound wave (second audio signal) in step S1005 has been described, but the present invention is not limited to this. For example, step S1005 in the processing procedure may be replaced with a process of lowering (suppressing) the output volume of the first audio signal correlated with the speech sound wave (second audio signal).

[0055] In the processing procedure shown in FIG. 5 , a determination process may be added immediately before step S1002 to determine whether the audio stream received from the server device 100 is an audio stream from a nearby terminal device 10. For example, identifiers of terminal devices 10 installed near the server device 100 may be stored in the storage unit 26, and the identifier of the terminal device 10 that is the sender included in the audio stream received from the server device 100 may be compared to determine whether the audio stream is from a nearby terminal device 10. Then, in the processing procedure shown in FIG. 5 , if it is determined that the received audio stream is an audio stream from a nearby terminal device 10, the processing from step S1002 onward is executed. Furthermore, in the processing procedure shown in FIG. 5 , if it is determined that the received audio stream is not an audio stream from a nearby terminal device 10, the processing from step S1002 onward is not executed. The first audio signal of the audio stream for which the processing from step S1002 onward is not executed is output from the audio output unit 20. By adding such a determination process, the processing procedure shown in FIG. 5 can improve the accuracy of the determination.

[0056] In the first embodiment, when another terminal device 10 is located near the terminal device 10 and a user is participating in a teleconference using a headset, the terminal device 10 receives a voice signal from the user using the other terminal device 10 from the server device 100. When the correlation value between the received voice signal (first voice signal) and the second voice signal is equal to or greater than a predetermined threshold, the terminal device 10 does not output the received voice signal from the voice output unit 20. This prevents the terminal device 10 from outputting a voice signal from a nearby participant participating in the teleconference within a distance where the voice signal can be directly heard, thereby making the voice of the nearby participant easier to hear. Using the correlation value of the voice signal makes it easier to hear the voice of the nearby participant than when controlling the output from the voice output unit 20 based solely on location information indicating whether the nearby participant is installed nearby. Furthermore, even if the voice output unit 20 is a speaker, the terminal device 10 can make the voice of the nearby participant easier to hear by not outputting the voice signal of the nearby participant. As a result, even if the terminal device 10 is used in the same electronic conference as a different terminal device 10 located nearby, the voices of other participants can be easily heard.

[0057] Second Embodiment (Electronic Conference System) An electronic conference system according to a second embodiment will be described with reference to Fig. 6. Fig. 6 is a diagram showing the flow of an audio signal and an identifier in the electronic conference system according to the second embodiment.

[0058] The electronic conference system 1 is a system for holding an electronic conference using the SFU method, and includes multiple terminal devices 10-1 and a server device 100. In the example shown in Figure 6, the electronic conference system 1 holds an electronic conference between adjacent terminal devices 10-1A and 10-1B and a remote terminal device 10-1C. The conference participants using the terminal devices 10-1A, 10-1B, and 10-1C participate in the electronic conference using headsets. In the following description, when there is no need to distinguish between the terminal devices 10-1A, 10-1B, and 10-1C, they will simply be referred to as the terminal device 10-1.

[0059] The electronic conference system 1 treats a series of signals such as data, audio, and images as a single stream. In the example shown in Fig. 6, the terminal device 10-1A adds its own identifier to a series of audio signals acquired by its microphone and transmits (uploads) the signals to the server device 100 as an audio stream STA. The terminal device 10-1B adds its own identifier to a series of audio signals acquired by its microphone and transmits the signals to the server device 100 as an audio stream STB. The terminal device 10-1C adds its own identifier to a series of audio signals acquired by its microphone and transmits the signals to the server device 100 as an audio stream STC.

[0060] The server device 100 transmits (transfers) to the terminal device 10-1A an audio stream STB including the identifier received from the terminal device 10-1B, and an audio stream STC including the identifier received from the terminal device 10-1C. The server device 100 transmits (transfers) to the terminal device 10-1B an audio stream STA including the identifier received from the terminal device 10-1A, and an audio stream STC received from the terminal device 10-1C. The server device 100 transmits (transfers) to the terminal device 10-1C an audio stream STA including the identifier received from the terminal device 10-1A, and an audio stream STB including the identifier received from the terminal device 10-1B. As a result, the terminal device 10-1A realizes a teleconference by outputting the audio signals of the audio stream STB and audio stream STC received from the server device 100. The terminal device 10-1B realizes a teleconference by outputting the audio signals of the audio stream STA and audio stream STC received from the server device 100. The terminal device 10-1C outputs the audio signals of the audio streams STA and STB received from the server device 100, thereby realizing an electronic conference.

[0061] Each of the terminal devices 10-1A, 10-1B, and 10-1C has a function of transmitting its own identifier to another terminal device 10-1 installed near the own device via short-range wireless communication. In the example shown in Figure 6, the terminal device 10-1A transmits its own identifier IDA to the terminal device 10-1B installed near the own device via short-range wireless communication. The terminal device 10-1B transmits its own identifier IDB to the terminal device 10-1A installed near the own device via short-range wireless communication.

[0062] (Terminal Device) An example of the configuration of the terminal device 10-1 according to the second embodiment will be described with reference to Fig. 7. Fig. 7 is a block diagram showing an example of the configuration of the terminal device 10-1 according to the second embodiment.

[0063] The terminal device 10-1 includes a camera 12, an operation unit 14, a sound collection unit 16, a display unit 18, an audio output unit 20, a first communication unit 22, a second communication unit 24, a memory unit 26, and a control unit 28.

[0064] The storage unit 26 stores, for example, identification information 26A, proximity information 26B, etc. The proximity information 26B includes information indicating a list of second identifiers related to other terminal devices 10-1 installed near the terminal device 10-1 (in the same room). The proximity information 26B indicates other terminal devices 10-1 installed nearby whose voices may be directly heard by the user of the terminal device 10-1, and indicates other terminal devices 10-1 that mute or lower the volume of the voices output from the audio output unit 20 of the terminal device 10-1. The proximity information 26B includes information indicating second identifiers related to other terminal devices 10 received via the second communication unit 24, second identifiers set by the administrator or user of the electronic conference, etc.

[0065] The control unit 28 includes a determination unit 28A and a transmission / reception unit 28B as functional blocks realized by the processing performed by the control unit 28.

[0066] When the second identifier is received by the second communication unit 24 from another terminal device 10-1 and stored in the proximity information 26B, the determination unit 28A of the control unit 28 determines whether the second identifier matches the first identifier related to the other terminal device 10-1 that transmitted the audio stream included in the audio stream received from the server device 100. If the first identifier matches the second identifier and the second identifier and the correlation value between the first audio signal and the second audio signal is equal to or greater than a predetermined threshold, the determination unit 28A determines not to output the audio signal included in the audio stream received from the server device 100 from the audio output unit 20, or to reduce the output volume of the audio signal.

[0067] The transmitter / receiver 28B transmits the identifier of its own device via the second communication unit 24 based on the identification information 26A in the storage unit 26. The transmitter / receiver 28B transmits the identifier of its own device to the outside, for example, at the timing when an audio signal of the voice spoken by the user of its own device and picked up by the sound pickup unit 16 is transmitted to the server device 100, at the timing when an electronic conference is started, etc. The transmitter / receiver 28B stores the identifier indicating the other terminal device 10 received via the second communication unit 24 in the proximity information 26B of the storage unit 26 as a second identifier.

[0068] (Sequence of electronic conference system) An example of a sequence of an audio signal in the SFU system of the electronic conference system 1 according to the second embodiment will be described with reference to Fig. 8. Fig. 8 is a diagram showing an example of a sequence of the electronic conference system 1 according to the second embodiment. In Fig. 8, in the electronic conference system 1, the terminal device 10-1A and the terminal device 10-1B are placed near each other in the same room, and the terminal device 10-1C is placed in a remote location. Each user of the terminal device 10-1A, the terminal device 10-1B, and the terminal device 10-1C is wearing a headset and participating in the electronic conference.

[0069] First, when the user of the terminal device 10-1C speaks, the terminal device 10-1C picks up the voice with the sound pickup unit 16, converts the picked-up voice signal MC into an audio stream, and starts transmitting it to the server device 100 (step S101).

[0070] The server device 100 transmits the audio signal MC received from the terminal device 10-1C as an audio stream to the terminal device 10-1A and the terminal device 10-1B (step S102). In this case, the server device 100 controls the destination of the audio signal MC.

[0071] Each of the terminal device 10-1A and the terminal device 10-1B outputs the audio signal MC received via the first communication unit 22 from the audio output unit 20. This allows the users of the terminal device 10-1A and the terminal device 10-1B to hear the speech of the user of the terminal device 10C from the audio output unit 20.

[0072] Next, when the user of terminal device 10-1B speaks while the user of terminal device 10-1C is speaking, terminal device 10-1B picks up the voice with sound pickup unit 16, converts the picked-up voice signal MB into an audio stream, and starts transmitting it to server device 100 (step S103). In this case, the speech sound waves WB uttered by the user of terminal device 10-1B can be directly heard by the user of terminal device 10-1A located near terminal device 10-1B. Therefore, terminal device 10-1B transmits its own identifier IDB from second communication unit 24 (step S110).

[0073] The server device 100 transmits the audio signal MB received from the terminal device 10-1B as an audio stream to the terminal device 10-1A and the terminal device 10-1C (step S104). In this case, the server device 100 controls the destinations of the audio signal MB and the audio signal MC, respectively.

[0074] The terminal device 10-1C outputs the voice signal MB received via the first communication unit 22 from the voice output unit 20. This allows the user of the terminal device 10-1C to hear from the voice output unit 20 the speech of the user of the terminal device 10B.

[0075] The terminal device 10-1A also receives an identifier IDB (second identifier) ​​from the terminal device 10-1B via the second communication unit 24 (step S110) and stores it in the proximity information 26B in the storage unit 26. The terminal device 10-1A then compares the identifier IDB (first identifier) ​​indicating the terminal device 10-1B included in the audio stream of the audio signal MB received from the server device 100 via the first communication unit 22 with the second identifier. If it is determined that both the first identifier and the second identifier are identifier IDB and match, the audio signal MB received from the server device 100 is not output from the audio output unit 20 (step S111). As a result, the user of the terminal device 10-1A can hear the speech sound waves WB uttered by the user of the terminal device 10-1B without interference between the audio of the audio signal MB and the speech sound waves WB, since the audio of the audio signal MB is not output from the audio output unit 20. At this time, if the user of the terminal device 10-1C is speaking, the voice output unit 20 continues to output the voice of the user of the terminal device 10-1C.

[0076] When the user of terminal device 10-1B finishes speaking, terminal device 10-1B ends transmission of the audio signal MB to server device 100 as an audio stream (step S106). Accordingly, when server device 100 finishes receiving the audio signal MB from terminal device 10-1B, it ends transmission of the audio signal MB to terminal device 10-1A and terminal device 10-1C (step S107). When terminal device 10-1C finishes receiving the audio signal MB via first communication unit 22, it ends output of the audio of the user of terminal device 10B from audio output unit 20. When terminal device 10-1A finishes receiving the audio signal MB via first communication unit 22, it mutes the output of audio signal MB and ends audio output.

[0077] When the user of terminal device 10-1C finishes speaking, terminal device 10-1C ends transmission of the audio signal MC to server device 100 via an audio stream (step S108). As a result, when server device 100 finishes receiving the audio signal MC from terminal device 10-1C, it ends transmission of the audio signal MC to terminal device 10-1A and terminal device 10-1B (step S109). Then, when each of terminal device 10-1A and terminal device 10-1B finishes receiving the audio signal MC via first communication unit 22, it ends output of the voice of the user of terminal device 10-1C from audio output unit 20.

[0078] Thereafter, in the electronic conference system 1, the terminal device 10-1 transmits a voice signal uttered by the user to the server device 100 in the same manner, and the server device 100 distributes the received voice signal to the other terminal devices 10-1.

[0079] Furthermore, in step S111, terminal device 10-1A receives an identifier IDB (second identifier) ​​from terminal device 10-1B via second communication unit 24, and if the identifier IDB (first identifier) ​​indicating terminal device 10-1B contained in the audio stream of audio signal MB received from server device 100 via first communication unit 22 does not match, terminal device 10-1A outputs audio signal MB received from server device 100 from audio output unit 20. This allows the user of terminal device 10-1A to hear the voice of the user of terminal device 10-1B output from audio output unit 20.

[0080] (Processing Procedure of Adjacent Terminal Device) An example of the processing procedure of the terminal device 10-1 installed adjacently according to the second embodiment will be described using Figures 9 and 10. Figure 9 is a flowchart showing the processing procedure when the terminal device 10-1 according to the second embodiment transmits an audio signal. Figure 10 is a flowchart showing the processing procedure when the terminal device according to the second embodiment receives an audio signal. The processing procedures shown in Figures 9 and 10 are realized by the control unit 28 of the terminal device 10-1 executing a program.

[0081] 9, the control unit 28 of the terminal device 10-1 transmitting the audio signal determines whether the first communication unit 22 is currently transmitting the audio signal (step S1101). For example, the control unit 28 of the terminal device 10-1B determines that the first communication unit 22 is currently transmitting the audio signal when the audio stream of the audio signal MB is being transmitted to the server device 100 via the first communication unit 22. If the control unit 28 determines that the first communication unit 22 is not currently transmitting the audio signal (step S1101; No), the control unit 28 proceeds to step S1103, which will be described later. If the control unit 28 determines that the first communication unit 22 is currently transmitting the audio signal (step S1101; Yes), the control unit 28 proceeds to step S1102.

[0082] The control unit 28 transmits the identifier of the own device through the second communication unit 24 (step S1102). For example, the control unit 28 generates a transmission signal including an identifier IDB indicating the own device based on the identification information 26A in the storage unit 26, and transmits the transmission signal from the second communication unit 24. When the process of step S1102 ends, the control unit 28 proceeds to step S1103.

[0083] The control unit 28 determines whether or not to terminate the process (step S1103). For example, the control unit 28 determines to terminate the process when the transmission signal transmitted by the control unit 28 in step S1102 is received by the other terminal device 10-1 and a response to the reception of the identifier is returned from the other terminal device 10-1, or when a termination condition is met, such as the identifier being repeatedly transmitted a predetermined number of times. If the control unit 28 determines not to terminate the process (step S1103; No), the control unit 28 returns the process to step S1101 already described and continues the process. On the other hand, if the control unit 28 determines to terminate the process (step S1103; Yes), the control unit 28 terminates the processing procedure shown in FIG. 9.

[0084] 10, the control unit 28 of the terminal device 10-1 receiving the voice signal determines whether the second communication unit 24 has received an identifier (step S1201). For example, the control unit 28 determines that the second communication unit 24 has received an identifier when a signal including the identifier is received via the second communication unit 24, or when the identifier is registered in the proximity information 26B of the storage unit 26. If the control unit 28 determines that the second communication unit 24 has not received an identifier (step S1201; No), the control unit 28 proceeds to step S1205, which will be described later. If the control unit 28 determines that the second communication unit 24 has received an identifier (step S1201; Yes), the control unit 28 proceeds to step S1202.

[0085] The control unit 28 determines whether or not speech sound waves are being collected (step S1202). For example, when sound waves are being collected by the sound collection unit 16, the control unit 28 determines that speech sound waves are being collected. When the control unit 28 determines that speech sound waves are not being collected (step S1202; No), the control unit 28 proceeds to step S1205, which will be described later. When the control unit 28 determines that speech sound waves are being collected (step S1202; Yes), the control unit 28 proceeds to step S1203.

[0086] The control unit 28 determines whether the identifiers match and whether there is a correlation between the audio signals (step S1203). For example, the control unit 28 determines that the identifiers match when the identifier of the other terminal device 10-1 received by the second communication unit 24 matches an identifier included in the audio stream of the audio signal received by the first communication unit 22, or when the identifier included in the audio stream of the audio signal received by the first communication unit 22 matches an identifier registered in the proximity information 26B of the storage unit 26. For example, the control unit 28 calculates a correlation value between the first audio signal received by the first communication unit 22 and the second audio signal collected by the audio collection unit 16, and determines that there is a correlation between the audio signals if the correlation value is equal to or greater than a predetermined threshold. If the control unit 28 determines that the identifiers match and there is a correlation between the audio signals (step S1203; Yes), the control unit 28 proceeds to step S1204.

[0087] The control unit 28 mutes the audio signal of the audio stream so that it is not output from the audio output unit 20 (step S1204). After completing the process of step S1204, the control unit 28 advances the process to step S1206, which will be described later.

[0088] Furthermore, if the control unit 28 determines in step S1203 that the identifiers match and that there is no correlation between the audio signals (step S1203; No), the control unit 28 proceeds to step S1205. The control unit 28 outputs the audio signal of the audio stream from the audio output unit 20 (step S1205). For example, if the output of the audio signal has been muted, the control unit 28 cancels the mute and causes the audio signal to be output from the audio output unit 20. When the process of step S1205 ends, the control unit 28 proceeds to step S1206. As a result, if the state of the second audio signal changes while the first audio signal is being received, for example, if the noise level in the room changes, the control unit 28 can monitor the state change as needed and control the output of the first audio signal.

[0089] The control unit 28 determines whether to terminate the process (step S1206). For example, the control unit 28 determines to terminate the process when a termination condition is met, such as when reception of the target audio stream is terminated or when a termination request from the user is received. If the control unit 28 determines not to terminate the process (step S1206; No), the control unit 28 returns the process to the already-described step S1201 and continues the process. On the other hand, if the control unit 28 determines to terminate the process (step S1206; Yes), the control unit 28 terminates the processing procedure shown in Fig. 10. When multiple audio streams are received simultaneously, the process of Fig. 10 is performed for each of the multiple audio streams, and it is determined whether or not to output each first audio signal from the audio output unit 20.

[0090] 10, the order of steps S1201 and S1202 may be reversed. That is, the processing procedure may be changed so that first, the "collection of speech sound waves" in the vicinity of the device itself is determined, and then the "reception of an identifier by the second communication unit" is determined.

[0091] In the second embodiment, when another terminal device 10-1 is located near the terminal device 10-1 and a user is participating in a teleconference using a headset, the terminal device 10-1 receives a voice signal from the user using the other terminal device 10-1 from the server device 100. When the correlation value between the received voice signal (first voice signal) and the second voice signal is equal to or greater than a predetermined threshold, and the identifier received by the second communication unit 24 matches, the terminal device 10-1 does not output the received voice signal from the voice output unit 20. This prevents the terminal device 10-1 from outputting the voice signal of a nearby participant participating in the teleconference within a distance where the voice wave can be directly heard from the voice output unit 20, making it easier to hear the voice of the nearby participant. Furthermore, by receiving the identifier via the second communication unit 24, the terminal device 10-1 can time the timing for collecting ambient sound and comparing it with the voice signal received by the first communication unit 22. When the terminal device 10-1 simultaneously receives multiple audio streams via the first communication unit 22, receiving an identifier via the second communication unit 24 makes it easier to identify the audio stream to be compared with the audio signal that captures ambient sound to calculate a correlation value. Furthermore, even if the audio output unit 20 is a speaker, the terminal device 10-1 can make the audio of nearby participants easier to hear by not outputting the audio signal of the nearby participants. As a result, even when the terminal device 10-1 is used in the same electronic conference as a different terminal device 10-1 located nearby, the audio of the other participants can be made easier to hear.

[0092] [Third embodiment] (Electronic conference system) An electronic conference system according to a third embodiment will be described with reference to Fig. 11. Fig. 11 is a diagram showing the flow of an audio signal in the electronic conference system according to the third embodiment.

[0093] The electronic conference system 1 shown in Fig. 11 is a system that holds electronic conferences using the MCU (Multi-point Control Unit) method. The electronic conference system 1 includes multiple terminal devices 10-1 and a server device 100-1. The MCU method is a method in which audio signals from devices other than the server device 100-1 are synthesized and sent to the terminal device 10-1 as a single audio stream.

[0094] The electronic conference system 1 treats a series of signals such as data, audio, and images as a single stream. In the example shown in FIG. 11 , the terminal device 10-1A adds its own identifier and the like to a series of audio signals acquired by its microphone and transmits (uploads) the result as an audio stream STA to the server device 100-1. The terminal device 10-1B adds its own identifier and the like to a series of audio signals acquired by its microphone and transmits the result as an audio stream STB to the server device 100-1. The terminal device 10-1C adds its own identifier and the like to a series of audio signals acquired by its microphone and transmits the result as an audio stream STC to the server device 100-1.

[0095] The server device 100-1 transmits to the terminal device 10-1A an audio signal obtained by combining the audio signal of the audio stream STB received from the terminal device 10-1B and the audio signal of the audio stream STC received from the terminal device 10-1C as the audio stream STBC. The server device 100-1 transmits to the terminal device 10-1B an audio signal obtained by combining the audio signal of the audio stream STA received from the terminal device 10-1A and the audio signal of the audio stream STC received from the terminal device 10-1C as the audio stream STAC. The server device 100-1 transmits to the terminal device 10-1C an audio signal obtained by combining the audio signal of the audio stream STA received from the terminal device 10-1A and the audio signal of the audio stream STB received from the terminal device 10-1B as the audio stream STAB. As a result, the terminal device 10-1A realizes a teleconference by outputting the audio signal of the audio stream STBC received from the server device 100. The terminal device 10-1B realizes the electronic conference by outputting the audio signal of the audio stream STAC received from the server device 100. The terminal device 10-1C realizes the electronic conference by outputting the audio signal of the audio stream STAB received from the server device 100.

[0096] The server device 100-1 further has a function of synthesizing the audio of the streams to be transmitted to each terminal device 10-1 based on the mute lists received from each terminal device 10-1. For example, in the example shown in FIG. 11 , the server device 100-1 stores the mute lists received from each terminal device 10-1 in the mute information. The mute list is information that lists, for example, the identifiers of the terminal devices 10-1 for which audio signal synthesis is to be suppressed (no synthesis or synthesis with reduced volume). For example, if the mute list received from the terminal device 10-1A includes the identifier of the terminal device 10-1B, the server device 100-1 does not synthesize (mute) the audio signal of the audio stream STB from the terminal device 10-1B, and transmits to the terminal device 10-1A an audio stream STBC containing only the audio signal of the audio stream STC from the terminal device 10-1C. For example, if the mute list received from terminal device 10-1B contains the identifier of terminal device 10-1A, server device 100-1 does not synthesize (mute) the audio signal of audio stream STA from terminal device 10-1A to terminal device 10-1B, but transmits audio stream STAC containing only the audio signal of audio stream STC from terminal device 10-1C.

[0097] (Server Device) A configuration example of the server device 100-1 according to the third embodiment will be described with reference to Fig. 12. Fig. 12 is a diagram showing a configuration example of the server device 100-1 according to the third embodiment.

[0098] The server device 100-1 acquires a series of video, audio, and other streams from multiple terminal devices 10-1 and distributes them to participants of the electronic conference. The server device 100 may be realized, for example, by the terminal device 10-1 of the organizer of the electronic conference.

[0099] The server device 100-1 includes a communication unit 102, a storage unit 104, and a control unit .

[0100] The communication unit 102 is a communication interface that performs communication between the server device 100-1 and an external device. The communication unit 102 executes communication between the server device 100-1 and the terminal device 10-1, for example. The communication unit 102 is realized by, for example, a wireless LAN or a wired LAN.

[0101] The storage unit 104 stores various types of information, such as the contents of calculations performed by the control unit 106 and information such as programs. The storage unit 104 includes at least one of, for example, a RAM, a main storage device such as a ROM, and an external storage device such as an HDD.

[0102] The storage unit 104 stores mute information 200, which includes a mute list. The mute information 200 includes information that identifies the other terminal devices 10-1 to be muted in the audio signals transmitted to each terminal device 10-1. In the example shown in FIG. 12, the mute information 200 includes information such as mute lists received from multiple terminal devices 10-1 and the identifiers of the terminal devices 10-1 that transmitted the mute lists. The mute information 200 associates each mute list with the identifiers of the terminal devices 10-1 that transmitted the mute lists. For example, the mute list contains the identifiers of the terminal devices 10-1 that are located near one or more terminal devices 10-1 that transmitted the mute lists to be muted.

[0103] The control unit 106 controls each unit of the server device 100. The control unit 106 includes, for example, an information processing device such as a CPU or an MPU, and a storage device such as a RAM or a ROM. The control unit 106 executes a program that controls the operation of the server device 100 according to the present invention. The control unit 106 may be realized by an integrated circuit such as an ASIC or an FPGA. The control unit 106 may be realized by a combination of hardware and software.

[0104] The control unit 106 includes an information acquisition unit 110 and a stream control unit 112 as functional blocks realized by processing by the control unit 106. Each of these functional blocks may be realized by the control unit 28 of the terminal device 10, or may be realized by sharing the processing between the control unit 106 and the control unit 28 of the terminal device 10-1.

[0105] The information acquisition unit 110 acquires a mute list indicating which audio streams of the electronic conference are to be muted from the multiple terminal devices 10-1. The information acquisition unit 110 creates and updates the mute information 200 based on the acquired mute list. For example, the information acquisition unit 110 acquires a mute list from the terminal device 10-1 via the communication unit 102, creates the mute information 200 based on the acquired mute list, and stores the mute information 200 in the storage unit 104. Note that the information acquisition unit 110 may also acquire the mute information 200 corresponding to the electronic conference from a database or the like and store it in the storage unit 104.

[0106] The stream control unit 112 controls the transmission of an audio stream of an audio signal obtained by synthesizing audio signals of audio streams from multiple terminal devices 10-1 (second terminal devices) different from the terminal device 10-1 (first terminal device) that is the transmission target. The stream control unit 112 excludes audio streams of second terminal devices indicated in the mute list corresponding to the terminal device 10-1 (first terminal device) from among the multiple terminal devices 10-1 (second terminal devices) and transmits a synthesized audio stream obtained by synthesizing audio streams of the remaining second terminal devices to the first terminal device. In other words, when the stream control unit 112 receives audio streams from one or more terminal devices 10-1, it mutes the audio signals of the audio streams from the terminal device 10-1 indicated in the mute list of the terminal device 10-1 that is the transmission target, and transmits an audio stream (synthesized audio stream) obtained by synthesizing audio signals of audio streams from terminal devices 10-1 other than the transmission target to the terminal device 10-1 that is the transmission target. Furthermore, if the stream control unit 112 receives audio streams from multiple terminal devices 10-1 but does not receive an audio stream from the terminal device 10-1 indicated in the mute list of the terminal device 10-1 to which the audio stream is to be sent, or if the mute list of the terminal device 10-1 to which the audio stream is to be sent is not registered in the mute information 200, the stream control unit 112 transmits an audio stream that combines the audio signals from all terminal devices 10-1 other than the terminal device 10-1 to which the audio stream is to be sent to the terminal device 10-1 to which the audio stream is to be sent.

[0107] (Terminal device) Similar to the terminal device 10-1 shown in FIG. 7, the terminal device 10-1 includes a camera 12, an operation unit 14, a sound collection unit 16, a display unit 18, an audio output unit 20, a first communication unit 22, a second communication unit 24, a memory unit 26, and a control unit 28.

[0108] The storage unit 26 stores, in the neighborhood information 26B, a mute list including identifiers of audio streams of other terminal devices 10-1 that the determination unit 28A has determined not to output audio from the audio output unit 20. The mute list may include identifiers of nearby terminal devices 10-1 received by the second communication unit 24.

[0109] The control unit 28 transmits the mute list data to the teleconference server device 100-1 via the first communication unit 22 and controls the server device 100-1 to request that the server device 100-1 not transmit audio streams containing audio signals of audio streams with identifiers included in the mute list. When the teleconference system 1 uses the MCU system, the control unit 28 controls the transmission of the mute list in the proximity information 26B to the server device 100-1 via the first communication unit 22. The timing of transmitting the mute list includes, for example, the start of the teleconference, the transmission and reception of audio streams, and the detection of another terminal device 10-1 that the determination unit 28A has determined not to output audio from the audio output unit 20 or to output audio from the audio output unit 20.

[0110] (Processing Procedure of Server Device) An example of the processing procedure of the server device 100-1 according to the third embodiment will be described with reference to Fig. 13. Fig. 13 is a flowchart showing the processing procedure of the server device 100-1 according to the third embodiment. The processing procedure shown in Fig. 13 is realized by the control unit 106 of the server device 100-1 executing a program.

[0111] The control unit 106 of the server device 100-1 determines whether or not a mute list has been acquired (step S2001). For example, the control unit 106 determines that a mute list has been acquired if the mute list of the terminal device 10-1 has been acquired from the terminal device 10-1, a database, or the like. If the control unit 106 determines that a mute list has not been acquired (step S2001; No), the control unit 106 proceeds to step S2003, which will be described later. If the control unit 106 determines that a mute list has been acquired (step S2001; Yes), the control unit 106 proceeds to step S2002.

[0112] The control unit 106 sets the acquired mute list in the mute information 200 (step S2002). For example, if the acquired mute list is not registered, the control unit 106 newly sets the mute list in the mute information 200 in the storage unit 104. If the acquired mute list is already registered, the control unit 106 updates the mute information 200 in the storage unit 104. After completing the process of step S2002, the control unit 106 proceeds to step S2003.

[0113] The control unit 106 determines whether to transmit an audio stream (step S2003). For example, the control unit 106 determines whether to transmit the audio signal of the audio stream received via the communication unit 102 to the terminal device 10-1 as an audio stream. If the control unit 106 determines not to transmit an audio stream (step S2003; No), the control unit 106 proceeds to step S2008, which will be described later. If the control unit 106 determines to transmit an audio stream (step S2003; Yes), the control unit 106 proceeds to step S2004.

[0114] The control unit 106 determines whether or not a mute list to be transmitted exists (step S2004). For example, the control unit 106 determines that a mute list to be transmitted exists when it extracts a mute list corresponding to the identifier of the terminal device 10-1 to be transmitted from the mute information 200 stored in the storage unit 104. If the control unit 106 determines that a mute list to be transmitted exists (step S2004; Yes), the control unit 106 proceeds to step S2005.

[0115] The control unit 106 excludes the audio stream received from the terminal device 10-1 indicated in the mute list and combines the audio signals of the audio streams received from the other terminal devices 10-1 (step S2005). For example, if the control unit 106 is receiving multiple audio streams, it excludes the audio stream from the terminal device 10-1 indicated in the mute list from the multiple audio streams and combines the audio signals of the audio streams from the other terminal devices 10-1 using publicly known voice synthesis software or the like. For example, if the control unit 106 is receiving multiple audio streams but the audio stream from the terminal device 10-1 indicated in the mute list is not included in the multiple audio streams being received, it combines the audio signals of all the audio streams from the terminal devices 10-1 being received using publicly known voice synthesis software or the like. After completing the process of step S2005, the control unit 106 proceeds to step S2007, which will be described later.

[0116] If the control unit 106 determines that there is no mute list to which transmission is to be made (step S2004; No), the control unit 106 proceeds to step S2006. The control unit 106 synthesizes the audio signals of all the received audio streams (step S2006). For example, the control unit 106 synthesizes the audio signals of the audio streams received from all the terminal devices 10-1 using publicly known voice synthesis software or the like. After completing step S2006, the control unit 106 proceeds to step S2007, which will be described later.

[0117] The control unit 106 transmits the synthesized audio stream to the terminal device 10-1 that is the transmission target (step S2007). For example, the control unit 106 controls the communication unit 102 to transmit the synthesized audio stream to the terminal device 10-1 that is the transmission target. When the process of step S2007 ends, the control unit 106 proceeds to step S2008.

[0118] The control unit 106 determines whether to terminate (step S2008). For example, the control unit 106 determines to terminate when a termination condition is met, such as the termination of the electronic conference or the receipt of a termination request from a user. If the control unit 106 determines not to terminate (step S2008; No), the control unit 106 returns the process to step S2001, which has already been described, and continues the process. If the control unit 106 determines to terminate (step S2008; Yes), the control unit 106 terminates the processing procedure shown in FIG. 13.

[0119] The server device 100-1 according to the third embodiment can exclude the audio stream received from the terminal device 10-1 indicated in the mute list corresponding to the terminal device 10-1 to which the audio is to be sent, and transmit a synthesized audio stream obtained by synthesizing the audio signals of the audio streams of the other terminal devices 10-1 to the terminal device 10-1 to which the audio is to be sent. By acquiring the mute list, the server device 100-1 can prevent the audio signal of a nearby participant from being output from the audio output unit 20 of a terminal device 10-1 that is located within a distance where the speech sound waves of the nearby participant participating in the same electronic conference can be directly heard, thereby making the audio of the nearby participant easier to hear. Furthermore, even if the audio output unit 20 of the terminal device 10-1 is a speaker, the server device 100-1 can prevent the audio signal of the nearby participant from being output, making the audio of the nearby participant easier to hear.

[0120] [Fourth embodiment] (Processing procedure of terminal device) An example of the processing procedure of a terminal device according to the fourth embodiment will be described using Fig. 14. Fig. 14 is a flowchart showing the processing procedure of a terminal device according to the fourth embodiment. The configuration of the terminal device according to the fourth embodiment is the same as the configuration of the terminal device 10, so the description will be omitted.

[0121] The process shown in FIG. 14 is the same as the process shown in FIG. 5 except for the processes in steps S1003A, S1004A, and S1010, and therefore a description thereof will be omitted.

[0122] In step S1003A, the control unit 28 determines the correlation between the first audio signal received from the server device 100 and the second audio signal collected by the audio collection unit 16 (step S1003A). For example, the control unit 28 determines the correlation between the first audio signal and the second audio signal, and stores the determined correlation value in the storage unit 26 in association with the identifier of the first audio signal. At this time, the control unit 28 calculates the correlation value for a predetermined time interval (correlation period), and stores the time at which the correlation value was calculated in association with the identifier of the first audio signal in the storage unit 26. For example, if the correlation period is 5 seconds, the control unit 28 calculates the correlation value of the audio signal for 5 seconds, and stores the time in association with the identifier of the first audio signal in the storage unit 26. A known autocorrelation function may be used to calculate the correlation value. For example, the autocorrelation function may be, but is not limited to, a power spectrum calculated by Fourier transforming the audio signal. When the process of step S1003 ends, the control unit 28 advances the process to S1010.

[0123] In step S1010, the control unit 28 determines whether the first audio signal is silent (step S1010). Specifically, the control unit 28 detects whether the first audio signal is currently silent. Silence refers to, for example, the sound pressure in the frequency range of human voice being equal to or lower than a predetermined sound pressure. For example, the control unit 28 detects the first audio signal as silent if the sound pressure in the frequency range of human voice is equal to or lower than a predetermined sound pressure continuously for a predetermined period (silence period). For example, the control unit 28 may detect the first audio signal as silent if a state in which human voice cannot be detected using known voice recognition technology continues for a silent period. If the first audio signal is determined to be silent (step S1010; Yes), the control unit 28 proceeds to step S1004A. If the first audio signal is not determined to be silent (step S1010; No), the control unit 28 returns to step S1002 and continues the correlation determination process. As a result, the output state of the first audio signal is not changed, and the state in which the first audio signal is output or not output from the audio output unit continues.

[0124] In step S1004A, the control unit 28 determines whether the first audio signal has a correlation (step S1004A). Specifically, the control unit 28 refers to a correlation value associated with a time prior to the silent period from the current time, and determines that the first audio signal has a correlation if the correlation value is equal to or greater than a predetermined value. If it is determined that the first audio signal has a correlation (step S1004A; Yes), the control unit 28 proceeds to step S1005. If it is not determined that the first audio signal has a correlation (step S1004A; No), the control unit 28 proceeds to step S1006.

[0125] In the fourth embodiment, the terminal device 10 switches the audio output of the first audio signal between on and off depending on whether the correlation value between the first audio signal and the second audio signal is equal to or greater than a predetermined value, starting from the point in time when it detects that the first audio signal is silent. As a result, in the fourth embodiment, the audio output is not switched between on and off while the user is speaking, and therefore it is possible to prevent the audio of the first audio signal from being output or not being output while the first audio signal is being spoken, and thus to prevent the user from finding the audio of the first audio signal unpleasant to the ear.

[0126] [Other Embodiments] The terminal device 10 and the terminal device 10-1 described above may be configured to determine whether to mute the audio stream based only on the identifier. That is, the terminal device 10 and the terminal device 10-1 may include a first communication unit 22 that transmits and receives an audio stream of an electronic conference, a second communication unit 24 that wirelessly transmits and receives identifiers of other electronic conference terminal devices, an audio pickup unit 16 that picks up audio, an audio output unit 20 that outputs the audio of the audio stream, and a determination unit 28A that determines which audio streams should not be output from the audio output unit 20. The determination unit 28A may receive a second identifier from another electronic conference terminal device via the second communication unit 24, acquire a first audio signal included in the audio stream received by the first communication unit 22, and, if it determines that the second identifier matches a preset first identifier, not output the first audio signal from the audio output unit 20 or reduce the volume of the first audio signal. Furthermore, the electronic conference terminal device may be equipped with a user interface that displays icons of participants in the electronic conference on the display unit 18 based on the identifier of the electronic conference terminal device that is transmitting the first audio signal received by the first communication unit 22, and accepts operations to prevent the selected first audio signal from being output from the audio output unit 20 or to lower the volume of the first audio signal by selecting an icon with the mouse of the operation unit 14.

[0127] The components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown, and all or part of the devices can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc. Note that this distribution and integration configuration may also be performed dynamically.

[0128] Although the embodiments of the present invention have been described above, the present invention is not limited to the contents of these embodiments. Furthermore, the above-described components include those that can be easily imagined by a person skilled in the art, those that are substantially the same, and those that are within the so-called equivalent range. Furthermore, the above-described components can be combined as appropriate. Furthermore, various omissions, substitutions, or modifications of the components can be made without departing from the spirit of the above-described embodiments.

[0129] The voice output determination device, electronic conference terminal device, electronic conference server device, and information processing method of the present disclosure can be used, for example, in an electronic conference system for holding online conferences.

[0130] 1 Electronic conference system 10, 10-1 Terminal device 12 Camera 14 Operation unit 16 Sound collection unit 18 Display unit 20 Audio output unit 22 First communication unit 24 Second communication unit 26 Storage unit 26A Identification information 26B Proximity information 28 Control unit 28A Determination unit 28B Transmitting / receiving unit 100, 100-1 Server device 102 Communication unit 104 Storage unit 106 Control unit 110 Information acquisition unit 112 Stream control unit 200 Mute information

Claims

1. An audio output determination device comprising: a first communication unit that transmits and receives an audio stream; a sound collection unit that collects audio; an audio output unit that outputs the audio of the audio stream; and a determination unit that determines which of the audio streams should not be output from the audio output unit, wherein the determination unit collects a second audio signal that is ambient sound using the sound collection unit, acquires a first audio signal included in the audio stream received by the first communication unit, and determines that the first audio signal should not be output from the audio output unit or that the output of the first audio signal should be reduced if the correlation value between the first audio signal and the second audio signal is equal to or greater than a predetermined threshold.

2. The audio output determination device according to claim 1, wherein the determination unit determines whether the first audio signal is silent or not, and if the first audio signal is silent and the correlation value between the first audio signal and the second audio signal is equal to or greater than a predetermined threshold, determines not to output the first audio signal from the audio output unit or to reduce the output of the first audio signal.

3. An electronic conference terminal device equipped with the audio output determination device described in claim 1 or 2, further comprising a second communication unit that wirelessly transmits and receives a second identifier that is identification information related to another electronic conference terminal device, wherein when the second communication unit receives the second identifier related to at least one other electronic conference terminal, if the first identifier related to the other electronic conference terminal that transmitted the first audio signal matches the second identifier and the correlation value between the first audio signal and the second audio signal is equal to or greater than the predetermined threshold, the determination unit does not output the first audio signal from the audio output unit.

4. The electronic conference terminal device of claim 3, wherein even if the judgment unit has judged that the audio of the first audio signal should not be output from the audio output unit, when the correlation value between the first audio signal and the second audio signal becomes less than the predetermined threshold, the judgment unit changes to a judgment that the audio of the first audio signal should be output from the audio output unit.

5. The electronic conference terminal device according to claim 3, further comprising a memory unit that stores a mute list capable of identifying the audio streams for which the determination unit has determined that audio should not be output from the audio output unit, and the first communication unit transmits data of the mute list to a server device of the electronic conference and requests the server device not to transmit the audio streams indicated by the mute list.

6. An electronic conference server device comprising: an information acquisition unit that acquires a mute list indicating audio streams to be muted from among audio streams of an electronic conference received from a plurality of terminal devices; and a stream control unit that controls transmission to the first terminal device of a synthesized audio stream obtained by synthesizing the audio streams from a plurality of second terminal devices different from the first terminal device to which the audio stream is to be transmitted; wherein the stream control unit excludes the audio streams of the second terminal devices indicated in the mute list corresponding to the first terminal device among the plurality of second terminal devices, and transmits to the first terminal device the synthesized audio stream obtained by synthesizing the audio streams of the remaining second terminal devices.

7. An information processing method executed by an audio output determination device having a first communication unit that transmits and receives an audio stream, a sound collection unit that collects audio, and an audio output unit that outputs the audio of the audio stream, the information processing method including the steps of: collecting a second audio signal, which is ambient sound, by the sound collection unit; acquiring the first audio signal included in the audio stream received by the first communication unit; and, if the correlation value between the first audio signal and the second audio signal is equal to or greater than a predetermined threshold, not outputting the first audio signal from the audio output unit or reducing the output of the first audio signal.

Citation Information

Patent Citations

  • System and method for muting audio associated with a source

    US20130044893A1

  • Chat terminal, chat system, and method for controlling chat system

    WO2024004006A1