Voice transmission and reception system

JP7909677B2Active Publication Date: 2026-08-21MAXELL LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025201738
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-11-21
Publication Date
2026-08-21
Estimated Expiration
2042-06-28

AI Technical Summary

Benefits of technology

【0010】 本発明によれば、同一の拠点から複数の参加者が各々のチャット端末を利用してチャットに参加する際に、近くにいる他の参加者が発した他者音声が配信音声と音声干渉をして聴きづらくなる不具合を解消することができる。上記した以外の目的、構成、効果については以下の実施形態において明らかにされる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007909677000001
    Figure 0007909677000001
  • Figure 0007909677000002
    Figure 0007909677000002
  • Figure 0007909677000003
    Figure 0007909677000003
Patent Text Reader

Abstract

To solve the problem that another person's voice uttered by another nearby participant interferes with a distributed voice to make it hard to hear.SOLUTION: An audio transmission and reception system (100), in which an information terminal (3) includes a microphone (12), a communication unit (15), an audio output unit (14), and a processor (16), the processor connecting to a new information terminal when finding the new information terminal by near field communication in a situation in which the information terminal is already connected to another information terminal by near field communication, and transmitting information relating to the new information terminal to the other information terminal, A distribution prohibition list (62) for prohibiting spoken voice collected by all other information terminals connected by short-range communication and a new information terminal searched by the other information terminals from being included in distribution voice is created and transmitted to a server (5), and distribution voice after voice reduction processing of the spoken voice based on the distribution prohibition list is received and output from a voice output device.SELECTED DRAWING: Figure 2A
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a voice transmission and reception system.

Background Art

[0002] Using a chat system implemented in a web conferencing system for business purposes etc., voice data is transmitted and received between remote locations and a remote conference is held. In the past, a chat was conducted in a form where one chat terminal per site shared the screen and voice with a plurality of participants within the site. However, in recent years, chat applications executed on personal computers and smartphones have come to be used, and a chat is conducted in a form where participants execute the chat application on their respective chat terminals even within the same site.

[0003] Regarding the voice processing of conferences between sites, there is a description in Patent Document 1 (Japanese Patent Laid-Open No. 8-237627). Patent Document 1 discloses a multipoint video conferencing system aimed at "preventing the speaker's own voice from being heard on the speaker's terminal (summary excerpt)". According to this multipoint video conferencing system, the speech voice is not distributed to and not output from the speaker's side terminal.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] On the other hand, if participants join the meeting using their own chat terminals even at the same location, the voice spoken by participant A for the remote meeting (referred to as an inter-location meeting) is picked up by the microphone on participant A's chat terminal, sent to the chat server, and then distributed to the chat terminals of participants at other locations and to the chat terminals of other participants at the same location (for example, participant B), and output through the chat terminal's speaker or earphones. As a result, participant B will hear both participant A's voice (other people's voices) directly and the distributed voice output from participant B's chat terminal. Hereafter in this specification, "other people's voices" refers to the voices of other people (people other than the chat terminal users) speaking at the same location that can be heard directly. "Distributed voices" refers to the voices output from the chat terminal.

[0006] Because the streamed audio has a delay due to passing through the chat server, when two audio streams overlap, the same audio is played twice with a time difference, resulting in a problem where it becomes very difficult to hear.

[0007] While Patent Document 1 describes how to prevent the speaker from hearing their own voice on their device, it does not address the issue of audio interference between other users' voices and the streamed audio that occurs when multiple chat devices are located at the same location, and therefore cannot solve the above problem.

[0008] This invention was made in view of the above points, and its purpose is to resolve the problem that occurs when multiple participants from the same location participate using their respective chat terminals, causing interference between the streamed audio and the voices of other nearby participants, making it difficult to hear. [Means for solving the problem]

[0009] To solve the above problems, the present invention comprises the configurations described in each claim. [Effects of the Invention]

[0010] According to the present invention, when multiple participants from the same location join a chat using their respective chat terminals, the problem of other participants' voices nearby interfering with the streamed audio and making it difficult to hear can be resolved. Other objectives, configurations, and effects will be clarified in the following embodiments. [Brief explanation of the drawing]

[0011] [Figure 1] This is a diagram illustrating the configuration of a web conferencing system. [Figure 2A] This is a hardware configuration diagram of a web conferencing terminal. [Figure 2B] This is a hardware configuration diagram of a web conferencing terminal. [Figure 3] This is a functional block diagram of a web conferencing terminal according to the first embodiment. [Figure 4] This is a functional block diagram showing the details of the correlation calculation unit. [Figure 5] This diagram illustrates the first example of voice reduction processing for other people near the user included in the streamed audio. [Figure 6] This is a flowchart showing the processing flow of the web conferencing system according to the first embodiment. [Figure 7] This diagram illustrates a second example of voice reduction processing for other people's voices. [Figure 8] This flowchart shows the processing flow of a web conferencing system, including a second audio reduction process for spoken audio (audio of other users broadcast on the system). [Figure 9] This is a diagram showing the mesh network configuration between web conferencing terminals within a single location. [Figure 10] This diagram illustrates the process of reducing the audio volume of spoken words based on a list of prohibited broadcasts. [Figure 11] This flowchart shows the processing flow of a web conferencing system that supports a third type of voice reduction processing for other users' voices. [Figure 12] This is a diagram illustrating the configuration of a web conferencing system according to the third embodiment. [Figure 13]It is a block diagram of a WEB conference terminal realized by an information processing device. [Figure 14] It is a functional block diagram of a WEB conference terminal according to a third embodiment. [Figure 15] It is a flowchart showing the processing flow of a WEB conference system corresponding to a serverless WEB conference system.

Embodiments for Carrying out the Invention

[0012] Hereinafter, embodiments of the present invention will be described with reference to the drawings. The same components and steps throughout the figures are denoted by the same reference numerals, and redundant descriptions are omitted.

[0013] The chat system according to the present invention is a system for transmitting and receiving voice data directly between a plurality of chat terminals or via a chat server. The chat system is applicable to a work support system that transmits and receives voice data, for example, between chat terminals worn by workers working at a work site and between the terminal of a management center located at a place away from the work site.

[0014] Also, the chat system according to the present invention is applicable to a voice chat system that transmits and receives voice data directly between chat terminals worn by each team member or via a chat server when a plurality of people play e-sports in a team. Furthermore, it is also applicable to an e-sports system or a game system incorporating a voice chat system.

[0015] In the following description, a WEB conference system incorporating the chat system according to the present invention will be described as an example. Since the present invention can be expected to bring diversification and technological improvement to labor-intensive industries, for example, it can be expected to contribute to 8.2 of the Sustainable Development Goals (SDGs) proposed by the United Nations (To increase the productivity of the economy through diversification, technological improvement, and innovation, mainly in industries that enhance the value of goods and services and labor-intensive industries).

[0016] [First Embodiment of the Invention] A first embodiment of the present invention will be described with reference to Figures 1 to 8.

[0017] Figure 1 is a diagram illustrating the configuration of a web conferencing system.

[0018] In Figure 1, the web conferencing system 100 is configured by connecting web conferencing terminals 3A-3F (equivalent to chat terminals; hereinafter sometimes simply referred to as "terminals") installed at web conferencing locations A, B, and C, respectively, and the web conferencing server 5 (equivalent to a chat server) to each other via network 4. Office AO is the interior of the office located at web conferencing location A.

[0019] The following explanation will use base A as an example, but the explanation for base A also applies to bases B and C.

[0020] Location A has three web conference participants: 2A, 2B, and 2C. The web conference terminals used by each participant (2A, 2B, and 2C) are terminals 3A, 3B, and 3C.

[0021] Even when participants 2A, 2B, and 2C at location A gather in the same room, such as a conference room, to conduct a web conference, participants 2A, 2B, and 2C will each use their own devices 3A, 3B, and 3C to participate in the web conference.

[0022] Participants 2A, 2B, and 2C are joining the web conference from location A. They access the web conference server 5 via network 4 using terminals 3A, 3B, and 3C respectively, and receive the web conference service. For example, participant A's image and voice (hereinafter referred to as "user voice") are collected by terminal A and sent to the web conference server 5.

[0023] The web conferencing server 5 receives the video and audio of all participants connected to the web conferencing service, generates the video and audio for web conferencing, and distributes them to each participant's terminal. For example, participant A's voice (user voice) is distributed as part of the audio for web conferencing to the terminals of participants at locations B and C (terminals D, E, and F).

[0024] However, the audio streamed by terminals 3B and 3C, operated by other participants located near participant A (in this example, participants B and C), does not include participant A's voice (user voice). This eliminates the problem where participants B and C hear participant A's voice (other people's voices) directly transmitted through the air within office AO, and participant A's voice (user voices) included in the audio streamed by terminals 3B and 3C, with a time lag between them. This is one of the features common to each embodiment of the present invention.

[0025] Figures 2A and 2B are hardware configuration diagrams for the web conferencing terminals. Since web conferencing terminals 3A to 3F have the same configuration, they are referred to as terminal 3 unless otherwise specified.

[0026] Terminal 3 is equipped with a camera 11, microphone 12, display 13, audio output 14, communication device 15, processor 16, first memory (RAM) 17, second memory (FROM) 18, input device 19, and sensor group 20, which are connected to each other by a bus 21. The camera 11 and display 13 are not mandatory for terminal 3; in that case, web conferencing will be conducted using only audio.

[0027] The processor 16 is composed of, for example, a CPU.

[0028] RAM17 is an example of volatile memory.

[0029] FROM18 is an example of non-volatile memory. FROM18 includes a basic operation program 30, a web conferencing application (abbreviated as "app" in the diagram) program 31, and data 32.

[0030] The camera 11 may be integrated with the terminal 3, or it may be a camera connected via a USB terminal.

[0031] Microphone 12 picks up the voice of the user of terminal 3 (user voice) as well as the voices of other participants speaking during the web conference at the same location (other participants' voices). If there is only one microphone 12 and it is not directional, it will pick up both the user voice and other participants' voices. When we refer to the voice picked up by microphone 12 without distinguishing between the user voice and other participants' voices, we simply say "microphone-picked voice".

[0032] Figure 2A shows a single microphone 12 (user microphone), illustrating a case where the same microphone picks up both the user's voice and the voices of others. However, separate microphones with directional characteristics suitable for picking up each type of voice may be provided. A microphone with directional characteristics suitable for picking up the user's voice from terminal 3 is called a user-dedicated microphone, and a microphone with directional characteristics suitable for picking up ambient sounds is called a shared microphone. A user-dedicated microphone is, for example, the microphone found in a headset. A shared microphone is, for example, a microphone placed on a table in a conference room that is suitable for omnidirectional sound pickup. As shown in Figure 2B, a microphone 12a dedicated to other voices (sometimes abbreviated as "dedicated microphone") may be connected to the bus 21, or a microphone 12b dedicated to other voices may be connected via Bluetooth® through a short-range wireless communication device 152. Other voices are the sounds picked up by the microphone (user microphone or microphone dedicated to picking up other voices) when the speaker is not speaking. Using the user microphone as the dedicated microphone is preferable because it eliminates the need to add a separate microphone 12b dedicated to other voices. Microphone 12 (the user's microphone) is muted when the user is not speaking (the function of Microphone 12 itself is running, but the audio from Microphone 12 is not broadcast as audio), so that any audio picked up during that time is processed as not being the user's voice, but rather someone else's voice.

[0033] The input device 19 is a keyboard or a touch sensor. In the case of a smartphone, the flat display (display 13) and the touch sensor are integrated, and the keyboard operates using the basic operation program 30.

[0034] The audio output device 14 is a device that outputs the streamed audio, and may be a speaker, earphones, headphones, headset, or audio output terminal.

[0035] The communication device 15 includes a LAN communication device 151 for exchanging data such as images and audio with the web conferencing server 5, as well as multiple communication methods and communication protocols, such as a Bluetooth® short-range wireless communication device 152 for use between terminals within the same location.

[0036] The sensor group 20 includes, for example, an illuminance sensor 201 and a motion sensor 202, and assists in the use of the terminal.

[0037] Figure 3 is a functional block diagram of the web conferencing terminal according to the first embodiment.

[0038] The web conferencing terminal 3 has a correlation calculation unit 161 and a sound reduction unit 162. The correlation calculation unit 161 and the sound reduction unit 162 are realized when the processor 16 loads the basic operation program 30 and the web conferencing application program 31 into RAM 17 and executes them. The data 32 contains the data necessary to execute the basic operation program 30 and the web conferencing application program 31, and the processor 16 reads it as appropriate when executing the web conferencing application program 31 and uses it for processing each unit.

[0039] The image of the terminal user captured by camera 11 is transmitted from LAN communication device 151 to web conferencing server 5 via network 4.

[0040] The LAN communication device 151 receives the video and audio streams for the web conference from the web conferencing server 5. The video streams are displayed on the display 13. The audio streams are supplied to the correlation calculation unit 161 and the audio reduction unit 162.

[0041] Furthermore, the user's voice and other people's voices (microphone-collected voices) picked up by the microphone 12 are transmitted from the LAN communication device 151 to the web conferencing server 5 via the network 4, and are also supplied to the correlation calculation unit 161 and the voice reduction unit 162.

[0042] The correlation calculation unit 161 takes the streamed audio and the user's voice and other people's voices from the microphone 12 as input, performs a correlation calculation, obtains the delay amount and correlation amount between the two voices, and sends them to the audio reduction unit 162.

[0043] The audio reduction unit 162 reduces the user voice and other voices from the streamed audio by subtracting them from the streamed audio by referring to the delay amount and correlation amount, and generates output audio for terminal 3.

[0044] The audio outputter 14 outputs the audio output from the audio reduction unit 162. This reduces the amount of microphone-collected audio (user voice and other people's voices) picked up by the terminal's microphone 12 that is output from the audio outputter 14 as distributed audio, thereby reducing interference with other people's voices that are directly audible.

[0045] Figure 4 is a functional block diagram showing the details of the correlation calculation unit.

[0046] The correlation calculation unit 161 includes a variable delay unit 161a, a delay amount setting unit 161b, a sum-of-products unit 161c, and an output processing unit 161d.

[0047] The microphone-collected audio (user voice and other people's voices) is input to the variable delay unit 161a. The delay time in the variable delay unit 161a is set by the delay amount setting unit 161b. The "spoken voice" input to the variable delay unit 161a is either the user's voice or another person's voice picked up while muted.

[0048] The delayed microphone-collected audio (user voice and other voices) and the streamed audio are input to the sum-of-products unit 161c, where a sum-of-products calculation is performed to obtain a correlation amount using the set delay time as a parameter. The sum-of-products unit 161c varies the delay time to obtain the delay time that maximizes the correlation amount, and sets this correlation amount to the delay amount associated with the stream.

[0049] The output processing unit 161d outputs the delay amount and correlation amount when the distributed audio is superimposed audio as shown in Figure 5, which will be described later. When the distributed audio is packet multiplexed audio as shown in Figure 7, which will be described later, the correlation amount is compared for each audio packet to be separated and the packet ID corresponding to the microphone-collected audio (user voice and other person's voice) is output.

[0050] (First example of audio reduction processing) Figure 5 illustrates the first example of voice reduction processing for other people near the user included in the streamed audio.

[0051] The web conferencing server 5 is equipped with an audio distribution unit 50. From the audio distribution unit 50, the distributed audio 53 is sent to the audio reduction unit 162.

[0052] In each terminal, as shown in Figure 5, the voice 51A from terminal A and the voices collected by other terminals (51E, 51D, and 51F in Figure 5) are superimposed and added together in the voice multiplexing unit 52 and distributed as the distributed voice 53.

[0053] The sound reduction unit 162a subtracts the collected sound from the distributed sound 53 by referring to the delay amount and correlation amount obtained by the correlation calculation unit 161, and subtracts the voice of another person near the user (other person's voice).

[0054] Figure 6 is a flowchart showing the processing flow of the web conferencing system according to the first embodiment.

[0055] When terminal 3 starts the web conferencing application program 31 (S10), terminal 3 logs in to the web conferencing service provided by the web conferencing server 5 (S11) and participates in the web conference.

[0056] Terminal 3 captures a camera image with the camera 11 (S12) and collects sound with the microphone 12 (S13).

[0057] Terminal 3 transmits the camera image and the microphone audio collected by the microphone 12 of Terminal 3 to the web conferencing server 5 (S14). The web conferencing server 5 receives the transmitted image and audio (S15).

[0058] If the microphone mute button on terminal 3 is pressed and terminal 3 is muted (S16: Yes), terminal 3 will determine that the voice picked up by the microphone is someone else's voice, because the terminal user has no intention of speaking.

[0059] By keeping the user's microphone active even when muted, it can be used as a microphone to pick up other people's voices. Alternatively, a separate microphone for picking up other people's voices may be provided in addition to the user's microphone. By placing the microphone for picking up other people's voices near the conference speaker who is close to the terminal user, other people's voices can be picked up more accurately, improving the accuracy of correlation calculations. When using the microphone for picking up other people's voices, the microphone audio pickup for S13 should be performed by the microphone for picking up other people's voices. When terminal 3 is muted, the microphone audio pickup for S13 should be performed by the microphone for picking up other people's voices.

[0060] If terminal 3 is muted (S16: Yes), the correlation calculation unit 161 performs a correlation calculation between the streamed audio and the other party's audio, calculates the delay amount and correlation amount, and outputs them to the audio reduction unit 162.

[0061] Specifically, the audio reduction unit 162 subtracts the audio picked up by the microphone (other people's voices) from the streamed audio (S17, S18), and the streamed audio with the other people's voices subtracted is output from the audio outputter 14 (S18, S19). The audio output from the audio outputter 14 is called "amplified audio".

[0062] If terminal 3 is muted (S16: No), the streaming audio does not include the user's voice (the user's voice has already been removed using the existing method), so the streaming audio is output directly from the audio outputter 14 as amplified audio (S19).

[0063] If the user does not log out (S21: NO), the process returns to step S12 and is repeated. If the user logs out (S21: YES), the web conferencing application program is terminated (S22).

[0064] (Second example of audio reduction processing) Figure 7 illustrates a second example of voice reduction processing for other people's voices.

[0065] In Figure 7, similar to Figure 5, the web conferencing server 5 is equipped with an audio distribution unit 50. From the audio distribution unit 50, the distributed audio 56 is sent to the audio reduction unit 162.

[0066] The audio from the web conference (audio 51A from terminal A in Figure 5) and the audio collected from other terminals D, E, and F (51D, 51E, and 51F in Figure 5) are combined in a packet multiplexing process in the packet multiplexing unit 55, where the audio from each terminal is stored in packets with different identification numbers (hereinafter referred to as IDs), and then distributed as the distributed audio 56.

[0067] The packet removal unit 57 of the audio reduction unit 162 separates the spoken audio (audio spoken by other users distributed on the system) from the distributed audio 56 using the packet ID obtained by the correlation calculation unit 161, and removes it in the packet removal unit 57. The terminal audio after removal is 51D, 51E, and 51F, which are then processed multiplexed in the audio multiplexing unit 58 and sent to the audio output device 14.

[0068] Figure 8 is a flowchart showing the processing flow of a web conferencing system, including a second audio reduction process for spoken audio (audio spoken by other users distributed through the system).

[0069] The audio reduction process follows the audio reduction method shown in Figure 7. Steps with the same function as those described in the first flowchart in Figure 6 are assigned the same numbers, and redundant explanations are omitted.

[0070] The difference between the flowchart in Figure 8 and the flowchart in Figure 6 lies in step S30, where the spoken audio (audio spoken by other users distributed on the system) is removed using the packet removal method described in Figure 7.

[0071] As described above, the web conferencing terminal, web conferencing application, and web conferencing system of the first embodiment of the present invention have the advantage of providing a web conferencing environment in which participants use their respective web conferencing terminals, with less interference between the participants' voices and the audio streamed on the web conferencing system, and making the voices easier to hear.

[0072] [Second Embodiment of the Invention] A second embodiment of the present invention will be described with reference to Figures 9 to 11.

[0073] Figure 9 shows a mesh network configuration between web conferencing terminals within a single location. In Figure 9, terminals A, B, and C are present at location A and are interconnected via proximity communication 36. Terminal H is then added to this configuration.

[0074] When terminal H enters base A, it searches for nearby terminals using proximity communication 36 and completes a connection with an available terminal C. Terminal C detects terminal H's new participation and informs terminals A and B of this, as well as relaying information about terminals A and B to terminal H. As a result, terminals A, B, C, and H obtain information about all terminals within base A and can create a list of prohibited audio streams that prohibit the inclusion of spoken audio collected by terminals within the same base in the streamed audio.

[0075] Figure 10 illustrates the audio reduction process for spoken words based on a distribution blacklist, and shows the audio distribution section of a web conferencing server.

[0076] The microphone-collected audio 51A, picked up by the microphone 12 of terminal A, and microphone-collected audio (51D, 51E, 51F) picked up from other terminals are input to the packet removal unit 60. The audio multiplexer 61 adds the data values ​​and distributes them as the distributed audio 63.

[0077] The audio distribution unit 50 of the web conferencing server 5 receives a distribution blacklist 62 from the participants' terminals. For example, terminal B's distribution blacklist 62 lists terminals A, C, and H, which are located at the same site. In this way, the distribution blacklist 62 specifies the audio that should be removed from the audio distributed by each terminal. The audio to be removed is defined by the name of the terminal (terminal A, C, H, etc.) to which the microphone that collected the audio was connected.

[0078] In generating audio for distribution to terminal B, the packet removal unit 60 removes audio packets that are on the distribution prohibition list 62 for each terminal based on the distribution prohibition list 62.

[0079] The audio multiplexing unit 61 adds (multiplexes) the audio that has passed through the packet removal unit 60 to generate the distributed audio 63 and distributes it to terminal B.

[0080] Figure 11 is a flowchart showing the processing flow of a web conferencing system that supports a third type of audio reduction processing for other people's voices.

[0081] In the flowchart of Figure 11, steps with the same function as those explained in the flowchart of Figure 6 are assigned the same numbers, and redundant explanations are omitted.

[0082] The flowchart in Figure 11 differs from the first flowchart in Figure 6 in steps S40, S41, and S42. In S40, a new proximity network is created or updated as explained in Figure 9. In S41, a new distribution ban list 62 is created or updated, and in S42, the distribution ban list 62 is sent to the web conferencing server 5.

[0083] In S15, the system receives the streamed video and audio from the web conferencing server 5. However, the received streamed audio does not include the voices of other participants at the same location, as explained in Figure 10.

[0084] As described above, the web conferencing terminal, web conferencing application, and web conferencing system of the second embodiment of the present invention have the same features as the first embodiment, and also have the feature of being able to reliably remove the voices of other participants who are at the same location.

[0085] [Third Embodiment of the Invention] A third embodiment of the present invention will be described with reference to Figures 12 to 14. In this embodiment, a web conference can be performed even without a web conferencing server 5.

[0086] Figure 12 is a diagram showing the configuration of a web conferencing system according to the third embodiment.

[0087] In Figure 12, the difference from the web conferencing system in Figure 1 is that it is a serverless system without a web conferencing server 5. For example, the camera image and microphone audio of participant 2A, captured and recorded by terminal 3A, are distributed to the terminals (terminals B to F) of all participants in the web conference.

[0088] Terminal 3A also receives images and audio from all terminals (Terminals B to F) and generates images and audio for the web conference within the terminal.

[0089] Figure 13 is a block diagram of a web conferencing terminal implemented using an information processing device, and is a web conferencing terminal that supports serverless web conferencing. In the web conferencing terminal in Figure 13, blocks that have the same functionality as the web conferencing terminal in Figure 3 are assigned the same numbers, and redundant explanations are omitted.

[0090] In terminal 3 of Figure 13, the web conferencing application program 31 included in FROM 18 includes a server program 33 and a client program 34. The server program 33 distributes the camera image and microphone audio of the terminal user to other terminals and receives images and audio from other terminals.

[0091] The client program 34 captures and collects camera images and microphone audio from the terminal user, and shares the terminal user's camera images, microphone audio, and camera images and microphone audio from other terminals with the server program 33.

[0092] The server program 33 generates images and audio for the web conference and outputs them to the display 13 and audio output device 14 via the client program 34. Note that the server program 33 does not need to be implemented on all terminals participating in the web conference; it is sufficient for at least one terminal to implement it for the web conference to function. In this case, the terminal implementing the server program 33 and the client programs 34 of the other terminals exchange images and audio via the communication unit 24.

[0093] Figure 14 is a functional block diagram of a web conferencing terminal according to the third embodiment.

[0094] Terminal 3 in Figure 14 is further equipped with a participant list creation unit 163 that creates a participant list based on the communication results of short-range communication 35 from a short-range wireless communication device 152, in addition to the terminal 3 in Figure 2.

[0095] Figure 15 is a flowchart showing the processing flow of a web conferencing system that supports serverless web conferencing systems.

[0096] The flowchart showing the processing flow of the web conferencing system in Figure 6 uses the same numbers for the same steps.

[0097] The program is started (S10). The flowchart showing the processing flow of the web conferencing system consists of a client process and a server process.

[0098] In the client process, a notification is sent indicating that the user is participating in the web conference (S50). The notification is sent to the terminals of the participant candidates listed in the participant candidate list obtained in advance.

[0099] When a camera image is captured (S12) and audio is collected by the microphone 12 (S13), the camera image and the audio collected by the microphone are shared with the server process.

[0100] Furthermore, S51 shares the images and audio output by the server process.

[0101] In S16, the system checks whether the microphone-collected audio shared in S51 contains the voices of other participants at the same location. If it is determined that there are voices of other participants (S16: YES), the correlation calculation unit 161 performs a correlation calculation between the output audio of the server process and the voices of other participants at the same location (S17), outputs parameters indicating the delay amount and correlation amount to the audio reduction unit 162, and the audio reduction unit 162 subtracts the voices of other participants spoken by the participants (S18) and outputs the amplified audio (S19). In addition, the image shared in step S51 is displayed on the display 13 (S20).

[0102] In the server process, upon receiving notifications from each terminal (S52), the participant list creation unit 163 creates or updates a list of participants who are actually attending the meeting from the pre-distributed list of candidate participants (S53).

[0103] In S54, the camera image and collected audio are shared with the client process, and in S55, camera images and audio are received from other terminals. In S56, the output image of the web conference is obtained from the camera images of all terminals.

[0104] In S57, it is checked whether a distribution ban list exists and whether the spoken audio is included in the distribution ban list. If a distribution ban list exists and the spoken audio is included in the distribution ban list (S57: YES), the other person's audio is removed (S58). The distribution ban list is constructed by adding a distribution ban item (flag) to the participant list. In this case, the list of participants in the participant list who are prohibited from distribution corresponds to the distribution ban list.

[0105] If there is no distribution blacklist or if the other party's voice is not included in the distribution blacklist (S57:NO), step S58 is skipped. Then, in step S59, the output audio is created and shared with the client process.

[0106] The output images and audio from the server process are equivalent to the streamed images and audio from a web conferencing system with a server.

[0107] As described above, the web conferencing terminal, web conferencing application, and web conferencing system of the third embodiment of the present invention have the same features as the first and second embodiments, and enable serverless web conferencing. This offers a cost advantage when running web conferencing with a small number of terminals.

[0108] [Fourth Embodiment of the Present Invention] When web conference participants use noise-canceling headphones (hereinafter referred to as NCH) during a web conference, it is possible to reduce the actual (speaker's) voice being spoken in real time using noise-canceling technology, without reducing the system audio output from the NCH. However, in that case, external sounds other than the speaker's voice are also reduced, which can lead to inconveniences such as not noticing phone rings or other people calling out to you during the web conference. Therefore, by enabling the noise-canceling function of the NCH only when the actual speaker's voice is present, the actual speaker's voice is reduced, and when the actual speaker's voice is not present, the noise-canceling function is disabled so that external sounds are not reduced, making it possible to distinguish other external sounds. Furthermore, by applying noise cancellation to external sounds only to the system audio's voice, it is possible to reduce only the voice, making it possible to distinguish other external sounds even when the voice is present.

[0109] Although each embodiment has been described using web conferencing as an example, the method of the present invention is effective not only for web conferencing but also for systems that use information terminals to conduct conversations between people in remote locations while there are participants nearby.

[0110] While embodiments of the present invention have been described above, it goes without saying that the configurations for realizing the technology of the present invention are not limited to the above embodiments, and various modifications are conceivable. For example, the embodiments described above are described in detail for the purpose of explaining the present invention in an easy-to-understand manner, and are not necessarily limited to those comprising all the described configurations. Furthermore, it is possible to replace a part of the configuration of one embodiment with the configuration of another embodiment, and it is also possible to add the configuration of another embodiment to the configuration of one embodiment. All of these fall within the scope of the present invention. In addition, the numbers and messages that appear in the text and figures are merely examples, and using different ones will not impair the effects of the present invention.

[0111] Furthermore, the programs described in each processing example may be independent programs, or multiple programs may constitute a single application program. The order in which each processing step is executed may also be changed.

[0112] The functions of the present invention described above may be implemented in hardware, in whole or in part, for example, by designing them as an integrated circuit. Alternatively, they may be implemented in software by a microprocessor unit, CPU, etc., interpreting and executing an operating program that realizes each function. Furthermore, the scope of software implementation is not limited, and hardware and software may be used in combination. In addition, some or all of each function may be implemented on a server. The server only needs to be able to execute functions in cooperation with other components via communication, and its form is not limited, for example, a local server, cloud server, edge server, network service, etc. Information such as programs, tables, files that realize each function may be stored in memory, a recording device such as a hard disk or SSD (Solid State Drive), or a recording medium such as an IC card, SD card, or DVD, or it may be stored on a device on a communication network.

[0113] Furthermore, the control lines and information lines shown in the diagram are those deemed necessary for explanation and do not necessarily represent all control lines and information lines on the product. In reality, it is reasonable to assume that almost all components are interconnected.

[0114] The above-mentioned embodiment includes the following forms.

[0115] (Note 1) It is a chat terminal, Mike and, A communication device that sends and receives data to and from the chat server, Audio output device, Equipped with a processor, The microphone collects the user's voice spoken by the terminal user, and the voices of other people present near the terminal user. The communication device transmits the user's voice to the chat server, and the chat server Received the audio stream, The processor determines the correlation between the streamed audio and the other party's audio, To reduce the amount of other people's voices included in the distributed audio, The distributed audio, with the aforementioned other party's voice reduced, is output to the audio output device. Chat terminal. (Note 2) It is a chat terminal, Mike and, A communication device that sends and receives data between other chat terminals, Audio output device, Equipped with a processor, The microphone collects the user's voice spoken by the terminal user, and the voices of other people present near the terminal user. The communication device transmits the user's voice to the other chat terminal and receives the transmitted voice from the other chat terminal. The processor determines the correlation between the streamed audio and the other party's audio, To reduce the amount of other people's voices included in the distributed audio, The distributed audio, with the aforementioned other party's voice reduced, is output to the audio output device. Chat terminal. (Note 3) A chat system that is configured by communicating between a chat terminal and a chat server. , The chat terminal is, Mike and, A communication device that sends and receives data to and from the chat server, Audio output device, Equipped with a processor, The aforementioned microphone captures the user's voice spoken by the terminal user, as well as the voices of other people in the vicinity of the terminal user. Collects the voices of others that are generated by the user. The communication device transmits the user's voice to the chat server, and the chat server Received the audio stream, The processor determines the correlation between the streamed audio and the other party's audio, To reduce the amount of other people's voices included in the distributed audio, The distributed audio, with the aforementioned other party's voice reduced, is output to the audio output device. Chat system. (Note 4) A method for controlling a chat system that consists of a communication connection between a chat terminal and a chat server. It is a law, The user's voice, spoken by the terminal user through a microphone connected to the chat terminal, and the terminal The process involves collecting sounds from other people nearby the user, The aforementioned user voice is sent to the chat server, and the streamed audio is received from the chat server. Step and, The steps include determining the correlation between the distributed audio and the other party's audio, A step of reducing the other party's voice included in the distributed audio, The aforementioned streamed audio, with the other party's voice reduced, is output from the audio output device connected to the chat terminal. The steps to output, A method for controlling a chat system, including the chat system itself. [Explanation of symbols]

[0116] 2A:Participant 2B:Participant 2C:Participant 3: Web conferencing terminal 3A: Web conferencing terminal 3B: Web conferencing terminal 3C: Web conferencing terminal 3D: Web conferencing terminal 3E: Web conferencing terminal 3F: Web conferencing terminals 4: Network 5: Web conferencing server 11: Camera 12: Mike 12a: Microphone for recording other people's voices only 12b: Microphone for recording other people's voices only 13: Display 14: Audio output device 15: Communication device 16: Processor 17: RAM 19: Input device 20: Sensor group 21: Bus 24: Communications Department 30: Basic Operation Program 31: Web conferencing application program 32: Data 33: Server Program 34: Client Program 35: Near field communication 36: Proximity Communication 50: Audio Distribution Department 51A: Microphone-collected audio 52: Audio multi-channel section 53: Streaming audio 55: Packet multiplexing section 56: Streaming audio 57: Packet Removal Unit 58: Audio multi-channel 60: Packet Removal Unit 61: Audio multi-channel 62: List of content banned from distribution 63: Streaming audio 100: Web conferencing system 151 :LAN communication device 152: Near field wireless communication device 161: Correlation Calculation Unit 161a: Variable delay section 161b: Delay amount setting section 161c: Sekiwabu 161d: Output processing unit 162: Voice Reduction Unit 162a: Subtraction section 163: Participant List Creation Department 201: Illuminance sensor 202: Motion sensor

Claims

1. A voice transmission and reception system configured by communicating between an information terminal and a server, Information terminals are, A microphone to collect spoken audio, A communication device that sends and receives data to and from a server, Audio output device, Equipped with a processor, The aforementioned processor, When already connected via proximity communication to at least one other information terminal, and searching for a new information terminal via proximity communication, the system connects to the new information terminal and transmits information about the new information terminal to the other information terminals. From another information terminal that is in close proximity, the other information terminal receives information about an information terminal that has newly established a close proximity connection. Create a list of prohibited audio distributions for the distributed audio, which prohibits the inclusion of spoken audio collected by all other information terminals connected via proximity communication, and any new information terminals discovered by those other information terminals, and send it to the server. The server receives the audio of the spoken words, after processing to reduce their volume based on the prohibited distribution list, and outputs it to the audio output device. Voice transmission and reception system.

2. A voice transmission and reception system according to claim 1, The aforementioned server, Packet removal unit, It comprises an audio multi-channel section, The collected audio transmitted from all of the aforementioned information terminals is input to the packet removal unit. The packet removal unit removes audio packets specified in the distribution prohibition list when generating audio to be distributed to a predetermined information terminal. The audio multiplexing unit adds the audio remaining after passing through the packet removal unit to generate the distributed audio, and distributes the generated distributed audio to the predetermined information terminal. Voice transmission and reception system.

Citation Information

Patent Citations

  • Multi-point video conference system

    JP1996237627A

  • Sound controller, sound control method, and sound control program

    JP2014131096A

  • Conference terminal, conference server, conference system and program

    JP2014165888A