Sound communication device

By using multiple input units and sound image localization processing in the voice communication device, the sound image position and background noise in the virtual space are simulated, which solves the problem of insufficient presence in remote meetings and achieves clearer sound image localization and enhanced immersion.

CN114173275BActive Publication Date: 2026-07-24SOCIONEXT INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SOCIONEXT INC
Filing Date
2021-07-15
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

In remote meetings and online dinners, participants find it difficult to achieve a strong sense of presence. Existing voice communication devices cannot effectively simulate the positional relationship of multiple speakers' voices in virtual space and background noise, resulting in unclear sound localization.

Method used

Using N input units, a sound image location determination unit, and a sound image localization unit, the sound image localization position in the virtual space is simulated through head transfer function and addition operation processing, so that the sound signal does not overlap between the first wall and the second wall, and the atmosphere of the virtual space is enhanced by combining background noise signal processing.

Benefits of technology

It enhances the sense of presence in remote meetings and online dinners, making it easier for listeners to distinguish the direction from which the speaker's voice is coming, improving the realism of the virtual space and the simulation of background noise, and enhancing the participants' immersion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114173275B_ABST
    Figure CN114173275B_ABST
Patent Text Reader

Abstract

A sound communication device improves the sense of presence in a teleconference. The sound communication device includes: a sound image position decision unit that decides a sound image positioning position in a virtual space having a first wall and a second wall for each of N sound signals; N sound image positioning units that perform sound image positioning processing in a manner that a sound image is positioned at the sound image positioning position, and output sound image positioning sound signals; and an addition unit that adds the N sound image positioning sound signals, and outputs an added sound image positioning sound signal. The sound image positioning unit performs the sound image positioning processing using a first head transfer function that simulates a direct arrival at a listener's binaural ears of a sound wave emitted from the sound image positioning position and a second head transfer function that simulates an arrival at the listener's binaural ears of a sound wave emitted from the sound image positioning position and reflected from a wall closer to the sound image positioning position than the other wall among the first wall and the second wall.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a voice communication device used by multiple speakers in a remote conference. Background Technology

[0002] Previously known are voice communication devices used by multiple speakers in remote conferences (e.g., see Patent Document 1).

[0003] (Existing technical literature)

[0004] (Patent Documents)

[0005] Patent Document 1: Japanese Patent Application Publication No. 2006-237841

[0006] (Non-patent literature)

[0007] Non-Patent Document 1: Yeons Brunellet, Masayuki Morimoto, and Toshiyuki Goto, “Spatial Sound”, Kashima Publishing.

[0008] In remote meetings and online dinners held using voice communication devices, the aim is to enhance the sense of presence experienced by participants. Summary of the Invention

[0009] Therefore, the purpose of this disclosure is to provide a voice communication device that can enhance the sense of presence experienced by participants in remote meetings, online dinners, and other similar events conducted using the voice communication device.

[0010] One aspect of this disclosure relates to a sound communication device comprising: N input units for inputting sound signals, where N is an integer of 2 or more; a sound image position determination unit for determining a sound image positioning position in a virtual space having a first wall and a second wall for each of the N sound signals input from the N input units; N sound image positioning units, each of the N sound image positioning units corresponding to each of the N input units, each of the N sound image positioning units performing sound image positioning processing and outputting a sound image positioning sound signal, the sound image positioning processing being a process of positioning the sound image at the sound image positioning position, the sound image positioning position being a position determined by the sound image position determination unit for the input unit corresponding to that sound image positioning unit; and an addition unit for adding the N sound image positioning sound signals output from the N sound image positioning units and outputting an added sound image positioning sound signal, wherein the sound image position determination unit performs the following... The sound image positioning positions of the N sound signals are determined such that the sound image positioning positions of the N sound signals are located between the first wall and the second wall, and are in non-overlapping positions when viewed from the listener's position between the first wall and the second wall. Each of the N sound image positioning units performs the sound image positioning processing using a first head transfer function and a second head transfer function. The first head transfer function is a function that simulates the sound waves emitted from the sound image positioning position directly reaching the ears of a virtual listener at the listener's position. This sound image positioning position is the position determined by the sound image positioning unit for this sound image positioning unit. The second head transfer function is a function that simulates the sound waves emitted from the sound image positioning position being reflected by the wall between the first wall and the second wall that is closer to the sound image positioning position and reaching the listener's ears.

[0011] One aspect of this disclosure relates to a voice communication device comprising: N input units for inputting voice signals, where N is an integer of 2 or more; a sound image position determination unit for determining a sound image positioning position in a virtual space for each of the N voice signals input from the N input units; N sound image positioning units, each of the N sound image positioning units corresponding to each of the N input units, each of the N sound image positioning units performing sound image positioning processing and outputting a sound image positioning voice signal, wherein the sound image positioning processing is a process of positioning the sound image at the sound image positioning position, the sound image positioning position being determined by the sound image position determination unit for the input unit corresponding to that sound image positioning unit; and an addition unit for adding the N sound image positioning voice signals output from the N sound image positioning units and outputting an added sound image positioning voice signal. The sound image location determination unit determines the sound image positioning positions of the N sound signals in the following manner: the sound image positioning positions of the N sound signals are located in positions that do not overlap when viewed from the listener's position, and when the frontal angle of the virtual listener at the listener's position is set to 0 degrees, the interval between sound image positioning positions that include 0 degrees or are adjacent to each other with 0 degrees between them is narrower than the interval between sound image positioning positions that do not include 0 degrees or are adjacent to each other with 0 degrees between them. Each of the N sound image positioning units performs the sound image positioning processing using a head transfer function, which is a function that simulates the sound waves emitted from the sound image positioning positions directly reaching the ears of the virtual listener at the listener's position. The sound image positioning positions are the positions determined by the sound image location determination unit for that sound image positioning unit.

[0012] One aspect of this disclosure relates to a voice communication device comprising: N input units for inputting voice signals, where N is an integer of 2 or more; a sound image position determination unit for determining a sound image positioning position in a virtual space for each of the N voice signals input from the N input units; N sound image positioning units, each of the N sound image positioning units corresponding to each of the N input units, each of the N sound image positioning units performing sound image positioning processing and outputting a sound image positioning voice signal, wherein the sound image positioning processing is a process of positioning the sound image at the sound image positioning position, the sound image positioning position being determined by the sound image position determination unit for the input unit corresponding to that sound image positioning unit; and a first addition unit for adding the N sound image positioning voice signals output from the N sound image positioning units and outputting a first phase... The system includes: a sound image positioning sound signal; a background noise signal storage unit that stores a background noise signal showing the background noise in the virtual space; and a second addition operation unit that adds the first added sound image positioning sound signal to the background noise signal and outputs a second added sound image positioning sound signal. The sound image position determination unit determines the sound image positioning positions of the N sound signals so that they are in non-overlapping positions when viewed from the listener's position. Each of the N sound image positioning units performs the sound image positioning processing using a head transfer function, which is a function that simulates the sound waves emitted from the sound image positioning position directly reaching the ears of the virtually existing listener at the listener's position. The sound image positioning position is the position determined by the sound image position determination unit for that sound image positioning unit.

[0013] The audio communication device disclosed herein can enhance the sense of presence experienced by participants in remote meetings, online dinners, and other events conducted using the audio communication device. Attached Figure Description

[0014] Figure 1 This is a schematic diagram illustrating an example of the configuration of the remote conferencing system according to Embodiment 1.

[0015] Figure 2 This is a schematic diagram illustrating an example of the configuration of the server device according to Embodiment 1.

[0016] Figure 3 This is a block diagram illustrating an example of the configuration of the voice communication device according to Embodiment 1.

[0017] Figure 4 This is a schematic diagram showing an example of how the sound image location determination unit according to Embodiment 1 determines the sound image positioning position.

[0018] Figure 5 This is a schematic diagram illustrating an example of how the acoustic image localization unit according to Embodiment 1 performs acoustic image localization processing.

[0019] Figure 6 This is a block diagram illustrating an example of the configuration of the voice communication device according to Embodiment 2.

[0020] Symbol Explanation

[0021] 1. Remote conferencing system

[0022] 10,10A Voice Communication Device

[0023] 11 Input Section

[0024] 11A First Input Section

[0025] 11B Second Input Section

[0026] 11C Third Input Section

[0027] 11D Fourth Input Section

[0028] 11E Fifth Input Section

[0029] 12. Sound and image position determination unit

[0030] 13. Sound and image localization unit

[0031] 13A First Sound Image Positioning Unit

[0032] 13B Second Sound Image Positioning Unit

[0033] 13C Third Sound Image Positioning Unit

[0034] 13D 4th Sound Image Positioning Unit

[0035] 13E Fifth Sound Image Positioning Unit

[0036] 14. Addition Operation Section

[0037] 15, 15A Output Section

[0038] 16. Second Addition Operation Section

[0039] 17 Background noise signal storage unit

[0040] 18 Selection Department

[0041] Terminals 20, 20A, 20B, 20C, 20D, 20E, and 20F

[0042] Microphones 21, 21A, 21B, 21C, 21D, 21E, 21F

[0043] Speakers 22, 22A, 22B, 22C, 22D, 22E, 22F

[0044] Users 23A, 23B, 23C, 23D, 23E, and 23F

[0045] 30 Network

[0046] 41 The First Wall

[0047] 42 The second wall

[0048] 50 listener positions

[0049] 51 First audio-visual position

[0050] 52 Second audio-visual position

[0051] 53 Third audio-visual position

[0052] 54. Fourth audio-visual position

[0053] 55. Fifth audio-visual position

[0054] 60 listeners

[0055] Speakers 71, 72, 73, 74, 75

[0056] 71A, 74A Speaker's mirror image

[0057] 90 Virtual Space

[0058] 100 server devices

[0059] 101 Input Device

[0060] 102 Output device

[0061] 103 CPU

[0062] 104 Built-in Memory

[0063] 105 RAM

[0064] 106 bus Detailed Implementation

[0065] (The process of obtaining one of the solutions disclosed herein)

[0066] Previously, with the increasing speed and capacity of the Internet and the high performance of server devices, voice communication devices for remote conferencing systems that could be participated in simultaneously from multiple locations became practical. In recent years, due to the impact of the COVID-19 pandemic, such remote conferencing systems have been used not only for business purposes but also widely for consumer purposes such as online dining.

[0067] With the widespread use of voice communication devices for remote meetings and online dining, there is a growing demand for enhancing the sense of presence experienced by participants in these events.

[0068] Therefore, in order to enhance the sense of presence for participants in remote meetings, online dinners, and other events held using voice communication devices, the inventors conducted thorough experiments and discussions. As a result, they devised the following voice communication devices.

[0069] One aspect of this disclosure relates to a sound communication device comprising: N input units for inputting sound signals, where N is an integer of 2 or more; a sound image position determination unit for determining a sound image positioning position in a virtual space having a first wall and a second wall for each of the N sound signals input from the N input units; N sound image positioning units, each of the N sound image positioning units corresponding to each of the N input units, each of the N sound image positioning units performing sound image positioning processing and outputting a sound image positioning sound signal, the sound image positioning processing being a process of positioning the sound image at the sound image positioning position, the sound image positioning position being a position determined by the sound image position determination unit for the input unit corresponding to that sound image positioning unit; and an addition unit for adding the N sound image positioning sound signals output from the N sound image positioning units and outputting an added sound image positioning sound signal, wherein the sound image position determination unit performs the following... The sound image positioning positions of the N sound signals are determined such that the sound image positioning positions of the N sound signals are located between the first wall and the second wall, and are in non-overlapping positions when viewed from the listener's position between the first wall and the second wall. Each of the N sound image positioning units performs the sound image positioning processing using a first head transfer function and a second head transfer function. The first head transfer function is a function that simulates the sound waves emitted from the sound image positioning position directly reaching the ears of a virtual listener at the listener's position. This sound image positioning position is the position determined by the sound image positioning unit for this sound image positioning unit. The second head transfer function is a function that simulates the sound waves emitted from the sound image positioning position being reflected by the wall between the first wall and the second wall that is closer to the sound image positioning position and reaching the listener's ears.

[0070] The aforementioned voice communication device creates an atmosphere where the voices of N speakers, input from N input units, appear to originate from a virtual space with a first wall and a second wall. Listeners hearing the voices of N speakers can more easily grasp the spatial relationships between the speakers and the walls in the virtual space. Therefore, they can more easily distinguish the directions from which the voices of the N speakers are coming. Consequently, the aforementioned voice communication device enhances the sense of presence experienced by participants in remote meetings, online gatherings, and similar events conducted using voice communication devices compared to previous methods.

[0071] Alternatively, each of the N acoustic image localization units may perform the acoustic image localization process in a manner that allows for free variation of at least one of the reflectivity of the sound waves from the first wall and the reflectivity of the sound waves from the second wall.

[0072] This allows for the free alteration of the echo level of a speaker's voice in virtual space.

[0073] Alternatively, each of the N acoustic positioning units may perform the acoustic positioning process in a manner that allows for free change of at least one of the positions of the first wall and the second wall.

[0074] This allows one to freely change the position of walls in virtual space.

[0075] One aspect of this disclosure relates to a voice communication device comprising: N input units for inputting voice signals, where N is an integer of 2 or more; a sound image position determination unit for determining a sound image positioning position in a virtual space for each of the N voice signals input from the N input units; N sound image positioning units, each of the N sound image positioning units corresponding to each of the N input units, each of the N sound image positioning units performing sound image positioning processing and outputting a sound image positioning voice signal, wherein the sound image positioning processing is a process of positioning the sound image at the sound image positioning position, the sound image positioning position being determined by the sound image position determination unit for the input unit corresponding to that sound image positioning unit; and an addition unit for adding the N sound image positioning voice signals output from the N sound image positioning units and outputting an added sound image positioning voice signal. The sound image location determination unit determines the sound image positioning positions of the N sound signals in the following manner: the sound image positioning positions of the N sound signals are located in positions that do not overlap when viewed from the listener's position, and when the frontal angle of the virtual listener at the listener's position is set to 0 degrees, the interval between sound image positioning positions that include 0 degrees or are adjacent to each other with 0 degrees between them is narrower than the interval between sound image positioning positions that do not include 0 degrees or are adjacent to each other with 0 degrees between them. Each of the N sound image positioning units performs the sound image positioning processing using a head transfer function, which is a function that simulates the sound waves emitted from the sound image positioning positions directly reaching the ears of the virtual listener at the listener's position. The sound image positioning positions are the positions determined by the sound image location determination unit for that sound image positioning unit.

[0076] Regarding the auditory sharpness of typical sound image localization, it is known that the more directly a listener is positioned, the more sensitive the sound becomes, while the more to the left or right of the listener, the less sensitive it becomes (see, for example, Non-Patent Document 1). With the aforementioned voice communication device, from the listener's perspective, the angle between speakers positioned to the left or right is larger than the angle between speakers positioned directly in front of them. Therefore, the listener can more easily distinguish the direction from which the voices of N speakers are coming. Consequently, with the aforementioned voice communication device, the sense of presence experienced by participants in remote meetings, online gatherings, etc., held using voice communication devices can be improved compared to previous methods.

[0077] One aspect of this disclosure relates to a voice communication device comprising: N input units for inputting voice signals, where N is an integer of 2 or more; a sound image position determination unit for determining a sound image positioning position in a virtual space for each of the N voice signals input from the N input units; N sound image positioning units, each of the N sound image positioning units corresponding to each of the N input units, each of the N sound image positioning units performing sound image positioning processing and outputting a sound image positioning voice signal, wherein the sound image positioning processing is a process of positioning the sound image at the sound image positioning position, the sound image positioning position being determined by the sound image position determination unit for the input unit corresponding to that sound image positioning unit; and a first addition unit for adding the N sound image positioning voice signals output from the N sound image positioning units and outputting a first phase... The system includes: a sound image positioning sound signal; a background noise signal storage unit that stores a background noise signal showing the background noise in the virtual space; and a second addition operation unit that adds the first added sound image positioning sound signal to the background noise signal and outputs a second added sound image positioning sound signal. The sound image position determination unit determines the sound image positioning positions of the N sound signals so that they are in non-overlapping positions when viewed from the listener's position. Each of the N sound image positioning units performs the sound image positioning processing using a head transfer function, which is a function that simulates the sound waves emitted from the sound image positioning position directly reaching the ears of the virtually existing listener at the listener's position. The sound image positioning position is the position determined by the sound image position determination unit for that sound image positioning unit.

[0078] The aforementioned voice communication device can provide an atmosphere where the voices of N speakers, each input from N input units, appear to be emanating from a virtual space filled with background noise. Therefore, the voice communication device enhances the sense of presence experienced by participants in remote meetings, online dinners, and other similar events conducted using voice communication devices compared to previous methods.

[0079] Alternatively, the background noise signal storage unit may store one or more background noise signals, and the voice communication device may further include a selection unit that selects one or more background noise signals from the one or more background noise signals stored in the background noise signal storage unit, and the second addition unit adds the first summed sound image positioning sound signal to the background noise signal selected by the selection unit, and outputs the second summed sound image positioning sound signal.

[0080] Therefore, the background noise can be selected according to the atmosphere of the virtual space that you want to present.

[0081] Alternatively, the selection unit may change the selected background noise signal over time.

[0082] Thus, the atmosphere of the virtual space can be changed over time.

[0083] The following description, with reference to the accompanying drawings, illustrates a specific example of a voice communication device according to one aspect of this disclosure. The embodiments shown herein are merely specific examples of this disclosure. The numerical values, shapes, constituent elements, arrangements and connection methods of the constituent elements, as well as the steps (processes) and their sequence shown in the following embodiments are illustrative and not intended to limit the scope of this disclosure. Furthermore, the figures are schematic diagrams and not rigorous illustrations.

[0084] Furthermore, the general or specific solutions disclosed herein can be implemented by systems, methods, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, or by any combination of systems, methods, integrated circuits, computer programs, and recording media.

[0085] (Implementation Method 1)

[0086] The following description refers to a remote conferencing system for meetings with multiple participants in different locations, with reference to the attached diagram.

[0087] Figure 1 This is a schematic diagram illustrating an example of the configuration of the remote conferencing system 1 according to Embodiment 1.

[0088] like Figure 1 As shown, the remote conferencing system 1 includes: a voice communication device 10, a network 30, and N+1 (N being an integer greater than 2) terminals 20 (connected to...). Figure 1 (Corresponding to terminals 20A to 20F), N+1 microphones 21 (and) Figure 1 Microphones 21A to 21F correspond to each other), and N+1 speakers 22 (corresponding to each other). Figure 1 (This corresponds to speakers 22A to 22F).

[0089] Microphones 21A to 21F are connected to terminals 20A to 20F respectively, and convert the voices of users 23A to 23F using terminals 20A to 20F into sound signals, which are then output to terminals 20A to 20F. These sound signals are electrical signals.

[0090] Microphones 21A to 21F may have the same function. Therefore, in this specification, unless it is necessary to distinguish between microphones 21A to 21F, they are referred to as microphone 21.

[0091] Speakers 22A to 22F are connected to terminals 20A to 20F respectively. The sound signals output from terminals 20A to 20F as electrical signals are converted into sound and output to the outside.

[0092] The functions of speakers 22A to 22F can be the same. Therefore, in this specification, unless it is necessary to distinguish between speakers 22A to 22F, they are referred to as speaker 22. Speaker 22 can be any speaker that has the function of converting electrical signals into sound, and does not need to be limited to a so-called loudspeaker. For example, it can also be headphones, over-ear headphones, etc.

[0093] Terminals 20A to 20F are respectively connected to microphones 21A to 21F, speakers 22A to 22F, and network 30. They have the following functions: transmitting sound signals output from the connected microphones 21A to 21F to external devices connected to network 30; and receiving sound signals from the external devices connected to network 30 and outputting the received sound signals to speakers 22A to 22F. The external devices connected to network 30 include a voice communication device 10.

[0094] Terminals 20A to 20F may have the same functions. Therefore, in this specification, unless it is necessary to distinguish between terminals 20A to 20F, they are referred to as terminal 20. Terminal 20 may be implemented by, for example, a computer or a smartphone.

[0095] Terminal 20, for example, may have the functionality of microphone 21. In this case, Figure 1 The diagram appears to show terminal 20 connected to microphone 21, but microphone 21 is actually included within terminal 20. Furthermore, terminal 20 may also function as a speaker 22. In this case, Figure 1 The diagram appears to show terminal 20 connected to speaker 22, but in reality, speaker 22 is included within terminal 20. Furthermore, terminal 20 may also include input / output devices such as a display, touchpad, and keyboard.

[0096] Conversely, microphone 21 could also function as terminal 20. In this case, Figure 1 The diagram appears to show terminal 20 connected to microphone 21, but in reality, terminal 20 is included within microphone 21. Alternatively, speaker 22 could also function as terminal 20. In this case, Figure 1 The diagram appears to show that terminal 20 is connected to speaker 22, but in reality, terminal 20 is included in speaker 22.

[0097] Network 30 is connected to multiple devices, including terminals 20A to 20F and voice communication device 10, and transmits signals between the connected devices. As described later, voice communication device 10 is implemented by server device 100. Therefore, network 30 is connected to server device 100 that implements voice communication device 10.

[0098] The voice communication device 10 is connected to the network 30 and is implemented by the server device 100.

[0099] Figure 2 This is a schematic diagram illustrating an example of the configuration of a server device 100 that implements the voice communication device 10.

[0100] like Figure 2 The server device 100 shown includes: an input device 101, an output device 102, a CPU (Central Processing Unit) 103, internal storage 104, RAM (Random Access Memory) 105, and a bus 106.

[0101] The input device 101 is a user interface device such as a keyboard, mouse, or touchpad, which accepts operations from the user using the server device 100. In addition to accepting touch operations from the user, the input device 101 may also accept operations via voice, remote control, or other methods.

[0102] The output device 102 is a device that serves as a user interface, such as a display, speaker, or output terminal, and outputs signals from the server device 100 to the outside.

[0103] Built-in memory 104 is a storage device such as flash memory, used to store programs executed by server device 100, data used by server device 100, etc.

[0104] RAM105 is a storage device called SRAM (Static RAM), DRAM (Dynamic RAM), etc., used as a temporary storage area when executing programs.

[0105] CPU 103 copies the program stored in built-in memory 104 to RAM 105, and reads and executes the commands contained in the copied program from RAM 105 in sequence.

[0106] Bus 106 is connected to input device 101, output device 102, CPU 103, internal memory 104, and RAM 105, and transmits signals between the connected components.

[0107] Although Figure 2Without illustration, server device 100 has communication capabilities. Server device 100 connects to network 30 via this communication capability.

[0108] The voice communication device 10 is implemented, for example, by having the CPU 103 copy the program stored in the built-in memory 104 to the RAM 105, and then sequentially reading and executing the commands contained in the copied program from the RAM 105.

[0109] Figure 3 This is a block diagram showing an example of the configuration of the voice communication device 10.

[0110] like Figure 3 As shown, the voice communication device 10 includes: N input units 11 (and... Figure 3 (corresponding to the first input section 11A to the fifth input section 11E), sound image position determination section 12, and N sound image positioning sections 13 (corresponding to the first input section 11A to the fifth input section 11E), and N sound image positioning sections 13 (corresponding to the first input section 11A to the fifth input section 11E). Figure 3 The first acoustic image positioning unit 13A to the fifth acoustic image positioning unit 13E correspond to the first acoustic image positioning unit 13A to the fifth acoustic image positioning unit 13E, the addition unit 14, and the output unit 15.

[0111] The first input units 11A to the fifth input units 11E are respectively connected to the first sound image positioning units 13A to the fifth sound image positioning units 13E, and are used to input sound signals output from any one of the terminals 20. Specifically, the first sound signal output from terminal 20A is input to the first input unit 11A, the second sound signal output from terminal 20B is input to the second input unit 11B, the third sound signal output from terminal 20C is input to the third input unit 11C, the fourth sound signal output from terminal 20D is input to the fourth input unit 11D, and the fifth sound signal output from terminal 20E is input to the fifth input unit 11E. Furthermore, it is explained here that the first audio signal includes an electrical signal that transforms the voice emitted by the user of the first terminal 20A (here, user 23A), the second audio signal includes an electrical signal that transforms the voice emitted by the user of the second terminal 20B (here, user 23B), the third audio signal includes an electrical signal that transforms the voice emitted by the user of the third terminal 20C (here, user 23C), the fourth audio signal includes an electrical signal that transforms the voice emitted by the user of the fourth terminal 20D (here, user 23D), and the fifth audio signal includes an electrical signal that transforms the voice emitted by the user of the fifth terminal 20E (here, user 23E).

[0112] The first input section 11A to the fifth input section 11E have the same function. Therefore, in this specification, when it is not necessary to distinguish between the first input section 11A to the fifth input section 11E, they are referred to as input section 11.

[0113] Output unit 15, connected to addition unit 14, outputs the summed sound image positioning sound signal (described later) output from addition unit 14 to any one of terminals 20. Specifically, output unit 15 outputs the summed sound image positioning sound signal to terminal 20F.

[0114] The sound image position determination unit 12 is connected to the first sound image positioning units 13A to the fifth sound image positioning units 13E, and is used to determine the position of N sound signals (and) input from the N input units 11. Figure 3 Each of the first to fifth sound signals in the diagram (corresponding to each other) determines the presence of the first wall 41 (see below). Figure 4 ) and the second wall 42 (see below) Figure 4 The location of sound and image in the virtual space.

[0115] Figure 4 This is a schematic diagram showing how the sound image location determination unit 12 determines the sound image positioning position in virtual space for each of N sound signals.

[0116] like Figure 4 The virtual space 90 shown includes: a first wall 41, a second wall 42, a first audio-visual position 51, a second audio-visual position 52, a third audio-visual position 53, a fourth audio-visual position 54, a fifth audio-visual position 55, and a listener position 50.

[0117] Wall 41 and Wall 42 are virtual walls that exist in the virtual space and reflect sound waves.

[0118] Listener position 50 is a virtual position of the listener who listens to the sounds indicated by the first to fifth sound signals.

[0119] The first sound image position 51 is the sound image position determined by the sound image position determination unit 12 for the first sound signal. The second sound image position 52 is the sound image position determined by the sound image position determination unit 12 for the second sound signal. The third sound image position 53 is the sound image position determined by the sound image position determination unit 12 for the third sound signal. The fourth sound image position 54 is the sound image position determined by the sound image position determination unit 12 for the fourth sound signal. The fifth sound image position 55 is the sound image position determined by the sound image position determination unit 12 for the fifth sound signal.

[0120] like Figure 4As shown, the sound image position determination unit 12 determines the sound image positioning positions of N sound image signals (here, the first sound image position 51 to the fifth sound image position 55) to be located between the first wall 41 and the second wall 42, and to be in a non-overlapping position when viewed from the listener position 50. More specifically, the sound image position determination unit 12 determines the sound image positioning positions of the N sound image signals in such a way that, when the front of the virtual listener at the listener position 50 is set to 0 degrees, the interval between sound image positioning positions that include 0 degrees or are adjacent to each other with 0 degrees between them is narrower than the interval between sound image positioning positions that do not include 0 degrees or are adjacent to each other without being adjacent to each other with 0 degrees between them.

[0121] Therefore, as Figure 4 As shown, when the angle between the first sound image position 51 and the second sound image position 52 as viewed from the listener position 50 is set as angle X, and the angle between the second sound image position 52 and the third sound image position 53 as viewed from the listener position 50 is set as angle Y, X>Y.

[0122] Back to Figure 3 Continuing with the description of the voice communication device 10.

[0123] The first sound image positioning unit 13A is connected to the first input unit 11A, the sound image position determination unit 12, and the addition unit 14. It performs sound image positioning processing and outputs a sound image positioning sound signal. This sound image positioning processing positions the sound image at the first sound image position 51 determined by the sound image position determination unit 12. The second sound image positioning unit 13B is connected to the second input unit 11B, the sound image position determination unit 12, and the addition unit 14. It performs sound image positioning processing and outputs a sound image positioning sound signal. This sound image positioning processing positions the sound image at the second sound image position 52 determined by the sound image position determination unit 12. The third sound image positioning unit 13C is connected to the third input unit 11C, the sound image position determination unit 12, and the addition unit 14. It performs sound image positioning processing and outputs a sound image positioning sound signal. This sound image positioning processing positions the sound image at the third sound image position 53 determined by the sound image position determination unit 12. The fourth sound image positioning unit 13D is connected to the fourth input unit 11D, the sound image position determination unit 12, and the addition unit 14. It performs sound image positioning processing and outputs a sound image positioning sound signal. This sound image positioning processing positions the sound image at the fourth sound image position 54 determined by the sound image position determination unit 12. The fifth sound image positioning unit 13E is connected to the fifth input unit 11E, the sound image position determination unit 12, and the addition unit 14. It performs sound image positioning processing and outputs a sound image positioning sound signal. This sound image positioning processing positions the sound image at the fifth sound image position 55 determined by the sound image position determination unit 12.

[0124] The functions of the first acoustic image positioning unit 13A to the fifth acoustic image positioning unit 13E are the same. Therefore, in this specification, when it is not necessary to distinguish between the first acoustic image positioning unit 13A to the fifth acoustic image positioning unit 13E, they are referred to as acoustic image positioning unit 13.

[0125] More specifically, the sound image localization unit 13 performs sound image localization processing using a first head-related transfer function (HRTF) and a second head transfer function. The first head transfer function is a function that simulates the sound waves emitted from the sound image position determined by the sound image position determination unit 12 directly reaching the ears of a virtual listener at the listener position 50. The second head transfer function is a function that simulates the sound waves emitted from the sound image position determined by the sound image position determination unit 12 being reflected by the wall closer to the sound image position between the first wall and the second wall, thus reaching the ears of a virtual listener at the listener position 50.

[0126] Figure 5 This is a schematic diagram showing how the sound image positioning unit 13 performs sound image positioning processing.

[0127] exist Figure 5 In this context, speaker 71 is a speaker virtually existing at the first sound image position 51, speaker 72 is a speaker virtually existing at the second sound image position 52, speaker 73 is a speaker virtually existing at the third sound image position 53, speaker 74 is a speaker virtually existing at the fourth sound image position 54, and speaker 75 is a speaker virtually existing at the fifth sound image position 55. Listener 60 is a listener virtually existing at listener position 50.

[0128] Speaker 71 is, for example, the icon of user 23A; speaker 72 is, for example, the icon of user 23B; speaker 73 is, for example, the icon of user 23C; speaker 74 is, for example, the icon of user 23D; speaker 75 is, for example, the icon of user 23E; and listener 60 is, for example, the icon of user 23F.

[0129] Furthermore, speaker 71A is a virtual mirror image of speaker 71 existing in a mirror position when the first wall 41 is a mirror, and speaker 74A is a virtual mirror image of speaker 74 existing in a mirror position when the second wall 42 is a mirror.

[0130] like Figure 5 As shown, in virtual space 90, for example, the sound emitted by the first speaker 71 reaches the ears of the listener 60 directly through the transmission path shown by the two solid lines. In addition, the sound emitted by the first speaker 71 reaches the ears of the listener by being reflected by the first wall 41 through the transmission path shown by the two dashed lines.

[0131] Therefore, in virtual space 90, for the sound emitted by the first speaker 71, two signals are generated by convolving the sound with the first head transfer function corresponding to each of the two transmission paths shown by solid lines, and two more signals are generated by convolving the sound with the second head transfer function corresponding to each of the two transmission paths shown by dashed lines. When these signals are summed, and the listener 60 listens, for example using headphones, the listener 60 hears what sounds like the sound emitted by the first speaker 71 at the first sound image position. At this time, because the listener 60 also hears the sound reflected from the first wall 41, the listener 60 can perceive that virtual space 90 is a virtual space with walls.

[0132] like Figure 5 As shown, in virtual space 90, for example, the sound emitted by the fourth speaker 74 reaches the ears of the listener 60 directly through the transmission path shown by the two solid lines. Furthermore, the sound emitted by the fourth speaker 74 reaches the listener's ears by being reflected off the second wall 42 through the transmission path shown by the two dashed lines.

[0133] Therefore, in virtual space 90, for the sound emitted by the fourth speaker 74, two signals are generated by convolving the first head transfer function corresponding to each of the two transmission paths shown by solid lines, and two more signals are generated by convolving the second head transfer function corresponding to each of the two transmission paths shown by dashed lines. When these signals are summed, for example, when the listener 60 listens using headphones, the listener 60 hears what sounds like the sound emitted by the fourth speaker 74 at the fourth sound image position. At this time, because the listener 60 also hears the sound reflected from the second wall 42, the listener 60 can perceive that virtual space 90 is a virtual space with walls.

[0134] At this time, the sound image localization unit 13 can perform sound image localization processing in a manner that allows free modification of at least one of the reflectivity of the sound waves from the first wall 41 and the reflectivity of the sound waves from the second wall 42. By changing the reflectivity, the echo intensity of the sound in the virtual space 90 can be changed.

[0135] Furthermore, at this time, the sound image positioning unit 13 can perform sound image positioning processing in a manner that allows free change of at least one of the positions of the first wall 41 and the second wall 42. By changing the position of the walls, the degree of spatial expansion in the virtual space 90 can be changed.

[0136] Furthermore, the sound image position determination unit 12 can of course further utilize the third head transfer function to perform sound processing. The third head transfer function is a function that simulates the sound waves emitted from the sound image position determined by the sound image position determination unit 12, which are reflected by the wall of the first wall 41 and the second wall 42 that is farthest from the sound image position and reach the ears of the listener 60.

[0137] Back to Figure 3 Continuing with the description of the voice communication device 10.

[0138] The addition unit 14 is connected to the N sound image positioning units 13 and the output unit 15, and adds the N sound image positioning sound signals output from the N sound image positioning units 13, and outputs the added sound image positioning sound signal.

[0139] The voice communication device 10 provides an atmosphere where the voices of N speakers (5 people in this case) inputting from each of the N (5 in this case) input units 11 appear to originate from a virtual space 90 with a first wall 41 and a second wall 42. Furthermore, the listener 60, hearing the voices of the N speakers through the voice communication device 10, can more easily grasp the positional relationship between the speakers and the walls in the virtual space 90. Therefore, the listener 60 can more easily distinguish the direction from which the voices of the N speakers originate. Thus, the voice communication device 10 enhances the sense of presence experienced by participants in remote meetings, online gatherings, and similar events conducted using voice communication devices compared to previous methods.

[0140] As mentioned above, regarding the auditory sharpness of typical sound image localization, it is known that the more directly a listener is perceived, the less sensitive the sound becomes, while the more to the left or right, the less sensitive it becomes. Through the voice communication device 10, from the listener 60's perspective, the angle between speakers positioned to the left or right is larger than the angle between speakers positioned directly in front. Therefore, the listener 60 can more easily distinguish the direction from which the voices of N speakers originate. Consequently, through the voice communication device 10, the sense of presence experienced by participants in remote meetings, online gatherings, and similar events conducted using voice communication devices can be significantly improved compared to previous methods.

[0141] (Implementation Method 2)

[0142] The following describes a voice communication device according to Embodiment 2, which is a modified version of the voice communication device 10 according to Embodiment 1.

[0143] In the following description of the voice communication device according to Embodiment 2, the same constituent elements as those of the voice communication device 10 are referred to as the already described parts, and are given the same reference numerals. Detailed descriptions are omitted, and the description focuses on the differences from the voice communication device 10.

[0144] Figure 6 This is a block diagram illustrating an example of the configuration of the voice communication device 10A according to Embodiment 2.

[0145] like Figure 6 As shown, the voice communication device 10A according to Embodiment 2 is configured as follows: a second addition unit 16, a background noise signal storage unit 17, and a selection unit 18 are added to the voice communication device 10. In addition, the output unit 15 is changed to the output unit 15A.

[0146] The background noise signal storage unit 17 is connected to the selection unit 18 and stores one or more background noise signals that show the background noise in the virtual space 90.

[0147] The background noise signal can represent background noise, such as pre-recorded background noise in a real conference room. Alternatively, it can represent pre-recorded loud noise in a real bar, izakaya, or concert hall. It can also represent jazz music played in a real jazz teahouse. Furthermore, the background noise signal can be an artificially synthesized signal, such as an artificial signal generated by synthesizing multiple pre-recorded loud noises in a real space.

[0148] The selection unit 18 is connected to the background noise signal storage unit 17 and the second addition unit 16, and selects one or more background noise signals from one or more background noise signals stored in the background noise signal storage unit 17.

[0149] The selection unit 18 can, for example, change the selected background noise signal over time.

[0150] The second addition unit 16 is connected to the addition unit 14, the selection unit 18, and the output unit 15A. It adds the summed sound image positioning sound signal output from the addition unit 14 and the background noise signal selected by the selection unit 18, and outputs the second summed sound image positioning sound signal.

[0151] Output unit 15A, connected to the second addition unit 16, outputs the second-added sound image positioning audio signal output from the second addition unit 16 to any one of the terminals 20. Here, output unit 15A is described as inputting the second-added sound image positioning audio signal to terminal 20F.

[0152] The voice communication device 10A can provide an atmosphere where the voices of N speakers (5 people in this case) input from each of the N (5 in this case) input units 11 appear to be coming from a virtual space 90 filled with background noise. Therefore, for example, if the selection unit 18 selects a background noise signal that is pre-recorded in a real meeting room, an atmosphere as if the virtual space 90 were a real meeting room can be created. Furthermore, for example, if the selection unit 18 selects a background noise signal that is pre-recorded in a real bar, izakaya, or concert hall, an atmosphere as if the virtual space 90 were a real bar, izakaya, or concert hall can be created. Furthermore, for example, if the selection unit 18 selects a background noise signal that is jazz music played in a real jazz teahouse, an atmosphere as if the virtual space 90 were a real jazz teahouse can be created. Thus, the voice communication device 10A can significantly enhance the sense of presence experienced by participants in remote meetings, online dinners, and other similar events conducted using voice communication devices.

[0153] Furthermore, the background noise can be selected according to the atmosphere of the desired virtual space 90 through the voice communication device 10A.

[0154] Furthermore, the atmosphere of the virtual space 90 can be changed over time via the aforementioned voice communication device 10A.

[0155] (Other implementation methods)

[0156] The above description of the voice communication device of this disclosure has been based on Embodiment 1 and Embodiment 2, but this disclosure is not limited to these embodiments. For example, other embodiments that combine any of the constituent elements described in this specification, or exclude some of the constituent elements, may also be embodiments of this disclosure. Furthermore, various modifications conceived by those skilled in the art, implemented with respect to the described embodiments without departing from the spirit of this disclosure, i.e., the meaning indicated by the text describing the technical solution, are also included in this disclosure.

[0157] (1) In Embodiment 1 and Embodiment 2, the voice communication device 10 and the voice communication device 10A are configuration examples where N is 5. However, the voice communication device disclosed herein does not necessarily have to be limited to a configuration example where N is 5, as long as N is an integer of 2 or more.

[0158] (2) In Embodiment 1, the voice communication device 10 is described as follows: a first voice signal to a fifth voice signal are input to terminals 20A to 20E respectively, and the summed audio-visual positioning signal is output to terminal 20F. The voice communication device 10 can be modified as follows: a first modified voice communication device to a fifth modified voice communication device. The first modified voice communication device is configured such that the first voice signal to a fifth voice signal are input from terminals 20B to 20F respectively, the summed audio-visual positioning signal is output to terminal 20A. The second modified voice communication device is configured such that the first voice signal to a fifth voice signal are input from terminals 20C to 20F and terminal 20A respectively, the summed audio-visual positioning signal is output to terminal 20B. The third modified voice communication device is configured such that the first voice signal to a fifth voice signal are input from terminals 20D to 20F and terminals 20A to 20B respectively, the summed audio-visual positioning signal is output to terminal 20C. The fourth deformable sound communication device is configured such that the first to fifth sound signals are input from terminals 20E to 20F and terminals 20A to 20C respectively, are added together to form a sound image positioning sound signal, and are output to terminal 20D. The fifth deformable sound communication device is configured such that the first to fifth sound signals are input from terminals 20F and terminals 20A to 20D respectively, are added together to form a sound image positioning sound signal, and are output to terminal 20E.

[0159] Furthermore, the voice communication device 10 and the first to fifth modified voice communication devices can be implemented simultaneously by the server device 100. For example, the server device 100 can implement the voice communication device 10 and the first to fifth modified voice communication devices simultaneously through time-division processing, or it can implement the voice communication device 10 and the first to fifth modified voice communication devices simultaneously through parallel processing.

[0160] Furthermore, a voice communication device can be implemented by server device 100, which can realize the functions obtained by simultaneously implementing voice communication device 10, the first modified voice communication device to the fifth modified voice communication device.

[0161] (3) In Embodiment 2, the voice communication device 10A is described as follows: the first to fifth voice signals are input from terminals 20B to 20E, respectively, and the second summed acoustic image positioning voice signal is output to terminal 20F. Therefore, the voice communication device 10A can be modified into the following sixth to tenth modified voice communication devices. The sixth modified voice communication device is configured such that the first to fifth voice signals are input from terminals 20B to 20F, respectively, and the second summed acoustic image positioning voice signal is output to terminal 20A. The seventh modified voice communication device is configured such that the first to fifth voice signals are input from terminals 20C to 20F and terminal 20A, respectively, and the second summed acoustic image positioning voice signal is output to terminal 20B. The eighth deformable voice communication device is configured such that the first to fifth voice signals are input from terminals 20D to 20F and terminals 20A to 20B, respectively, and a second summed acoustic image positioning voice signal is output to terminal 20C. The ninth deformable voice communication device is configured such that the first to fifth voice signals are input from terminals 20E to 20F and terminals 20A to 20C, respectively, and a second summed acoustic image positioning voice signal is output to terminal 20D. The tenth deformable voice communication device is configured such that the first to fifth voice signals are input from terminals 20F and terminals 20A to 20D, respectively, and a second summed acoustic image positioning voice signal is output to terminal 20E.

[0162] Furthermore, the voice communication device 10A and the sixth to fifth modified voice communication devices can be simultaneously implemented by the server device 100. For example, the server device 100 can implement the voice communication device 10A and the sixth to tenth modified voice communication devices simultaneously through time-sharing processing, or it can implement the voice communication device 10A and the sixth to tenth modified voice communication devices simultaneously through parallel processing. In this case, the selection unit 18 included in the voice communication device 10A and the sixth to tenth modified voice communication devices can be configured to select the same background noise signal. As a result, the sense of presence experienced by participants in remote meetings, online dinners, etc., held using voice communication devices can be improved more than before.

[0163] The server device 100 can then implement such a voice communication device, which can achieve the functions obtained by simultaneously implementing the voice communication device 10A and the sixth to tenth modified voice communication devices.

[0164] (4) Some or all of the constituent elements of the voice communication device 10 and the voice communication device 10A may be constituted by a single system LSI (Large Scale Integration). A system LSI is a multi-functional LSI manufactured by integrating multiple components onto a single chip; specifically, it is a computer system comprising a microprocessor, ROM (Read Only Memory), RAM (Random Access Memory), etc. The ROM records the computer program. The microprocessor operates according to the computer program, thereby enabling the system LSI to perform its functions.

[0165] Furthermore, while referred to here as a system LSI, it is also known as an IC, LSI, VLSI, or extra-large LSI, depending on the level of integration. Moreover, the method of integrated circuitization is not limited to LSI; it can be implemented using dedicated circuits or general-purpose processors. Alternatively, FPGAs (Field Programmable Gate Arrays) that are programmable after LSI manufacturing, or reconfigurable LSIs with interconnected circuitry and configured reconfigurable processors can be used.

[0166] Furthermore, with advancements in semiconductor technology or the emergence of other derived technologies, when integrated circuit technologies emerge that can replace LSIs, these technologies can certainly be used for the integration of functional blocks. This may also be applicable to technologies such as biotechnology.

[0167] (5) Each component of the voice communication device 10 and the voice communication device 10A can be constructed by dedicated hardware, and can be implemented by a program execution unit such as a CPU or processor, which reads and executes a software program recorded in a recording medium such as a hard disk or semiconductor memory.

[0168] This disclosure can be widely used in remote conferencing systems, etc.

Claims

1. A voice communication device, comprising: N input sections are used to input audio signals, where N is an integer greater than or equal to 2; The sound image location determination unit determines the sound image location position in a virtual space having a first wall and a second wall for each of the N sound signals input from the N input units. There are N sound image localization units, each of which corresponds to each of the N input units. Each of the N sound image localization units performs sound image localization processing and outputs a sound image localization sound signal. The sound image localization processing is the process of localizing the sound image at the sound image localization position. The sound image localization position is the position determined by the sound image position determination unit for the input unit corresponding to that sound image localization unit. The addition unit adds the N sound image positioning signals output from the N sound image positioning units and outputs the added sound image positioning signal. The sound image location determination unit determines the sound image positioning positions of the N sound signals in such a way that the sound image positioning positions of the N sound signals are located between the first wall and the second wall, and are in non-overlapping positions when viewed from the listener's position between the first wall and the second wall. Each of the N sound image localization units performs the sound image localization processing using a first head transfer function and a second head transfer function. The first head transfer function is a function that simulates how sound waves emitted from the sound image localization position directly reach the ears of a virtual listener at the listener's position. The sound image localization position is a position determined by the sound image position determination unit for the sound image localization unit. The second head transfer function is a function that simulates how sound waves emitted from the sound image localization position are reflected from the wall closer to the sound image localization position between the first wall and the second wall, and reach the listener's ears.

2. The voice communication device as described in claim 1, Each of the N acoustic image localization units performs the acoustic image localization process in a manner that allows for free variation of at least one of the reflectivity of the sound waves from the first wall and the reflectivity of the sound waves from the second wall.

3. The voice communication device as described in claim 1 or 2, Each of the N acoustic image positioning units performs the acoustic image positioning process in a manner that allows for free change of at least one of the positions of the first wall and the second wall.