Sound field reproduction program, apparatus, and method using position information of a sound source and a sound receiver

The sound field reproduction program and apparatus address the challenge of reproducing sound fields with movable sources and receivers by using position-based microphone and speaker identification, ensuring accurate and immersive sound reproduction.

JP7705228B2Active Publication Date: 2025-07-09KDDI CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022108891
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-07-06
Publication Date
2025-07-09
Estimated Expiration
2042-07-06

AI Technical Summary

Technical Problem

Existing sound field reproduction technologies struggle to accurately reproduce the sound field when the sound source is at an arbitrary or movable position within the sound collection space, or when the sound receiver is at an arbitrary or movable position within the reproduction space, due to disturbances caused by the presence and movement of listeners.

Method used

A sound field reproduction program and apparatus that utilizes position information acquisition, input acoustic signal determination, and output acoustic signal generation to recreate the sound field by identifying appropriate microphones and speakers based on the positions of the sound source and receiver, even when they are movable, using a combination of omnidirectional and highly directional microphones, and adjusting amplitude and phase to ensure accurate sound reproduction.

Benefits of technology

Enables accurate reproduction of the sound field regardless of the positions or movements of the sound source and receiver, providing a high sense of presence and fidelity in sound reproduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007705228000001
    Figure 0007705228000001
  • Figure 0007705228000002
    Figure 0007705228000002
  • Figure 0007705228000003
    Figure 0007705228000003
Patent Text Reader

Abstract

To provide a sound field reproduction program with which it is possible to reproduce the sound field of a sound collection space for a sound-receiving medium even when a sound source and / or the sound-receiving medium is present at a discretionary position or movable.SOLUTION: The present program causes a computer to function as: position information acquisition means for acquiring measured position information at a sound source and a sound-receiving medium; input acoustic signal determination means for identifying, on the basis of the acquired position information, a microphone located in a direction leading to the corresponding position of the sound-receiving medium or a direction close by a prescribed distance or more to said direction in a sound-collection space, as seen from the sound source, and acquiring an acoustic signal that pertains to the sound collected by the identified microphone to determine an input acoustic signal; and output acoustic signal generation means for generating, on the basis of the acquired position information, an output acoustic signal, using the input acoustic signal, the output acoustic signal causing a sound equivalent to the collected sound to be outputted toward the sound-receiving medium from a direction leading toward the corresponding position of the sound source or a direction close by a prescribed distance or more to said direction in the reproduction space, as seen from the sound-receiving medium.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a sound field reproduction technology including sound field reproduction.

Background Art

[0002] In recent years, due to the spread of Web conferencing applications, remote session applications, etc., services that provide a suitable or desired acoustic environment using microphones and speakers have been widely used. Also, the utilization of high-immersion remote conferencing systems and online concerts is being actively promoted. Here, in the provision of such services, improvement of the sound field reproduction (sound field reproduction) technology for reproducing (reproducing) the sound field related to the sound source at the reproduction site is an important issue.

[0003] As the basis of this sound field reproduction (sound field reproduction) technology, for example, Non-Patent Document 1 introduces a sound field control method using a wavefront synthesis method for spatially generating wavefronts. Specifically, for example, the principle of boundary sound field control that applies the inverse system theory to the Kirchhoff-Helmholtz integral equation, which is a mathematical expression of the Huygens principle, is explained.

[0004] Also, as a sound field reproduction method based on sound pressure gradient, impulse response, etc. as disclosed in this Non-Patent Document 1, for example, Patent Document 1 discloses a high-immersion sound field reproduction information transmission device that enables high-immersion sound field reproduction with a balance between recorded sound and reflected sound on the receiving side. Specifically, this device is a device that transmits a dry source sound signal, which is a signal acquired from a sound source in an anechoic environment, and sound field reproduction information, which is information for reproducing the dry source sound signal corresponding to the environment on the receiving side.

[0005] Here, the sound source has conventionally been positioned outside the control area where the sound is to be reproduced. In contrast, for example, in Non-Patent Document 2, attempts have been made to reproduce the sound field when the sound source is located within the control area. Specifically, a sound field reproduction method has been proposed in which a virtual sound source signal is estimated from the acoustic signals picked up by a microphone array composed of a plurality of microphones using an inverse filter, and a virtual reality system with fewer constraints on the sound source is constructed.

[0006] Furthermore, for example, Non-Patent Document 3 discloses a sound field reproduction method that exhibits high robustness against such disturbances in situations where disturbances are inevitably generated, such as when the listener himself / herself is present in the reproduction field where the sound is to be reproduced. Here, generally, sound field reproduction is considered to be completed by making the propagation of sound in the pickup field and the reproduction field match. However, for example, even if it becomes possible to reproduce the propagation of sound in the pickup field and the reproduction field when there is no person, the propagation of sound will change when a person enters the reproduction field. That is, it is inevitable that disturbances will occur due to the presence of a person in the reproduction field.

[0007] Also, Non-Patent Document 3 proposes a sound pickup method using 24 highly directional microphones and a signal processing method for reproducing the signals generated by such sound pickup. According to this signal processing method, even in the case of simply picking up sound and outputting it, the impulse response is measured to directly remove the direct sound component, and it is said that it is possible to express spatial acoustic characteristics such that, for example, the movement of the sound source can be felt.

[0008] Furthermore, for example, Patent Document 2 discloses an acoustic field reproduction device that adopts a rigid sphere model including a sound source and a rigid sphere whose position is known, and reproduces the acoustic field at a listening position on the side opposite to the sound source with respect to the rigid sphere. Specifically, this device includes a transfer function storage unit for supplying transfer functions obtained in advance at a plurality of locations on a predetermined line having a predetermined relationship with respect to the line connecting the rigid sphere and the sound source in the space, a position estimation unit for determining information regarding the position where the acoustic field is to be reproduced, and an output signal synthesis unit that receives the supply of the corresponding transfer function from the transfer function storage unit based on the information regarding the determined position, converts the input sound source signal according to this transfer function, synthesizes the acoustic signals, and outputs them. Here, it is assumed that such a functional configuration enables real-time reproduction of the acoustic field in response to changes in the arrangement of objects in the space.

Prior Art Documents

Patent Documents

[0009]

Patent Document 1

Patent Document 2

Non-Patent Documents

[0010]

Non-Patent Document 1

Non-Patent Document 2

Non-Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0011] For example, the sound field reproduction method disclosed in Non-Patent Document 1 and Patent Document 1, which is performed based on a previously measured sound pressure gradient, impulse response, etc., is a basic method aimed at improving reproduction accuracy. However, the sound pressure gradient and impulse response are easily affected by the presence of the listener and also change greatly depending on the movement of the listener. Therefore, with such a sound field reproduction method, it becomes difficult to enhance the robustness against disturbances caused by the presence of people in the reproduction field.

[0012] Also, the sound field reproduction technology disclosed in Non-Patent Document 2, which uses an acoustic signal picked up by a microphone array, is a method that can handle the situation where a sound source is placed within the control area, but it does not generate a virtual sound source signal corresponding to the movement of the listener. That is, similar to the technologies disclosed in Non-Patent Document 1 and Patent Document 1 above, sound field reproduction corresponding to the presence and movement of the listener still remains difficult.

[0013] On the other hand, the sound field reproduction technology disclosed in Non-Patent Document 3 uses 24 highly directional microphones for sound pickup and is a method with high robustness against the presence of the listener. However, the listening position in the reproduction field is limited to the position corresponding to the position where the highly directional microphone array is installed in the sound pickup field. That is, since the auditory characteristics are different at other positions in the reproduction field, ultimately, it has to be said that sound field reproduction corresponding to the movement of the listener is still difficult.

[0014] On the other hand, although the sound field reproduction apparatus disclosed in Patent Document 2 certainly reproduces changes in the characteristics in sound perception due to changes in the listening position, in this sound field reproduction apparatus, it is a major premise that the input sound source is a pre-recorded dry sound source and its position is known.

[0015] Therefore, an object of the present invention is to provide a sound field reproduction program, apparatus, and method capable of reproducing the sound field of a sound collection space for a sound receiving body even when the sound source exists at an arbitrary position in the sound collection space or is movable within the sound collection space and / or when the sound receiving body exists at an arbitrary position in the reproduction space or is movable within the reproduction space.

Means for Solving the Problems

[0016] According to the present invention, there is provided a sound field reproduction program for reproducing, for a sound receiving body in a reproduction space having a correspondence relationship with a sound collection space with respect to a position in the space, a sound field related to a sound source in the sound collection space, a position information acquisition means for acquiring information related to the measured or set positions of the sound source and the sound receiving body; an input acoustic signal determination means for specifying at least one microphone located in a direction from the sound source in the sound collection space toward the corresponding position of the sound receiving body or in a direction close to the direction until a predetermined condition is satisfied in the direction, and acquiring an acoustic signal related to the sound picked up by the specified microphone to determine an input acoustic signal; an output acoustic signal generation means for generating, based on the information related to the acquired position, an output acoustic signal that outputs a sound corresponding to or corresponding to the picked-up sound toward the sound receiving body from a direction from the sound receiving body in the reproduction space toward the corresponding position of the sound source or in a direction close to the direction until a predetermined condition is satisfied in the direction, using the input acoustic signal; and a sound field reproduction program for causing a computer to function is provided.

[0017] As a preferred embodiment of the sound field reproduction program according to the present invention, the position information acquisition means also acquires information related to the measured or set sound output direction in the sound source, The input acoustic signal determination means identifies at least one other microphone located in the sound collection space in the output direction or in a direction close to the output direction until a predetermined condition is satisfied as viewed from the sound source based on the acquired information related to the position and the information related to the output direction, acquires another acoustic signal related to the sound picked up by the identified other microphone, and determines another input acoustic signal, It is also preferable that the output acoustic signal generation means generates the output acoustic signal by combining the other input acoustic signal and the input acoustic signal related to the sound picked up by the microphone.

[0018] Also, in the above embodiment, it is also preferable that the other microphone has higher directivity than the microphone.

[0019] Furthermore, as another embodiment of the sound field reproduction program according to the present invention, the reproduction space is a real space, The output acoustic signal generation means identifies at least one speaker located in the reproduction space in the direction toward the corresponding position of the sound source as viewed from the sound receiving body or in a direction close to the direction until a predetermined condition is satisfied based on the acquired information related to the position, and it is also preferable to generate the output acoustic signal to be supplied to the identified speaker.

[0020] Also, in the above other embodiment, it is also preferable that the output acoustic signal generation means also identifies at least one other speaker that is closer to the position of the sound receiving body than the identified speaker or is located more inside as viewed from the surrounding boundary, and generates the output acoustic signal to be supplied to the identified other speaker.

[0021] Furthermore, in the above other embodiment, the sound source and the sound receiving body are each movable within the sound collection space and the reproduction space, which are real spaces, Based on the information related to the position measured and obtained at one point in time by the position information acquisition means, the input acoustic signal determination means identifies the microphone at the one point in time and determines the input acoustic signal related to the one point in time. It is also preferable that the output acoustic signal generation means identifies the speaker at the one point in time based on the information related to the position obtained at the one point in time and generates the output acoustic signal related to the one point in time.

[0022] Furthermore, as still another embodiment of the sound field reproduction program according to the present invention, the reproduction space is a virtual space related to sound image localization in a stereophonic headset worn on the sound receiver. It is also preferable that the output acoustic signal generation means generates the output acoustic signal to be supplied to the stereophonic headset, which outputs a sound corresponding to or corresponding to the picked-up sound, from a direction toward the corresponding position of the sound source as seen from the sound receiver in the reproduction space or a direction close to the direction until a predetermined condition is satisfied, toward the sound receiver.

[0023] Also, in the embodiment using the above-mentioned "another microphone", it is also preferable that the output acoustic signal generation means performs an amplitude adjustment process of attenuating the amplitude according to the distance between the position of the sound source and the position of the sound receiver on the determined another input acoustic signal, and generates an output acoustic signal using the another input acoustic signal subjected to the process.

[0024] Furthermore, in the embodiment using the above-mentioned "another microphone", the output acoustic signal generation means performs a filtering process of emphasizing a predetermined high frequency band according to the proximity between the position of the sound source and the position of the sound receiver on the determined another input acoustic signal, or determines an impulse response between the position of the sound source and the position of the sound receiver and performs a convolution process in the frequency domain related to the impulse response, and generates an output acoustic signal using the another input acoustic signal subjected to the process.

[0025] Furthermore, in the embodiment using the "another microphone" described above, it is also preferable that the output acoustic signal generation means determines the phase difference between the another input acoustic signal and the input acoustic signal based on the information related to the position, and after delaying or advancing one of the phases so as to cancel the determined phase difference, combines the another input acoustic signal and the input acoustic signal to generate the output acoustic signal.

[0026] According to the present invention, there is also provided an acoustic field reproduction program for reproducing the sound related to a sound source in a sound collection space for a sound receiver in a reproduction space that has a corresponding relationship with the sound collection space with respect to the position in the space, position information acquisition means for acquiring information related to the measured or set positions of the sound source and the sound receiver, and information related to the measured or set sound output direction of the sound source; (a) identifying at least one microphone located in the sound collection space in a direction toward the corresponding position of the sound receiver as seen from the sound source or in a direction close to the direction until a predetermined condition is satisfied in the direction, acquiring an acoustic signal related to the sound picked up by the identified microphone to determine an input acoustic signal, and (b) identifying at least one another microphone located in the sound collection space in the output direction or in a direction close to the output direction until a predetermined condition is satisfied as seen from the sound source, based on the acquired information related to the position and the information related to the output direction, and acquiring another acoustic signal related to the sound picked up by the identified another microphone to determine another input acoustic signal; output acoustic signal generation means for combining the input acoustic signal and the another input acoustic signal to generate an output acoustic signal that outputs a sound corresponding to or corresponding to the picked-up sound toward the sound receiver; An acoustic field reproduction program for causing a computer to function is provided.

[0028] According to the present invention, there is further provided an acoustic field reproduction apparatus for reproducing the acoustic field related to a sound source in a sound collection space for a sound receiver in a reproduction space that has a corresponding relationship with the sound collection space with respect to the position in the space, Position information acquisition means for acquiring information related to the measured or set positions of the sound source and the sound receiver; Input acoustic signal determination means for identifying at least one microphone positioned in a direction from the sound source in the sound collection space toward the corresponding position of the sound receiver or in a direction close to the direction until a predetermined condition is satisfied in that direction, based on the acquired information related to the position, and acquiring an acoustic signal related to the sound picked up by the identified microphone to determine an input acoustic signal; Output acoustic signal generation means for generating, based on the acquired information related to the position, an output acoustic signal that will output a sound equivalent or corresponding to the picked-up sound toward the sound receiver from a direction from the sound receiver in the playback space toward the corresponding position of the sound source or in a direction close to the direction until a predetermined condition is satisfied in that direction, using the input acoustic signal; A sound field reproduction apparatus having the above is provided.

[0029] According to the present invention, there is also provided a sound field reproduction system for reproducing a sound field related to a sound source in a sound collection space for a sound receiver in a playback space that has a correspondence relationship with the sound collection space with respect to positions in the space, Position information acquisition means for acquiring information related to the measured or set positions of the sound source and the sound receiver; Input acoustic signal determination means for identifying at least one microphone positioned in a direction from the sound source in the sound collection space toward the corresponding position of the sound receiver or in a direction close to the direction until a predetermined condition is satisfied in that direction, based on the acquired information related to the position, and acquiring an acoustic signal related to the sound picked up by the identified microphone to determine an input acoustic signal; Output acoustic signal generation means for generating, based on the acquired information related to the position, an output acoustic signal that will output a sound equivalent or corresponding to the picked-up sound toward the sound receiver from a direction from the sound receiver in the playback space toward the corresponding position of the sound source or in a direction close to the direction until a predetermined condition is satisfied in that direction, using the input acoustic signal; A sound field reproduction system having the above is provided.

[0030] According to the present invention, there is further provided an acoustic field reproduction method implemented by a computer for reproducing an acoustic field related to a sound source in a sound collection space with respect to a sound receiver in a reproduction space having a correspondence relationship with the sound collection space with respect to a position in the space, comprising: obtaining information related to the measured or set positions of the sound source and the sound receiver; identifying at least one microphone located in a direction from the sound source in the sound collection space toward the corresponding position of the sound receiver or in a direction close to the direction until a predetermined condition is satisfied in the direction, and obtaining an acoustic signal related to the sound picked up by the identified microphone to determine an input acoustic signal; generating an output acoustic signal that will output a sound equivalent to or corresponding to the picked-up sound toward the sound receiver from a direction from the sound receiver in the reproduction space toward the corresponding position of the sound source or in a direction close to the direction until a predetermined condition is satisfied in the direction, based on the obtained information related to the positions, using the input acoustic signal; An acoustic field reproduction method having the above steps is provided.

Advantages of the Invention

[0031] According to the acoustic field reproduction program, apparatus, and method of the present invention, even when the sound source exists at an arbitrary position in the sound collection space or is movable within the sound collection space, and / or when the sound receiver exists at an arbitrary position in the reproduction space or is movable within the reproduction space, it is possible to reproduce the acoustic field of the sound collection space with respect to the sound receiver.

Brief Description of the Drawings

[0032]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

[0033] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.

[0034] [Sound Field Reproduction Device] FIG. 1 is a functional block diagram showing a functional configuration in an embodiment of a sound field reproduction device according to the present invention. In the same figure, an embodiment of the sound collection space and the playback space related to this device is also shown.

[0035] The sound field reproduction device 1 as an embodiment of the present invention shown in FIG. 1 is a device capable of reproducing the sound field related to a sound source (a speaker in this embodiment) in the sound collection space for a sound receiver (a listener in this embodiment) in the playback space.

[0036] Here, the sound collection space and the playback space are spaces that have a corresponding relationship with respect to the positions within the space (spaces having one-to-one corresponding positions within the space). Therefore, (a) The position of the sound receiver (listener) in the sound collection space, and the relative position and the direction (azimuth, orientation) of the sound receiver (listener) as seen from the sound source (speaker), and (b) The position of the sound source (speaker) in the playback space, and the relative position and the direction (azimuth, orientation) of the sound source (speaker) as seen from the sound receiver (listener) can be determined.

[0037] Here, in this embodiment, in order to further improve the reproduction accuracy of the sound field, the distance between the sound source (speaker) and the sound receiver (listener), and the direction (azimuth, orientation) formed by the sound source (speaker) and the sound receiver (listener) are also set to be the same between the sound collection space and the sound reproduction space. However, the size and shape of the sound collection space (as the area partitioned by the microphone group described later) and the size and shape of the sound reproduction space (as the area partitioned by the speaker group described later) may be different from each other.

[0038] Also, in this embodiment, the sound collection space and the sound reproduction space are respectively a sound collection room and a sound reproduction room, and a plurality of microphones (hereinafter abbreviated as microphones) (2, 3) and a plurality of speakers (5) are installed on the boundary of the room (in the horizontal plane). Further, in this embodiment, both the sound collection space and the sound reproduction space are provided with at least one (two in this embodiment that can measure the depth in opposite directions) depth sensor 4 on the boundary of the room, and the position of the sound source (speaker) in the sound collection space and the position of the sound receiver (listener) in the sound reproduction space can be measured.

[0039] Furthermore, in this embodiment, both the sound source (speaker) and the sound receiver (listener) are movable (for example, walking), and the depth sensor 4 can measure their positions in real time. However, of course, it is also possible for one or both of the sound source (speaker) and the sound receiver (listener) to be in a set fixed position without moving (for example, sitting at a predetermined position). In this case, the device 1 can acquire the information related to this set fixed position (for example, by device input by the setter).

[0040] Under such circumstances, the sound field reproduction device 1, in order to perform sound field reproduction processing, (A) A position information acquisition unit 111 that acquires the measured (or set) "information related to the position" of the sound source (speaker) and the sound receiver (listener), (B) Based on the obtained "information related to the position", at least one microphone located in the direction from the sound source (speaker) to the corresponding position of the sound receiver (listener) in the sound collection space or in a direction close to this direction until a predetermined condition is satisfied is identified, and an acoustic signal related to the "collected sound" is obtained with the identified microphone to determine the "input acoustic signal" by the input acoustic signal determination unit 112. (C) From the direction from the sound receiver (listener) to the corresponding position of the sound source (speaker) in the playback space or in a direction close to this direction until a predetermined condition is satisfied, an "output acoustic signal" in which a sound equivalent or corresponding to the "collected sound" is output toward the sound receiver (listener) is generated using the "input acoustic signal" based on the obtained "information related to the position" by the output acoustic signal generation unit 113. It is characterized by having the following.

[0041] Here, the identified microphones in the above (B) are two microphones 3a and 3b located in the sound collection space of the present embodiment shown in FIG. 1 in a direction close to the direction from the speaker to the corresponding position of the listener until a predetermined condition is satisfied (the closest direction and the second closest direction in FIG. 1). That is, in the present embodiment, two microphones 3a and 3b are selected (identified) from a plurality of microphones installed so as to surround the speaker based on the "information related to the position of the speaker" and the "information related to the position of the listener", and the "input acoustic signal" is determined from the acoustic signals related to the sounds collected by these microphones 3a and 3b.

[0042] Furthermore, as will be described later, the "output acoustic signal" in the above (C) is an acoustic signal to be supplied to two speakers 5f and 5g selected (identified) based on the "information related to the position of the listener" and the "information related to the position of the speaker" from a plurality of speakers (5) installed so as to surround the listener in the playback space of the present embodiment shown in FIG. 1.

[0043] Here, as will be described in detail later, these two speakers 5f and 5g are speakers positioned in a direction that is close to satisfying a predetermined condition in the direction from the listener towards the corresponding position of the speaker. So to speak, they have a positional relationship of facing each other with the speaker and the listener in between, with respect to the two specified microphones 3a and 3b. Therefore, from these speakers 5f and 5g that receive the "output acoustic signal", it is possible to output a sound that corresponds to or is equivalent to the "sound picked up" by the specified microphones 3a and 3b.

[0044] Incidentally, this is also an embodiment that will be described in detail later. Regarding the reproduction space as a virtual space related to sound image localization in a stereo phone (e.g., headphones) worn by the listener, the output acoustic signal generation unit 113 in (C) above may generate an "output acoustic signal" to be supplied to the worn stereo phone such that a sound corresponding to or equivalent to the "sound picked up" is output towards the listener from the direction towards the corresponding position of the speaker as seen by the listener in this reproduction space or from a direction close to this direction that satisfies a predetermined condition.

[0045] In any case, according to the sound field reproduction device 1, based on the acquired "information related to the position of the speaker" and "information related to the position of the listener", the microphones are specified (selected) to determine the "input acoustic signal", and further the "output acoustic signal" is generated. Therefore, even when the sound receiver (listener) exists at an arbitrary position within the reproduction space or is movable within the reproduction space, the sound field of the sound pickup space can be suitably reproduced for this sound receiver (listener), for example, reproduced with a high sense of presence. Also, it is understood that even when the sound source (speaker) exists at an arbitrary position within the sound pickup space or is movable within the sound pickup space, the sound field of the sound pickup space can similarly be suitably reproduced for the sound receiver (listener), for example, reproduced with a high sense of presence.

[0046] Incidentally, as described above, when the sound source (speaker) and the sound receiver (listener) are each movable within the sound pickup space, which is the real space, and within the reproduction space, (a) The input acoustic signal determination unit 112 of (B) above determines the microphone at this one point in time and determines the "input acoustic signal" related to this one point in time based on the "information related to the position" measured and obtained at one point in time by the position information acquisition unit 111 of (A) above. (b) It is also preferable that the output acoustic signal generation unit 113 of (C) above specifies the speaker at this one point in time and generates the "output acoustic signal" related to this one point in time based on the "information related to the position" obtained at this one point in time. In this case, for example, it is possible to reproduce, almost in real time, the sound field of the sound collection space that changes moment by moment due to the movement of the sound source (speaker) for the sound receiver (listener) moving within the reproduction space.

[0047] Also, in this embodiment, the microphones (2, 3) and the speakers (5) are each installed with their positions fixed at the boundaries of the sound collection space and the reproduction space, but they may not be fixed and their positions may change. That is, at least one of the plurality of microphones (2, 3) may be movable, for example, on the boundary of the sound collection space, or may be a microphone attached to a robot movable within the sound collection space. Furthermore, at least one of the plurality of speakers (5) may also be movable, for example, on the boundary of the reproduction space, or may be a speaker attached to a robot movable within the reproduction space.

[0048] In this case, the input acoustic signal determination unit 112 in (B) above identifies, from among the microphone group including the moving microphone, a microphone located in the direction from the sound source (speaker) toward the corresponding position of the sound receiver (listener) at a certain point in time, or in a direction close to this direction until a predetermined condition is satisfied. Here, for example, a microphone near the position corresponding to the above condition may be identified after moving it to the corresponding position. Further, the output acoustic signal generation unit 113 in (C) above identifies, for example, a speaker located in the direction from the sound receiver (listener) toward the corresponding position of the sound source (speaker) at this certain point in time, or in a direction close to this direction until a predetermined condition is satisfied. Here, for example, a speaker near the position corresponding to the above condition may be identified after moving it to the corresponding position.

[0049] Furthermore, in the present invention, it is also possible to adopt a form in which the input acoustic signal determination unit 112 in (B) above, which is a component of the sound field reproduction device 1, and the output acoustic signal generation unit 113 in (C) above are provided in different devices. For example, (a) A device provided in or near the sound collection space, equipped with the position information acquisition unit 111 and the input acoustic signal determination unit 112, and (b) A device provided in or near the reproduction space, equipped with the position information acquisition unit 111 and the output acoustic signal generation unit 113, and capable of exchanging information with the device in (a) above through communication may constitute a sound field reproduction system according to the present invention.

[0050] [Device Configuration, Sound Field Reproduction Program / Method] Hereinafter, the functional configuration of the sound field reproduction device 1 as an embodiment of the present invention will be described in more detail. Also in the functional block diagram of FIG. 1, the sound field reproduction device 1 includes a communication interface 101, a keyboard (KB) / display (DP) 102, and a processor / memory (an arithmetic processing system having a memory function). Here, the processor / memory stores the sound field reproduction program according to the present invention, and also has a computer function, and performs sound field reproduction processing by executing this sound field reproduction program.

[0051] Also, from this, the sound field reproduction device 1 may be a device dedicated to sound field reproduction processing, but it may also be a general-purpose cloud server or non-cloud type server equipped with the sound field reproduction program according to the present invention. Furthermore, it may be a personal computer (PC), a notebook or tablet computer, a smartphone, or even a wearable device such as a head-mounted display (HMD).

[0052] Also, the processor-memory, as a functional component, includes a position information acquisition unit 111, an omnidirectional microphone identification unit 112a, an input mixing unit 112b, a directional microphone identification unit 112c, and an input acoustic signal determination unit 112 including an input mixing unit 112d; an amplitude adjustment unit 113a, a timbre adjustment unit 113b, an output mixing unit 113c, a speaker (SP) identification panning unit 113d, and an output acoustic signal generation unit 113 including a near speaker (SP) identification panning unit 113e; a communication control unit 121; and an input / output control unit 122. Note that these functional components can be regarded as functions realized by executing the sound field reproduction program according to the present invention stored in the processor-memory. Also, the processing flow shown by connecting the functional components of the sound field reproduction device 1 in the functional block diagram of FIG. 1 with arrows is also understood as one embodiment of the sound field reproduction method according to the present invention.

[0053] <Position information acquisition means> Similarly, in the functional block diagram of FIG. 1, the position information acquisition unit 111 in the present embodiment (a) Information related to the measured or set positions of the speaker as the sound source and the listener as the sound receiver, which is "position coordinate information" in the present embodiment, and (b) Information related to the measured or set sound output direction of the speaker, which is "face azimuth angle information" related to the orientation (azimuth) of the speaker's face in the present embodiment is acquired.

[0054] Specifically, in this embodiment, both the sound collection room (sound collection space) and the playback room (playback space) are equipped with two depth sensors 4 capable of measuring depths in opposite directions within their own horizontal planes (to reliably obtain the information (a) and (b) in the room). Here, by performing image recognition processing capable of recognizing people and faces on the depth images obtained by these depth sensors 4, the position coordinate information and face azimuth angle information of the speaker in the sound collection room and the position coordinate information of the listener in the playback room are generated. The position information acquisition unit 111 may acquire this information from an image recognition processing device connected to the depth sensor 4, or may perform such image recognition processing on the depth images obtained from the depth sensor 4 by itself to generate and acquire this information.

[0055] As a further modification, instead of or together with the depth sensor 4, a visible light (and infrared) camera may be used, and by performing image recognition processing capable of recognizing people and faces on the images obtained by these cameras, the position coordinate information and face azimuth angle information of the speaker in the sound collection room and the position coordinate information of the listener in the playback room may be generated. Also, for example, it is possible to adopt Microsoft's Azure Kinect DK as the depth sensor 4. Azure Kinect DK is equipped with a ToF-type depth sensor and an RGB camera, and by cooperating with an image recognition cloud service, it is possible to generate this position coordinate information and face azimuth angle information. As an even further modification, the speaker or the listener may carry a terminal equipped with a magnetic sensor, an acceleration sensor, or an electrostatic sensor capable of determining the position in the room, and it is also possible to acquire the information related to the position from this terminal.

[0056] In this embodiment, in the sound collection room (sound collection space), a plurality of (eight in FIG. 1) omnidirectional microphones 2a to 2h and a plurality of (eight in FIG. 1) highly directional microphones 3a to 3h are installed so as to surround the inside of the space at the boundary in their own horizontal planes. Here, the highly directional microphones 3a to 3h are "different microphones" from the omnidirectional microphones 2a to 2h, and of course, they have a higher directivity than the omnidirectional microphones 2a to 2h. For example, the omnidirectional microphones 2a to 2h may be omnidirectional ambient microphones, and the highly directional microphones 3a to 3h may be gun microphones.

[0057] Also, in this embodiment, one omnidirectional microphone 2 and one highly directional microphone 3 are paired and installed at one position, and the installation positions of these pairs form a quadrangular boundary as a whole. However, of course, it is not limited to such a microphone arrangement. For example, the omnidirectional microphone 2 and the highly directional microphone 3 may be installed at different positions from each other, for example, at alternating positions. Also, the installation positions of the microphones may form a boundary such as a circle, an ellipse, or a polygon (including a triangle) as a whole.

[0058] Furthermore, in this embodiment, the acoustic signals related to the sounds collected by these microphones, and the position coordinate information and face azimuth angle information (or depth image information) generated using the above-described depth sensor 4 are received by the communication interface 101 of the sound field reproduction device 1 via a communication network from, for example, an acoustic signal / position information management device (not shown) installed in or near the sound collection room (sound collection space), and are provided to the input acoustic signal determination unit 112 and the position information acquisition unit 111 by the communication control unit 121 of the device 1. Additionally, the position information acquisition unit 111 also receives information related to the set (or measured) positions of the omnidirectional microphones 2 and the highly directional microphones 3, and information related to the set (or measured) position of the microphone 5 from the sound collection room side or the reproduction room side (or by direct input to the device 1), and provides this information to the input acoustic signal determination unit 112 and the output acoustic signal generation unit 113.

[0059] <Input acoustic signal determination means> Also in the functional block diagram of FIG. 1, the omnidirectional microphone specifying section 112a of the input acoustic signal determining section 112 (a) Based on the acquired position coordinate information, at least one, in this embodiment two, omnidirectional microphones 2 located in the direction from the speaker (sound source) in the sound collection room (sound collection space) toward the corresponding position of the listener (sound receiver), or in a direction close to this direction until a predetermined condition is satisfied, are specified (selected), and acoustic signals related to the sound picked up by the two specified omnidirectional microphones 2 (2a and 2b in FIG. 1) are acquired.

[0060] Here, the omnidirectional microphone 2 is a microphone that captures the acoustics (sound field) of the sound collection space itself. Therefore, the omnidirectional microphones 2 (2a and 2b) specified as above can pick up the sound corresponding to the acoustics (sound field) that the listener (sound receiver) perceives through hearing.

[0061] Also, the directional microphone specifying section 112c of the input acoustic signal determining section 112 (b) Based on the acquired position coordinate information and face azimuth angle information, at least one, in this embodiment two, directional microphones 3 located in the direction the speaker's face is facing from the speaker (sound source) in the sound collection room (sound collection space), or in a direction close to this direction until a predetermined condition is satisfied, are specified (selected), and acoustic signals related to the sound picked up by the specified directional microphones 3 (3c and 3d in FIG. 1) are acquired.

[0062] Here, the directional microphone 3 is a microphone that captures the sound itself from the sound source. Therefore, the directional microphones 3 (3c and 3d) specified as above can surely pick up the sound (voice) from the speaker (sound source) to be delivered to the listener (sound receiver).

[0063] Incidentally, information specifying the identified microphone may be notified to the sound collection room side, and only the acoustic signals related to the sounds collected by the identified omnidirectional microphones 2 and the directional microphones 3 may be received from the sound collection room side, or after receiving all the acoustic signals related to the sounds collected by these microphones from the sound collection room side, it is also possible to select the acoustic signals related to the identified microphones therefrom.

[0064] FIG. 2 is a schematic diagram related to a sound collection space for explaining an embodiment of the microphone identification process according to the present invention. Here, in the sound collection room which is the sound collection space, an xy position coordinate system that defines positions in the horizontal plane is set.

[0065] As shown in FIG. 2, in this embodiment, the omnidirectional microphone identification unit 112a (FIG. 1) calculates (a) the intersection point AM (x_amb, y_amb) of the line segment connecting the position S (x_s, y_s) of the speaker to the position R (x_r, y_r) of the listener and the boundary at the extension of the line segment, and (b) identifies (selects) the omnidirectional microphone 2b whose direction from the position S (x_s, y_s) to its installation position is closest (the angle formed therebetween is the smallest) to the direction from the position S (x_s, y_s) to the position AM (x_amb, y_amb), and the omnidirectional microphone 2a that is the second closest (the angle formed therebetween is the second smallest).

[0066] Also, in this embodiment, the directional microphone identification unit 112c (FIG. 1) calculates (c) the intersection point GM (x_gun, y_gun) of the line segment extending in the face azimuth angle θ_s direction from the position S (x_s, y_s) of the speaker and the boundary, and (d) identifies (selects) the directional microphone 3d whose direction from the position S (x_s, y_s) to its installation position is closest (the angle formed therebetween is the smallest) to the direction from S (x_s, y_s) to GM (x_gun, y_gun), and the directional microphone 3c that is the second closest (the angle formed therebetween is the second smallest).

[0067] Here, for both the omnidirectional microphone 2 and the highly directional microphone 3, it is also possible to identify (select) only the closest one, for example. However, in the present embodiment, in order to reproduce a natural sound field even when the position of the speaker changes moment by moment, that is, so that the change in the sound field due to the movement of the speaker is reproduced naturally, two such microphones are identified (selected) as described above. Of course, it is also possible to identify (select) three or more for the same reason.

[0068] Returning to the functional block diagram of FIG. 1, the input mixing unit 112b of the input acoustic signal determination unit 112 combines (for example, mixes) the acoustic signals related to the omnidirectional microphones (2a and 2b in FIG. 1) selected by the omnidirectional microphone identification unit 112a to generate a first input acoustic signal. Also, the input mixing unit 112d of the input acoustic signal determination unit 112 combines (for example, mixes) the acoustic signals related to the highly directional microphones (3c and 3d in FIG. 1) selected by the highly directional microphone identification unit 112c to generate a second input acoustic signal.

[0069] FIG. 3 is a schematic diagram related to a sound collection space for explaining an embodiment of the input acoustic signal determination process (acoustic signal mixing process) according to the present invention. Incidentally, the microphone settings and the states of the speaker / listener in the sound collection room (sound collection space) shown in this figure are the same as those in the sound collection room shown in FIG. 2.

[0070] First, using FIG. 3(A), the process of generating the first input acoustic signal (amplitude intensity) t_amb(p) by mixing the acoustic signal (amplitude intensity) t_2a(p) related to the selected omnidirectional microphone 2a and the acoustic signal (amplitude intensity) t_2b(p) related to the selected omnidirectional microphone 2b, which is performed by the input mixing unit 112b (FIG. 1), will be described. Here, p is a parameter representing a time point or a sampling index corresponding to the time point.

[0071] As shown in FIG. 3(A), (a) Let the angle formed by the line segment connecting the position S of the speaker and the position P2a of the omnidirectional microphone 2a and the line segment connecting the position S of the speaker and the position P2b of the omnidirectional microphone 2b be α, (b) When the angle formed by the line segment connecting the position S of the speaker and the position P2a of the omnidirectional microphone 2a and the line segment connecting the position S of the speaker and the position AM (explained in FIG. 2) is β, The first input acoustic signal t_amb(p) is calculated by calculating these α and β, and then using the following equation (1) t_amb(p)=t_2a(p)×(sin(α / 2)+sin(α / 2-β)) +t_2b(p)×(sin(α / 2)-sin(α / 2-β)) is derived.

[0072] In the above equation (1), the coefficients of t_2a(p) and t_2b(p) in the first term on the right side are respectively the mixing ratio of the acoustic signal related to the omnidirectional microphone 2a and the mixing ratio of the acoustic signal related to the omnidirectional microphone 2b. Incidentally, when β = 0 and β = α, the first input acoustic signal t_amb(p) becomes a constant multiple of t_2a(p) and a constant multiple of t_2b(p) respectively. Substantially, although two microphones are specified (selected) in advance, the acoustic signal related to one (a single) microphone is used as the input acoustic signal. That is, in this embodiment, when the position S of the speaker, the position R of the listener, and the position of the microphone are on the same straight line in this order, it may be set to specifically (select) that (a single) microphone.

[0073] Next, also using FIG. 3(A), the process of mixing the acoustic signal t_3c(p) related to the specified (selected) sharp directional microphone 3c and the acoustic signal t_3d(p) related to the specified (selected) sharp directional microphone 3d in the input mixing unit 112d (FIG. 1) to generate the second input acoustic signal t_gun(p) will be described.

[0074] As shown in FIG. 3(A), (a) The angle formed by the line segment connecting the position S of the speaker and the position P3c of the sharp directional microphone 3c and the line segment connecting the position S of the speaker and the position P3d of the sharp directional microphone 3d is δ, (b) Let the angle formed by the line segment connecting the position S of the speaker and the position P3c of the highly directional microphone 3c and the line segment connecting the position S of the speaker and the position GM (described in FIG. 2) be γ. After calculating these δ and γ, the second input acoustic signal t_gun(p) is given by the following equation (2) t_gun(p)=t_3c(p)×(sin(δ / 2)+sin(δ / 2 - γ)) +t_3d(p)×(sin(δ / 2)-sin(δ / 2 - γ)) and is derived therefrom. Here, regarding the number of specified (selected) microphones in the generation of the second input acoustic signal t_gun(p), the situation is the same as that described in the above equation (1).

[0075] <Output acoustic signal generation means> Returning to the functional block diagram of FIG. 1, the amplitude adjustment unit 113a of the output acoustic signal generation unit 113 performs an amplitude adjustment process on the second input acoustic signal t_gun(p) according to the distance between the position of the speaker (sound source) and the position of the listener (sound receiver) in this embodiment. Note that the reason for setting the processing target to the second input acoustic signal t_gun(p) (related to the highly directional microphone 3) is that for the sound (voice) from the speaker to be delivered to the listener, which is picked up by the highly directional microphones 3 (3c and 3d), in reality, the closer the speaker and the listener are to each other, the louder the volume, and conversely, the farther they are, the smaller the volume. By adjusting the amplitude of the input acoustic signal in this way, the reproduction accuracy of the sound field is improved and the sense of presence is enhanced. Of course, it is also possible not to perform such an amplitude adjustment process in order to shorten the calculation time.

[0076] Here, generally, the distance attenuation y (dB) at the position of the distance d (from the sound source) in the acoustic signal is obtained by y = 20log 10 (d_0 / d) with the reference distance being d_0. Also, generally, the amplitude of the acoustic signal and the decibel dB are related by (amplitude) = 10 dB / 20 . Therefore, specifically, the amplitude adjustment unit 113a adjusts the amplitude-adjusted second input acoustic signal tj_gun(p) according to the following equation (3) tj_gun(p) = t_gun(p) × amp amp = 10 y / 20 y = 20 log 10 (d_SR / d_SG) can be generated by. In the above formula (3), the (reference) distance d_SR and the distance d_SG are, as shown in FIG. 3(B), the distance between the position S of the speaker and the position R of the listener, and the distance between the position S of the speaker and the position GM, respectively.

[0077] Also in the functional block diagram of FIG. 1, the timbre adjustment unit 113b of the output acoustic signal generation unit 113 performs timbre adjustment processing on the second input acoustic signal related to the highly directional microphone 3 that picks up the sound from the speaker (sound source) to be delivered to the listener (sound receiving body) so that the timbre will be the one that the listener (sound receiving body) will actually receive. Here, generally, it is known that sound attenuates high-frequency components as it propagates (as it moves away from the sound source). Therefore, for example, the closer the listener is to the speaker, the more the distance sense of being closer is expressed by emphasizing a predetermined high-frequency band in the input acoustic signal, and thereby, it becomes possible to improve the reproduction accuracy of the sound field and enhance the sense of presence.

[0078] As such timbre adjustment processing, the timbre adjustment unit 113b performs, on the second input acoustic signal tj_gun(p) that has undergone amplitude adjustment processing, (a: Biquadratic filtering processing in time information) performs biquadratic filtering processing to emphasize a predetermined high-frequency band according to the proximity between the position of the speaker (sound source) and the position of the listener (sound receiving body) to generate a second input acoustic signal tf_gun(p) that has undergone filtering processing, or (b: Convolution processing in the frequency domain) determines the impulse response between the position of the speaker (sound source) and the position of the listener (sound receiving body) and performs convolution processing in the frequency domain related to this impulse response, and performs this convolution processing to generate a second input acoustic signal tc_gun(p) that has undergone convolution processing. Incidentally, of course, it is also possible not to perform such tone color adjustment processing in order to shorten the calculation time.

[0079] (Biquadratic filtering process in time information) First, a specific explanation will be given for the above (a). For example, a high shelf filter is adopted as a biquad filter, and the sound output from the speaker is regarded as speech sound including the consonant component in the 4 kHz band. The settings of this filter are · Freq (threshold value for boosting frequency bands higher than this value) = 4000 (Hz) · Q (width of frequency band peak) = 1.0 · Gain (gain of the filter) = amp, where amp = 10 y / 20 , y = 20log 10 (d_SR / d_SG) and so on. Then, from here, the biquad transfer function: (4) H(z)=(b0 + b1×z -1 + b2×z -2 ) / (a0 + a1×z -1 + a2×z -2 ) Calculate the coefficients a0, a1, a2, b0, b1, and b2 of the denominator and numerator in the biquad transfer function, and obtain the second input acoustic signal tf_gun(p) after biquadratic filtering processing using the following formula (5) tf_gun(p)=(1 / a0)×(b0×tj_gun(p)+b1×tj_gun(p - 1)+ b2×tj_gun(p - 2)-a1×tf_gun(p - 1)-a2×tj_gun(p - 2)) That is, it is generated.

[0080] Here, the calculation methods for the coefficients a0, a1, a2, b0, b1, and b2 in various filters are disclosed, for example, in the non-patent document: "Cookbook formulae for audio equalizer biquad filter coefficients", [online], [searched on June 27, Reiwa 4], Internet <URL: https: / / webaudio.github.io / Audio-EQ-Cookbook / audio-eq-cookbook.html>, and the non-patent document: "++C++; / / Unconfirmed Flight C", [online], [searched on June 27, Reiwa 4], Internet <URL: https: / / ufcpp.net / study / sp / digital_filter / biquad / >. Incidentally, the selection of the filter as described above and the filter settings such as Freq, Q, and Gain are preferably appropriately performed according to the sound field to be reproduced and the situation in the sound pickup space and reproduction space.

[0081] (Convolution processing in the frequency domain) Next, a specific explanation will be given for the above (i). First, prepare an impulse response that reflects the absorption characteristics of air or an impulse response of a filter with a plurality of peaks simulating the absorption characteristics of air between the position S of the speaker and the position R of the listener, and let this be i(p). Such an impulse response i(p) is a quantity that depends on the distance d_SR (Fig. 3(B)). Here, it is also preferable to appropriately set the time length N of the impulse response i(p) according to the capabilities of the computer. For example, typically, N = 1024 can be set, but when using a PC with a low processing speed, N = 2048 or N = 4096 may also be used.

[0082] Next, the impulse response i(p) is Fourier-transformed to derive the transfer function I(ω) (ω is the frequency), and the second input acoustic signal tj_gun(p) cut out with the set time length N is Fourier-transformed to calculate the second input acoustic signal tj_gun(ω) after Fourier transformation. Thereby, the second input acoustic signal tc_gun(p) after convolution processing in the frequency domain is given by the following equation (6) tc_gun(p)=IF[tc_gun(ω)] tc_gun(ω)=I(ω)*tj_gun(ω) It can be generated by. Here, IF[] is an inverse Fourier transform operator, and * is a convolution integral operator.

[0083] As described above, two processes (a) and (i) have been explained as the timbre adjustment process. However, in a situation where the real-time performance of sound field reproduction is emphasized, the process with relatively less calculation time (a: biquadratic filtering process in time information) is adopted. On the other hand, in a situation where a complex filtering process (for example, a process of emphasizing a plurality of specific different high-frequency bands) is performed to improve sound field reproduction and enhance the sense of presence, (i: convolution process in the frequency domain) may be adopted. That is, it is also preferable that a setting is made such that either one of the timbre adjustment processes can be selected and implemented according to the situation.

[0084] Similarly, in the functional block diagram of FIG. 1, the output mixing unit 113c of the output acoustic signal generation unit 113 combines (mixes) the generated first input acoustic signal (t_amb(p)) and the generated second input acoustic signals (tf_gun(p), tc_gun(p)) to generate an output acoustic signal t_out(p). That is, the following equation (7) t_out(p)=t_amb(p)+tf_gun(p) or t_out(p)=t_amb(p)+tc_gun(p) is used to derive the output acoustic signal t_out(p).

[0085] Here, as one aspect, the output mixing unit 113c determines the phase difference between the first input acoustic signal (t_amb(p)) and the second input acoustic signal (tf_gun(p), tc_gun(p)) based on the position S of the speaker and the position R of the listener, and after delaying or advancing one of the phases so as to eliminate the determined phase difference, combines (mixes) the first input acoustic signal (t_amb(p)) and the second input acoustic signal (tf_gun(p), tc_gun(p)) to generate an output acoustic signal t_out(p). This is also preferable.

[0086] Incidentally, the above phase difference occurs when (a) the distance d_SA (Fig. 3(B)) between the position S of the speaker related to the first input acoustic signal and the position AM and (b) the distance d_SG (Fig. 3(B)) between the position S of the speaker related to the second input acoustic signal and the position GM are different, which causes interference when mixing the two signals. Of course, it is also possible to shorten the calculation time by not performing such a phase difference adjustment process that requires time-consuming appropriate adjustment.

[0087] Specifically, the output mixing unit 113c first calculates the phase difference p_th using the following equation (8) p_th = fs×(d_SA - d_SG) / c where c (m / s) is the speed of sound and fs (Hz) is the sampling frequency.

[0088] Next, when d_SA ≥ d_SG, the output mixing unit 113c adds a delay to the second input acoustic signal (tf_gun(p), tc_gun(p)) related to the distance d_SG (position GM) to eliminate the phase difference, and then generates the output acoustic signal t_out(p). That is, the following equation (when d_SA ≥ d_SG) (9) t_out(p) = t_amb(p) + tf_gun(p + p_th) or t_out(p) = t_amb(p) + tc_gun(p + p_th) Based on this, the output acoustic signal t_out(p) is derived. On the other hand, when d_SA < d_SG, a delay is added to the first input acoustic signal (t_amb(p)) related to the distance d_SA (position AM) to eliminate the phase difference, and then the output acoustic signal t_out(p) is generated. That is, the following equation (when d_SA < d_SG) (9’) t_out(p) = t_amb(p + p_th) + tf_gun(p) or t_out(p) = t_amb(p + p_th) + tc_gun(p) is used to derive the output acoustic signal t_out(p).

[0089] Also in the functional block diagram of FIG. 1, the speaker identification panning unit 113d of the output acoustic signal generation unit 113 (a) At least one speaker 5 (in FIG. 1, two speakers 5f and 5g) located in the direction from the listener (sound receiver) to the corresponding position of the speaker (sound source) as seen from the listener in the playback room (playback space) or in a direction close to this direction until a predetermined condition is satisfied is identified based on the acquired position coordinate information (information related to the position) of the listener (sound receiver) and the speaker (sound source), (b) Generate the output acoustic signal to be supplied to the identified speakers 5 (5f and 5g). Hereinafter, with reference to FIG. 4, specific descriptions of the speaker identification process in (a) above and the speaker panning process in (b) above will be given.

[0090] FIG. 4 is a schematic diagram related to a playback space for explaining an embodiment of the speaker identification process and the speaker panning process according to the present invention. Here, in the playback room which is the playback space, an xy position coordinate system for defining positions in the horizontal plane is set.

[0091] As shown in FIG. 4, the speaker identification panning unit 113d (FIG. 1) first in this embodiment (a1) Calculate the intersection point SP(x_sp, y_sp) of the line segment connecting the position R of the listener and the position S of the speaker and the boundary at the extension of the line segment, (A2) Identify (select) the speaker 5g whose direction from its installation position toward the listener's position R is closest (the angle formed therewith is the smallest) to the direction from the position SP(x_sp, y_sp) toward the listener's position R, and the speaker 5f that is second closest (the angle formed therewith is second smallest).

[0092] Next, in this embodiment, the speaker identification panning unit 113d (FIG. 1) (i) Based on the acquired position coordinate information of the listener and the speaker, generate output acoustic signals for speakers 5f and 5g from which sounds corresponding to or corresponding to "the sound picked up (by the microphone)" will be output toward the listener, using the output acoustic signal t_out(p) (generated using the first and second input acoustic signals).

[0093] Specifically, as shown in FIG. 4, (a) Let ρ be the angle formed by the line segment connecting the position P5f of speaker 5f and the position R of the listener and the line segment connecting the position P5g of speaker 5g and the position R of the listener. (b) When the angle formed by the line segment connecting the position P5f of speaker 5f and the position R of the listener and the line segment connecting the position SP and the position R of the listener is σ, The output acoustic signal t5f_out(p) for speaker 5f (to be supplied to speaker 5f) and the output acoustic signal t5g_out(p) for speaker 5g (to be supplied to speaker 5g) are calculated based on these ρ and σ, and then by the following equations (10) t5f_out(p)=t_out(p)×(sin(ρ / 2)+sin(ρ / 2 - σ)) t5g_out(p)=t_out(p)×(sin(ρ / 2)-sin(ρ / 2 - σ)) are derived.

[0094] Here, the above equation (10) implements what is called (amplitude) panning, which uses speakers 5f and 5g to achieve sound image localization by the volume difference (difference in amplitude intensity) between the two speakers (channels).

[0095] The above has described the process of reproducing the sound field of the sound collection space (recording room) including the sound source (speaker) for the sound receiver (listener) existing in the reproduction space (reproduction room) with reference to FIGS. 1 to 4. Here, the process in this embodiment will be briefly summarized. In this embodiment, (a) Based on the position coordinate information of the sound source (speaker) and the sound receiver (listener) from the multi-channel omnidirectional microphone 2, two channels of omnidirectional microphones 2 are selected to determine a one-channel (mono) first input acoustic signal, (b) Based on the position coordinate information and face azimuth angle information of the sound source (speaker) from the multi-channel highly directional microphone 3, two channels of highly directional microphones 3 are selected to determine a one-channel (mono) second input acoustic signal, (c) After mixing these mono (a total of two channels) input acoustic signals to generate a one-channel (mono) output acoustic signal, based on the position coordinate information of the sound source (speaker) and the sound receiver (listener) from the multi-channel speaker 5, two channels of speakers 5 are selected, and an output acoustic signal capable of sound image localization by the selected speaker 5 is generated and provided.

[0096] Returning to the functional block diagram of FIG. 1, the near speaker identification panning unit 113e of the output acoustic signal generation unit 113 also identifies at least one "other speaker" that is closer to the position of the listener (sound receiver) or located more inward (when viewed from the room wall) compared to the identified speakers 5 (5f and 5g in FIG. 1), and also generates the output acoustic signal to be supplied to the identified "other speaker" (speaker 5_cell in FIG. 1). Thereby, it becomes possible to reproduce or recreate a sound field (acoustics) that cannot be achieved only by the surround speakers 5a to 5h, for example, a sound field (acoustics) that makes the listener feel that the sound source is present in the vicinity, that is, a sound field (acoustics) that localizes the sound image closer.

[0097] Here, in the embodiment shown in FIG. 1, the speaker 5_cell installed at the center of the ceiling of the playback room is a "different speaker" from the surround speakers 5a to 5h arranged in the playback room, and in this playback room, it is a speaker capable of performing more inner acoustic presentation. In this case, the near speaker identification panning unit 113e determines the output acoustic signal tcell_out(p) for the speaker 5_cell using any one of the following equations (11) tcell_out(p)=tf_gun(p) (12) tcell_out(p)=tc_gun(p) (13) tcell_out(p)=t_out(p) . Hereinafter, an embodiment in which there are a plurality of such "different speakers" will be described with reference to FIG. 5.

[0098] FIG. 5 is a schematic diagram related to a playback space for explaining an embodiment of the ceiling speaker identification process and the ceiling speaker panning process according to the present invention.

[0099] According to FIG. 5, in the playback space, as "different speakers" from the surround speakers 5a to 5h (omitted in FIG. 5), a plurality of positions (surrounding the interior of the space) at the ceiling boundary and the position at the center of the ceiling are provided with (a total of nine in FIG. 5) ceiling speakers. Here, the near speaker identification panning unit 113e (FIG. 1) identifies (selects) a predetermined number of (three in FIG. 5) ceiling speakers based on the position coordinate information (information related to the position) of the listener, and generates an output acoustic signal to be supplied to each ceiling speaker so that panning can be performed by these ceiling speakers.

[0100] Specifically, in FIG. 5, three ceiling speakers 5_cell1, 5_cell2, and 5_cell3 located at the vertices of the smallest "triangle" with the ceiling speaker as the vertex and including the ceiling position (the cross mark in FIG. 5) above the listener's head are specified (selected). The near speaker specifying panning unit 113e (FIG. 1) generates output acoustic signals for these specified ceiling speakers 5_cell1, 5_cell2, and 5_cell3 so that sound can be heard from the ceiling position (cross mark) above the listener's head, for example, by the VBAP (Vector Based Amplitude Panning) method.

[0101] Here, the VBAP method is a typical control method for sound image localization in three-dimensional multi-channel acoustics (sound field), and is described in detail, for example, in the non-patent literature: Akio Ando, "High Immersion Technology of Acoustics", Journal of the Society of Motion Picture and Television Engineers, Vol. 66, No. 8, pp. 671-677, 2012.

[0102] In this embodiment, the ceiling speakers are fixedly installed on the ceiling of the playback room (playback space), but they may not be fixed and their positions may change. That is, at least one of the plurality of ceiling speakers may be movable on the ceiling of the playback room, for example. Also, as "another speaker", a speaker attached to a robot movable in the playback room can be adopted. In this case, for example, in order to realize sound image localization in the vicinity for the listener, the robot equipped with the speaker may move to the vicinity of the listener as appropriate. Furthermore, as "another speaker", it is also possible to use the dodecahedron speaker 5p installed at the center of the floor of the playback room or movable on the floor.

[0103] The sound field reproduction process using the speakers installed in the playback room (playback space) by the output acoustic signal generation unit 113 (FIG. 1) has been described above. Hereinafter, the sound field reproduction process in the stereo on worn by the listener (sound receiving body), in this embodiment, the headphones will be described. Here, in this case, the playback space is not a physical space, but a virtual space (sound image space) related to sound image localization in the headphones worn by the listener.

[0104] Returning to the functional block diagram of FIG. 1, the output acoustic signal generation unit 113 uses the output acoustic signal t_out(p) generated by the output mixing unit 113c, and from the direction towards the corresponding position of the speaker (sound source) as seen by the listener (sound receiver) in the playback space (sound image space) of the headphones or from a direction close to this direction until a predetermined condition is satisfied, towards the listener, an output acoustic signal th_out(p) to be supplied to the headphones is generated, in which a sound corresponding or corresponding to the "collected sound" is output.

[0105] Specifically, the head-related transfer function (HRTF) of the headphones that reflects the position of the speaker as seen by the listener and the orientation of the speaker with respect to both ears of the listener is convolved (performs convolution integration processing) with the output acoustic signal t_out(p) in the frequency domain, whereby the output acoustic signal th_out(p) can be generated. Regarding the HRTF, for example, in the non-patent literature: Takanao Nishino, "Head-Related Transfer Function Database", [online], [searched on June 27, Reiwa 4], Internet <URL: https: / / sites.google.com / site / takanorinishinomu / research / hrtf / database-j>, it is explained in detail.

[0106] Also, when generating the output acoustic signal th_out(p) for headphones as described above, instead of the output acoustic signal t_out(p), · The output acoustic signal tan_out(p) generated by convolving (performing convolution integration processing) the inverse filter generated from the room transfer function (RTF) of the recording room (recording space) with this output acoustic signal t_out(p) may be used. Thereby, the listener can also enjoy a clear sound with, for example, the reverberation of the recording room removed by these headphones. Of course, this output acoustic signal tan_out(p) may also be used in other acoustic output systems such as speakers in the recording room, in addition to headphones.

[0107] As yet another embodiment, the output acoustic signal generation unit 113 can also cause the output acoustic signal t_out(p) output from the output mixing unit 113d to be provided to the outside of the apparatus 1 as it is (for example, without performing panning processing). Even in this case, this output acoustic signal t_out(p) can be used as an acoustic signal that corresponds to or is equivalent to the "sound picked up" toward the sound receiver outside the apparatus 1 and is output as a sound.

[0108] Furthermore, as yet a further other embodiment, the input acoustic signal determination unit 112 may obtain an acoustic signal related to the sound picked up by at least one pre-specified microphone without particularly performing the omnidirectional microphone identification process or the highly directional microphone identification process as described above, and determine the input acoustic signal. Even in this case, the output acoustic signal generation unit 113 identifies at least one speaker 5 positioned in the direction toward the corresponding position of the sound source (speaker) as seen from the sound receiver (listener) in the playback space (playback room) or in a direction close to this direction until a predetermined condition is satisfied, and generates an output acoustic signal to be supplied to the identified speaker 5 based on the information related to the positions of the obtained sound receiver (listener) and sound source (speaker).

[0109] Also in the functional block diagram of FIG. 1, the output acoustic signal generated by the output acoustic signal generation unit 113 as described above is transmitted from the communication interface 101 via the communication control unit 121, through the communication network, to each speaker installed (identified) in the playback space, or to an acoustic signal distribution control device installed (not shown) in or near the playback room (playback space). Here, it is also preferable that the output acoustic signal to be transmitted is provided with identification information of the (identified) speaker to which the signal is to be supplied.

[0110] Also, the sound field reproduction process as described above may be implemented, for example, by the input / output control unit 122 outputting a process execution instruction input from the keyboard (KB) 102 to the corresponding functional component. Also, for example, the input / output control unit 122 may appropriately acquire information on the specified microphone and speaker, as well as information related to the positions of the speaker and the listener, from the corresponding functional component, and display, on the display (DP) 102, an image showing the momentary situation in the sound collection room and the reproduction room, such as that shown above in FIG. 1.

[0111] As described in detail above, according to the present invention, based on the information related to the positions of the acquired sound source and the sound receiver, the microphone is specified to determine the input acoustic signal, and further the output acoustic signal is generated. Therefore, even when the sound receiver exists at an arbitrary position within the reproduction space or is movable within the reproduction space, the sound field of the sound collection space can be reproduced for this sound receiver. Also, even when the sound source exists at an arbitrary position within the sound collection space or is movable within the sound collection space, it is similarly possible to reproduce the sound field of the sound collection space for the sound receiver.

[0112] Furthermore, the present invention can greatly contribute to faithfully reproducing, for the provider of the sound field, a sound field corresponding to or corresponding to the sound field in the sound collection space, for example, in web conferences, remote sessions, remote location conferences, and even online concerts, which have been widely used in recent years.

[0113] Moreover, according to the present invention, for example, (a) music teaching materials for singers and musicians who perform singing and playing while moving on a stage, (b) physical education teaching materials that record voice and movement sounds of athletes who move around, and further (c) language education teaching materials related to conversations while walking side by side with native speakers or while performing various collaborative tasks, and (d) social study teaching materials that record various sounds generated in various locations in facilities such as factories and exhibition halls, etc. are created with high sound field reproducibility, and excellent music education, physical education, language education, social education, etc. that utilize such high-quality teaching materials can be provided not only to children and learners in urban areas but also in local areas, for example, online, and in some cases, in real time. That is, according to the present invention, it is also possible to contribute to Goal 4 of the Sustainable Development Goals (SDGs) led by the United Nations, "Provide inclusive and equitable quality education for all and promote lifelong learning opportunities."

[0114] Regarding the various embodiments of the present invention described above, various changes, modifications, and omissions within the scope of the technical idea and perspective of the present invention can be easily made by those skilled in the art. The foregoing description is merely an example and is not intended to impose any restrictions. The present invention is limited only by the scope of the claims and their equivalents.

Explanation of Reference Numerals

[0115] 1 Sound field reproduction device 101 Communication interface 102 Keyboard (KB)·Display (DP) 111 Position information acquisition unit 112 Input acoustic signal determination unit 112a Omnidirectional microphone identification unit 112b, 112d Input mixing unit 112c Directional microphone identification unit 113 Output acoustic signal generation unit 113 113a Amplitude adjustment unit 113b Tone color adjustment unit 113c Output mixing unit 113d Speaker (SP) identification panning unit 113e Near Speaker (SP) Specific Panning Section 121 Communication Control Section 122 Input / Output Control Section 2, 2a, 2b, 2c, 2d, 2e, 2f, 2g, 2h Omnidirectional Microphone 3, 3a, 3b, 3c, 3d, 3e, 3f, 3g, 3h Sharply Directional Microphone 4 Depth Sensor 5, 5a, 5b, 5c, 5d, 5e, 5f, 5g, 5h Speaker 5_cell1, 5_cell2, 5_cell3 Ceiling Speaker 5p Dodecahedron Speaker< / url:>

Claims

1. A sound field reproduction program for reproducing a sound field related to a sound source in a sound collection space for a sound receiving body in a reproduction space that has a correspondence relationship with the sound collection space with respect to positions in the space, a position information acquisition means for acquiring information related to the measured or set positions of the sound source and the sound receiving body, identifying at least one microphone located in the sound collection space in a direction from the sound source toward the corresponding position of the sound receiving body or in a direction close to the direction until a predetermined condition is satisfied in the direction, and acquiring an acoustic signal related to the sound picked up by the identified microphone to determine an input acoustic signal as input acoustic signal determination means, output acoustic signal generation means for generating, based on the information related to the acquired position, an output acoustic signal that should output a sound equivalent to or corresponding to the picked-up sound toward the sound receiving body from a direction from the sound receiving body toward the corresponding position of the sound source in the reproduction space or in a direction close to the direction until a predetermined condition is satisfied in the direction, using the input acoustic signal characterized in that the computer is caused to function. A sound field reproduction program.

2. The position information acquisition means also acquires information related to the measured or set sound output direction of the sound source, the input acoustic signal determination means identifies at least one other microphone located in the sound collection space in the output direction or in a direction close to the direction until a predetermined condition is satisfied in the direction from the sound source, based on the information related to the acquired position and the information related to the output direction, and acquires another acoustic signal related to the sound picked up by the identified other microphone to determine another input acoustic signal, the output acoustic signal generation means generates the output acoustic signal by combining the other input acoustic signal and the input acoustic signal related to the sound picked up by the microphone The sound field reproduction program according to claim 1, characterized in that.

3. The sound field reproduction program according to claim 2, characterized in that the other microphone has higher directivity than the microphone.

4. The reproduction space is a real space, the output acoustic signal generation means identifies at least one speaker located in the reproduction space in a direction from the sound receiving body toward the corresponding position of the sound source or in a direction close to the direction until a predetermined condition is satisfied in the direction, based on the information related to the acquired position, and generates the output acoustic signal to be supplied to the identified speaker The sound field reproduction program according to any one of claims 1 to 3, characterized in that...

5. The output acoustic signal generation means also identifies at least one other speaker that is closer to the position of the sound receiver than the identified speaker or is located more inside when viewed from the surrounding boundary, and generates the output acoustic signal to be supplied to the identified other speaker. The sound field reproduction program according to claim 4, characterized in that...

6. The sound source and the sound receiver are each movable within the sound collection space and the reproduction space, which are real spaces. The input acoustic signal determination means identifies the microphone at the one point in time based on the information related to the position measured and acquired at one point in time by the position information acquisition means, and determines the input acoustic signal related to the one point in time. The output acoustic signal generation means identifies the speaker at the one point in time based on the information related to the position acquired at the one point in time, and generates the output acoustic signal related to the one point in time. The sound field reproduction program according to claim 4, characterized in that...

7. The reproduction space is a virtual space related to sound image localization in a stereo headphone worn by the sound receiver. The output acoustic signal generation means generates the output acoustic signal to be supplied to the stereo headphone, such that a sound corresponding to or corresponding to the collected sound is output toward the sound receiver from a direction toward the corresponding position of the sound source as viewed from the sound receiver in the reproduction space or from a direction close to the direction until a predetermined condition is satisfied in the direction. The sound field reproduction program according to any one of claims 1 to 3, characterized in that...

8. The output acoustic signal generation means performs an amplitude adjustment process of attenuating the amplitude according to the distance between the position of the sound source and the position of the sound receiver on the determined other input acoustic signal, and generates an output acoustic signal using the other input acoustic signal subjected to the process. The sound field reproduction program according to claim 2 or 3, characterized in that...

9. The output acoustic signal generation means performs filtering processing to emphasize a predetermined high-frequency band according to the proximity between the position of the sound source and the position of the sound receiving body on the determined other input acoustic signal, or determines an impulse response between the position of the sound source and the position of the sound receiving body and performs convolution processing in the frequency domain related to the impulse response, and generates an output acoustic signal using the processed other input acoustic signal. The sound field reproduction program according to claim 2 or 3, characterized in that.

10. The output acoustic signal generation means determines a phase difference between the other input acoustic signal and the input acoustic signal based on the information related to the position, and after delaying or advancing one phase so as to eliminate the determined phase difference, combines the other input acoustic signal and the input acoustic signal to generate the output acoustic signal. The sound field reproduction program according to claim 2 or 3, characterized in that.

11. A sound field reproduction program for reproducing a sound field related to a sound source in a sound collection space for a sound receiving body in a reproduction space that has a correspondence relationship with the sound collection space with respect to the position in the space, Position information acquisition means for acquiring information related to the measured or set positions of the sound source and the sound receiving body, and information related to the measured or set sound output direction of the sound source; (a) Based on the acquired information related to the position, at least one microphone located in the direction from the sound source in the sound collection space toward the corresponding position of the sound receiving body or in a direction close to the direction until a predetermined condition is satisfied is specified, and an acoustic signal related to the sound picked up by the specified microphone is acquired to determine an input acoustic signal. Also, (b) based on the acquired information related to the position and the information related to the output direction, at least one other microphone located in the output direction or in a direction close to the output direction until a predetermined condition is satisfied as seen from the sound source in the sound collection space is specified, and another acoustic signal related to the sound picked up by the specified other microphone is acquired to determine another input acoustic signal. Input acoustic signal determination means; Output acoustic signal generation means for combining the input acoustic signal and the other input acoustic signal to generate an output acoustic signal for which a sound corresponding or corresponding to the picked-up sound is output toward the sound receiving body; A sound field reproduction program characterized by causing a computer to function.

12. A sound field reproduction device for reproducing a sound field related to a sound source in a sound collection space for a sound receiving body in a reproduction space having a correspondence relationship with the sound collection space with respect to a position in the space, a position information acquisition means for acquiring information related to measured or set positions of the sound source and the sound receiving body, an input acoustic signal determination means for specifying at least one microphone located in a direction from the sound source in the sound collection space toward the corresponding position of the sound receiving body or in a direction close to the direction until a predetermined condition is satisfied, acquiring an acoustic signal related to the sound picked up by the specified microphone, and determining an input acoustic signal, an output acoustic signal generation means for generating, based on the acquired information related to the position, an output acoustic signal in which a sound corresponding to or corresponding to the picked-up sound is output toward the sound receiving body from a direction from the sound receiving body in the reproduction space toward the corresponding position of the sound source or in a direction close to the direction until a predetermined condition is satisfied, using the input acoustic signal, characterized by comprising the above.

13. A sound field reproduction system for reproducing a sound field related to a sound source in a sound collection space for a sound receiving body in a reproduction space having a correspondence relationship with the sound collection space with respect to a position in the space, a position information acquisition means for acquiring information related to measured or set positions of the sound source and the sound receiving body, an input acoustic signal determination means for specifying at least one microphone located in a direction from the sound source in the sound collection space toward the corresponding position of the sound receiving body or in a direction close to the direction until a predetermined condition is satisfied, acquiring an acoustic signal related to the sound picked up by the specified microphone, and determining an input acoustic signal, an output acoustic signal generation means for generating, based on the acquired information related to the position, an output acoustic signal in which a sound corresponding to or corresponding to the picked-up sound is output toward the sound receiving body from a direction from the sound receiving body in the reproduction space toward the corresponding position of the sound source or in a direction close to the direction until a predetermined condition is satisfied, using the input acoustic signal, characterized by comprising the above.

14. A sound field reproduction method implemented by a computer for reproducing a sound field related to a sound source in a sound collection space for a sound receiving body in a reproduction space having a correspondence relationship with the sound collection space with respect to a position in the space, A step of obtaining information related to the measured or set positions of the sound source and the sound receiver; Based on the obtained information related to the positions, identifying at least one microphone positioned in a direction from the sound source toward the corresponding position of the sound receiver in the sound collection space or in a direction close to the direction until a predetermined condition is satisfied, and obtaining an acoustic signal related to the sound picked up by the identified microphone to determine an input acoustic signal; Based on the obtained information related to the positions, generating an output acoustic signal that will output a sound equivalent to or corresponding to the picked-up sound toward the sound receiver from a direction from the sound receiver toward the corresponding position of the sound source in the sound reproduction space or in a direction close to the direction until a predetermined condition is satisfied, using the input acoustic signal; A sound field reproduction method, characterized by including the above steps.

Citation Information

Patent Citations

  • Method for reproducing sound field

    JP1999262097A

  • Sound field reproducing device

    JP2001142471A

  • Information recording device and information reproducing device

    JP2001169309A

  • High presence sound field reproduction information transmitter, high presence sound field reproduction information transmitting program, high presence sound field reproduction information transmitting method and high presence sound field reproduction information receiver, high presence sound field reproduction information receiving program, high presence sound field reproduction information receiving method

    JP2005086537A

  • Video conference system

    JP2020088516A