Sound field reproduction device, sound field reproduction method, and sound field reproduction system

The sound field reproduction device uses low-order Ambisonics components to accurately reproduce sound fields outside the recording space by generating high-order signals for multiple speakers, addressing processing delays and localization errors in existing systems.

JP7808792B2Active Publication Date: 2026-01-30PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2022129434
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-08-15
Publication Date
2026-01-30
Estimated Expiration
2042-08-15

AI Technical Summary

Technical Problem

Existing scene-based 3D sound reproduction systems face challenges in reproducing sound fields accurately when the listener is positioned outside the sound collection target space, and increasing the number of microphone elements for higher-order sound field components leads to processing delays and sound source localization errors.

Method used

A sound field reproduction device and method that utilizes low-order sound field components recorded using an Ambisonics microphone, employing a sound source extraction direction control unit, a re-encoding unit, and a sound field reproduction unit to generate high-order basis acoustic signals for output through multiple speakers, allowing accurate sound field reproduction outside the recording space.

Benefits of technology

Suppresses sound source localization errors and enables accurate sound field reproduction in a different space by using low-order components, reducing processing load and output delays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007808792000050
    Figure 0007808792000050
  • Figure 0007808792000051
    Figure 0007808792000051
  • Figure 0007808792000052
    Figure 0007808792000052
Patent Text Reader

Abstract

To suppress an increase in a sound source localization error within a sound field reproduction space by using a low-order sound field component recorded using an ambisonics microphone.SOLUTION: A sound field reproduction device includes a sound source extraction direction control unit that receives designation of a sound source extraction direction in a sound field recording space in which a recording device is placed, a re-encoding unit that generates a high-order base acoustic signal by re-encoding the acoustic signal corresponding to the sound source extraction direction among the low-order base acoustic signals based on the encoding process using the recorded signal by the recording device, and a sound field reproduction unit that outputs a signal based on a high-order fundamental acoustic signal from each of a plurality of speakers provided in a sound field reproduction space different from the sound field recording space.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a sound field reproduction device, a sound field reproduction method, and a sound field reproduction system. [Background technology]

[0002] Recently, scene-based 3D sound reproduction technology has been attracting attention for its ability to reproduce sound fields in real time. Scene-based 3D sound reproduction technology is a method of reproducing a 3D sound field in real time using speakers arranged to surround the listening environment (space) by processing multi-channel signals recorded using an Ambisonics microphone, which has multiple directional microphone elements arranged on a rigid or hollow sphere. This reproduces a 3D sound field in real time, as if the listener were actually present at the location where the Ambisonics microphones were installed.

[0003] Known prior art related to sound field reproduction is, for example, Patent Document 1. Patent Document 1 discloses a signal processing device that acquires a plurality of collected sound signals based on sound collected by a plurality of sound collecting units that are installed together in a target sound collection space and that are installed in a plurality of different orientations according to the position of a sound source and the position of an object that reflects the sound emitted from the sound source, and generates an acoustic signal corresponding to a designated listening point in the target sound collection space based on the acquired plurality of collected sound signals. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Publication No. 2019-192975 Summary of the Invention [Problem to be solved by the invention]

[0005] The configuration of Patent Document 1 is premised on the existence of a listening point within a sound collection target space where multiple sound collection units are arranged. Therefore, even if a scene-based 3D sound system is constructed using Patent Document 1, the listener must be located within the sound collection target space where the sound collection units are arranged. In other words, if the listener is located in a position other than the sound collection target space, it is difficult to reproduce the sound field so that the sound signals collected within the sound collection target space can be heard within the sound collection target space.

[0006] On the other hand, it is known that an Ambisonics microphone can synthesize higher-order Ambisonics components (sound field components) as the number of microphone elements arranged on a spherical surface increases, thereby improving directional resolution during recording and reproduction. However, synthesizing higher-order sound field components for real-time live streaming of events such as public viewing requires increasing the number of microphone elements arranged in the recording target space to synthesize higher-order sound field components. Therefore, depending on the installation space of the microphone elements, there are physical constraints on microphone element placement. Furthermore, the increase in the number of transmission channels associated with an increase in the number of microphone elements increases the processing load for signal processing and synthesis processes, resulting in output delays for sound field reproduction. Therefore, it is desirable to reproduce sound fields using lower-order sound field components without unnecessarily increasing the number of microphone elements. However, in this case, sound source localization errors increase significantly as the microphone element moves away from the center position of the surrounding speakers, making it impossible to achieve the expected sound field reproduction.

[0007] The present disclosure has been devised in view of the above-described conventional situation, and aims to provide a sound field reproduction device, a sound field reproduction method, and a sound field reproduction system that utilize low-order sound field components recorded using an Ambisonics microphone and suppress an increase in sound source localization errors within a sound field reproduction space. [Means for solving the problem]

[0008] The present disclosure provides a sound field reproduction device comprising: a sound source extraction direction control unit that receives a designation of a sound source extraction direction within a sound field recording space in which a recording device is placed; a re-encoding unit that generates a high-order basis acoustic signal by re-encoding an acoustic signal that corresponds to the sound source extraction direction among low-order basis acoustic signals based on an encoding process using a signal recorded by the recording device; and a sound field reproduction unit that outputs a signal based on the high-order basis acoustic signal from each of a plurality of speakers provided in a sound field reproduction space different from the sound field recording space.

[0009] The present disclosure also provides a sound field reproduction method comprising the steps of: receiving a designation of a sound source extraction direction within a sound field recording space in which a recording device is placed; generating high-order basis acoustic signals by re-encoding acoustic signals corresponding to the sound source extraction direction among low-order basis acoustic signals based on an encoding process using signals recorded by the recording device; and outputting signals based on the high-order basis acoustic signals from each of a plurality of speakers provided in a sound field reproduction space different from the sound field recording space.

[0010] The present disclosure also provides a sound field reproduction system comprising: a sound field recording apparatus having a recording device capable of recording a sound source within a sound field recording space; and a sound field reproduction device that reproduces the acoustic signal recorded by the recording device in a sound field reproduction space different from the sound field recording space, wherein the sound field reproduction device comprises: a sound source extraction direction control unit that receives a designation of a sound source extraction direction within the sound field recording space; a re-encoding unit that re-encodes acoustic signals that correspond to the sound source extraction direction among low-order basis acoustic signals based on an encoding process using the signal recorded by the recording device to generate high-order basis acoustic signals; and a sound field reproduction unit that outputs signals based on the high-order basis acoustic signals from each of a plurality of speakers provided in the sound field reproduction space.

[0011] These comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a recording medium, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium. [Effects of the Invention]

[0012] According to the present disclosure, it is possible to suppress an increase in sound source localization errors within a sound field reproduction space by utilizing low-order sound field components recorded using an Ambisonics microphone. [Brief explanation of the drawings]

[0013] [Figure 1] A schematic diagram showing the concept of scene-based 3D sound reproduction using Ambisonics microphones, from sound field recording to sound field reproduction. [Figure 2] A diagram showing an example of a basis of Ambisonics components based on spherical harmonic expansion for degree n and degree m. [Figure 3] A block diagram showing an example of the system configuration of a sound field reproduction system according to the first embodiment. [Figure 4] FIG. 10 is a diagram showing an example of an outline of operations from sound field recording to sound field reproduction according to the first embodiment; [Figure 5] 1 is a flowchart showing an example of a procedure for reproducing a sound field by the sound field reproduction device according to the first embodiment in chronological order. [Figure 6] A block diagram showing an example of the system configuration of a sound field reproduction system according to a second embodiment. [Figure 7] FIG. 10 is a diagram showing an example of an outline of operations from sound field recording to sound field reproduction according to the second embodiment. [Figure 8] 10 is a flowchart showing an example of the operation procedure for reproducing a sound field by the sound field reproduction device according to the second embodiment in chronological order. DETAILED DESCRIPTION OF THE INVENTION

[0014] Hereinafter, with appropriate reference to the drawings, detailed descriptions will be given of embodiments specifically disclosing a sound field reproduction device, a sound field reproduction method, and a sound field reproduction system according to the present disclosure. However, unnecessary detailed descriptions may be omitted. For example, detailed descriptions of well-known matters and redundant descriptions of substantially identical configurations may be omitted. This is to avoid unnecessary redundancy in the following description and to facilitate understanding by those skilled in the art. Note that the accompanying drawings and the following description are provided to enable those skilled in the art to fully understand the present disclosure, and are not intended to limit the subject matter of the claims.

[0015] In the following embodiments, we will explain a scene-based stereophonic reproduction technology that uses an Ambisonics microphone as a recording device for recording sound source signals such as sounds, music, and human voices in a sound field recording space (e.g., a live music venue). In the scene-based stereophonic reproduction technology using an Ambisonics microphone, signals (recorded signals) recorded by multiple microphone elements constituting the Ambisonics microphone or point sound sources are represented (encoded) as an intermediate representation ITMR1 (see FIG. 1) or B-format signal using spherical harmonic functions, thereby handling the sound field arriving from all directions in a unified manner in the Ambisonics signal domain (see below). Furthermore, speaker drive signals are generated by decoding this intermediate representation, thereby realizing the desired sound field reproduction in a sound field reproduction space (e.g., a satellite venue).

[0016] (Embodiment 1) First, the concept of scene-based stereophonic sound reproduction technology will be explained with reference to Fig. 1. Fig. 1 is a diagram that schematically shows the concept from sound field recording to sound field reproduction in scene-based stereophonic sound reproduction technology using Ambisonics microphones 11. The Ambisonics microphones 11 are placed in a sound field recording space such as a live venue LV1. At the live venue LV1, a performance is held using multiple sound sources (for example, in the case of a band performance by multiple people, various sound sources such as vocals, bass, guitar, and drums), and the sounds of the performance are recorded by the Ambisonics microphones 11.

[0017] The Ambisonics microphone 11, which is an example of a recording device, includes four microphone elements Mc1, Mc2, Mc3, and Mc4. When the direction Dr1 is the front direction, the microphone elements Mc1 to Mc4 are arranged in midair so as to face the four vertices from the center of the cube CB1 in FIG. 1, and have unidirectionality with respect to each vertex direction. The microphone element Mc1 faces the front left up (FLU) direction of the Ambisonics microphone 11 and records sound in the front left up (FLU) direction. The microphone element Mc2 faces the front right down (FRD) direction of the Ambisonics microphone 11 and records sound in the front right down (FRD) direction. The microphone element Mc3 faces the back left down (BLD) direction of the Ambisonics microphone 11 and records sound in the back left down (BLD) direction. The microphone element Mc4 faces the back right up (BRU) of the Ambisonics microphone 11 and records sound in the direction of the back right up.

[0018] The recorded signals of sounds from these four directions (i.e., FLU, FRD, BLD, BRU) are called A-format signals. A-format signals cannot be used as they are and are converted into B-format signals as intermediate representations ITMR1 having directional characteristics (directivity). B-format signals include, for example, a B-format signal W for sounds from all directions (omnidirectional), a B-format signal X for sounds in the front and rear directions, a B-format signal Y for sounds in the left and right directions, and a B-format signal Z for sounds in the up and down directions. A-format signals are converted into B-format signals using the following conversion formula:

[0019] W=FLU+FRD+BLD+BRU X=FLU+FRD-BLD-BRU Y=FLU-FRD+BLD-BRU Z=FLU-FRD-BLD+BRU

[0020] By combining the B-format signals W, X, Y, and Z, sound signals in all directions (front / back, left / right, up / down) can be obtained. By changing and combining the signal levels of the B-format signals W, X, Y, and Z, sound signals with any directional characteristics in any direction (front / back, left / right, up / down) can be generated. For example, as shown in Figure 1, consider a 3D coordinate system similar to that of a sound field recording space (e.g., a live venue LV1) (i.e., the front / back, left / right, and up / down directions are parallel or parallel). Eight speakers (SPk1, SPk2, SPk3, SPk4, SPk5, SPk6, SPk7, and SPk8) are placed at the vertices of a sound field reproduction space (e.g., a satellite venue STL1) modeled as a cube.

[0021] The positions of the speakers SPk1 to SPk8 are determined by a predetermined distance and angle (azimuth angle θ i and elevation angle φ i ) where i is a variable that indicates the speaker placed in the sound field reproduction space (for example, satellite venue STL1), and takes an integer between 1 and 8 in the example in Figure 1.

[0022] Assume that a listener (user) is present at the center position LSP1 of a sound field reproduction space (for example, satellite venue STL1) and is facing the front direction (Front). Under such circumstances, the sound field in the sound field recording space (for example, live venue LV1) can be freely reproduced in the sound field reproduction space (for example, satellite venue STL1) based on the data of B format signals W, X, Y, and Z obtained by encoding processing based on the A format signal recorded in the sound field recording space (for example, live venue LV1) and the respective directions of the speakers SPk1 to SPk8 in the sound field reproduction space (for example, satellite venue STL1). In other words, when a listener (user) is present in the sound field reproduction space (for example, satellite venue STL1), the front direction of the listener is set as the reference direction, and any three-dimensional direction (for example, a sound source presentation direction θ (described later)) from the reference direction can be freely reproduced in the sound field reproduction space (for example, satellite venue STL1). target ) sound can be reproduced and output.

[0023] Next, a basis of Ambisonics components based on spherical harmonic function expansion for order n and degree m will be described with reference to Fig. 2. Fig. 2 is a diagram showing an example of a basis of Ambisonics components based on spherical harmonic function expansion for order n and degree m.

[0024] The horizontal axis (m) in Figure 2 indicates the degree, and the vertical axis (n) in Figure 2 indicates the order. The degree m ranges from -n to +n. The sum of the spherical harmonics up to n = Nth order is (N + 1). 2 It includes bases. For example, when n=N=0, one basis (i.e., an omnidirectional B format signal W) is obtained. Also, for example, when n=N=1, four bases (i.e., an omnidirectional B format signal W corresponding to (n, m)=(0, 0), a front-to-back B format signal X corresponding to (n, m)=(1, -1), a top-to-bottom B format signal Z corresponding to (n, m)=(1, 0), and a left-to-right B format signal Y corresponding to (n, m)=(1, 1)) are obtained. Note that the same applies when n=N=2 and onward, and therefore description thereof will be omitted.

[0025] It is known that spherical harmonics have the property that their spatial periodicity increases as n and m increase. This makes it possible to express B-format signals with different directional patterns (directional characteristics) depending on the combination of n and m. If the dimension for the order n and degree m is defined as K = n(n + 1) + m based on Ambisonics Channel Numbering (ACN), spherical harmonics can be expressed in vector form as shown in equation (1). In equation (1), the superscript T indicates transposition.

[0026]

number

[0027] [ka]

[0028]

number

[0029]

number

[0030] Next, the system configuration and operation overview of the sound field reproduction system 100 according to embodiment 1 will be described with reference to Fig. 3 and Fig. 4. Fig. 3 is a block diagram showing an example of the system configuration of the sound field reproduction system 100 according to embodiment 1. Fig. 4 is a diagram showing an example of the operation overview from sound field recording to sound field reproduction according to embodiment 1.

[0031] The sound field reproduction system 100 includes a sound field recording device 1 and a sound field reproduction device 2. The sound field recording device 1 and the sound field reproduction device 2 are connected to each other via a network NW1 so as to be able to communicate data with each other. The network NW1 may be a wired network or a wireless network. The wired network corresponds to at least one of a wired local area network (LAN), a wired wide area network (WAN), and power line communication (PLC), for example, and may also be other network configurations capable of wired communication. On the other hand, the wireless network corresponds to at least one of a wireless LAN such as Wi-Fi (registered trademark), a wireless WAN, a short-range wireless communication such as Bluetooth (registered trademark), and a mobile communication network such as 4G or 5G, and may also be other network configurations capable of wireless communication.

[0032] The sound field recording device 1 is placed in, for example, a sound field recording space (for example, a live venue LV1), and includes an Ambisonics microphone 11, an A / D conversion unit 12, an encoding unit 13, and a microphone element direction designation unit 14. It is sufficient that the sound field recording device 1 has at least the Ambisonics microphone 11, and the A / D conversion unit 12, the encoding unit 13, and the microphone element direction designation unit 14 may be provided in the sound field reproduction device 2. In other words, the Ambisonics microphone 11 may be provided outside the sound field reproduction device 2.

[0033] The Ambisonics microphone 11 includes four microphone elements Mc1, Mc2, Mc3, and Mc4. The microphone element Mc1 records sound from the upper left direction in front (see FIG. 1), the microphone element Mc2 records sound from the lower right direction in front (see FIG. 1), and the microphone element Mc3 records sound from the lower left direction behind (see FIG. 1) and the upper right direction behind (see FIG. 1). The Ambisonics microphone 11 may include more unidirectional microphone elements than the four microphone elements Mc1, Mc2, Mc3, and Mc4 arranged in midair, or may include omnidirectional microphone elements arranged on a rigid sphere. Using an Ambisonics microphone with multiple microphone elements enables the encoding unit 13 to synthesize second- or higher-order Ambisonics signals. The signals (recorded signals) recorded by each microphone element constituting the Ambisonics microphone 11 are input to the A / D conversion unit 12.

[0034] The A / D conversion unit 12, the encoding unit 13, and the microphone element direction designation unit 14 are configured by a semiconductor chip or dedicated hardware that implements at least one of electronic devices such as a CPU (Central Processing Unit), a DSP (Digital Signal Processor), a GPU (Graphical Processing Unit), and an FPGA (Field Programmable Gate Array).

[0035] The A / D converter 12 converts analog signals recorded from each microphone element constituting the Ambisonics microphone 11 into digital signals and sends them to the encoder 13 .

[0036] The encoding unit 13 encodes the recorded signal converted by the A / D conversion unit 12 and the direction vector θ from the microphone element direction designation unit 14. m The recorded signal converted by the A / D conversion unit 12 is encoded using the above to generate a low-order basis acoustic signal (for example, a first-order Ambisonics signal). The encoding process by the encoding unit 13 will be described in detail later.

[0037] [ka]

[0038] Here, the encoding process by the encoding unit 13 will be described in detail.

[0039] Generally, it is known that the sound pressure p observed (recorded) at a position of radius r for any angle (θ, φ) on a sphere can be expanded as equation (4) using the spherical harmonic function of equation (2) as a basis for wave number k as a solution to the internal problem in the spherical harmonic domain of the wave equation. In equation (4), A m n is the expansion coefficient, and R n (kr) is a radial function term. In addition, an infinite sum with respect to order n is approximated by truncating it at a finite order N, and the accuracy of sound field reproduction changes depending on this truncation order N. Hereinafter, the truncation order will be expressed as N.

[0040]

number

[0041] [ka]

[0042]

number

[0043]

number

[0044] In equation (6), i is the imaginary unit and j n (kr) is the nth order spherical Bessel function, j ’ n (kr) is its derivative. In this disclosure, the expansion coefficient vector γ for this plane wave m n is treated as a B-format signal (intermediate representation) that is the output of the encoding process by the encoding unit 13. Hereinafter, this expansion coefficient vector may be referred to as an Ambisonics domain signal or simply as an Ambisonics signal.

[0045] More specifically, in the encoding process by the encoding unit 13, the recorded signal, which is a time domain signal after conversion by the A / D conversion unit 12, is converted into an Ambisonics signal (e.g., a first-order Ambisonics signal), and this Ambisonics signal (e.g., a first-order Ambisonics signal) is decoded by each of the first decoding unit 25 and the second decoding unit 26 of the sound field reproduction device 2 and converted into a speaker drive signal.

[0046] [ka]

[0047]

number

[0048]

number

[0049]

number

[0050] [ka]

[0051] The sound field reproduction device 2 is placed, for example, in a sound field reproduction space (for example, satellite venue STL1), and includes a sound source extraction direction control unit 21, a sound source presentation direction control unit 22, a re-encoding unit 23, a speaker direction designation unit 24, a first decoding unit 25, a second decoding unit 26, a signal mixing unit 27, a sound field reproduction unit 28, and speakers SPk1, SPk2, ..., SPk8. Note that in the following description, the number of speakers arranged is eight as an example, but it goes without saying that it is not limited to eight as long as it is an integer equal to or greater than two.

[0052] [ka]

[0053] [ka]

[0054] [ka]

[0055] [ka]

[0056] [ka]

[0057] [ka]

[0058] The signal mixer 27 mixes the speaker drive signals corresponding to the high-order base acoustic signals from the first decoding unit 25 and the speaker drive signals corresponding to the low-order base acoustic signals from the second decoding unit 26 so as to correspond to each speaker, and sends the mixed signals to the sound field reproduction unit 28. Note that the configuration of the signal mixer 27 may be omitted from the sound field reproduction device 2, in which case only the high-order base acoustic signals from the first decoding unit 25 are output from each of the speakers SPk1 to SPk8 via the sound field reproduction unit 28.

[0059] The sound field reproduction unit 28 converts the digital speaker drive signals for each speaker after mixing by the signal mixer 27 into analog speaker drive signals, amplifies the signals, and outputs (reproduces) them from the corresponding speakers.

[0060] Each of the speakers SPk1, SPk2, ..., SPk8 is arranged at a vertex of a sound field reproduction space (e.g., satellite venue STL1) modeled as a cube, and reproduces (reproduces) the sound field based on a speaker drive signal from the sound field reproduction unit 28. The number of speakers installed may vary depending on the sound field to be reproduced. For example, sound field reproduction may be performed using fewer than eight speakers when no specific direction is required, or by combining a commonly known virtual sound image generation method such as a transaural system or a VBAP (Vector Based Amplitude Panning) method. Conversely, sound field reproduction may be performed using more than eight speakers. Furthermore, the speakers may be installed at locations other than the vertices of the sound field reproduction space (e.g., satellite venue STL1) as long as they surround a reference position (e.g., center position LSP1) of satellite venue STL1. Instead of speakers, the sound field reproduction unit 28 may output signals to a binaural reproduction device, such as headphones or earphones, worn by the listener (user). When supplying signals to a binaural reproduction device (e.g., the above-mentioned headphones or earphones) worn by the listener (user), the sound field reproduction unit 28 may generate reproduction signals corresponding to azimuth angles of +-90° using a decoding process described below. Alternatively, the sound field reproduction unit 28 may generate virtual sound images for multiple directions surrounding the head and generate reproduction signals by multiplying the virtual sound images for the corresponding directions in the frequency domain or convolving them in the time domain with transfer characteristics, such as Head Related Transfer Functions (HRTFs), corresponding to the multiple angles, that allow the user to perceive a three-dimensional sound image. This allows the sound field to be reproduced not only from the speakers SPk1, SPk2, ..., SPk8 located at the satellite venue STL1, but also from a reproduction device (e.g., the above-mentioned headphones or earphones) worn by the listener (user) located at the satellite venue STL1.

[0061] Here, the re-encoding process by the re-encoding unit 23 and the processes by the first decoding unit 25 and the second decoding unit 26 will be described in detail.

[0062] [ka]

[0063]

change

[0064]

number

[0065]

change

[0066]

number

[0067]

change

[0068]

number

[0069]

change

[0070]

number

[0071]

change

[0072]

number

[0073] Next, the operation procedure of sound field reproduction by the sound field reproduction device 2 will be described with reference to Fig. 5. Fig. 5 is a flowchart showing an example of the operation procedure of sound field reproduction by the sound field reproduction device 2 according to embodiment 1 in chronological order. In the following explanation, the processes of steps St1 and St2 will be described as being executed within the sound field recording device 1, but the process of step St2 may be executed by the sound field reproduction device 2 when components other than the Ambisonics microphone 11 of the sound field recording device 1 are provided within the sound field reproduction device 2.

[0074] [ka]

[0075] After the processing of step St2, the sound field reproduction device 2 executes a series of processing of steps St3 to St6 (i.e., re-encoding processing for generating a higher-order basis acoustic signal) and the processing of step St7 (i.e., decoding processing for generating a lower-order basis acoustic signal) in parallel.

[0076] [ka]

[0077] [ka]

[0078] The signal mixer 27 of the sound field reproduction device 2 mixes the speaker drive signals (an example of the output of the first decoding process) corresponding to the high-order basis acoustic signals from the first decoding unit 25 in step St6 with the speaker drive signals (an example of the output of the second decoding process) corresponding to the low-order basis acoustic signals from the second decoding unit 26 in step St7 so as to correspond to each speaker (step St8). The sound field reproduction unit 28 of the sound field reproduction device 2 converts the digital speaker drive signals for each speaker after mixing by the signal mixer 27 in step St8 into analog speaker drive signals, amplifies the signals, and outputs (plays) them from the corresponding speakers SPk1 to SPk8 (step St9).

[0079] [ka]

[0080] [ka]

[0081] [ka]

[0082] [ka]

[0083] [ka]

[0084] [ka]

[0085] [ka]

[0086] The recording device is configured with an Ambisonics microphone 11, in which multiple microphone elements Mc1 to Mc4 are arranged three-dimensionally so that each faces in a different direction. This allows the sound field recording device 1 to three-dimensionally record the atmospheric sounds of a performance or the like produced by multiple sound sources in a sound field recording space (live venue LV1).

[0087] [ka]

[0088] First, the system configuration and operation overview of a sound field reproduction system 100A according to embodiment 2 will be described with reference to Fig. 6 and Fig. 7. Fig. 6 is a block diagram showing an example of the system configuration of the sound field reproduction system 100A according to embodiment 2. Fig. 7 is a diagram showing an example of the operation overview from sound field recording to sound field reproduction according to embodiment 2. In the description of Figs. 6 and 7, the same reference numerals will be used to simplify or omit the description of the configurations and operations that overlap with those of the corresponding Figs. 3 and 4, and only the different contents will be described.

[0089] The sound field reproduction system 100A includes a sound field recording device 1 and a sound field reproduction device 2A. The configuration of the sound field recording device 1 is the same as that in the first embodiment, and therefore a description thereof will be omitted.

[0090] The sound field reproduction device 2A is placed, for example, in a sound field reproduction space (for example, satellite venue STL1), and includes a sound source extraction direction control unit 21, a sound source presentation direction control unit 22, a re-encoding unit 23, a speaker direction designation unit 24, a first decoding unit 25, a sound source acquisition unit 29, a second encoding unit 30, a second signal mixing unit 31, a second decoding unit 32, a signal mixing unit 27, a sound field reproduction unit 28, and speakers SPk1, SPk2, ..., SPk8.

[0091] The sound source acquisition unit 29 acquires acoustic signals s1[n], ..., sb[n] of multiple sound sources (e.g., various sound sources such as vocals, bass, guitar, and drums) to be presented in a sound field reproduction space (e.g., satellite venue STL1) and sends them to the second encoding unit 30. Each acoustic signal s1[n], ..., sb[n] can be expressed as a point sound source. n indicates a discrete time, and b indicates the number of sound sources. These sound sources may be recorded individually in the sound field recording space (live venue Lv1), or may be sound sources unrelated to the sound field recording space.

[0092] [ka]

[0093]

number

[0094] The second signal mixer 31 mixes the high-order basis acoustic signals (for example, N-th order Ambisonics signals) for each sound source obtained by the encoding process by the second encoder 30, and sends the mixed signals to the second decoder 32.

[0095] [ka]

[0096]

number

[0097] Next, the operational procedure for sound field reproduction by the sound field reproduction device 2A will be described with reference to Fig. 8. Fig. 8 is a flowchart showing an example of the operational procedure for sound field reproduction by the sound field reproduction device 2A according to embodiment 2 in chronological order. In the description of Fig. 8, the same step numbers will be assigned to processes that overlap with the description of Fig. 5, and the description will be simplified or omitted, and only different contents will be described.

[0098] 8, the sound source acquisition unit 29 of the sound field reproduction device 2A acquires acoustic signals s1[n], ..., sb[n] (examples of point sound source signals) of multiple sound sources (for example, various sound sources such as vocals, bass, guitar, drums, etc.) to be presented in a sound field reproduction space (for example, satellite venue STL1) (step St11). The second encoding unit 30 of the sound field reproduction device 2A encodes direction vectors θ of the b point sound sources. b is read from a memory (not shown) or acquired based on a specification from a user interface (not shown) (step St12).

[0099] [ka]

[0100] As described above, the sound field reproduction device 2A according to the second embodiment further includes a second encoding unit 30 that encodes each of a plurality of sound source signals (e.g., sound signals from various sound sources such as vocals, bass, guitar, and drums) to be presented in the sound field reproduction space (satellite venue STL1) to generate second higher-order basis acoustic signals (Nth-order Ambisonics signals), and a second signal mixing unit 31 that mixes the second higher-order basis acoustic signals for each sound source signal. As a result, the sound field reproduction device 2A according to the second embodiment can output atmospheric sounds from sound sources that are to be uniquely presented in the sound field reproduction space (satellite venue STL1), which is different from the sound field recording space (live venue LV1), with high directional resolution due to the high-order basis.

[0101] [ka]

[0102] [ka]

[0103] Although the embodiments have been described above with reference to the accompanying drawings, the present disclosure is not limited to such examples. It is clear that a person skilled in the art can conceive of various modifications, alterations, substitutions, additions, deletions, and equivalents within the scope of the claims, and it is understood that these also fall within the technical scope of the present disclosure. Furthermore, the components of the above-described embodiments may be combined in any manner without departing from the spirit of the invention. [Industrial Applicability]

[0104] The present disclosure is useful as a sound field reproduction device, a sound field reproduction method, and a sound field reproduction system that utilizes low-order sound field components recorded using an Ambisonics microphone and suppresses an increase in sound source localization errors within a sound field reproduction space. [Explanation of symbols]

[0105] 1. Sound field recording device 2. 2A sound field reproduction device 11 Ambisonics Microphone 12 A / D conversion section 13 Encoding section 14 Microphone element direction specification section 21 Sound source extraction direction control unit 22 Sound source presentation direction control unit 23 Re-encoding section 24 Speaker direction specification section 25 First Decoding Unit 26, 32 Second Decoding Unit 27 Signal mixing section 28 Sound field reproduction section 29 Sound source acquisition section 30 Second encoding section 31 2nd signal mixing section 100, 100A sound field reproduction system SPk1, SPk2, SPk3, SPk4, SPk5, SPk6, SPk7, SPk8 speakers

Claims

1. a sound source extraction direction control unit that receives a designation of a sound source extraction direction within a sound field recording space in which a recording device is placed; a re-encoding unit that re-encodes an acoustic signal corresponding to the sound source extraction direction among low-order basis acoustic signals based on an encoding process using a signal recorded by the recording device to generate a high-order basis acoustic signal; a sound field reproduction unit that outputs a signal based on the high-order base acoustic signal from each of a plurality of speakers provided in a sound field reproduction space different from the sound field recording space, Sound field reproduction device.

2. a sound source presentation direction control unit that receives a designation of a sound source presentation direction that is the same as or different from the sound source extraction direction and is a direction of emphasizing sound field reproduction in a sound field reproduction space that is different from the sound field recording space; the re-encoding unit performs the re-encoding using the acoustic signal and the sound source presentation direction to generate the high-order basis acoustic signal. The sound field reproduction device according to claim 1 .

3. a first decoding unit that generates a first sound field drive signal having a high-order basis component for each of the speakers by using the high-order basis acoustic signal and arrangement information of each of the plurality of speakers; The sound field reproduction device according to claim 1 .

4. a second decoding unit that generates a second sound field drive signal having a low-order basis component for each of the speakers by using the low-order basis acoustic signal and arrangement information of each of the plurality of speakers, The sound field reproduction device according to claim 3.

5. a signal mixer that mixes the first sound field drive signal and the second sound field drive signal for each speaker; the sound field reproduction unit outputs the mixed signal by the signal mixer to each of the speakers as a signal based on the high-order basis acoustic signal. The sound field reproduction device according to claim 4.

6. The sound source extraction direction is specified as a three-dimensional direction from a reference position in the sound field recording space. The sound field reproduction device according to claim 1 .

7. The sound source presentation direction is specified as a three-dimensional direction from a reference position in the sound field recording space. The sound field reproduction device according to claim 2.

8. a second encoding unit that encodes each of a plurality of sound source signals to be presented in the sound field reproduction space to generate second higher-order basis acoustic signals; a second signal mixer that mixes the second higher-order basis acoustic signals for each of the sound source signals, The sound field reproduction device according to claim 3.

9. a second decoding unit that generates a third sound field drive signal having a high-order basis component for each speaker by using the second high-order basis acoustic signal for each sound source signal after mixing by the second signal mixer and arrangement information of each of the plurality of speakers, The sound field reproduction device according to claim 8.

10. a signal mixing unit that mixes the first sound field drive signal and the third sound field drive signal for each speaker, the sound field reproduction unit outputs the mixed signal by the signal mixer to each of the speakers as a signal based on the high-order basis acoustic signal. The sound field reproduction device according to claim 9.

11. receiving a designation of a sound source extraction direction within a sound field recording space in which a recording device is placed; generating high-order basis acoustic signals by re-encoding acoustic signals corresponding to the sound source extraction direction among low-order basis acoustic signals based on an encoding process using the signals recorded by the recording device; outputting a signal based on the high-order base acoustic signal from each of a plurality of speakers provided in a sound field reproduction space different from the sound field recording space, Sound field reproduction method.

12. a sound field recording device having a recording device capable of recording a sound source in a sound field recording space; a sound field reproduction device that reproduces the acoustic signal recorded by the recording device in a sound field reproduction space that is different from the sound field recording space, The sound field reproduction device comprises: a sound source extraction direction control unit that receives a designation of a sound source extraction direction within the sound field recording space; a re-encoding unit that re-encodes an acoustic signal corresponding to the sound source extraction direction among low-order basis acoustic signals based on an encoding process using a signal recorded by the recording device to generate a high-order basis acoustic signal; a sound field reproduction unit that outputs a signal based on the high-order base acoustic signal from each of a plurality of speakers provided in the sound field reproduction space, Sound field reproduction system.

13. The recording device is configured by an Ambisonics microphone in which a plurality of microphone elements are arranged three-dimensionally so that each of the microphone elements faces in a different direction. A sound field reproduction system according to claim 12.

Citation Information

Patent Citations

  • Method and apparatus for encoding and optimally reproducing a three-dimensional sound field

    JP2012514358A

  • Method and apparatus for enhancing the directivity of a first-order ambisonics signal

    JP2016517033A

  • Signal processing device, signal processing method, and program

    JP2019192975A

  • Spatial audio signal format generation from microphone array using adaptive capture

    JP2019530389A

  • Apparatus and method for encoding a spatial audio representation or decoding an audio signal encoded using transport metadata, and related computer program

    JP2022518744A