Information processing apparatus and method, and program
The described technology addresses the issue of reduced sound field control accuracy in enclosed spaces by calculating correction coefficients to match attenuation parameters, enabling precise sound field control and continuous reproduction in specified areas.
Patent Information
- Application Number
- PCT/JP2024/014550
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-05
- Filing Date
- 2024-04-10
- Publication Date
- 2025-08-14
AI Technical Summary
Conventional sound field control methods based on wave field synthesis and high-order Ambisonics fail to accurately control sound fields in enclosed spaces with elastic walls due to mismatched attenuation parameters, leading to reduced reproduction accuracy.
An information processing device and method that calculates correction coefficients based on attenuation parameters of both the desired and controlled sound fields, adjusting drive signals to match these parameters and perform continuous sound field control in specified areas using a speaker and microphone array.
Enables precise sound field control in enclosed spaces by correcting attenuation characteristics, allowing for continuous control in specified areas and improving reproduction accuracy.
Smart Images

Figure JP2024014550_14082025_PF_FP_ABST
Abstract
Description
Information processing device, method, and program
[0001] The present technology relates to an information processing device, method, and program, and in particular to an information processing device, method, and program that enable sound field control with higher accuracy.
[0002] Sound field control (spatial acoustic reproduction) within a vehicle cabin or other enclosed space is one of the key technologies in in-car and home entertainment.
[0003] For example, in the technical field of sound field control, commonly used methods include wave field synthesis based on Rayleigh integrals (see, for example, Non-Patent Document 1) and high-order Ambisonics methods based on spherical harmonic expansion or circular harmonic expansion (see, for example, Non-Patent Document 2).
[0004] These spatial control methods (sound field control techniques) are often based on the Kirchhoff-Helmholtz integral equation in a free sound field (a space without obstacles or attenuation characteristics), and involve controlling the entire space by controlling the sound pressure and particle velocity on a closed boundary.
[0005] Furthermore, as a technology for controlling a sound field, for example, a technology has been proposed in which two speakers with different distance attenuation rates are used to separate a sound field into a non-sound-reduced area and a sound-reduced area (see, for example, Patent Document 1).
[0006] Japanese Patent Application Laid-Open No. 2006-270409
[0007] AJ Berkhout, D. de Vries, P. Vogel, "Acoustic control by wave field synthesis," The Journal of the Acoustical Society of America 93(5), 1993: 2764-2778.MA Poletti, "Three-dimensional surround sound systems based on spherical harmonics," Journal of the Audio Engineering Society 53(11), 2005: 1004-1025.
[0008] Incidentally, the sound field control methods using the wave field synthesis method and the high-order Ambisonics method, etc., described above, are basically premised on a free sound field. However, the sound field in a vehicle cabin or other enclosed space has elastic walls and attenuation characteristics, and is not an appropriate control target for the above sound field control methods, resulting in low control accuracy in enclosed spaces.
[0009] Specifically, in a room with elastic walls or air damping, the attenuation parameters of the sound wave propagation in the desired sound field and the controlled sound field may differ. In such cases, the Kirchhoff-Helmholtz integral equation does not hold, and the sound field control methods described above either cannot control the sound field or the reproduction accuracy of the sound field is significantly reduced.
[0010] The present technology has been made in view of such circumstances, and makes it possible to control a sound field with higher precision.
[0011] An information processing device according to one aspect of the present technology includes a correction coefficient calculation unit that calculates a correction coefficient based on a first attenuation parameter indicating the attenuation characteristics of a desired sound field to be reproduced, a second attenuation parameter indicating the attenuation characteristics of a control sound field, and information indicating a control area of the control sound field, and a drive signal calculation unit that calculates a drive signal for a speaker to form the control sound field based on an acoustic signal of the desired sound field, the correction coefficient, and a transfer function indicating the transfer characteristics of the control sound field.
[0012] An information processing method or program according to one aspect of the present technology includes the steps of calculating a correction coefficient based on a first attenuation parameter indicating the attenuation characteristics of a desired sound field to be reproduced, a second attenuation parameter indicating the attenuation characteristics of a controlled sound field, and information indicating a control area of the controlled sound field, and calculating a drive signal for a speaker to form the controlled sound field based on an acoustic signal of the desired sound field, the correction coefficient, and a transfer function indicating the transfer characteristics of the controlled sound field.
[0013] In one aspect of the present technology, a correction coefficient is calculated based on a first attenuation parameter indicating the attenuation characteristics of a desired sound field to be reproduced, a second attenuation parameter indicating the attenuation characteristics of a controlled sound field, and information indicating a control area of the controlled sound field, and a drive signal for a speaker to form the controlled sound field is calculated based on the acoustic signal of the desired sound field, the correction coefficient, and a transfer function indicating the transfer characteristics of the controlled sound field.
[0014] FIG. 1 is a diagram showing a reproduction error of sound field control. FIG. 2 is a diagram showing an example of the configuration of a sound field control system. FIG. 3 is a diagram explaining a control area. FIG. 4 is a flowchart explaining sound field control processing. FIG. 5 is a diagram showing a reproduction error of sound field control with correction method 1. FIG. 6 is a diagram explaining sound field reproduction with correction method 1. FIG. 7 is a diagram explaining sound field reproduction with correction method 2. FIG. 8 is a diagram explaining sound field reproduction with correction method 2. FIG. 9 is a diagram showing an example of the configuration of a computer.
[0015] Hereinafter, embodiments to which the present technology is applied will be described with reference to the drawings.
[0016] First Embodiment Overview of the Present Technology As described above, in a commonly used sound field control method, when the attenuation parameters of sound wave propagation in a desired sound field and a controlled sound field are different, the Kirchhoff-Helmholtz integral equation does not hold, and therefore the sound field cannot be controlled or the reproduction accuracy of the sound field is significantly reduced.
[0017] The results of a computer simulation of sound field control using a general sound field control method are shown in Figure 1. In particular, Figure 1 shows the results for a two-dimensional horizontal plane including the center of the sound field in a three-dimensional space.
[0018] In FIG. 1, the upper part shows an example in which the attenuation characteristics (attenuation parameters) of sound wave propagation in the desired sound field and the controlled sound field are different (do not match).
[0019] That is, the part indicated by arrow Q11 shows the desired sound field to be reproduced, and the part indicated by arrow Q12 shows the controlled sound field obtained by reproducing the desired sound field using a sound field control technique. In the part indicated by arrow Q11 and the part indicated by arrow Q12, the vertical and horizontal directions indicate directions in space, and the shading at each position indicates the sound pressure level at that position.
[0020] The portion indicated by arrow Q13 shows the reproduction error (Root Mean Squared Error (RMSE)) between the desired sound field indicated by arrow Q11 and the controlled sound field indicated by arrow Q12.
[0021] In the figure, the lower part shows an example in which the attenuation characteristics (attenuation parameters) of sound wave propagation in the desired sound field and the controlled sound field match.
[0022] That is, the area indicated by arrow Q14 shows the desired sound field to be reproduced, and the area indicated by arrow Q15 shows the controlled sound field obtained by reproducing the desired sound field using a sound field control technique. In the areas indicated by arrows Q14 and Q15, the vertical and horizontal directions indicate spatial directions, and the shading at each position indicates the sound pressure level at that position.
[0023] The portion indicated by arrow Q16 shows the reproduction error (RMSE) between the desired sound field indicated by arrow Q14 and the controlled sound field indicated by arrow Q15.
[0024] In Figure 1, the crosses "+" indicate the positions of the speakers used to form the controlled sound field, i.e., to control the sound field, and the dots indicate the positions of the microphones used to record (collect) the desired sound field. Here, the area in the space where the microphones are located, i.e., the area in the center of the space, is the control target area where we want to form a controlled sound field with minimal reproduction error.
[0025] In the figure, comparing the example in which the damping characteristics match on the bottom with the example in which the damping characteristics do not match on the top, it is clear that the reproduction error (RMSE) is small in the example in which the damping characteristics match, whereas the reproduction error increases in the example in which the damping characteristics do not match.
[0026] Therefore, in this technology, even if the spatial attenuation characteristics (attenuation parameters) of the desired sound field to be reproduced and the reproduced control sound field differ, it is possible to reproduce the sound field with higher accuracy, i.e., it is possible to control the sound field with higher accuracy.
[0027] In this technology, in sound field control (spatial control of sound) using a speaker array and a microphone array, the signals recorded by the microphone array, the attenuation parameters of the physical space to be controlled, and information indicating the control area are input, and wave number correction processing is performed based on the attenuation parameters.
[0028] By doing so, whereas in the past it was not possible to control a sound field if the attenuation parameters of the desired sound field and the control sound field did not match, with this technology it is possible to control a sound field in a closed space in a specified control area by performing a correction process that takes into account the attenuation parameters of the space including that control area.
[0029] Attenuation parameters are parameters that represent the attenuation characteristics of sound in a specific space or propagation medium. The definition of attenuation in this specification is different from the simple attenuation of sound due to the spread of the sound source, more specifically, the spread of sound from the sound source (radiation characteristics), and refers to attenuation (damping) due to acoustic modes in space and the viscosity of the propagation medium. For example, attenuation due to acoustic modes in space is significant in closed spaces with glass or elastic walls.
[0030] Furthermore, in the present technology, the control area, which is the region to be controlled by sound field control, is specified by, for example, a single radius, a set of multiple radii, or a range of radii.
[0031] The wave number is the number of waves per unit length. Generally, the wave number k is defined as k = ω / c, where ω is the angular frequency and c is the sound speed. However, in a broad sense, the wave number k' can be defined as k' = (1 - jα)k, where α is the attenuation parameter, j is the imaginary unit, and k is the wave number.
[0032] When the attenuation parameter α and the wave number k are not 0, the wave number k′ is a complex number, and therefore, hereinafter, this broad wave number k′ will also be referred to as a complex wave number.
[0033] Furthermore, the above-mentioned correction process is a process for performing correction so as to match the attenuation parameters of the control sound field and the desired sound field in a specified control area, for example, and is performed in a spherical harmonic domain or a circular harmonic domain.
[0034] The sound field control method of this technology is characterized by its ability to control sound fields in closed spaces, which conventional sound field control methods cannot handle, and to correct the attenuation characteristics of the space.
[0035] Specifically, this technology is characterized by taking into account the attenuation characteristics of a space in sound field control using physical quantities that can express attenuation characteristics, such as complex wave numbers, and by introducing a correction term that matches the attenuation parameters of the controlled sound field and the desired sound field for a specific control area.
[0036] Furthermore, in conventional sound field control methods, when the attenuation parameters of the desired sound field and the controlled sound field are different, discrete control is performed at the positions of individual control points rather than the entire space. Note that when the control points are arranged on the same spherical surface, continuous control is performed near the surface of the sphere, i.e., control is performed over a continuous area.
[0037] In contrast, the sound field control method of this technology makes it possible to specify the radius of a sphere where continuous control is possible as the control area, and furthermore, makes continuous control possible not only in the vicinity of the sphere but also in part of the space.
[0038] For example, when playing highly immersive spatial audio content in a closed space such as a car interior, the attenuation parameters of the car interior and the content being played may differ. Even in such cases, the sound field control method of this technology makes it possible to adjust the controllable area (sweet spot), providing a highly immersive stereophonic experience.
[0039] <Configuration Example of Sound Field Control System> FIG. 2 shows a configuration example of an embodiment of a sound field control system to which the present technology is applied.
[0040] The sound field control system shown in FIG. 2 includes a microphone array 11, a signal acquisition unit 12, a transfer function calculation unit 13, a correction coefficient calculation unit 14, a signal correction unit 15, a transfer function correction unit 16, a drive signal calculation unit 17, and a speaker array 18.
[0041] The signal acquisition unit 12 to the drive signal calculation unit 17 may be realized by one information processing device such as a computer, or may be realized by multiple information processing devices. Also, the sound field control system does not necessarily need to be provided with the microphone array 11.
[0042] The microphone array 11 is made up of a plurality of microphones arranged in a space where a primary sound field, which is a desired sound field to be reproduced (to be reproduced), is formed. Specifically, the microphone array 11 is made up of, for example, a spherical microphone array in which a plurality of microphones are arranged in a spherical shape.
[0043] It is assumed here that the microphone array 11 is configured with M microphones. The microphone array 11 supplies observation signals obtained by collecting ambient sounds to the signal acquisition unit 12 as acoustic signals of a primary sound field.
[0044] The signal acquisition unit 12 acquires an acoustic signal (observation signal) of the primary sound field and supplies it to the signal correction unit 15. For example, the acoustic signal of the primary sound field supplied to the signal correction unit 15 is a signal in the spherical harmonic domain of K channels. Here, the number of channels K is determined according to the maximum order N of the spherical harmonic expansion, and K = (N + 1) 2 In other words, the number of channels K is the number of combinations of the degree n and the order m of the spherical harmonic domain.
[0045] For example, the signal acquisition unit 12 performs spherical harmonic expansion on the time-domain acoustic signal supplied from the microphone array 11 based on microphone coordinates indicating the positions of the microphones constituting the microphone array 11, to obtain an acoustic signal in the spherical harmonic domain of the primary sound field. This acoustic signal of the primary sound field is a desired signal that serves as an input for calculating a drive signal for forming a controlled sound field.
[0046] The acoustic signals in the spherical harmonic domain of the primary sound field are not limited to those calculated from the observation signals obtained by the microphone array 11, but may also be theoretical values calculated from the positions of virtual microphones, i.e., the positions of the observation points in the primary sound field or the positions of the sound sources in the primary sound field.
[0047] The transfer function calculation unit 13 calculates the transfer function of a secondary sound field, which is a controlled sound field formed by the speaker array 18, based on speaker coordinates and the like indicating the arrangement position of each speaker that constitutes the speaker array 18, and supplies the calculated transfer function to the transfer function correction unit 16. The transfer function of the secondary sound field here refers to a transfer function that indicates the transfer characteristics from the speakers that constitute the speaker array 18 to control points (positions corresponding to observation points) in the space in which the secondary sound field is formed.
[0048] Here, the speaker array 18 is assumed to consist of L speakers, and the speaker coordinates are coordinate information indicating the positions of the speakers in the space where the secondary sound field is formed. The secondary sound field is the sound field to be controlled, and ideally, the secondary sound field coincides with the primary sound field in a predetermined control area.
[0049] For example, the transfer function calculation unit 13 calculates the transfer function of the secondary sound field for each speaker that makes up the speaker array 18 by theoretical value calculation based on the speaker coordinates, and obtains the transfer function in the spherical harmonic domain by performing spherical harmonic expansion on the obtained transfer function.
[0050] The transfer function calculation unit 13 supplies the transfer function for each combination of the order n and the order m of the spherical harmonic domain obtained for each speaker, that is, the transfer function of the K×L channel second-order sound field, to the transfer function correction unit 16.
[0051] The correction coefficient calculation unit 14 receives as input the attenuation characteristics (attenuation parameters) of the primary sound field, the attenuation characteristics (attenuation parameters) of the secondary sound field, the frequency of the sound to be controlled (hereinafter also referred to as the control frequency), and the control radius, and calculates the correction coefficients of the primary sound field and the secondary sound field based on these inputs.
[0052] The control radius is information indicating a control area, which is an area targeted for control of sound pressure, etc. in a space where a secondary sound field is formed (hereinafter also referred to as a reproduction space). During sound field control, sound pressure, etc. within the control area are controlled so that the secondary sound field in the control area becomes as close as possible to the primary sound field (so that the error is minimized). Here, the control area is assumed to be a spherical area (space) specified by the control radius, more specifically, an area on the surface of the sphere, or an area inside the sphere including the surface of the sphere.
[0053] The correction coefficients for the primary sound field and the secondary sound field are coefficients calculated for each order n of the spherical harmonic domain. For channels determined by a combination of order n and order m, the same correction coefficients are used for channels with the same order n, regardless of the order m. The correction coefficient calculation unit 14 supplies the correction coefficients for the spherical harmonic domain of the primary sound field to the signal correction unit 15, and supplies the correction coefficients for the spherical harmonic domain of the secondary sound field to the transfer function correction unit 16.
[0054] The signal correction unit 15 corrects the acoustic signal of the primary sound field supplied from the signal acquisition unit 12 based on the correction coefficient of the primary sound field supplied from the correction coefficient calculation unit 14, and supplies the resulting corrected acoustic signal (corrected acoustic signal) to the drive signal calculation unit 17.
[0055] The transfer function correction unit 16 corrects the transfer function of the secondary sound field supplied from the transfer function calculation unit 13 based on the correction coefficient of the secondary sound field supplied from the correction coefficient calculation unit 14, and supplies the resulting corrected transfer function (corrected transfer function) to the drive signal calculation unit 17.
[0056] The drive signal calculation unit 17 calculates a drive signal for forming a secondary sound field in the reproduction space based on the corrected acoustic signal of the primary sound field supplied from the signal correction unit 15 and the corrected transfer function of the secondary sound field supplied from the transfer function correction unit 16, and supplies the drive signal to the speaker array 18.
[0057] The drive signals are speaker drive signals that drive the speakers that make up the speaker array 18 to output sounds that form a secondary sound field. For example, when calculating the drive signals, a cost function based on the acoustic signals of the primary sound field and the transfer function of the secondary sound field is minimized to calculate (generate) drive signals for driving each of the L speakers that make up the speaker array 18, i.e., drive signals for the L channels. For example, the drive signals may be signals that output sounds with a different wave number from the input acoustic signals of the primary sound field.
[0058] Since the drive signal calculation unit 17 uses the corrected acoustic signal and transfer function, it can also be said that the drive signal is calculated based on the correction coefficient of the primary sound field, the correction coefficient of the secondary sound field, the acoustic signal of the primary sound field before correction, and the transfer function of the secondary sound field before correction.
[0059] The speaker array 18 is made up of L speakers arranged in the reproduction space, and forms a secondary sound field in the reproduction space by outputting sounds based on the drive signals supplied from the drive signal calculation unit 17. For example, the speaker array 18 is made up of a spherical speaker array (spherical speaker array) in which multiple speakers are arranged in a spherical shape.
[0060] The space in which the primary sound field is formed (hereinafter also referred to as the recording space) may be an open space or any closed space including the interior of a vehicle, etc. Similarly, the reproduction space in which the secondary sound field is formed may be an open space or a closed space.
[0061] In the sound field control system shown in Figure 2, when controlling the sound field of a reproduction space such as a closed space, in addition to the acoustic signal (desired signal) of the primary sound field recorded by a microphone array 11 or the like and the transfer function from the speaker to the control point, attenuation parameters (or wave numbers) of the primary sound field (desired sound field) and secondary sound field (control sound field) and a control radius indicating the control area are also input.
[0062] Then, for the control area specified by the input, a correction process is performed to match the attenuation characteristics of the secondary sound field with those of the primary sound field. In other words, an extrapolation process or an interpolation process is performed so that the position indicated by the control radius becomes the position (control point) of the control target such as sound pressure. As a result, a speaker drive signal is generated (calculated) so that the primary sound field is reproduced (spatially) continuously in the control area.
[0063] The wavenumbers of the primary sound field and the secondary sound field may be the same or different. Furthermore, as the control radius, one radius may be input to specify an area on the surface of the sphere that will become the control area, or multiple radii or a range of control radii may be input to specify an area inside the sphere that will become the control area. For example, when multiple control radii or a range of control radii are input, the control radii may be thinned out with respect to the frequency (wavenumber). Alternatively, the recording space or the reproduction space may be a closed space such as a vehicle cabin, or an open space. For example, one of the recording space and the reproduction space may be a closed space, and the other an open space.
[0064] This technology will be described in further detail below.
[0065] <About the Present Technology> (A. Definition of Attenuation Characteristics and Attenuation Parameter in Closed Space) The attenuation parameter α in this specification is a parameter that represents acoustic energy loss (sound attenuation) in a specific space or propagation medium.
[0066] (1. Damping characteristics in closed spaces) Acoustic modes exist in closed spaces, and acoustic damping (damping) can occur for each mode due to the influence of boundary materials, etc. Sound absorption in closed spaces can also be approximately modeled as modal damping (see, for example, the following references 1 and 2). This damping is particularly noticeable in closed spaces that have elastic walls (not rigid walls) such as glass, and where coupling between sound and vibration exists.
[0067] The above modeling is described in detail in, for example, "H. Kuttruff, Room acoustics (sixth edition), CRC Press, 2016." (referred to as Reference 1) and "ME Johnson, et al., "An equivalent source technique for calculating the sound field inside an enclosure containing scattering objects," The Journal of the Acoustical Society of America, 104(3), 1998: 1221-1231." (referred to as Reference 2).
[0068] (2. Attenuation characteristics in propagation media) In viscous propagation media, there is a phenomenon in which some of the acoustic energy during the propagation of vibrations (sound vibrations) is converted into thermal energy as viscous loss. There is also loss due to particle-level effects. This loss of acoustic energy during the propagation process is called acoustic damping. This attenuation occurs more significantly in liquid media such as water, but in the case of air, it is omitted in most research and technology by setting α = 0.
[0069] In the case of air, if the propagation is long distance or if there is a reflected wave that travels back and forth, it may be considered as "air absorption" or "air attenuation."
[0070] (3. Others) All acoustic attenuation phenomena that can be modeled with complex wavenumbers, as explained in "C. Speaker-microphone arrays and their signals," below, such as the synergistic effects of attenuation explained in "1. Attenuation characteristics in closed spaces" and "2. Attenuation characteristics in propagation media" above, and energy losses due to scattering, diffusion, refraction, etc., are considered attenuation characteristics.
[0071] In addition, for point sound sources, line sound sources, and other sources that have radiation characteristics in which sound waves spread from the source position, the energy is dispersed in space as it propagates, and the observed sound pressure attenuates in proportion to the propagation distance, a phenomenon known as "distance attenuation." Although this type of distance attenuation is sometimes abbreviated to "attenuation" in conventional technology, it is different from the attenuation characteristics described above because it is not a loss of acoustic energy in space.
[0072] In the field of architectural acoustics, terms such as "decay time (energy decay time)" and "decay curve (energy decay curve)" are often used to describe the phenomenon caused by the combined action of distance attenuation and room wall reflections. However, this "decay" is a statistical expression and differs from the attenuation characteristics described in this section, which describe the physical phenomenon.
[0073] (B. Complex Wave Numbers) In the strict sense (real number) of wave numbers, the wave number k is generally defined as k = 2πf / c = ω / c = 1 / λ, where π is the constant of the circumference of a circle, f is frequency, c is the speed of sound, ω is angular frequency, and λ is wavelength.
[0074] The complex wave number k', which takes into account the attenuation parameter α described in "A. Definition of attenuation characteristics and attenuation parameters for closed spaces," can be defined as k' = (1 - jα)k, where j is the imaginary unit (see, for example, Reference 2).
[0075] When the attenuation parameter α=0, k′=k, and the complex wave number k′ is equal to a general wave number (real wave number). When the attenuation parameter α≠0 and K≠0 exists, the complex wave number k′ becomes a complex number.
[0076] Here, a physical phenomenon expressed by a complex wave number will be described using a point sound source as an example.
[0077] If the distance from the sound source to the sound receiving point is x, the sound pressure observed at the sound receiving point can be calculated using the transfer function g.
[0078] Transfer function g of a point sound source in three-dimensional space using real wave number k x,k can be expressed by the following equation (1).
[0079]
[0080] On the other hand, when a complex wave number k' is used, the transfer function g of a point sound sourcex,k’ can be rearranged as in the following equation (2).
[0081]
[0082] Transfer function g x,k’ is the acoustic propagation -αkx It can be considered that the propagation distance x, the wave number k determined by the frequency and the speed of sound, and the attenuation term dependent on the subtraction parameter α are introduced. In the following explanation, since the explanation is the same whether it is a complex number or a real number, the complex wave number will also be referred to as the wave number, and k' will also be written simply as k.
[0083] (C. Speaker / Microphone Array and Its Signals) Sound field control and reproduction is a technology that uses microphones to observe (collect) sound waves in a specified space (recording space) and then uses speakers to reproduce similar sound waves in a target space (reproduction space). In particular, the sound observed in the recording space is called the primary sound field, and the sound reproduced in the reproduction space is called the secondary sound field.
[0084] Since sound field control and sound field reproduction technologies aim to reproduce sound throughout an entire space (sound field), rather than just one or a few points in the space, they often use multiple speakers and multiple microphones arranged in the space. Generally, what is obtained by arranging multiple speakers is called a speaker array, and what is obtained by arranging multiple microphones is called a microphone array.
[0085] Sound field control is a process of controlling the driving signal d of a speaker so that the primary sound field signal p(x, k), which is the acoustic signal of the primary sound field, and the secondary sound field signal p'(x, k), which is the acoustic signal of the secondary sound field, match at all observation point positions x in the control area Ω, as shown in the following equation (3). l The secondary sound field signal p'(x,k) is an acoustic signal observed at the observation point x in the reproduction space.
[0086]
[0087] Here, the transfer function from the speaker in the secondary sound field (reproduction space) to the observation point is g l (x, k), and the speaker drive signal input to the speaker array is dl (k), the secondary sound field signal p'(x,k) when sound field control is performed is expressed by the following equation (4).
[0088]
[0089] Transfer function g l (x, k) and driving signal d l (k) is the transfer function and drive signal for the l-th speaker that makes up the speaker array. Although it may be passed to the processing block implicitly, the primary sound field signal p(x,k) and the secondary sound field speaker transfer function g are essential as inputs for sound field control. l (x,k).
[0090] The input of the primary sound field signal and the secondary sound field transfer function has the following variations:
[0091] (Acquisition method) (1) Sound signal picked up by a microphone array (measurement and sound recording in real space) (2) Virtual sound source signal (theoretical value) that can be calculated from spatial information including observation point coordinates and sound source location
[0092] (Data format) (1) Time domain signal (2) Frequency domain signal (3) Spherical harmonic domain signal (4) Other signals from which spherical harmonic domain signals can be calculated
[0093] For example, an example of acquisition method (1) is a case where an acoustic signal of a primary sound field is acquired by the microphone array 11 in the sound field control system shown in Fig. 2. In contrast to this, acquisition method (2) is a case where a theoretical value of an acoustic signal of a primary sound field calculated in advance is input to the signal acquisition unit 12.
[0094] Furthermore, when the acoustic signal of the primary sound field is a time domain signal or a frequency domain signal, the signal acquisition unit 12 converts the acquired (input) acoustic signal of the primary sound field into an acoustic signal in the spherical harmonic domain and supplies it to the signal correction unit 15. Similarly, the transfer function of the secondary sound field is also converted into a transfer function in the spherical harmonic domain by the transfer function calculation unit 13 as appropriate and supplied to the transfer function correction unit 16.
[0095] The spherical harmonic domain signal can be obtained by performing a spatial Fourier transform on the time domain or frequency domain signal. Alternatively, as will be described later, the acoustic signal of the primary sound field and the transfer function of the secondary sound field may be transformed into a signal in the circular harmonic domain (cylindrical harmonic domain) instead of the spherical harmonic domain signal.
[0096] The variations (data formats) of the drive signal that is the output of the drive signal calculation unit 17 can be two types: a time domain signal and a frequency domain signal.
[0097] Furthermore, in sound field control, the same processing method is basically used for any spatial arrangement of speakers and microphones.
[0098] 3, a primary sound field is collected (recorded) by a microphone array MK11, and a secondary sound field is formed by a speaker array SP11. Note that the sound field control system of the present technology can control the sound field for any area within the area surrounded by the speaker array SP11. In other words, any area within the area surrounded by the speaker array SP11 can be set as a control area.
[0099] However, the principles of sound field control require that there are no sound sources or obstacles within the control area Ω (both in the primary and secondary sound fields), as shown for example in FIG.
[0100] (D. Sound Field Control Processing) In the sound field control system shown in FIG. 2, the correction coefficient calculation unit 14 to the transfer function correction unit 16 perform characteristic processing according to the present technology that has not been performed in the past.
[0101] In contrast to this, the signal acquisition unit 12, transfer function calculation unit 13, and drive signal calculation unit 17 in the sound field control system perform processing similar to that of the conventional sound field control technology.
[0102] For example, when the signal acquisition unit 12 acquires a time domain signal as an acoustic signal of the primary sound field (primary sound field signal), it performs a Fourier transform on the acquired acoustic signal to obtain a primary sound field signal p(x, k), which is an acoustic signal of the primary sound field in the frequency domain.
[0103] The signal acquisition unit 12 performs a spatial Fourier transform shown in the following equation (5) based on the primary sound field signal p(x, k) and the positions of the observation points of the primary sound field, i.e., microphone coordinates (coordinate information) indicating the arrangement positions of the microphones of the microphone array 11. In other words, a spherical harmonic expansion is performed.
[0104]
[0105] In equation (5), x indicates the position of the observation point (microphone), and the position x is expressed by coordinates in a polar coordinate system with the origin being the center of the area surrounded by the multiple microphones (observation points) that make up the microphone array 11. For example, the position x is expressed as (r, θ, φ) using a radius r that indicates the distance from the origin to the position x, an azimuth angle φ that indicates the horizontal position of the position x, and a depression angle θ that indicates the vertical position of the position x.
[0106] Also, in formula (5), Y nm (θ,φ) is a spherical harmonic function, and n and m indicate the degree and order of the spherical harmonic domain. nm (r, k) represents the primary sound field signal in the spherical harmonic domain, and can be calculated by the spherical surface integral of the primary sound field signal p(x, k) or by the least squares method.
[0107] The signal acquisition unit 12 acquires the primary sound field signal P nm (r, k) is supplied to the signal correction unit 15 .
[0108] If the condition that there is no sound source within the control area, as explained in "C. Speaker-microphone array and its signal", is satisfied, the primary sound field can be considered to be an internal sound field, and the primary sound field signal P nm (r, k) can be expressed by the following equation (6): nm (r, k) is the sound field coefficient p nm (k) and the spherical Bessel function, i.e., the wave function j n (kr) and can be separated.
[0109]
[0110] Similarly, the transfer function of the second-order sound field gl The spherical harmonic expansion (spatial Fourier transform) of (x, k) is given by the following equation (7). l nm (r,k) denotes the transfer function of the spherical harmonic domain of the second-order sound field.
[0111]
[0112] Transfer function g l By substituting (x, k) into the above equation (4), the acoustic signal in the spherical harmonic domain of the secondary sound field, i.e., the secondary sound field signal p'(x, k), can be expressed as shown in the following equation (8).
[0113]
[0114] In addition, the transfer function G in the spherical harmonic domain l nm Since (r, k) is an internal sound field like the primary sound field, the coefficient g independent of the radius r is used as shown in the following equation (9). l nm (k) and the spherical Bessel function, i.e., the wave function j n (kr) and can be separated.
[0115]
[0116] The transfer function calculation unit 13 calculates a transfer function g from the speaker coordinates of each speaker that constitutes the speaker array 18. l (x, k) and then calculate the transfer function g l From (x, k), the transfer function G of the spherical harmonic domain of the second-order sound field is calculated by equation (7). l nm (r, k) is calculated and supplied to the transfer function correction unit 16 .
[0117] Incidentally, when the wave numbers of the primary sound field and the secondary sound field are the same, sound field control becomes possible by using the mode matching method (see, for example, the following document 3) to match the primary sound field and the secondary sound field at each order n and order m of the spherical harmonic region by truncating the maximum order of the spherical harmonic expansion at N. The mode matching method is described in detail, for example, in "MA Poletti, "Three-dimensional surround sound systems based on spherical harmonics," Journal of the Audio Engineering Society 53(11), 2005: 1004-1025." (hereinafter referred to as document 3).
[0118] For example, when the wave numbers of the primary sound field and the secondary sound field are the same, the drive signal for the speaker can be calculated by minimizing the cost function J shown in the following equation (10).
[0119]
[0120] The cost function J is often minimized by the least squares method.
[0121] Sound field control is possible by matching the primary and secondary sound fields in the spherical harmonic domain because signals in the spherical harmonic domain can be freely interpolated and extrapolated in the internal sound field region.
[0122] For example, if the primary and secondary sound fields are perfectly aligned at the position of radius r, that is, P nm (r, k)=ΣG l nm (r,k)d l In the case where (k), when the secondary sound field is observed at a position of a predetermined radius r'≠r, the following equation (11) holds.
[0123]
[0124] From this equation (11), we can see that when the wave numbers are equal in the primary sound field and the secondary sound field, the secondary sound field signal p'(r',θ,φ) observed at the position of radius r' in the spherical harmonic domain matches the primary sound field signal p(r',θ,φ).
[0125] (E. Correction Processing) In the sound field control system shown in FIG. 2, the correction coefficient calculation unit 14 to the transfer function correction unit 16 perform a new correction processing that has not been performed in the past.
[0126] In the conventional sound field control method described in "D. Sound field control processing", the wave number k of the primary sound field 1 and the wave number k of the secondary sound field 2 It is assumed that and are the same.
[0127] However, the physical space where the primary sound field is observed (recording space) and the physical space where the secondary sound field is formed (reproduction space) each have their own attenuation characteristics. 1 and the attenuation parameter α of the secondary sound field (reproduced space) 2 If different from 1 and wave number k 2 will have different values. 1 ≠α 2 If so, then k1=(1−jα1)k≠(1−jα2)k=k2.
[0128] In such a case, the sound field control by the interpolation and extrapolation mentioned above becomes impossible except at the observation point of radius r. Similarly, when the primary and secondary sound fields completely coincide at the position of radius r, that is, P nm (r, k1)=ΣG l nm (r,k2)d l Assuming that (k2), when observation is performed at a radius r'≠r, the result is as shown in the following equation (12).
[0129]
[0130] In the example of equation (12), the Kirchhoff-Helmholtz integral equation, which is the theoretical background of sound field control, does not hold, so there is no theoretical solution for controlling the entire space. That is, at the position of radius r', the secondary sound field signal p'(r',θ,φ) does not match the primary sound field signal p(r',θ,φ). Therefore, as explained with reference to Figure 1, the reproduction error becomes large (increased).
[0131] Therefore, in the present technology, the correction coefficient calculation unit 14 to the transfer function correction unit 16 calculate the attenuation parameter α of the primary sound field. 1 and the attenuation parameter α of the secondary sound field 2 , i.e., the wave number k of the primary sound field 1 (complex wave number k 1 ) and the wave number k of the secondary sound field 2 (complex wave number k 2 ) and correction processing is performed based on this.
[0132] Here, the following two correction methods will be described as correction processes for the acoustic signals of the primary sound field and the transfer functions of the secondary sound field, that is, as correction methods.
[0133] (Correction Method 1: Adjustment of Control Radius) As mentioned above, the wave number k of the primary sound field 1 and the wave number k of the secondary sound field 2 When the distances are different, there is no theoretical solution for controlling the entire space, but control for a single radius is possible. However, in conventional methods, the radius is determined by the spatial arrangement of the control points (microphones), and the control radius cannot be adjusted.
[0134] In contrast, in the correction method 1 of the present technology, first, the attenuation parameter α of the primary sound field input in advance is calculated. 1 and the attenuation parameter α of the secondary sound field 2 and the complex wave number k of the primary sound field based on the control frequency f and the sound speed c (or the real wave number k). 1 and the complex wave number k of the secondary sound field 2 Here, k1=(1-jα1)k and k2=(1-jα2)k.
[0135] Next, the position of the control point, i.e., the radius r indicating the position of the observation point of the primary sound field, the newly input control radius r', and the calculated wave number k 1 and wave number k 2 The correction coefficient (hereinafter also referred to as the sound field correction coefficient) is calculated using the above. n (k1,r') and the correction coefficient γ for the secondary sound field n (k2,r') is calculated.
[0136] Specifically, for the order n of the spherical harmonic domain, the wave number k 1and the wave function j depends on the control radius r'. n (k1r') (spherical Bessel function) and wave number k 1 and the wave function j depends on the radius r n The ratio of (k1r) to the correction coefficient γ n (k1, r'). That is, the correction coefficient γ n (k1,r')=j n (k1r') / j n (k1r).
[0137] Similarly, the wave number k 2 and the wave function j depends on the control radius r'. n (k2r') and wave number k 2 and the wave function j, which depends on the radius r n The ratio of (k2r) to the correction coefficient γ n (k2, r'). That is, the correction coefficient γ n (k2,r')=j n (k2r') / j n It is said to be (k2r).
[0138] These sound field correction coefficients are wave number correction terms for interpolation or extrapolation. In other words, the sound field correction coefficients are used to realize sound field control that reproduces, by interpolation or extrapolation, a sound field that is the same as the primary sound field as a secondary sound field at a position of an arbitrary control radius r', that is, sound field control that minimizes reproduction errors at the position of the control radius r'. In particular, in correction method 1, a control area is represented by one control radius r', and the control area is defined as a region on the surface of a sphere with the control radius r' as its radius.
[0139] The signal correction unit 15 calculates the correction coefficient γ n Based on (k1,r'), the primary sound field signal P nm (r, k1) is corrected. That is, the correction coefficient γ n (k1,r') is the primary sound field signal P nm (r, k1) is multiplied to obtain the corrected primary sound field signal γ n (k1,r')P nm It is assumed to be (r, k1).
[0140] In addition, the transfer function correction unit 16 calculates the correction coefficient γ nBased on (k2,r'), the transfer function G of the second-order sound field l nm (r, k2) is corrected. That is, the correction coefficient γ n (k2,r') is the transfer function G l nm (r, k2) is multiplied to obtain the corrected secondary sound field transfer function γ n (k2,r')G l nm It is assumed to be (r, k2).
[0141] Finally, the drive signal calculation unit 17 calculates the corrected primary sound field signal γ n (k1,r')P nm (r, k1) and the corrected transfer function γ n (k2,r')G l nm Based on (r, k2), the cost function J shown in the following equation (13) is minimized to obtain the drive signal d of each speaker of the speaker array 18. l (k2) is required.
[0142]
[0143] In equation (13), the driving signal d is used to realize sound field control that minimizes the reproduction error at the position of the control radius r' by interpolation or extrapolation. l (k2) is obtained. The driving signal d l (k2) is the wave number of the sound output, i.e., the wave number k of the secondary sound field. 2 is the wave number k of the primary sound field 1 In the above correction method 1, radius information indicating one control radius r′ that indicates (specifies) the control area is input as information indicating the control area of the controlled sound field (secondary sound field).
[0144] (Correction Method 2: Control by Numerical Minimization of Subspace) Wave number k of the primary sound field 1 and the wave number k of the secondary sound field 2 When the values are different, there is no theoretical solution for controlling the entire space, but it is possible to numerically minimize the control error in a part of the space (reproduction space).
[0145] Specifically, in correction method 2, the sound field correction coefficients explained in correction method 1 are set for multiple control radii r', and the cost function for each control radius r' is minimized simultaneously, thereby making it possible to reduce the control error for all control radii r', i.e., in some regions (spaces) of the reproduction space.
[0146] In this case, the control area is represented by multiple control radii r', and the control area is the region inside a sphere (including the surface of the sphere) whose radius is the largest control radius r' among the multiple control radii r'. More specifically, if the smallest control radius r' is relatively large, the control area is the region inside the sphere whose radius is the largest control radius r' and outside the sphere whose radius is the smallest control radius r'.
[0147] In the correction method 2, as in the correction method 1, first, the attenuation parameter α 1 and the damping parameter α 2 Based on this, the wave number k of the primary sound field 1 and the wave number k of the secondary sound field 2 Then, for a plurality of control radii r', a sound field correction coefficient for each control radius r' is calculated.
[0148] For example, let the control radius r' be r' = [r1', r2', ..., r Q A total of Q control radii r ′] are input. When inputting the control radii, each of the Q control radii r q ' (q=1,2,...,Q) can be directly specified, or the control radius r q ' can take, and each of the Q control radii r q That is, in the correction method 2, the control radius r q ' radius information that indicates the radius itself, and the control radius r q Multiple control radii r indicating the control area, such as range information indicating the range of values that ' can take q The control radius r qWhen the range of ' is input (specified), the correction coefficient calculation unit 14 calculates a plurality of control radii r q ', that is, a plurality of control radii r for calculating the sound field correction coefficient q In this case, some radii within the input range are determined as the control radius r q ' is selected.
[0149] In addition, the signal correction unit 15 calculates the control radius r q The same correction as in the correction method 1 is performed on the primary sound field signal for each radii r q For each ', the same correction as in correction method 1 is performed on the transfer function of the secondary sound field.
[0150] Finally, the drive signal calculation unit 17 calculates all the Q control radii r q ', the cost function J is simultaneously minimized, and the drive signal d l (k2) is calculated. Specifically, each control radius r q The corrected primary sound field signal γ n (k1,r q ')P nm (r, k1) and the corrected transfer function γ n (k2,r q ')G l nm (r, k2), the cost function J shown in the following equation (14) is minimized. In other words, q The driving signal for the speaker is calculated based on the sound field correction coefficients of the two-dimensional sound field and the transfer function of the two-dimensional sound field.
[0151]
[0152] In equation (14), each control radius r q The driving signal d is used to realize sound field control that minimizes reproduction error at the position of l (k2) is required.
[0153] In addition, each control radius r q ' directly or by using the control radius rq When entering the range of ', the adjacent control radii r for the wavelength λ = 1 / k of the control frequency f are q If the interval between ' is extremely small, the amount of processing may become large.
[0154] Therefore, the correction coefficient calculation unit 14 calculates the control radius r based on the control frequency f, i.e., the wavelength λ corresponding to the control frequency f. q ' is thinned out, and the multiple control radii r q The radius set consisting of the final control radius r q In such a case, for example, the adjacent control radii r q The minimum spacing between ' (control radius r q It is conceivable to perform thinning so that the interval between adjacent pixels (interval between adjacent pixels) becomes λ / 4 or λ / 2.
[0155] (F. Explanation of Sound Field Control Processing) The flow of processing by the sound field control system will be explained with reference to Fig. 4. That is, the sound field control processing by the sound field control system will be explained below with reference to the flowchart of Fig. 4.
[0156] In step S11 , the signal acquisition unit 12 acquires an acoustic signal of the primary sound field (primary sound field signal) and supplies it to the signal correction unit 15 .
[0157] For example, when the signal acquisition unit 12 acquires a primary sound field signal in the time domain or frequency domain obtained by sound collection by the microphone array 11, it performs a spherical harmonic expansion similar to the above-mentioned equation (5) and supplies the acoustic signal in the spherical harmonic domain of the primary sound field (primary sound field signal) obtained as a result to the signal correction unit 15. Note that the signal acquisition unit 12 may also acquire the acoustic signal in the spherical harmonic domain of the primary sound field obtained by theoretical value calculation or the like and supply it to the signal correction unit 15 as is.
[0158] In step S12, the transfer function calculation unit 13 calculates the transfer function of the secondary sound field (reproduction space) based on the speaker coordinates of each speaker of the speaker array 18, and supplies the transfer function to the transfer function correction unit 16. For example, in step S12, a spherical harmonic expansion similar to the above-mentioned equation (7) is also performed, and the resulting transfer function of the spherical harmonic domain of the secondary sound field is supplied to the transfer function correction unit 16.
[0159] In step S13, the correction coefficient calculation unit 14 calculates the wave numbers (complex wave numbers) of the primary sound field and the secondary sound field.
[0160] For example, the correction coefficient calculation unit 14 includes an attenuation parameter α 1 and the attenuation parameter α, which indicates the attenuation characteristics of the secondary sound field. 2 and a control frequency f, which is the frequency of the sound to be controlled. Note that a real wave number k may be input instead of the control frequency f.
[0161] The correction coefficient calculation unit 14 calculates the input attenuation parameter α 1 Based on the control frequency f and the known sound speed c, the complex wave number k of the primary sound field is 1 =(1-jα1)k (where k=2πf / c). Similarly, the correction coefficient calculation unit 14 calculates the attenuation parameter α 2 Based on the control frequency f and the known sound speed c, the complex wave number k of the secondary sound field is calculated. 2 Calculate =(1-jα2)k.
[0162] In step S14, the correction coefficient calculation unit 14 calculates correction coefficients for the primary sound field and the secondary sound field based on the wave number calculated in step S13 and the input control radius r', and supplies the correction coefficient for the primary sound field to the signal correction unit 15 and the correction coefficient for the secondary sound field to the transfer function correction unit 16.
[0163] For example, when one control radius r′ is input (specified), that is, when the correction process is performed by the above-described correction method 1, the correction coefficient calculation unit 14 calculates the complex wave number k 1 , the control radius r′, and the correction coefficient γ of the primary sound field based on the radius r n (k1,r')=j n (k1r') / j nCalculate (k1r).
[0164] Here, the radius r is the radius r in the polar coordinates (r, θ, φ) that represent the position x of the microphone that constitutes the microphone array 11, and indicates the distance from the origin of the polar coordinate system to the microphone (control point). The position indicated by the radius r is the observation point (observation position) of the primary sound field.
[0165] Similarly, the correction coefficient calculation unit 14 calculates the complex wave number k 2 , the control radius r′, and the correction coefficient γ of the secondary sound field based on the radius r n (k2,r')=j n (k2r') / j n Calculate (k2r).
[0166] Also, for example, the control radius r q ' directly or by using the control radius r q By specifying the range of ', multiple control radii r q When the control radius r′ is input, that is, when the correction process is performed by the above-described correction method 2, the correction coefficient calculation unit 14 calculates the control radius r q That is, the correction coefficient calculation unit 14 calculates a correction coefficient for each control radius r q ', the correction factor γ n (k1,r q ') and the correction factor γ for the secondary sound field n (k2,r q ') is calculated.
[0167] As described above, the correction coefficient calculation unit 14 calculates the control radius r used to calculate the sound field correction coefficient based on the control frequency f (wavelength λ of the secondary sound field). q In such a case, the correction coefficient calculation unit 14 calculates the control radius r q ', the correction coefficient γ n (k1,r q ') and the correction factor γ n (k2,r q ') is calculated.
[0168] In step S15, the signal correction unit 15 corrects the acoustic signal of the primary sound field (primary sound field signal) supplied from the signal acquisition unit 12 based on the correction coefficient of the primary sound field supplied from the correction coefficient calculation unit 14, and supplies the corrected acoustic signal of the primary sound field to the drive signal calculation unit 17.
[0169] For example, when the correction process is performed by the correction method 1, the signal correction unit 15 calculates the correction coefficient γ n (k1,r') is the primary sound field signal P nm (r, k1) is multiplied to obtain the corrected primary sound field signal γ n (k1,r')P nm Let (r, k1).
[0170] Furthermore, when correction processing is performed using correction method 2, for example, the signal correction unit 15 calculates the control radius r q That is, the signal corrector 15 corrects the primary sound field signal for each of the correction coefficients γ n (k1,r q ') into the primary sound field signal P nm (r, k1) is multiplied to obtain the corrected primary sound field signal γ n (k1,r q ')P nm Let (r, k1).
[0171] In step S16, the transfer function correction unit 16 corrects the transfer function of the secondary sound field supplied from the transfer function calculation unit 13 based on the correction coefficient of the secondary sound field supplied from the correction coefficient calculation unit 14, and supplies the corrected transfer function to the drive signal calculation unit 17.
[0172] For example, when the correction process is performed by the correction method 1, the transfer function correction unit 16 calculates the correction coefficient γ n (k2,r') is the transfer function G of the second-order sound field l nm (r, k2) and the corrected secondary sound field transfer function γ n (k2,r')G l nm Let (r, k2).
[0173] Furthermore, when correction processing is performed using correction method 2, for example, the transfer function correction unit 16 calculates the control radius r q That is, the transfer function correction unit 16 corrects the transfer function for each correction coefficient γn (k2,r q ') into the transfer function G l nm (r, k2) and the corrected secondary sound field transfer function γ n (k2,r q ')G l nm Let (r, k2).
[0174] In step S17, the drive signal calculation unit 17 calculates the drive signal for each speaker of the speaker array 18 by interpolation or extrapolation based on the corrected primary sound field signal supplied from the signal correction unit 15 and the corrected secondary sound field transfer function supplied from the transfer function correction unit 16.
[0175] For example, when the correction process is performed using the correction method 1, the drive signal calculation unit 17 calculates the corrected primary sound field signal γ n (k1,r')P nm (r, k1) and the corrected transfer function γ n (k2,r')G l nm (r, k2) and the cost function J of Equation (13) is minimized to obtain the driving signal d l Calculate (k2).
[0176] Furthermore, when correction processing is performed using correction method 2, for example, the drive signal calculation unit 17 calculates the control radius r q The primary sound field signal γ after correction of n (k1,r q ')P nm (r, k1) and each control radius r q The transfer function γ of the second-order sound field after the correction of n (k2,r q ')G l nm (r, k2) and the cost function J of Equation (14) is minimized to obtain the driving signal d l Calculate (k2).
[0177] In step S18, the drive signal calculation unit 17 outputs the drive signal obtained in step S17 to each speaker of the speaker array 18. In this case, the drive signal calculation unit 17 appropriately converts the drive signal obtained in step S17 into a time domain drive signal before outputting it.
[0178] The speakers of the speaker array 18 output sounds based on the supplied drive signals, thereby forming a secondary sound field in the reproduction space. Once the secondary sound field is formed, the sound field control process ends.
[0179] In this way, the sound field control system calculates a sound field correction coefficient based on the attenuation parameters of the primary and secondary sound fields, the control frequency, and the control radius, and corrects the transfer functions of the primary sound field signal and the secondary sound field based on the sound field correction coefficient to calculate a drive signal.
[0180] By doing this, even if the attenuation characteristics of the recording space and the reproduction space are different, the control error in the control area indicated by the control radius can be reduced, and sound field control can be performed with higher precision.
[0181] (G. Effects) According to this technology, a new sound field control method (space control method) is realized by introducing a wave number correction term called a sound field correction coefficient separately for the primary sound field and the secondary sound field. This makes it possible to reproduce a sound field with higher accuracy in a specific control area even if the attenuation parameters of the target sound field (primary sound field) and the sound field to be controlled (secondary sound field) are different.
[0182] As an example of application of this technology, it is possible to reproduce the sound field of an open space, such as an anechoic chamber or outdoors, in a closed space such as a car cabin or a general room, and conversely, it is also possible to reproduce the sound field of a closed space in an open space. In these cases, highly immersive spatial audio reproduction with 6DoF (Degrees of Freedom) becomes possible, which was not possible with conventional technology.
[0183] Also, consider the case where, for example, when correction processing is performed using correction method 1, minimization is performed so that the cost function J in equation (13) becomes 0, that is, the case where the following equation (15) holds. In this case, at a specified control radius r', the secondary sound field signal p'(r', θ, φ) becomes as shown in the following equation (16).
[0184]
[0185]
[0186] In equation (16), at the position of the control radius r', the secondary sound field signal p'(r', θ, φ) matches the primary sound field signal p(r', θ, φ), and it can be seen that the primary sound field is accurately reproduced in the reproduction space.
[0187] Figure 5 shows the spatial control error (RMSE) of the two-dimensional horizontal plane including the center of the sound field in three-dimensional space when the control radius r' is set to r' = 0.2 m, r' = 0.4 m, and r' = 0.6 m, respectively.
[0188] The example in Figure 5 shows the reproduction error (spatial control error) when the microphone array 11 is a spherical microphone array with 32 channels (M = 32) and the radius of the spherical microphone array, i.e., the radius r indicating the position of the control point, is 0.3 m. In Figure 5, the crosses "+" indicate the positions of the speakers in the speaker array 18, and the dots indicate the positions of the microphones in the microphone array 11.
[0189] In Figure 5, the reproduction error when the control radius r' = 0.2 m is shown on the left side of the figure, the reproduction error when the control radius r' = 0.4 m is shown in the center of the figure, and the reproduction error when the control radius r' = 0.6 m is shown on the right side of the figure.
[0190] In these examples, it can be seen that the area with small reproduction error (control area) is spherical, and the radius of the sphere matches the specified control radius r'.
[0191] Therefore, correction method 1 makes it possible to adjust the control radius r' that is different from the radius r of the microphone array 11 placed in the space (recording space) during sound collection.
[0192] This makes it possible to reproduce a sound field in an appropriate control radius r', that is, a control area, depending on the number of listeners listening to sound in the reproduction space, as shown in FIG.
[0193] For example, as shown by arrow Q41 in Figure 6, in the conventional sound field control method, if the wave numbers (attenuation characteristics) of the primary sound field and the secondary sound field do not match, sound field control can only be performed around the position of radius r of the microphone array 11, i.e., around the control point. Therefore, if the radius r is inappropriate, for example, when the listener is not near the control point, the reproduction error at the listener's position will be large.
[0194] In contrast, with correction method 1 of the present technology, even if the radius r of the microphone array 11 is not appropriate, it is possible to reduce the reproduction error at the listener's position, as in the example of the portion indicated by arrow Q42 and the example of the portion indicated by arrow Q43. In other words, it is possible to reproduce the sound field near the listener's ear with high accuracy.
[0195] In the example indicated by arrow Q42, the region W11 inside the microphone array 11 is the control area (the area subject to control of sound pressure, etc.) determined by the control radius r'. In this example, the control radius r' is reduced for one listener, improving the control accuracy near the listener's ear.
[0196] In the example indicated by arrow Q43, the control area defined by control radius r' is the area W12 outside the microphone array 11. In this example, the control radius r' is expanded for multiple listeners so that the heads of all listeners are included (covered) in the control area.
[0197] Furthermore, whereas conventional techniques only allowed control on a spherical surface of a single radius, correction method 2 makes it possible to minimize the control error (reproduction error) over the entire specified control space (control area) by numerically minimizing the control error (reproduction error).
[0198] Figure 7 shows the spatial control error (RMSE) of the two-dimensional horizontal plane including the center of the sound field in three-dimensional space when the control radius r' is set to r' = [0.1, 0.2, 0.3] and r' = [0.2, 0.4, 0.6] (unit: m).
[0199] The example in Fig. 7 shows the reproduction error (spatial control error) when the microphone array 11 is a spherical microphone array with 32 channels (M = 32) and the radius of the spherical microphone array, i.e., the radius r indicating the position of the control point, is 0.3 m. In Fig. 7, the crosses "+" indicate the positions of the speakers in the speaker array 18, and the dots indicate the positions of the microphones in the microphone array 11.
[0200] In FIG. 7, the left side of the figure shows the reproduction error when the control radius r′=[0.1, 0.2, 0.3], and the right side of the figure shows the reproduction error when the control radius r′=[0.2, 0.4, 0.6].
[0201] In these examples, the reproduction error is larger than the conventional theoretical solution shown in Figure 1, but the area with small reproduction error (control area) is spherical (including the internal volume), and it can be seen that the reproduction error is kept below -10 dB for the entire space of the set control area.
[0202] Therefore, correction method 2 can secure a wider control area than the conventional sound field control methods and correction method 1, as shown in Figure 8, for example, and allows more listeners to simultaneously experience the secondary sound field.
[0203] An example of sound field control using correction method 1 is shown on the left side of Fig. 8. In this example, the area W31 inside the microphone array 11 is the control area determined by the control radius r'. Therefore, each of the multiple listeners must listen with their head positioned within area W31.
[0204] Furthermore, with conventional sound field control methods, if the wave numbers (attenuation characteristics) of the primary sound field and the secondary sound field do not match, sound field control can only be performed around the position (control point) of radius r of the microphone array 11.
[0205] In contrast, in correction method 2 of the present technology, as shown on the right side of the figure, a spherical region determined by multiple control radii r' can be used as the control area. Here, a spherical region W32 including the surface and interior of the sphere is used as the control area, and it can be seen that this control area is wider than the control area in correction method 1. Therefore, correction method 2 allows more listeners to simultaneously listen to the secondary sound field.
[0206] <Modifications> (Processing of Time Domain Signals) The above has been described as a case where correction processing is performed in the frequency domain, i.e., a case where correction processing is performed on a signal obtained by performing a spatial Fourier transform (spherical harmonic expansion) on a frequency domain signal. However, the present invention is not limited to this, and processing can also be performed in the time domain or other domains as long as the cost function can be minimized.
[0207] For example, when performing correction processing in the time domain, in correction method 1, it is necessary to take into consideration the cost function J shown in the following equation (17): In equation (17), * indicates a convolution operation.
[0208]
[0209] In addition, in the formula (17), P nm (t,r,k1) is the time waveform of the primary sound field, that is, the primary sound field signal obtained by performing a spatial Fourier transform (spherical harmonic expansion) on the time domain acoustic signal, and t indicates the time index. l nm (t,r,k2) are the transfer functions of the second-order sound field obtained by performing a spatial Fourier transform on the impulse response of the second-order sound field. These spatial Fourier transforms can be calculated by time-domain convolution of a filter bank.
[0210] In equation (17), γ n (t,k1,r') and γ n (t, k2, r') is the correction coefficient γ for each frequency (control frequency) n (k1,r') and correction coefficient γ n After (k2, r') is obtained, a time domain filter that can be calculated by inverse Fourier transform is shown.
[0211] In this example, a time domain driving signal d for each speaker of the speaker array 18 is calculated by a method such as the conjugate gradient method so that the cost function J is minimized. l (t, k2) can be calculated.
[0212] (Spatial active noise control device) In this technology, by making the primary sound field an antiphase field of the noise and performing sound field control in the same manner as in the example described above, the sound field control system shown in Fig. 2 can also function as a signal processing device (signal processing system) that performs spatial active noise control. In the case of noise control, performance improvement can be expected by using the above-mentioned time domain signal processing technique.
[0213] (Circular Harmonic Expansion) In the above, an example has been described in which analysis and experiments were carried out on a three-dimensional space based on spherical harmonic expansion for a three-dimensional sound field.
[0214] However, since the acoustic field in a two-dimensional space has properties very similar to those in the case described above, by replacing the spherical harmonic expansion with a circular harmonic expansion (also called a cylindrical harmonic expansion) and the spherical Bessel function with a Bessel function, the correction processes by the above-mentioned correction method 1 and correction method 2 can also be applied to a two-dimensional space. That is, the drive signal calculation unit 17 can calculate the drive signal for the secondary sound field based on the primary sound field signal in the circular harmonic domain and the transfer function of the circular harmonic domain of the secondary sound field.
[0215] (Moving and Deforming the Control Area Using the Addition Theorem) The above has been described as an example of controlling a sound field on a concentric sphere when the control radius r' is adjusted with the center of the control point (observation point) as the origin of the coordinate system. However, because the addition theorem of the spherical harmonic domain (circular harmonic domain) allows for the transformation (movement) of the coordinate center, when this addition theorem is used, it is also possible to specify, as the control area, a spherical surface or a sphere whose center position is shifted from the center of the control point.
[0216] (Mismatch of wave numbers due to influences other than attenuation parameters) In principle, the correction method (correction process) of the present technology can be applied not only to mismatch of wave numbers due to different attenuation parameters between the primary sound field and the secondary sound field, but also to mismatch of frequency and sound speed. In other words, the sound field control system of the present technology can accurately control the sound field even if at least one of the attenuation characteristics, sound frequency, and sound speed differs between the primary sound field (desired sound field) and the secondary sound field (control sound field). If the frequency and sound speed, i.e., the real wave numbers, differ between the primary sound field and the secondary sound field, the complex wave numbers will also differ (mismatch).
[0217] For example, when a frequency mismatch occurs, if noise is present in a specific frequency component during recording of the primary sound field and the recorded data for that frequency component cannot be used, then the recorded data for that frequency component in the secondary sound field can be used to reproduce that frequency component after correction. Also, for example, a mismatch in sound speed can occur when reproducing underwater sounds in an ordinary room, or when dealing with changes in the speed of sound due to temperature and humidity.
[0218] According to the present technology as described above, for example, in sound field control in a closed space having attenuation characteristics, the accuracy of sound field control can be improved by correcting the discrepancy in attenuation parameters due to the space using the sound field correction coefficient.
[0219] Furthermore, for example, in this technology, by introducing complex wave numbers into sound field control, it is possible to realize sound field control that takes into account not only spatial attenuation parameters but also attenuation parameters due to the viscosity of the propagation medium, such as air attenuation.
[0220] Furthermore, with this technology, even when the wave numbers of the desired sound field (primary sound field) and the control sound field (secondary sound field) are different, the discrepancy between these wave numbers can be corrected to improve the accuracy of sound field control. In other words, sound field control can be performed even when the frequencies and sound velocities of the primary sound field and the secondary sound field are different.
[0221] In addition, while in conventional sound field control methods the control area is determined by the placement position of the control points, with this technology it is possible to adjust the position and size of the control area by inputting information indicating the control area, such as the control radius.
[0222] <Example of Computer Configuration> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs constituting the software are installed on a computer. Here, the computer includes a computer built into dedicated hardware, and a general-purpose personal computer, for example, that can execute various functions by installing various programs.
[0223] FIG. 9 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.
[0224] In the computer, a CPU (Central Processing Unit) 501 , a ROM (Read Only Memory) 502 , and a RAM (Random Access Memory) 503 are interconnected by a bus 504 .
[0225] An input / output interface 505 is further connected to the bus 504. An input unit 506, an output unit 507, a recording unit 508, a communication unit 509, and a drive 510 are connected to the input / output interface 505.
[0226] The input unit 506 includes a keyboard, a mouse, a microphone array, an image sensor, etc. The output unit 507 includes a display, a speaker array, etc. The recording unit 508 includes a hard disk, a non-volatile memory, etc. The communication unit 509 includes a network interface, etc. The drive 510 drives a removable recording medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0227] In a computer configured as described above, the CPU 501 loads, for example, a program recorded in the recording unit 508 into the RAM 503 via the input / output interface 505 and the bus 504, and executes the program, thereby performing the above-described series of processes.
[0228] The program executed by the computer (CPU 501) can be provided by being recorded on a removable recording medium 511 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.
[0229] In a computer, a program can be installed in the recording unit 508 via the input / output interface 505 by inserting a removable recording medium 511 into the drive 510. The program can also be received by the communication unit 509 via a wired or wireless transmission medium and installed in the recording unit 508. Alternatively, the program can be installed in the ROM 502 or the recording unit 508 in advance.
[0230] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.
[0231] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.
[0232] For example, the present technology can be configured as a cloud computing system in which a single function is shared and processed collaboratively by a plurality of devices via a network.
[0233] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.
[0234] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0235] Furthermore, the present technology can also be configured as follows.
[0236] (1) An information processing device comprising: a correction coefficient calculation unit that calculates a correction coefficient based on a first attenuation parameter indicating the attenuation characteristics of a desired sound field to be reproduced, a second attenuation parameter indicating the attenuation characteristics of a controlled sound field, and information indicating a control area of the controlled sound field, and a drive signal calculation unit that calculates a drive signal for a speaker to form the controlled sound field based on an acoustic signal of the desired sound field, the correction coefficient, and a transfer function indicating the transfer characteristics of the controlled sound field. (2) The information processing device described in (1), wherein the correction coefficient calculation unit calculates the correction coefficient based on a first complex wave number of the desired sound field based on the first attenuation parameter, a second complex wave number of the control sound field based on the second attenuation parameter, and information indicating the control area. (3) The information processing device according to (2), wherein the correction coefficient calculation unit calculates a first correction coefficient as the correction coefficient based on the first complex wave number and information indicating the control area, and calculates a second correction coefficient as the correction coefficient based on the second complex wave number and information indicating the control area, and the drive signal calculation unit calculates the drive signal based on the acoustic signal of the desired sound field corrected by the first correction coefficient and the transfer function corrected by the second correction coefficient. (4) The information processing device according to (3), wherein the first correction coefficient is a ratio of a wave function that depends on the first complex wave number and information indicating the control area to a wave function that depends on the first complex wave number and the position of an observation point of the desired sound field, and the second correction coefficient is a ratio of a wave function that depends on the second complex wave number and information indicating the control area to a wave function that depends on the second complex wave number and the position of the observation point. (5) The information processing device according to any one of (1) to (4), wherein the desired sound field and the controlled sound field are different in at least one of attenuation characteristics, frequency, and sound speed. (6) The information processing device according to any one of (1) to (5), wherein the information indicating the control area of the controlled sound field is radius information indicating one radius that specifies the control area.(7) The information processing device described in any one of (1) to (5), wherein the information indicating the control area of the controlled sound field is information capable of identifying multiple radii indicating the control area, the correction coefficient calculation unit calculates the correction coefficient for each radius based on the first attenuation parameter, the second attenuation parameter, and the radius, and the drive signal calculation unit calculates the drive signal based on an acoustic signal of the desired sound field, the correction coefficient for each of the multiple radii, and the transfer function. (8) The information processing device described in (7), wherein the information indicating the control area of the controlled sound field is range information indicating a range of the radii, and the correction coefficient calculation unit determines the multiple radii used in calculating the correction coefficient based on the range information. (9) The information processing device described in (7) or (8), wherein the correction coefficient calculation unit thins out the radii used in calculating the correction coefficient based on a wavelength of the controlled sound field. (10) The information processing device according to any one of (1) to (9), wherein the drive signal calculation unit calculates the drive signal by interpolation or extrapolation. (11) The information processing device according to any one of (1) to (10), wherein the drive signal calculation unit calculates the drive signal based on an acoustic signal of the desired sound field in a spherical harmonic domain or a circular harmonic domain, the correction coefficient, and the transfer function of the spherical harmonic domain or the circular harmonic domain. (12) The information processing device according to any one of (1) to (11), wherein one of the space in which the desired sound field is formed and the space in which the control sound field is formed is a closed space, and the other is an open space. (13) An information processing method in which an information processing device calculates a correction coefficient based on a first attenuation parameter indicating the attenuation characteristics of a desired sound field to be reproduced, a second attenuation parameter indicating the attenuation characteristics of a controlled sound field, and information indicating a control area of the controlled sound field, and calculates a drive signal for a speaker to form the controlled sound field based on the acoustic signal of the desired sound field, the correction coefficient, and a transfer function indicating the transfer characteristics of the controlled sound field.(14) A program that causes a computer to execute a process including the steps of: calculating a correction coefficient based on a first attenuation parameter indicating the attenuation characteristics of a desired sound field to be reproduced, a second attenuation parameter indicating the attenuation characteristics of a controlled sound field, and information indicating a control area of the controlled sound field; and calculating a drive signal for a speaker to form the controlled sound field based on the acoustic signal of the desired sound field, the correction coefficient, and a transfer function indicating the transfer characteristics of the controlled sound field.
[0237] REFERENCE SIGNS LIST 11 microphone array, 12 signal acquisition unit, 13 transfer function calculation unit, 14 correction coefficient calculation unit, 15 signal correction unit, 16 transfer function correction unit, 17 drive signal calculation unit, 18 speaker array
Claims
1. An information processing device comprising: a correction coefficient calculation unit that calculates a correction coefficient based on a first attenuation parameter that indicates the attenuation characteristics of a desired sound field to be reproduced, a second attenuation parameter that indicates the attenuation characteristics of a controlled sound field, and information that indicates a control area of the controlled sound field; and a drive signal calculation unit that calculates a drive signal for a speaker to form the controlled sound field based on the acoustic signal of the desired sound field, the correction coefficient, and a transfer function that indicates the transfer characteristics of the controlled sound field.
2. The information processing device according to claim 1, wherein the correction coefficient calculation unit calculates the correction coefficient based on a first complex wave number of the desired sound field based on the first attenuation parameter, a second complex wave number of the control sound field based on the second attenuation parameter, and information indicating the control area.
3. The information processing device described in claim 2, wherein the correction coefficient calculation unit calculates a first correction coefficient as the correction coefficient based on the first complex wave number and information indicating the control area, and calculates a second correction coefficient as the correction coefficient based on the second complex wave number and information indicating the control area, and the drive signal calculation unit calculates the drive signal based on the acoustic signal of the desired sound field corrected by the first correction coefficient and the transfer function corrected by the second correction coefficient.
4. The information processing device of claim 3, wherein the first correction coefficient is a ratio of a wave function that depends on the first complex wave number and information indicating the control area to a wave function that depends on the first complex wave number and the position of the observation point of the desired sound field, and the second correction coefficient is a ratio of a wave function that depends on the second complex wave number and information indicating the control area to a wave function that depends on the second complex wave number and the position of the observation point.
5. The information processing device according to claim 1, wherein the desired sound field and the controlled sound field differ from each other in at least one of attenuation characteristics, frequency, and sound speed.
6. The information processing device according to claim 1, wherein the information indicating the control area of the controlled sound field is radius information indicating one radius that identifies the control area.
7. The information processing device of claim 1, wherein the information indicating the control area of the controlled sound field is information capable of identifying multiple radii indicating the control area, the correction coefficient calculation unit calculates the correction coefficient for each radius based on the first attenuation parameter, the second attenuation parameter, and the radius, and the drive signal calculation unit calculates the drive signal based on the acoustic signal of the desired sound field, the correction coefficient for each of the multiple radii, and the transfer function.
8. The information processing device according to claim 7, wherein the information indicating the control area of the controlled sound field is range information indicating the range of the radii, and the correction coefficient calculation unit determines the multiple radii to be used in calculating the correction coefficient based on the range information.
9. The information processing device according to claim 7, wherein the correction coefficient calculation unit thins out the radii used to calculate the correction coefficients based on the wavelength of the control sound field.
10. The information processing device according to claim 1, wherein the drive signal calculation unit calculates the drive signal by interpolation or extrapolation.
11. The information processing device according to claim 1, wherein the drive signal calculation unit calculates the drive signal based on an acoustic signal of the desired sound field in a spherical harmonic domain or a circular harmonic domain, the correction coefficient, and the transfer function in the spherical harmonic domain or the circular harmonic domain.
12. The information processing device according to claim 1, wherein one of the space in which the desired sound field is formed and the space in which the control sound field is formed is a closed space, and the other is an open space.
13. An information processing method in which an information processing device calculates a correction coefficient based on a first attenuation parameter indicating the attenuation characteristics of a desired sound field to be reproduced, a second attenuation parameter indicating the attenuation characteristics of a controlled sound field, and information indicating a control area of the controlled sound field, and calculates a drive signal for a speaker to form the controlled sound field based on the acoustic signal of the desired sound field, the correction coefficient, and a transfer function indicating the transfer characteristics of the controlled sound field.
14. A program that causes a computer to execute processing including the steps of: calculating a correction coefficient based on a first attenuation parameter that indicates the attenuation characteristics of a desired sound field to be reproduced, a second attenuation parameter that indicates the attenuation characteristics of a controlled sound field, and information that indicates the control area of the controlled sound field; and calculating a drive signal for a speaker to form the controlled sound field based on the acoustic signal of the desired sound field, the correction coefficient, and a transfer function that indicates the transfer characteristics of the controlled sound field.
Citation Information
Patent Citations
Sound field control system and method
JP2005215250A
Signal processing device, method, and program
WO2021251182A1