Signal processing system, signal processing method, and microphone array

The modified multi-ring microphone array with a two-stage mixing method effectively reduces computational complexity, enabling real-time high-spatial-resolution stereophonic sound processing, suitable for VR and AR applications.

WO2025150204A1PCT designated stage expired Publication Date: 2025-07-17KANEKO SHOKEN ECHHART
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/000698
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-12
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently processing multi-channel audio signals from microphone arrays due to high computational costs, particularly in achieving real-time signal processing for high-quality stereophonic sound with high spatial resolution, which is essential for applications like VR and AR, where interactivity and personalization are crucial.

Method used

A signal processing system utilizing a microphone array with a modified multi-ring configuration and a two-stage mixing method, including local and local-to-global mixing processes, reduces computational complexity from O(N^2) to O(N^1.5) by employing matrix decomposition and complementary angular frequency components, allowing for efficient calculation of expansion coefficients.

Benefits of technology

This approach enables real-time processing of high-spatial-resolution stereophonic sound, addressing issues of computational cost and spatial aliasing, and facilitates interactivity and personalization in applications like VR and AR.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024000698_17072025_PF_FP_ABST
    Figure JP2024000698_17072025_PF_FP_ABST
Patent Text Reader

Abstract

A signal processing system according to one embodiment of the present invention includes: a reception means for receiving input of a sound signal of N channels outputted from a microphone array that includes a plurality of microphone elements having a predetermined arrangement; a calculation means for calculating, by an algorithm with less calculation cost than the order NM, M expansion coefficients of a basis function for expressing a three-dimensional sound field corresponding to the product of a conversion matrix determined according to arrangement of the plurality of microphone elements and the vector of the sound signal of the plurality of channels; and an output means for outputting information on the expansion coefficients.
Need to check novelty before this filing date? Find Prior Art

Description

Signal processing system, signal processing method, and microphone array

[0001] The present invention relates to a system for providing stereophonic sound.

[0002] Techniques related to spatial sound are known. For example, Patent Document 1 discloses a system for separating sound sources using spatial frequency masks to reduce computational costs when separating sound sources from multi-channel audio signals obtained by a microphone array. Patent Document 2 discloses a technology for determining spherical harmonic coefficients of an audio signal based on a captured signal and a spherical harmonic transfer function, and generating an audio signal based on the determined spherical harmonic coefficients. Patent Document 3 discloses a system for compressing and decoding Higher Order Ambisonics (HOA) signals. Patent Document 4 discloses a method for recording and reconstructing a three-dimensional sound field. In this method, a microphone array is placed in a three-dimensional sound field, a sound source within the sound field is tracked and identified, and a corresponding sound source signal is obtained. A plurality of control points are set within the three-dimensional sound field to be reconstructed, and these control points are used to establish relationships between the sound source signal, the three-dimensional sound field, the reconstructed sound field, and the reconstructed sound source signal.

[0003] Patent No. 6807029 U.S. Patent No. 11218807 U.S. Patent No. 10176814 U.S. Patent No. 9510098

[0004] The above-mentioned conventional techniques have had problems when expressing stereophonic sound from multi-channel audio signals picked up by a microphone array. One of the problems that the present invention aims to solve is to improve the processing of multi-channel audio signals picked up by a microphone array.

[0005] One aspect of the present disclosure provides a signal processing system having: a receiving means for receiving, as input audio signals, N-channel audio signals output from a microphone array including a plurality of microphone elements having a predetermined arrangement; a calculating means for calculating M expansion coefficients of basis functions for expressing a three-dimensional sound field, which corresponds to the product of a transformation matrix determined according to the arrangement of the plurality of microphone elements and a vector of the input audio signals, using an algorithm with a lower computational cost than order N-M; and an output means for outputting information regarding the expansion coefficients.

[0006] Another aspect of the present disclosure provides a signal processing system having: a receiving means for receiving, as input audio signals, multi-channel audio signals output from a microphone array including a plurality of microphone elements having a predetermined arrangement; a calculating means for calculating, by a method based on matrix decomposition of a transformation matrix determined according to the arrangement of the plurality of microphone elements, expansion coefficients of basis functions for expressing a three-dimensional sound field, the expansion coefficients corresponding to the product of the vector of the input audio signals and a transformation matrix determined according to the arrangement of the plurality of microphone elements; and an output means for outputting information relating to the expansion coefficients.

[0007] The signal processing system may further comprise a generating means for generating a stereophonic signal for expressing a stereophonic field from the expansion coefficients, and the output means may output the recording signal as information relating to the expansion coefficients.

[0008] The determined arrangement may be an arrangement along a plurality of virtual rings, and the calculation means may include first processing means for performing a first process of obtaining an angular Fourier spectrum for each of the plurality of rings from the audio signals of the multiple channels by mixing audio signals within the ring independently of other rings; interpolation means for interpolating data on angular frequency components that are missing in the angular Fourier spectra corresponding to the plurality of rings compared to the spectrum with the largest number of angular frequency components obtained, so that the number of angular frequency components is the same for all rings; and second processing means for performing a second process of obtaining the expansion coefficients by mixing the angular Fourier spectra of the plurality of rings for each angular frequency component on the interpolated plurality of angular Fourier spectra.

[0009] The microphone elements may be equally spaced in the rings, and the first processing may include a dense matrix vector product or a Fourier type transform.

[0010] The microphone elements may be non-equidistantly spaced in the rings, and the first processing may include a dense matrix vector product, a non-uniform fast Fourier transform, or a matrix decomposition method.

[0011] The second processing may include a plurality of dense matrix-vector products or matrix-vector products by matrix decomposition.

[0012] The multiple dense matrix vector products may be calculated for each of multiple angular frequency components on the angular Fourier spectrum of the multiple rings independently of other angular frequency components.

[0013] The matrix decomposition may involve a process of singular value decomposition or multi-stage decomposition of a transformation matrix.

[0014] The calculation means may include calculations for obtaining the expansion coefficients by a matrix decomposition method.

[0015] The matrix decomposition method may include a process of multi-stage decomposition and multi-stage application of a transformation matrix.

[0016] The microphone array may be included.

[0017] The plurality of microphone elements may be arranged along the surface of a virtual or physical sphere or spheroid.

[0018] The rings may be parallel on the surface of the sphere or spheroid.

[0019] In each of the plurality of rings, a plurality of microphone elements may be arranged at equal intervals.

[0020] The plurality of rings may include two adjacent rings having the same number of microphone elements.

[0021] The basis functions may be spherical harmonics and the expansion coefficients may be used to calculate an Ambisonic signal.

[0022] Another aspect of the present disclosure provides a signal processing method including the steps of: receiving, as input audio signals, N-channel audio signals output from a microphone array including a plurality of microphone elements having a predetermined arrangement; calculating M expansion coefficients of basis functions for expressing a three-dimensional sound field, which correspond to the product of a transformation matrix determined according to the arrangement of the plurality of microphone elements and a vector of the N-channel audio signals, using an algorithm with a lower computational cost than order N-M; and outputting information related to the expansion coefficients.

[0023] Another aspect of the present disclosure provides a signal processing method including the steps of: receiving, as input audio signals, multi-channel audio signals output from a microphone array including a plurality of microphone elements having a predetermined arrangement; calculating expansion coefficients of basis functions for expressing a three-dimensional sound field, which correspond to the product of a transformation matrix determined according to the arrangement of the plurality of microphone elements and a vector of the input audio signals, using a method based on matrix decomposition of the transformation matrix; and outputting information related to the expansion coefficients.

[0024] Another aspect of the present disclosure provides a microphone array having a plurality of microphone elements arranged along a plurality of virtual rings formed on the surface of a virtual or physical sphere or spheroid, the plurality of rings being parallel to isolatitudinal lines in a defined coordinate system.

[0025] In each of the plurality of rings, a plurality of microphone elements may be arranged at equal intervals.

[0026] The plurality of rings may include two adjacent rings having the same number of microphone elements.

[0027] Among the multiple rings, the number of microphone elements belonging to rings included in a first range where the latitude θ in the coordinate system is θ < θth1 or θth2 < θ may be smaller than the number of microphone elements belonging to a second range where θth1≦θ≦θth2.

[0028] The number of microphone elements belonging to each of the plurality of rings included in the second range may all be equal.

[0029] The set of longitudes of the microphone elements belonging to each of the plurality of rings included in the second range may be the same for adjacent rings.

[0030] According to the present invention, it is possible to improve the processing of multi-channel audio signals picked up by a microphone array.

[0031] 1 is a diagram illustrating the shape of a spherical harmonic function. 2 is a diagram schematically illustrating an overview of Ambisonics. 3 is a diagram illustrating an arrangement of measurement points on a spherical surface. 4 is a diagram illustrating another example of an arrangement of measurement points on a spherical surface. 5 is a diagram illustrating the functional configuration of a stereophonic sound system 1 according to an embodiment. 6 is a diagram illustrating an arrangement of microphone elements in a microphone array 10. 7 is a diagram illustrating the hardware configuration of a signal processing device 20. 8 is a flowchart illustrating the operation of the signal processing device 20. 9 is a schematic diagram illustrating audio signals acquired from the microphone array 10. 10 is a schematic diagram illustrating the result of local mixing processing. 11 is a diagram illustrating an angular Fourier spectrum after interpolation processing. 12 is a schematic diagram of local-global mixing processing. 13 is a diagram illustrating an overview of a microphone array 10 according to a modified example.

[0032] 1. Overview Spatial sound is important in technologies such as VR (Virtual Reality) and AR (Augmented Reality). For example, hearing sounds from a 360-degree space can immerse users more deeply in VR / AR. Alternatively, changing sounds according to the user's movements or actions can provide users with a more dynamic and interactive experience. A spatial sound system includes two elements: signal processing and a microphone array.

[0033] 1-1. Signal Processing Signal processing in 3D sound is the process of calculating a signal that represents 3D sound (hereinafter referred to as "3D sound signal") from multi-channel audio signals picked up by a microphone array. A 3D sound signal is typically a multi-channel signal consisting of weights of a set of basis functions related to direction or space. This signal processing usually involves mixing the signals of all input channels for all output channels. This is mathematically formulated as a dense matrix vector product (DMVP). When the number of input channels (i.e., the number of microphone elements) is N and the number of output channels is M, the computational cost is on the order of N × M (hereinafter referred to as O(NM)). Since it is reasonable to set M to a number on the same order as N, assuming O(M) = O(N), the computational cost is O(N 2 ) For several reasons, it is desirable to have a large number of microphone elements N in a high-quality 3D sound system, but increasing N results in O(N 2 ) is extremely large, which means that the calculation cost becomes extremely high. Therefore, when the number of microphone elements N becomes large to a certain extent, real-time signal processing becomes impossible.

[0034] This signal processing is a weighted mixing of multi-channel signals. Mathematically, if the input multi-channel signal is x, the output multi-channel signal is y, and the mixing matrix is ​​A, then This A is specifically called the "Global Mixing Matrix (GMM)." ​​This signal processing is also called "global mixing processing." Note that the signals, signal vectors, and matrix elements referred to in this paper can all be real or complex numbers.

[0035] Ambisonics is known as a technology for recording and reproducing a three-dimensional sound field. Ambisonics involves two processes: encoding and decoding. Encoding is a process of processing signals obtained from a microphone to obtain expansion coefficients. Decoding is a process of processing the expansion coefficients to obtain audio signals for outputting sound from speakers. First, encoding will be explained in more detail.

[0036] In Ambisonics, a three-dimensional sound field (where sound pressure is determined as a function of position) is expressed by the following function: where p(r) is the sound field at position r, are the expansion coefficients, is the spherical Bessel function, is a spherical harmonic function. n is the order and m is the degree. The product of a spherical Bessel function and a spherical harmonic function can be called a spherical (spherical wave) basis function. In other words, the sound field at any position can be expressed by the expansion of the spherical wave basis function.

[0037] In Ambisonics, the process of estimating the expansion coefficients of a sound field from a microphone signal is called encoding. In other words, encoding in Ambisonics refers to the process of converting an audio signal into coefficients. These expansion coefficients are sometimes called Ambisonics coefficients, Ambisonics signals, or B signals.

[0038] In equation (2), mathematically, if Nc = ∞, the original sound field can be perfectly reproduced, but in real products, the calculation is terminated at a finite Nc. The simplest example is when Nc = 1. This is called first-order Ambisonics or B-format. First-order Ambisonics can record and output four-channel audio signals. First-order Ambisonics is in practical use, and microphone arrays compatible with first-order Ambisonics are in practical use.

[0039] Figure 1 is a diagram illustrating the shape of a spherical harmonic function. When n = 0, the signal is omnidirectional, and when n = 1, the directivity extends in three directions (x, y, and z). When n = 2, the directivity spreads out in even more directions. In other words, the larger the order n, the better the directional resolution, or spatial resolution.

[0040] One of the requirements for high-quality 3D sound is high spatial resolution. Achieving high spatial resolution requires a larger number of microphone elements, which in turn requires efficient and highly accurate signal processing and a microphone array design that enables this. Here, Ambisonics with Nc > 2 is called higher-order Ambisonics. In principle, increasing the order allows for more detailed sound fields to be reproduced, but this also increases the number of channels required for recording. For example, second-order Ambisonics requires nine signals, and third-order Ambisonics requires 16 signals. A microphone array typically requires a slightly larger number of microphone elements than these channels.

[0041] Figure 2 is a diagram that shows a schematic overview of Ambisonics. The upper half of the diagram shows encoding, and the lower half shows decoding. In this diagram, audio signals s1, s2, ..., si, ..., sK from K microphone elements are input to the system. Encoding can be broadly divided into two stages. The first stage is a process of spatially mixing the input audio signals. In one example, this mixing is implemented as a spherical harmonic transform of the input audio signals. More specifically, this process calculates the matrix-vector product of the encoding matrix Y+ and the vector of the input audio signals. The encoding matrix Y+ is a matrix obtained by the arrangement of the microphone elements. Note that this mixing process can also be calculated using the least squares method.

[0042] The latter process is a process of temporally mixing the solutions obtained by the former process. In one example, this process is a process of temporally mixing by filtering, i.e., convolution operation, which is independent for each spherical harmonic basis (i.e., without spatial mixing). More specifically, this process applies a radial function (or radial distribution function) Wn(kR) to the solution vector obtained by the former process to obtain the expansion coefficients B. Wn(kR) is a radial distribution function with a radius R of the microphone array and an Ambisonics order n.

[0043] Generally, among these processes, as the order n increases, the amount of calculation, particularly in the pre-stage processing, becomes enormous, and the calculation cost becomes a problem. One of the problems to be solved in this embodiment is to reduce the calculation cost in this pre-stage processing. Therefore, in the following description, the input audio signal corresponds to x in equation (1), and the encoding matrix Y+ corresponds to matrix A in equation (1). Conceptually, the calculation to find Ambisonics coefficients when a certain audio signal is given includes a process of multiplying the audio signal vector by an inverse matrix determined by the arrangement of microphone elements.

[0044] Decoding is conceptually the reverse process of encoding. While detailed explanations are omitted, Fn(kr) shown in Figure 2 is a radial distribution function with the radius r of the speaker array and the Ambisonics order n. Since sound information and directional information are expanded at the listening point during encoding, the playback system does not require directional information from the time of recording. In other words, the orientation of the microphone elements and the orientation of the speakers in the recording and playback systems do not need to match. For more information on higher-order Ambisonics, see, for example, Iwatani et al., "Sound Field Representation by Spherical Harmonic Analysis," Journal of the Acoustical Society of Japan, Vol. 67, No. 11, pp. 544-549 (this document is incorporated herein by reference).

[0045] 1-2. Microphone Array Next, we will provide an overview of microphone arrays. As a microphone array for higher-order Ambisonics, a microphone array mounted on a spherical scatterer (hereinafter referred to as a "hard sphere array") is typically used for sound collection. For example, a 64-channel hard sphere array is known. When designing a microphone array and a signal processing algorithm, attention should be paid to, for example, the following (1) to (5), but to date, no realistic solution that satisfies all of these has existed. The arrangement of microphone elements in a microphone array typically prioritizes the accuracy of the output signal and spatial aliasing. (1) Spatial aliasing (2) Accuracy of the output signal (3) Spatial resolution (4) Computational complexity (i.e., computational cost) (5) Feasibility of implementing microphone elements (i.e., physical constraints in the real world)

[0046] The sound signal picked up by the rigid sphere array is processed based on equation (2), that is, the Ambisonics coefficients are estimated. As the signal processing algorithm, there are known algorithms based on the least squares method and numerical integration. In either case, the calculation cost is O(N 2 )

[0047] Generally, 3D sound recording faces the following challenges (a) to (d): (a) interactivity, (b) personalization, (c) spatial resolution, and (d) the width of the reproducible sound field. In addition to Ambisonics, binaural recording (or binaural sound recording) is also known as a 3D sound recording technique, and we will consider these challenges in conjunction with binaural recording.

[0048] (a) Interactivity (or real-time capability) Interactivity refers to the ability to change the position and orientation of the sound field acquisition in response to user instructions. For example, in a situation where an avatar moves freely within a virtual space, the sound field experienced by the avatar changes depending on the avatar's position and facial orientation. Binaural recording uses a dummy head (i.e., a microphone) placed at a specific position and orientation within the space, making it difficult to achieve interactivity. To achieve interactivity, an omnidirectional microphone array and low-latency real-time processing are required.

[0049] (b) Personalization Personalization refers to changing the playback signal according to the individual characteristics of the user (e.g., ear and head shape). As with interactivity, this is difficult to achieve with binaural recording. To achieve personalization, an omnidirectional microphone array is required.

[0050] (c) Spatial resolution It has been reported that the angular discrimination ability of human hearing reaches 1° in the front direction (for example, Kurozumi Koichi, "2-2 Mechanisms in Hearing," Journal of the Institute of Television Engineers, 45.4 (1991): 438-445). To ensure interactivity, high resolution is required in all directions. For example, calculations have shown that tens of thousands of microphone elements are required to collect sound with a resolution of 1° in all directions. Here, when Nc is several thousand or more, the computation time becomes O(N2 ) requires a huge amount of calculation, and calculations of this order are virtually impossible to achieve.

[0051] (d) Width of reproducible sound field The sound field that can be reproduced accurately is the radius O(N 1 / 2 ) range. In order to accurately represent the sound field of a large space, a large number of microphone elements are required.

[0052] In summary, achieving high-quality stereophonic sound requires an omnidirectional microphone array using a large number of microphone elements and its signal processing. However, with existing methods, computational costs are a bottleneck, making it difficult to achieve real-time (i.e., low-latency) processing. This embodiment provides an improved technology to address this issue.

[0053] As explained above, there is no practical solution that satisfies the requirements for a microphone array and signal processing algorithm, but several ideas (for example, for computer simulation) are known that disregard physical implementation. The problem of how to arrange microphone elements in a microphone array can be viewed as the problem of arranging measurement points (or sampling points) on a spherical surface.

[0054] Figure 3 is a diagram illustrating the arrangement of measurement points on a sphere. Figure 3 shows an example of the arrangement of measurement points according to a type of spherical t design (Delsarte, P., Goethals, JM, and Seidel, JJ, Spherical codes and designs, Geometriae Dedicata, 6 (1977), no. 2, 363-388.). The spherical t design is a method for approximately uniformly arranging points on a sphere.

[0055] A microphone array can be obtained by placing microphone elements at points on a sphere determined according to the spherical t design. This microphone array is thought to have no problems with spatial aliasing, accuracy of the output signal, feasibility of implementation, and efficiency (expansion order / number of microphone elements). However, the computational complexity is O(N 2 ) which is difficult to achieve due to the computational cost.

[0056] FIG. 4 is a diagram showing another example of the arrangement of measurement points on a spherical surface. FIG. 4 shows an example of the arrangement of measurement points based on a tensor product. In this arrangement, coordinates (e.g., θ, φ) are discretized independently, and the point group is defined by the tensor product of the two coordinates. In this example, the same number of measurement points are arranged on multiple equilatitude lines. This arrangement makes it possible to use calculations based on spherical harmonic transformations. Several specific calculation algorithms based on spherical harmonic transformations are known, and although the calculation cost varies depending on the calculation algorithm, it is generally O(N 1.5 )~O(Nlog a N), which is advantageous in terms of computational cost compared to arrangements based on spherical t designs. Therefore, this arrangement is often used in computer simulations.

[0057] However, this arrangement has two fatal flaws. The first is the problem of spatial aliasing. For example, using a typical θ discretization method (e.g., Gauss-Legendre integration points or equiangular points, which can be used with the Driscoll-Healy algorithm) that is advantageous for numerical calculations results in significantly non-uniform microphone element distribution. Specifically, the density of microphone elements is significantly higher near the North Pole (θ = 0) and South Pole (θ = π), while the density is significantly lower near the equator. This arrangement results in differences in the density of measurement points depending on latitude. Measurement points are dense near the North Pole and South Pole, but sparse near the equator. These sparse areas in the measurement point distribution lead to spatial aliasing. The other problem is feasibility. While computer simulations show no problems, the extremely dense density of measurement points near the North Pole and South Pole poses the challenge of installing such a large number of microphone elements in a small space when attempting to implement a physical device.

[0058] 2. Configuration Fig. 5 is a diagram illustrating the functional configuration of a stereophonic system 1 according to one embodiment. The stereophonic system 1 includes a microphone array 10 and a signal processing device 20. The stereophonic system 1 is an example of a system that performs encoding in Ambisonics, i.e., a signal processing system. That is, in the stereophonic system 1, the signal processing device 20 processes the audio signal obtained by the microphone array 10 to obtain an Ambisonics signal. The Ambisonics signal thus obtained can be reproduced using any output system (e.g., a speaker system or a headphone system). A description of decoding is omitted here.

[0059] The stereophonic sound system 1 has a receiving means 21, a calculating means 22, and an output means 23. The receiving means 21 receives input of audio signals of multiple channels output from the microphone array 10. The microphone array 10 has a plurality of microphone elements arranged in a predetermined arrangement. In one example, the microphone array 10 has a multi-ring arrangement (arrangement along multiple rings) described below. The calculating means 22 performs global mixing processing. That is, the calculating means 22 calculates the expansion coefficients of the basis functions for expressing the stereophonic sound field in a time-series manner in O(N 2 ) using an algorithm with lower computational cost. In other words, the calculation means 22 calculates the expansion coefficients using a method other than a method of calculating the entire linear transformation from the input vector (i.e., the input audio signal) to the output vector (i.e., the expansion coefficients) as a single dense matrix vector product. That is, the calculation means 22 calculates the output vector y using a method other than a method of calculating the product of the matrix A and the input vector x as a single dense matrix vector product in the transformation expressed by equation (1). The expansion coefficients correspond to the product of the transformation matrix determined depending on the arrangement of the multiple microphone elements and the vectors of the audio signals of the multiple channels. The output means 23 outputs information related to the expansion coefficients.

[0060] In this example, the stereophonic system 1 further includes a generating means 24. The generating means 24 generates a stereophonic signal (in one example, an Ambisonics signal) for expressing a stereophonic sound field from the expansion coefficients. The output means 23 outputs the generated recording signal as information related to the expansion coefficients.

[0061] In this example, the calculation means 22 performs global mixing processing using two-step mixing (TSM). The two-step mixing refers to a technique for mixing signals in two stages, the first and second processes described herein. Therefore, the calculation means 22 includes a first processing means 25, a complementing means 26, and a second processing means 27. The first processing means 25 performs a first process for each of a plurality of rings of multi-channel audio signals, mixing the audio signals in each ring independently from the other rings to obtain an angular Fourier spectrum. In this specification, "mixing" of signals refers to weighted mixing. That is, "mixing" of signals refers to a process in which each channel is multiplied by a real or complex weighting coefficient before adding the signals of the plurality of channels. The complementing means 26 complements data of missing angular frequency components in the angular Fourier spectra corresponding to the plurality of rings compared to the spectrum with the largest number of angular frequency components obtained, so that the number of angular frequency components is the same for all rings. The second processing means 27 performs a second process of mixing the interpolated angular Fourier spectra. The second process is a process of mixing the angular Fourier spectra of the multiple rings for each angular frequency component to obtain a signal at the front stage of the Ambisonics signal.

[0062] The stereophonic sound system 1 further includes a generating means 24. The generating means 24 performs post-Ambisonics processing on the signal output from the second processing means 27 to obtain an Ambisonics signal. The output means 23 outputs the obtained Ambisonics signal.

[0063] The stereophonic sound system 1 further includes processing means 28 and storage means 29. The processing means 28 performs various processes. The storage means 29 stores various data and programs. Note that the "output" performed by the output means 23 is a concept that includes at least one of output from the signal processing device 20 to an external device and writing to the storage means 29.

[0064] 6 is a diagram illustrating an example of the arrangement of microphone elements in the microphone array 10. The microphone array 10 has a housing 101 and a plurality of microphone elements 102. The housing 101 functions as a scatterer when the microphone elements 102 pick up sound. The plurality of microphone elements 102 are fixed to the housing 101. That is, in the microphone array 10, the positions of the plurality of microphone elements 102 are defined.

[0065] To explain the arrangement of the microphone element 102, a coordinate system with latitude θ and longitude φ is introduced. The definition of the coordinate system is in accordance with ISO 80000-2:2019. A point P on the surface of the housing 101 is N is defined as the origin of latitude θ = 0. For convenience of explanation, point P N is called the "North Pole," and point P N The symmetric point P S The point P on the surface of the housing 101 is called the "South Pole," and the line (θ = π / 2) that is the set of midpoints between the North Pole and the South Pole is called the "Equator." N Another point P a A point P N and point P a The straight line passing through is defined as the line of longitude φ = 0. For convenience of explanation, this line is called the "prime meridian." Note that the equator and the prime meridian are not shown in Figure 6.

[0066] In this example, the microphone elements 102 are divided into multiple groups. The microphone elements 102 in each group are located at the same latitude θ. If the microphone elements 102 in the same group are connected by an imaginary line, a ring is formed, and this group is called a "ring R." The microphone array 10 has I rings R (I≧2). In other words, the microphone array 10 has a multiple ring arrangement. The i-th ring R in ascending order of latitude is called ring R[i]. The ring R closest to the North Pole is ring R[1], and the farthest ring R (i.e., the ring R closest to the South Pole) is ring R[I].

[0067] In each ring R, the microphone elements 102 belonging to that ring R are arranged at equal intervals. The arrangement of the microphone elements 102 in the microphone array 10 according to this embodiment offers greater flexibility than the conventional arrangement based on the spherical t-design and the tensor product arrangement. In one example, the microphone elements 102 are arranged symmetrically with respect to the equator (i.e., the microphone elements 102 are arranged in the same manner in the northern and southern hemispheres). Furthermore, within the latitude θ range of θth1≦θ≦θth2, the microphone elements 102 are arranged in the same manner as the tensor product arrangement. In this example, (π / 2 - θth1) = (θth2 - π / 2), and the angles from the equator to θth1 and θth2 are equal. In other words, this range can be considered a fixed range based on the equator (hereinafter referred to as the "equatorial reference range"; an example of a second range). Within the equatorial reference range, the number of microphone elements in each ring R is the same. In the range where the latitude θ is θ < θth1 and θth2 < θ, i.e., a certain range based on the North Pole and the South Pole (hereinafter referred to as the "polar reference range"; an example of the first range), the arrangement of the microphone elements 102 is an arrangement with fewer microphone elements 102 than the tensor product arrangement. That is, the number of microphone elements 102 in the ring R belonging to the polar reference range is fewer than the number of microphone elements 102 in the ring R belonging to the equatorial reference range. In the ring R belonging to the polar reference range, the number of microphone elements 102 belonging to the ring R may be the same, some may be different, or all may be different. In any case, the number of microphone elements 102 in at least the ring R near the pole is fewer than the ring R near the equator. In other words, the arrangement of the microphone elements 102 in this example can be said to be an arrangement in which the microphone elements 102 belonging to the ring R[i] included in the polar reference range in the tensor product arrangement are thinned out. The number of microphone elements 102 belonging to the ring R[i] is represented as Ni. It should be noted that θth1 and θth2 are not limited to angles equal to each other from the equator, and may be asymmetric with respect to the equator.

[0068] In the ring R[i], the jth microphone element 102 in ascending longitude order is referred to as microphone element 102[i,j]. The position of microphone element 102[i,j] is expressed as coordinates (θi, φj). The arrangement of microphone elements 102 in this example is referred to as a "modified multi-ring configuration (MMR)." Here, "modified" refers to a modification from the arrangement based on the tensor product in Figure 4. In other words, while the example in Figure 4 placed the same number of measurement points on multiple equilatitude lines, this has been modified to "place different numbers of measurement points at equal intervals (at least in rings near the poles) compared to rings near the equator." The advantages of this arrangement will be discussed later.

[0069] The discrete values ​​of latitude θ for the multiple rings R are selected, for example, based on Gauss-Legendre integral points. Alternatively, the discrete values ​​of latitude θ may be selected using the method of Swarztrauber and Spotz (1999). In this example, the longitude coordinate values ​​of the microphone elements 102 belonging to the multiple rings R included in the equatorial reference range may all be the same. That is, the sets of longitudes (coordinate values) of the microphone elements 102 belonging to each ring R are the same for the multiple rings R included in the equatorial reference range. The sets of longitudes do not necessarily have to be the same for all the multiple rings R included in the equatorial reference range; some rings R may have sets of longitudes different from those of the other rings R. That is, the sets of longitudes may be the same for at least two adjacent rings R. At least one of the polar reference range and the equatorial reference range may include three or more rings R.

[0070] In this example, the microphone array 10 can be said to be a microphone array in which the microphone elements 102 are arranged at the positions of the points of a point cloud that does not coincide with the set of vertices of any one of a regular polyhedron, a semiregular polyhedron, a Catalan solid, a Goldberg polyhedron, or a geodesic polyhedron, or with a rotated set of vertices. The microphone array 10 may have microphone elements 102 arranged in a different manner that satisfies these conditions, other than the example in FIG. 6.

[0071] In one example, the housing 101 is a spherical rigid body, and is formed of, for example, resin, metal, or a composite material thereof. More specifically, a plurality of microphone elements 102 are fixed to the surface of the housing 101. For example, adhesive or fasteners are used for fixing. The fasteners include, for example, at least one of screws, bolts, nuts, rivets, clips, and pins. Alternatively, a plurality of holes or recesses may be formed in the surface of the housing 101, and the microphone elements 102 may be embedded and fixed in these holes or recesses. The housing 101 may be hollow or solid.

[0072] Instead of or in addition to a spherical rigid body, the housing 101 may have a plurality of frames fixed to each other. The microphone elements 102 are placed in the frames so as to be positioned on an imaginary spherical surface.

[0073] 6, the microphone array 10 also has cables for outputting signals from each microphone element 102. In order to output signals from the N microphone elements 102, the microphone array 10 has N cables. A plurality of holes or recesses may be formed in the surface of the housing 101, through which these cables may be passed.

[0074] FIG. 7 is a diagram illustrating an example of the hardware configuration of the signal processing device 20. The signal processing device 20 includes an audio interface 201, a CPU (Central Processing Unit) 202, a memory 203, a storage 204, a communication interface 205, and an output interface 206. The audio interface 201 receives an output signal from the microphone array 10. The audio interface 201 includes at least N connectors (not shown) for connecting N cables. The audio interface 201 converts the signal received from the microphone array 10 into data. The CPU 202 is a processing device that processes various data according to a program. The memory 203 is a main storage device that functions as a work area when the CPU 202 executes a program, and includes, for example, a RAM (Random Access Memory). The storage 204 is an auxiliary storage device that stores various data and programs, and includes, for example, a solid state drive (SSD) or a hard disk drive (HDD). The communication interface 205 is a device that communicates with other devices according to a predetermined communication standard. The signal processing device 20 can communicate with other devices, for example, via the Internet, via the communication interface 205. The output interface 206 outputs signals to an output device (for example, a display and / or a speaker).

[0075] In this example, the storage 204 stores a program (hereinafter referred to as the "signal processing program") for causing the computer device to function as the signal processing device 20. When the CPU 202 is executing the signal processing program, the audio interface 201 is an example of a receiving means 21, the CPU 202 is an example of a calculating means 22, a generating means 24, a first processing means 25, a complementing means 26, and a second processing means 27, at least one of the communication interface 205 and the output interface 206 is an example of an output means 23, and at least one of the memory 203 and the storage 204 is an example of a storage means 29.

[0076] 8 is a flowchart showing the operation of the signal processing device 20. The following flow is started, for example, when a user starts a signal processing program in the signal processing device 20. The following processing is realized by the CPU 202 executing the signal processing program working in cooperation with other hardware elements such as the memory 203.

[0077] In step S1 , the signal processing device 20 acquires multi-channel audio signals from the microphone array 10 .

[0078] FIG. 9 is a schematic diagram showing an audio signal acquired from the microphone array 10. In this example, the microphone array 10 has L rings R. Let Ni be the number of microphone elements 102 belonging to the ring R[i]. The signal from the microphone elements 102[i, j] at a certain moment (i.e., a certain time t) is represented as S[i, j]. When referring collectively without specifying individual microphone elements, it is simply referred to as signal S. At a certain moment, the signal S[i, j] is a scalar. In FIG. 9, the signals S from the microphone elements 102 belonging to a certain ring R[i] are illustrated lined up in a single line. In this figure, the signal S from the microphone elements 102 belonging to the ring R closer to the North Pole is located higher. In each ring R, the signal S from the microphone elements 102 closer to the prime meridian is located further to the left. The rings R closer to the North Pole and South Pole are shorter because they have fewer microphone elements 102. The rings R near the equator are longer.

[0079] 8 again, in step S2, the signal processing device 20 performs global mixing processing, which in this example is performed based on a two-stage mixing method and is divided into steps S21 to S23.

[0080] In step S21, the signal processing device 20 performs local mixing processing (an example of a first processing). The local mixing processing is processing for mixing the audio signal S for each ring R. The local mixing processing is an example of processing for each of multiple rings R, mixing the audio signals in the ring R independently from the other rings R to obtain an angular Fourier spectrum. Specifically, the signal processing device 20 performs a discrete Fourier transform (DFT) on the signal S for each ring R. Note that the arrangement of the microphone elements 102 in the microphone array 10 is known to the signal processing device 20. The signal processing device 20 performs the discrete Fourier transform using parameters of the arrangement of the microphone elements 102 (such as the number of rings, latitude θ, and longitude φ).

[0081] FIG. 10 is a schematic diagram showing the results of the local blending process. Because the DFT is performed for each ring R, the DFT results are also divided into L lines. The DFT results for each ring are called the angular Fourier spectrum. In FIG. 10, the horizontal axis represents the spatial frequency ω [rad-1] (i.e., angular frequency). In this example, the microphone elements 102 are spaced equally apart in each ring R, so the angle formed by two adjacent microphone elements 102 when viewed from the center of the sphere is larger closer to the north and south poles and smaller closer to the equator. Therefore, in the DFT results, the closer to the equator there are high-frequency components, and the closer to the north and south poles there are high-frequency components, the more they are lost (disappear).

[0082] The results of this DFT are divided into multiple discrete frequencies (i.e., frequency components). The kth lowest frequency component in the ring R[i] is represented as ω[i,k]. The discrete frequencies are common to all rings R. That is, ω[i,k] = ω[j,k]. In the ring R[i], the DFT results are obtained for the Mi lowest discrete frequencies. There are higher frequency components closer to the equator, and the higher frequency components disappear closer to the North and South Pole, so Mi becomes smaller closer to the North and South Pole and larger closer to the equator.

[0083] Referring again to FIG. 8 , in step S22, the signal processing device 20 interpolates the angular Fourier spectrum. This process is an example of interpolating data on angular frequency components missing from angular Fourier spectra corresponding to multiple rings R compared to the spectrum with the largest number of obtained angular frequency components so that the number of angular frequency components is the same for all rings R. Here, Nmax is the maximum value of Ni, i.e., the number of microphone elements 102 belonging to the ring[i] with the largest number of microphone elements 102. The signal processing device 20 interpolates discrete frequency components at high frequencies so that the number of discrete frequencies (having component values) is the same (i.e., Nmax or greater) for all rings R[i]. In one example, the signal processing device 20 interpolates discrete frequency components with no value as constants (specifically, zero, for example). Alternatively, the signal processing device 20 may interpolate using values ​​of high angular frequency components estimated using any super-resolution method. When interpolating using zeros, calculations may be omitted because these components do not contribute to the signal in actual implementation.

[0084] 11 shows an example of the angular Fourier spectrum after the interpolation process. The missing data in the local mixing process of step S2 is interpolated, and the number of frequency components is the same for all rings R (P in this example).

[0085] Referring again to Fig. 8, in step S23, the signal processing device 20 performs local-to-global mixing (L2G mixing, an example of second processing). The local-to-global mixing is a process of mixing signals of multiple rings R[i], and the weights of the spherical harmonic basis, i.e., the coefficients of the front stage of the Ambisonics signal, are calculated by linearly transforming the signal vectors belonging to each angular frequency from the angular Fourier spectrum interpolated for each ring R[i].

[0086] Figure 12 is a schematic diagram of the local-global blending process. In the local-global blending process, the output signal y (in this example, the result of operation with the encoding matrix Y+ in Figure 2) is obtained by blending each angular frequency component of the set of interpolated angular Fourier spectra obtained for each ring R[i]. This method reduces the computational cost of the global blending process to O(N 2 ), whereas, for example, the number of rings is O(N 0.5 ), the computational cost of the entire signal processing for mixing is at most O(N 1.5 ) For example, I = O(N (1-a) ) and 0≦a≦1, and if both local and local-global mixing are performed by dense matrix-vector products, the total asymptotic computational cost is at most O(N (1+a) ) + O(N (2-a) ) This asymptotic calculation cost is smallest when a = 1 / 2, and the total asymptotic calculation cost is at most O(N 1.5 ) This computational cost can be further reduced by using less computationally expensive methods in both the local and local-global blending steps.

[0087] Here, we will provide additional information on the computational cost of the mixing process. Let N be the number of microphone elements 102 in the microphone array 10, Np be the number of microphone elements in the tensor product arrangement corresponding to the microphone array 10, and Mp be the number of basis functions. The tensor product arrangement corresponding to the microphone array 10 refers to an arrangement in which the number of microphone elements in all rings is the same, and is the same as the number of microphone elements 102 in the ring R closest to the equator in the microphone array 10. The computational complexity of the conventional method (a method of simply calculating a dense matrix-vector product) is O(N 2 ), the calculation cost in this embodiment is O(N 1.5), but this is the calculation cost when it is assumed that O(N) = O(Np) = O(Mp). Strictly speaking, the calculation cost of the conventional method is O(NM), and the calculation cost of this embodiment is considered to be, more strictly speaking, a more complicated formula including Np and Mp. In an actual product, the difference between N and Np cannot be made extremely large (it is not possible to reduce the number of microphone elements by an extremely large number), so it can be understood that the assumption that O(N) = O(Np) = O(Mp) is valid.

[0088] It can be said that in the signal processing device 20, the calculation means 22 calculates M expansion coefficients of the basis functions for expressing a three-dimensional sound field, which correspond to the product of a transformation matrix and the vectors of N-channel audio signals, using an algorithm with a computational cost lower than O(NM).

[0089] For example, the local-global blending process may be implemented by dense matrix-vector multiplication, or may use a matrix-vector multiplication calculation method using matrix decomposition, such as an algorithm for multi-stage decomposition and multi-stage application of a transformation matrix.

[0090] The computational cost can be further reduced by using a faster algorithm in each step of the local mixing process and the local-global mixing process, for example, a fast Fourier transform or a process of multi-stage decomposition and multi-stage application of a transformation matrix for the local mixing process, and a process of multi-stage decomposition and multi-stage application of a transformation matrix for the local-global mixing process.

[0091] Referring again to Figure 8, in step S3, the signal processing device 20 outputs the generated Ambisonics signal. In one example, this output is stored in the storage 204 or transmitted to another device (e.g., a server on the Internet) via the communication interface 205. The processes of steps S2 to S4 are performed on data of a limited moment of the input audio signal, but by continuously repeating this process, an Ambisonics signal can be obtained for the entire input audio signal.

[0092] In this embodiment, the use of a multi-ring arrangement in the microphone array 10 simultaneously solves the problem of signal processing computational costs and problems resulting from the uniformity of microphone element arrangement (for example, spatial aliasing and implementability). The problem of computational costs, which was a fatal issue in existing technologies that were based solely on naive dense matrix-vector product calculations, has been solved in an implementable manner by this embodiment, making it possible for the first time to capture high-spatial-resolution stereophonic sound using an ultra-multichannel microphone array.

[0093] 4. Modifications The present invention is not limited to the above-described embodiment, and various modifications are possible. Some modifications will be described below. Some of the matters described below may be applied in combination with some of the above-described embodiment or other modifications.

[0094] 4-1. Spacing of Microphone Elements The arrangement of microphone elements in a microphone array is not limited to that illustrated in the embodiment. For example, in the embodiment, an example has been described in which the spacing between two adjacent microphone elements 102 within a ring R is the same in all rings R. However, the spacing between two adjacent microphone elements 102 within a ring R may be equal in one ring R but different from the spacing between two adjacent microphone elements 102 in other rings R. That is, the spacing d[i] between two adjacent microphone elements 102 in a ring R[i] may satisfy d[i]≠d[i+1] for at least some of i. For example, d[i] > d[i+1] may be satisfied in a ring R[i] relatively close to the equator and a ring R[i+1] far from the equator.

[0095] The number of microphone elements 102 belonging to two adjacent rings R does not need to be different and may be the same. That is, N[i] = N[i+1] may be set for some two adjacent rings R[i] and R[i+1] among the L rings R that make up the microphone array 10.

[0096] Alternatively, the microphone elements 102 may be arranged at non-equidistant intervals in each ring R. When the microphone elements 102 are arranged at non-equidistant intervals, local mixing processing can be performed using a calculation method such as a non-uniform fast Fourier transform (NUFFT). In this case, the signal processing device 20 may switch the calculation algorithm used in the local mixing processing depending on whether the microphone elements 102 are arranged at equal intervals or non-equidistant intervals in each ring R of the microphone array 10.

[0097] 4-2. Shape of the Microphone Array The shape of the housing 101 in the microphone array 10 is not limited to a sphere. The housing 101 may have a spheroidal shape, which is a sphere stretched in a specific direction. Spheroids include oblate spheroids and prolate spheroids, and the housing 101 may have either shape. Even if the housing 101 has a spheroidal shape, i.e., if the microphone elements 102 are arranged along a spheroid rather than a sphere, the same calculation method as in the case where the housing 101 is spherical can be used. Furthermore, the spheres and spheroids referred to in this paper do not necessarily have to be mathematically strict spheres and spheroids; they may have shapes that are distorted from spheres and spheroids within acceptable limits. "Within acceptable limits" means that even if signal processing is performed assuming the shape to be a sphere or spheroid, the difference in the output signal is so slight that it can be ignored. For example, the surface of the sphere or spheroid may have slight irregularities for designs, screw holes, supports, etc., one or more parts of the sphere or spheroid may be flat, or it may have any three-dimensional shape that approximates a sphere or spheroid.

[0098] In another example, the housing 101 may have any shape or topology. In this case, signal processing is performed by matrix decomposition of the global mixing matrix and its application to realize mixing calculation. As the matrix decomposition, singular value decomposition or multi-stage decomposition of a transformation matrix can be used.

[0099] 4-3. Arrangement of microphone elements The arrangement of microphone elements in a microphone array is not limited to that exemplified in the embodiment. For example, the microphone array does not need to have a bicyclic arrangement. In this case, signal processing is performed using matrix decomposition of a global mixing matrix and its application to realize mixing calculations. As an argument for matrix decomposition, singular value decomposition or multi-stage decomposition of a transformation matrix can be used.

[0100] Even if the microphone array has a multiple-ring arrangement, the first microphone element 102 does not have to be positioned at φ = 0. For example, the longitude φj of the j-th microphone element 102 in the ring R[i] to which Ni microphone elements 102 belong can be calculated using an arbitrary constant c as follows: because any phase difference can be cancelled out in the Fourier domain by multiplying it with a correction factor.

[0101] Furthermore, two microphone array arrangements that can be transformed into each other by a general transformation that is a combination of a rotational transformation, a translational transformation, and an additional rotational transformation can be said to be equivalent arrangements. In other words, an arrangement that matches the multi-ring arrangement of this embodiment by a general transformation is equivalent to the multi-ring arrangement.

[0102] 4-4. Structure of the microphone array The specific structure of the microphone array is not limited to that exemplified in the embodiment. In one example, the microphone array 10 may be formed by bonding flexible printed circuits (FPCs) to a housing 101. One or more microphone elements 102 are formed on one flexible substrate.

[0103] However, even when flexible substrates are used, as the number of microphone elements 102 increases, the number of required flexible substrates also increases, making assembly more difficult. Therefore, a microphone array may be formed by molding rigid substrates into a three-dimensional shape and assembling these. To form a substrate having a three-dimensional shape, techniques such as molded interconnect devices (MID), injection molded structural electronics (IMSE), or 3D printing are used.

[0104] FIG. 13 is a diagram illustrating an overview of a microphone array 10 according to this modification. The example in FIG. 13 shows a microphone array 10 with a multi-ring arrangement mounted on the surface of a spherical scatterer. Black circles represent microphone elements 102, and black lines indicate the boundaries of the allocation of the microphone elements 102 to the substrates. In this example, each region is formed by a rigid, curved circuit board. Two microphone elements 102 are formed on each circuit board. In this way, it is preferable to use the multi-ring arrangement exemplified in the embodiment when determining the allocation of microphone elements 102 to individual substrates. In the multi-ring arrangement, the microphone elements 102 are divided into groups that share a latitude parameter, and allocating all or part of each ring R to a single substrate facilitates circuit design, mechanical design, and assembly.

[0105] The calculation for obtaining the angular Fourier spectrum in the local mixing process is not limited to the discrete Fourier transform. Any calculation may be used as long as it is a discrete Fourier transform, such as a fast Fourier transform (FFT), a discrete cosine transform (DCT), a discrete sine transform (DST), a real discrete Fourier transform (RDFT), or a real fast Fourier transform (FFT).

[0106] The basis functions used in the global mixing process are not limited to spherical harmonic functions. Any basis function that can express a three-dimensional sound field, such as a spherical wave basis, a plane wave basis, or a spheroidal wave function basis, may be used.

[0107] 4-7. Factorized Global Mixing Method The method by which the calculation means 22 performs global mixing processing is not limited to the two-stage mixing method exemplified in the embodiment. Instead of the two-stage mixing method, the calculation means 22 may use factorized global mixing (FGM). The factorized global mixing method is a method based on matrix decomposition of the global mixing matrix. The calculation cost of the global mixing processing can be dramatically reduced by using a high-speed matrix-vector product calculation method based on matrix decomposition of the global mixing matrix. Typically, this is achieved by decomposing the global mixing matrix and quickly calculating a dense matrix-vector product using the calculated decomposition. As a matrix decomposition method, for example, singular value decomposition or multi-stage decomposition and multi-stage application of a transformation matrix can be used. The matrix-vector product calculation method using singular value decomposition is a method of reducing the calculation cost of the matrix-vector product by performing singular value decomposition on the global mixing matrix and ignoring singular values ​​with small absolute values ​​to perform low-rank approximation of the original matrix with a matrix of a smaller rank (rank K). Here, K is a constant that is independent of N and M and is smaller than both N and M. The multi-stage decomposition and multi-stage application process of a transformation matrix is ​​a calculation method that performs matrix decomposition in multiple stages and multiplies the input vector by the decomposed matrices in sequence, and this process can be understood as an extension of the matrix-vector product calculation method using singular value decomposition. This process is a numerical calculation method that makes it possible to quickly calculate a matrix-vector product for any given matrix by performing multiple approximations with matrices that have a lower rank than the original matrix. The multi-stage decomposition and multi-stage application process of a transformation matrix allows for high-speed calculations by allowing for some error in some cases. The calculation cost of these algorithms is O(N 2 ) less than

[0108] As is well known, the global mixing matrix for a microphone array with an arbitrary microphone arrangement can be calculated exactly or approximately from a matrix representing the inverse system of the global mixing matrix obtained by physical measurement or physical simulation. The signal processing device 20 can calculate this global mixing matrix and its decomposition in advance. Alternatively, the signal processing device 20 can obtain the decomposition of the global mixing matrix calculated in advance in another device.

[0109] These calculation methods are particularly useful when the microphone elements in a microphone array are arranged in a configuration other than a compound ring. The decomposed global mixing method can be used not only when the microphone elements are arranged in a compound ring, but also when the microphone elements are arranged at arbitrary positions along the surface of a sphere or spheroid, and even when the microphone elements are arranged at arbitrary positions along a curved surface of any shape or topology other than a sphere or spheroid.

[0110] 4-8. Method Based on Matrix Decomposition of Transformation Matrix The two-stage mixing method described in the embodiment can also be interpreted mathematically as a method of decomposing the entire transformation matrix A into the product of a local mixing matrix A1 and a local global mixing matrix A2 and applying them sequentially, i.e., as a type of matrix decomposition method. Based on this interpretation, both the two-stage mixing method described in the embodiment and the decomposed global mixing method described in the modified example can also be said to be methods based on matrix decomposition of a transformation matrix.

[0111] 4-9. Output from Output Means The signal output by the output means 23 is not limited to the Ambisonic signal generated by the generation means 24. The output means 23 may output the result of calculation by the calculation means 22 without processing by the generation means 24 (for example, as is). In other words, the information output by the output means 23 is not limited to the Ambisonic signal, but may be information at a stage preceding the Ambisonic signal. Even information at a stage preceding the Ambisonic signal can be used for spatial sound field analysis, and is therefore useful.

[0112] 4-10. Signal Processing The signal processing in the present invention relates to spatial signal mixing, and does not depend on whether the signal is a signal in the time domain or a signal in the frequency domain. Signals in either domain may be processed.

[0113] In the case of signal mixing using the decomposition global mixing method, the output signal may be a general signal. In other words, the output signal does not need to be a signal associated with the weights of the basis functions, and can be applied to any signal processing represented by a global mixing matrix. In other words, even if the physical meaning of the signal mixing is unknown, signal mixing using the decomposition global mixing method can be applied as long as the global mixing matrix is ​​given.

[0114] The method for recording described in the embodiment can also be applied to signal processing for reproduction, i.e., signal processing for calculating drive signals for a loudspeaker array (this loudspeaker array may be physically real or may be computationally virtual) by mixing weight signals of basis functions. In one example, a global mixing process is performed in the order of a global local mixing process and a local mixing process.

[0115] 4-12. Target of Signal Processing The signal processing device 20 may perform the signal processing described in the embodiments on signals of only some of the channels of the multi-channel audio signals output from the microphone array 10. For signals of channels other than the target, calculations may be performed using, for example, dense matrix vector product or the least squares method. In other words, the microphone array may have additional microphone elements to which other signal processing methods are applied, in addition to the microphone elements to which the signal processing methods described in the embodiments are applied.

[0116] 4-13. System Configuration The specific system configuration, i.e., hardware configuration, of the stereophonic sound system 1 is not limited to that exemplified in the embodiment. For example, a signal transmission path (e.g., a computer network such as the Internet) may exist between the microphone array 10 and the signal processing device 20. For example, the microphone array 10 and the signal processing device 20 may be located in physically remote locations separated by the Internet or the like. Alternatively, the stereophonic sound system 1 may include a signal processing device for outputting audio from speakers or headphones, and the signal processing device 20 and the output signal processing device may be located in physically remote locations separated by the Internet or the like.

[0117] Some or all of the functions implemented in the signal processing device 20 in the embodiment may be implemented in a server on the Internet. This server may be a physical server or a virtual server (including a so-called cloud).

[0118] 4-14. Functional Configuration Some of the functional configurations described in the embodiments may be omitted. Furthermore, the correspondence between the functional configurations and hardware configurations described in the embodiments is merely an example, and each functional element may be implemented in any hardware element. In this document, the term "system" refers to a device (or group) consisting of one or more pieces of hardware and software for realizing a specific function.

[0119] The stereophonic system 1 according to this embodiment can be used for the following tasks, for example: decoding a stereophonic expression signal into a speaker array driving signal, decoding into a binaural signal, beamforming, sound source localization, noise suppression, sound enhancement, speech recognition, speaker estimation, acoustic scene analysis, acoustic scene recognition, emotion recognition, music information processing, sound source separation, voice conversion, audio compression, or application of audio effects.

[0120] 4-16. Programs In the embodiments, the programs executed by a processor such as the CPU 202 may be distributed in a state recorded on a non-transitory recording medium such as a DVD-ROM, or may be distributed by being made available for download from a server on a computer network.

[0121] 1...stereophonic sound system, 10...microphone array, 20...signal processing device, 21...receiving means, 22...calculating means, 23...output means, 24...generating means, 25...first processing means, 26...complementing means, 27...second processing means, 28...processing means, 29...storage means, 101...casing, 102...microphone element, 202...CPU, 203...memory, 204...storage, 205...communication interface, 206...output interface

Claims

1. A receiving means for receiving, as an input audio signal, an N-channel audio signal output from a microphone array including a plurality of microphone elements having a determined arrangement; a calculation means for calculating, by an algorithm with a calculation cost less than the order of NM, M expansion coefficients of basis functions for expressing a stereophonic field, which corresponds to the product of a conversion matrix determined according to the arrangement of the plurality of microphone elements and a vector of the input audio signal; and an output means for outputting information regarding the expansion coefficients. A signal processing system having the above.

2. A receiving means for receiving, as an input audio signal, a multi-channel audio signal output from a microphone array including a plurality of microphone elements having a determined arrangement; a calculation means for calculating expansion coefficients of basis functions for expressing a stereophonic field, which corresponds to the product of a conversion matrix determined according to the arrangement of the plurality of microphone elements and a vector of the input audio signal, by a method based on matrix decomposition of the conversion matrix; and an output means for outputting information regarding the expansion coefficients. A signal processing system having the above.

3. A signal processing system according to claim 1 or 2, further comprising a generating means for generating a stereophonic audio signal for expressing a stereophonic field from the expansion coefficients, wherein the output means outputs the stereophonic audio signal as information regarding the expansion coefficients.

4. The determined arrangement is an arrangement along a plurality of virtual rings, and the calculation means includes: a first processing means for performing a first process of obtaining an angular Fourier spectrum by mixing audio signals within each of the plurality of rings independently from other rings from the input audio signal; a complementing means for complementing data of angular frequency components missing in comparison with a spectrum having the largest number of obtained angular frequency components in the angular Fourier spectra corresponding to the plurality of rings so that the number of angular frequency components is common for all rings; and a second processing means for performing a second process of obtaining the expansion coefficients by mixing the angular Fourier spectra of the plurality of rings for each angular frequency component on the complemented plurality of angular Fourier spectra. A signal processing system according to claim 1 or 2 having the above.

5. A signal processing system according to claim 4, wherein a plurality of microphone elements are arranged at equal intervals in the plurality of rings, and the first process includes a dense matrix vector product or a Fourier-type transform.

6. In the plurality of rings, a plurality of microphone elements are arranged at unequal intervals, and the first process includes a dense matrix vector product, a non-uniform fast Fourier transform, or a matrix decomposition method. The signal processing system according to claim 4.

7. The second process includes a matrix vector product by a plurality of dense matrix vector products or a matrix decomposition method. The signal processing system according to claim 4.

8. For each of the plurality of angular frequency components, the plurality of dense matrix vector products are calculated for the angular Fourier spectra of the plurality of rings independently of other angular frequency components. The signal processing system according to claim 6.

9. The matrix decomposition method includes singular value decomposition or multi-stage decomposition of the transformation matrix. The signal processing system according to claim 6.

10. The calculation means includes a calculation for obtaining the expansion coefficients by a matrix decomposition method. The signal processing system according to claim 1 or 2.

11. The matrix decomposition method includes a process of multi-stage decomposition and multi-stage application of the transformation matrix. The signal processing system according to claim 10.

12. The signal processing system according to claim 1 or 2, having the microphone array.

13. The plurality of microphone elements are arranged along the surface of a virtual or physical sphere or ellipsoid of revolution. The signal processing system according to claim 12.

14. The plurality of rings are parallel on the surface of the sphere or ellipsoid of revolution. The signal processing system according to claim 13.

15. In each of the plurality of rings, a plurality of microphone elements are arranged at equal intervals. The signal processing system according to claim 13.

16. The plurality of rings include two adjacent rings having the same number of microphone elements. The signal processing system according to claim 13.

17. The basis function is a spherical harmonic function, and the expansion coefficients are used for the calculation of an ambisonic signal. The signal processing system according to claim 1 or 2.

18. A step of receiving an N-channel audio signal output from a microphone array including a plurality of microphone elements having a determined arrangement as an input audio signal; a step of calculating M expansion coefficients of basis functions for expressing a stereophonic sound field, which correspond to a product of a conversion matrix determined according to the arrangement of the plurality of microphone elements and a vector of the input audio signal, by an algorithm with a calculation cost less than that of order NM; and a step of outputting information regarding the expansion coefficients.

19. A step of receiving a multi-channel audio signal output from a microphone array including a plurality of microphone elements having a determined arrangement as an input audio signal; a step of calculating expansion coefficients of basis functions for expressing a stereophonic sound field, which correspond to a product of a conversion matrix determined according to the arrangement of the plurality of microphone elements and a vector of the input audio signal, by a method based on matrix decomposition of the conversion matrix; and a step of outputting information regarding the expansion coefficients.

20. A microphone array having a plurality of microphone elements arranged along a plurality of virtual rings formed on the surface of a virtual or physical sphere or ellipsoid of revolution, the plurality of rings being parallel to lines of equal latitude in a defined coordinate system.

21. The microphone array according to claim 20, wherein in each of the plurality of rings, a plurality of microphone elements are arranged at equal intervals.

22. The microphone array according to claim 20, wherein the plurality of rings include two adjacent rings having the same number of microphone elements.

23. The microphone array according to claim 20, wherein the number of microphone elements belonging to a ring included in a first range where the latitude θ in the coordinate system satisfies θ < θth1 or θth2 < θ is less than the number of microphone elements belonging to a second range where θth1 ≦ θ ≦ θth2.

24. The microphone array according to claim 23, wherein the number of microphone elements belonging to each of the plurality of rings included in the second range is all equal.

25. The microphone array according to claim 23, wherein the set of longitudes of the microphone elements belonging to each of the plurality of rings included in the second range is the same for adjacent rings.

Citation Information

Patent Citations

  • AUDIO DATA PROCESSING METHOD AND SOUND COLLECTOR FOR IMPLEMENTING THIS METHOD

    JP2006506918A