Binauralization method and device using virtual microphones

The encoding device generates virtual stereo microphones using extrapolation from multiple actual microphones, addressing spatial representation issues and simplifying data conversion, ensuring accurate audio capture and perception.

WO2025159083A1PCT designated stage expired Publication Date: 2025-07-31PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/001766
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-12-10
Filing Date
2025-01-21
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing methods for capturing audio signals using multiple microphones on devices like smartphones struggle to accurately represent spatial characteristics due to non-horizontal microphone arrangements, leading to incorrect inter-channel time differences and difficulties in converting microphone positions between devices.

Method used

An encoding device uses the phase components of signals from at least three microphones arranged in a two-dimensional plane to generate virtual microphones at specific positions, performing extrapolation to create virtual stereo microphones with an inter-aural distance, allowing for accurate spatial representation and simplified data conversion across devices.

Benefits of technology

This approach enables accurate capture and representation of spatial audio characteristics without increasing the number of microphones, ensuring correct perception of sound direction and simplifying data conversion on receiving devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025001766_31072025_PF_FP_ABST
    Figure JP2025001766_31072025_PF_FP_ABST
Patent Text Reader

Abstract

In the present invention, an encoding device is provided with a generation unit that, by using at least the phase components of signals of at least three microphones arranged in a two-dimensional plane, generates signals at a first position and a second position on a straight line connecting two virtual microphones constituting a virtual stereo microphone, arranged along a first direction of the two-dimensional plane, and that generates signals of the two virtual microphones positioned at both ends, the first position and the second position, by performing extrapolation processing with respect to the signals at the first position and the second position.
Need to check novelty before this filing date? Find Prior Art

Description

Binaural generation method and device using virtual microphone

[0001] The present disclosure relates to an encoding device and an encoding method using a binauralization method and device using a virtual microphone.

[0002] For example, there is a signal processing technique related to capturing an audio signal or an acoustic signal (hereinafter referred to as an "audio / acoustic signal") (see, for example, Patent Document 1).

[0003] International Publication No. 2012 / 061151

[0004] Ryoga Jinzai, Kouei Yamaoka, Mitsuo Matsumoto, Takeshi Yamada, Shoji Makino, “Microphone Position Realignment by Extrapolation of Virtual Microphone”, Proceedings, APSIPA Annual Summit and Conference 2018

[0005] There is room for further study on how to capture audio signals.

[0006] Non-limiting embodiments of the present disclosure contribute to providing an encoding device and an encoding method that can appropriately capture an audio signal.

[0007] An encoding device according to one embodiment of the present disclosure includes a generation unit that uses at least phase components of signals from at least three microphones arranged on a two-dimensional plane to generate signals for a first position and a second position on a line connecting two virtual microphones that constitute a virtual stereo microphone arranged in a first direction on the two-dimensional plane, and generates signals for the two virtual microphones located at both ends of the first position and the second position by extrapolating the signals from the first position and the second position; and an encoding unit that performs stereo encoding using the signals from the two virtual microphones.

[0008] These comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a recording medium, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0009] According to an embodiment of the present disclosure, audio and sound signals can be appropriately captured.

[0010] Further advantages and benefits of one embodiment of the present disclosure will become apparent from the specification and drawings. Such advantages and / or benefits may be provided by some embodiments and features described in the specification and drawings, respectively, but not necessarily all of them may be provided to obtain one or more identical features.

[0011] FIG. 1 is a block diagram showing an example of the configuration of an audio / acoustic signal encoding device. FIG. 2 is a diagram showing an example of capturing an audio / acoustic signal. FIG. 3 is a diagram showing an example of capturing an audio / acoustic signal. FIG. 4 is a diagram showing an example of capturing an audio / acoustic signal. FIG. 5 is a diagram showing an example of capturing an audio / acoustic signal. FIG. 6 is a diagram showing an example of capturing an audio / acoustic signal. FIG. 7 is a diagram showing an example of capturing an audio / acoustic signal. FIG.

[0012] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings.

[0013] For example, a mobile terminal (also referred to as a terminal), such as a smartphone or a tablet, is capable of communicating audio signals (or audio data), and the terminal is equipped with multiple microphones (hereinafter referred to as "mics").

[0014] Types of microphones include, for example, Micro Electro Mechanical Systems microphones (hereinafter referred to as "MEMS microphones"), condenser microphones, etc. For example, there are stereo microphones and multi-channel microphones that are configured using multiple MEMS microphones.

[0015] For example, when a terminal simultaneously captures camera video data and stereo audio data, the orientation (or tilt) of the terminal may change (for example, from portrait to landscape, or from landscape to portrait).

[0016] For example, when inputting camera images, a device may have a function that automatically switches the orientation of the camera image capture depending on the orientation of the device.

[0017] Similarly, when inputting audio signals, it is expected that the orientation for capturing the audio signals (e.g., microphone position) will be appropriately switched depending on the orientation of the terminal. For example, a method has been proposed for switching microphones (e.g., stereo microphones) used to capture audio signals (e.g., stereo audio) so that a pair of microphones positioned horizontally or nearly horizontally is selected depending on the orientation of the terminal (see, for example, Patent Document 1). The reason for switching to a horizontal or nearly horizontal microphone pair is that it is effective for stereo microphones to be positioned horizontally when expressing spatial characteristics using stereo microphones.

[0018] For example, when capturing audio signals (e.g., stereo audio data), there may be a method for selecting (or switching) a microphone pair. However, when capturing stereo audio data, when switching microphone pairs using multiple microphones provided on a terminal, the microphone pairs are not necessarily arranged horizontally. If the microphone pairs are not arranged horizontally, capturing using the microphone pairs may not properly obtain the L channel (left channel or L-ch) signal and the R channel (right channel or R-ch) signal included in the stereo audio data. Furthermore, for example, to support all orientations of the terminal (e.g., to cover all orientations of 360 degrees), the number of microphones provided on the terminal would increase, which is unrealistic.

[0019] Furthermore, for example, stereo microphones mounted on a terminal such as a smartphone may be placed close to each other at a distance (e.g., several centimeters) that is extremely short compared to the interaural distance due to the size of the terminal. Therefore, when a receiver listens to (plays back) stereo audio captured by the stereo microphones mounted on the terminal using headphones, for example, the inter-channel time difference (ITD) is likely to be smaller than the time difference at the interaural distance, and the direction (or orientation) of the captured sound may not be correctly perceived in a transmitting device (e.g., an encoding device).

[0020] In response to this, for example, there may be a method in which the receiving device (e.g., a decoding device) converts the decoded stereo audio data into audio data corresponding to the inter-aural distance. However, in this method, since the microphone positions mounted on the transmitting device may differ from device to device, it is not easy for the receiving device to uniformly convert audio data from the transmitting device at any inter-microphone distance into audio data corresponding to the inter-aural distance unless the microphone position information on the transmitting side is transmitted to the receiving side.

[0021] Also, a method has been proposed in which original audio data is converted into stereo audio data (e.g., binaural sound) having an interaural time difference using a virtual microphone (hereinafter referred to as a "virtual microphone") obtained by interpolating or extrapolating a microphone (hereinafter referred to as a "real microphone") implemented in a transmitting device (e.g., a terminal) (see, for example, Non-Patent Document 1). However, this method extends the virtual microphone in one direction (e.g., one-dimensional direction, horizontal axis direction) of the line connecting the real microphones included in the stereo microphone, and it is difficult to extend the virtual microphone in accordance with the orientation of the terminal (e.g., extension in two dimensions). For example, when the orientation of the terminal changes, it is difficult to obtain appropriate sound extended to binaural distance.

[0022] In one non-limiting example embodiment of the present disclosure, a method for appropriately capturing an audio signal depending on the orientation of a terminal is described.

[0023] [Configuration Example of a Transmission System for Audio and Audio Signals] A transmission system for audio and audio signals may include, for example, a coding device (or a transmitting device or an audio capturing device) that codes audio and audio signals, and a decoding device (or a receiving device) that decodes the audio and audio signals coded by the coding device. Note that the decoding device may have the configuration of an existing decoding device, and in the following description, configuration and operation examples of the decoding device will be omitted.

[0024] [Configuration Example of Encoding Device] A configuration example of an encoding device will be described below.

[0025] 1 is a block diagram showing an example configuration of an encoding device 10. The encoding device 10 may be installed in, for example, a terminal (for example, a smartphone).

[0026] The encoding device 10 shown in FIG. 1 may include, for example, an input unit 11, an A / D conversion unit 12, a tilt sensing unit 13, a weighting coefficient determination unit 14, a virtual microphone signal generation unit 15, a stereo encoding unit 16, and a multiplexing unit 17.

[0027] The input unit 11 converts, for example, an input audio signal (for example, air vibration) into an electric signal (for example, an analog signal) and outputs the analog signal to the A / D conversion unit 12. The input unit 11 may be, for example, a MEMS microphone or a condenser microphone.

[0028] The A / D conversion unit 12 converts, for example, an analog signal input from the input unit 11 into a digital signal, and outputs the digital signal to the weighting coefficient determination unit 14 and the virtual microphone signal generation unit 15 .

[0029] In the encoding device 10, at least one of the input unit 11 and the A / D conversion unit 12 may be provided in plural (for example, two) units in order to handle stereo signals.

[0030] The tilt sensing unit 13 acquires information (tilt sensing information) relating to the tilt (or orientation) of a terminal (e.g., a smartphone) equipped with the encoding device 10, and outputs the information to the weighting coefficient determination unit 14. The tilt sensing unit 13 may acquire the tilt of the terminal (e.g., angle information indicating the tilt relative to the horizontal and vertical axes) using, for example, a tilt sensor.

[0031] The weighting coefficient determination unit 14 receives configuration information of the input unit 11 (e.g., a microphone or a real microphone) included in the terminal and tilt sensing information acquired by the tilt sensing unit 13. The real microphone configuration information may include, for example, information such as the position of the microphone physically implemented in the terminal or the number of microphones. The weighting coefficient determination unit 14 determines a weighting coefficient to be applied to a signal (e.g., an input signal of a real microphone) input from the A / D conversion unit 12 based on the real microphone configuration information and the tilt sensing information of the terminal. For example, the weighting coefficient determination unit 14 may select a real microphone to be used in generating a virtual microphone input signal (e.g., a pseudo input signal of a virtual microphone) in the virtual microphone signal generation unit 15 (described later) and determine a weighting coefficient to be applied to the input signal of the selected real microphone. Examples of a method for selecting a real microphone and a method for determining a weighting coefficient to be applied to a real microphone signal will be described later.

[0032] The virtual microphone signal generation unit 15 generates an input signal for a virtual microphone (for example, a virtual stereo microphone) by applying weighting to the signal (input signal of the real microphone) input from the A / D conversion unit 12 using the weighting coefficient input from the weighting coefficient determination unit 14. The virtual microphone signal generation unit 15 outputs the generated virtual microphone signal (for example, a "virtual stereo signal" including a "virtual L channel signal" and a "virtual R channel signal") to the stereo encoding unit 16.

[0033] The stereo encoding unit 16 performs stereo encoding using the virtual microphone signal (e.g., a virtual stereo signal) input from the virtual microphone signal generation unit 15, and outputs the encoding result to the multiplexing unit 17. For example, the audio codec used in the stereo encoding unit 16 may be Codec for Immersive Voice and Audio Services (IVAS) defined by the 3rd Generation Partnership Project (3GPP). Furthermore, the stereo encoding unit 16 is not limited to IVAS, and may include various stereo audio codecs standardized by standardization organizations such as the Moving Picture Experts Group (MPEG), 3GPP, or the International Telecommunication Union Telecommunication Standardization Sector (ITU-T).

[0034] The multiplexing unit 17 multiplexes the coded information (e.g., referred to as stereo coded information or coded stereo signal) input from the stereo coding unit 16 with other coding parameters (e.g., stereo cues such as inter-channel time difference (ITD)), and transmits the multiplexed coded information (e.g., referred to as coded data) to the decoding device via a communication network or a storage medium (not shown).

[0035] [Example of Virtual Microphone Settings] Hereinafter, an example of setting a virtual microphone (for example, an example of generating a virtual microphone signal) in the encoding device 10 (for example, the weighting coefficient determination unit 14 and the virtual microphone signal generation unit 15) will be described.

[0036] For example, in a terminal equipped with a stereo microphone (2ch) or a multi-channel microphone (3ch or more), the encoding device 10 sets a "virtual microphone L" (L-channel microphone) that virtually captures the L-channel signal of a stereo signal, and a "virtual microphone R" (R-channel microphone) that virtually captures the R-channel signal of a stereo signal. For example, the encoding device 10 multiplies each input signal of the real microphone by a weighting coefficient and adds them together to generate an input signal of the virtual microphone (virtual stereo microphone) composed of the input of the virtual microphone L and the signal of the virtual microphone R.

[0037] For example, the encoding device 10 may use the input signal of the real microphone to generate an input signal of a virtual microphone (virtual microphone L and virtual microphone R). Note that the input signal of the real microphone may be transformed into the frequency domain for each time-frequency slot, and represented by a phase component and an amplitude component. The encoding device 10 may generate a signal of a virtual microphone located at a distance between the ears by processing (e.g., interpolation processing or extrapolation processing) at least the phase component of the input signal of the real microphone. For example, a weighting coefficient for the input signal of the real microphone may be determined according to the positional relationship between the real microphone and the virtual microphone as a coefficient used in the interpolation processing or extrapolation processing.

[0038] For example, the virtual microphones L and R may be set to be arranged in the horizontal direction regardless of the inclination of the terminal 1.

[0039] By setting the virtual stereo microphones (virtual microphones L and L), the encoding device 10 can virtually generate horizontal (e.g., left-right) audio and acoustic signals and express spatial characteristics corresponding to microphones at a horizontal distance between the ears, thereby realizing immersive audio and acoustic communication.

[0040] [Setting Example 1] Fig. 2 shows an example in which three microphones A, B, and C (real microphones) are arranged on the same plane. In the two-dimensional plane in which microphones A, B, and C are arranged as shown in Fig. 2, microphone A is arranged at "position A," microphone B is arranged at "position B," and microphone C is arranged at "position C."

[0041] The microphone arrangement shown in FIG. 2 is an example, and microphones A, B, and C may be arranged at any positions on the same plane.

[0042] The encoding device 10 generates input signals for virtual microphones L and R, which are arranged horizontally on the same plane as microphones A, B, and C, based on the input signals from the three microphones A, B, and C. In Fig. 2 , virtual microphone L is set to "position L," and virtual microphone R is set to "position R."

[0043] For example, the encoding device 10 uses at least the phase components of the signals from three microphones A, B, and C arranged on a two-dimensional plane to generate signals at position P (corresponding to a first position) and position Q (corresponding to a second position) on a line connecting two virtual microphones L and R that make up a virtual stereo microphone arranged horizontally on the two-dimensional plane, and generates signals from the two virtual microphones L and R located at both ends of position P and position Q by extrapolating the signals from position P and position Q.

[0044] For example, in Fig. 2, a "virtual microphone P" is placed at position P, which is the intersection of a line connecting microphone A and microphone B with a line connecting virtual microphone L and virtual microphone R (the line connecting the two virtual microphones that make up the virtual stereo microphone). Also, for example, in Fig. 2, a "virtual microphone Q" is placed at position Q, which is the intersection of a line connecting microphone A and microphone C with a line connecting virtual microphone L and virtual microphone R. Also, in Fig. 2, point M is the midpoint of the line connecting virtual microphone P and virtual microphone Q.

[0045] Here, the ratio between line segments AP and PB is AP:PB=1-α:α, the ratio between line segments AQ and QC is AQ:QC=1-β:β, the ratio between line segments MQ and MR is MQ:MR=γ:1, and the ratio between line segments MP and ML is MP:ML=γ:1. The values ​​of α, β, and γ can vary depending on, for example, the positions of microphones A, B, and C and the positions of virtual microphones L and R.

[0046] Furthermore, the signal from microphone A is "Sa," the signal from microphone B is "Sb," the signal from microphone C is "Sc," the signal from virtual microphone L is "Sl," and the signal from virtual microphone R is "Sr." In this case, the encoding device 10 may calculate the signal Sl from virtual microphone L and the signal Sr from virtual microphone R according to the following procedure.

[0047] The signal from each microphone may be represented by, for example, phase information (phase component) "Φ" and amplitude information (amplitude component) "G." For example, the signal Sa from microphone A may be represented by phase information "Φa" and amplitude information "Ga." The same is true for the signals from the other microphones. The phase information Φ and amplitude information G are, for example, functions of time and frequency, and the phase and amplitude are respectively defined for a frequency f at time t.

[0048] <Step 1> For example, the encoding device 10 calculates (generates) a signal "Sp" (phase information Φp, amplitude information Gp) of the virtual microphone P based on the signal Sa and the signal Sb. Φp=αΦa+(1-α)Φb Gp=αGa+(1-α)Gb

[0049] As with the phase information, the amplitude information Gp may be approximated by calculation using an interpolation process (e.g., dividing the line segment AB) on the signals Sa and Sb, as shown in the above formula. Alternatively, the amplitude information Gp may be calculated as the average value of the amplitude information of the signals Sa and Sb, or may be approximated based on the amplitude information of either the signals Sa or Sb. For example, the amplitude information of the signal that is closer to the signal Sa or Sb may be set as the amplitude information Gp.

[0050] <Step 2> For example, the encoding device 10 calculates (generates) the signal “Sq” (phase information Φq, amplitude information Gq) of the virtual microphone Q based on the signal Sa and the signal Sc: Φq=βΦa+(1−β)Φc Gq=βGa+(1−β)Gc

[0051] Note that, similarly to the phase information, the amplitude information Gq may be approximated by calculation using an interpolation process (e.g., a process of dividing the line segment AC) on the signals Sa and Sc, as shown in the above formula. Alternatively, the amplitude information Gq may be calculated as the average value of the signals Sa and Sc (amplitude information), or may be approximated based on the amplitude information of either the signals Sa or Sc. For example, the amplitude information of the signal Sa or the signal Sb that is closer to the signal Sa may be set as the amplitude information Gq.

[0052] <Step 3> The encoding device 10 calculates (generates) phase information Φl of signal Sl and phase information Φq of signal Sr based on signals Sp and Sq. For example, the encoding device 10 may calculate phase information Φl and phase information Φq by extrapolating the phase information of signals Sp and Sq as follows: Φl=0.5(Φp+Φq)+0.5(1 / γ)(Φp-Φq) Φr=0.5(Φp+Φq)+0.5(1 / γ)(Φq-Φp)

[0053] Furthermore, for example, the encoding device 10 calculates (generates) amplitude information Gl of signal Sl and amplitude information Gr of signal Sr based on signal Sp and signal Sq. For example, the encoding device 10 may approximate the amplitude information Gl and Gr based on amplitude information of microphones (real microphones or virtual microphones P and Q) that are closer to positions L and R, respectively. For example, the encoding device 10 may set amplitude information Gp of position P (virtual microphone P) that is closer to position L among positions P and Q as the amplitude information Gl of signal Sl (Gl = Gp). Similarly, for example, the encoding device 10 may set amplitude information Gq of position Q (virtual microphone Q) that is closer to position R among positions P and Q as the amplitude information Gr of signal Sr (Gr = Gq).

[0054] As described above, the method of approximating the amplitude information Gl, Gr of the signals Sl, Sr using the amplitude information Gp, Gq at positions P, Q is not limited to the above. For example, the amplitude information Gl, Gr of the signals Sl, Sr may be set to a value calculated by interpolation processing of the amplitude information Gp, Gq at positions P, Q, or the average value of the amplitude information Gp, Gq at positions P, Q. Alternatively, for example, the amplitude information Gl, Gr of the signals Sl, Sr may be set to a value calculated by extrapolation processing of the amplitude information Gp, Gq at positions P, Q, as follows: In the extrapolation processing of the amplitude information Gp, Gq, for example, the value of "1 / γ" may be set to a value different from the value used in the external division of the phase information (e.g., a value closer to 1). This makes it possible to suppress a decrease in approximation accuracy even when the accuracy of approximation by external division of the amplitude information is expected to be low. Gl=0.5(Gp+Gq)+0.5(1 / γ)(Gp-Gq) Gr=0.5(Gp+Gq)+0.5(1 / γ)(Gq-Gp)

[0055] The processing of steps 1 to 3 has been explained above.

[0056] For example, by substituting the equations for Φp and Φq into the equations for Φl and Φr, the processing of steps 1 to 3 can be performed in one step, as shown in the following equation. Φl=0.5(α+β)Φa + 0.5(1-α)Φb + 0.5(1-β)Φc + (0.5 / γ)(α-β)Φa + (0.5 / γ)(1-α)Φb - (0.5 / γ)(1-β)Φc =0.5(α+β+α / γ-β / γ)Φa+0.5(1-α+1 / γ-α / γ)Φb+0.5(1-β-1 / γ+β / γ)Φc =0.5Φa(α(1+γ)-β(1-γ)) / γ+ 0.5Φb(1+γ)(1-α) / γ-0.5Φc(1-γ)(1-β) / γ =(0.5 / γ){ (α(1+γ)-β(1-γ))Φa+(1+γ)(1-α)Φb-(1-γ)(1-β)Φc} Φr=0.5(α+β)Φa + 0.5(1-α)Φb + 0.5(1-β)Φc - (0.5 / γ)(α-β)Φa - (0.5 / γ)(1-α)Φb + (0.5 / γ)(1-β)Φc =0.5(α+β-α / γ+β / γ)Φa+0.5(1-α-1 / γ+α / γ)Φb+0.5(1-β+1 / γ-β / γ)Φc =-0.5Φa(α(1-γ)-β(1+γ)) / γ-0.5Φb(1-γ)(1-α) / γ+0.5Φc(1+γ)(1-β) / γ =(0.5 / γ){ (-α(1-γ)+β(1+γ))Φa-(1-γ)(1-α)Φb+(1+γ)(1-β)Φc}

[0057] In this way, the encoding device 10 generates signals for virtual microphones P and Q, which are arranged on a horizontal line along which the virtual stereo microphones are arranged, using each real microphone input signal represented by phase information and amplitude information for each time and frequency slot. The encoding device 10 also generates signals for two virtual microphones L and R, which are located at an arbitrary distance (e.g., interaural distance), by extrapolating (or interpolating) at least the phase information of the microphone input signals for virtual microphones P and Q. In FIG. 2 , for example, the encoding device 10 can generate input signals for virtual stereo microphones (virtual microphone L and virtual microphone R) located at two arbitrary points symmetrically arranged around a certain point (e.g., point M) based on the input signals from the three real microphones.

[0058] FIG. 3 shows an example of generating input signals for virtual stereo microphones (virtual microphone L and virtual microphone R) located at a distance between the ears when three microphones A, B, and C on the same plane shown in FIG. 2 are placed (implemented) on a terminal 1 (e.g., a smartphone).

[0059] In FIG. 3, the origin O is the position of the center of gravity when the centers of microphones A, B, and C are taken as three points, and FIG. 3 shows vertical and horizontal axes passing through the origin O.

[0060] As shown in FIG. 3, virtual microphones P and Q, and virtual microphones L and R are set on a horizontal axis passing through origin O (center of gravity of multiple real microphones) on a two-dimensional plane.

[0061] For example, in the equation for calculating the signal of virtual microphone P described above, α may be set so that virtual microphone P is located at the intersection of line AB and a horizontal axis passing through origin O (e.g., a line connecting two virtual microphones L and R). Similarly, in the equation for calculating the signal of virtual microphone Q described above, β may be set so that virtual microphone Q is located at the intersection of line AC and a horizontal axis passing through origin O (e.g., a line connecting two virtual microphones L and R). This allows encoding device 10 to calculate the signals (amplitude information G and phase information Φ) of virtual microphone P and virtual microphone Q in FIG. 3 , as in FIG. 2 . Then, encoding device 10 may generate signals of virtual microphones L and R located at any distance (e.g., interaural distance) based on the signals of virtual microphone P and virtual microphone Q.

[0062] In this way, the encoding device 10 can extend the stereo representation to signals from the virtual microphones L and R at a binaural distance, for example, by setting the virtual microphones L and L (e.g., a desired pair of virtual microphones) at positions where any two points are at a binaural distance, even if the distance between the real microphones (or the size of the terminal 1) is shorter than the binaural distance. Thus, a pseudo-binaural signal can be generated from a stereo signal captured by the encoding device 10, and when a receiver listens (plays) the signal through headphones, the receiver can correctly perceive the direction (or orientation) of the sound captured by the transmitting device (e.g., the encoding device 10).

[0063] [Variation 1 of Setting Example 1] FIG. 4 shows another example of generating input signals for virtual stereo microphones (virtual microphone L and virtual microphone R) located at a distance between the ears when three microphones A, B, and C arranged on the same plane as shown in FIG. 2 are arranged (implemented) on terminal 1 (e.g., a smartphone).

[0064] 4, the midpoint M is equal to the midpoint of the line segment PQ connecting the virtual microphone P (position P) and the virtual microphone Q (position Q) shown in FIG. 3, and corresponds to the midpoint M shown in FIG.

[0065] 4, the encoding device 10 generates signals from two virtual microphones L and R that are arranged on both sides of a midpoint M in the horizontal direction in which the virtual microphones L and R are arranged, and are equidistant from the midpoint M. For example, in the formula for calculating the signals from the virtual microphones L and R described above, γ may be set so that the line segments LM and MR are equidistant and the line segment LR is the distance between the ears.

[0066] As a result, in FIG. 4, the encoding device 10 can generate signals from virtual microphones L and R located at any distance (for example, the distance between the ears) using the same equations as in FIG.

[0067] [Variation 2 of Setting Example 1] In Setting Example 1 and Variation 1, the virtual microphones L and R are placed at any two points that are symmetrically arranged around a certain point (for example, the origin O or the midpoint M), but the positions of the virtual microphones L and R are not limited to this.

[0068] FIG. 5 shows an example in which the virtual microphone R is placed at a position spaced apart from the midpoint M (for example, the midpoint of the line segment PQ) shown in FIG. 4 by an inter-aural distance to the right.

[0069] 5, the encoding device 10 generates a signal (amplitude information and phase information) of a virtual microphone R, for example, using a signal of a virtual microphone M placed at a midpoint M (position M) and a signal of a virtual microphone Q (position Q). For example, the encoding device 10 may use the signal of the virtual microphone M and the signal of the virtual microphone Q to extrapolate the virtual microphone R to a position where the distance MR is the interaural distance.

[0070] For example, the encoding device 10 may calculate the amplitude information Gm of the virtual microphone M using the amplitude information Gp of the virtual microphone P and the amplitude information Gq of the virtual microphone Q as Gm=(Gp+Gq) / 2.

[0071] Furthermore, for example, the encoding device 10 may calculate the phase information Φm of the virtual microphone M using the phase information Φp of the virtual microphone P and the phase information Φq of the virtual microphone Q as Φm=(Φp+Φq) / 2.

[0072] For example, when the ratio of the distance MQ to the distance MR is 1:ζ, the encoding device 10 may calculate the phase information Φr of the virtual microphone R as Φr=Φm+ζ(Φq−Φm) by extrapolation.

[0073] Furthermore, for example, the encoding device 10 may approximate the amplitude information Gr of the virtual microphone R based on the signal (amplitude information) of a virtual microphone (or a real microphone) that is closer to the position R. In the example of FIG. 5 , the encoding device 10 may set the amplitude information Gq of the virtual microphone Q to the amplitude information Gr of the virtual microphone R.

[0074] For example, the amplitude information Gr is not limited to approximation based on the amplitude Gq, but may be set to a value calculated by interpolation processing of the amplitude information Gp and Gq, or may be set to an average value of the amplitude information Gp and Gq (e.g., amplitude information Gm). Alternatively, for example, the amplitude information Gr may be set to a value calculated by extrapolation processing of the amplitude information Gm and Gq. For example, if the ratio of the distance MQ to the distance MR is 1:ζ, the encoding device 10 may calculate the amplitude information Gr of the virtual microphone R as Gr = Gm + ζ(Gq - Gm).

[0075] In the example of FIG. 5, the encoding device 10 may generate a signal from a virtual microphone L using a signal from a virtual microphone M placed at the midpoint M.

[0076] As shown in Figure 6, similar to Figure 5 (the case of generating a signal from virtual microphone R), the encoding device 10 may use virtual microphone M and virtual microphone P to extrapolate virtual microphone L to a position where distance ML is the distance between the ears.

[0077] In the example of Figure 6, when the ratio of the distance MP to the distance ML is 1:ζ, the encoding device 10 may calculate the phase information Φl of the virtual microphone L as Φl = Φm + ζ (Φp - Φm) by extrapolation processing.

[0078] Furthermore, for example, the encoding device 10 may approximate the amplitude information Gr of the virtual microphone L based on the signal (amplitude information) of a virtual microphone (or a real microphone) that is closer to the position L. In the example of FIG. 6 , the encoding device 10 may set the amplitude information Gp of the virtual microphone P to the amplitude information Gl of the virtual microphone L.

[0079] For example, the amplitude information Gl is not limited to approximation based on the amplitude Gp, and may be set to a value calculated by interpolation processing of the amplitude information Gp and Gq, or may be set to an average value of the amplitude information Gp and Gq (e.g., amplitude information Gm). Alternatively, for example, the amplitude information Gl may be set to a value calculated by extrapolation processing of the amplitude information Gm and Gp. For example, if the ratio of the distance MP to the distance ML is 1:ζ, the encoding device 10 may calculate the amplitude information Gl of the virtual microphone L as Gl = Gm + ζ(Gp - Gm).

[0080] In the example of FIG. 6, the encoding device 10 may generate a signal for a virtual microphone R using a signal for a virtual microphone M placed at the midpoint M.

[0081] As described above, in this variation, the encoding device 10 uses the signal of virtual microphone M, which is located at midpoint M of line segment PQ connecting virtual microphone P and virtual microphone Q, and the signal of either virtual microphone P or Q, to generate a signal of one of the two virtual microphones L and R (virtual microphone R in FIG. 5 and virtual microphone L in FIG. 6), and uses the signal of virtual microphone M to generate a signal of the other of the two virtual microphones L and R (virtual microphone L in FIG. 5 and virtual microphone R in FIG. 6). According to this variation, the encoding device 10 calculates the signal of one of virtual microphones L and R without having to calculate the signal of the other (for example, the signal of virtual microphone M may be used), thereby enabling virtual binaural sound pickup with a reduced amount of calculation (for example, extrapolation processing).

[0082] [Variation 3 of Setting Example 1] In Setting Example 1 and Variation 1, the virtual microphones L and R are placed at any two points that are symmetrically arranged around a certain point (for example, the origin O or the midpoint M), but the positions of the virtual microphones L and R are not limited to this.

[0083] FIG. 7 shows an example in which the virtual microphone R is placed at a position inter-aural to the right of the position P shown in FIG. 4 as the starting point.

[0084] 5, the encoding device 10 uses, for example, a signal from a virtual microphone P (position P) and a signal from a virtual microphone Q (position Q) to generate a signal (amplitude information and phase information) from a virtual microphone R. For example, the encoding device 10 may use the signal from the virtual microphone P and the signal from the virtual microphone Q to extrapolate the virtual microphone R to a position where the distance PR is the interaural distance.

[0085] For example, when the ratio of the distance PQ to the distance PR is 1:η, the encoding device 10 may calculate the phase information Φr of the virtual microphone R as Φr=Φp+η(Φq−Φp) by extrapolation.

[0086] Furthermore, for example, the encoding device 10 may approximate the amplitude information Gr of the virtual microphone R based on the signal (amplitude information) of a virtual microphone (or a real microphone) that is closer to the position R. In the example of FIG. 7 , the encoding device 10 may set the amplitude information Gq of the virtual microphone Q to the amplitude information Gr of the virtual microphone R.

[0087] For example, the amplitude information Gr is not limited to approximation based on the amplitude Gq, but may be set to a value calculated by interpolation processing of the amplitude information Gp and Gq, or may be set to the average value of the amplitude information Gp and Gq. Alternatively, for example, the amplitude information Gr may be set to a value calculated by extrapolation processing of the amplitude information Gp and Gq. For example, if the ratio of the distance PQ to the distance PR is 1:η, the encoding device 10 may calculate the amplitude information Gr of the virtual microphone R as Gr = Gp + η(Gq - Gp).

[0088] In the example of FIG. 7, the encoding device 10 may generate a signal for the virtual microphone L using a signal for the virtual microphone P.

[0089] As shown in Figure 8, similar to Figure 7 (the case of generating a signal from virtual microphone R), the encoding device 10 may use virtual microphone P and virtual microphone Q to extrapolate virtual microphone L to a position where distance QL is the distance between the two ears.

[0090] Also, in the example of Figure 8, if the ratio of the distance QP to the distance QL is 1:η, the encoding device 10 may calculate the phase information Φl of the virtual microphone L as Φl = Φq + η(Φp - Φq) by extrapolation processing.

[0091] Furthermore, for example, the encoding device 10 may approximate the amplitude information Gr of the virtual microphone L based on the signal (amplitude information) of a virtual microphone (or a real microphone) that is closer to the position L. In the example of FIG. 8 , the encoding device 10 may set the amplitude information Gp of the virtual microphone P to the amplitude information Gl of the virtual microphone L.

[0092] For example, the amplitude information Gl is not limited to approximation based on the amplitude Gp, and may be set to a value calculated by interpolation processing of the amplitude information Gp and Gq, or may be set to the average value of the amplitude information Gp and Gq. Alternatively, for example, the amplitude information Gl may be set to a value calculated by extrapolation processing of the amplitude information Gq and Gp. For example, if the ratio of the distance QP to the distance QL is 1:η, the encoding device 10 may calculate the amplitude information Gl of the virtual microphone L as Gl = Gq + η(Gp - Gq).

[0093] In the example of FIG. 8, the encoding device 10 may generate a signal for the virtual microphone R using a signal for the virtual microphone Q.

[0094] As described above, in this variation, the encoding device 10 uses the signal of the virtual microphone P and the signal of the virtual microphone Q to generate a signal for one of the two virtual microphones L and R (virtual microphone R in FIG. 7 and virtual microphone L in FIG. 8 ), and uses the signal of the virtual microphone (virtual microphone P in FIG. 7 and virtual microphone Q in FIG. 8 ) that is farther from the generated one of the virtual microphones P and Q (virtual microphone R in FIG. 7 and virtual microphone L in FIG. 8 ) to generate a signal for the other of the two virtual microphones L and R (virtual microphone L in FIG. 7 and virtual microphone R in FIG. 8 ). According to this variation, the encoding device 10 calculates the signal for one of the virtual microphones L and R, but does not need to calculate the signal for the other, and does not need to define the midpoint M between P and Q. This reduces the amount of calculation and enables virtual binaural sound pickup.

[0095] 7 has described an example in which virtual microphone P approximates virtual microphone L that is aurally distant from virtual microphone R, but the present invention is not limited to this, and the virtual microphone (position) that approximates virtual microphone L that is aurally distant from virtual microphone R may be virtual microphone Q. That is, encoding device 10 may generate virtual microphone R that is positioned aurally distant from virtual microphone Q. The same applies to FIG. 8.

[0096] [Variation 4 of Setting Example 1] In the example described in Setting Example 1, the origin O shown in FIG. 3 may be set to the center of the terminal 1.

[0097] Fig. 9 shows an example in which the origin O is set to the center of terminal 1. For example, in Fig. 9, the positions of the centers of gravity of antennas A, B, and C are set to the center of terminal 1 in the two-dimensional plane on which antennas A, B, and C are arranged.

[0098] 9, auxiliary lines are drawn passing through the midpoints of the vertical and horizontal sides of the terminal 1 to indicate that the origin O is set at the center of the terminal 1. Also, in FIG. 9, vertical and horizontal axis lines are drawn passing through the origin O.

[0099] As shown in FIG. 9, the virtual microphones P and Q, and the virtual microphones L and R are set (placed) on a horizontal axis passing through the origin O.

[0100] In the example shown in Figure 9, the encoding device 10 can extrapolate virtual microphones L and R located at a distance between the ears using virtual microphones P and Q, as in setting example 1 or its variations.

[0101] Furthermore, according to this variation, the virtual microphones are set evenly on the left and right sides on a horizontal axis based on the center of the terminal 1 (e.g., a smartphone), making it easier for the user to grasp the position of the virtual microphones. This makes it easier for the user to intuitively position the virtual microphones, and when used simultaneously with a camera, it is possible to pick up sound from a position that matches the video.

[0102] [Variation 5 of Setting Example 1] The placement of the real microphones is not limited to the placement shown in the setting example and other variations described above, and other placements may also be used.

[0103] For example, FIG. 10 shows an example in which microphones A, B, and C are positioned in the upper half of the horizontal axis passing through the origin O set in the same manner as in FIG.

[0104] As shown in FIG. 10 , if there is no intersection on the horizontal axis passing through the origin O at a position that divides the line connecting the real microphones internally, the encoding device 10 may generate the virtual microphones P and Q by extrapolation.

[0105] 10 , for example, the encoding device 10 uses microphone A and microphone B to extrapolate virtual microphone P to the intersection of line AB and the horizontal axis passing through origin O. Similarly, the encoding device 10 uses microphone A and microphone C to extrapolate virtual microphone Q to the intersection of line AC and the horizontal axis passing through origin O.

[0106] Note that at least the phase information (phase components) of the signals of the virtual microphones P and Q may be generated by extrapolation. For example, the amplitude information of the signals of the virtual microphones P and Q may be set to a value calculated by interpolation processing of the amplitude information of the real microphones used to set each of the virtual microphones P and Q, or the average value of the amplitude information of the real microphones used to set each of the virtual microphones P and Q may be set, or the amplitude information of any real microphone (for example, the real microphone closest to the positions of the virtual microphones P and Q) may be set. Alternatively, for example, the amplitude information of the virtual microphones P and Q may be set to a value calculated by extrapolation processing of the amplitude information of the real microphones used to set each of the virtual microphones P and Q.

[0107] The encoding device 10 may use the generated signals of virtual microphone P and virtual microphone Q to generate (e.g., extrapolate) signals of virtual microphone L and virtual microphone R in a manner similar to the method described above.

[0108] 10 , encoding device 10 can determine virtual microphones L and R by generating virtual microphones P and Q by extrapolating microphones A, B, and C. That is, if there is no intersection on the horizontal axis passing through origin O at a position that divides the line connecting the real microphones internally, encoding device 10 may generate at least one of virtual microphones P and Q by extrapolation, and if there is an intersection on the horizontal axis passing through origin O at a position that divides the line connecting the real microphones internally, encoding device 10 may generate at least one of virtual microphones P and Q by interpolation.

[0109] According to this variation, the encoding device 10 can set a virtual microphone at a specific position from three microphones at any positions.

[0110] [Setting Example 2] In setting example 2, for example, as shown in FIG. 11 , if the direction of a line connecting any two of the multiple real microphones arranged on terminal 1 is parallel to the horizontal axis of terminal 1 (or the direction of a line connecting virtual microphones L and R), the encoding device 10 generates signals for virtual microphones L and R using the signals of the two real microphones (for example, extrapolates virtual microphones L and R).

[0111] In addition, in FIG. 11, a horizontal axis line passing through the center of the terminal 1 (for example, a smartphone) is drawn.

[0112] For example, the example of Fig. 11 shows an example in which microphones A, B, and C are arranged (implemented) on terminal 1 (e.g., a smartphone), similar to Fig. 3. Fig. 11 differs from Fig. 3 in that the direction of line BC is parallel to the horizontal axis of terminal 1.

[0113] For example, as shown in Fig. 11 , virtual microphones L and R are arranged on a straight line (e.g., line BC) that passes through microphones B and C. In the example of Fig. 11 , encoding device 10 may extrapolate virtual microphones L and R using signals from real microphones B and C, similar to the method of extrapolating virtual microphones L and R using signals from virtual microphones P and Q in setting example 1 (or its variation).

[0114] In FIG. 11, the midpoint M of the line segment connecting microphone B and microphone C passes through the position B of microphone B and the position C of microphone C, and is on a straight line parallel to the horizontal axis of terminal 1.

[0115] For example, the encoding device 10 uses signals from microphones B and C to extrapolate virtual microphones L and R to positions where line segments LM and MR are equidistant and line segment LR is the distance between the ears.

[0116] For example, the encoding device 10 may calculate the amplitude information Gm of the virtual microphone M placed at the midpoint M as Gm=(Gb+Gc) / 2, where Gb represents the amplitude information of the microphone B, and Gc represents the amplitude information of the microphone C.

[0117] Furthermore, for example, the encoding device 10 may calculate the phase information Φm of the virtual microphone M placed at the midpoint M as Φm=(Φb+Φc) / 2, where Φb represents the phase information of the microphone B, and Φc represents the phase information of the microphone C.

[0118] 11, when the ratio of the distances MC and MR is 1:ε, the encoding device 10 may calculate the phase information Φr of the virtual microphone R as Φr=Φm+ε(Φc-Φm) by extrapolation. Similarly, the phase information Φl of the virtual microphone L may be calculated as Φl=Φm+ε(Φb-Φm) by extrapolation.

[0119] Furthermore, for example, the encoding device 10 calculates (generates) amplitude information Gl of signal Sl and amplitude information Gr of signal Sr based on signal Sb and signal Sc. For example, the encoding device 10 may approximate the amplitude information Gl and Gr based on amplitude information of microphones closer to positions L and R, respectively. For example, the encoding device 10 may set amplitude information Gb of position B (microphone B) closer to position L of positions B and C as the amplitude information Gl of signal Sl (Gl = Gb). Similarly, for example, the encoding device 10 may set amplitude information Gc of position C (microphone C) closer to position R of positions B and C as the amplitude information Gr of signal Sr (Gr = Gc).

[0120] For example, the amplitude information Gl and Gr are not limited to approximations based on the amplitudes Gb and Gc. Values ​​calculated by interpolation of the amplitude information Gb and Gc may be set, or the average value of the amplitude information Gb and Gc (e.g., amplitude information Gm) may be set. Alternatively, for example, values ​​calculated by extrapolation of the amplitude information Gb and Gc may be set as the amplitude information Gl and Gr. For example, if the ratio of the distance MC to the distance MR is 1:ε, the encoding device 10 may calculate the amplitude information Gr of the virtual microphone R as Gr = Gm + ε(Gc - Gm). Similarly, the amplitude information Gl of the virtual microphone L may be calculated as Gl = Gm + ε(Gb - Gm).

[0121] Thus, in setting example 2, when at least two of the real microphones arranged on terminal 1 are arranged parallel to the horizontal axis of terminal 1, virtual microphones L and R (e.g., a desired pair of virtual microphones) are set using these two real microphones. For example, as shown in FIG. 11 , the encoding device 10 generates virtual microphones L and R by extrapolating to both sides of real microphones B and C. In setting example 2, when setting virtual microphones L and R, it is not necessary to use signals from real microphones other than the pair of real microphones arranged parallel to the horizontal axis of terminal 1. This allows the encoding device 10 to calculate the outputs of two virtual microphones with a smaller amount of calculation than, for example, calculating the outputs of two virtual microphones from three real microphones.

[0122] Furthermore, according to setting example 2, for example, it is possible to set virtual microphones evenly spaced on the left and right sides of the horizontal axis based on the center of terminal 1 (e.g., a smartphone), making it easier for the user to grasp the position of the virtual microphone. This makes it easier for the user to intuitively position the virtual microphone, and when used simultaneously with a camera, it is possible to pick up sound from a position that matches the video.

[0123] Furthermore, according to setting example 2, the virtual microphones L and R are generated based on the signals (phase information and amplitude information) of real microphones (e.g., microphones B and C). Therefore, compared to generating the signals of the virtual microphones L and R based on the virtual microphones P and Q generated from the signals of the real microphones, for example, errors due to calculations can be reduced, which may improve the quality of the virtual microphones L and R.

[0124] In addition, if there is no pair of real microphones arranged parallel to the horizontal axis of the terminal 1, the encoding device 10 may generate a virtual stereo microphone, for example, according to setting example 1 (or setting example 3 described below).

[0125] [Variation 1 of Setting Example 2] In Setting Example 2, the virtual microphones L and R are placed at any two points that are symmetrically arranged around a certain point (for example, the midpoint M), but the positions of the virtual microphones L and R are not limited to this.

[0126] FIG. 12 shows an example in which a virtual microphone R is placed at a position spaced apart from the ears on the right side of position M (for example, the midpoint of line segment BC) shown in FIG. 11 as a starting point.

[0127] 12 , the encoding device 10 generates a signal (amplitude information and phase information) of a virtual microphone R using, for example, a signal of a virtual microphone M placed at a midpoint M and a signal of a microphone C. For example, the encoding device 10 may use the signal of the virtual microphone M and the signal of the microphone C to extrapolate the virtual microphone R to a position where the distance MR is the interaural distance.

[0128] For example, the encoding device 10 may calculate the amplitude information Gm of the virtual microphone M using the amplitude information Gb of the microphone B and the amplitude information Gc of the microphone C as Gm=(Gb+Gc) / 2.

[0129] Furthermore, for example, the encoding device 10 may calculate the phase information Φm of the virtual microphone M using the phase information Φb of the microphone B and the phase information Φc of the microphone C as Φm=(Φb+Φc) / 2.

[0130] For example, when the ratio of the distance MC to the distance MR is 1:κ, the encoding device 10 may calculate the phase information Φr of the virtual microphone R as Φr=Φm+κ(Φc−Φm) by extrapolation.

[0131] Furthermore, for example, the encoding device 10 may approximate the amplitude information Gr of the virtual microphone R based on the signal (amplitude information) of a microphone closer to the position R. In the example of FIG. 12 , the encoding device 10 may set the amplitude information Gq of the virtual microphone Q to the amplitude information Gr of the virtual microphone R.

[0132] For example, the amplitude information Gr is not limited to approximation based on the amplitude Gc, but may be set to a value calculated by interpolation processing of the amplitude information Gb and Gc, or may be set to an average value of the amplitude information Gb and Gc (e.g., amplitude information Gm). Alternatively, for example, the amplitude information Gr may be set to a value calculated by extrapolation processing of the amplitude information Gm and Gc. For example, if the ratio of the distance MC to the distance MR is 1:κ, the encoding device 10 may calculate the amplitude information Gr of the virtual microphone R as Gr = Gm + κ(Gc - Gm).

[0133] In the example of FIG. 12, the encoding device 10 may generate a signal from a virtual microphone L using a signal from a virtual microphone M placed at the midpoint M.

[0134] 12 , a virtual microphone L may be generated at a position separated by an interaural distance to the left of the midpoint M (not shown). In this case, for example, if the ratio of the distance MB to the distance ML is 1:κ, the encoding device 10 may calculate the phase information Φl of the virtual microphone L as Φl = Φm + κ(Φb - Φm). The encoding device 10 may also set the amplitude information Gb of the microphone B to the amplitude information Gl of the virtual microphone L. The encoding device 10 may also generate a signal for the virtual microphone R using a signal for the virtual microphone M located at the midpoint M.

[0135] As described above, in this variation, the encoding device 10 generates a signal for one of the two virtual microphones L and R (virtual microphone R in FIG. 12 ) using a signal from virtual microphone M located at midpoint M of line segment BC connecting two microphones B and C and a signal from either of the two microphones B and C, and generates a signal for the other of the two virtual microphones L and R (virtual microphone L in FIG. 12 ) using the signal from virtual microphone M. According to this variation, the encoding device 10 calculates the signal from one of virtual microphones L and R without having to calculate the signal from the other (for example, the signal from virtual microphone M may be used), thereby reducing the amount of calculation (for example, extrapolation processing) and enabling virtual binaural sound pickup.

[0136] [Variation 2 of Setting Example 2] In Setting Example 2, the virtual microphones L and R are placed at any two points that are symmetrically arranged around a certain point (for example, the midpoint M), but the positions of the virtual microphones L and R are not limited to this.

[0137] FIG. 13 shows an example in which the virtual microphone R is placed at a position at an inter-aural distance to the right of position B shown in FIG. 11 as the starting point.

[0138] 13 , the encoding device 10 generates a signal (amplitude information and phase information) of a virtual microphone R using, for example, a signal from microphone B and a signal from microphone C. For example, the encoding device 10 may use the signal from microphone B and the signal from microphone C to extrapolate the virtual microphone R to a position where the distance BR is the interaural distance.

[0139] For example, when the ratio of the distance BC to the distance BR is 1:λ, the encoding device 10 may calculate the phase information Φr of the virtual microphone R as Φr=Φb+λ(Φc−Φb) by extrapolation.

[0140] Furthermore, for example, the encoding device 10 may approximate the amplitude information Gr of the virtual microphone R based on the signal (amplitude information) of a microphone closer to the position R. In the example of Fig. 13 , the encoding device 10 may set the amplitude information Gc of the microphone C to the amplitude information Gr of the virtual microphone R.

[0141] For example, the amplitude information Gr is not limited to approximation based on the amplitude Gc, but may be set to a value calculated by interpolation processing of the amplitude information Gb and Gc, or may be set to the average value of the amplitude information Gb and Gc. Alternatively, for example, the amplitude information Gr may be set to a value calculated by extrapolation processing of the amplitude information Gb and Gc. For example, if the ratio of the distance BC to the distance BR is 1:λ, the encoding device 10 may calculate the amplitude information Gr of the virtual microphone R as Gr = Gb + λ(Gc - Gb).

[0142] In the example of FIG. 13, the encoding device 10 may generate a signal for the virtual microphone L using a signal for the microphone B.

[0143] 13 , a virtual microphone L may be generated at a position a distance between the ears to the left of position C (not shown). In this case, for example, if the ratio of distance CB to distance CL is 1:λ, the encoding device 10 may calculate phase information Φl of virtual microphone L as Φl=Φc+λ(Φb-Φc). The encoding device 10 may also set amplitude information Gb of microphone B to amplitude information Gl of virtual microphone L. The encoding device 10 may also generate a signal for virtual microphone R using a signal from microphone C.

[0144] As described above, in this variation, the encoding device 10 uses signals from two microphones B and C to generate a signal for one of the two virtual microphones L and R (virtual microphone R in FIG. 13 ), and uses a signal from the microphone (microphone B in FIG. 13 ) that is farther from the generated one of the two virtual microphones B and C (virtual microphone R in FIG. 13 ) to generate a signal for the other of the two virtual microphones L and R (virtual microphone L in FIG. 13 ). According to this variation, the encoding device 10 calculates the signal for one of the virtual microphones L and R, but does not need to calculate the signal for the other, and does not need to define the midpoint M between B and C. This reduces the amount of calculation and enables virtual binaural sound pickup.

[0145] 13 has described an example in which the virtual microphone L, which is located at an aural distance from the virtual microphone R, is approximated by the microphone B. However, the present invention is not limited to this, and the microphone (position) that approximates the virtual microphone L, which is located at an aural distance from the virtual microphone R, may be the microphone C. In other words, the encoding device 10 may generate the virtual microphone R that is located at a position that is located at an aural distance from the microphone C.

[0146] [Setting Example 3] In setting example 3, for example, as shown in FIG. 14 , a case will be described in which microphones A, B, and C are arranged (implemented) on terminal 1 (e.g., a smartphone) and positioned on a straight line and on a vertical axis (or in a direction perpendicular to the direction in which virtual microphones L and R are arranged).

[0147] In FIG. 14, microphone A is placed at position A, microphone B is placed at position B, and microphone C is placed at position C.

[0148] 14 may be the center of the terminal 1. Alternatively, the origin O may be the midpoint of any two points (for example, the positions of two of the microphones A, B, and C) or the center of gravity of three points (for example, the positions of the microphones A, B, and C).

[0149] For example, the encoding device 10 may repeatedly perform interpolation or extrapolation of virtual microphones based on any two of the three microphones A, B, and C, to generate a virtual microphone Ov that is located at the origin O (for example, the intersection of the line connecting virtual microphones L and R with the line connecting microphones A, B, and C). In this case, the virtual microphone Ov is equivalent to the virtual microphone P and the virtual microphone Q (for example, expressed as Ov=P=Q).

[0150] Alternatively, the encoding device 10 may generate a virtual microphone Ov at the origin O using signals at any two points obtained based on the microphones A, B, and C. Note that the method of interpolating the virtual microphone may be the same as the setting example or variation described above.

[0151] 14, the encoding device 10 may generate virtual microphones L and R at a distance between the ears by extrapolating the virtual microphone Ov to both sides in the horizontal direction. For example, the encoding device 10 may set the signal of the virtual microphone Ov to the signals of the virtual microphones L and R (e.g., expressed as L=R=Ov). That is, when the microphones A, B, and C are arranged on the vertical axis, the encoding device 10 performs mono downmix processing using the signals of the microphones A, B, and C.

[0152] Note that in setting example 3, the method of generating the virtual microphones L and R is not limited to the above-described method. For example, when the orientation of the terminal 1 changes over time, the encoding device 10 may set a phase difference between the virtual microphones L and R based on a temporal change in the signal (e.g., phase information) of the virtual microphone Ov. This allows the encoding device 10 to generate signals of the virtual microphones L and R having a phase difference according to the orientation of the terminal 1, using the signal of the virtual microphone Ov.

[0153] According to setting example 3, when multiple real microphones are arranged in a vertical direction (a direction perpendicular to the direction in which virtual microphones L and R are arranged), encoding device 10 generates virtual microphones L and R using virtual microphone Ov. For example, when the direction in which multiple real microphones are arranged (e.g., line AC in FIG. 14 ) is vertical, the left-right (or horizontal) spatial characteristics obtainable from the real microphones are reduced, and the effect of improving the spatial characteristics that can be represented by virtual microphones is reduced. Therefore, when the direction in which multiple real microphones are arranged is vertical, encoding device 10 can reduce the amount of calculation in encoding device 10 by obtaining a downmixed monaural signal.

[0154] For example, the encoding device 10 may periodically (e.g., periodically) acquire (track) the phase difference between the virtual microphone Ov and the virtual microphone L, and the phase difference between the virtual microphone Ov and the virtual microphone R. When the orientation of the terminal 1 rotates to the state shown in FIG. 14 , the encoding device 10 may generate phase information for the virtual microphones L and R from the signal of the virtual microphone Ov in the state shown in FIG. 14 using past (e.g., immediately preceding) phase difference information for the virtual microphones Ov and L and past (e.g., immediately preceding) phase difference information for the virtual microphones Ov and R. This allows the encoding device 10 to virtually generate horizontal stereo microphone signals even at times when horizontal (left-right) information cannot be obtained, as shown in FIG. 14 . For example, when the orientation of the terminal 1 changes and passes through the state shown in FIG. 14 , the direction of the sound may change instantaneously in the state shown in FIG. 14 (a state where horizontal information cannot be obtained), which may result in discontinuity in the sound direction. In contrast, according to the method described above, even if the orientation of the terminal 1 changes and passes through the state shown in FIG. 14, the encoding device 10 can maintain the continuity of the sound direction from the virtual stereo microphone, thereby maintaining the stereophonic effect.

[0155] Note that if the angle θ between the vertical axis and a line (e.g., line AC) connecting microphones A, B, and C is equal to or less than a threshold (e.g., 15 degrees or less), the encoding device 10 may assume that line AC is vertical (e.g., the state shown in FIG. 14 ) and calculate the virtual microphones L and R using the method described above. For example, the closer line AC is to vertical, the more the left-right (or horizontal) spatial characteristics obtainable from real microphones A, B, and C are reduced, reducing the effect of improving the spatial characteristics representable by the virtual microphones. Therefore, if line AC is close to vertical (i.e., if angle θ is equal to or less than a threshold), the encoding device 10 may obtain a downmixed mono signal by assuming that line AC is vertical (or that real microphones A, B, and C are positioned on the vertical axis). This reduces the amount of calculations performed by the encoding device 10.

[0156] Furthermore, even if the tilt of terminal 1 rotates left and right on the vertical axis and wobbles due to hand shake or the like, encoding device 10 performs stereo encoding assuming that microphones A, B, and C are positioned on the vertical axis (i.e., positioned in fixed positions), thereby suppressing variation in the stereo encoding results and improving sound stability.

[0157] The above describes an example of setting up a virtual microphone.

[0158] Thus, in one embodiment of the present disclosure, the encoding device 10 generates signals from two virtual microphones (e.g., virtual microphones L and R) that constitute a virtual stereo microphone arranged horizontally on a two-dimensional plane on which multiple real microphones are arranged, based on the signals from the multiple real microphones, and performs stereo encoding using the signals from the two virtual microphones L and R.

[0159] As a result, even if the tilt of terminal 1 changes, the encoding device 10 can generate (or convert, or estimate) a signal from a virtual stereo microphone arranged horizontally from a signal from a real microphone captured when terminal 1 is tilted, and can therefore appropriately capture an audio signal according to the orientation of terminal 1.

[0160] In this way, by introducing a virtual microphone, the encoding device 10 can obtain immersive sound using an audio signal captured from a real microphone located at any position. For example, the encoding device 10 can generate a signal from a virtual microphone located at any position, thereby obtaining a wide range of stereo audio and sound.

[0161] Furthermore, since the encoding device 10 can generate signals from virtual stereo microphones arranged horizontally regardless of the orientation of the terminal 1, it can properly capture audio and sound signals without increasing the number of microphones that the terminal 1 has.

[0162] Therefore, according to an embodiment of the present disclosure, audio and sound signals can be captured appropriately.

[0163] Furthermore, for example, by generating a virtual microphone signal in the encoding device 10 (transmitting device) where the microphone distance is the distance between both ears, i.e., by converting the microphone distance corresponding to the stereo signal to be stereo encoded into the distance between both ears, the data format is unified on the transmitting side, so that the receiving side does not need to convert the data from the transmitting side individually, thereby simplifying the processing on the receiving side.

[0164] The embodiments of the present disclosure have been described above.

[0165] In the above-described embodiment, a method for generating virtual stereo microphones L and R using three real microphones arranged on the same plane has been described. For example, although examples in which virtual microphones are extrapolated to the distance between the ears have been shown in Figs. 3 to 14, the present invention is not limited to this. The encoding device 10 can set the virtual stereo microphones L and R at any distance different from the distance between the ears by setting (e.g., changing) a weighting factor. This allows the encoding device 10 to generate virtual stereo microphones at any position and distance depending on the weighting factor setting, thereby enabling a wide range of stereo audio signals to be acquired.

[0166] Furthermore, in the above-described embodiment, the virtual stereo microphones (e.g., virtual microphones L and R) are arranged horizontally, but the direction in which the virtual microphones constituting the virtual stereo microphone are arranged is not limited to the horizontal direction and may be other directions.

[0167] Furthermore, in the above-described embodiment, interpolation may be applied in addition to extrapolation to set the virtual stereo microphones. For example, if the distance between virtual microphone P and virtual microphone Q is greater than the distance between the ears, encoding device 10 can determine virtual microphones L and R by interpolating using virtual microphones P and Q. In this way, encoding device 10 can also generate virtual microphones at a distance shorter than the distance between virtual microphones P and Q.

[0168] Furthermore, the positions of the real microphones and the virtual microphones described in the above embodiments are merely examples, and the real microphones may be placed at other positions within the mobile terminal, and the virtual microphones may be set at other positions. Furthermore, the number of real microphones is not limited to three, and may be four or more.

[0169] Furthermore, in the above embodiments, if the signal of a virtual microphone located at a desired distance cannot be obtained directly from the input signal of a real microphone, encoding device 10 may, for example, generate a virtual microphone signal from the input signal of any real microphone, and multiply the generated virtual microphone signal by a weighting factor in the same manner as in the above embodiments, thereby generating the signal of a virtual microphone located at the desired distance. Furthermore, encoding device 10 may, for example, generate a different virtual microphone signal from the input signal of any real microphone and the generated virtual microphone signal, and multiply the generated virtual microphone signal by a weighting factor in the same manner as in the above embodiments, thereby generating the signal of a virtual microphone located at the desired distance.

[0170] Although the above embodiment has been described with reference to a virtual stereo microphone (e.g., a two-channel microphone), the virtual microphone generated from the real microphone may be a virtual multi-channel microphone with three or more channels. The encoding device 10 may generate a virtual multi-channel microphone signal for each virtual microphone constituting the virtual multi-channel microphone in the same manner as for the virtual microphones L and R described above.

[0171] Although various embodiments have been described above with reference to the drawings, it goes without saying that the present disclosure is not limited to such examples. Furthermore, the components in the above-described embodiments may be combined in any manner.

[0172] Furthermore, the notation "... section" in the above-described embodiments may be replaced with other notations such as "... circuitry," "... device," "... unit," or "... module."

[0173] The present disclosure can be realized by software, hardware, or software in conjunction with hardware. Each functional block used in the description of the above embodiments may be partially or entirely realized as an LSI, which is an integrated circuit, and each process described in the above embodiments may be partially or entirely controlled by a single LSI or a combination of LSIs. The LSI may be composed of individual chips, or may be composed of a single chip that includes some or all of the functional blocks. The LSI may have data input and output. Depending on the degree of integration, the LSI may also be called an IC, system LSI, super LSI, or ultra LSI.

[0174] The integrated circuit method is not limited to LSI, and may be realized by a dedicated circuit, a general-purpose processor, or a dedicated processor. Also, a field programmable gate array (FPGA) that can be programmed after LSI manufacturing, or a reconfigurable processor that can reconfigure the connections and settings of circuit cells within the LSI, may be used. The present disclosure may be realized as digital processing or analog processing.

[0175] Furthermore, if an integrated circuit technology that can replace LSI emerges due to advances in semiconductor technology or other derivative technologies, it is natural that such technology may be used to integrate functional blocks. The application of biotechnology, etc. is also a possibility.

[0176] The present disclosure may be implemented in any type of apparatus, device, or system (collectively referred to as a communications apparatus) that has a communications function. The communications apparatus may include a radio transceiver and processing / control circuitry. The radio transceiver may include a receiver and a transmitter, or both functions. The radio transceiver (transmitter and receiver) may include a radio frequency (RF) module and one or more antennas. The RF module may include an amplifier, an RF modulator / demodulator, or the like. Non-limiting examples of communication devices include telephones (e.g., cell phones, smartphones), tablets, personal computers (PCs) (e.g., laptops, desktops, notebooks), cameras (e.g., digital still / video cameras), digital players (e.g., digital audio / video players), wearable devices (e.g., wearable cameras, smartwatches, tracking devices), game consoles, digital book readers, telehealth / telemedicine devices, communication-enabled vehicles or mobile transportation (e.g., cars, airplanes, ships), and combinations of the above devices.

[0177] The communication devices are not limited to portable or mobile devices, but also include any kind of non-portable or fixed equipment, devices, and systems, such as smart home devices (such as home appliances, lighting equipment, smart meters or measuring devices, control panels, etc.), vending machines, and any other "things" that may exist on an IoT (Internet of Things) network.

[0178] Communications include data communications using cellular systems, wireless local area networks (LANs), communication satellite systems, and the like, as well as data communications using combinations of these.

[0179] A communications apparatus also includes devices such as controllers and sensors connected or coupled to a communications device that performs the communications functions described in this disclosure, such as controllers and sensors that generate control and data signals used by the communications device to perform the communications functions of the communications apparatus.

[0180] The communication apparatus also includes infrastructure facilities, such as base stations, access points, and any other apparatus, device, or system that communicates with or controls the various apparatuses listed above, but are not limited to these.

[0181] An encoding device according to one embodiment of the present disclosure includes a generation unit that uses at least phase components of signals from at least three microphones arranged on a two-dimensional plane to generate signals for a first position and a second position on a line connecting two virtual microphones that constitute a virtual stereo microphone arranged in a first direction on the two-dimensional plane, and generates signals for the two virtual microphones located at both ends of the first position and the second position by extrapolating the signals from the first position and the second position; and an encoding unit that performs stereo encoding using the signals from the two virtual microphones.

[0182] In one embodiment of the present disclosure, the at least three microphones include a first microphone, a second microphone, and a third microphone, the first position is an intersection of a line connecting the two virtual microphones and a line connecting the first microphone and the second microphone, and the second position is an intersection of a line connecting the two virtual microphones and a line connecting the first microphone and the third microphone.

[0183] In one embodiment of the present disclosure, the generation unit generates the signals for the first position and the second position by interpolation or extrapolation of at least the phase components of the at least three microphones.

[0184] In one embodiment of the present disclosure, the generation unit generates signals of the two virtual microphones located equidistant from a midpoint of a line segment connecting the first position and the second position in the first direction.

[0185] In one embodiment of the present disclosure, the generation unit generates a signal for one of the two virtual microphones using a signal at a third position, which is the midpoint of a line segment connecting the first position and the second position, and a signal at either the first position or the second position, and generates a signal for the other of the two virtual microphones using the signal at the third position.

[0186] In one embodiment of the present disclosure, the generation unit generates a signal for one of the two virtual microphones using a signal at the first position and a signal at the second position, and generates a signal for the other of the two virtual microphones using a signal at a position farther from the one of the virtual microphones from among the first position and the second position.

[0187] In one embodiment of the present disclosure, the line connecting the two virtual microphones is a line passing through the centers of gravity of the at least three microphones.

[0188] In one embodiment of the present disclosure, the position of the center of gravity is the center of a terminal equipped with the encoding device in the two-dimensional plane.

[0189] In one embodiment of the present disclosure, the distance between the two virtual microphones is set to the interaural distance.

[0190] In one embodiment of the present disclosure, the first direction is a horizontal direction.

[0191] In an encoding method according to one embodiment of the present disclosure, an encoding device uses at least phase components of signals from at least three microphones arranged on a two-dimensional plane to generate signals at a first position and a second position on a line connecting two virtual microphones that constitute a virtual stereo microphone arranged in a first direction on the two-dimensional plane, performs an extrapolation process on the signals at the first position and the second position to generate signals for the two virtual microphones located at both ends of the first position and the second position, and performs stereo encoding using the signals from the two virtual microphones.

[0192] The disclosures of the U.S. provisional application No. 63 / 623,364 filed on January 22, 2024, and the specification, drawings, and abstract of the Japanese patent application No. 2024-215338 filed on December 10, 2024, are incorporated herein by reference in their entirety.

[0193] An embodiment of the present disclosure is useful for coding systems and the like.

[0194] REFERENCE SIGNS LIST 10 Encoding device 11 Input unit 12 A / D conversion unit 13 Tilt sensing unit 14 Weighting coefficient determination unit 15 Virtual microphone signal generation unit 16 Stereo encoding unit 17 Multiplexing unit

Claims

1. A coding device comprising: a generation unit that generates signals at a first position and a second position on a straight line connecting two virtual microphones that constitute a virtual stereo microphone arranged in a first direction of the two-dimensional plane, using at least the phase components of signals of at least three microphones arranged in the two-dimensional plane, and generates signals of the two virtual microphones located at both ends of the first position and the second position by an extrapolation process for the signals at the first position and the second position; and a coding unit that performs stereo coding using the signals of the two virtual microphones.

2. The coding device according to claim 1, wherein the at least three microphones include a first microphone, a second microphone, and a third microphone, the first position is an intersection of the straight line connecting the two virtual microphones and the straight line connecting the first microphone and the second microphone, and the second position is an intersection of the straight line connecting the two virtual microphones and the straight line connecting the first microphone and the third microphone.

3. The coding device according to claim 1, wherein the generation unit generates the signals at the first position and the second position by an interpolation process or an extrapolation process for the at least phase components of the at least three microphones.

4. The coding device according to claim 1, wherein the generation unit generates signals of the two virtual microphones that are equidistant from the midpoint of the line segment connecting the first position and the second position in the first direction.

5. The coding device according to claim 1, wherein the generation unit generates a signal of one of the two virtual microphones using a signal at a third position that is the midpoint of the line segment connecting the first position and the second position and a signal at either the first position or the second position, and generates a signal of the other of the two virtual microphones using the signal at the third position.

6. The coding device according to claim 1, wherein the generation unit generates a signal of one of the two virtual microphones using the signal at the first position and the signal at the second position, and generates a signal of the other of the two virtual microphones using a signal at a position farther from one of the two virtual microphones among the first position and the second position.

7. The straight line connecting the two virtual microphones is a straight line passing through the centroid of the at least three microphones. The encoding device according to claim 1.

8. The position of the centroid is the center of the terminal provided with the encoding device in the two-dimensional plane. The encoding device according to claim 7.

9. The distance between the two virtual microphones is set to the inter-aural distance. The encoding device according to claim 1.

10. The first direction is the horizontal direction. The encoding device according to claim 1.

11. An encoding method, wherein the encoding device uses at least the phase components of the signals of at least three microphones arranged in a two-dimensional plane to generate signals at a first position and a second position on a straight line connecting two virtual microphones that constitute a virtual stereo microphone arranged in a first direction of the two-dimensional plane, generates signals of the two virtual microphones located at both ends of the first position and the second position by an extrapolation process for the signals at the first position and the second position, and performs stereo encoding using the signals of the two virtual microphones.

Citation Information

Patent Citations

  • Decoding and encoding apparatus and corresponding methods

    EP3312833A1

  • Sound field estimation device, method and program therefor

    JP2017112415A

  • Distance panning with near / far rendering

    JP2019523913A

  • Apparatus, program, and method of mixing picked-up sound signals from plurality of microphones

    JP2021132261A

  • Method of Rendering One or More Captured Audio Soundfields to a Listener

    US20160029144A1